Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event Perception

Junyu GaoMengyuan ChenChangsheng Xu

article2023CVPR63 citations

Proposes an evidential learning framework that extracts presence evidence from the primary modality and absence evidence from the complementary modality to accurately localize unsynchronized audio and visual events using only video-level labels.

Listen

Real-world video analysis relies heavily on interpreting both sight and sound, yet annotating precise timestamps for when events occur in audio and visual streams is prohibitively expensive. As a result, automated systems must learn from weak supervision, relying solely on broad video-level category tags without snippet-level time boundaries. Existing approaches typically assume audio and visual cues occur simultaneously or fail to effectively leverage one sensory stream to assist the other, resulting in noisy, overconfident predictions when audio and visual events are unsynchronized.

The article aims to design and validate a unified framework called Cross-Modal Presence-Absence Evidence (CMPAE) to accurately localize and categorize audible, visible, and combined audio-visual events using only video-level supervision.

The authors develop a framework grounded in evidential deep learning and Subjective Logic theory, which quantifies predictive uncertainty rather than producing standard classification probabilities. The approach introduces a presence-absence evidence collector that derives positive evidence of an event directly from its own modality (e.g., audio signals predicting audio events) while using the complementary modality (e.g., visual context) as a reference selector to establish absence evidence. A joint-modal mutual learning module then dynamically calibrates individual modal predictions against fused cross-modal evidence based on predictive uncertainty. The method was evaluated on standard video parsing and event localization benchmarks containing thousands of video clips spanning dozens of categories.

The experimental findings show that the proposed framework consistently surpasses existing state-of-the-art weakly-supervised methods. On the standard video parsing benchmark (Look, Listen, and Parse), the framework achieved event-level visual and audio performance gains of 3.8% and 2.7% over the leading baseline, and outpaced prior models on visual detection by up to 9.3%. On synchronized event localization benchmarks, the model achieved a top accuracy of 74.8%, nearing fully-supervised performance. On an expanded 39-category combined dataset, the method maintained robust performance, showing improvements of 3.6% in audio and 6.1% in visual event-level detection. Ablation analyses confirmed that both the cross-modal absence evidence collector and uncertainty-guided mutual calibration were essential to these gains.

These results demonstrate that explicitly modeling both presence and absence while accounting for predictive uncertainty substantially mitigates label noise in multi-sensory video analysis. For decision-makers and technology leaders, adopting uncertainty-aware evidential models lowers the operational costs and timelines associated with manual data annotation while improving system reliability across non-synchronized multimedia streams.

Organizations developing automated video indexing, surveillance, or content moderation systems should consider integrating uncertainty-calibrated evidential frameworks into their multimodal pipelines to reduce annotation overhead. Future development should focus on deploying and testing these models on larger, more diverse video datasets to evaluate generalizability across broader operational domains.

The primary limitation noted in the article is the reliance on existing public benchmark datasets, where certain synchronized localization tasks appear to be approaching a performance ceiling. Nevertheless, given the consistent empirical improvements across diverse metrics, there is high confidence in the framework's effectiveness for weakly-supervised audio-visual perception.

  • Paper: ReconBoost: Boosting Can Achieve Modality Reconcilement, Cong Hua et al. (2024). It advances multimodal learning beyond joint uncertainty calibration by introducing a dynamic boosting reconcilement framework that resolves modality competition in audio-visual event localization.
  • Paper: PMR: Prototypical Modal Rebalance for Multimodal Learning, Yunfeng Fan et al. (2023). It builds on the challenge of uneven multimodal convergence by developing prototypical rebalancing techniques that dynamically accelerate slower-learning modalities during audio-visual perception.
  • Paper: Egocentric Audio-Visual Object Localization, Chao Huang et al. (2023). It extends audio-visual localization to challenging egocentric video environments where camera motion and out-of-view sound sources introduce severe temporal and spatial misalignment.
  • Paper: ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos, Jr-Jen Chen et al. (2024). It evaluates multimodal reasoning across non-simultaneous, temporally separated video events, offering a comprehensive benchmark suite for evaluating models that handle asynchronous sensory evidence.
Cover for Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event Perception

Abstract

With only video-level event labels, this paper targets at the task of weakly-supervised audio-visual event perception (WS-AVEP), which aims to temporally localize and categorize events belonging to each modality. Despite the recent progress, most existing approaches either ignore the unsynchronized property of audio-visual tracks or discount the complementary modality for explicit enhancement. We argue that, for an event residing in one modality, the modality itself should provide ample presence evidence of this event, while the other complementary modality is encouraged to afford the absence evidence as a reference signal. To this end, we propose to collect Cross-Modal Presence-Absence Evidence (CMPAE) in a unified framework. Specifically, by leveraging uni-modal and cross-modal representations, a presence-absence evidence collector (PAEC) is designed under Subjective Logic theory. To learn the evidence in a reliable range, we propose a joint-modal mutual learning (JML) process, which calibrates the evidence of diverse audible, visible, and audi-visible events adaptively and dynamically. Extensive experiments show that our method surpasses state-of-the-arts (e.g., absolute gains of 3.6% and 6.1% in terms of event-level visual and audio metrics). Code is available in github.com/MengyuanChen21/CVPR2023-CMPAE.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Approach
  • 3.1. Notations and Preliminaries
  • 3.2. Presence-Absence Evidence Collector
  • 3.3. Joint-modal Mutual Learning
  • 3.4. Learning and Inference
  • 4. Experimental Results
  • 4.1. Experimental Setup
  • 4.2. Comparison with State-of-the-art Methods
  • 4.3. Further Remarks
  • 5. Conclusions
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Task Formulation of Weakly-Supervised Audio-Visual Event Perception

    definition

    Weakly-Supervised Audio-Visual Event Perception (WS-AVEP) aims to temporally localize and categorize events across sensory tracks (audio, visual, and synchronized audio-visual) using only video-level event category labels during training.

    A video VV is partitioned into TT non-overlapping temporal snippets with extracted audio and visual feature sequences {xta,xtv}t=1T⊂RD\{x_t^a, x_t^v\}_{t=1}^T \subset \mathbb{R}^D. While fine-grained snippet-level labels are unavailable during training, the video is associated with multi-hot category vectors:

    • ya∈{0,1}Cy^a \in \{0, 1\}^C for audio events,
    • yv∈{0,1}Cy^v \in \{0, 1\}^C for visual events, and
    • yav∈{0,1}Cy^{av} \in \{0, 1\}^C for synchronized audio-visual events, where CC is the total number of event categories. Under the weakly-supervised setting, training only accesses a modality-agnostic video-level label y∈{0,1}Cy \in \{0, 1\}^C, where yc=1y_c = 1 indicates that category cc occurs in at least one modality, and yc=0y_c = 0 indicates absence across all modalities.
  2. Knowl 2 — Cross-Modal Presence-Absence Evidence Collector

    model/method

    To avoid overconfidence and explicitly quantify predictive uncertainty for individual modalities m∈{a,v}m \in \{a, v\} and category c∈{1,…,C}c \in \{1, \dots, C\}, the Presence-Absence Evidence Collector (PAEC) separates the generation of presence evidence et,cme_{t,c}^m and absence evidence e^t,cm\hat{e}_{t,c}^m at snippet tt:

    1. Presence Evidence: Derived purely from the uni-modal representation xtm∈RDx_t^m \in \mathbb{R}^D to ensure self-reliability and encourage modality-specific discrimination: et,cm=g(fc(xtm;θ1))e_{t,c}^m = g(f_c(x_t^m; \theta_1)) where fc(⋅;θ1)f_c(\cdot; \theta_1) is a deep neural network shared across audio and visual modalities, and g(⋅)g(\cdot) is an activation function ensuring non-negativity (e.g., the exponential function g(z)=exp⁡(z)g(z) = \exp(z)).

    2. Absence Evidence: Constructed using the complementary modality m^\hat{m} (where m^=v\hat{m} = v if m=am = a, and m^=a\hat{m} = a if m=vm = v) as a cross-modal context selector: e^t,cm=g(f^c(x^tm;θ2))\hat{e}_{t,c}^m = g(\hat{f}_c(\hat{x}_t^m; \theta_2)) x^tm=xtm⊙(hm(xtm+xtm^;θ3m)+1)\hat{x}_t^m = x_t^m \odot \left(h^m(x_t^m + x_t^{\hat{m}}; \theta_3^m) + \mathbf{1}\right) where hm(⋅;θ3m)h^m(\cdot; \theta_3^m) is a fully-connected layer mapping the fused features xtm+xtm^x_t^m + x_t^{\hat{m}} to channel-wise weights, ⊙\odot denotes the element-wise (Hadamard) product, 1\mathbf{1} is an all-ones vector for residual scaling, and f^c(⋅;θ2)\hat{f}_c(\cdot; \theta_2) is a shared absence evidence collector.

    3. Temporal Aggregation: Snippet-level evidence is aggregated into video-level presence evidence evid,cme_{vid,c}^m and absence evidence e^vid,cm\hat{e}_{vid,c}^m via learned modality-aware temporal attention weights At,cmA_{t,c}^m: {evid,cm,e^vid,cm}=∑t=1TAt,cm{et,cm,e^t,cm}\{e_{vid,c}^m, \hat{e}_{vid,c}^m\} = \sum_{t=1}^T A_{t,c}^m \{e_{t,c}^m, \hat{e}_{t,c}^m\}

  3. Knowl 3 — Evidential Beta Distribution and Classification Loss

    equation

    Under Subjective Logic theory, the video-level presence evidence evid,cme_{vid,c}^m and absence evidence e^vid,cm\hat{e}_{vid,c}^m for class c∈{1,…,C\Itemc \in \{1, \dots, C\Item in modality m∈{a,v,av}m \in \{a, v, av\} parameterize a Beta distribution over the class probability pc∈[0,1]p_c \in [0, 1]: Beta(pc∣αc,βc)=1B(αc,βc)pcαc−1(1−pc)βc−1\mathrm{Beta}(p_c \mid \alpha_c, \beta_c) = \frac{1}{\mathrm{B}(\alpha_c, \beta_c)} p_c^{\alpha_c - 1} (1 - p_c)^{\beta_c - 1} where B(αc,βc)=Γ(αc)Γ(βc)Γ(αc+βc)\mathrm{B}(\alpha_c, \beta_c) = \frac{\Gamma(\alpha_c)\Gamma(\beta_c)}{\Gamma(\alpha_c + \beta_c)}, Γ(⋅)\Gamma(\cdot) is the Gamma function, and the distribution parameters αc\alpha_c and βc\beta_c relate to the collected evidence by: αc=evid,cm+1,βc=e^vid,cm+1\alpha_c = e_{vid,c}^m + 1, \quad \beta_c = \hat{e}_{vid,c}^m + 1

    The classification loss LclsmL_{cls}^m is formulated as the Bayes risk with respect to the binary cross-entropy loss: Lclsm=∑c=1C∫01−ycmlog⁡(pc) Beta(pc∣αc,βc) dpc=∑c=1C[ψ(αc+βc)−ψ(ycmαc+(1−ycm)βc)]L_{cls}^m = \sum_{c=1}^C \int_0^1 -y_c^m \log(p_c) \, \mathrm{Beta}(p_c \mid \alpha_c, \beta_c) \, dp_c = \sum_{c=1}^C \left[ \psi(\alpha_c + \beta_c) - \psi(y_c^m \alpha_c + (1 - y_c^m)\beta_c) \right] where ψ(⋅)\psi(\cdot) denotes the digamma function, and ycm∈{0,1}y_c^m \in \{0, 1\} is the video-level binary label for category cc in track mm. In numerical implementation for training stability, the negative logarithm of the marginal likelihood can be employed by replacing ψ(⋅)\psi(\cdot) with a logarithm function.

  4. Knowl 4 — Evidential Probabilities, Uncertainties, and Uni-Modal Prediction Fusion

    model/method

    From the Subjective Logic formulation with parameters αc=ecm+1\alpha_c = e_c^m + 1 and βc=e^cm+1\beta_c = \hat{e}_c^m + 1, the expected classification probability pcmp_c^m and predictive uncertainty ucmu_c^m for track m∈{a,v,av}m \in \{a, v, av\} and category c∈{1,…,C}c \in \{1, \dots, C\} are inferred as: pcm=ecm+1ecm+e^cm+2,ucm=2ecm+e^cm+2p_c^m = \frac{e_c^m + 1}{e_c^m + \hat{e}_c^m + 2}, \quad u_c^m = \frac{2}{e_c^m + \hat{e}_c^m + 2}

    Because video-level weak labels y∈{0,1}Cy \in \{0, 1\}^C do not specify which modality contains positive events, a fused uni-modal probability pcunip_c^{uni} and uncertainty ucuniu_c^{uni} are computed via a dynamic selector δ(c)\delta(c): {ucuni,pcuni}=δ(c){uca,pca}+(1−δ(c)){ucv,pcv}\{u_c^{uni}, p_c^{uni}\} = \delta(c)\{u_c^a, p_c^a\} + (1 - \delta(c))\{u_c^v, p_c^v\} where: δ(c)={1,if pca>pcv and yc=1,0,if pca≤pcv and yc=1,12,if yc=0.\delta(c) = \begin{cases} 1, & \text{if } p_c^a > p_c^v \text{ and } y_c = 1, \\ 0, & \text{if } p_c^a \le p_c^v \text{ and } y_c = 1, \\ \frac{1}{2}, & \text{if } y_c = 0. \end{cases} For target classes (yc=1y_c = 1), the selector acts as a max-operator taking the more confident modality's predictions; for non-target classes (yc=0y_c = 0), it takes the mean of both modalities.

  5. Knowl 5 — Joint-Modal Mutual Learning with Uncertainty Calibration

    model/method

    Joint-Modal Mutual Learning (JML) coordinates uni-modal and joint-modal predictions by calibrating their mutual interaction with predictive uncertainties.

    1. Global Joint-Modal Evidence: Joint-modal presence and absence evidence are generated using fused snippet features xta+xtvx_t^a + x_t^v: evid,cav,e^vid,cav=∑t=1TAt,cav⋅g(fcav(xta+xtv;θ4))e_{vid,c}^{av}, \hat{e}_{vid,c}^{av} = \sum_{t=1}^T A_{t,c}^{av} \cdot g(f_c^{av}(x_t^a + x_t^v; \theta_4)) where At,cav=12(At,ca+At,cv)A_{t,c}^{av} = \frac{1}{2}(A_{t,c}^a + A_{t,c}^v) is the joint temporal attention score, fcav(⋅;θ4)f_c^{av}(\cdot; \theta_4) is a dedicated evidence collector, and gg is the non-negative activation function (e.g., Exp). The video-level joint loss LclsavL_{cls}^{av} is computed using the evidential classification loss with video labels yy.

    2. Mutual Learning Loss Objectives: Letting uav=1C∑c=1Cucavu^{av} = \frac{1}{C}\sum_{c=1}^C u_c^{av} denote the mean global joint uncertainty, s(⋅)s(\cdot) denote gradient truncation (stop-gradient), and l(⋅)l(\cdot) denote a distance metric (such as squared L2L_2-norm), mutual learning objectives are defined as: Ljml,1=∑c=1C(1−uav)(1−ucuni)⋅l(s(pcav),pcuni)L_{jml, 1} = \sum_{c=1}^C (1 - u^{av}) (1 - u_c^{uni}) \cdot l(s(p_c^{av}), p_c^{uni}) Ljml,2=∑c=1Cuav(1−ucuni)⋅l(pcav,s(pcuni))L_{jml, 2} = \sum_{c=1}^C u^{av} (1 - u_c^{uni}) \cdot l(p_c^{av}, s(p_c^{uni})) When joint-modal uncertainty uavu^{av} is low (1−uav1 - u^{av} is high), Ljml,1L_{jml,1} guides the uni-modal prediction pcunip_c^{uni} using the stable joint prediction pcavp_c^{av}. Conversely, when joint uncertainty uavu^{av} is high, Ljml,2L_{jml,2} allows confident uni-modal predictions to guide the joint branch.

    3. Alternating Training: Optimization alternates between the two total losses across training iterations: Li=∑m∈{a,v,av}Lclsm+Ljml,i,i∈{1,2}L_i = \sum_{m \in \{a, v, av\}} L_{cls}^m + L_{jml, i}, \quad i \in \{1, 2\}

  6. Knowl 6 — Experimental Setup for Audio-Visual Event Perception Benchmarks

    experimental setup

    The Cross-Modal Presence-Absence Evidence (CMPAE) framework is evaluated across three benchmarks:

    • Look, Listen, and Parse (LLP) for Audio-Visual Video Parsing (AVVP): Contains 11,849 10-second video clips (10,000 train, 649 validation, 1,200 test) covering 25 event categories from AudioSet with an average of 1.64 event categories per video. Events on audio and visual tracks can be unsynchronized.
    • Audio-Visual Event (AVE): Comprises 4,143 video clips (3,339 train, 402 validation, 402 test) across 28 categories, focusing strictly on synchronized audio-visual events.
    • AVEP: Combines LLP and non-overlapping AVE categories, yielding 11,581 training, 840 validation, and 1,391 testing videos across 39 categories.

    Feature Extraction and Architecture:

    • Audio: 512-dimensional features (D=512D=512) extracted via pre-trained VGGish over T=10T=10 snippets per video.
    • Visual: ResNet152 and R(2+1)D features for LLP and AVEP; VGG-19 features for AVE.
    • Evidence collectors fc,f^cf_c, \hat{f}_c consist of the backbone (built on JoMoLD) with two fully-connected layers activated by LeakyReLU, with evidence function g=Expg = \mathrm{Exp}.

    Optimization:

    • Optimizer: Adam, batch size 128, initial learning rate 5×10−45 \times 10^{-4} decayed by a factor of 0.25 every 6 epochs, trained for 25 epochs on an NVIDIA RTX 3090 GPU.

    Metrics:

    • Segment-level and event-level F-scores for audio (A), visual (V), audio-visual (AV), modality-averaged (Type), and modality-agnostic (Event) events. Event-level F-scores require segment predictions with temporal mean Intersection over Union (mIoU\mathrm{mIoU}) ≥0.5\ge 0.5. For AVE, overall classification accuracy is reported.
  7. Knowl 7 — Audio-Visual Video Parsing Performance on the LLP Dataset

    data/table

    On the Look, Listen, and Parse (LLP) dataset, Cross-Modal Presence-Absence Evidence (CMPAE) outperforms existing weakly-supervised audio-visual parsing methods across all segment-level and event-level metrics:

    Methods Segment-level (%) Event-level (%)
    A V AV Type Event A V AV Type Event
    AVE (ECCV 2018) 49.9 37.3 37.0 41.4 43.6 43.6 32.4 32.6 36.2 37.4
    AVSDN (ICASSP 2019) 47.8 52.0 37.1 45.7 50.8 34.1 46.3 26.5 35.6 37.7
    HAN (ECCV 2020) 60.1 52.9 48.9 54.0 55.4 51.3 48.9 43.0 47.7 48.0
    CVCMS (NeurIPS 2021) 60.8 63.5 57.0 60.5 59.5 53.8 58.9 49.5 54.0 52.1
    MA (CVPR 2021) 59.8 57.5 52.6 56.6 56.6 52.1 54.4 45.8 50.8 49.4
    DHHN (ACM MM 2022) 61.4 63.4 56.8 60.5 59.5 54.6 60.8 51.1 55.5 53.3
    MM-Pyramid (ACM MM 2022) 61.1 60.3 55.8 59.7 59.1 53.8 56.7 49.4 54.1 51.2
    CMBS (CVPR 2022) 60.2 54.3 50.0 54.8 55.7 51.1 50.8 43.7 48.5 48.3
    JoMoLD (ECCV 2022) 61.3 63.8 57.2 60.8 59.9 53.9 59.9 49.6 54.5 52.5
    CMPAE (Ours) 64.2 66.4 59.2 63.3 62.8 56.6 63.7 51.8 57.4 55.7

    Compared to the JoMoLD baseline, CMPAE achieves absolute gains of +2.9% on segment-level audio (A), +2.6% on segment-level visual (V), +2.5% on segment-level Type, +3.8% on event-level visual (V, reaching 63.7%), and +2.9% on event-level Type (reaching 57.4%).

  8. Knowl 8 — Performance Comparison on AVE and AVEP Datasets

    data/table

    Evaluations on the Audio-Visual Event (AVE) dataset for synchronized event localization and the combined AVEP dataset (39 categories) demonstrate that CMPAE surpasses prior weakly-supervised methods across both settings:

    Audio-Visual Event (AVE) Accuracy:

    Methods Accuracy (%)
    AVEL (ECCV 2018) 66.7
    AVRB (WACV 2020) 68.9
    CMRAN (ACM MM 2020) 72.9
    PSP (CVPR 2021) 73.5
    CMAN (AAAI 2022) 70.4
    MM-Pyramid (ACM MM 2022) 73.2
    CMBS (CVPR 2022) 74.2
    DPNet (ECCV 2022) 74.5
    JoMoLD (ECCV 2022) 71.8
    CMPAE (Ours) 74.8
    CMBS (fully-supervised) 79.3

    CMPAE achieves 74.8% accuracy on AVE, exceeding PSP (73.5%), CMBS (74.2%), and JoMoLD (71.8%).

    AVEP Dataset Performance:

    Methods Segment-level (%) Event-level (%)
    A V AV Type Event A V AV Type Event
    CMBS (CVPR 2022) 58.0 56.2 52.3 55.5 54.8 51.5 53.6 46.4 50.5 49.4
    JoMoLD (ECCV 2022) 60.6 58.9 54.5 58.0 57.7 53.6 55.8 48.6 52.7 51.0
    CMPAE (Ours) 64.1 64.4 58.8 62.4 62.2 57.2 61.9 52.3 57.1 55.6

    On AVEP, CMPAE improves over JoMoLD by +4.4% in segment-level Type (62.4% vs. 58.0%) and +4.4% in event-level Type (57.1% vs. 52.7%), with an event-level visual F-score gain of +6.1% (61.9% vs. 55.8%).

  9. Knowl 9 — Ablation Study of PAEC and Joint-Modal Mutual Learning Components

    empirical result

    Ablation experiments on the LLP (AVVP) and AVEP datasets evaluate the contributions of Evidential Deep Learning (EDL), the Presence-Absence Evidence Collector (PAEC), and Joint-Modal Mutual Learning (JML):

    EDL PAEC JML Seg-level Type (%) Eve-level Type (%)
    AVVP AVEP AVVP AVEP
    60.8 58.0 54.5 52.7
    ✓ 61.0 58.9 54.9 53.8
    ✓ ✓ 61.9 61.5 56.1 55.9
    ✓ ✓ 61.4 60.8 55.3 54.6
    ✓ ✓ ✓ 63.3 62.4 57.4 57.1

    Further design investigations show:

    • Evidence Feature Sourcing: Deriving both presence and absence evidence from uni-modal features yields 55.2% event-level Type on AVVP, while using cross-modal features for both gives 55.7%. Inverting the configuration by deriving presence evidence from cross-modal features and absence evidence from uni-modal features achieves 56.4%, all lower than CMPAE's 57.4%.
    • Uncertainty Terms: Removing joint-modal uncertainty uavu^{av} reduces event-level Type on AVVP to 56.5%; removing uni-modal uncertainty uuniu^{uni} reduces it to 56.4%.
    • Uni-Modal Selector: Replacing the conditional selector δ(c)\delta(c) with a uniform mean operator drops event-level Type on AVVP to 56.3% and on AVEP to 56.5%.
  10. Knowl 10 — Limitations of WS-AVEP Benchmarking and Dataset Diversity

    limitation

    Although combining the LLP and AVE datasets into the AVEP benchmark expands evaluation to 39 categories with varied modality-synchronization profiles, the performance of weakly-supervised models on AVE is near that of fully-supervised methods (e.g., 74.8% for CMPAE vs. 79.3% for fully-supervised CMBS), suggesting a ceiling effect under current evaluation regimes. The authors state that larger-scale datasets with greater real-world diversity and complexity are required in the future to more comprehensively assess the effectiveness of weakly-supervised audio-visual event perception methods.

Coverage note — No substantial contributed material was omitted. All primary methodology components (PAEC, evidential loss, JML), experimental evaluations on LLP, AVE, and AVEP, ablation studies, and limitations are fully covered.

References

  1. 1.Alexander Amini, Wilko Schwarting, Ava Soleimany, and Daniela Rus. Deep evidential regression. In NeurIPS, 2020. 2, 3
  2. 2.Relja Arandjelovic and Andrew Zisserman. Look, listen and learn. In ICCV, 2017. 2
  3. 3.Wentao Bao, Qi Yu, and Yu Kong. Evidential deep learning for open set action recognition. In ICCV (ICCV), 2021. 2, 3, 4
  4. 4.Abhijit Bendale and Terrance E Boult. Towards open set deep networks. In CVPR, 2016.
  5. 5.Mengyuan Chen, Junyu Gao, Shicai Yang, and Changsheng Xu. Dual-evidential learning for weakly-supervised temporal action localization. In ECCV, 2022. 2, 3, 4
  6. 6.Haoyue Cheng, Zhaoyang Liu, Hang Zhou, Chen Qian, Wayne Wu, and Limin Wang. Joint-modal label denoising for weakly-supervised audio-visual video parsing. ECCV, 2022. 2, 3, 4, 5, 6, 7
  7. 7.Junyu Gao, Mengyuan Chen, and Changsheng Xu. Fine-grained temporal contrastive learning for weakly-supervised temporal action localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19999–20009, June 2022. 3
  8. 8.Junyu Gao and Changsheng Xu. Fast video moment retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1523–1532, 2021. 2
  9. 9.Junyu Gao, Tianzhu Zhang, and Changsheng Xu. Graph convolutional tracking. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4649–4659, 2019. 2
  10. 10.Junyu Gao, Tianzhu Zhang, and Changsheng Xu. I know the relationships: Zero-shot action recognition via two-stream graph convolutional networks and knowledge graphs. In AAAI, 2019. 2
  11. 11.Junyu Gao, Tianzhu Zhang, and Changsheng Xu. Smart: Joint sampling and regression for visual tracking. IEEE Transactions on Image Processing, 28(8):3923–3935, 2019. 1
  12. 12.Junyu Gao, Tianzhu Zhang, and Changsheng Xu. Learning to model relationships for zero-shot video classification. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(10):3476–3491, 2021. 2
  13. 13.Ruohan Gao, Rogerio Feris, and Kristen Grauman. Learning to separate object sounds by watching unlabeled video. In ECCV, 2018. 2
  14. 14.Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, and Lorenzo Torresani. Listen to look: Action recognition by previewing audio. In CVPR, 2020. 2
  15. 15.Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. Audio set: An ontology and human-labeled dataset for audio events. In ICASSP, 2017. 6
  16. 16.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In ICML. PMLR, 2017. 3
  17. 17.Zongbo Han, Changqing Zhang, Huazhu Fu, and Joey Tianyi Zhou. Trusted multi-view classification. In ICLR, 2020. 2, 3, 4
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 6
  19. 19.Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al. Cnn architectures for large-scale audio classification. In ICASSP, 2017. 6
  20. 20.Di Hu, Feiping Nie, and Xuelong Li. Deep multimodal clustering for unsupervised audiovisual learning. In CVPR, 2019. 2
  21. 21.Yibo Hu, Yuzhe Ou, Xujiang Zhao, Jin-Hee Cho, and Feng Chen. Multidimensional uncertainty-aware evidential neural networks. In AAAI, 2021. 3
  22. 22.Xun Jiang, Xing Xu, Zhiguo Chen, Jingran Zhang, Jingkuan Song, Fumin Shen, Huimin Lu, and Heng Tao Shen. Dhhn: Dual hierarchical hybrid network for weakly-supervised audio-visual video parsing. In ACM MM, 2022. 2, 3, 7
  23. 23.Audun Jsang. Subjective logic: A formalism for reasoning under uncertainty. Springer Verlag, 2016. 2, 3, 4, 5
  24. 24.Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. Epic-fusion: Audio-visual temporal binding for egocentric action recognition. In ICCV, 2019. 2
  25. 25.Bruno Korbar, Du Tran, and Lorenzo Torresani. Scsampler: Sampling salient clips from video for efficient action recognition. In ICCV, 2019. 2
  26. 26.Jun-Tae Lee, Mihir Jain, Hyoungwoo Park, and Sungrack Yun. Cross-attentional audio-visual fusion for weakly-supervised action localization. In ICLR, 2020. 2
  27. 27.Bolian Li, Zongbo Han, Haining Li, Huazhu Fu, and Changqing Zhang. Trustworthy long-tailed classification. In CVPR, 2022. 2, 3, 4
  28. 28.Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, and Di Hu. Learning to answer questions in dynamic audio-visual scenarios. In CVPR, 2022. 3
  29. 29.Yan-Bo Lin, Yu-Jhe Li, and Yu-Chiang Frank Wang. Dual-modality seq2seq network for audio-visual event localization. In ICASSP, 2019. 3, 7
  30. 30.Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin, and Ming-Hsuan Yang. Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing. NeurIPS, 2021. 1, 2, 3, 6, 7
  31. 31.Yan-Bo Lin and Yu-Chiang Frank Wang. Audiovisual transformer with instance attention for audio-visual event localization. In ACCV, 2020. 3
  32. 32.Shuo Liu, Weize Quan, Yuan Liu, and Dong-Ming Yan. Bi-directional modality fusion network for audio-visual event localization. In ICASSP, 2022. 3
  33. 33.Wei Liu, Xiaodong Yue, Yufei Chen, and Thierry Denoeux. Trusted deep learning with opinion aggregation. In AAAI, 2022. 3
  34. 34.Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song. Active contrastive learning of audio-visual video representations. In ICLR, 2020. 2
  35. 35.Tanvir Mahmud and Diana Marculescu. Ave-clip: Audioclip-based multi-window temporal transformer for audio visual event localization. arXiv preprint arXiv:2210.05060, 2022. 3
  36. 36.Andrey Malinin and Mark Gales. Predictive uncertainty estimation via prior networks. In NeurIPS, 2018. 2, 3, 4
  37. 37.Otniel-Bogdan Mercea, Lukas Riesch, A Koepke, and Zeynep Akata. Audio-visual generalised zero-shot learning with cross-modal attention and language. In CVPR, 2022. 3
  38. 38.Shentong Mo and Yapeng Tian. Multi-modal grouping network for weakly-supervised audio-visual video parsing. In Advances in Neural Information Processing Systems, 2022. 3
  39. 39.Shentong Mo and Yapeng Tian. Semantic-aware multi-modal grouping for weakly-supervised audio-visual video parsing. In ECCV Workshop, 2022. 3
  40. 40.Pedro Morgado, Ishan Misra, and Nuno Vasconcelos. Robust audio-visual instance discrimination. In CVPR, 2021. 2
  41. 41.Pedro Morgado, Nuno Vasconcelos, and Ishan Misra. Audio-visual instance discrimination with cross-modal agreement. In CVPR, 2021. 2
  42. 42.Jonathan Munro and Dima Damen. Multi-modal domain adaptation for fine-grained action recognition. In CVPR, 2020. 3
  43. 43.Kazuki Osawa, Siddharth Swaroop, Mohammad Emtiyaz E Khan, Anirudh Jain, Runa Eschenhagen, Richard E Turner, and Rio Yokota. Practical deep learning with bayesian principles. In NeurIPS, 2019. 2
  44. 44.Andrew Owens and Alexei A Efros. Audio-visual scene analysis with self-supervised multisensory features. In ECCV, 2018. 2
  45. 45.Deep Shankar Pandey and Qi Yu. Multidimensional belief quantification for label-efficient meta-learning. In CVPR, 2022. 3
  46. 46.Janani Ramaswamy. What makes the sound?: A dual-modality interacting network for audio-visual event localization. In ICASSP. IEEE, 2020. 3
  47. 47.Janani Ramaswamy and Sukhendu Das. See the sound, hear the pixels. In WACV, 2020. 1, 3, 7
  48. 48.Varshanth Rao, Md Ibrahim Khalil, Haoda Li, Peng Dai, and Juwei Lu. Dual perspective network for audio-visual event localization. In ECCV, 2022. 3, 7
  49. 49.Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon. Learning to localize sound source in visual scenes. In CVPR, 2018. 2
  50. 50.Murat Sensoy, Lance Kaplan, and Melih Kandemir. Evidential deep learning to quantify classification uncertainty. In NeurIPS, 2018. 2, 3, 4, 7
  51. 51.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 6
  52. 52.Yapeng Tian, Dingzeyu Li, and Chenliang Xu. Unified multisensory perception: Weakly-supervised audio-visual video parsing. In ECCV, 2020. 1, 2, 3, 5, 6, 7
  53. 53.Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. Audio-visual event localization in unconstrained videos. In ECCV, 2018. 1, 2, 3, 6, 7
  54. 54.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In CVPR, 2018. 6
  55. 55.Dennis Ulmer. A survey on evidential deep learning for single-pass uncertainty estimation. arXiv preprint arXiv:2110.03051, 2021. 2, 3
  56. 56.Yake Wei, Di Hu, Yapeng Tian, and Xuelong Li. Learning in audio-visual context: A review, analysis, and new perspective. arXiv preprint arXiv:2208.09579, 2022. 1
  57. 57.Yu Kong Wentao Bao, Qi Yu. Opental: Towards open set temporal action localization. In CVPR, 2022. 2, 3
  58. 58.Yu Wu and Yi Yang. Exploring heterogeneous clues for weakly-supervised audio-visual video parsing. In CVPR, 2021. 2, 3, 4, 7
  59. 59.Yiling Wu, Xinfeng Zhang, Yaowei Wang, and Qingming Huang. Span-based audio-visual localization. In ACM MM, 2022. 3
  60. 60.Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang. Dual attention matching for audio-visual event localization. In ICCV, 2019. 1, 2, 3
  61. 61.Yan Xia and Zhou Zhao. Cross-modal background suppression for audio-visual event localization. In CVPR, 2022. 1, 2, 3, 6, 7
  62. 62.Haoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan, and Chuang Gan. Cross-modal relation-aware networks for audio-visual event localization. In ACM MM, 2020. 1, 3, 7
  63. 63.Hanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang, and Yan Yan. Cross-modal attention network for temporal inconsistent audio-visual event localization. In AAAI, 2020. 1, 3, 7
  64. 64.Ronald R Yager and Liping Liu. Classic works of the Dempster-Shafer theory of belief functions, volume 219. Springer, 2008. 2, 3
  65. 65.Jiashuo Yu, Ying Cheng, Rui-Wei Zhao, Rui Feng, and Yuejie Zhang. Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video parsing. In ACM MM, 2022. 3, 7
  66. 66.Yunhua Zhang, Hazel Doughty, Ling Shao, and Cees GM Snoek. Audio-adaptive activity recognition across video domains. In CVPR, 2022. 1, 2, 3, 8
  67. 67.Ying Zhang, Tao Xiang, Timothy M Hospedales, and Huchuan Lu. Deep mutual learning. In CVPR, 2018. 8
  68. 68.Xujiang Zhao, Xuchao Zhang, Wei Cheng, Wenchao Yu, Yuncong Chen, Haifeng Chen, and Feng Chen. Seed: Sound event early detection via evidential uncertainty. In ICASSP, 2022. 4
  69. 69.Hang Zhou, Ziwei Liu, Xudong Xu, Ping Luo, and Xiaogang Wang. Vision-infused deep audio inpainting. In ICCV, 2019. 2
  70. 70.Jinxing Zhou, Liang Zheng, Yiran Zhong, Shijie Hao, and Meng Wang. Positive sample propagation along the audio-visual event line. In CVPR, 2021. 3, 7
  71. 71.Yonggang Zhu, Chao Tian, Zhuqing Jiang, Aidong Men, Haiying Wang, and Qingchao Chen. Mixed in time and modality: Curse or blessingƒ cross-instance data augmentation for weakly supervised multimodal temporal fusion. In ICASSP, 2022. 3

Citation

MLA
Gao, J., et al. “Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event Perception”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 18827–36, https://doi.org/10.1109/CVPR52729.2023.01805.
APA
Gao, J., Chen, M., & Xu, C. (2023). Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event Perception. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18827–18836. https://doi.org/10.1109/CVPR52729.2023.01805
Chicago
Gao, J., M. Chen, and C. Xu. 2023. “Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event Perception”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18827–36. https://doi.org/10.1109/CVPR52729.2023.01805.
Harvard
Gao, J., Chen, M. and Xu, C. (2023) “Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event Perception”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 18827–18836. Available at: https://doi.org/10.1109/CVPR52729.2023.01805.
Vancouver
1. Gao J, Chen M, Xu C (2023) Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event Perception. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 18827–18836

BibTeX

@inproceedings{Gao_2023, title={Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event Perception}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01805}, DOI={10.1109/cvpr52729.2023.01805}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Gao, Junyu and Chen, Mengyuan and Xu, Changsheng}, year={2023}, month=June, pages={18827–18836} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE