Collecting Cross-Modal Presence-Absence Evidence for Weakly-Supervised Audio- Visual Event Perception
Junyu GaoMengyuan ChenChangsheng Xu
Proposes an evidential learning framework that extracts presence evidence from the primary modality and absence evidence from the complementary modality to accurately localize unsynchronized audio and visual events using only video-level labels.
Real-world video analysis relies heavily on interpreting both sight and sound, yet annotating precise timestamps for when events occur in audio and visual streams is prohibitively expensive. As a result, automated systems must learn from weak supervision, relying solely on broad video-level category tags without snippet-level time boundaries. Existing approaches typically assume audio and visual cues occur simultaneously or fail to effectively leverage one sensory stream to assist the other, resulting in noisy, overconfident predictions when audio and visual events are unsynchronized.
The article aims to design and validate a unified framework called Cross-Modal Presence-Absence Evidence (CMPAE) to accurately localize and categorize audible, visible, and combined audio-visual events using only video-level supervision.
The authors develop a framework grounded in evidential deep learning and Subjective Logic theory, which quantifies predictive uncertainty rather than producing standard classification probabilities. The approach introduces a presence-absence evidence collector that derives positive evidence of an event directly from its own modality (e.g., audio signals predicting audio events) while using the complementary modality (e.g., visual context) as a reference selector to establish absence evidence. A joint-modal mutual learning module then dynamically calibrates individual modal predictions against fused cross-modal evidence based on predictive uncertainty. The method was evaluated on standard video parsing and event localization benchmarks containing thousands of video clips spanning dozens of categories.
The experimental findings show that the proposed framework consistently surpasses existing state-of-the-art weakly-supervised methods. On the standard video parsing benchmark (Look, Listen, and Parse), the framework achieved event-level visual and audio performance gains of 3.8% and 2.7% over the leading baseline, and outpaced prior models on visual detection by up to 9.3%. On synchronized event localization benchmarks, the model achieved a top accuracy of 74.8%, nearing fully-supervised performance. On an expanded 39-category combined dataset, the method maintained robust performance, showing improvements of 3.6% in audio and 6.1% in visual event-level detection. Ablation analyses confirmed that both the cross-modal absence evidence collector and uncertainty-guided mutual calibration were essential to these gains.
These results demonstrate that explicitly modeling both presence and absence while accounting for predictive uncertainty substantially mitigates label noise in multi-sensory video analysis. For decision-makers and technology leaders, adopting uncertainty-aware evidential models lowers the operational costs and timelines associated with manual data annotation while improving system reliability across non-synchronized multimedia streams.
Organizations developing automated video indexing, surveillance, or content moderation systems should consider integrating uncertainty-calibrated evidential frameworks into their multimodal pipelines to reduce annotation overhead. Future development should focus on deploying and testing these models on larger, more diverse video datasets to evaluate generalizability across broader operational domains.
The primary limitation noted in the article is the reliance on existing public benchmark datasets, where certain synchronized localization tasks appear to be approaching a performance ceiling. Nevertheless, given the consistent empirical improvements across diverse metrics, there is high confidence in the framework's effectiveness for weakly-supervised audio-visual perception.
- Paper: A Closer Look at Weakly-Supervised Audio-Visual Source Localization, Shentong Mo et al. (2022). It analyzes the critical failure modes of weakly-supervised audio-visual localization under silent objects and off-screen sounds, establishing the core motivation for explicitly modeling cross-modal absence.
- Paper: Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks, Nan Wu et al. (2022). It provides the foundational study on how greedy learning and modality competition lead to unbalanced multimodal optimization, which the source addresses through uncertainty-guided mutual calibration.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). It introduces cross-modal attention mechanisms for unaligned multimodal sequences, establishing fundamental principles for modeling non-synchronized temporal interactions across audio and visual streams.
- Paper: Exploring Denoised Cross-video Contrast for Weakly-supervised Temporal Action Localization, Jingjing Li et al. (2022). It demonstrates effective contrastive denoising and background separation strategies for weakly-supervised temporal video localization using only video-level annotations.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). It formalizes the decomposition of multimodal inputs into modality-invariant and modality-specific latent spaces, laying key conceptual groundwork for separating positive self-evidence from cross-modal reference evidence.
- Paper: ReconBoost: Boosting Can Achieve Modality Reconcilement, Cong Hua et al. (2024). It advances multimodal learning beyond joint uncertainty calibration by introducing a dynamic boosting reconcilement framework that resolves modality competition in audio-visual event localization.
- Paper: PMR: Prototypical Modal Rebalance for Multimodal Learning, Yunfeng Fan et al. (2023). It builds on the challenge of uneven multimodal convergence by developing prototypical rebalancing techniques that dynamically accelerate slower-learning modalities during audio-visual perception.
- Paper: Egocentric Audio-Visual Object Localization, Chao Huang et al. (2023). It extends audio-visual localization to challenging egocentric video environments where camera motion and out-of-view sound sources introduce severe temporal and spatial misalignment.
- Paper: ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos, Jr-Jen Chen et al. (2024). It evaluates multimodal reasoning across non-simultaneous, temporally separated video events, offering a comprehensive benchmark suite for evaluating models that handle asynchronous sensory evidence.
