Egocentric Audio-Visual Object Localization
Chao HuangYapeng TianAnurag KumarChenliang Xu
Introduces a self-supervised framework for egocentric audio-visual object localization that explicitly accounts for camera motion via geometry-aware temporal aggregation and separates out-of-view audio distractors to accurately pinpoint sounding objects in first-person video.
Wearable devices and first-person cameras are increasingly deployed across robotics, augmented reality, and healthcare. However, automated systems struggle to accurately understand egocentric video due to constant camera wearer movement (egomotion) and limited visual fields of view, which frequently cause sound sources to shift out of view or create visual distortions. The article evaluates and demonstrates a self-supervised computational framework designed to overcome these challenges by accurately localizing sounding objects in egocentric recordings.
To address egomotion and out-of-view audio without relying on expensive manual data labeling, the authors developed a dual-module framework trained through audio-visual temporal synchronization. They introduced a geometry-aware temporal aggregation module that computes geometric transformations between frames to align visual features across changing viewpoints. Simultaneously, they incorporated a cascaded feature enhancement module using a sound separation task to disentangle visible sounds from background audio mixtures and direct visual attention to relevant regions. To evaluate the framework, the authors created the Epic Sounding Object benchmark by annotating 3,172 video clips across 30 sounding object classes from the Epic-Kitchens dataset and tested cross-scenario generalization on the Ego4D dataset.
The experimental findings show that the proposed framework substantially outperforms existing methods. On the benchmark's primary localization metric (Consensus Intersection over Union at a 0.2 threshold), the proposed method achieved a score of 38.71%, compared to 26.01% for the best-performing prior approach and 16.51% for a central-focus baseline. The overall Area Under Curve improved to 18.38% versus 15.39% for the previous state of the art. Ablation studies confirmed that geometric alignment contributed the largest single performance gain (increasing CIoU@0.2 from 27.41% to 37.38%), while sound disentanglement and soft visual attention provided vital additional accuracy gains.
These results demonstrate that explicitly accounting for physical camera motion and separating off-screen audio noise are critical for robust first-person perception. In practical terms, this capability provides a path toward lower data-labeling costs through self-supervised training and enhances the reliability of applications such as audio-queried memory retrieval in smart glasses, object state tracking during human-environment interactions, and robotic trajectory forecasting.
Organizations developing egocentric vision and wearable artificial intelligence systems should integrate geometric motion alignment and audio disentanglement into their perception pipelines. Before full commercial deployment, teams should conduct further development to address operational limitations: current geometric estimation relies on image alignment techniques that can fail during extreme camera motion or sudden illumination shifts. Enhancing geometric estimation robustness across diverse real-world lighting and motion conditions represents the key next step for operational deployment.
- Paper: Scaling Egocentric Vision: The EPIC-KITCHENS Dataset, Dima Damen et al. (2018). It introduces the EPIC-KITCHENS benchmark dataset, from which the source paper derives and constructs its primary Epic Sounding Object evaluation suite.
- Paper: Ego4D: Around the World in 3,000 Hours of Egocentric Video, Kristen Grauman et al. (2021). It provides the foundational large-scale egocentric dataset and audio-visual benchmarks used by the source paper to test cross-scenario generalization.
- Paper: A Closer Look at Weakly-Supervised Audio-Visual Source Localization, Shentong Mo et al. (2022). It establishes key principles for weakly-supervised visual sound source localization under realistic challenges like off-screen sounds and silent objects, upon which the source's dual-module approach builds.
- Paper: Unsupervised Learning of Depth and Ego-Motion from Video, Tinghui Zhou et al. (2017). It lays the groundwork for unsupervised learning of camera ego-motion and geometric frame warping, which underpins the source paper's geometry-aware temporal aggregation module.
- Paper: HourVideo: 1-Hour Video-Language Understanding, Keshigeyan Chandrasegaran et al. (2024). It scales up egocentric multimodal understanding to long-horizon, hour-long video reasoning, extending the first-person perception challenges addressed in the source paper.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). It advances long-form egocentric and third-person video understanding through temporal grounding and multimodal LLM tuning, building beyond localized audio-visual alignment.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). It develops feed-forward visual geometry grounding transformers to overcome camera motion and estimate 3D scene attributes directly from video streams.
