Egocentric Audio-Visual Object Localization

Chao HuangYapeng TianAnurag KumarChenliang Xu

article2023CVPR54 citations

Introduces a self-supervised framework for egocentric audio-visual object localization that explicitly accounts for camera motion via geometry-aware temporal aggregation and separates out-of-view audio distractors to accurately pinpoint sounding objects in first-person video.

Listen

Wearable devices and first-person cameras are increasingly deployed across robotics, augmented reality, and healthcare. However, automated systems struggle to accurately understand egocentric video due to constant camera wearer movement (egomotion) and limited visual fields of view, which frequently cause sound sources to shift out of view or create visual distortions. The article evaluates and demonstrates a self-supervised computational framework designed to overcome these challenges by accurately localizing sounding objects in egocentric recordings.

To address egomotion and out-of-view audio without relying on expensive manual data labeling, the authors developed a dual-module framework trained through audio-visual temporal synchronization. They introduced a geometry-aware temporal aggregation module that computes geometric transformations between frames to align visual features across changing viewpoints. Simultaneously, they incorporated a cascaded feature enhancement module using a sound separation task to disentangle visible sounds from background audio mixtures and direct visual attention to relevant regions. To evaluate the framework, the authors created the Epic Sounding Object benchmark by annotating 3,172 video clips across 30 sounding object classes from the Epic-Kitchens dataset and tested cross-scenario generalization on the Ego4D dataset.

The experimental findings show that the proposed framework substantially outperforms existing methods. On the benchmark's primary localization metric (Consensus Intersection over Union at a 0.2 threshold), the proposed method achieved a score of 38.71%, compared to 26.01% for the best-performing prior approach and 16.51% for a central-focus baseline. The overall Area Under Curve improved to 18.38% versus 15.39% for the previous state of the art. Ablation studies confirmed that geometric alignment contributed the largest single performance gain (increasing CIoU@0.2 from 27.41% to 37.38%), while sound disentanglement and soft visual attention provided vital additional accuracy gains.

These results demonstrate that explicitly accounting for physical camera motion and separating off-screen audio noise are critical for robust first-person perception. In practical terms, this capability provides a path toward lower data-labeling costs through self-supervised training and enhances the reliability of applications such as audio-queried memory retrieval in smart glasses, object state tracking during human-environment interactions, and robotic trajectory forecasting.

Organizations developing egocentric vision and wearable artificial intelligence systems should integrate geometric motion alignment and audio disentanglement into their perception pipelines. Before full commercial deployment, teams should conduct further development to address operational limitations: current geometric estimation relies on image alignment techniques that can fail during extreme camera motion or sudden illumination shifts. Enhancing geometric estimation robustness across diverse real-world lighting and motion conditions represents the key next step for operational deployment.

Cover for Egocentric Audio-Visual Object Localization

Abstract

Humans naturally perceive surrounding scenes by unifying sound and sight from a first-person view. Likewise, machines are advanced to approach human intelligence by learning with multisensory inputs from an egocentric perspective. In this paper, we explore the challenging egocentric audio-visual object localization task and observe that 1) egomotion commonly exists in first-person recordings, even within a short duration; 2) The out-of-view sound components can be created when wearers shift their attention. To address the first problem, we propose a geometry-aware temporal aggregation module that handles the egomotion explicitly. The effect of egomotion is mitigated by estimating the temporal geometry transformation and exploiting it to update visual representations. Moreover, we propose a cascaded feature enhancement module to overcome the second issue. It improves cross-modal localization robustness by disentangling visually-indicated audio representation. During training, we take advantage of the naturally occurring audio-visual temporal synchronization as the “free” self-supervision to avoid costly labeling. We also annotate and create the Epic Sounding Object dataset for evaluation purposes. Extensive experiments show that our method achieves state-of-the-art localization performance in egocentric videos and can be generalized to diverse audio-visual scenes. Code is available at https://github.com/WikiChao/Ego-AV-Loc.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Problem Formulation and Method Overview
  • 3.2. Feature Extraction
  • 3.3. Cascaded Feature Enhancement
  • 3.4. Geometry-Aware Temporal Modeling
  • 3.5. Training Objective
  • 4. The Epic Sounding Object Dataset
  • 5. Experiment
  • 5.1. Results
  • 5.1.1 Experimental Comparison
  • 5.1.2 Ablation Study
  • 5.1.3 Generalization to More Scenarios
  • 6. Discussions and Conclusions
  • References

Knowls

  1. Knowl 1 — Egocentric Audio-Visual Object Localization Formulation and Architecture

    model/method

    Egocentric audio-visual object localization aims to predict binary location maps O={Oi}i=1T\mathcal{O} = \{O_i\}_{i=1}^T indicating visible sound sources in a first-person video clip V={Ii}i=1TV = \{I_i\}_{i=1}^T of TT frames paired with an audio stream ss. In first-person footage, captured audio is often a mixture s=∑n=1Nsns = \sum_{n=1}^N s_n containing out-of-view acoustic events alongside visible ones, and visual frames undergo frequent viewpoint changes due to wearer egomotion.

    The framework processes inputs through two parallel feature extractors:

    1. Visual Feature Extraction: A visual encoder EvE_v (a pre-trained Dilated ResNet without the final classification layer) extracts feature maps vi=Ev(Ii)∈Rc×hv×wvv_i = E_v(I_i) \in \mathbb{R}^{c \times h_v \times w_v} for each frame IiI_i, where cc is the channel dimension and hv×wvh_v \times w_v is the spatial grid resolution.
    2. Audio Feature Extraction: The raw audio stream ss is converted via Short-Time Fourier Transform (STFT) into a magnitude spectrogram XX. A 2D CNN encoder EaE_a extracts an audio representation a=Ea(X)∈Rc×ha×waa = E_a(X) \in \mathbb{R}^{c \times h_a \times w_a} in the Time-Frequency (T-F) domain.

    Extracted features are passed through a cascaded feature enhancement module (to separate visible audio sources from out-of-view distractors and compute soft spatial attention) followed by a geometry-aware temporal aggregation module (to align visual features across viewpoints before temporal pooling).

  2. Knowl 2 — Cascaded Feature Enhancement via Sound Disentanglement and Soft Localization

    model/method

    To prevent out-of-view sounds and background visual clutter from corrupting audio-visual associations, features are updated through a two-stage cascaded enhancement pipeline:

    1. Visually-Guided Audio Disentanglement: During training, an artificial mixture s~=s(1)+s(2)\tilde{s} = s^{(1)} + s^{(2)} is constructed by mixing the target audio s(1)s^{(1)} with an audio clip s(2)s^{(2)} sampled from a different video, yielding mixed spectrogram X~\tilde{X} and mixed audio features a=Ea(X~)a = E_a(\tilde{X}). A global visual vector gv∈Rcg_v \in \mathbb{R}^c is computed by applying spatial average pooling over each visual map viv_i and temporal max pooling across all TT frames. The vector gvg_v is tiled across spatial dimensions ha×wah_a \times w_a and concatenated with aa along the channel dimension. A disentanglement subnetwork f(⋅)f(\cdot) (consisting of two 1×11 \times 1 convolution layers) extracts the visually-indicated audio features: a^=f(Concat[a,Tile(gv)])∈Rc×ha×wa\hat{a} = f(\text{Concat}[a, \text{Tile}(g_v)]) \in \mathbb{R}^{c \times h_a \times w_a} An audio decoder DaD_a predicts a binary separation mask Mpred=Da(a^)M_{\text{pred}} = D_a(\hat{a}). The model is supervised by an ℓ2\ell_2 mask reconstruction loss: Ldis=∥Mpred−Mgt∥22\mathcal{L}_{\text{dis}} = \|M_{\text{pred}} - M_{\text{gt}}\|_2^2 where Mgt(u,v)=[X(1)(u,v)≥X~(u,v)]M_{\text{gt}}(u, v) = [X^{(1)}(u, v) \ge \tilde{X}(u, v)] indicates whether the target audio is dominant at time-frequency bin (u,v)(u, v). At inference, the unmixed stream s=s(1)s = s^{(1)} is used directly.

    2. Soft Visual Localization: An audio vector ga^∈Rcg_{\hat{a}} \in \mathbb{R}^c is produced via max pooling over the time and frequency axes of a^\hat{a}. An audio-visual correlation map SiS_i is calculated by cosine similarity at each spatial location (x,y)(x, y): Si(x,y)=vi(x,y)⋅ga^∥vi(x,y)∥2∥ga^∥2S_i(x, y) = \frac{v_i(x, y) \cdot g_{\hat{a}}}{\|v_i(x, y)\|_2 \|g_{\hat{a}}\|_2} The visual features are then weighted by soft spatial attention to suppress sound-irrelevant regions: v^i=Softmax(Si)⋅vi\hat{v}_i = \text{Softmax}(S_i) \cdot v_i

  3. Knowl 3 — Geometry-Aware Temporal Aggregation Module

    model/method

    To leverage temporal context across neighboring frames while compensating for rapid egocentric viewpoint shifts and object deformations, the framework employs Geometry-Aware Temporal Modeling (GATM):

    1. Geometry Modeling: Between a query frame IiI_i and each supporting frame IjI_j (j∈{1,…,T}j \in \{1, \dots, T\}), a 3×33 \times 3 homography matrix Hji\mathcal{H}_{ji} with 8 degrees of freedom is estimated using SIFT feature matching and RANSAC: Hji=h(Ij,Ii)j→i\mathcal{H}_{ji} = h(I_j, I_i)_{j \to i} After downsampling Hji\mathcal{H}_{ji} to match the visual feature map resolution, the enhanced feature map v^j\hat{v}_j is warped into the coordinate frame of IiI_i via a warping operator ⊗\otimes: v^ji=Hji⊗v^j\hat{v}_{ji} = \mathcal{H}_{ji} \otimes \hat{v}_j

    2. Temporal Aggregation: For query frame IiI_i, the aligned features across all time steps are concatenated along the temporal axis into v^=[v^1i;… ;v^Ti]∈RT×c×hv×wv\boldsymbol{\hat{v}} = [\hat{v}_{1i}; \dots; \hat{v}_{Ti}] \in \mathbb{R}^{T \times c \times h_v \times w_v}. At each spatial location (x,y)(x, y), temporal context is aggregated using a scaled dot-product attention mechanism with residual connection: zi(x,y)=v^i(x,y)+Softmax(v^i(x,y)v^(x,y)Td)v^(x,y)z_i(x, y) = \hat{v}_i(x, y) + \text{Softmax}\left(\frac{\hat{v}_i(x, y) \boldsymbol{\hat{v}}(x, y)^T}{\sqrt{d}}\right) \boldsymbol{\hat{v}}(x, y) where d=cd = c is the feature channel dimension and (⋅)T(\cdot)^T denotes transposition.

  4. Knowl 4 — Multiple-Instance Self-Supervised Audio-Visual Contrastive Objective

    equation

    Given temporally aggregated visual features ziz_i and audio feature vector ga^g_{\hat{a}}, an audio-visual similarity map is computed as Si(x,y)=CosineSim(zi(x,y),ga^)S_i(x, y) = \text{CosineSim}(z_i(x, y), g_{\hat{a}}). A soft sounding objectness map is derived via differential thresholding: Oi=sigmoid((Si−ϵ)/τ)O_i = \text{sigmoid}((S_i - \epsilon)/\tau), where ϵ=0.5\epsilon = 0.5 is the threshold and τ=0.03\tau = 0.03 is the sharpness temperature.

    To account for temporal dynamics across the clip, attention maps S=[S1;… ;ST]\boldsymbol{S} = [S_1; \dots; S_T] are aggregated using soft Multiple-Instance Learning (MIL) pooling: S‾=∑t=1T(Wt⋅S)[:,:,t],where Wt[x,y,:]=Softmax(S[x,y,:])\overline{S} = \sum_{t=1}^T (W_t \cdot \boldsymbol{S})[:, :, t], \quad \text{where } W_t[x, y, :] = \text{Softmax}(\boldsymbol{S}[x, y, :]) The aggregated objectness map O‾\overline{O} is derived identically from S‾\overline{S}.

    For each video sample kk in a training batch of size BB, positive matching energy PkP_k and negative matching energy NkN_k are defined as: Pk=1∣O‾∣⟨O‾,S‾⟩,Nk=1hvwv⟨1,Sneg⟩P_k = \frac{1}{|\overline{O}|} \langle \overline{O}, \overline{S} \rangle, \quad N_k = \frac{1}{h_v w_v} \langle \mathbf{1}, S_{\text{neg}} \rangle where ⟨⋅,⋅⟩\langle \cdot, \cdot \rangle is the Frobenius inner product, 1\mathbf{1} is an all-ones tensor, and SnegS_{\text{neg}} is computed by pairing the visual features with an audio stream from a different video clip.

    The localization loss is given by: Lloc=−1B∑k=1Blog⁡exp⁡(Pk)exp⁡(Pk)+exp⁡(Nk)\mathcal{L}_{\text{loc}} = -\frac{1}{B} \sum_{k=1}^B \log \frac{\exp(P_k)}{\exp(P_k) + \exp(N_k)} The total multi-task learning objective is: L=Lloc+λLdis\mathcal{L} = \mathcal{L}_{\text{loc}} + \lambda \mathcal{L}_{\text{dis}} where λ=5\lambda = 5 balances the audio source separation loss Ldis\mathcal{L}_{\text{dis}}.

  5. Knowl 5 — Epic Sounding Object Dataset

    definition

    The Epic Sounding Object dataset is an egocentric benchmark for evaluating audio-visual sounding object localization, derived from the test set of Epic-Kitchens.

    Dataset creation details:

    • Source and Preprocessing: From 13,000 action recognition videos, silent recordings are filtered out based on full-scale decibel thresholds, and remaining videos are trimmed to central 1-second clips, yielding 5,089 candidate clips.
    • Candidate Proposals: Object candidate bounding boxes across three uniformly selected frames per clip are generated using a Mask R-CNN pre-trained on MS-COCO and an egocentric hand-object detector.
    • Crowdsourced Verification: Amazon Mechanical Turk annotators identify whether out-of-view sounds occur and label the bounding boxes emitting sound. A majority voting rule requiring agreement between at least two annotators filters the candidate pool.
    • Final Statistics: The finalized dataset comprises 3,172 videos, 9,196 annotated frames across 30 sounding object noun classes, evenly split into validation and test partitions. Over 80% of sounding object bounding boxes cover less than 20% of the image area, and approximately 17% of videos contain out-of-view sounds.
  6. Knowl 6 — Quantitative Evaluation on Epic Sounding Object Benchmark

    data/table

    Performance of sounding object localization models evaluated on the Epic Sounding Object dataset using Consensus Intersection over Union (CIoU) at thresholds {0.2, 0.3, 0.4} and Area Under Curve (AUC). All comparative models were re-trained from scratch on the Epic-Kitchens training split under identical self-supervised conditions.

    Method [email protected] [email protected] [email protected] AUC
    Attention 7.12 - - 6.42
    STM 12.10 7.64 4.01 8.87
    Hardway 24.51 13.55 6.10 13.38
    SSPL 13.62 8.10 4.45 9.56
    Mix 26.01 15.25 9.90 15.39
    Ours 38.71 19.42 10.51 18.38

    The proposed framework outperforms the strongest third-person baseline (Mix) by +12.70% on [email protected] and +2.99% on AUC, validating the importance of addressing egomotion and out-of-view sound.

  7. Knowl 7 — Ablation of Temporal Modeling, Soft Localization, and Audio Disentanglement

    data/table

    Ablation study evaluating the individual contributions of temporal aggregation variants (Average, Max, Geometry-Aware Temporal Modeling), Soft Localization (SL), and the audio disentanglement objective (Ldis\mathcal{L}_{\text{dis}}) on the Epic Sounding Object dataset.

    Model TM (Avg) TM (Max) TM (GA) SL Ldis\mathcal{L}_{\text{dis}} [email protected] AUC
    a 27.41 15.36
    b ✓ 31.84 15.79
    c ✓ 33.29 16.10
    d ✓ 37.38 16.59
    f ✓ ✓ 38.21 17.92
    g (Full) ✓ ✓ ✓ 38.71 18.38

    Key takeaways:

    • Incorporating standard temporal aggregation (models b and c) improves [email protected] over the single-frame baseline (model a) by 4.43% to 5.88%.
    • Geometry-Aware Temporal Modeling (model d) provides a significant additional jump (+4.09% [email protected] over Max pooling), demonstrating that geometric alignment explicitly mitigates egomotion distortions.
    • Soft localization (model f) and audio disentanglement via Ldis\mathcal{L}_{\text{dis}} (model g) provide incremental gains up to 38.71% [email protected] and 18.38% AUC by filtering spatial noise and suppressing out-of-view acoustic distractors.
    • A naive center-box Gaussian baseline achieves a [email protected] of only 16.51%.
  8. Knowl 8 — Failure Modes of Geometry-Aware Homography Estimation

    limitation

    The Geometry-Aware Temporal Aggregation Module relies on SIFT feature extraction and RANSAC-based homography matrix computation to model camera movement between frames. In egocentric scenes with severe lighting variations, low-texture surfaces, fast non-planar head rotations, or large parallax caused by 3D translation, homography estimation can fail to establish accurate point correspondences. Under such failures, feature alignment degrades to vanilla unaligned temporal modeling.

Coverage note — None was omitted. All key architectural contributions (GATM, cascaded enhancement, multi-task MIL objective), the Epic Sounding Object dataset, primary benchmarking results, component ablations, and stated limitations are fully represented.

References

  1. 1.Yazan Abu Farha, Alexander Richard, and Juergen Gall. When will you do what?-anticipating temporal occurrences of activities. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5343–5352, 2018.
  2. 2.Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman. Self-supervised learning of audio-visual objects from video. In European Conference on Computer Vision, pages 208–224. Springer, 2020.
  3. 3.Relja Arandjelovic and Andrew Zisserman. Look, listen and learn. In Proceedings of the IEEE International Conference on Computer Vision, pages 609–617, 2017.
  4. 4.Relja Arandjelovic and Andrew Zisserman. Objects that sound. In Proceedings of the European conference on computer vision (ECCV), pages 435–451, 2018.
  5. 5.Yusuf Aytar, Carl Vondrick, and Antonio Torralba. Soundnet: Learning sound representations from unlabeled video. Advances in neural information processing systems, 29, 2016.
  6. 6.David A Bulkin and Jennifer M Groh. Seeing sounds: visual and auditory interactions in the brain. Current opinion in neurobiology, 16(4):415–419, 2006.
  7. 7.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015.
  8. 8.Minjie Cai, Kris M Kitani, and Yoichi Sato. Understanding hand-object manipulation with grasp types and object attributes. In Robotics: Science and Systems, volume 3. Ann Arbor, Michigan;, 2016.
  9. 9.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017.
  10. 10.Alejandro Cartas, Jordi Luque, Petia Radeva, Carlos Segura, and Mariella Dimiccoli. How much does audio matter to recognize egocentric object interactions? arXiv preprint arXiv:1906.00634, 2019.
  11. 11.Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman. Localizing visual sounds the hard way. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16867–16876, 2021.
  12. 12.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the European Conference on Computer Vision (ECCV), pages 720–736, 2018.
  13. 13.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Rescaling egocentric vision. arXiv preprint arXiv:2006.13256, 2020.
  14. 14.Dima Damen, Teesid Leelasawassuk, and Walterio MayolCuevas. You-do, i-learn: Egocentric unsupervised discovery of objects and their modes of interaction towards videobased guidance. Computer Vision and Image Understanding, 149:98–112, 2016.
  15. 15.Jacob Donley, Vladimir Tourbabin, Jung-Suk Lee, Mark Broyles, Hao Jiang, Jie Shen, Maja Pantic, Vamsi Krishna Ithapu, and Ravish Mehra. Easycom: An augmented reality dataset to support algorithms for easy communication in noisy environments. arXiv preprint arXiv:2107.04174, 2021.
  16. 16.Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein. Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. arXiv preprint arXiv:1804.03619, 2018.
  17. 17.Alireza Fathi, Xiaofeng Ren, and James M Rehg. Learning to recognize objects in egocentric activities. In CVPR 2011, pages 3281–3288. IEEE, 2011.
  18. 18.Martin A Fischler and Robert C Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6):381–395, 1981.
  19. 19.Antonino Furnari and Giovanni Maria Farinella. Rollingunrolling lstms for action anticipation from first-person video. IEEE transactions on pattern analysis and machine intelligence, 43(11):4021–4036, 2020.
  20. 20.Chuang Gan, Deng Huang, Hang Zhao, Joshua B Tenenbaum, and Antonio Torralba. Music gesture for visual sound separation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10478– 10487, 2020.
  21. 21.Ruohan Gao, Rogerio Feris, and Kristen Grauman. Learning to separate object sounds by watching unlabeled video. In Proceedings of the European Conference on Computer Vision (ECCV), pages 35–53, 2018.
  22. 22.Ruohan Gao and Kristen Grauman. 2.5 d visual sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 324–333, 2019.
  23. 23.Ruohan Gao and Kristen Grauman. Co-separating sounds of visual objects. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3879–3888, 2019.
  24. 24.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. arXiv preprint arXiv:2110.07058, 2021.
  25. 25.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
  26. 26.Di Hu, Feiping Nie, and Xuelong Li. Deep multimodal clustering for unsupervised audiovisual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9248–9257, 2019.
  27. 27.Di Hu, Rui Qian, Minyue Jiang, Xiao Tan, Shilei Wen, Errui Ding, Weiyao Lin, and Dejing Dou. Discriminative sounding objects localization via self-supervised audiovisual matching. Advances in Neural Information Processing Systems, 33, 2020.
  28. 28.Xixi Hu, Ziyang Chen, and Andrew Owens. Mix and localize: Localizing sound sources in mixtures. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10483–10492, 2022.
  29. 29.Robert A Jacobs and Chenliang Xu. Can multisensory training aid visual learning? a computational investigation. Journal of vision, 19(11):1–1, 2019.
  30. 30.Hao Jiang and Kristen Grauman. Seeing invisible poses: Estimating 3d body pose from egocentric video. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3501–3509. IEEE, 2017.
  31. 31.J Adam Jones, J Edward Swan, Gurjot Singh, Eric Kolstad, and Stephen R Ellis. The effects of virtual reality, augmented reality, and motion parallax on egocentric depth perception. In Proceedings of the 5th symposium on Applied perception in graphics and visualization, pages 9–14, 2008.
  32. 32.Kazuhiko Kawamura, A Bugra Koku, D Mitchell Wilkes, Richard Alan Peters, and Ali Sekmen. Toward egocentric navigation. International Journal of Robotics and Automation, 17(4):135–145, 2002.
  33. 33.Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. Epic-fusion: Audio-visual temporal binding for egocentric action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5492–5501, 2019.
  34. 34.Daekyum Kim, Brian Byunghyun Kang, Kyu Bum Kim, Hyungmin Choi, Jeesoo Ha, Kyu-Jin Cho, and Sungho Jo. Eyes are faster than hands: A soft wearable robot learns user intention from the egocentric view. Science Robotics, 4(26):eaav2949, 2019.
  35. 35.Bruno Korbar, Du Tran, and Lorenzo Torresani. Cooperative learning of audio and video models from self-supervised synchronization. Advances in Neural Information Processing Systems, 31, 2018.
  36. 36.Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman. Discovering important people and objects for egocentric video summarization. In 2012 IEEE conference on computer vision and pattern recognition, pages 1346–1353. IEEE, 2012.
  37. 37.Cheng Li and Kris M Kitani. Pixel-level hand detection in ego-centric videos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3570– 3577, 2013.
  38. 38.Sizhe Li, Yapeng Tian, and Chenliang Xu. Space-time memory network for sounding object localization in videos. arXiv preprint arXiv:2111.05526, 2021.
  39. 39.Yin Li, Alireza Fathi, and James M Rehg. Learning to predict gaze in egocentric video. In Proceedings of the IEEE international conference on computer vision, pages 3216–3223, 2013.
  40. 40.Yin Li, Miao Liu, and James M Rehg. In the eye of beholder: Joint learning of gaze and actions in first person video. In Proceedings of the European conference on computer vision (ECCV), pages 619–635, 2018.
  41. 41.Yanghao Li, Tushar Nagarajan, Bo Xiong, and Kristen Grauman. Ego-exo: Transferring visual representations from third-person to first-person videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6943–6953, 2021.
  42. 42.Yin Li, Zhefan Ye, and James M Rehg. Delving into egocentric actions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 287–295, 2015.
  43. 43.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  44. 44.Yan-Bo Lin, Yu-Jhe Li, and Yu-Chiang Frank Wang. Dual-modality seq2seq network for audio-visual event localization. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2002–2006. IEEE, 2019.
  45. 45.Miao Liu, Siyu Tang, Yin Li, and James M Rehg. Forecasting human-object interaction: joint prediction of motor attention and actions in first person video. In European Conference on Computer Vision, pages 704–721. Springer, 2020.
  46. 46.David G Lowe. Distinctive image features from scale-invariant keypoints. International journal of computer vision, 60(2):91–110, 2004.
  47. 47.Oded Maron and Tomas Lozano-Pérez. A framework for multiple-instance learning. Advances in neural information processing systems, 10, 1997.
  48. 48.Roberto Martin-Martin, Mihir Patel, Hamid Rezatofighi, Abhijeet Shenoi, JunYoung Gwak, Eric Frankel, Amir Sadeghian, and Silvio Savarese. Jrdb: A dataset and benchmark of egocentric robot visual perception of humans in built environments. IEEE transactions on pattern analysis and machine intelligence, 2021.
  49. 49.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2630–2640, 2019.
  50. 50.Himangi Mittal, Pedro Morgado, Unnat Jain, and Abhinav Gupta. Learning state-aware visual representations from audible interactions. In Proceedings of the European conference on computer vision (ECCV), 2022.
  51. 51.Shentong Mo and Pedro Morgado. A closer look at weakly-supervised audio-visual source localization. arXiv preprint arXiv:2209.09634, 2022.
  52. 52.Shentong Mo and Pedro Morgado. Localizing visual sounds the easy way. arXiv preprint arXiv:2203.09324, 2022.
  53. 53.Francesca Morganti, Stefano Stefanini, and Giuseppe Riva. From allo-to egocentric spatial ability in early alzheimer’s disease: a study with virtual reality spatial tasks. Cognitive neuroscience, 4(3-4):171–180, 2013.
  54. 54.Jonathan Munro and Dima Damen. Multi-modal domain adaptation for fine-grained action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 122–132, 2020.
  55. 55.Tushar Nagarajan, Christoph Feichtenhofer, and Kristen Grauman. Grounded human-object interaction hotspots from video. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8688–8697, 2019.
  56. 56.Evonne Ng, Donglai Xiang, Hanbyul Joo, and Kristen Grauman. You2me: Inferring body pose in egocentric video via first and second person interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9890–9900, 2020.
  57. 57.Curtis Northcutt, Shengxin Zha, Steven Lovegrove, and Richard Newcombe. Egocom: A multi-person multi-modal egocentric communications dataset. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020.
  58. 58.Andrew Owens and Alexei A Efros. Audio-visual scene analysis with self-supervised multisensory features. In Proceedings of the European Conference on Computer Vision (ECCV), pages 631–648, 2018.
  59. 59.Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba. Ambient sound provides supervision for visual learning. In European conference on computer vision, pages 801–816. Springer, 2016.
  60. 60.Hyun Soo Park, Jyh-Jing Hwang, Yedong Niu, and Jianbo Shi. Egocentric future localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4697–4705, 2016.
  61. 61.Ivan Poupyrev, Tadao Ichikawa, Suzanne Weghorst, and Mark Billinghurst. Egocentric object manipulation in virtual environments: empirical evaluation of interaction techniques. In Computer graphics forum, volume 17, pages 41–52. Wiley Online Library, 1998.
  62. 62.Rui Qian, Di Hu, Heinrich Dinkel, Mengyue Wu, Ning Xu, and Weiyao Lin. Multiple sound sources localization from coarse to fine. In European Conference on Computer Vision, pages 292–308. Springer, 2020.
  63. 63.Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh McDermott, and Antonio Torralba. Self-supervised audio-visual co-segmentation. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2357–2361. IEEE, 2019.
  64. 64.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4104–4113, 2016.
  65. 65.Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon. Learning to localize sound source in visual scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4358–4366, 2018.
  66. 66.Silvia Serino, Francesca Morganti, Fabio Di Stefano, and Giuseppe Riva. Detecting early egocentric and allocentric impairments deficits in alzheimer’s disease: An experimental study with virtual reality. Frontiers in aging neuroscience, 7:88, 2015.
  67. 67.Ladan Shams and Aaron R Seitz. Benefits of multisensory learning. Trends in cognitive sciences, 12(11):411–417, 2008.
  68. 68.Dandan Shan, Jiaqi Geng, Michelle Shu, and David F Fouhey. Understanding human hands in contact at internet scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9869–9878, 2020.
  69. 69.Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. Charades-ego: A large-scale dataset of paired third and first person videos. arXiv preprint arXiv:1804.09626, 2018.
  70. 70.Krishna Kumar Singh, Kayvon Fatahalian, and Alexei A Efros. Krishnacam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks. In 2016 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 1–9. IEEE, 2016.
  71. 71.Zengjie Song, Yuxi Wang, Junsong Fan, Tieniu Tan, and Zhaoxiang Zhang. Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3222–3231, 2022.
  72. 72.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.
  73. 73.Charles Spence and Sarah Squire. Multisensory integration: maintaining the perception of synchrony. Current Biology, 13(13):R519–R521, 2003.
  74. 74.Yu-Chuan Su and Kristen Grauman. Detecting engagement in egocentric video. In European Conference on Computer Vision, pages 454–471. Springer, 2016.
  75. 75.J Edward Swan, Adam Jones, Eric Kolstad, Mark A Livingston, and Harvey S Smallman. Egocentric depth judgments in optical, see-through augmented reality. IEEE transactions on visualization and computer graphics, 13(3):429–442, 2007.
  76. 76.Yapeng Tian, Di Hu, and Chenliang Xu. Cyclic co-learning of sounding object visual grounding and sound separation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2745–2754, 2021.
  77. 77.Yapeng Tian, Dingzeyu Li, and Chenliang Xu. Unified multisensory perception: Weakly-supervised audio-visual video parsing. In European Conference on Computer Vision, pages 436–454. Springer, 2020.
  78. 78.Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. Audio-visual event localization in unconstrained videos. In Proceedings of the European Conference on Computer Vision (ECCV), pages 247–263, 2018.
  79. 79.Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. Audio-visual event localization in the wild. In IEEE Computer Society Conference on Computer Vision and Pattern Recognition workshops, 2019.
  80. 80.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  81. 81.Sudheendra Vijayanarasimhan, Susanna Ricco, Cordelia Schmid, Rahul Sukthankar, and Katerina Fragkiadaki. Sfm-net: Learning of structure and motion from video. arXiv preprint arXiv:1704.07804, 2017.
  82. 82.Xiaohan Wang, Linchao Zhu, Heng Wang, and Yi Yang. Interactive prototype learning for egocentric action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8168–8177, 2021.
  83. 83.Yu Wu and Yi Yang. Exploring heterogeneous clues for weakly-supervised audio-visual video parsing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1326–1335, 2021.
  84. 84.Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang. Dual attention matching for audio-visual event localization. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6292–6300, 2019.
  85. 85.Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer. Audiovisual slowfast networks for video recognition. arXiv preprint arXiv:2001.08740, 2020.
  86. 86.Fisher Yu, Vladlen Koltun, and Thomas Funkhouser. Dilated residual networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 472–480, 2017.
  87. 87.Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. The sound of pixels. In Proceedings of the European conference on computer vision (ECCV), pages 570–586, 2018.
  88. 88.Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu. Sep-stereo: Visually guided stereophonic audio generation by associating source separation. In European Conference on Computer Vision, pages 52–69. Springer, 2020.
  89. 89.Yipin Zhou and Tamara L Berg. Temporal perception and prediction in ego-centric video. In Proceedings of the IEEE International Conference on Computer Vision, pages 4498–4506, 2015.

Citation

MLA
Huang, C., et al. “Egocentric Audio-Visual Object Localization”. arXiv, 2023, http://arxiv.org/abs/2303.13471v1.
APA
Huang, C., Tian, Y., Kumar, A., & Xu, C. (2023). Egocentric Audio-Visual Object Localization. arXiv. http://arxiv.org/abs/2303.13471v1
Chicago
Huang, C., Y. Tian, A. Kumar, and C. Xu. 2023. “Egocentric Audio-Visual Object Localization”. arXiv. http://arxiv.org/abs/2303.13471v1.
Harvard
Huang, C. et al. (2023) “Egocentric Audio-Visual Object Localization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.13471v1.
Vancouver
1. Huang C, Tian Y, Kumar A, Xu C (2023) Egocentric Audio-Visual Object Localization. arXiv

BibTeX

@article{huang2023egocentric,
  title = {Egocentric Audio-Visual Object Localization},
  author = {Huang, Chao and Tian, Yapeng and Kumar, Anurag and Xu, Chenliang},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.13471v1},
  eprint = {2303.13471}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE