A Closer Look at Weakly-Supervised Audio-Visual Source Localization

Shentong MoPedro Morgado

article2022NeurIPS85 citations

Exposes critical flaws in standard audio-visual localization benchmarks by introducing negative-sample evaluation protocols alongside SLAVC, a framework combining momentum encoders and extreme visual dropout to prevent overfitting and eliminate reliance on annotated early stopping.

Listen

Artificial intelligence systems that identify the visual source of a sound in video footage are vital for applications in automated surveillance, multimedia indexing, and robotics. Because manually annotating bounding boxes for sounding objects across massive video libraries is prohibitively expensive, the field has turned to weakly-supervised methods that learn directly from natural audio-visual pairings. However, existing development and evaluation practices rely on two unrealistic assumptions: models are permitted to use fully annotated validation sets to stop training early before performance collapses, and benchmarks assume that a visible sound source is present in every video frame. In realistic deployments where off-screen sounds and silent objects are common, these assumptions conceal severe performance flaws.

The article demonstrates that standard weakly-supervised visual sound source localization models suffer from extreme overfitting and high false-positive rates when visible sound sources are absent. To resolve these issues, the article establishes a rigorous evaluation benchmark and introduces a novel training framework called Simultaneous Localization and Audio-Visual Correspondence (SLAVC).

To create a realistic evaluation testbed, the researchers expanded two standard benchmarks—Flickr SoundNet and VGG-Sound Sources—by adding negative samples consisting of off-screen sounds, silent objects, and mismatched audio-video pairs. They eliminated early stopping, requiring models to train to full convergence. The article then evaluated prior methods alongside the proposed SLAVC framework, which incorporates extreme visual feature dropout and slow-moving momentum target encoders to prevent overfitting, while simultaneously evaluating regional localization and cross-instance correspondence to suppress false positive detections.

The findings show that prior state-of-the-art models rapidly deteriorate during training without manual early stopping, with localization accuracy dropping sharply after just two to three epochs. When evaluated on datasets containing negative samples, previous approaches produced high false-positive rates, frequently hallucinating sound sources when none existed. In contrast, the SLAVC framework achieved stable convergence without early stopping and set a new performance benchmark. On the extended VGG-Sound Sources dataset, SLAVC achieved an Average Precision of 32.95% (increasing to 34.46% when combined with object-guided localization), compared to 24.55% for the prior leading method and near-zero scores for earlier baselines.

These results demonstrate that standard weakly-supervised audio-visual localization models are ill-suited for real-world deployment unless explicitly regularized against spurious visual alignments. By eliminating the hidden dependency on annotated validation subsets for early stopping, the SLAVC framework lowers deployment costs and reduces false-positive risks in automated monitoring systems where off-screen noise is prevalent.

Organizations developing or deploying audio-visual localization systems should adopt evaluation protocols that include negative samples and measure performance at full convergence. Engineering teams should integrate heavy visual dropout and momentum encoders into their training pipelines while using dual-branch correspondence objectives to guard against false alarms. Further research and development are recommended to improve the detection of small objects and achieve high-precision bounding-box quality, where all evaluated systems still experience noticeable accuracy degradation.

Confidence in these findings is high, supported by systematic evaluations across large datasets totaling over 144,000 training pairs and diverse test splits. However, stakeholders should note that the current approach relies on vision backbones pre-trained on object recognition data and that localization precision drops substantially when attempting to pinpoint small or highly precise object boundaries.

Cover for A Closer Look at Weakly-Supervised Audio-Visual Source Localization

Abstract

Audio-visual source localization is a challenging task that aims to predict the location of visual sound sources in a video. Since collecting ground-truth annotations of sounding objects can be costly, a plethora of weakly-supervised localization methods that can learn from datasets with no bounding-box annotations have been proposed in recent years, by leveraging the natural co-occurrence of audio and visual signals. Despite significant interest, popular evaluation protocols have two major flaws. First, they allow for the use of a fully annotated dataset to perform early stopping, thus significantly increasing the annotation effort required for training. Second, current evaluation metrics assume the presence of sound sources at all times. This is of course an unrealistic assumption, and thus better metrics are necessary to capture the model's performance on (negative) samples with no visible sound sources. To accomplish this, we extend the test set of popular benchmarks, Flickr SoundNet and VGG-Sound Sources, in order to include negative samples, and measure performance using metrics that balance localization accuracy and recall. Using the new protocol, we conducted an extensive evaluation of prior methods, and found that most prior works are not capable of identifying negatives and suffer from significant overfitting problems (rely heavily on early stopping for best results). We also propose a new approach for visual sound source localization that addresses both these problems. In particular, we found that, through extreme visual dropout and the use of momentum encoders, the proposed approach combats overfitting effectively, and establishes a new state-of-the-art performance on both Flickr SoundNet and VGG-Sound Source. Code and pre-trained models are available at https://github.com/stoneMo/SLAVC.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 3 Visual Source Localization
  • 3.1 Preliminaries: Weakly-Supervised Visual Source Localization
  • 3.2 Simultaneous Localization and Audio-Visual Correspondence (SLAVC)
  • 4 Benchmarking Visual Source Localization
  • 5 Experiments
  • 5.1 Experimental setup
  • 5.2 Main results
  • 5.3 Analysis
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Simultaneous Localization and Audio-Visual Correspondence Framework

    model/method

    The Simultaneous Localization and Audio-Visual Correspondence (SLAVC) framework is a weakly-supervised visual sound source localization (VSL) architecture designed to eliminate training overfitting and reject false positive detections when sounding objects are absent from a video frame.

    SLAVC addresses spurious audio-visual alignments and false positive detections via three core mechanisms:

    1. Subspace Factorization: The model projects audio representations aia_i and visual feature maps viv_i into two separate subspaces via linear projection heads: a localization subspace (gaextloc,gvextloc)(g_a^{ ext{loc}}, g_v^{ ext{loc}}) and an audio-visual correspondence subspace (gaextavc,gvextavc)(g_a^{ ext{avc}}, g_v^{ ext{avc}}). Localization produces spatial probabilities over image coordinates to isolate candidate sound regions, while correspondence produces instance probabilities across the mini-batch to suppress visual regions better explained by other audio samples.
    2. Extreme Visual Dropout: A high dropout rate (p=0.9p = 0.9) is applied exclusively to the visual encoder feature maps fv(vi)f_v(v_i), while no dropout is applied to the audio encoder fa(ai)f_a(a_i). This prevents the multiple-instance contrastive objective from locking onto spurious spatial artifacts by forcing spatial representations to be redundant across image locations.
    3. Momentum Target Encoders: Exponential moving average (EMA) target encoders f^a\hat{f}_a and f^v\hat{f}_v with parameter momentum m=0.999m = 0.999 compute smooth target representations a^i\hat{a}_i and v^ixy\hat{v}_i^{xy}, stabilizing contrastive learning and preventing representation collapse during extended training.
  2. Knowl 2 — SLAVC Probability Factorization and Contrastive Optimization Objective

    equation

    Let aia_i denote the global audio feature vector for sample ii, and let vjxyv_j^{xy} denote the visual feature vector at spatial coordinates (x,y)(x, y) for image jj. Let s(u,w)=uopw∥u∥2∥w∥2s(u, w) = \frac{u^ op w}{\|u\|_2 \|w\|_2} be the cosine similarity, and τ\tau be the temperature hyperparameter.

    The localization probability map Ploc(ai,vjxy)P^{\text{loc}}(a_i, v_j^{xy}) applies a spatial softmax ρxy\rho_{xy} over spatial coordinates x,yx, y:

    Ploc(ai,vjxy)=exp⁡(1τs(galoc(ai),gvloc(vjxy)))∑x′,y′exp⁡(1τs(galoc(ai),gvloc(vjx′y′)))P^{\text{loc}}(a_i, v_j^{xy}) = \frac{\exp\left(\frac{1}{\tau} s\left(g_a^{\text{loc}}(a_i), g_v^{\text{loc}}(v_j^{xy})\right)\right)}{\sum_{x', y'} \exp\left(\frac{1}{\tau} s\left(g_a^{\text{loc}}(a_i), g_v^{\text{loc}}(v_j^{x'y'})\right)\right)}

    The audio-visual correspondence probability map Pavc(ai,vjxy)P^{\text{avc}}(a_i, v_j^{xy}) applies an instance softmax ρi\rho_i over all BB batch elements:

    Pavc(ai,vjxy)=exp⁡(1τs(gaavc(ai),gvavc(vjxy)))∑k=1Bexp⁡(1τs(gaavc(ak),gvavc(vjxy)))P^{\text{avc}}(a_i, v_j^{xy}) = \frac{\exp\left(\frac{1}{\tau} s\left(g_a^{\text{avc}}(a_i), g_v^{\text{avc}}(v_j^{xy})\right)\right)}{\sum_{k=1}^B \exp\left(\frac{1}{\tau} s\left(g_a^{\text{avc}}(a_k), g_v^{\text{avc}}(v_j^{xy})\right)\right)}

    The joint prediction map P(ai,vjxy)P(a_i, v_j^{xy}) is the point-wise product of localization and correspondence:

    P(ai,vjxy)=Ploc(ai,vjxy)⋅Pavc(ai,vjxy)P(a_i, v_j^{xy}) = P^{\text{loc}}(a_i, v_j^{xy}) \cdot P^{\text{avc}}(a_i, v_j^{xy})

    Using momentum target representations a^i=f^a(ai)\hat{a}_i = \hat{f}_a(a_i) and v^jxy=f^v(vjxy)\hat{v}_j^{xy} = \hat{f}_v(v_j^{xy}), the model optimizes the bidirectional multi-instance contrastive loss Lifull\mathcal{L}_i^{\text{full}} per sample ii:

    Lifull=−log⁡max⁡xyP(ai,v^ixy)∑j=1Bmax⁡xyP(ai,v^jxy)−log⁡max⁡xyP(a^i,vixy)∑k=1Bmax⁡xyP(a^k,vixy)\mathcal{L}_i^{\text{full}} = - \log \frac{\max_{xy} P(a_i, \hat{v}_i^{xy})}{\sum_{j=1}^B \max_{xy} P(a_i, \hat{v}_j^{xy})} - \log \frac{\max_{xy} P(\hat{a}_i, v_i^{xy})}{\sum_{k=1}^B \max_{xy} P(\hat{a}_k, v_i^{xy})}

  3. Knowl 3 — SLAVC Inference Map and Negative Detection Confidence

    model/method

    During inference, SLAVC computes localization using momentum audio representations a^\hat{a} and momentum visual feature maps v^\hat{v}. For an input audio-visual pair (v,a)(v, a), the spatial audio-visual localization map SAVLxyS_{\text{AVL}}^{xy} at spatial coordinates (x,y)(x, y) is evaluated by summing similarities across both localization and correspondence heads:

    SAVLxy=s(galoc(a^),gvloc(v^xy))+s(gaavc(a^),gvavc(v^xy))S_{\text{AVL}}^{xy} = s\left(g_a^{\text{loc}}(\hat{a}), g_v^{\text{loc}}(\hat{v}^{xy})\right) + s\left(g_a^{\text{avc}}(\hat{a}), g_v^{\text{avc}}(\hat{v}^{xy})\right)

    where s(u,w)=u⊤w∥u∥2∥w∥2s(u, w) = \frac{u^\top w}{\|u\|_2 \|w\|_2} denotes cosine similarity.

    To determine whether a sounding source is visible in the frame (distinguishing positive from negative samples), the model generates a global confidence score did_i defined as the maximum spatial similarity response:

    di=max⁡xySAVLxyd_i = \max_{xy} S_{\text{AVL}}^{xy}

    If an object prior map is available via Object-Guided Localization (OGL), SAVLxyS_{\text{AVL}}^{xy} is linearly combined with the OGL prior map to yield enhanced localization maps.

  4. Knowl 4 — Extended Flickr-SoundNet and Extended VGG-Sound Sources Test Benchmarks

    experimental setup

    Standard VSL benchmarks assume visible sounding objects in every test frame. To evaluate false positive rejection, Extended Flickr-SoundNet and Extended VGG-Sound Sources (Extended VGG-SS) test sets augment standard benchmarks with non-visible sound/silent frame negative samples:

    1. Real Negatives: Test clips manually verified to contain non-audible audio or off-screen/non-visible sound sources (42 samples for Flickr-SoundNet; 379 samples for VGG-SS).
    2. Automated Easy Negatives: Unrelated audio-visual pairs created by pairing frames with audio tracks from completely different semantic categories (169 samples for Flickr-SoundNet; 3,594 samples for VGG-SS).
    3. Automated Hard Negatives: Audio-visual pairs mismatched from distinct videos sharing the same class label (39 samples for Flickr-SoundNet; 1,185 samples for VGG-SS).

    Extended Flickr-SoundNet contains 500 total test samples (250 positive, 250 negative). Extended VGG-SS contains 10,316 total test samples (5,158 positive, 5,158 negative).

    Samples are categorized by ground-truth bounding box area (in pixels) into four size tiers: Small (1−3221 - 32^2), Medium (322−96232^2 - 96^2), Large (962−144296^2 - 144^2), and Huge (1442−2242144^2 - 224^2).

  5. Knowl 5 — Evaluation Protocol and Metrics for VSL with Negative Audio-Visual Pairs

    definition

    Given an audio-visual sample ii with binary ground truth ci∈{0,1}c_i \in \{0, 1\} (11 indicating a visible sound source and 00 indicating no visible sound source), prediction confidence di=max⁡xySAVLxyd_i = \max_{xy} S_{\text{AVL}}^{xy}, consensus Intersection-over-Union uiu_i between the predicted localization bounding box and human annotations, confidence threshold δ\delta, and cIoU threshold γ\gamma (default γ=0.5\gamma = 0.5):

    • True Positives (TP): TP(δ)={i∣ci=1,di>δ,ui>γ}\text{TP}(\delta) = \{i \mid c_i = 1, d_i > \delta, u_i > \gamma\}
    • False Positives (FP): FP(δ)={i∣ci=1,di>δ,ui≤γ}∪{i∣ci=0,di>δ}\text{FP}(\delta) = \{i \mid c_i = 1, d_i > \delta, u_i \le \gamma\} \cup \{i \mid c_i = 0, d_i > \delta\}
    • False Negatives (FN): FN(δ)={i∣ci=1,di≤δ}\text{FN}(\delta) = \{i \mid c_i = 1, d_i \le \delta\}

    From these sets, Precision(δ)=∣TP(δ)∣∣TP(δ)∣+∣FP(δ)∣(\delta) = \frac{|\text{TP}(\delta)|}{|\text{TP}(\delta)| + |\text{FP}(\delta)|} and Recall(δ)=∣TP(δ)∣∣TP(δ)∣+∣FN(δ)∣(\delta) = \frac{|\text{TP}(\delta)|}{|\text{TP}(\delta)| + |\text{FN}(\delta)|} are determined.

    Evaluation is performed using three metrics:

    1. Average Precision (AP): Area under the Precision-Recall curve obtained by sweeping δ\delta.
    2. max-F1: Peak F1 score across all confidence thresholds: max⁡δ2⋅Precision(δ)⋅Recall(δ)Precision(δ)+Recall(δ)\max_\delta \frac{2 \cdot \text{Precision}(\delta) \cdot \text{Recall}(\delta)}{\text{Precision}(\delta) + \text{Recall}(\delta)}.
    3. Localization Accuracy (LocAcc): Proportion of positive instances (ci=1c_i = 1) achieving ui>γu_i > \gamma.
  6. Knowl 6 — Localization Accuracy With and Without Early Stopping

    data/table

    Prior weakly-supervised VSL methods rely heavily on early stopping based on annotated validation data, masking severe overfitting during training. When trained to convergence (20 epochs on VGG-Sound 144k) without early stopping, performance in prior models degrades sharply. In contrast, SLAVC avoids overfitting and improves with prolonged training.

    Method Flickr-SoundNet VGG-SS
    Early Stop NO Early Stop Early Stop NO Early Stop
    Attention10k 42.26 34.16 18.52 14.04
    CoarsetoFine – 47.20 – 21.93
    DMC 55.60 52.80 23.90 22.63
    AVObject – – 29.70 –
    DSOL 74.00 72.91 29.91 26.87
    LVS 71.60 19.60 33.36 10.43
    HardPos 76.80 – 34.60 –
    EZ-VSL 79.60 66.40 34.28 31.58
    SLAVC (ours) 83.20 83.60 37.22 37.79
    EZ-VSL + OGL 83.94 72.80 38.85 37.86
    SLAVC (ours) + OGL 86.40 86.00 39.67 39.80

    Without early stopping, LVS drops from 71.60% to 19.60% LocAcc on Flickr-SoundNet and from 33.36% to 10.43% on VGG-SS. SLAVC increases performance from 83.20% to 83.60% on Flickr-SoundNet and 37.22% to 37.79% on VGG-SS when trained to convergence.

  7. Knowl 7 — Benchmark Evaluation on Extended Flickr-SoundNet and Extended VGG-SS

    data/table

    Evaluation of models trained on VGG-Sound 144k without early stopping across the Extended Flickr-SoundNet and Extended VGG-SS benchmarks under AP, max-F1, and LocAcc:

    Method Extended Flickr-SoundNet Extended VGG-SS
    AP max-F1 LocAcc AP max-F1 LocAcc
    Center Prior – – 67.60 – – 34.16
    CoarsetoFine 0.00 38.20 47.20 0.00 19.80 21.93
    LVS 9.80 17.90 19.60 5.15 9.90 10.43
    Attention10k 15.98 24.00 34.16 6.70 13.10 14.04
    DMC 25.56 41.80 52.80 11.53 20.30 22.63
    DSOL 38.32 49.40 72.91 16.84 25.60 26.87
    OGL 40.20 55.70 77.20 18.73 30.90 36.58
    EZ-VSL 46.30 54.60 66.40 24.55 30.90 31.58
    SLAVC (ours) 51.63 59.10 83.60 32.95 40.00 37.79
    EZ-VSL + OGL 48.75 56.80 72.80 27.71 34.60 37.86
    SLAVC (ours) + OGL 52.15 60.10 86.00 34.46 41.50 39.80

    SLAVC outperforms prior methods on both datasets across all metrics. SLAVC + OGL establishes the top results with 52.15% AP / 60.10% max-F1 / 86.00% LocAcc on Extended Flickr-SoundNet and 34.46% AP / 41.50% max-F1 / 39.80% LocAcc on Extended VGG-SS.

  8. Knowl 8 — Ablation of Dropout, Target Momentum, and Subspace Decomposition

    empirical result

    Ablation experiments conducted on VGG Sound Sources (144k training subset) demonstrate the individual contributions of visual dropout, audio dropout, momentum target encoders, and dual-subspace decomposition:

    1. Visual Dropout (pvdropp_{\text{vdrop}}): Heavy visual dropout is essential to prevent overfitting. Setting pvdrop=0.90p_{\text{vdrop}} = 0.90 yields 32.95% AP, 40.00% max-F1, and 37.79% LocAcc, compared to 26.03% AP, 32.00% max-F1, and 36.45% LocAcc at pvdrop=0p_{\text{vdrop}} = 0.
    2. Audio Dropout (padropp_{\text{adrop}}): Applying dropout to audio encoder outputs harms performance. Increasing padropp_{\text{adrop}} from 00 to 0.950.95 causes AP to collapse monotonically from 32.95% to 7.63% and LocAcc from 37.79% to 10.64%.
    3. Momentum Encoder Rate (mm): Target momentum m=0.999m = 0.999 achieves optimal representation stability (32.95% AP, 40.00% max-F1, 37.79% LocAcc), outperforming training without momentum target encoders (m=0m = 0: 25.03% AP, 32.50% max-F1, 32.45% LocAcc).
    4. SLAVC Factorization: Joint training with both localization (AVLoc) and audio-visual correspondence (AVC) heads and using both during inference achieves 32.95% AP / 40.00% max-F1 / 37.79% LocAcc, whereas inference using only AVLoc yields 22.01% AP / 23.65% LocAcc, and inference using only AVC yields 17.29% AP / 34.72% LocAcc.
  9. Knowl 9 — Negative Rejection Performance Across Sample Difficulties

    data/table

    Max-F1 performance evaluated on balanced subsets pairing positive test samples with specific subsets of negative samples (Real Negatives, Automated Easy Negatives, and Automated Hard Negatives) on Extended Flickr-SoundNet and Extended VGG-SS:

    Method Extended Flickr-SoundNet Extended VGG-SS
    Real Neg Automated Easy Neg Automated Hard Neg Real Neg Automated Easy Neg Automated Hard Neg
    LVS 14.50 17.80 14.30 8.30 8.80 7.60
    Attention10k 27.10 28.10 27.00 9.10 7.40 7.90
    CoarsetoFine 36.90 39.40 35.80 22.10 19.80 18.90
    DMC 48.80 41.40 43.70 20.70 19.40 21.90
    EZ-VSL 52.60 55.70 54.20 32.80 36.10 31.40
    SLAVC (ours) 63.50 57.40 63.30 40.30 43.20 33.00

    SLAVC consistently attains the highest max-F1 across all negative difficulty types on both datasets, showing balanced discrimination between visible sound sources and real or artificially mismatched audio-visual pairs.

  10. Knowl 10 — Weakly-Supervised VSL Limitations on Small Objects and Strict IoU Thresholds

    limitation

    Existing weakly-supervised visual sound localization methods, including SLAVC, exhibit two severe performance limitations:

    1. Small Sound Source Failure: When evaluated on objects categorized by area, all existing methods (Center Prior, OGL, EZ-VSL, and SLAVC) fail completely on the Small subset (ground-truth bounding box area ≤322\le 32^2 pixels), achieving a peak Localization Accuracy of 0%0\%.
    2. High-Precision IoU Degradation: As the required consensus IoU evaluation threshold γ\gamma increases from 0.30.3 to 0.750.75, Average Precision (AP) drops monotonically across all methods. At γ=0.75\gamma = 0.75, the AP of all evaluated methods approaches zero, indicating that current models cannot produce tight, high-precision bounding box localizations under weak supervision.

Coverage note — None was omitted. All primary contributions—including the SLAVC framework, mathematical equations, inference mechanism, extended benchmarks, negative-aware evaluation protocol, empirical results with/without early stopping, ablations, and diagnostic limitations—are fully captured in self-contained knowls.

References

  1. 1.Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon. Learning to localize sound source in visual scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4358–4366, 2018.
  2. 2.Di Hu, Feiping Nie, and Xuelong Li. Deep multimodal clustering for unsupervised audiovisual learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 9248–9257, 2019.
  3. 3.Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman. Self-supervised learning of audio-visual objects from video. In Proceedings of European Conference on Computer Vision (ECCV), pages 208–224, 2020.
  4. 4.Rui Qian, Di Hu, Heinrich Dinkel, Mengyue Wu, Ning Xu, and Weiyao Lin. Multiple sound sources localization from coarse to fine. In Proceedings of European Conference on Computer Vision (ECCV), pages 292–308, 2020.
  5. 5.Shentong Mo and Pedro Morgado. Localizing visual sounds the easy way. arXiv preprint arXiv:2203.09324, 2022.
  6. 6.Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman. Localizing visual sounds the hard way. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16867–16876, 2021.
  7. 7.Di Hu, Rui Qian, Minyue Jiang, Xiao Tan, Shilei Wen, Errui Ding, Weiyao Lin, and Dejing Dou. Discriminative sounding objects localization via self-supervised audiovisual matching. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), pages 10077–10087, 2020.
  8. 8.Xian Liu, Rui Qian, Hang Zhou, Di Hu, Weiyao Lin, Ziwei Liu, Bolei Zhou, and Xiaowei Zhou. Visual sound localization in the wild by cross-modal interference erasing. In Proceedings of 36th AAAI Conference on Artificial Intelligence, 2022.
  9. 9.Arda Senocak, Hyeonggon Ryu, Junsik Kim, and In So Kweon. Learning sound localization better from semantically similar samples. In Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022.
  10. 10.Virginia R de Sa. Learning classification with unlabeled data. Advances in neural information processing systems, pages 112–112, 1994.
  11. 11.Yusuf Aytar, Carl Vondrick, and Antonio Torralba. Soundnet: Learning sound representations from unlabeled video. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2016.
  12. 12.Andrew Owens, Jiajun Wu, Josh H. McDermott, William T. Freeman, and Antonio Torralba. Ambient sound provides supervision for visual learning. In Proceedings of the European Conference on Computer Vision (ECCV), pages 801–816, 2016.
  13. 13.Limin Wang, Yuanjun Xiong, Dahua Lin, and Luc Van Gool. Untrimmednets for weakly supervised action recognition and detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4325–4334, 2017.
  14. 14.Pedro Morgado, Nuno Nvasconcelos, Timothy Langlois, and Oliver Wang. Self-supervised generation of spatial audio for 360°video. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2018.
  15. 15.Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. The sound of pixels. In Proceedings of the European Conference on Computer Vision (ECCV), pages 570–586, 2018.
  16. 16.Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba. The sound of motions. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1735–1744, 2019.
  17. 17.Yapeng Tian, Dingzeyu Li, and Chenliang Xu. Unified multisensory perception: Weakly-supervised audio-visual video parsing. In Proceedings of European Conference on Computer Vision (ECCV), page 436–454, 2020.
  18. 18.Chuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum, and Antonio Torralba. Music gesture for visual sound separation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10478–10487, 2020.
  19. 19.Pedro Morgado, Nuno Vasconcelos, and Ishan Misra. Audio-visual instance discrimination with cross-modal agreement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12475–12486, June 2021.
  20. 20.Relja Arandjelovic and Andrew Zisserman. Look, listen and learn. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 609–617, 2017.
  21. 21.Pedro Morgado, Ishan Misra, and Nuno Vasconcelos. Robust audio-visual instance discrimination. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12934–12945, 2021.
  22. 22.Bruno Korbar, Du Tran, and Lorenzo Torresani. Cooperative learning of audio and video models from self-supervised synchronization. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2018.
  23. 23.Andrew Owens and Alexei A. Efros. Audio-visual scene analysis with self-supervised multisensory features. In Proceedings of the European Conference on Computer Vision (ECCV), pages 631–648, 2018.
  24. 24.Pedro Morgado, Yi Li, and Nuno Nvasconcelos. Learning representations from audio-visual spatial alignment. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), pages 4733–4744, 2020.
  25. 25.Ruohan Gao, Rogerio Feris, and Kristen Grauman. Learning to separate object sounds by watching unlabeled video. In Proceedings of the European Conference on Computer Vision (ECCV), pages 35–53, 2018.
  26. 26.Ruohan Gao and Kristen Grauman. Co-separating sounds of visual objects. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 3879–3888, 2019.
  27. 27.Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, and Lorenzo Torresani. Listen to look: Action recognition by previewing audio. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10457–10467, 2020.
  28. 28.Yapeng Tian, Di Hu, and Chenliang Xu. Cyclic co-learning of sounding object visual grounding and sound separation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2745–2754, 2021.
  29. 29.Ruohan Gao and Kristen Grauman. Visualvoice: Audio-visual speech separation with cross-modal consistency. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15495–15505, 2021.
  30. 30.Ruohan Gao and Kristen Grauman. 2.5d visual sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 324–333, 2019.
  31. 31.Changan Chen, Unnat Jain, Carl Schissler, S. V. A. Garí, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman. Soundspaces: Audio-visual navigation in 3d environments. In Proceedings of European Conference on Computer Vision (ECCV), pages 17–36, 2020.
  32. 32.Ziyang Chen, David F. Fouhey, and Andrew Owens. Sound localization by self-supervised time delay estimation. arXiv preprint arXiv:2204.12489, 2022.
  33. 33.John Hershey and Javier Movellan. Audio vision: Using audio-visual synchrony to locate sounds. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 1999.
  34. 34.John W Fisher III, Trevor Darrell, William Freeman, and Paul Viola. Learning joint statistical models for audio-visual fusion and segregation. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), 2000.
  35. 35.Einat Kidron, Yoav Y Schechner, and Michael Elad. Pixels that sound. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2005.
  36. 36.Efthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey, Tal Remez, Dan Ellis, and John R. Hershey. Into the wild with audioscope: Unsupervised audio-visual separation of on-screen sounds. In International Conference on Learning Representations, 2021.
  37. 37.Efthymios Tzinis, Scott Wisdom, Tal Remez, and John R Hershey. Audioscopev2: Audio-visual attention architectures for calibrated open-domain on-screen sound separation. arXiv preprint arXiv:2207.10141, 2022.
  38. 38.Hakan Bilen and Andrea Vedaldi. Weakly supervised deep detection networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2846–2854, 2016.
  39. 39.Krishna Kumar Singh and Yong Jae Lee. Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization. In Proceedings of the IEEE International Conference on Computer Vision, pages 3524–3533, 2017.
  40. 40.Yanzhao Zhou, Yi Zhu, Qixiang Ye, Qiang Qiu, and Jianbin Jiao. Weakly supervised instance segmentation using class peak response. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3791–3800, 2018.
  41. 41.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2921–2929, 2016.
  42. 42.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014.
  43. 43.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. Advances in neural information processing systems, 30, 2017.
  44. 44.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729–9738, 2020.
  45. 45.Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. Vggsound: A large-scale audio-visual dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 721–725. IEEE, 2020.
  46. 46.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016.
  47. 47.Jia Deng, Wei Dong, Richard Socher, Li-Jia. Li, Kai Li, and Li Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 248–255, 2009.
  48. 48.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  49. 49.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS), pages 8026–8037, 2019.

Citation

MLA
Mo, S., and P. Morgado. “A Closer Look at Weakly-Supervised Audio-Visual Source Localization”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 37524–36, https://proceedings.neurips.cc/paper_files/paper/2022/file/f3f2ff9579ba6deeb89caa2fe1f0b99c-Paper-Conference.pdf.
APA
Mo, S., & Morgado, P. (2022). A Closer Look at Weakly-Supervised Audio-Visual Source Localization. Advances in Neural Information Processing Systems, 35, 37524–37536. https://proceedings.neurips.cc/paper_files/paper/2022/file/f3f2ff9579ba6deeb89caa2fe1f0b99c-Paper-Conference.pdf
Chicago
Mo, S., and P. Morgado. 2022. “A Closer Look at Weakly-Supervised Audio-Visual Source Localization”. Advances in Neural Information Processing Systems 35: 37524–36. https://proceedings.neurips.cc/paper_files/paper/2022/file/f3f2ff9579ba6deeb89caa2fe1f0b99c-Paper-Conference.pdf.
Harvard
Mo, S. and Morgado, P. (2022) “A Closer Look at Weakly-Supervised Audio-Visual Source Localization”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 37524–37536. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/f3f2ff9579ba6deeb89caa2fe1f0b99c-Paper-Conference.pdf.
Vancouver
1. Mo S, Morgado P (2022) A Closer Look at Weakly-Supervised Audio-Visual Source Localization. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 37524–37536

BibTeX

@inproceedings{mo2022closer,
  title = {A Closer Look at Weakly-Supervised Audio-Visual Source Localization},
  author = {Mo, Shentong and Morgado, Pedro},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {37524-37536},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/f3f2ff9579ba6deeb89caa2fe1f0b99c-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors