Conditional Generation of Audio from Video via Foley Analogies

Yuexi DuZiyang ChenJustin SalamonBryan C. RussellAndrew Owens

article2023CVPR71 citations

Proposes a self-supervised video-to-audio framework that synthesizes synchronized soundtracks for silent videos by transferring the acoustic timbre of user-provided reference clips to match on-screen actions.

Listen

In film and multimedia production, sound designers—known as Foley artists—rarely rely on actual recorded set sounds. Instead, they manipulate unrelated audio sources to achieve a desired artistic effect while precisely matching on-screen movements, such as using coconut shells to mimic horse hooves. While prior automated video-to-sound systems predict original scene audio, they fail to grant creators artistic control. The article addresses this limitation by introducing "conditional Foley generation," a framework designed to synthesize a customized audio track for a silent input video based on an exemplary audio-visual clip provided by a user to specify what the scene should sound like.

The article demonstrates a generative approach using a self-supervised training pretext task. Natural video often features repeated, self-similar actions. Exploiting this property, the model trains on pairs of clips extracted from different timestamps within the same video, learning to extract action timing from the input video while deriving acoustic timbre from the conditional clip. The architecture combines a discrete spectrogram representation (spectrogram VQGAN), a decoder-only transformer to autoregressively predict audio codes, and a neural vocoder (MelGAN) to produce the final waveform. At inference, the system generalizes to cross-video conditions and enhances synchronization by generating multiple candidate outputs and re-ranking them using an automated audio-visual synchronization model. The approach was evaluated on the physical-interaction Greatest Hits dataset and diverse in-the-wild video from CountixAV.

The primary findings demonstrate that the proposed model successfully synthesizes audio that matches the physical material properties of the conditional clip while maintaining accurate temporal alignment with the input video. Automated evaluation shows the re-ranked model achieved a 44.0% material accuracy and a 66.7% action accuracy, substantially outperforming prior unconditional video-to-audio benchmarks (27.2% and 62.5%, respectively). While a naive non-generative onset-transfer baseline scored high on material matching by directly copying audio, it failed to adjust sound dynamics when actions mismatched, scoring only 52.9% on action accuracy. Human perceptual studies confirmed these advantages: participants preferred the re-ranked model over the base model 54.3% of the time for material resemblance and 53.8% for synchronization, while unconditional baselines were preferred less than 20% of the time.

These results demonstrate a viable pathway toward semi-automated, user-in-the-loop sound design. Automating the labor-intensive task of adjusting timing and timbre reduces post-production timelines and costs while keeping artistic control in human hands. Organizations exploring automated creative tooling should consider deploying generative Foley frameworks as assistive plugins for sound editors. To maximize output quality, pipelines should incorporate multi-sample generation with automated synchronization re-ranking. Future initiatives must explore scaling this capability to non-repetitive scenes, background ambience, and complex multi-source audio tracks.

Decision-makers should note certain limitations: the re-ranking synchronization model occasionally suffered performance degradation due to domain shifts between training and test distributions. Furthermore, the technology presents potential misuse risks regarding deceptive video synthesis, emphasizing the ongoing necessity of pairing generative creative tools with robust audio-visual forensic defenses.

arXiv: 2304.08490
  • Paper: VideoGPT: Video Generation using VQ-VAE and Transformers, Wilson Yan et al. (2021). VideoGPT introduces the foundational two-stage generative modeling framework combining discrete token quantization via VQ-VAE with autoregressive transformer prediction over code sequences that directly underpins the source paper's spectrogram VQGAN architecture.
  • Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). HiFi-GAN establishes high-fidelity neural vocoding from intermediate time-frequency representations, providing essential context for the source's neural audio synthesis pipeline that inverts spectrogram representations to raw waveforms.
  • Paper: AST: Audio Spectrogram Transformer, Yuan Gong et al. (2021). AST demonstrates the effectiveness of processing audio spectrogram patches using standard visual transformer encoders, establishing key architectural principles for modeling time-frequency audio representations.
  • Paper: CNN architectures for large-scale audio classification, Shawn Hershey et al. (2016). This foundational study explores processing spectrogram representations directly with visual convolutional architectures, providing historical grounding for the visual-spectrogram translation paradigm used in the source.
  • Paper: Contrastive Audio-Visual Masked Autoencoder, Yuan Gong et al. (2023). CAV-MAE extends multimodal audio-visual representation learning by combining contrastive matching and masked autoencoding, providing advanced self-supervised cross-modal alignment mechanisms relevant to Foley synchronization.
  • Paper: Wan: Open and Advanced Large-Scale Video Generative Models, Ang Wang et al. (2025). Wan scales video foundation architectures to include synchronized audio generation across complex video-generation tasks, advancing the generative audio-visual synthesis explored in the source.
  • Paper: Pengi: An Audio Language Model for Audio Tasks, Soham Deshmukh et al. (2023). Pengi generalizes audio generation and comprehension into a unified language-modeled framework, broadening the discrete token conditioning paradigms utilized in conditional Foley generation.
Cover for Conditional Generation of Audio from Video via Foley Analogies

Abstract

The sound effects that designers add to videos are designed to convey a particular artistic effect and, thus, may be quite different from a scene’s true sound. Inspired by the challenges of creating a soundtrack for a video that differs from its true sound, but that nonetheless matches the actions occurring on screen, we propose the problem of conditional Foley. We present the following contributions to address this problem. First, we propose a pretext task for training our model to predict sound for an input video clip using a conditional audio-visual clip sampled from another time within the same source video. Second, we propose a model for generating a soundtrack for a silent input video, given a user-supplied example that specifies what the video should “sound like”. We show through human studies and automated evaluation metrics that our model successfully generates sound from videos, while varying its output according to the content of a supplied example. Project site: https://xypb.github.io/CondFoleyGen.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Pretext task for conditional prediction
  • 3.2. Conditional sound prediction architecture
  • 4. Experiments
  • 4.1. Experiment Setup
  • 4.2. Automated Timbre Evaluation
  • 4.3. Automated Onset Evaluation
  • 4.4. Perceptual Study
  • 4.5. Qualitative Results
  • 5. Discussion
  • References

Knowls

  1. Knowl 1 — Task Formulation for Conditional Foley Generation

    definition

    Conditional Foley generation via analogy is defined as the task of synthesizing an audio track for a silent input video vqv_q, conditioned on an exemplary audio-visual clip (vc,ac)(v_c, a_c) provided by a user. The objective is to learn a mapping Fθ(vq,vc,ac)F_\theta(v_q, v_c, a_c) parameterized by θ\theta such that the generated soundtrack conveys the physical and timbral properties of the conditional audio-visual pair (vc,ac)(v_c, a_c) (e.g., material acoustic resonance) while temporally aligning with and reflecting the specific physical interactions and actions depicted in the query video vqv_q (e.g., hitting, scratching, and onset timing).

  2. Knowl 2 — Self-Supervised Pretext Task via Intra-Video Clip Pairing

    model/method

    To train the conditional generator FθF_\theta without human-annotated analog pairs, a self-supervised pretext task is constructed by sampling two distinct temporal segments from the same continuous source video. For a video containing repeated or self-similar physical actions, two clips centered at timestamps tt and t+Δtt + \Delta t are extracted: one serves as the conditioning pair (vc,ac)(v_c, a_c), and the other serves as the query silent video vqv_q with corresponding ground-truth audio target aga_g.

    The training objective minimizes a prediction loss:

    L(ag,Fθ(vq,vc,ac))\mathcal{L}(a_g, F_\theta(v_q, v_c, a_c))

    Because both clips originate from the same recording, they frequently share physical materials and environmental acoustics. However, because their specific motion timing differs, the model cannot trivially copy aca_c directly and must learn to transfer the conditioning timbre while synchronizing sound events to the visual actions of vqv_q.

  3. Knowl 3 — Conditional Autoregressive Transformer Architecture for Spectrogram Codes

    model/method

    The conditional vision-to-sound model predicts discrete audio codes conditioned on visual and acoustic token streams.

    Video sequences vqv_q and vcv_c are processed with a ResNet (2+1)D-18 network from which temporal striding is removed to retain the native video frame rate. Spatial average pooling produces per-frame token sequences Tv(vq)T_v(v_q) and Tv(vc)T_v(v_c).

    The conditioning audio waveform aca_c is converted to a log mel spectrogram MSTFT(ac)∈RT×F\text{MSTFT}(a_c) \in \mathbb{R}^{T \times F} and encoded via a 2D CNN encoder EE into a latent grid z^∈RT′×F′×d\hat{z} \in \mathbb{R}^{T' \times F' \times d}. Vector quantization maps each spatial-frequency vector z^tf\hat{z}_{tf} to the nearest codebook vector ck∈{c1,…,cK}c_k \in \{c_1, \dots, c_K\}:

    ztf=q(z^tf)=arg⁡min⁡ck∥z^tf−ck∥z_{tf} = q(\hat{z}_{tf}) = \arg\min_{c_k} \|\hat{z}_{tf} - c_k\|

    Flattening zz in raster-scan order yields the conditional audio token sequence Ta(ac)T_a(a_c).

    The concatenated sequence S=[Tv(vc),Tv(vq),Ta(ac)]S = [T_v(v_c), T_v(v_q), T_a(a_c)] is passed as context to a decoder-only GPT-2 style Transformer pθp_\theta. The target discrete code sequence s∈{0,1,…,K−1}T′×F′s \in \{0, 1, \dots, K-1\}^{T' \times F'} is predicted autoregressively in raster-scan order via cross-entropy loss:

    pθ(s∣vq,vc,ac)=∏ipθ(si∣s<i,vq,vc,ac)p_\theta(s \mid v_q, v_c, a_c) = \prod_{i} p_\theta(s_i \mid s_{<i}, v_q, v_c, a_c)

  4. Knowl 4 — Audio-Visual Synchronization Re-ranking Procedure

    algorithm

    To enhance the temporal synchronization of generated sounds at inference time, a multi-candidate generation and synchronization re-ranking strategy is employed using an external pre-trained audio-visual synchronization model.

    Input: Query video vqv_q, conditional pair (vc,ac)(v_c, a_c), generative model FθF_\theta, synchronization model MsyncM_\text{sync}, candidate count N=100N = 100, offset tolerance τ\tau
    Output: Selected soundtrack a∗a^*
    Initialize candidate set A=∅A = \emptyset
    for j=1j = 1 to NN do
        Sample code sequence s(j)∼pθ(s∣vq,vc,ac)s^{(j)} \sim p_\theta(s \mid v_q, v_c, a_c)
        Synthesize audio waveform a(j)=MelGAN(D(s(j)))a^{(j)} = \text{MelGAN}(D(s^{(j)}))
        Add a(j)a^{(j)} to AA
    end for
    for each candidate a(j)∈Aa^{(j)} \in A do
        Predict temporal alignment offset tjt_j and confidence cjc_j using Msync(vq,a(j))M_\text{sync}(v_q, a^{(j)})
    end for
    Compute minimum absolute offset tmin=min⁡j∣tj∣t_\text{min} = \min_{j} |t_j|
    Filter candidates to form valid subset Avalid={a(j)∈A∣∣tj∣≤tmin+τ}A_\text{valid} = \{a^{(j)} \in A \mid |t_j| \le t_\text{min} + \tau\}
    Select a∗=arg⁡max⁡a(j)∈Avalidcja^* = \arg\max_{a^{(j)} \in A_\text{valid}} c_j
    return a∗a^*
  5. Knowl 5 — Automated Evaluation of Material, Action, and Onset Synchronization on Greatest Hits

    data/table

    The performance of conditional sound generation models is evaluated on the Greatest Hits dataset across material classification accuracy (17 classes via a fine-tuned VGGish classifier matching the conditional clip's material), action classification accuracy (hit vs. scratch matching the query video), onset count accuracy (exact match of number of onsets), and onset synchronization average precision (AP within 0.1 seconds of ground truth).

    Model Material Acc (%) Action Acc (%) Onset
    match mismatch overall match mismatch overall # onset Acc (%) onset sync AP (%)
    Style transfer∗^* 30.0 33.5 32.3 20.8 36.6 31.3 19.1 46.9
    Onset transfer 54.8 51.4 52.6 69.0 44.7 52.9 24.8 71.9
    Chance 5.9 5.9 5.9 50.0 50.0 50.0 – –
    SpecVQGAN 25.4 26.8 26.1 52.3 43.1 46.2 11.3 51.0
    SpecVQGAN - finetuned 29.9 25.7 27.2 70.6 58.4 62.5 25.8 59.3
    Ours - No cond. 21.3 24.9 23.7 61.4 55.1 57.2 24.6 59.3
    Ours - Base 41.1 41.6 41.4 67.5 59.2 62.0 26.5 60.0
    Ours - w/ re-rank 43.4 45.2 44.0 78.2 61.3 66.7 25.3 54.3

    Here, ∗* denotes an oracle model given ground-truth query audio. The results show that unconditional baselines fail to achieve high material accuracy, whereas the full conditional model achieves high material fidelity while retaining action alignment. Synchronization re-ranking further boosts action accuracy (from 62.0% to 66.7%) and material accuracy (from 41.4% to 44.0%).

  6. Knowl 6 — Perceptual Evaluation of Audio-Visual Synchronization and Material Timbre

    data/table

    A human perceptual study conducted on Amazon Mechanical Turk with 376 participants evaluated generated sound quality across two criteria: temporal synchronization with the video, and acoustic resemblance to the conditional material/object.

    Model Variation Task
    Material Chosen (%) ↑\uparrow Sync. Chosen (%) ↑\uparrow
    Style transfer∗^* – 9.9 (±2.3\pm 2.3) 10.6 (±2.4\pm 2.4)
    Onset transfer – 64.7 (±3.8\pm 3.8) 57.3 (±3.9\pm 3.9)
    SpecVQGAN – 16.3 (±2.9\pm 2.9) 18.0 (±3.0\pm 3.0)
    Ours base 50.0 (±0.0\pm 0.0) 50.0 (±0.0\pm 0.0)
    Ours - cond. 35.3 (±3.8\pm 3.8) 40.1 (±4.0\pm 4.0)
    Ours - cond. video 46.1 (±4.0\pm 4.0) 45.0 (±4.0\pm 4.0)
    Ours - augment 51.3 (±3.9\pm 3.9) 49.7 (±4.0\pm 4.0)
    Ours w/ rand. cond. 45.5 (±4.0\pm 4.0) 47.5 (±4.0\pm 4.0)
    Ours + re-rank 54.3 (±3.4\pm 3.4) 53.8 (±3.4\pm 3.4)

    Values represent the preference rate relative to the base model (at 50.0%) reported with 95% confidence intervals. The re-ranked model achieved the highest preference among the generative methods (54.3% material, 53.8% sync). While the non-generative Onset Transfer baseline received higher perceptual scores on clean, discrete hits, it cannot adapt sounds when action types change.

  7. Knowl 7 — Mel-Spectrogram Decoding and Neural Waveform Synthesis

    model/method

    Once the sequence of discrete tokens ss is generated by the transformer, the corresponding codebook vectors are gathered into a quantized latent tensor z∈RT′×F′×dz \in \mathbb{R}^{T' \times F' \times d}. A 2D CNN-based codebook decoder DD reconstructs the log mel spectrogram S^=D(z)\hat{S} = D(z).

    To synthesize the final audible time-domain waveform from S^\hat{S}, the model utilizes a pre-trained MelGAN neural vocoder, which maps mel spectrograms directly to waveforms and avoids the acoustic artifacts associated with deterministic phase reconstruction algorithms like Griffin-Lim.

  8. Knowl 8 — Rule-Based Onset Transfer Baseline

    model/method

    The Onset Transfer baseline provides a non-generative benchmark that directly transfers audio segments from the conditioning clip to the query video. A ResNet (2+1)D model is trained to detect audio onsets from video frames in both query and conditioning videos. Audio snippets centered at random detected onset locations in aca_c are extracted and pasted into the query timeline at the timestamps of detected visual onsets in vqv_q.

    While this approach guarantees high material fidelity when individual sound events are well-separated impacts, it lacks the ability to modulate timbre according to different action types (such as adapting a hit sound into a continuous scratch).

  9. Knowl 9 — Limitations of Conditional Foley Generation

    limitation

    The conditional Foley framework exhibits several limitations:

    1. Synchronization Domain Shift: The pre-trained synchronization model used for re-ranking was trained on general video datasets (e.g., VGGSound), causing a domain shift that occasionally degrades fine-grained onset timing precision on specialized physical interaction datasets.
    2. Action Continuity and Complexity: When applied to complex, in-the-wild videos with overlapping, non-discrete sound events, rule-based baseline approaches fail completely, and generative modeling requires dense visual-acoustic alignment.
    3. Potential for Disinformation: The ability to plausibly synthesize and alter soundtracks to match video actions presents forensic and disinformation risks if applied maliciously to create deceptive multimedia.

Coverage note — None was omitted; all key contributions including problem formulation, self-supervised pretext training, architecture, re-ranking algorithm, baselines, quantitative/perceptual results, and limitations are fully represented.

References

  1. 1.Vanessa Theme Ament. The Foley grail: The art of performing sound for film, games, and animation. Routledge, 2014. 1, 2
  2. 2.Relja Arandjelovic and Andrew Zisserman. Objects that sound. In Proceedings of the European conference on computer vision (ECCV), pages 435–451, 2018. 2
  3. 3.Changan Chen, Ruohan Gao, Paul Calamia, and Kristen Grauman. Visual acoustic matching. arXiv preprint arXiv:2202.06875, 2022. 2
  4. 4.Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. Vggsound: A large-scale audio-visual dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 721–725. IEEE, 2020. 7
  5. 5.Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu. Talking-head generation with rhythmic head motion. In European Conference on Computer Vision, pages 35–51. Springer, 2020. 2
  6. 6.Ziyang Chen, Xixi Hu, and Andrew Owens. Structure from silence: Learning scene structure from ambient sound. In 5th Annual Conference on Robot Learning, 2021. 2
  7. 7.Joon Son Chung, Bong-Jin Lee, and Icksang Han. Who said that?: Audio-visual speaker diarisation of real-world meetings. arXiv preprint arXiv:1906.10042, 2019. 2
  8. 8.Joon Son Chung and Andrew Zisserman. Out of time: automated lip sync in the wild. In Asian conference on computer vision, pages 251–263. Springer, 2016. 2, 4
  9. 9.Chenye Cui, Yi Ren, Jinglin Liu, Rongjie Huang, and Zhou Zhao. Varietysound: Timbre-controllable video to sound generation via unsupervised information disentanglement. arXiv preprint arXiv:2211.10666, 2022. 2
  10. 10.Abe Davis and Maneesh Agrawala. Visual rhythm and beat. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 2532–2535, 2018. 2
  11. 11.Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein. Looking to listen at the cocktail party: A speakerindependent audio-visual model for speech separation. SIGGRAPH, 2018. 2
  12. 12.Ariel Ephrat and Shmuel Peleg. Vid2speech: speech reconstruction from silent video. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5095–5099. IEEE, 2017. 2
  13. 13.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12873–12883, 2021. 2, 3, 4, 1
  14. 14.Chuang Gan, Deng Huang, Peihao Chen, Joshua B Tenenbaum, and Antonio Torralba. Foley music: Learning to generate music from videos. In European Conference on Computer Vision, pages 758–775. Springer, 2020. 2
  15. 15.Ruohan Gao, Rogerio Feris, and Kristen Grauman. Learning to separate object sounds by watching unlabeled video. In Proceedings of the European Conference on Computer Vision (ECCV), pages 35–53, 2018. 2
  16. 16.Ruohan Gao and Kristen Grauman. 2.5 d visual sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 324–333, 2019. 2
  17. 17.Rishabh Garg, Ruohan Gao, and Kristen Grauman. Geometryaware multi-task learning for binaural audio generation from video. arXiv preprint arXiv:2111.10882, 2021. 2
  18. 18.Leon A Gatys, Alexander S Ecker, and Matthias Bethge. A neural algorithm of artistic style. arXiv preprint arXiv:1508.06576, 2015. 5, 7, 8, 2
  19. 19.Leon A Gatys, Alexander S Ecker, and Matthias Bethge. Image style transfer using convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2414–2423, 2016. 2
  20. 20.Sanchita Ghose and John Jeffrey Prevost. Autofoley: Artificial synthesis of synchronized sound tracks for silent videos with deep learning. IEEE Transactions on Multimedia, 23:1895–1907, 2020. 2
  21. 21.Sanchita Ghose and John J Prevost. Foleygan: Visually guided generative adversarial network-based synchronous sound generation in silent videos. arXiv preprint arXiv:2107.09262, 2021. 2
  22. 22.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems, pages 2672–2680, 2014. 2
  23. 23.Daniel Griffin and Jae Lim. Signal estimation from modified short-time fourier transform. IEEE Transactions on acoustics, speech, and signal processing, 32(2):236–243, 1984. 4
  24. 24.John Hershey and Michael Casey. Audio-visual sound separation via hidden markov models. Advances in Neural Information Processing Systems, 14, 2001. 2
  25. 25.Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al. Cnn architectures for large-scale audio classification. arXiv preprint arXiv:1609.09430, 2016. 6, 2
  26. 26.Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, Malcolm Slaney, Ron Weiss, and Kevin Wilson. Cnn architectures for largescale audio classification. In International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2017. 6, 2
  27. 27.Aaron Hertzmann, Charles E Jacobs, Nuria Oliver, Brian Curless, and David H Salesin. Image analogies. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques, pages 327–340, 2001. 1, 2, 8
  28. 28.Sicong Huang, Qiyang Li, Cem Anil, Xuchan Bao, Sageev Oore, and Roger B Grosse. Timbretron: A wavenet (cyclegan (cqt (audio))) pipeline for musical timbre transfer. arXiv preprint arXiv:1811.09620, 2018. 2
  29. 29.Vladimir Iashin and Esa Rahtu. Taming visually guided sound generation. arXiv preprint arXiv:2110.08791, 2021. 1, 2, 3, 4, 5, 7, 8
  30. 30.Vladimir Iashin, Weidi Xie, Esa Rahtu, and Andrew Zisserman. Sparse in space and time: Audio-visual synchronisation with trainable selectors. arXiv preprint arXiv:2210.07055, 2022. 2, 4, 7
  31. 31.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017. 3, 1
  32. 32.Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution, 2016. 3, 1
  33. 33.Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. International Conference on Learning Representation, 2015. 1, 2
  34. 34.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. 1
  35. 35.A Sophia Koepke, Olivia Wiles, Yael Moses, and Andrew Zisserman. Sight to sound: An end-to-end approach for visual piano transcription. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1838–1842. IEEE, 2020. 2
  36. 36.Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brebisson, ´ Yoshua Bengio, and Aaron C Courville. Melgan: Generative adversarial networks for conditional waveform synthesis. Advances in neural information processing systems, 32, 2019. 4
  37. 37.Timothy R Langlois and Doug L James. Inverse-foley animation: synchronizing rigid-body motions to sound. ACM Transactions on Graphics (TOG), 33(4):41, 2014. 2
  38. 38.Seung Hyun Lee, Wonseok Roh, Wonmin Byeon, Sang Ho Yoon, Chan Young Kim, Jinkyu Kim, and Sangpil Kim. Sound-guided semantic image manipulation. arXiv preprint arXiv:2112.00007, 2021. 2
  39. 39.Tingle Li, Yichen Liu, Andrew Owens, and Hang Zhao. Learning visual styles from audio-visual associations. arXiv, 2022. 2
  40. 40.Javier Nistal, Stefan Lattner, and Gael Richard. Drumgan: Synthesis of drum sounds with timbral feature conditioning using generative adversarial networks. arXiv preprint arXiv:2008.12073, 2020. 2
  41. 41.Andrew Owens and Alexei A Efros. Audio-visual scene analysis with self-supervised multisensory features. European Conference on Computer Vision (ECCV), 2018. 2, 4
  42. 42.Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman. Visually indicated sounds. CVPR, 2016. 1, 2, 5, 6, 7, 8
  43. 43.Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman. Visually indicated sounds. In Computer Vision and Pattern Recognition (CVPR), 2016. 7
  44. 44.KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. Learning individual speaking styles for accurate lip to speech synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13796–13805, 2020. 2
  45. 45.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. 2019. 4
  46. 46.Colin Raffel, Brian McFee, Eric J Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, Daniel PW Ellis, and C Colin Raffel. Mir eval: A transparent implementation of common mir metrics. In ISMIR, pages 367–372, 2014. 2
  47. 47.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021. 4
  48. 48.Tim Sainburg, Marvin Thielk, and Timothy Q Gentner. Finding, visualizing, and quantifying latent structure across diverse animal vocal repertoires. PLoS computational biology, 16(10):e1008228, 2020. 1
  49. 49.Eli Shechtman and Michal Irani. Matching local selfsimilarities across images and videos. In 2007 IEEE Conference on Computer Vision and Pattern Recognition, pages 1–8. IEEE, 2007. 3
  50. 50.Kun Su, Xiulong Liu, and Eli Shlizerman. Multiinstrumentalist net: Unsupervised generation of music from body movements. arXiv preprint arXiv:2012.03478, 2020. 2
  51. 51.Kun Su, Xiulong Liu, and Eli Shlizerman. How does it sound? Advances in Neural Information Processing Systems, 34, 2021. 2
  52. 52.D´ıdac Sur´ıs, Carl Vondrick, Bryan Russell, and Justin Salamon. It’s time for artistic correspondence in music and video. CVPR, 2022. 2
  53. 53.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6450–6459, 2018. 4, 5, 2
  54. 54.Efthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey, Tal Remez, Daniel PW Ellis, and John R Hershey. Into the wild with audioscope: Unsupervised audio-visual separation of on-screen sounds. arXiv preprint arXiv:2011.01143, 2020. 2
  55. 55.Dmitry Ulyanov. Audio texture synthesis and style transfer. https : / / dmitryulyanov . github . io / audio - texture - synthesis - and - style - transfer, 2016. 2, 5, 7, 8
  56. 56.Kees Van Den Doel, Paul G Kry, and Dinesh K Pai. Foleyautomatic: physically-based sound effects for interactive simulation and animation. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques, pages 537–544. ACM, 2001. 2
  57. 57.Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017. 2, 3, 4, 1
  58. 58.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017. 2
  59. 59.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Neural Information Processing Systems (NIPS), 2017. 4
  60. 60.Prateek Verma and Julius O Smith. Neural style transfer for audio spectograms. arXiv preprint arXiv:1801.01589, 2018. 2
  61. 61.Yu Wang, Nicholas J Bryan, Justin Salamon, Mark Cartwright, and Juan Pablo Bello. Who calls the shots? rethinking fewshot learning for audio. In 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pages 36–40. IEEE, 2021. 2
  62. 62.Yu Wang, Justin Salamon, Nicholas J Bryan, and Juan Pablo Bello. Few-shot sound event detection. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 81–85. IEEE, 2020. 2
  63. 63.Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin. Visually informed binaural audio generation without binaural audios. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15485–15494, 2021. 2
  64. 64.Karren Yang, Bryan Russell, and Justin Salamon. Telling left from right: Learning spatial correspondence of sight and sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9932–9941, 2020. 2
  65. 65.Richard Zhang, Jun-Yan Zhu, Phillip Isola, Xinyang Geng, Angela S Lin, Tianhe Yu, and Alexei A Efros. Real-time userguided image colorization with learned deep priors. arXiv preprint arXiv:1705.02999, 2017. 2
  66. 66.Yunhua Zhang, Ling Shao, and Cees GM Snoek. Repetitive activity counting by sight and sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14070–14079, 2021. 2, 3, 5, 6, 8, 1
  67. 67.Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. The sound of pixels. In Proceedings of the European conference on computer vision (ECCV), pages 570–586, 2018. 2
  68. 68.Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg. Visual to sound: Generating natural sound for videos in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3550–3558, 2018. 1, 2
  69. 69.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. 2017. 2

Citation

MLA
Du, Y., et al. “Conditional Generation of Audio from Video via Foley Analogies”. arXiv, 2023, http://arxiv.org/abs/2304.08490v1.
APA
Du, Y., Chen, Z., Salamon, J., Russell, B., & Owens, A. (2023). Conditional Generation of Audio from Video via Foley Analogies. arXiv. http://arxiv.org/abs/2304.08490v1
Chicago
Du, Y., Z. Chen, J. Salamon, B. Russell, and A. Owens. 2023. “Conditional Generation of Audio from Video via Foley Analogies”. arXiv. http://arxiv.org/abs/2304.08490v1.
Harvard
Du, Y. et al. (2023) “Conditional Generation of Audio from Video via Foley Analogies”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.08490v1.
Vancouver
1. Du Y, Chen Z, Salamon J, Russell B, Owens A (2023) Conditional Generation of Audio from Video via Foley Analogies. arXiv

BibTeX

@article{du2023conditional,
  title = {Conditional Generation of Audio from Video via Foley Analogies},
  author = {Du, Yuexi and Chen, Ziyang and Salamon, Justin and Russell, Bryan and Owens, Andrew},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.08490v1},
  eprint = {2304.08490}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE