Conditional Generation of Audio from Video via Foley Analogies
Yuexi DuZiyang ChenJustin SalamonBryan C. RussellAndrew Owens
Proposes a self-supervised video-to-audio framework that synthesizes synchronized soundtracks for silent videos by transferring the acoustic timbre of user-provided reference clips to match on-screen actions.
In film and multimedia production, sound designers—known as Foley artists—rarely rely on actual recorded set sounds. Instead, they manipulate unrelated audio sources to achieve a desired artistic effect while precisely matching on-screen movements, such as using coconut shells to mimic horse hooves. While prior automated video-to-sound systems predict original scene audio, they fail to grant creators artistic control. The article addresses this limitation by introducing "conditional Foley generation," a framework designed to synthesize a customized audio track for a silent input video based on an exemplary audio-visual clip provided by a user to specify what the scene should sound like.
The article demonstrates a generative approach using a self-supervised training pretext task. Natural video often features repeated, self-similar actions. Exploiting this property, the model trains on pairs of clips extracted from different timestamps within the same video, learning to extract action timing from the input video while deriving acoustic timbre from the conditional clip. The architecture combines a discrete spectrogram representation (spectrogram VQGAN), a decoder-only transformer to autoregressively predict audio codes, and a neural vocoder (MelGAN) to produce the final waveform. At inference, the system generalizes to cross-video conditions and enhances synchronization by generating multiple candidate outputs and re-ranking them using an automated audio-visual synchronization model. The approach was evaluated on the physical-interaction Greatest Hits dataset and diverse in-the-wild video from CountixAV.
The primary findings demonstrate that the proposed model successfully synthesizes audio that matches the physical material properties of the conditional clip while maintaining accurate temporal alignment with the input video. Automated evaluation shows the re-ranked model achieved a 44.0% material accuracy and a 66.7% action accuracy, substantially outperforming prior unconditional video-to-audio benchmarks (27.2% and 62.5%, respectively). While a naive non-generative onset-transfer baseline scored high on material matching by directly copying audio, it failed to adjust sound dynamics when actions mismatched, scoring only 52.9% on action accuracy. Human perceptual studies confirmed these advantages: participants preferred the re-ranked model over the base model 54.3% of the time for material resemblance and 53.8% for synchronization, while unconditional baselines were preferred less than 20% of the time.
These results demonstrate a viable pathway toward semi-automated, user-in-the-loop sound design. Automating the labor-intensive task of adjusting timing and timbre reduces post-production timelines and costs while keeping artistic control in human hands. Organizations exploring automated creative tooling should consider deploying generative Foley frameworks as assistive plugins for sound editors. To maximize output quality, pipelines should incorporate multi-sample generation with automated synchronization re-ranking. Future initiatives must explore scaling this capability to non-repetitive scenes, background ambience, and complex multi-source audio tracks.
Decision-makers should note certain limitations: the re-ranking synchronization model occasionally suffered performance degradation due to domain shifts between training and test distributions. Furthermore, the technology presents potential misuse risks regarding deceptive video synthesis, emphasizing the ongoing necessity of pairing generative creative tools with robust audio-visual forensic defenses.
- Paper: VideoGPT: Video Generation using VQ-VAE and Transformers, Wilson Yan et al. (2021). VideoGPT introduces the foundational two-stage generative modeling framework combining discrete token quantization via VQ-VAE with autoregressive transformer prediction over code sequences that directly underpins the source paper's spectrogram VQGAN architecture.
- Paper: HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis, Jungil Kong et al. (2020). HiFi-GAN establishes high-fidelity neural vocoding from intermediate time-frequency representations, providing essential context for the source's neural audio synthesis pipeline that inverts spectrogram representations to raw waveforms.
- Paper: AST: Audio Spectrogram Transformer, Yuan Gong et al. (2021). AST demonstrates the effectiveness of processing audio spectrogram patches using standard visual transformer encoders, establishing key architectural principles for modeling time-frequency audio representations.
- Paper: CNN architectures for large-scale audio classification, Shawn Hershey et al. (2016). This foundational study explores processing spectrogram representations directly with visual convolutional architectures, providing historical grounding for the visual-spectrogram translation paradigm used in the source.
- Paper: Contrastive Audio-Visual Masked Autoencoder, Yuan Gong et al. (2023). CAV-MAE extends multimodal audio-visual representation learning by combining contrastive matching and masked autoencoding, providing advanced self-supervised cross-modal alignment mechanisms relevant to Foley synchronization.
- Paper: Wan: Open and Advanced Large-Scale Video Generative Models, Ang Wang et al. (2025). Wan scales video foundation architectures to include synchronized audio generation across complex video-generation tasks, advancing the generative audio-visual synthesis explored in the source.
- Paper: Pengi: An Audio Language Model for Audio Tasks, Soham Deshmukh et al. (2023). Pengi generalizes audio generation and comprehension into a unified language-modeled framework, broadening the discrete token conditioning paradigms utilized in conditional Foley generation.
