AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis

Susan LiangChao HuangYapeng TianAnurag KumarChenliang Xu

article2023NeurIPS77 citations

Presents a multimodal Neural Radiance Field framework that synthesizes synchronized novel-view video frames and binaural spatial audio along arbitrary camera trajectories by integrating sound propagation physics and 3D visual geometry.

Listen

Immersive virtual experiences in augmented and virtual reality require synchronized, realistic visuals and spatial audio. While recent advancements in neural radiance fields allow machines to render realistic visual viewpoints from novel positions, existing systems struggle to generate corresponding spatial sound in real-world settings. Previous acoustic synthesis approaches relied on synthetic impulse responses, discrete positions, or ground-truth imagery, limiting their ability to produce continuous, realistic audio-visual environments along arbitrary paths.

The article demonstrates AV-NeRF, a neural field framework designed to synthesize both video frames and matching binaural audio along arbitrary, novel camera trajectories in real-world environments. The system evaluates whether neural representations can integrate spatial acoustics with visual cues to generate consistent, multi-view audio-visual scenes directly from raw recordings.

The approach integrates visual scene rendering with an acoustic generation branch informed by environmental physics. The framework uses a visual neural field to capture 3D geometry and surface appearance, passing rendered depth and color features through an audio-visual mapping module to inform acoustic predictions. A specialized coordinate transformation calculates the listener's head orientation relative to the sound source rather than using absolute coordinates. To train and evaluate the framework, the authors assembled a real-world dataset comprising 3.8 hours of synchronized video and binaural audio recorded across indoor and outdoor environments, supplemented by experiments on a synthetic acoustic benchmark.

Evaluation shows that the proposed approach consistently outperforms existing methods across diverse environments. On the real-world dataset, the system achieved a 34% to 39% relative improvement in time-frequency magnitude accuracy and an 8% to 11.6% improvement in time-domain envelope accuracy compared to prior neural acoustic representations. Ablation experiments confirmed that both the audio-visual feature mapping and relative direction encoding were essential to performance, with their removal significantly degrading generation quality. On the synthetic benchmark, the framework delivered a 21% improvement in reverberation decay estimation over previous state-of-the-art baselines. Additional testing showed that the model successfully extended to complex environments with multiple sound sources.

These findings demonstrate that machines can generate coherent, high-fidelity multimodal environments directly from standard recordings without requiring expensive, specialized acoustic simulation pipelines. This capability lowers the cost and complexity of developing immersive simulations, digital twins, and virtual reality content. However, the system requires careful data capture, as performance degrades in environments with heavy ambient background noise or visually uniform surfaces that lack clear geometric features. Additionally, the framework does not explicitly model complex reverberation effects, and models must be trained separately for each individual scene.

Decision-makers investing in multimodal media pipelines should consider adopting neural field frameworks that jointly model visual geometry and spatial acoustics. Before deploying these models commercially, engineering teams should conduct pilot implementations to develop noise-filtering pipelines and test cross-scene generalization to avoid retraining from scratch for every new space. Future development must also establish ethical guidelines and watermarking protocols to mitigate the risk of creating deceptive synthetic media.

arXiv: 2302.02088
Cover for AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis

Abstract

Can machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task—real-world audio-visual scene synthesis—and a first-of-its-kind NeRF-based approach for multimodal learning. Concretely, given a video recording of an audio-visual scene, the task is to synthesize new videos with spatial audios along arbitrary novel camera trajectories in that scene. We propose an acoustic-aware audio generation module that integrates prior knowledge of audio propagation into NeRF, in which we implicitly associate audio generation with the 3D geometry and material properties of a visual environment. Furthermore, we present a coordinate transformation module that expresses a view direction relative to the sound source, enabling the model to learn sound source-centric acoustic fields. To facilitate the study of this new task, we collect a high-quality Real-World Audio-Visual Scene (RWAVS) dataset. We demonstrate the advantages of our method on this real-world dataset and the simulation-based SoundSpaces dataset. We recommend that readers visit our project page for convincing comparisons: https://liangsusan-git.github.io/project/avnerf/.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Task Definition
  • 4 Method
  • 4.1 V-NeRF
  • 4.2 A-NeRF
  • 4.3 AV-Mapper
  • 4.4 Coordinate Transformation
  • 4.5 Learning Objective
  • 5 Experiments
  • 5.1 Datasets
  • 5.2 Results on RWAVS Dataset
  • 5.3 Results on SoundSpaces Dataset
  • 6 Discussion
  • References
  • A Additional Visualization Results
  • B Failure Cases
  • C Rationality of AV-Mapper
  • D Necessity of Distance and Direction Coordinates
  • E Multiple Sound Sources
  • F Architectures
  • G Implementation Details
  • H Setup of RWAVS Dataset

Knowls

  1. Knowl 1 — Real-World Audio-Visual Scene Synthesis Task Definition

    definition

    The real-world audio-visual scene synthesis task requires generating synchronized novel visual frames and corresponding binaural audios along arbitrary continuous camera trajectories within a physical environment.

    Formally, let EE be a static environment with a known sound source location (Sx,Sy)(S_x, S_y), and let O={O1,O2,…,ON}\mathcal{O} = \{O_1, O_2, \dots, O_N\} denote a set of NN training observations. Each observation Oi=(pi,as,at,Ii)O_i = (p_i, a_s, a_t, I_i) contains the listener's camera pose p=(x,y,z,θ,ϕ)p = (x, y, z, \theta, \phi), an input single-channel (mono) source audio clip asa_s, the recorded two-channel binaural audio clip ata_t, and the captured visual image II. The objective is to learn a mapping function ff that synthesizes a novel binaural audio at∗a_t^* and a novel viewpoint image I∗I^* for any unseen target camera pose query p∗=(x∗,y∗,z∗,θ∗,ϕ∗)p^* = (x^*, y^*, z^*, \theta^*, \phi^*) and arbitrary source audio as∗a_s^*:

    (at∗,I∗)=f(p∗,as∗∣O,E)(a_t^*, I^*) = f(p^*, a_s^* \mid \mathcal{O}, E)

    During inference, raw training observations O\mathcal{O} are inaccessible, requiring ff to rely entirely on learned neural scene and acoustic fields.

  2. Knowl 2 — A-NeRF Neural Acoustic Field and Spatial Audio Synthesis Pipeline

    model/method

    A-NeRF models acoustic propagation as a continuous neural field that maps 2D spatial listener coordinates (x,y)(x, y), frequency queries f∈[0,F]f \in [0, F] (where FF is the number of Short-Time Fourier Transform frequency bins), and a relative viewing angle θ′\theta' into two acoustic masks: a mixture mask mm∈Rm_m \in \mathbb{R} and a channel difference mask md∈[−1,1]m_d \in [-1, 1]. Height coordinate zz and pitch angle ϕ\phi are omitted based on empirical observation that spatial audio variation is dominant along the horizontal plane.

    A-NeRF comprises two Multi-Layer Perceptrons (MLPs):

    1. The first MLP takes the position (x,y)(x, y) and frequency query ff (encoded via positional encodings up to frequency 10) along with visual context embeddings from AV-Mapper, outputting a mixture mask mmm_m for frequency ff and a latent scene embedding vector in Rc\mathbb{R}^c (where c=128c=128 for real-world data and c=256c=256 for synthetic data).
    2. The second MLP concatenates this feature vector with the relative orientation embedding Θ′∈Rc\Theta' \in \mathbb{R}^c and predicts the binaural channel difference mask md∈[−1,1]m_d \in [-1, 1] using a Sigmoid and linear scaling layer.

    Synthesizing the output binaural audio ata_t from an input mono audio asa_s follows these steps:

    1. Apply Short-Time Fourier Transform (STFT) on asa_s to extract spectrogram magnitude ss∈RF×Ws_s \in \mathbb{R}^{F \times W} (where WW is the number of time frames) and source phase.
    2. Compute mixture magnitude: sm=ss⊙mms_m = s_s \odot m_m.
    3. Compute channel difference magnitude: sd=sm⊙mds_d = s_m \odot m_d.
    4. Derive raw left and right magnitude spectrograms: sl=sm+sds_l = s_m + s_d and sr=sm−sds_r = s_m - s_d.
    5. Refine sls_l and srs_r through 2D convolutional layers.
    6. Apply Inverse STFT using the refined magnitudes and the source audio phase to produce the time-domain binaural audio ata_t.
  3. Knowl 3 — Sound Source-Centric Relative Coordinate Transformation

    model/method

    Standard Neural Radiance Fields parameterize viewing direction (θ,ϕ)(\theta, \phi) in an absolute world coordinate frame. In spatial acoustics, however, human auditory localization depends on the listener's head orientation relative to the sound source position rather than absolute world coordinates.

    Given the listener's 2D position (x,y)(x, y), absolute viewing direction angle θ\theta, and known 2D sound source position (Sx,Sy)(S_x, S_y), the relative coordinate transformation defines two vectors in R2\mathbb{R}^2:

    • The source direction vector: V1=(Sx−x,Sy−y)V_1 = (S_x - x, S_y - y)
    • The listener orientation vector: V2=(cos⁡θ,sin⁡θ)V_2 = (\cos\theta, \sin\theta)

    The sound source-centric relative direction is calculated as the signed angular difference:

    θ′=∠(V1,V2)\theta' = \angle(V_1, V_2)

    To inject θ′∈[0∘,360∘)\theta' \in [0^\circ, 360^\circ) into the acoustic network without relying on standard high-frequency sinusoidal positional encodings, a learnable codebook Θ∈R4×c\Theta \in \mathbb{R}^{4 \times c} is defined for four canonical basis directions (0∘,90∘,180∘,270∘0^\circ, 90^\circ, 180^\circ, 270^\circ). The continuous angle θ′\theta' is projected by linearly interpolating between the nearest canonical embeddings to yield direction representation Θ′∈Rc\Theta' \in \mathbb{R}^c, which is provided to the linear layers of A-NeRF.

  4. Knowl 4 — AV-Mapper for Integrating Visual Geometry and Material Priors into A-NeRF

    model/method

    Acoustic wave propagation is physically determined by environmental 3D geometry (which dictates reflection and diffraction paths) and material properties (which govern absorption and acoustic impedance). AV-Mapper extracts this information from the visual neural radiance field (V-NeRF) to guide acoustic synthesis in A-NeRF.

    For a target camera pose p=(x,y,z,θ,ϕ)p = (x, y, z, \theta, \phi):

    1. V-NeRF renders an RGB image IrgbI_{\text{rgb}} and a depth map IdepthI_{\text{depth}} using volumetric rendering: C(r)=∫tntfT(t)σ(r(t))c(r(t),d)dt,D(r)=∫tntfT(t)σ(r(t))tdtC(r) = \int_{t_n}^{t_f} T(t) \sigma(r(t)) c(r(t), d) dt, \quad D(r) = \int_{t_n}^{t_f} T(t) \sigma(r(t)) t dt where T(t)=exp⁡(−∫tntσ(r(s))ds)T(t) = \exp\left(-\int_{t_n}^t \sigma(r(s)) ds\right), σ\sigma is volume density, and cc is view-dependent color.
    2. IrgbI_{\text{rgb}} and IdepthI_{\text{depth}} are resized to 256×256256 \times 256, center-cropped to 224×224224 \times 224, and passed through a frozen ImageNet-pretrained ResNet-18 feature extractor, producing two 512-dimensional feature vectors.
    3. The 1024-dimensional concatenated visual vector is processed by AV-Mapper, a 3-layer MLP with ReLU activations, projecting it to an environment embedding e∈Rce \in \mathbb{R}^c.
    4. The embedding ee is added directly to the input of each linear layer in A-NeRF.
  5. Knowl 5 — Learning Objectives for AV-NeRF

    equation

    The visual branch V-NeRF is trained using photometric rendering loss identical to vanilla NeRF:

    LV=∑r∈R∥C(r)−C^(r)∥2\mathcal{L}_V = \sum_{r \in \mathcal{R}} \|C(r) - \hat{C}(r)\|^2

    where C(r)C(r) is the ground-truth pixel color along ray rr, C^(r)\hat{C}(r) is the volume-rendered color, and R\mathcal{R} is the batch of sampled camera rays.

    Following the convergence of V-NeRF, A-NeRF and AV-Mapper are trained jointly via an L2L_2 reconstruction loss over the STFT spectrogram magnitudes:

    LA=∥sm−s^m∥2+∥sl−s^l∥2+∥sr−s^r∥2\mathcal{L}_A = \|s_m - \hat{s}_m\|^2 + \|s_l - \hat{s}_l\|^2 + \|s_r - \hat{s}_r\|^2

    where s^m,s^l,s^r∈RF×W\hat{s}_m, \hat{s}_l, \hat{s}_r \in \mathbb{R}^{F \times W} represent the ground-truth spectrogram magnitudes for the channel mixture, left audio channel, and right audio channel, respectively, and sm,sl,srs_m, s_l, s_r are their corresponding network predictions.

  6. Knowl 6 — Real-World Audio-Visual Scene (RWAVS) Benchmark Performance

    data/table

    Performance comparison on the real-world RWAVS dataset across four distinct indoor and outdoor environments. Audio synthesis quality is evaluated using Magnitude Distance (MAG, measuring time-frequency domain error via L2L_2 spectrogram magnitude distance) and Envelope Distance (ENV, measuring time-domain signal envelope error via Hilbert transformation ∥hilbert(aprd)−hilbert(agt)∥2\|\text{hilbert}(a_{\text{prd}}) - \text{hilbert}(a_{\text{gt}})\|^2). Lower is better for both metrics.

    Methods Office House Apartment Outdoors Overall
    MAG ENV MAG ENV MAG ENV MAG ENV MAG ENV
    Mono-Mono 9.269 0.411 11.889 0.424 15.120 0.474 13.957 0.470 12.559 0.445
    Mono-Energy 1.536 0.142 4.307 0.180 3.911 0.192 1.634 0.127 2.847 0.160
    Stereo-Energy 1.511 0.139 4.301 0.180 3.895 0.191 1.612 0.124 2.830 0.159
    INRAS 1.405 0.141 3.511 0.182 3.421 0.201 1.502 0.130 2.460 0.164
    NAF 1.244 0.137 3.259 0.178 3.345 0.193 1.284 0.121 2.283 0.157
    ViGAS 1.049 0.132 2.502 0.161 2.600 0.187 1.169 0.121 1.830 0.150
    Ours (AV-NeRF) 0.930 0.129 2.009 0.155 2.230 0.184 0.845 0.111 1.504 0.145

    AV-NeRF outperforms all baseline and prior methods across all individual environments and overall metrics, achieving an overall MAG reduction of 39% relative to INRAS and 18% relative to ViGAS, alongside superior envelope tracking.

  7. Knowl 7 — Ablation Analysis of AV-NeRF Components and Parameterizations

    data/table

    Ablation studies on the RWAVS benchmark evaluating the architectural components of AV-NeRF across overall Magnitude Distance (MAG) and Envelope Distance (ENV).

    Component Ablation Multimodal Fusion Direction Encoding
    Methods MAG ENV Methods MAG ENV Methods MAG ENV
    Baseline 2.287 0.157 Concat Input 1.507 0.145 Absolute Direction 1.701 0.149
    Ours w/o AV 1.791 0.150 Add Input 1.504 0.145 Relative Direction 1.508 0.145
    Ours w/o CT 1.701 0.149 Add All Layers 1.505 0.145 Relative Embedding 1.504 0.145
    Ours (Full) 1.504 0.145

    Key takeaways:

    1. Removing the AV-Mapper (AV) degrades MAG from 1.504 to 1.791 (16% degradation); removing Coordinate Transformation (CT) degrades MAG to 1.701.
    2. Adding the visual embedding ee to the input layer performs marginally better than concatenating or adding to all linear layers.
    3. Relative learnable directional embeddings outshine standard sinusoidal positional encodings applied to either relative angles (MAG 1.508) or absolute angles (MAG 1.701).
  8. Knowl 8 — Room Impulse Response Estimation on the SoundSpaces Dataset

    data/table

    Evaluation of AV-NeRF adapted for Room Impulse Response (RIR) synthesis on the synthetic SoundSpaces dataset across six simulated indoor environments. Evaluated metrics are:

    • T60T_{60} error (%): relative error in the time required for sound energy to decay by 60 dB, defined as ∣T60(aprd)−T60(agt)∣/T60(agt)|T_{60}(a_{\text{prd}}) - T_{60}(a_{\text{gt}})| / T_{60}(a_{\text{gt}}).
    • C50C_{50} error (dB): absolute error in early-to-late energy clarity ratio, ∣C50(aprd)−C50(agt)∣|C_{50}(a_{\text{prd}}) - C_{50}(a_{\text{gt}})|.
    • Early Decay Time (EDT) error (sec): absolute error in initial reverberation decay rate, ∣EDT(aprd)−EDT(agt)∣|\text{EDT}(a_{\text{prd}}) - \text{EDT}(a_{\text{gt}})|.
    Methods T60 (%) ↓\downarrow C50 (dB) ↓\downarrow EDT (sec) ↓\downarrow
    Opus-nearest 10.10 3.58 0.115
    Opus-linear 8.64 3.13 0.097
    AAC-nearest 9.35 1.67 0.059
    AAC-linear 7.88 1.68 0.057
    NAF 3.18 1.06 0.031
    INRAS 3.14 0.60 0.019
    Ours (AV-NeRF) 2.47 0.57 0.016

    AV-NeRF achieves state-of-the-art results on all RIR metrics, attaining a 21% relative error reduction in T60T_{60}, 5% in C50C_{50}, and 16% in EDT compared to the prior state of the art (INRAS).

  9. Knowl 9 — Real-World Audio-Visual Scene (RWAVS) Dataset Setup

    experimental setup

    The Real-World Audio-Visual Scene (RWAVS) dataset is a multimodal benchmark designed for novel-view audio-visual synthesis across continuous trajectories in real-world spaces.

    Dataset specifications and capture setup:

    • Hardware: A 3Dio Free Space XLR binaural microphone for stereo audio, a TASCAM DR-60DMKII field recorder, a GoPro Max video camera, and an LG XBOOM 360 omnidirectional speaker playing music continuously.
    • Diversity: Collected across four categories—office, house, apartment, and outdoors—with varied speaker placements per scene, totaling 232 minutes (3.8 hours) of recording covering continuous 360∘360^\circ viewer orientations.
    • Pre-processing: Video keyframes are extracted at 1 frame per second and paired with 1-second synchronized binaural and mono clips. Camera poses are recovered via COLMAP structure-from-motion. Audio clips undergo background noise suppression via Adobe Audition.
    • Splits: 9,850 training samples and 2,469 validation samples (an 80%/20% split).
  10. Knowl 10 — AV-NeRF Extension to Multi-Source Acoustic Environments

    empirical result

    AV-NeRF can be extended to synthesize scenes containing multiple independent sound sources by stacking parallel A-NeRF branches, where each branch separately parameterizes the acoustic field and generates source-centric acoustic masks for one specific sound source.

    In experiments conducted on two RWAVS multi-source scenes containing two simultaneous active speakers, AV-NeRF achieves superior audio quality compared to state-of-the-art baselines:

    Methods Multi-source MAG ↓\downarrow Multi-source ENV ↓\downarrow
    Mono-Mono 1.949 0.172
    Mono-Energy 0.533 0.075
    Stereo-Energy 0.527 0.073
    INRAS 0.472 0.078
    NAF 0.401 0.080
    Ours (AV-NeRF) 0.282 0.063

    AV-NeRF improves multi-source MAG to 0.282 (a 29.7% improvement over NAF) and ENV to 0.063 (a 21.3% improvement over NAF).

  11. Knowl 11 — Limitations of AV-NeRF

    limitation

    AV-NeRF possesses several notable limitations:

    1. Scene-Specific Optimization: Similar to original NeRF, a distinct AV-NeRF model must be trained from scratch for every individual audio-visual scene, lacking cross-scene generalization or zero-shot transfer.
    2. Static Sound Sources: The core formulation assumes static sound sources with known 2D coordinates, restricting direct applicability to moving or unlocalized sound sources without explicit multi-branch replication.
    3. Vulnerability to Ambient Acoustic Noise: Real-world background noises (e.g., HVAC units, refrigerators, outdoor ambient sounds) recorded in the training targets corrupt the learned acoustic neural masks.
    4. Non-Informative Visual Failure Cases: When camera viewpoints capture untextured, featureless surfaces (e.g., blank whiteboards or plain walls) or suffer heavy occlusions, AV-Mapper fails to extract meaningful geometric and material features, causing erroneous acoustic mask predictions.
    5. Lack of Explicit Reverberation Synthesis: AV-NeRF primarily predicts distance energy attenuation and binaural channel differences in real-world audio, without explicitly modeling late reverberation impulse responses.

Coverage note — None was omitted; all key contributions including problem definition, A-NeRF architecture, relative coordinate transformation, AV-Mapper, loss formulations, RWAVS dataset details, benchmark results, SoundSpaces evaluations, ablations, multi-source extensions, and limitations are fully covered.

References

  1. 1.Andrew Owens and Alexei A Efros. Audio-visual scene analysis with self-supervised multisensory features. In ECCV, pages 631–648, 2018.
  2. 2.Shentong Mo and Pedro Morgado. Localizing visual sounds the easy way. In Shai Avidan, Gabriel J. Brostow, Moustapha Cisse, Giovanni Maria Farinella, and Tal Hassner, editors, ECCV, volume 13697, pages 218–234, 2022.
  3. 3.Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. Audio-visual event localization in unconstrained videos. In ECCV, pages 247–263, 2018.
  4. 4.Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman. Soundspaces: Audio-visual navigation in 3d environments. In ECCV, 2020.
  5. 5.Changan Chen, Ziad Al-Halah, and Kristen Grauman. Semantic audio-visual navigation. In CVPR, pages 15516–15525, 2021.
  6. 6.Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. The sound of pixels. In ECCV, pages 570–586, 2018.
  7. 7.Ruohan Gao and Kristen Grauman. 2.5d visual sound. In CVPR, 2019.
  8. 8.Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu. Sep-stereo: Visually guided stereophonic audio generation by associating source separation. In ECCV, pages 52–69. Springer, 2020.
  9. 9.Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In ICCV, pages 5784–5794, 2021.
  10. 10.Yuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu, Zhongqin Wu, Shiguang Shan, and Xilin Chen. Unicon: Unified context network for robust active speaker detection. In ACM MM, pages 3964–3972, 2021.
  11. 11.Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong. Audio-visual segmentation. In ECCV, 2022.
  12. 12.Andrew Luo, Yilun Du, Michael J Tarr, Joshua B Tenenbaum, Antonio Torralba, and Chuang Gan. Learning neural acoustic fields. NeurIPS, 2022.
  13. 13.Kun Su, Mingfei Chen, and Eli Shlizerman. Inras: Implicit neural representation for audio scenes. In NeurIPS, 2022.
  14. 14.Yilun Du, M. Katherine Collins, B. Joshua Tenenbaum, and Vincent Sitzmann. Learning signal-agnostic manifolds of neural fields. In NeurIPS, 2021.
  15. 15.Changan Chen, Alexander Richard, Roman Shapovalov, Vamsi Krishna Ithapu, Natalia Neverova, Kristen Grauman, and Andrea Vedaldi. Novel-view acoustic synthesis. arXiv preprint arXiv:2301.08730, 2023.
  16. 16.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
  17. 17.Vincent Sitzmann, Michael Zollhofer, and Gordon Wetzstein. Scene representation networks: Continuous 3d-structure-aware neural scene representations. NeurIPS, 32, 2019.
  18. 18.Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger. Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision. In CVPR, pages 3504–3515, 2020.
  19. 19.Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and Yaron Lipman. Multiview neural surface reconstruction by disentangling geometry and appearance. NeurIPS, 33:2492–2502, 2020.
  20. 20.Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Justin Kerr, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, David McAllister, and Angjoo Kanazawa. Nerfstudio: A modular framework for neural radiance field development. arXiv preprint arXiv:2302.04264, 2023.
  21. 21.Carl Schissler, Christian Loftin, and Dinesh Manocha. Acoustic classification and optimization for multi-modal rendering of real-world scenes. IEEE TVCG, 24(3):1246–1259, 2017.
  22. 22.Dingzeyu Li, Timothy R. Langlois, and Changxi Zheng. Scene-aware audio for 360° videos. ACM TOG, 37(4), 2018.
  23. 23.Zhenyu Tang, Rohith Aralikatti, Anton Jeran Ratnarajah, and Dinesh Manocha. Gwa: A large high-quality acoustic dataset for audio processing. In SIGGRAPH, pages 1–9, 2022.
  24. 24.Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duckworth. NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections. In CVPR, 2021.
  25. 25.Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthesis of dynamic scenes. In CVPR, 2021.
  26. 26.Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. Dynamic view synthesis from dynamic monocular video. In ICCV, 2021.
  27. 27.Tianye Li, Mira Slavcheva, Michael Zollhofer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard A. Newcombe, and Zhaoyang Lv. Neural 3d video synthesis from multi-view video. In CVPR, pages 5511–5521, 2022.
  28. 28.Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu. Sep-stereo: Visually guided stereophonic audio generation by associating source separation. In ECCV, 2020.
  29. 29.Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin. Visually informed binaural audio generation without binaural audios. In CVPR, 2021.
  30. 30.Pedro Morgado, Nuno Vasconcelos, Timothy R. Langlois, and Oliver Wang. Self-supervised generation of spatial audio for 360° video. In NeurIPS, pages 360–370, 2018.
  31. 31.Changan Chen, Ruohan Gao, Paul Calamia, and Kristen Grauman. Visual acoustic matching. In CVPR, 2022.
  32. 32.V.R. Algazi, R.O. Duda, D.M. Thompson, and C. Avendano. The cipic hrtf database. In Proceedings of the 2001 IEEE Workshop on the Applications of Signal Processing to Audio and Acoustics (Cat. No.01TH8575), pages 99–102, 2001. doi: 10.1109/ASPAA.2001.969552.
  33. 33.Changan Chen, Carl Schissler, Sanchit Garg, Philip Kobernik, Alexander Clegg, Paul Calamia, Dhruv Batra, Philip W Robinson, and Kristen Grauman. Soundspaces 2.0: A simulation platform for visual-acoustic learning. arXiv, 2022.
  34. 34.Zhenyu Tang, Nicholas J Bryan, Dingzeyu Li, Timothy R Langlois, and Dinesh Manocha. Scene-aware audio rendering via deep acoustic analysis. IEEE TVCG, 2020.
  35. 35.Anton Ratnarajah, Shi-Xiong Zhang, Meng Yu, Zhenyu Tang, Dinesh Manocha, and Dong Yu. Fast-rir: Fast neural diffuse room impulse response generator. In ICASSP, pages 571–575, 2022.
  36. 36.Anton Ratnarajah, Zhenyu Tang, and Dinesh Manocha. Ir-gan: Room impulse response generator for far-field speech recognition. In INTERSPEECH, pages 286–290, 2021.
  37. 37.Anton Ratnarajah, Zhenyu Tang, Rohith Aralikatti, and Dinesh Manocha. Mesh2ir: Neural acoustic impulse response generator for complex 3d scenes. In ACM MM, pages 924–933, 2022.
  38. 38.Alexander Richard, Dejan Markovic, Israel D Gebru, Steven Krenn, Gladstone Butler, Fernando de la Torre, and Yaser Sheikh. Neural synthesis of binaural speech from mono audio. In ICLR, 2021.
  39. 39.Alexander Richard, Peter Dodds, and Vamsi Krishna Ithapu. Deep impulse responses: Estimating and parameterizing filters with deep networks. In ICASSP, pages 3209–3213, 2022.
  40. 40.Sagnik Majumder, Changan Chen, Ziad Al-Halah, and Kristen Grauman. Few-shot audio-visual learning of environment acoustics. In NeurIPS, 2022.
  41. 41.James T Kajiya and Brian P Von Herzen. Ray tracing volume densities. SIGGRAPH, 18(3): 165–174, 1984.
  42. 42.Nelson Max. Optical models for direct volume rendering. IEEE TVCG, 1(2):99–108, 1995.
  43. 43.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  44. 44.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009.
  45. 45.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, pages 5998–6008, 2017.
  46. 46.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In CVPR, pages 4104–4113, 2016.
  47. 47.Adobe Inc. Adobe audition. Software, 2023. URL https://www.adobe.com/products/audition.html.
  48. 48.Andrew Szot, Alex Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra. Habitat 2.0: Training home assistants to rearrange their habitat. In NeurIPS, 2021.
  49. 49.Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A Platform for Embodied AI Research. In ICCV, 2019.
  50. 50.Xiph.Org Foundation. Xiph opus. https://opus-codec.org/, 2012.
  51. 51.International Organization for Standardization. Advanced audio coding (aac). ISO/IEC 13818-7:2006, 2006.
  52. 52.Manik Varma and Andrew Zisserman. Classifying images of materials: Achieving viewpoint and illumination independence. In ECCV, pages 255–271, 2002.
  53. 53.Zhoutong Zhang, Jiajun Wu, Qiujia Li, Zhengjia Huang, James Traer, Josh H McDermott, Joshua B Tenenbaum, and William T Freeman. Generative modeling of audible shapes for object perception. In ICCV, pages 1251–1260, 2017.
  54. 54.Changan Chen, Wei Sun, David Harwath, and Kristen Grauman. Learning audio-visual dereverberation. In ICASSP, pages 1–5, 2023.
  55. 55.Nikhil Singh, Jeff Mentch, Jerry Ng, Matthew Beveridge, and Iddo Drori. Image2reverb: Cross-modal reverb impulse response synthesis. In ICCV, pages 286–295, 2021.
  56. 56.Zirui Wang, Shangzhe Wu, Weidi Xie, Min Chen, and Victor Adrian Prisacariu. NeRF−−: Neural radiance fields without known camera parameters. arXiv preprint arXiv:2102.07064, 2021.
  57. 57.Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey. Barf: Bundle-adjusting neural radiance fields. In ICCV, 2021.
  58. 58.Thomas Muller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM TOG, 41(4):1–15, 2022.
  59. 59.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. NeurIPS, 32, 2019.
  60. 60.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Yoshua Bengio and Yann LeCun, editors, ICLR, 2015.
  61. 61.Julius Orion Smith. Mathematics of the discrete Fourier transform (DFT): with audio applications. 2008.

Citation

MLA
Liang, S., et al. “AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 37472–90, https://proceedings.neurips.cc/paper_files/paper/2023/file/760dff0f9c0e9ed4d7e22918c73351d4-Paper-Conference.pdf.
APA
Liang, S., Huang, C., Tian, Y., Kumar, A., & Xu, C. (2023). AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis. Advances in Neural Information Processing Systems, 36, 37472–37490. https://proceedings.neurips.cc/paper_files/paper/2023/file/760dff0f9c0e9ed4d7e22918c73351d4-Paper-Conference.pdf
Chicago
Liang, S., C. Huang, Y. Tian, A. Kumar, and C. Xu. 2023. “AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis”. Advances in Neural Information Processing Systems 36: 37472–90. https://proceedings.neurips.cc/paper_files/paper/2023/file/760dff0f9c0e9ed4d7e22918c73351d4-Paper-Conference.pdf.
Harvard
Liang, S. et al. (2023) “AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 37472–37490. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/760dff0f9c0e9ed4d7e22918c73351d4-Paper-Conference.pdf.
Vancouver
1. Liang S, Huang C, Tian Y, Kumar A, Xu C (2023) AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 37472–37490

BibTeX

@inproceedings{liang2023nerf,
  title = {AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis},
  author = {Liang, Susan and Huang, Chao and Tian, Yapeng and Kumar, Anurag and Xu, Chenliang},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {37472-37490},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/760dff0f9c0e9ed4d7e22918c73351d4-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors