DiffRF: Rendering-Guided 3D Radiance Field Diffusion

Norman MüllerYawar SiddiquiLorenzo PorziSamuel Rota BulòPeter KontschiederMatthias Nießner

article2023CVPR246 citations

Introduces a 3D diffusion framework that operates directly on explicit volumetric radiance fields and integrates a rendering loss to synthesize high-quality, view-consistent 3D assets while naturally enabling conditional tasks like shape completion.

Listen

Generating high-quality 3D digital assets is critical for industries such as gaming, virtual and augmented reality, mapping, and simulation. While neural radiance fields represent 3D scenes with photo-realistic detail from 2D images, most models are limited to fitting individual scenes rather than generating new, diverse 3D objects. Meanwhile, alternative generative adversarial networks often struggle with training instability and view-inconsistent geometric artifacts. The article introduces and evaluates DiffRF, a novel generative diffusion model that directly creates 3D volumetric radiance fields to produce geometrically accurate, view-consistent 3D shapes with detailed appearance.

To evaluate this capability, the authors trained a 3D denoising model operating directly on explicit voxel grid representations of 3D radiance fields. Because ground-truth volumetric data often contains reconstruction errors and floating artifacts, the authors paired the standard diffusion noise prediction objective with a rendering-guided 2D loss. This rendering loss forces the model to prioritize final image quality and suppress fitting errors. The approach was tested across standard benchmarks, including the PhotoShape Chairs dataset (15,576 samples) and the Amazon Berkeley Objects Tables dataset (1,676 samples), assessing both unconditional 3D asset generation and conditional tasks such as masked 3D object completion and single-image 3D reconstruction.

The findings show that DiffRF outperforms leading generative adversarial networks in both visual quality and geometric fidelity. On the PhotoShape Chairs benchmark, DiffRF achieved a better image synthesis score (Fréchet Inception Distance of 15.95 compared to EG3D's 16.54) and improved geometric quality (Minimum Matching Distance of 4.42 compared to 5.62). The experiments confirmed that adding the 2D rendering loss is essential; removing it degraded the image quality score from 15.95 to 18.27 on chairs and from 27.06 to 35.89 on tables. Furthermore, unlike prior networks that require task-specific retraining, DiffRF successfully performs conditional inference at test time. For masked 3D completion, it reliably preserved known object geometry while generating coherent missing parts, achieving an average peak signal-to-noise ratio of 27.53 compared to 24.82 for the baseline across various masking levels.

These results demonstrate that direct volumetric diffusion provides a stable, unified framework for both generating and editing 3D content without sacrificing 3D geometric consistency. By overcoming the shape distortions and training instability common to adversarial approaches, this method offers a viable path toward automating 3D content creation workflows. The findings show that volumetric representations can be directly synthesized if rendering guidance is incorporated during training to filter out volumetric fitting artifacts.

For technical teams seeking to adopt volumetric diffusion models, the article indicates that training pipelines must incorporate multi-view rendering supervision rather than relying purely on 3D grid regression. However, practitioners should note several constraints before deployment: the approach currently exhibits slower generation sampling times compared to feed-forward networks and requires substantial training-time memory, which limited the experimental grid resolution to 32 cubed. Future development should focus on integrating faster diffusion sampling techniques, memory-efficient sparse grids, and scalable factorized radiance field representations before deploying the framework to high-resolution, production-scale pipelines.

arXiv: 2212.01206
Cover for DiffRF: Rendering-Guided 3D Radiance Field Diffusion

Abstract

We introduce DiffRF, a novel approach for 3D radiance field synthesis based on denoising diffusion probabilistic models. While existing diffusion-based methods operate on images, latent codes, or point cloud data, we are the first to directly generate volumetric radiance fields. To this end, we propose a 3D denoising model which directly operates on an explicit voxel grid representation. However, as radiance fields generated from a set of posed images can be ambiguous and contain artifacts, obtaining ground truth radiance field samples is non-trivial. We address this challenge by pairing the denoising formulation with a rendering loss, enabling our model to learn a deviated prior that favours good image quality instead of trying to replicate fitting errors like floating artifacts. In contrast to 2D-diffusion models, our model learns multi-view consistent priors, enabling free-view synthesis and accurate shape generation. Compared to 3D GANs, our diffusion-based approach naturally enables conditional generation such as masked completion or single-view 3D synthesis at inference time.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Method
  • 3.1. Radiance Fields
  • 3.2. Generating Radiance Fields
  • 3.3. Training Objective
  • 4. Experiments
  • 4.1. Unconditional Radiance Field Synthesis
  • 4.2. Conditional Generation
  • 4.3. Limitations
  • 5. Conclusions
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Rendering-Guided 3D Radiance Field Diffusion Objective

    model/method

    DiffRF trains a 3D denoising model directly over explicit voxel grid radiance fields. Because ground-truth radiance fields obtained from multi-view image fitting frequently contain reconstruction artifacts (such as floaters and fitting noise), training solely with a 3D voxel-space Euclidean loss forces the generative model to replicate these flaws. DiffRF pairs the standard 3D diffusion loss LRF\mathcal{L}_{\text{RF}} with an explicit 2D volumetric rendering loss LRGB\mathcal{L}_{\text{RGB}} evaluated against ground-truth multi-view training images IvI_v.

    To avoid the computational cost of simulating multi-step reverse sampling during training, DiffRF approximates the fully denoised clean radiance field using an analytical single-step estimate f~0t(ϵ,θ)\tilde{f}_0^t(\epsilon, \theta):

    f~0t(ϵ,θ):=f0+1−αˉtαˉt(ϵ−ϵθ(ft,t))\tilde{f}_0^t(\epsilon, \theta) := f_0 + \frac{\sqrt{1-\bar{\alpha}_t}}{\sqrt{\bar{\alpha}_t}}\left(\epsilon - \epsilon_\theta(f_t, t)\right)

    where f0∈RH×W×D×Cf_0 \in \mathbb{R}^{H \times W \times D \times C} is the pre-activated radiance field training sample, ft=αˉtf0+1−αˉtϵf_t = \sqrt{\bar{\alpha}_t}f_0 + \sqrt{1-\bar{\alpha}_t}\epsilon is its corrupted state at timestep t∈{1,…,T}t \in \{1, \dots, T\} with noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I), and ϵθ\epsilon_\theta is a 3D-UNet denoiser.

    The time-conditioned rendering loss is defined as:

    LRGBt(f0∣θ):=wtEv∼ψ,ϵ∼N(0,I)[∥Iv−R(v,f~0t(ϵ,θ))∥2]\mathcal{L}_{\text{RGB}}^t(f_0|\theta) := w_t \mathbb{E}_{v \sim \psi, \epsilon \sim \mathcal{N}(0, I)}\left[\left\|I_v - R(v, \tilde{f}_0^t(\epsilon, \theta))\right\|^2\right]

    where R(v,f)R(v, f) denotes the 2D image obtained by volumetrically rendering radiance field ff from camera viewpoint vv, ψ\psi is the viewpoint distribution, and wt:=αˉt2w_t := \bar{\alpha}_t^2 is a decaying weight schedule that down-weights rendering supervision at large noise levels tt.

    The combined training objective per sample f0f_0 is:

    L(θ)=Et∼U(1,T)[LRFt(f0∣θ)+λRGBLRGBt(f0∣θ)]\mathcal{L}(\theta) = \mathbb{E}_{t \sim \mathcal{U}(1, T)}\left[\mathcal{L}_{\text{RF}}^t(f_0|\theta) + \lambda_{\text{RGB}} \mathcal{L}_{\text{RGB}}^t(f_0|\theta)\right]

    where LRFt(f0∣θ)=Eϵ[∥ϵ−ϵθ(ft,t)∥2]\mathcal{L}_{\text{RF}}^t(f_0|\theta) = \mathbb{E}_{\epsilon}\left[\|\epsilon - \epsilon_\theta(f_t, t)\|^2\right] and λRGB\lambda_{\text{RGB}} balances the two objectives.

  2. Knowl 2 — Explicit Voxel Grid Radiance Field Representation and Volumetric Rendering

    model/method

    DiffRF models a 3D object as an explicit volumetric grid f∈RH×W×D×Cf \in \mathbb{R}^{H \times W \times D \times C} spanning a bounded 3D spatial domain X⊂R3\mathcal{X} \subset \mathbb{R}^3. The channels index a pre-activated density field σ:X→R\sigma: \mathcal{X} \to \mathbb{R} and a pre-activated RGB color field ξ:X→R3\xi: \mathcal{X} \to \mathbb{R}^3. Operating in pre-activation linear space ensures that additive Gaussian noise corruptions preserve valid representations, while nonlinear activation functions are applied during ray casting.

    Values at continuous 3D coordinates rs∈Xr_s \in \mathcal{X} along a ray rr are retrieved via trilinear interpolation of the explicit grid vertices. The rendered RGB color crc_r along camera ray r(s)r(s) parametrized by arc length s≥0s \ge 0 with unit direction is computed via numerical integration of the volume rendering equation:

    cr:=∫0∞τs(r)σ(rs)ξ(rs)dsc_r := \int_0^\infty \tau_s(r) \sigma(r_s) \xi(r_s) ds

    where τs(r)\tau_s(r) denotes the accumulated transmittance probability from the origin to distance ss along ray rr:

    τs(r):=exp⁡(−∫0sσ(s′)ds′)\tau_s(r) := \exp\left(-\int_0^s \sigma(s') ds'\right)

    Rendering an entire image R(v,f)R(v, f) from camera pose vv corresponds to integrating across all rays cast from the camera origin through the image plane pixels.

  3. Knowl 3 — 3D Radiance Field Denoising Diffusion Probabilistic Model

    model/method

    DiffRF formalizes 3D radiance field generation as a Denoising Diffusion Probabilistic Model (DDPM) over explicit flattened 4D tensors f∈Ff \in \mathcal{F}. The forward diffusion process corrupts a ground-truth radiance field f0∼q(f0)f_0 \sim q(f_0) over discrete steps t∈{1,…,T}t \in \{1, \dots, T\} with variance schedule β1,…,βT∈(0,1)\beta_1, \dots, \beta_T \in (0, 1):

    q(ft∣ft−1):=N(ft;αtft−1,βtI)q(f_t | f_{t-1}) := \mathcal{N}\left(f_t; \sqrt{\alpha_t} f_{t-1}, \beta_t I\right)

    where αt:=1−βt\alpha_t := 1 - \beta_t. The distribution of ftf_t conditioned directly on f0f_0 is:

    q(ft∣f0)=N(ft;αˉtf0,(1−αˉt)I)q(f_t | f_0) = \mathcal{N}\left(f_t; \sqrt{\bar{\alpha}_t} f_0, (1 - \bar{\alpha}_t) I\right)

    where αˉt:=∏i=1tαi\bar{\alpha}_t := \prod_{i=1}^t \alpha_i.

    The generative reverse denoising process parameterizes the Gaussian transitions pθ(ft−1∣ft):=N(ft−1;μθ(ft,t),Σt)p_\theta(f_{t-1} | f_t) := \mathcal{N}(f_{t-1}; \mu_\theta(f_t, t), \Sigma_t) using a 3D-UNet ϵθ(ft,t)\epsilon_\theta(f_t, t) that predicts the applied noise:

    μθ(ft,t):=1αt(ft−βt1−αˉtϵθ(ft,t))\mu_\theta(f_t, t) := \frac{1}{\sqrt{\alpha_t}}\left(f_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}} \epsilon_\theta(f_t, t)\right)

    with fixed covariance schedule Σt:=βt22αt(1−αˉt)I\Sigma_t := \frac{\beta_t^2}{2 \alpha_t (1 - \bar{\alpha}_t)} I.

    In unconditional sampling, generation begins from isotropic Gaussian noise fT∼N(0,I)f_T \sim \mathcal{N}(0, I) and proceeds iteratively down to f0f_0, yielding an explicit 3D radiance field.

  4. Knowl 4 — Masked 3D Radiance Field Completion

    algorithm

    Masked radiance field completion synthesizes missing regions within a 3D object from a partially observed input radiance field finf^{\text{in}} and a binary voxel mask m∈{0,1}H×W×Dm \in \{0, 1\}^{H \times W \times D} (where m=1m=1 indicates masked/unknown regions and m=0m=0 indicates known regions). The procedure requires no task-specific retraining and guides unconditional reverse diffusion by fusing the known observed field at each sampling timestep.

    Input: Observed radiance field finf^{\text{in}}, binary 3D mask mm, total timesteps TT, variance schedule α1,…,αT\alpha_1, \dots, \alpha_T, cumulative factors αˉ1,…,αˉT\bar{\alpha}_1, \dots, \bar{\alpha}_T, trained 3D denoiser ϵθ\epsilon_\theta
    Output: Completed 3D radiance field f0f_0
    Sample initial noise state fT∼N(0,I)f_T \sim \mathcal{N}(0, I)
    for t=Tt = T down to 11 do
        Predict noise ϵ^←ϵθ(ft,t)\hat{\epsilon} \leftarrow \epsilon_\theta(f_t, t)
        Compute clean field estimate f~0t←1αˉt(ft−1−αˉtϵ^)\tilde{f}_0^t \leftarrow \frac{1}{\sqrt{\bar{\alpha}_t}}\left(f_t - \sqrt{1 - \bar{\alpha}_t} \hat{\epsilon}\right)
        Fuse known and generated regions: f0t−1←αˉt(m⊙f~0t+(1−m)⊙fin)f_0^{t-1} \leftarrow \sqrt{\bar{\alpha}_t}\left(m \odot \tilde{f}_0^t + (1 - m) \odot f^{\text{in}}\right)
        if t>1t > 1 then
            Sample noise z∼N(0,I)z \sim \mathcal{N}(0, I)
            ft−1←f0t−1+1−αˉtzf_{t-1} \leftarrow f_0^{t-1} + \sqrt{1 - \bar{\alpha}_t} z
        else
            f0←m⊙f~01+(1−m)⊙finf_0 \leftarrow m \odot \tilde{f}_0^1 + (1 - m) \odot f^{\text{in}}
        end if
    end for
    return f0f_0
  5. Knowl 5 — Single-View Image-to-Volume Synthesis via Classifier Guidance

    model/method

    DiffRF performs single-view 3D reconstruction by steering the unconditional diffusion reverse sampling trajectory using classifier guidance on the rendering loss. Given a single posed RGB image ItargetI_{\text{target}} from viewpoint vv along with a foreground object mask MM, the sampling trajectory of the 3D-UNet denoiser is guided at each step by the gradient of the photometric and masked rendering error ∇ft∥M⊙(Itarget−R(v,f~0t))∥2\nabla_{f_t} \|M \odot (I_{\text{target}} - R(v, \tilde{f}_0^t))\|^2, where f~0t\tilde{f}_0^t is the single-step prediction of the clean radiance field and R(v,⋅)R(v, \cdot) is the differentiable volume renderer. This guides the sampling process to produce diverse 3D radiance fields whose volume renderings match the input observation from the specified camera pose.

  6. Knowl 6 — Quantitative Evaluation of Unconditional 3D Synthesis

    data/table

    Unconditional generation is evaluated on PhotoShape Chairs (15,576 shapes rendered from 200 views) and ABO Tables (1,676 shapes rendered from 91 views). Training grids have spatial resolution 32332^3. Evaluated metrics include Fréchet Inception Distance (FID ↓\downarrow), Kernel Inception Distance (KID ×103↓\times 10^3 \downarrow), Coverage Score based on Chamfer Distance (COV %↑\% \uparrow), and Minimum Matching Distance based on Chamfer Distance (MMD ×103↓\times 10^3 \downarrow), all computed at 128×128128 \times 128 image resolution.

    PhotoShape Chairs FID ↓\downarrow KID ×103↓\times 10^3 \downarrow COV %↑\% \uparrow MMD ×103↓\times 10^3 \downarrow
    π\pi-GAN 52.71 13.64 39.92 7.387
    EG3D 16.54 8.412 47.55 5.619
    DiffRF w/o 2D 18.27 9.263 59.20 4.543
    DiffRF (Full) 15.95 7.935 58.93 4.416
    ABO Tables FID ↓\downarrow KID ×103↓\times 10^3 \downarrow COV %↑\% \uparrow MMD ×103↓\times 10^3 \downarrow
    π\pi-GAN 41.67 13.81 44.23 10.92
    EG3D 31.18 11.67 48.15 9.327
    DiffRF w/o 2D 35.89 13.94 63.46 8.013
    DiffRF (Full) 27.06 10.03 61.54 7.610

    DiffRF outperforms 3D GAN baselines in both image synthesis quality (lower FID/KID) and geometric quality/diversity (higher COV, lower MMD). Ablating the 2D rendering loss ("DiffRF w/o 2D") causes FID to degrade by 2.322.32 on PhotoShape and by 8.838.83 on ABO Tables, demonstrating the importance of rendering-guided supervision in eliminating voxel fitting artifacts.

  7. Knowl 7 — Quantitative Evaluation of Masked Radiance Field Completion

    data/table

    Masked completion performance is benchmarked on PhotoShape Chairs across four levels of volumetric masking (20%20\%, 40%40\%, 60%60\%, and 80%80\%). Evaluation is performed over 200 randomly masked test objects using masked Peak Signal-to-Noise Ratio (mPSNR ↑\uparrow) to measure the photometric preservation of unmasked regions and FID ↓\downarrow (rendered from 10 viewpoints) to evaluate the visual plausibility of completed objects. DiffRF is compared against EG3D equipped with masked GAN inversion using Global Latent Optimization (GLO).

    Metric / Method 20% 40% 60% 80% Avg
    mPSNR ↑\uparrow
    EG3D 23.71 24.86 24.92 25.79 24.82
    DiffRF 24.85 26.66 28.23 30.38 27.53
    FID ↓\downarrow
    EG3D 25.91 29.41 33.06 34.31 30.67
    DiffRF 22.36 27.74 31.16 29.84 27.78

    DiffRF achieves a higher average mPSNR (27.5327.53 dB vs. 24.8224.82 dB) and lower average FID (27.7827.78 vs. 30.6730.67) than EG3D. Because EG3D relies on optimizing a single global latent code, it struggles to strictly preserve non-masked visual features, whereas DiffRF directly preserves unmasked voxels while generating coherent completions.

  8. Knowl 8 — Limitations of DiffRF

    limitation

    DiffRF has three primary limitations:

    1. Data Requirement: Unlike 2D-supervised 3D GANs that train directly on unposed or single-view 2D image collections, DiffRF requires a sufficiently dense set of posed multi-view images per training instance to pre-optimize initial explicit 3D radiance fields.
    2. Sampling Speed: Iterative diffusion denoising across T=1000T = 1000 steps is substantially slower at inference time than single forward-pass GAN generators.
    3. Memory and Resolution Constraints: Directly running 3D convolutions and 3D attention operators over dense voxel grids creates large memory overhead during training, constraining the explicit radiance field resolution to 32332^3 voxels.

Coverage note — None was omitted; all key contributions (formulation of 3D radiance field diffusion, rendering loss, masked completion algorithm, single-view guidance, quantitative benchmarks, ablations, and stated limitations) are represented.

References

  1. 1.Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas Guibas. Learning representations and generative models for 3d point clouds. In International conference on machine learning, pages 40–49. PMLR, 2018.
  2. 2.Titas Anciukevicius, Zexiang Xu, Matthew Fisher, Paul Henderson, Hakan Bilen, Niloy J. Mitra, and Paul Guerrero. RenderDiffusion: Image diffusion for 3D reconstruction, inpainting and generation. arXiv, 2022.
  3. 3.Miguel Angel Bautista, Pengsheng Guo, Samira Abnar, Walter Talbott, Alexander Toshev, Zhuoyuan Chen, Laurent Dinh, Shuangfei Zhai, Hanlin Goh, Daniel Ulbricht, et al. Gaudi: A neural architect for immersive 3d scene generation. arXiv preprint arXiv:2207.13751, 2022.
  4. 4.Mikołaj Bińkowski, Dougal J. Sutherland, Michael Arbel, and Arthur Gretton. Demystifying MMD GANs. In International Conference on Learning Representations, 2018.
  5. 5.Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018.
  6. 6.Ruojin Cai, Guandao Yang, Hadar Averbuch-Elor, Zekun Hao, Serge J. Belongie, Noah Snavely, and Bharath Hariharan. Learning gradient fields for shape generation. CoRR, abs/2008.06520, 2020.
  7. 7.Eric Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In arXiv, 2020.
  8. 8.Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3D generative adversarial networks. In CVPR, 2022.
  9. 9.Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. arXiv preprint arXiv:2203.09517, 2022.
  10. 10.Kevin Chen, Christopher B Choy, Manolis Savva, Angel X Chang, Thomas Funkhouser, and Silvio Savarese. Text2shape: Generating shapes from natural language by learning joint embeddings. In Asian conference on computer vision, pages 100–116. Springer, 2018.
  11. 11.Jasmine Collins, Shubham Goel, Kenan Deng, Achleshwar Luthra, Leon Xu, Erhan Gundogdu, Xi Zhang, Tomas F Yago Vicente, Thomas Dideriksen, Himanshu Arora, et al. Abo: Dataset and benchmarks for real-world 3d object understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21126–21136, 2022.
  12. 12.Blender Online Community. Blender - a 3D modelling and rendering package. Blender Foundation, Stichting Blender Foundation, Amsterdam, 2018.
  13. 13.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017.
  14. 14.Giannis Daras, Mauricio Delbracio, Hossein Talebi, Alexandros G Dimakis, and Peyman Milanfar. Soft diffusion: Score matching for general corruptions. arXiv preprint arXiv:2209.05442, 2022.
  15. 15.Terrance DeVries, Miguel Angel Bautista, Nitish Srivastava, Graham W. Taylor, and Joshua M. Susskind. Unconstrained scene generation with locally conditioned radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (ICCV), 2021.
  16. 16.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.
  17. 17.Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5501–5510, 2022.
  18. 18.Xiao Fu, Shangzhan Zhang, Tianrun Chen, Yichong Lu, Lanyun Zhu, Xiaowei Zhou, Andreas Geiger, and Yiyi Liao. Panoptic nerf: 3d-to-2d label transfer for panoptic urban scene segmentation. In International Conference on 3D Vision (3DV), 2022.
  19. 19.Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, and Sanja Fidler. Get3d: A generative model of high quality 3d textured shapes learned from images. arXiv preprint arXiv:2209.11163, 2022.
  20. 20.Jiatao Gu, Lingjie Liu, Peng Wang, and Christian Theobalt. Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis. arXiv preprint arXiv:2110.08985, 2021.
  21. 21.Philipp Henzler, Niloy J Mitra, and Tobias Ritschel. Escaping plato’s cave: 3d shape from adversarial rendering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9984–9993, 2019.
  22. 22.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  23. 23.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. arXiv preprint arxiv:2006.11239, 2020.
  24. 24.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. arXiv preprint arXiv:2204.03458, 2022.
  25. 25.Emiel Hoogeboom and Tim Salimans. Blurring diffusion models. arXiv preprint arXiv:2209.05557, 2022.
  26. 26.Chin-Wei Huang, Jae Hyun Lim, and Aaron C Courville. A variational perspective on diffusion-based generative models and score matching. Advances in Neural Information Processing Systems, 34:22863–22876, 2021.
  27. 27.Animesh Karnewar, Tobias Ritschel, Oliver Wang, and Niloy Mitra. Relu fields: The little non-linearity that could. In ACM SIGGRAPH 2022 Conference Proceedings, SIGGRAPH ’22, New York, NY, USA, 2022. Association for Computing Machinery.
  28. 28.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. In Proc. NeurIPS, 2022.
  29. 29.Zhifeng Kong and Wei Ping. On fast sampling of diffusion probabilistic models. arXiv preprint arXiv:2106.00132, 2021.
  30. 30.Abhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi, Caroline Pantofaru, Leonidas J Guibas, Andrea Tagliasacchi, Frank Dellaert, and Thomas Funkhouser. Panoptic neural fields: A semantic object-aware neural scene representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  31. 31.Gang Li, Heliang Zheng, Chaoyue Wang, Chang Li, Changwen Zheng, and Dacheng Tao. 3ddesigner: Towards photorealistic 3d object generation and editing with text-guided diffusion models. arXiv preprint arXiv:2211.14108, 2022.
  32. 32.Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11461–11471, June 2022.
  33. 33.Shitong Luo and Wei Hu. Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2837–2845, 2021.
  34. 34.Shitong Luo, Chence Shi, Minkai Xu, and Jian Tang. Predicting molecular conformation via dynamic graph score matching. Advances in Neural Information Processing Systems, 34:19784–19795, 2021.
  35. 35.Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duckworth. NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  36. 36.Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. SDEdit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022.
  37. 37.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European conference on computer vision, pages 405–421. Springer, 2020.
  38. 38.Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957, 2018.
  39. 39.Norman Müller, Andrea Simonelli, Lorenzo Porzi, Samuel Rota Bulò, Matthias Nießner, and Peter Kontschieder. Autorf: Learning 3d object radiance fields from single view observations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  40. 40.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. CoRR, abs/2112.10741, 2021.
  41. 41.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pages 8162–8171. PMLR, 2021.
  42. 42.Michael Niemeyer and Andreas Geiger. Giraffe: Representing scenes as compositional generative neural feature fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11453–11464, 2021.
  43. 43.Anton Obukhov, Maximilian Seitzer, Po-Wei Wu, Semen Zhydenko, Jonathan Kyl, and Elvis Yu-Jing Lin. High-fidelity performance metrics for generative models in pytorch, 2020. Version: 0.3.0, DOI: 10.5281/zenodo.4957738.
  44. 44.Keunhong Park, Konstantinos Rematas, Ali Farhadi, and Steven M. Seitz. Photoshape: Photorealistic materials for large-scale shape collections. ACM Trans. Graph., 37(6), Nov. 2018.
  45. 45.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022.
  46. 46.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  47. 47.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  48. 48.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2021.
  49. 49.Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. Graf: Generative radiance fields for 3d-aware image synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  50. 50.Katja Schwarz, Axel Sauer, Michael Niemeyer, Yiyi Liao, and Andreas Geiger. Voxgraf: Fast 3d-aware image synthesis with sparse voxel grids. In Advances in Neural Information Processing Systems.
  51. 51.Yawar Siddiqui, Justus Thies, Fangchang Ma, Qi Shan, Matthias Nießner, and Angela Dai. Texturify: Generating textures on 3d shape surfaces. arXiv preprint arXiv:2204.02411, 2022.
  52. 52.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015.
  53. 53.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.
  54. 54.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021.
  55. 55.Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon. Maximum likelihood training of score-based diffusion models. Advances in Neural Information Processing Systems, 34:1415–1428, 2021.
  56. 56.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. CoRR, abs/1907.05600, 2019.
  57. 57.Yang Song and Stefano Ermon. Improved techniques for training score-based generative models. Advances in neural information processing systems, 33:12438–12448, 2020.
  58. 58.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  59. 59.Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In CVPR, 2022.
  60. 60.Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. Resolution-robust large mask inpainting with fourier convolutions. arXiv preprint arXiv:2109.07161, 2021.
  61. 61.Towaki Takikawa, Joey Litalien, Kangxue Yin, Karsten Kreis, Charles Loop, Derek Nowrouzezahrai, Alec Jacobson, Morgan McGuire, and Sanja Fidler. Neural geometric level of detail: Real-time rendering with implicit 3D shapes. 2021.
  62. 62.Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul Srinivasan, Jonathan T. Barron, and Henrik Kretzschmar. Block-NeRF: Scalable large scene neural view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  63. 63.Brian L Trippe, Jason Yim, Doug Tischer, Tamara Broderick, David Baker, Regina Barzilay, and Tommi Jaakkola. Diffusion probabilistic modeling of protein backbones in 3d for the motif-scaffolding problem. arXiv preprint arXiv:2206.04119, 2022.
  64. 64.Haithem Turki, Deva Ramanan, and Mahadev Satyanarayanan. Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  65. 65.Peng-Shuai Wang, Yang Liu, Yu-Xiao Guo, Chun-Yu Sun, and Xin Tong. O-CNN: Octree-based Convolutional Neural Networks for 3D Shape Analysis. ACM Transactions on Graphics (SIGGRAPH), 36(4), 2017.
  66. 66.Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P. Srinivasan, Howard Zhou, Jonathan T. Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser. Ibrnet: Learning multi-view image-based rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  67. 67.Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, et al. Rodin: A generative model for sculpting 3d digital avatars using diffusion. arXiv preprint arXiv:2212.06135, 2022.
  68. 68.Jiajun Wu, Chengkai Zhang, Tianfan Xue, Bill Freeman, and Josh Tenenbaum. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. Advances in neural information processing systems, 29, 2016.
  69. 69.Minkai Xu, Lantao Yu, Yang Song, Chence Shi, Stefano Ermon, and Jian Tang. Geodiff: A geometric diffusion model for molecular conformation generation. arXiv preprint arXiv:2203.02923, 2022.
  70. 70.Guandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu, Serge Belongie, and Bharath Hariharan. Pointflow: 3d point cloud generation with continuous normalizing flows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4541–4550, 2019.
  71. 71.Ruihan Yang, Prakhar Srivastava, and Stephan Mandt. Diffusion probabilistic modeling for video generation. arXiv preprint arXiv:2203.09481, 2022.
  72. 72.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4578–4587, 2021.
  73. 73.Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. Generative image inpainting with contextual attention. arXiv preprint arXiv:1801.07892, 2018.
  74. 74.Xiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, and Karsten Kreis. Lion: Latent point diffusion models for 3d shape generation. arXiv preprint arXiv:2210.06978, 2022.
  75. 75.Linqi Zhou, Yilun Du, and Jiajun Wu. 3d shape generation and completion through point-voxel diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5826–5835, 2021.
  76. 76.Peng Zhou, Lingxi Xie, Bingbing Ni, and Qi Tian. Cips-3d: A 3d-aware generator of gans based on conditionally-independent pixel synthesis. arXiv preprint arXiv:2110.09788, 2021.

Citation

MLA
Müller, N., et al. “DiffRF: Rendering-Guided 3D Radiance Field Diffusion”. arXiv, 2022, http://arxiv.org/abs/2212.01206v2.
APA
Müller, N., Siddiqui, Y., Porzi, L., Bulò, S. R., Kontschieder, P., & Nießner, M. (2022). DiffRF: Rendering-Guided 3D Radiance Field Diffusion. arXiv. http://arxiv.org/abs/2212.01206v2
Chicago
Müller, N., Y. Siddiqui, L. Porzi, S. R. Bulò, P. Kontschieder, and M. Nießner. 2022. “DiffRF: Rendering-Guided 3D Radiance Field Diffusion”. arXiv. http://arxiv.org/abs/2212.01206v2.
Harvard
Müller, N. et al. (2022) “DiffRF: Rendering-Guided 3D Radiance Field Diffusion”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.01206v2.
Vancouver
1. Müller N, Siddiqui Y, Porzi L, Bulò SR, Kontschieder P, Nießner M (2022) DiffRF: Rendering-Guided 3D Radiance Field Diffusion. arXiv

BibTeX

@article{muller2022diffrf,
  title = {DiffRF: Rendering-Guided 3D Radiance Field Diffusion},
  author = {Müller, Norman and Siddiqui, Yawar and Porzi, Lorenzo and Bulò, Samuel Rota and Kontschieder, Peter and Nießner, Matthias},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.01206v2},
  eprint = {2212.01206}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE