DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis

Yinghao XuMenglei ChaiZifan ShiSida PengIvan SkorokhodovAliaksandr SiarohinCeyuan YangYujun ShenHsin-Ying LeeBolei Zhou

article2023CVPR80 citations

Proposes a 3D-aware generative model that uses simple 3D bounding boxes as layout priors to spatially disentangle multi-object scenes into composable radiance fields from single-view 2D images, enabling high-fidelity scene synthesis and interactive object-level editing.

Listen

Generating realistic and editable three-dimensional scenes from standard two-dimensional images remains a major challenge in artificial intelligence and computer vision. While existing generative models excel at synthesizing single, isolated objects such as individual faces or cars, they struggle to generate complex, multi-object environments like indoor rooms or outdoor street views. Most current frameworks either treat the entire scene as a single fused volume, preventing users from editing specific objects, or fail to produce photorealistic results and proper spatial depth when dealing with intricate backgrounds and diverse object arrangements.

The article introduces DisCoScene, a 3D-aware generative framework designed to produce high-quality, complex scenes while enabling interactive, object-level editing directly from 2D image collections. The primary objective is to demonstrate that incorporating a lightweight layout prior—specifically simple 3D bounding boxes without semantic category labels—enables the model to disentangle objects from the background and each other without requiring ground-truth 3D camera sequences.

To achieve this, the approach breaks the entire scene down into spatially disentangled, object-centric radiance fields shared within a single generator, combined with a separate background generator. The framework integrates an efficient neural rendering pipeline that uses ray-box intersections to sample points only within valid bounding boxes, substantially reducing computational overhead. To ensure high visual quality during training, the authors implement a dual global-local discrimination strategy: a global discriminator critiques full scene coherence, while a local object discriminator evaluates cropped individual object patches. The model was evaluated across three diverse benchmarks representing synthetic primitives (CLEVR, 80,000 images), complex indoor environments (3D-FRONT, 80,000 images), and real-world autonomous driving footage (WAYMO, 70,000 images).

The evaluation yields several critical findings. First, DisCoScene achieved state-of-the-art visual fidelity among 3D-aware methods across all datasets, achieving Frechet Inception Distance (FID) scores of 3.5 on CLEVR, 13.8 on 3D-FRONT, and 16.0 on WAYMO, matching or approaching top-tier 2D generative models. Second, it demonstrated drastic performance gains on complex real-world data over prior multi-object baselines; for instance, on WAYMO, DisCoScene improved FID from GIRAFFE's 175.7 down to 16.0. Third, the local object discriminator proved vital for preventing the background model from collapsing or entangling with foreground objects, improving object-level FID on WAYMO from 95.1 to 16.3. Finally, the framework successfully supports diverse user-controlled operations at inference time, including object rotation, translation, insertion, removal, restyling, camera movement, and real-image inversion editing.

These findings imply that complex 3D environments can be synthesized efficiently without requiring expensive, labor-intensive scene graphs or dense 3D scans. By relying on simple bounding boxes, organizations can build interactive simulation, virtual staging, and content creation pipelines from unannotated 2D imagery at lower operational and computational costs. Furthermore, the ability to seamlessly compose and edit individual objects reduces the risk of visual artifacts and spatial inconsistencies in downstream applications like autonomous vehicle simulation and virtual reality.

For future implementation, stakeholders should consider combining this framework with automated monocular 3D object detection to infer layout priors directly from raw real-world images in an end-to-end manner. When scaling to massive urban environments, integrating large-scale neural radiance field techniques will be necessary to expand model capacity. The primary limitation of the current method is its reliance on pre-existing bounding box priors and limited capacity when modeling massive global streetscapes. Nonetheless, the reported empirical results provide high confidence in DisCoScene's capability to deliver controllable, photorealistic scene synthesis across diverse indoor and outdoor domains.

arXiv: 2212.11984
Cover for DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis

Abstract

Existing 3D-aware image synthesis approaches mainly focus on generating a single canonical object and show limited capacity in composing a complex scene containing a variety of objects. This work presents DisCoScene: a 3D-aware generative model for high-quality and controllable scene synthesis. The key ingredient of our method is a very abstract object-level representation (i.e., 3D bounding boxes without semantic annotation) as the scene layout prior, which is simple to obtain, general to describe various scene contents, and yet informative to disentangle objects and background. Moreover, it serves as an intuitive user control for scene editing. Based on such a prior, the proposed model spatially disentangles the whole scene into object-centric generative radiance fields by learning on only 2D images with the global-local discrimination. Our model obtains the generation fidelity and editing flexibility of individual objects while being able to efficiently compose objects and the background into a complete scene. We demonstrate state-of-the-art performance on many scene datasets, including the challenging Waymo outdoor dataset. Project page can be found here.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Abstract Layout Prior
  • 3.2. Spatially Disentangled Radiance Fields
  • 3.3. Efficient Rendering Pipeline
  • 3.4. Local & Global Discrimination
  • 3.5. Training and Inference
  • 4. Experiments
  • 4.1. Settings
  • 4.2. Main Results
  • 4.3. Controllable Scene Generation
  • 4.4. Ablation Study
  • 5. Discussion and Conclusion
  • References

Knowls

  1. Knowl 1 — DisCoScene Architecture for 3D-Aware Scene Synthesis

    model/method

    DisCoScene is a 3D-aware generative model designed to synthesize complex multi-object scenes with spatial disentanglement and controllable camera and object attributes from 2D image collections.

    The framework represents a scene using an abstract 3D layout prior comprising unannotated 3D bounding boxes. Scene generation consists of three main stages:

    1. Spatially Disentangled Field Generation: A shared canonical object generator models individual object radiance fields conditioned on per-object latent codes and layout parameters (position and scale), while a global background generator models the background radiance field conditioned on camera viewing direction and a background latent code.

    2. Bounded Neural Volume Rendering: Rays cast from a virtual camera are evaluated using ray-box intersection against the 3D bounding boxes. Points sampled along the ray segments within object boxes are sorted by depth and volume-rendered into a low-resolution foreground feature map, which is alpha-composited with the global background feature map and upsampled to full image resolution via a 2D convolutional upsampler network.

    3. Global-Local Adversarial Training: The generative radiance fields are trained using a scene discriminator evaluating the whole rendered image alongside a local object discriminator evaluating cropped 2D object patches extracted from the rendered and real images based on projected bounding boxes.

  2. Knowl 2 — Abstract 3D Layout Prior Parameterization

    definition

    The scene layout prior is defined as a set of NN unannotated 3D bounding boxes B={Bi∣i∈[1,N]}\mathcal{B} = \{B_i \mid i \in [1, N]\}. Each bounding box BiB_i is parameterized by 9 spatial parameters:

    Bi=[ai,ti,si]B_i = [a_i, t_i, s_i]

    where:

    • ai=[ax,ay,az]∈R3a_i = [a_x, a_y, a_z] \in \mathbb{R}^3 denotes three Euler rotation angles, which form a 3D rotation matrix Ri∈SO(3)R_i \in \mathrm{SO}(3);
    • ti=[tx,ty,tz]∈R3t_i = [t_x, t_y, t_z] \in \mathbb{R}^3 denotes the 3D translation vector;
    • si=[sx,sy,sz]∈R3s_i = [s_x, s_y, s_z] \in \mathbb{R}^3 denotes the 3D scale dimensions.

    Any bounding box BiB_i is mapped from a canonical unit cube C=[−0.5,0.5]3C = [-0.5, 0.5]^3 centered at the coordinate origin via the affine transformation bi(C)b_i(C):

    Bi=bi(C)=Ri⋅diag(si)⋅C+tiB_i = b_i(C) = R_i \cdot \mathrm{diag}(s_i) \cdot C + t_i

    where diag(si)\mathrm{diag}(s_i) is a diagonal matrix containing the scale factors sis_i.

  3. Knowl 3 — Spatially Conditioned Generative Radiance Fields

    model/method

    DisCoScene divides radiance field generation into a shared object generator and a background generator.

    Object Generator: To model diverse objects with a single network while maintaining spatial disentanglement, the object generator GobjG_{\text{obj}} operates in canonical object space and is conditioned on both a latent code zi∼N(0,I)z_i \sim \mathcal{N}(0, I) and the object's spatial configuration (translation tit_i and scale sis_i):

    (ci,σi)=Gobj(bi−1(γ(x)),concat(zi,γ(ti),γ(si)))(c_i, \sigma_i) = G_{\text{obj}}\left(b_i^{-1}(\gamma(x)), \mathrm{concat}(z_i, \gamma(t_i), \gamma(s_i))\right)

    where x∈R3x \in \mathbb{R}^3 is a 3D spatial coordinate in world space, bi−1b_i^{-1} is the inverse transformation mapping xx to the canonical box coordinate system of box BiB_i, γ(⋅)\gamma(\cdot) is a sinusoidal positional encoding function mapping coordinates to Fourier features, ci∈RCc_i \in \mathbb{R}^C is the emitted feature/color, and σi∈R\sigma_i \in \mathbb{R} is the volume density. The object generator is not conditioned on viewing direction vv, leaving view-dependent effects to the convolutional upsampler.

    Background Generator: The background is modeled in global coordinates by a generator GbgG_{\text{bg}} conditioned on 3D coordinate xx, camera viewing direction v∈S2v \in \mathbb{S}^2, and background latent code zbg∼N(0,I)z_{\text{bg}} \sim \mathcal{N}(0, I):

    (cbg,σbg)=Gbg(x,v,zbg)(c_{\text{bg}}, \sigma_{\text{bg}}) = G_{\text{bg}}(x, v, z_{\text{bg}})

  4. Knowl 4 — Ray-Box Bounded Volume Rendering and Scene Composition

    model/method

    Rendering an S×SS \times S feature map from the spatially disentangled radiance fields uses axis-aligned bounding box (Ray-AABB) intersections in canonical coordinates to avoid sampling empty 3D space:

    1. Ray Intersection and Sampling: For each ray rjr_j (j∈[1,S2]j \in [1, S^2]) intersecting box BlB_l, near and far intersection depths (dj,l,n,dj,l,f)(d_{j,l,n}, d_{j,l,f}) are computed. An intersection matrix M∈{0,1}N×S2M \in \{0, 1\}^{N \times S^2} records active intersections. Exactly NdN_d points are sampled equidistantly in [dj,l,n,dj,l,f][d_{j,l,n}, d_{j,l,f}].

    2. Foreground Depth Sorting: For a ray rjr_j intersecting nj≥1n_j \ge 1 boxes, the collected njNdn_j N_d points are sorted along the ray by depth, producing an ordered point sequence Xjs={xj,sk∣sk∈[1,njNd],dj,sk≤dj,sk+1}\mathcal{X}_j^s = \{x_{j, s_k} \mid s_k \in [1, n_j N_d], d_{j, s_k} \le d_{j, s_{k+1}}\}.

    3. Volume Rendering: The foreground feature f(rj)f(r_j) and foreground feature map FF are computed via numerical quadrature:

    f(rj)=∑k=1njNdTj,kαj,kc(xj,sk),Tj,k=exp⁡(−∑o=1k−1σ(xj,so)δj,so),αj,k=1−exp⁡(−σ(xj,sk)δj,sk)f(r_j) = \sum_{k=1}^{n_j N_d} T_{j,k} \alpha_{j,k} c(x_{j, s_k}), \quad T_{j,k} = \exp\left(-\sum_{o=1}^{k-1} \sigma(x_{j, s_o}) \delta_{j, s_o}\right), \quad \alpha_{j,k} = 1 - \exp\left(-\sigma(x_{j, s_k}) \delta_{j, s_k}\right)

    Fj={f(rj),if ∃m∈M:,j such that m=10,otherwiseF_j = \begin{cases} f(r_j), & \text{if } \exists m \in M_{:, j} \text{ such that } m = 1 \\ 0, & \text{otherwise} \end{cases}

    where δj,sk=dj,sk+1−dj,sk\delta_{j, s_k} = d_{j, s_{k+1}} - d_{j, s_k}.

    1. Composition with Background: The global background map NN is rendered by sampling background points (using fixed depths for indoor scenes or inverse depth sampling for unbounded outdoor scenes). The composite image feature map InI_n is obtained by alpha-blending:

    In=F+∏k=1njNd(1−αj,k)⊙NI_n = F + \prod_{k=1}^{n_j N_d} (1 - \alpha_{j,k}) \odot N

    A 2D convolutional upsampler then processes InI_n to generate the final RGB output.

  5. Knowl 5 — Global-Local Adversarial Training Formulation

    model/method

    DisCoScene trains its generators G={Gobj,Gbg}G = \{G_{\text{obj}}, G_{\text{bg}}\} and two discriminators (a global scene discriminator DsD_s and a local object discriminator DobjD_{\text{obj}}) using non-saturating GAN losses with R1R_1-style gradient penalties.

    Let If=G(B,Z,ξ)I_f = G(\mathcal{B}, \mathcal{Z}, \xi) denote an image synthesized from layout B\mathcal{B}, latent codes Z\mathcal{Z}, and camera pose ξ∼pξ\xi \sim p_\xi. Let IrI_r denote a real image. Object patches PIfP_{I_f} and PIrP_{I_r} are extracted by projecting 3D bounding boxes to 2D boxes Bi2DB_i^{2D} and cropping the corresponding regions: PI={Pi∣Pi=crop(I,Bi2D)}P_I = \{P_i \mid P_i = \mathrm{crop}(I, B_i^{2D})\}.

    The generator loss LG\mathcal{L}_G and discriminator loss LD\mathcal{L}_D are:

    min⁡GLG=E[f(−Ds(If))]+λ1E[f(−Dobj(PIf))]\min_G \mathcal{L}_G = \mathbb{E}[f(-D_s(I_f))] + \lambda_1 \mathbb{E}[f(-D_{\text{obj}}(P_{I_f}))]

    min⁡DLD=E[f(−Ds(Ir))]+E[f(Ds(If))]+λ1(E[f(−Dobj(PIr))]+E[f(Dobj(PIf))])+λ2∥∇IrDs(Ir)∥22+λ3∥∇PIrDobj(PIr)∥22\min_D \mathcal{L}_D = \mathbb{E}[f(-D_s(I_r))] + \mathbb{E}[f(D_s(I_f))] + \lambda_1 \left(\mathbb{E}[f(-D_{\text{obj}}(P_{I_r}))] + \mathbb{E}[f(D_{\text{obj}}(P_{I_f}))]\right) + \lambda_2 \|\nabla_{I_r} D_s(I_r)\|_2^2 + \lambda_3 \|\nabla_{P_{I_r}} D_{\text{obj}}(P_{I_r})\|_2^2

    where f(t)=log⁡(1+exp⁡(t))f(t) = \log(1 + \exp(t)) is the softplus function, λ1\lambda_1 controls the balance between object and scene discrimination (set to 1.01.0), and λ2,λ3\lambda_2, \lambda_3 weight the gradient regularization terms (both set to 1.01.0).

  6. Knowl 6 — Quantitative Evaluation of Scene Generation Quality

    data/table

    DisCoScene was evaluated against 2D GAN (StyleGAN2) and 3D-aware GAN baselines (EpiGRAF, VolumeGAN, EG3D, GIRAFFE, GSN) on three multi-object benchmarks: CLEVR (80K diagnostic images), 3D-FRONT (80K indoor bedroom images), and WAYMO (70K outdoor street front-view images) at 256×256256 \times 256 resolution. Evaluation metrics are Frechet Inception Distance (FID) and Kernel Inception Distance (KID ×103\times 10^3), evaluated over 50K generated samples against all real images. Training cost (TR.) is in V100 GPU days; inference latency (INF.) is in milliseconds per image on a single V100 GPU over 1K samples.

    Model CLEVR 3D-FRONT WAYMO
    FID ↓\downarrow KID ↓\downarrow TR. ↓\downarrow INF. ↓\downarrow FID ↓\downarrow KID ↓\downarrow FID ↓\downarrow KID ↓\downarrow
    StyleGAN2 4.5 3.0 13.3 44 12.5 4.3 15.1 8.3
    EpiGRAF 10.4 8.3 16.0 114 107.2 102.3 27.0 26.1
    VolumeGAN 7.5 5.1 15.2 90 52.7 38.7 29.9 18.2
    EG3D 4.1 12.7 25.8 55 19.7 13.5 26.0 45.4
    GIRAFFE 78.5 61.5 5.2 62 56.5 46.8 175.7 212.1
    GSN – – – – 130.7 87.5 – –
    DisCoScene 3.5 2.1 18.1 95 13.8 7.4 16.0 8.4

    DisCoScene achieves the best FID and KID scores among all 3D-aware methods across all three datasets, approaching the fidelity of the 2D baseline StyleGAN2 while enabling explicit 3D camera control and object-level manipulation.

  7. Knowl 7 — Ablation Analysis of Local Object Discriminator

    data/table

    The contribution of the local object discriminator DobjD_{\text{obj}} was evaluated across CLEVR, 3D-FRONT, and WAYMO by comparing scene-level FID and object-level fidelity FIDobj\text{FID}_{\text{obj}}. FIDobj\text{FID}_{\text{obj}} is computed by cropping 2D object patches using projected bounding boxes from generated and real images.

    Metric Setting CLEVR 3D-FRONT WAYMO
    FID ↓\downarrow w/o DobjD_{\text{obj}} 5.0 18.6 19.5
    w/ DobjD_{\text{obj}} 3.5 13.8 16.0
    FIDobj↓\text{FID}_{\text{obj}} \downarrow w/o DobjD_{\text{obj}} 19.1 33.7 95.1
    w/ DobjD_{\text{obj}} 5.6 19.5 16.3

    Without DobjD_{\text{obj}}, the background generator tends to overfit and entangle foreground objects into a single global field (evidenced by the sharp degradation of FIDobj\text{FID}_{\text{obj}} to 95.1 on WAYMO). DobjD_{\text{obj}} enforces clear spatial separation between individual foreground objects and the background, substantially improving object and full-scene synthesis quality.

  8. Knowl 8 — Unsupervised Semantic Alignment via Spatial Conditioning

    empirical result

    Conditioning the canonical object generator GobjG_{\text{obj}} on Fourier-encoded spatial parameters (translation tit_i and scale sis_i) allows the model to learn semantically consistent object assignments without category annotations.

    On the 3D-FRONT indoor bedroom benchmark:

    • Adding spatial conditioning improves overall image FID from 15.215.2 to 13.813.8 and object-level FIDobj\text{FID}_{\text{obj}} from 23.223.2 to 19.519.5.
    • Qualitatively, spatial conditioning causes the model to consistently generate large center-of-room objects as beds and perimeter objects as nightstands or wardrobes, matching natural room layout distributions. Without spatial conditioning, the model generates semantically unnatural configurations (e.g., placing small tables or nightstands in the center of the room).
  9. Knowl 9 — Interactive 3D Scene Manipulation Operations

    model/method

    Once trained, DisCoScene supports multiple interactive 3D scene editing capabilities at inference time via direct manipulation of the layout prior B\mathcal{B} and latent codes:

    1. Object Rearrangement: Users can modify the 3D translation tit_i and rotation aia_i of individual bounding boxes to reposition or rotate objects while preserving their appearance, identity, and handling mutual occlusions correctly.

    2. Object Removal and Cloning: Removing a bounding box eliminates the corresponding object, with the background seamlessly inpainted without having been trained on clean background-only images. Cloning is achieved by replicating a box and its latent code at a new 3D location.

    3. Object Restyling: By performing style-mixing across different layers of the modulated fully-connected layers in GobjG_{\text{obj}} using independently sampled latent codes ziz_i, object shape and appearance can be controlled independently.

    4. Supersampling Anti-Aliasing (SSAA): To prevent edge aliasing artifacts during dynamic object manipulation caused by coarse 64×6464 \times 64 ray marching, the foreground feature rendering is evaluated at a temporary 128×128128 \times 128 resolution and downsampled to 64×6464 \times 64 before entering the 2D neural upsampler, increasing inference time only marginally (from 95 ms to 105 ms per image).

  10. Knowl 10 — Limitations of DisCoScene

    limitation

    DisCoScene has two primary limitations:

    1. Dependence on 3D Bounding Box Layout Priors: The method relies on predefined 3D bounding box layouts B\mathcal{B}. For unannotated in-the-wild datasets, it requires an external monocular 3D object detector (e.g., FCOS3D) to extract pseudo-layouts during preprocessing rather than estimating layouts end-to-end.

    2. Global Space Street Scene Scaling: Representing expansive, unconstrained global environments (such as complex driving scenes in WAYMO) with a single global background generator is limited by model capacity, requiring future integration with multi-scale or tiled radiance field representations.

Coverage note — Deliberately omitted minor qualitative visualization figures, network layer-dimension breakdowns, and standard training hyperparameter lists that follow conventional StyleGAN2 and BlenderProc setups, retaining all core architecture, mathematical formulations, quantitative data, and experimental ablations.

References

  1. 1.Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Int. Conf. Comput. Vis., pages 5855–5864, 2021.
  2. 2.Daniel Bear, Chaofei Fan, Damian Mrowca, Yunzhu Li, Seth Alter, Aran Nayebi, Jeremy Schwartz, Li F Fei-Fei, Jiajun Wu, Josh Tenenbaum, et al. Learning physical graph representations from visual scenes. Adv. Neural Inform. Process. Syst., 2020.
  3. 3.Mikołaj Bińkowski, Danica J Sutherland, Michael Arbel, and Arthur Gretton. Demystifying mmd gans. arXiv preprint arXiv:1801.01401, 2018.
  4. 4.Eric R Chan, Connor Z Lin, Matthew A Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J Guibas, Jonathan Tremblay, Sameh Khamis, et al. Efficient geometry-aware 3d generative adversarial networks. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  5. 5.Eric R Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In IEEE Conf. Comput. Vis. Pattern Recog., 2021.
  6. 6.Steve Cunningham and Michael J Bailey. Lessons from scene graphs: using scene graphs to teach hierarchical modeling. Computers & Graphics, 2001.
  7. 7.Yu Deng, Jiaolong Yang, Jianfeng Xiang, and Xin Tong. Gram: Generative radiance manifolds for 3d-aware image generation. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  8. 8.Terrance DeVries, Miguel Angel Bautista, Nitish Srivastava, Graham W. Taylor, and Joshua M. Susskind. Unconstrained scene generation with locally conditioned radiance fields. In Int. Conf. Comput. Vis., 2021.
  9. 9.Dave Epstein, Taesung Park, Richard Zhang, Eli Shechtman, and Alexei A. Efros. Blobgan: Spatially disentangled scene representations. Eur. Conf. Comput. Vis., 2022.
  10. 10.Huan Fu, Bowen Cai, Lin Gao, Ling-Xiao Zhang, Jiaming Wang, Cao Li, Qixun Zeng, Chengyue Sun, Rongfei Jia, Binqiang Zhao, et al. 3d-front: 3d furnished rooms with layouts and semantics. In IEEE Conf. Comput. Vis. Pattern Recog., 2021.
  11. 11.Huan Fu, Rongfei Jia, Lin Gao, Mingming Gong, Binqiang Zhao, Steve Maybank, and Dacheng Tao. 3d-future: 3d furniture shape with texture. Int. J. Comput. Vis., 2021.
  12. 12.Raghudeep Gadde, Qianli Feng, and Aleix M Martinez. Detail me more: Improving gan’s photo-realism of complex scenes. In IEEE Conf. Comput. Vis. Pattern Recog., 2021.
  13. 13.Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, and Sanja Fidler. Get3d: A generative model of high quality 3d textured shapes learned from images. arXiv preprint arXiv:2209.11163, 2022.
  14. 14.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Adv. Neural Inform. Process. Syst., 2014.
  15. 15.Jiatao Gu, Lingjie Liu, Peng Wang, and Christian Theobalt. Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis. arXiv preprint arXiv:2110.08985, 2021.
  16. 16.Michelle Guo, Alireza Fathi, Jiajun Wu, and Thomas Funkhouser. Object-centric neural scene rendering. arXiv preprint arXiv:2012.08503, 2020.
  17. 17.Kamal Gupta, Justin Lazarow, Alessandro Achille, Larry S Davis, Vijay Mahadevan, and Abhinav Shrivastava. Layouttransformer: Layout generation and completion with self-attention. In Int. Conf. Comput. Vis., 2021.
  18. 18.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Adv. Neural Inform. Process. Syst., 2017.
  19. 19.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In IEEE Conf. Comput. Vis. Pattern Recog., 2017.
  20. 20.Justin Johnson, Agrim Gupta, and Li Fei-Fei. Image generation from scene graphs. In IEEE Conf. Comput. Vis. Pattern Recog., 2018.
  21. 21.Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In IEEE Conf. Comput. Vis. Pattern Recog., 2017.
  22. 22.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. In Int. Conf. Learn. Represent., 2018.
  23. 23.Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-free generative adversarial networks. In Adv. Neural Inform. Process. Syst., 2021.
  24. 24.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In IEEE Conf. Comput. Vis. Pattern Recog., 2019.
  25. 25.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of StyleGAN. In IEEE Conf. Comput. Vis. Pattern Recog., 2020.
  26. 26.Jianan Li, Jimei Yang, Aaron Hertzmann, Jianming Zhang, and Tingfa Xu. Layoutgan: Generating graphic layouts with wireframe discriminators. arXiv preprint arXiv:1901.06767, 2019.
  27. 27.Alexander Majercik, Cyril Crassin, Peter Shirley, and Morgan McGuire. A ray-box intersection algorithm and efficient dynamic voxel rendering. Journal of Computer Graphics Techniques Vol, 2018.
  28. 28.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In IEEE Conf. Comput. Vis. Pattern Recog., 2019.
  29. 29.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In Eur. Conf. Comput. Vis., 2020.
  30. 30.Thu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt, and Yong-Liang Yang. Hologan: Unsupervised learning of 3d representations from natural images. In Int. Conf. Comput. Vis., 2019.
  31. 31.Thu H Nguyen-Phuoc, Christian Richardt, Long Mai, Yongliang Yang, and Niloy Mitra. Blockgan: Learning 3d object-aware scene representations from unlabelled images. Adv. Neural Inform. Process. Syst., 2020.
  32. 32.Michael Niemeyer and Andreas Geiger. Giraffe: Representing scenes as compositional generative neural feature fields. In IEEE Conf. Comput. Vis. Pattern Recog., 2021.
  33. 33.Roy Or-El, Xuan Luo, Mengyi Shan, Eli Shechtman, Jeong Joon Park, and Ira Kemelmacher-Shlizerman. Stylesdf: High-resolution 3d-consistent image and geometry generation. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  34. 34.Julian Ost, Fahim Mannan, Nils Thuerey, Julian Knodt, and Felix Heide. Neural scene graphs for dynamic scenes. In IEEE Conf. Comput. Vis. Pattern Recog., pages 2856–2865, 2021.
  35. 35.Xingang Pan, Xudong Xu, Chen Change Loy, Christian Theobalt, and Bo Dai. A shading-guided generative implicit model for shape-accurate 3d-aware image synthesis. In Adv. Neural Inform. Process. Syst., 2021.
  36. 36.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In IEEE Conf. Comput. Vis. Pattern Recog., 2019.
  37. 37.Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. Semantic image synthesis with spatially-adaptive normalization. In IEEE Conf. Comput. Vis. Pattern Recog., 2019.
  38. 38.Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In IEEE Conf. Comput. Vis. Pattern Recog., pages 9054–9063, 2021.
  39. 39.Daniel Roich, Ron Mokady, Amit H Bermano, and Daniel Cohen-Or. Pivotal tuning for latent-based editing of real images. ACM Trans. Graph., 2022.
  40. 40.Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. Graf: Generative radiance fields for 3d-aware image synthesis. In Adv. Neural Inform. Process. Syst., 2020.
  41. 41.Katja Schwarz, Axel Sauer, Michael Niemeyer, Yiyi Liao, and Andreas Geiger. Voxgraf: Fast 3d-aware image synthesis with sparse voxel grids. Adv. Neural Inform. Process. Syst., 2022.
  42. 42.Yujun Shen, Ceyuan Yang, Xiaoou Tang, and Bolei Zhou. Interfacegan: Interpreting the disentangled face representation learned by gans. IEEE Trans. Pattern Anal. Mach. Intell., 2020.
  43. 43.A. Sherrod. Game Graphic Programming. Course Technology PTR game development series. Course Technology/Charles River Media/Cengage Learning, 2008.
  44. 44.Zifan Shi, Yujun Shen, Jiapeng Zhu, Dit-Yan Yeung, and Qifeng Chen. 3d-aware indoor scene synthesis with depth priors. In Eur. Conf. Comput. Vis., 2022.
  45. 45.Zifan Shi, Yinghao Xu, Yujun Shen, Deli Zhao, Qifeng Chen, and Dit-Yan Yeung. Improving 3d-aware image synthesis with a geometry-aware discriminator. Adv. Neural Inform. Process. Syst., 2022.
  46. 46.Ivan Skorokhodov, Sergey Tulyakov, Yiqun Wang, and Peter Wonka. Epigraf: Rethinking training of 3d gans. In Adv. Neural Inform. Process. Syst., 2022.
  47. 47.Henry Sowizral. Scene graphs in the new millennium. IEEE Computer Graphics and Applications, 2000.
  48. 48.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In IEEE Conf. Comput. Vis. Pattern Recog., 2020.
  49. 49.Zhentao Tan, Dongdong Chen, Qi Chu, Menglei Chai, Jing Liao, Mingming He, Lu Yuan, Gang Hua, and Nenghai Yu. Efficient semantic image synthesis via class-adaptive normalization. IEEE Trans. Pattern Anal. Mach. Intell., 2021.
  50. 50.Zhentao Tan, Qi Chu, Menglei Chai, Dongdong Chen, Jing Liao, Qiankun Liu, Bin Liu, Gang Hua, and Nenghai Yu. Semantic probability distribution modeling for diverse semantic image synthesis. IEEE Trans. Pattern Anal. Mach. Intell., 2022.
  51. 51.Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul P Srinivasan, Jonathan T Barron, and Henrik Kretzschmar. Block-nerf: Scalable large scene neural view synthesis. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  52. 52.Zhuowen Tu, Xiangrong Chen, Alan L Yuille, and Song-Chun Zhu. Image parsing: Unifying segmentation, detection, and recognition. IJCV, 2005.
  53. 53.Jianyuan Wang, Ceyuan Yang, Yinghao Xu, Yujun Shen, Hongdong Li, and Bolei Zhou. Improving gan equilibrium by raising spatial awareness. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  54. 54.Tai Wang, Xinge Zhu, Jiangmiao Pang, and Dahua Lin. Fcos3d: Fully convolutional one-stage monocular 3d object detection. In Int. Conf. Comput. Vis. Worksh., 2021.
  55. 55.Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. High-resolution image synthesis and semantic manipulation with conditional gans. In IEEE Conf. Comput. Vis. Pattern Recog., 2018.
  56. 56.Qianyi Wu, Xian Liu, Yuedong Chen, Kejie Li, Chuanxia Zheng, Jianfei Cai, and Jianmin Zheng. Object-compositional neural implicit surfaces. In Eur. Conf. Comput. Vis., 2022.
  57. 57.Yuanbo Xiangli, Linning Xu, Xingang Pan, Nanxuan Zhao, Anyi Rao, Christian Theobalt, Bo Dai, and Dahua Lin. Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering. In Eur. Conf. Comput. Vis., 2022.
  58. 58.Linning Xu, Yuanbo Xiangli, Anyi Rao, Nanxuan Zhao, Bo Dai, Ziwei Liu, and Dahua Lin. Blockplanner: City block generation with vectorized graph representation. In Int. Conf. Comput. Vis., 2021.
  59. 59.Xudong Xu, Xingang Pan, Dahua Lin, and Bo Dai. Generative occupancy fields for 3d surface-aware image synthesis. In Adv. Neural Inform. Process. Syst., 2021.
  60. 60.Yinghao Xu, Sida Peng, Ceyuan Yang, Yujun Shen, and Bolei Zhou. 3d-aware image synthesis via learning structural and textural representations. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  61. 61.Yinghao Xu, Yujun Shen, Jiapeng Zhu, Ceyuan Yang, and Bolei Zhou. Generative hierarchical features from synthesizing images. In IEEE Conf. Comput. Vis. Pattern Recog., 2021.
  62. 62.Yang Xue, Yuheng Li, Krishna Kumar Singh, and Yong Jae Lee. Giraffe hd: A high-resolution 3d-aware generative model. In CVPR, 2022.
  63. 63.Bangbang Yang, Yinda Zhang, Yinghao Xu, Yijin Li, Han Zhou, Hujun Bao, Guofeng Zhang, and Zhaopeng Cui. Learning object-compositional neural radiance field for editable scene rendering. In Int. Conf. Comput. Vis., pages 13779–13788, 2021.
  64. 64.Ceyuan Yang, Yujun Shen, and Bolei Zhou. Semantic hierarchy emerges in deep generative representations for scene synthesis. IJCV, 2021.
  65. 65.Chen Zhang, Yinghao Xu, and Yujun Shen. Decorating your own bedroom: Locally controlling image generation with generative adversarial networks. IEEE Conf. Comput. Vis. Pattern Recog. Worksh., 2021.
  66. 66.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020.
  67. 67.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Int. Conf. Comput. Vis., 2017.
  68. 68.Jun-Yan Zhu, Zhoutong Zhang, Chengkai Zhang, Jiajun Wu, Antonio Torralba, Joshua B. Tenenbaum, and William T. Freeman. Visual object networks: Image generation with disentangled 3D representations. In Adv. Neural Inform. Process. Syst., 2018.

Citation

MLA
Xu, Y., et al. “DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis”. arXiv, 2022, http://arxiv.org/abs/2212.11984v1.
APA
Xu, Y., Chai, M., Shi, Z., Peng, S., Skorokhodov, I., Siarohin, A., Yang, C., Shen, Y., Lee, H.-Y., Zhou, B., & Tulyakov, S. (2022). DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis. arXiv. http://arxiv.org/abs/2212.11984v1
Chicago
Xu, Y., M. Chai, Z. Shi, et al. 2022. “DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis”. arXiv. http://arxiv.org/abs/2212.11984v1.
Harvard
Xu, Y. et al. (2022) “DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.11984v1.
Vancouver
1. Xu Y, Chai M, Shi Z, et al (2022) DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis. arXiv

BibTeX

@article{xu2022discoscene,
  title = {DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis},
  author = {Xu, Yinghao and Chai, Menglei and Shi, Zifan and Peng, Sida and Skorokhodov, Ivan and Siarohin, Aliaksandr and Yang, Ceyuan and Shen, Yujun and Lee, Hsin-Ying and Zhou, Bolei and Tulyakov, Sergey},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.11984v1},
  eprint = {2212.11984}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE