DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis
Yinghao XuMenglei ChaiZifan ShiSida PengIvan SkorokhodovAliaksandr SiarohinCeyuan YangYujun ShenHsin-Ying LeeBolei Zhou
Proposes a 3D-aware generative model that uses simple 3D bounding boxes as layout priors to spatially disentangle multi-object scenes into composable radiance fields from single-view 2D images, enabling high-fidelity scene synthesis and interactive object-level editing.
Generating realistic and editable three-dimensional scenes from standard two-dimensional images remains a major challenge in artificial intelligence and computer vision. While existing generative models excel at synthesizing single, isolated objects such as individual faces or cars, they struggle to generate complex, multi-object environments like indoor rooms or outdoor street views. Most current frameworks either treat the entire scene as a single fused volume, preventing users from editing specific objects, or fail to produce photorealistic results and proper spatial depth when dealing with intricate backgrounds and diverse object arrangements.
The article introduces DisCoScene, a 3D-aware generative framework designed to produce high-quality, complex scenes while enabling interactive, object-level editing directly from 2D image collections. The primary objective is to demonstrate that incorporating a lightweight layout prior—specifically simple 3D bounding boxes without semantic category labels—enables the model to disentangle objects from the background and each other without requiring ground-truth 3D camera sequences.
To achieve this, the approach breaks the entire scene down into spatially disentangled, object-centric radiance fields shared within a single generator, combined with a separate background generator. The framework integrates an efficient neural rendering pipeline that uses ray-box intersections to sample points only within valid bounding boxes, substantially reducing computational overhead. To ensure high visual quality during training, the authors implement a dual global-local discrimination strategy: a global discriminator critiques full scene coherence, while a local object discriminator evaluates cropped individual object patches. The model was evaluated across three diverse benchmarks representing synthetic primitives (CLEVR, 80,000 images), complex indoor environments (3D-FRONT, 80,000 images), and real-world autonomous driving footage (WAYMO, 70,000 images).
The evaluation yields several critical findings. First, DisCoScene achieved state-of-the-art visual fidelity among 3D-aware methods across all datasets, achieving Frechet Inception Distance (FID) scores of 3.5 on CLEVR, 13.8 on 3D-FRONT, and 16.0 on WAYMO, matching or approaching top-tier 2D generative models. Second, it demonstrated drastic performance gains on complex real-world data over prior multi-object baselines; for instance, on WAYMO, DisCoScene improved FID from GIRAFFE's 175.7 down to 16.0. Third, the local object discriminator proved vital for preventing the background model from collapsing or entangling with foreground objects, improving object-level FID on WAYMO from 95.1 to 16.3. Finally, the framework successfully supports diverse user-controlled operations at inference time, including object rotation, translation, insertion, removal, restyling, camera movement, and real-image inversion editing.
These findings imply that complex 3D environments can be synthesized efficiently without requiring expensive, labor-intensive scene graphs or dense 3D scans. By relying on simple bounding boxes, organizations can build interactive simulation, virtual staging, and content creation pipelines from unannotated 2D imagery at lower operational and computational costs. Furthermore, the ability to seamlessly compose and edit individual objects reduces the risk of visual artifacts and spatial inconsistencies in downstream applications like autonomous vehicle simulation and virtual reality.
For future implementation, stakeholders should consider combining this framework with automated monocular 3D object detection to infer layout priors directly from raw real-world images in an end-to-end manner. When scaling to massive urban environments, integrating large-scale neural radiance field techniques will be necessary to expand model capacity. The primary limitation of the current method is its reliance on pre-existing bounding box priors and limited capacity when modeling massive global streetscapes. Nonetheless, the reported empirical results provide high confidence in DisCoScene's capability to deliver controllable, photorealistic scene synthesis across diverse indoor and outdoor domains.
- Paper: GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields, Michael Niemeyer et al. (2021). GIRAFFE pioneered compositional generative neural feature fields for multi-object scenes, serving as the foundational compositional baseline that DisCoScene directly improves upon.
- Paper: Efficient Geometry-aware 3D Generative Adversarial Networks, Eric R. Chan et al. (2022). EG3D established high-efficiency 3D-aware generative adversarial networks using tri-plane representations and dual discrimination, inspiring core volumetric rendering and discrimination strategies in DisCoScene.
- Paper: Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations, V. Sitzmann et al. (2019). Scene Representation Networks laid the groundwork for continuous coordinate-based implicit 3D scene representations and differentiable ray marching.
- Paper: D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video, Tianhao Wu et al. (2022). D2NeRF introduced methods to cleanly decouple dynamic foreground objects from static background neural radiance fields.
- Paper: CityDreamer: Compositional Generative Model of Unbounded 3D Cities, Haozhe Xie et al. (2024). CityDreamer extends compositional 3D generative radiance fields from bounded object-centric scenes to unbounded, large-scale urban layouts.
- Paper: XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies, Xuanchi Ren et al. (2024). XCube scales 3D generative modeling of large-scale outdoor driving scenes using hierarchical sparse voxel diffusion beyond bounding-box implicit fields.
- Paper: DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes, Xiaoyu Zhou et al. (2024). DrivingGaussian builds upon foreground-background scene decomposition for dynamic autonomous driving scenes by replacing implicit radiance fields with efficient composite Gaussian splatting.
- Paper: DiffRF: Rendering-Guided 3D Radiance Field Diffusion, Norman Müller et al. (2023). DiffRF explores 3D radiance field generation using 3D diffusion models with rendering-guided losses as an alternative to adversarial radiance field frameworks.
- Paper: MIME: Human-Aware 3D Scene Generation, Hongwei Yi et al. (2023). MIME advances 3D scene layout synthesis by utilizing human motion and interaction dynamics as priors for populating indoor environments.
