CityDreamer: Compositional Generative Model of Unbounded 3D Cities
Haozhe XieZhaoxi ChenFangzhou HongZiwei Liu
Proposes a compositional 3D generative framework that disentangles building instances from background elements using tailored neural fields to generate unbounded, editable urban scenes with realistic geometry and diverse appearances.
Generating expansive, realistic 3D virtual cities is critical for industries such as video game development, film production, urban planning, and simulations. However, scaling automated 3D city generation has proven difficult because urban environments contain vast architectural diversity. Unlike natural landscapes where objects such as trees share consistent textures, buildings feature varied facade patterns, sharp edges, and unique structures. Existing generative models often treat all buildings as a single object class, resulting in severe visual distortion and unnatural architectural geometry. The article introduces CityDreamer, a generative framework designed to produce unbounded, highly realistic 3D cities with accurate geometry and consistent multi-view perspectives.
To address structural complexity, the approach separates the generation process into distinct specialized components. An unbounded layout generator first creates expandable city maps and building heights. The system then processes the environment through two separate neural rendering paths: one tailored for background terrain such as roads, water, and greenery, and another specifically dedicated to individual building instances and facade textures. Finally, a compositor merges these elements into a unified 3D scene. To train and validate the system, the authors built extensive reference datasets containing geographical layouts from 80 global cities and 24,000 real-world aerial trajectory images of New York City, benchmarking the framework against four leading 3D generation models on visual realism, depth accuracy, multi-view consistency, and human perceptual ratings.
The findings show that CityDreamer substantially outperforms existing state-of-the-art methods across all evaluated metrics. The system achieved a visual quality score of 97.38, reducing standard image distortion metrics by more than 20% to 65% compared to competing baselines. Camera trajectory consistency improved significantly, reducing tracking errors to 0.06 compared to 0.186 or higher in previous models. Ablation analyses demonstrated that separating building instances from background terrain was essential to performance, as omitting dedicated building instance processing more than doubled visual distortion. Furthermore, in controlled user evaluations, human raters consistently scored CityDreamer higher in image quality, 3D realism, and viewpoint consistency than all baseline approaches.
These results indicate that modular, instance-level separation is a viable solution to the structural distortion problems that have limited large-scale automated 3D environment synthesis. For creative and technical industries, this methodology offers a practical pathway to dramatically reduce the time, labor, and production costs required to construct massive digital cities. Additionally, because the architecture isolates individual buildings, it enables precise, localized editing within generated urban environments without requiring full scene regeneration.
Decision-makers adopting automated 3D content pipelines should consider instance-aware architectures to achieve commercial-grade visual fidelity. However, leaders should account for certain operational constraints: the system currently generates building heights through vertical extrusion, meaning it cannot model concave geometries such as tunnels, bridges, or caves. Furthermore, rendering individual buildings sequentially increases computation time during deployment. Future development should focus on optimizing computational efficiency and expanding structural modeling capabilities to handle non-vertical overhangs and subterranean features. Nonetheless, the experimental evidence supports high confidence in the model's reliability for generating expansive, high-fidelity urban environments.
- Paper: GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields, Michael Niemeyer et al. (2021). Learn how compositional neural feature fields decouple individual objects from backgrounds to understand the core compositional neural rendering framework adapted by CityDreamer for unbounded urban scenes.
- Paper: High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs, Ting-Chun Wang et al. (2018). Read this paper to explore high-resolution conditional image synthesis and instance-level boundary manipulation from semantic layouts, which directly informs CityDreamer's layout-to-scene generation pipeline.
- Paper: Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields, Jonathan T. Barron et al. (2022). Understand the foundational techniques for contracting coordinates and handling scale in unbounded 360-degree neural radiance fields prior to examining CityDreamer's unbounded urban scene synthesis.
- Paper: Efficient Geometry-aware 3D Generative Adversarial Networks, Eric R. Chan et al. (2022). Examine how hybrid neural radiance field architectures and tri-plane representations enable efficient 3D-aware generative modeling before exploring large-scale compositional urban generation.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Master the foundational volume rendering and coordinate-based neural radiance field representation that underpins modern 3D neural scene synthesis methods.
- Paper: Image-to-Image Translation with Conditional Adversarial Networks, Phillip Isola et al. (2017). Review the baseline principles of conditional adversarial networks for translating semantic layout maps into photorealistic imagery that informs semantic urban scene generation.
- Paper: Structured 3D Latents for Scalable and Versatile 3D Generation, Jianfeng Xiang et al. (2025). Explore how structured 3D latents and rectified flow transformers advance beyond domain-specific compositional rendering to achieve versatile, unified 3D asset generation across multiple output formats.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). See how multi-role conditioning, background canvases, and real-world dynamics scale compositional generation and localized editing across video and image modalities.
- Paper: DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision, Lu Ling et al. (2024). Discover how large-scale real-world multi-view benchmarks facilitate comprehensive evaluation and training for complex outdoor novel view synthesis and 3D reconstruction.
