CityDreamer: Compositional Generative Model of Unbounded 3D Cities

Haozhe XieZhaoxi ChenFangzhou HongZiwei Liu

article2024CVPR89 citations

Proposes a compositional 3D generative framework that disentangles building instances from background elements using tailored neural fields to generate unbounded, editable urban scenes with realistic geometry and diverse appearances.

Listen

Generating expansive, realistic 3D virtual cities is critical for industries such as video game development, film production, urban planning, and simulations. However, scaling automated 3D city generation has proven difficult because urban environments contain vast architectural diversity. Unlike natural landscapes where objects such as trees share consistent textures, buildings feature varied facade patterns, sharp edges, and unique structures. Existing generative models often treat all buildings as a single object class, resulting in severe visual distortion and unnatural architectural geometry. The article introduces CityDreamer, a generative framework designed to produce unbounded, highly realistic 3D cities with accurate geometry and consistent multi-view perspectives.

To address structural complexity, the approach separates the generation process into distinct specialized components. An unbounded layout generator first creates expandable city maps and building heights. The system then processes the environment through two separate neural rendering paths: one tailored for background terrain such as roads, water, and greenery, and another specifically dedicated to individual building instances and facade textures. Finally, a compositor merges these elements into a unified 3D scene. To train and validate the system, the authors built extensive reference datasets containing geographical layouts from 80 global cities and 24,000 real-world aerial trajectory images of New York City, benchmarking the framework against four leading 3D generation models on visual realism, depth accuracy, multi-view consistency, and human perceptual ratings.

The findings show that CityDreamer substantially outperforms existing state-of-the-art methods across all evaluated metrics. The system achieved a visual quality score of 97.38, reducing standard image distortion metrics by more than 20% to 65% compared to competing baselines. Camera trajectory consistency improved significantly, reducing tracking errors to 0.06 compared to 0.186 or higher in previous models. Ablation analyses demonstrated that separating building instances from background terrain was essential to performance, as omitting dedicated building instance processing more than doubled visual distortion. Furthermore, in controlled user evaluations, human raters consistently scored CityDreamer higher in image quality, 3D realism, and viewpoint consistency than all baseline approaches.

These results indicate that modular, instance-level separation is a viable solution to the structural distortion problems that have limited large-scale automated 3D environment synthesis. For creative and technical industries, this methodology offers a practical pathway to dramatically reduce the time, labor, and production costs required to construct massive digital cities. Additionally, because the architecture isolates individual buildings, it enables precise, localized editing within generated urban environments without requiring full scene regeneration.

Decision-makers adopting automated 3D content pipelines should consider instance-aware architectures to achieve commercial-grade visual fidelity. However, leaders should account for certain operational constraints: the system currently generates building heights through vertical extrusion, meaning it cannot model concave geometries such as tunnels, bridges, or caves. Furthermore, rendering individual buildings sequentially increases computation time during deployment. Future development should focus on optimizing computational efficiency and expanding structural modeling capabilities to handle non-vertical overhangs and subterranean features. Nonetheless, the experimental evidence supports high confidence in the model's reliability for generating expansive, high-fidelity urban environments.

Cover for CityDreamer: Compositional Generative Model of Unbounded 3D Cities

Abstract

3D city generation is a desirable yet challenging task, since humans are more sensitive to structural distortions in urban environments. Additionally, generating 3D cities is more complex than 3D natural scenes since buildings, as objects of the same class, exhibit a wider range of appearances compared to the relatively consistent appearance of objects like trees in natural scenes. To address these challenges, we propose CityDreamer, a compositional generative model designed specifically for unbounded 3D cities. Our key insight is that 3D city generation should be a composition of different types of neural fields: 1) various building instances, and 2) background stuff, such as roads and green lands. Specifically, we adopt the bird’s eye view scene representation and employ a volumetric render for both instance-oriented and stuff-oriented neural fields. The generative hash grid and periodic positional embedding are tailored as scene parameterization to suit the distinct characteristics of building instances and background stuff. Furthermore, we contribute a suite of CityGen Datasets, which comprises a vast amount of real-world city imagery to enhance the realism of the generated 3D cities both in their layouts and appearances. CityDreamer achieves state-of-the-art performance not only in generating realistic 3D cities but also in localized editing within the generated cities.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Approach
  • 3.1. Unbounded City Layout Generator
  • 3.2. City Background Generator
  • 3.3. Building Instance Generator
  • 3.4. Compositor
  • 4. CityGen Datasets
  • 5. Experiments
  • 5.1. Evaluation Protocols
  • 5.2. Implementation Details
  • 5.3. Main Results
  • 5.4. Ablation Study
  • 5.5. Further Discussions
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Compositional Architecture and Rendering Pipeline of CityDreamer

    model/method

    CityDreamer is a compositional generative framework designed to synthesize unbounded, photorealistic, and 3D-consistent urban environments by decoupling the generation of background terrain ("stuff") from individual building architectures ("instances").

    The generation pipeline operates in four stages:

    1. Unbounded City Layout Generation: An autoregressive transformer operating on vector-quantized height fields and semantic maps generates an arbitrarily large 3D scene layout volume L\mathbf{L}.
    2. City Background Generation: A volumetric neural renderer parameterized by a generative hash grid renders the background stuff image I^G\hat{\mathbf{I}}_\text{G} and its binary mask MG\mathbf{M}_\text{G} covering roads, green lands, water bodies, and terrain.
    3. Building Instance Generation: An instance-specific volumetric neural renderer parameterized by periodic positional encodings renders appearance images {I^Bi}i=1n\{\hat{\mathbf{I}}_{\text{B}_i}\}_{i=1}^n and masks {MBi}i=1n\{\mathbf{M}_{\text{B}_i}\}_{i=1}^n for each of the nn individual building instances in the scene, conditioned on style latent codes.
    4. Composition: The rendered background image and instance images are composited into a coherent single multi-view consistent image IC\mathbf{I}_\text{C} according to their respective binary visibility masks: IC=I^GMG+∑i=1nI^BiMBi\mathbf{I}_\text{C} = \hat{\mathbf{I}}_\text{G} \mathbf{M}_\text{G} + \sum_{i=1}^n \hat{\mathbf{I}}_{\text{B}_i} \mathbf{M}_{\text{B}_i}
  2. Knowl 2 — Unbounded 3D City Layout Representation and Generation

    model/method

    The 3D city layout L∈{0,1,…,C}H×W×D\mathbf{L} \in \{0, 1, \dots, C\}^{H \times W \times D} is defined as a feature-free voxel volume constructed by vertically extruding 2D semantic maps S∈{1,…,C}H×W\mathbf{S} \in \{1, \dots, C\}^{H \times W} up to heights specified by a 2D height field H∈RH×W\mathbf{H} \in \mathbb{R}^{H \times W}: L(i,j,k)={S(i,j)if k≤H(i,j)0otherwise\mathbf{L}_{(i, j, k)} = \begin{cases} \mathbf{S}_{(i, j)} & \text{if } k \le \mathbf{H}_{(i, j)} \\ 0 & \text{otherwise} \end{cases} where 00 denotes empty voxel space, and semantic classes include roads, buildings, green lands, construction sites, water areas, and others.

    To generate extendable layouts, patches of paired (H,S)(\mathbf{H}, \mathbf{S}) are tokenized into a discrete codebook C={ck∈RD}k=1K\mathcal{C} = \{\mathbf{c}_k \in \mathbb{R}^D\}_{k=1}^K using a VQ-VAE trained via the loss: ℓVQ=λR∥H^p−Hp∥1+λSS(H^p,Hp)+λEE(S^p,Sp)\ell_\text{VQ} = \lambda_\text{R} \|\hat{\mathbf{H}}_p - \mathbf{H}_p\|_1 + \lambda_\text{S} \mathcal{S}(\hat{\mathbf{H}}_p, \mathbf{H}_p) + \lambda_\text{E} \mathcal{E}(\hat{\mathbf{S}}_p, \mathbf{S}_p) where H^p\hat{\mathbf{H}}_p and S^p\hat{\mathbf{S}}_p are generated patches, Hp\mathbf{H}_p and Sp\mathbf{S}_p are ground-truth patches, S\mathcal{S} is an edge-preserving smoothness loss to enforce sharp building boundaries, E\mathcal{E} is cross-entropy loss, and hyperparameter weights are set to λR=10\lambda_\text{R} = 10, λS=10\lambda_\text{S} = 10, λE=1\lambda_\text{E} = 1.

    During inference, layout tokens are predicted autoregressively with a MaskGIT transformer using a sliding-window strategy with a 25% window overlap to enable infinite layout extrapolation.

  3. Knowl 3 — City Background Generator with Generative Neural Hash Grid

    model/method

    The City Background Generator synthesizes background stuff (such as roads, greenery, and water) from a local layout window LGLocal\mathbf{L}_\text{G}^\text{Local} defined by local height field HGLocal\mathbf{H}_\text{G}^\text{Local} and semantic map SGLocal\mathbf{S}_\text{G}^\text{Local}.

    A global encoder EGE_\text{G} first compresses the local window into a compact scene-level feature vector fG∈RdG\mathbf{f}_\text{G} \in \mathbb{R}^{d_\text{G}} (with dG=2d_\text{G} = 2): fG=EG(HGLocal,SGLocal)\mathbf{f}_\text{G} = E_\text{G}(\mathbf{H}_\text{G}^\text{Local}, \mathbf{S}_\text{G}^\text{Local}) To enable cross-scene generalization and 3D consistency for irregular natural textures, a 3D position p=(px,py,pz)∈R3\mathbf{p} = (p_x, p_y, p_z) \in \mathbb{R}^3 and scene feature fG\mathbf{f}_\text{G} are indexed into a multi-resolution generative neural hash grid H:R3+dG→RNGC\mathcal{H}: \mathbb{R}^{3 + d_\text{G}} \to \mathbb{R}^{N_\text{G}^\text{C}} across NHL=16N_\text{H}^L = 16 resolution levels: fGp=H(p,fG)=(⨁i=1dGfGiπi⨁j=13pjπj) mod T\mathbf{f}_\text{G}^{\mathbf{p}} = \mathcal{H}(\mathbf{p}, \mathbf{f}_\text{G}) = \left( \bigoplus_{i=1}^{d_\text{G}} f_\text{G}^i \pi^i \bigoplus_{j=1}^3 p^j \pi^j \right) \bmod T where ⊕\oplus is the bitwise XOR operator, {πi,πj}\{\pi^i, \pi^j\} are unique large primes (1,2654435761,805459861,3674653429,20971920371, 2654435761, 805459861, 3674653429, 2097192037), T=219T = 2^{19} is the table size per level, and NGC=8N_\text{G}^\text{C} = 8 is the feature channel dimension.

    Volumetric rendering along camera rays r(t)=o+tv\mathbf{r}(t) = \mathbf{o} + t\mathbf{v} yields pixel color values: C(r)=∫0∞T(t)c(fGr(t),l(r(t)))σ(fGr(t))dtC(\mathbf{r}) = \int_0^\infty T(t) \mathbf{c}\left(\mathbf{f}_\text{G}^{\mathbf{r}(t)}, l(\mathbf{r}(t))\right) \boldsymbol{\sigma}\left(\mathbf{f}_\text{G}^{\mathbf{r}(t)}\right) dt where T(t)=exp⁡(−∫0tσ(fGr(s))ds)T(t) = \exp\left(-\int_0^t \boldsymbol{\sigma}(\mathbf{f}_\text{G}^{\mathbf{r}(s)}) ds\right), l(p)l(\mathbf{p}) is the semantic label at p\mathbf{p}, c\mathbf{c} denotes color, and σ\boldsymbol{\sigma} denotes volume density.

    The generator is trained using a composite loss restricted to background pixels: ℓG=λL1∥I^G−IG∥1+λPP(I^G,IG)+λGG(I^G,SG)\ell_\text{G} = \lambda_\text{L1} \|\hat{\mathbf{I}}_\text{G} - \mathbf{I}_\text{G}\|_1 + \lambda_\text{P} \mathcal{P}(\hat{\mathbf{I}}_\text{G}, \mathbf{I}_\text{G}) + \lambda_\text{G} \mathcal{G}(\hat{\mathbf{I}}_\text{G}, \mathbf{S}_\text{G}) where IG\mathbf{I}_\text{G} is the ground-truth background image, SG\mathbf{S}_\text{G} is the perspective semantic map, P\mathcal{P} is perceptual loss, G\mathcal{G} is GAN loss, and loss weights are λL1=10,λP=10,λG=0.5\lambda_\text{L1} = 10, \lambda_\text{P} = 10, \lambda_\text{G} = 0.5.

  4. Knowl 4 — Building Instance Generator with Periodic Positional Encoding and Style Modulation

    model/method

    The Building Instance Generator renders individual building instances Bi\text{B}_i using an instance-centric local coordinate representation and periodic positional encodings tailored to the repetitive structure of building façades.

    Individual buildings are segmented from the semantic map using connected component analysis. Within each local building volume LBiLocal\mathbf{L}_{\text{B}_i}^\text{Local} centered at 2D coordinate (cBix,cBiy)(c_{\text{B}_i}^x, c_{\text{B}_i}^y), distinct semantic labels are assigned to the roof (top-most voxel layer) and façade voxels, while other instances are masked to null.

    A local encoder EBE_\text{B} extracts a pixel-level 2D feature map fBi∈RNBH×NBW×NBC\mathbf{f}_{\text{B}_i} \in \mathbb{R}^{N_\text{B}^H \times N_\text{B}^W \times N_\text{B}^C} (with NBC=63N_\text{B}^C = 63 channels): fBi=EB(HBiLocal,SBiLocal)\mathbf{f}_{\text{B}_i} = E_\text{B}(\mathbf{H}_{\text{B}_i}^\text{Local}, \mathbf{S}_{\text{B}_i}^\text{Local}) For any 3D coordinate p=(px,py,pz)\mathbf{p} = (p_x, p_y, p_z), the point feature fBip\mathbf{f}_{\text{B}_i}^\mathbf{p} is obtained by concatenating the 2D local feature at (px,py)(p_x, p_y) with height coordinate pzp_z, followed by sinusoidal positional encoding O(⋅)\mathcal{O}(\cdot): fBip=O(Concat(fBi(px,py),pz))\mathbf{f}_{\text{B}_i}^\mathbf{p} = \mathcal{O}\left(\text{Concat}(\mathbf{f}_{\text{B}_i}^{(p_x, p_y)}, p_z)\right) where O(x)={sin⁡(2kπx),cos⁡(2kπx)}k=0NPL−1\mathcal{O}(x) = \{\sin(2^k \pi x), \cos(2^k \pi x)\}_{k=0}^{N_\text{P}^L - 1} with NPL=10N_\text{P}^L = 10.

    Volumetric rendering is modulated by a building style latent vector z\mathbf{z} to capture diverse building appearances under normalized ray coordinates centered at (cBix,cBiy,0)(c_{\text{B}_i}^x, c_{\text{B}_i}^y, 0): C(r)=∫0∞T(t)c(fBir(t),z,l(r(t)))σ(fBir(t))dtC(\mathbf{r}) = \int_0^\infty T(t) \mathbf{c}\left(\mathbf{f}_{\text{B}_i}^{\mathbf{r}(t)}, \mathbf{z}, l(\mathbf{r}(t))\right) \boldsymbol{\sigma}\left(\mathbf{f}_{\text{B}_i}^{\mathbf{r}(t)}\right) dt The building instance generator is trained exclusively with an adversarial GAN loss ℓB=G(I^Bi,SBi)\ell_\text{B} = \mathcal{G}(\hat{\mathbf{I}}_{\text{B}_i}, \mathbf{S}_{\text{B}_i}) on building-masked pixels.

  5. Knowl 5 — CityGen Datasets: OSM and GoogleEarth

    model/method

    To train unbounded 3D city generation models, two complementary datasets are constructed:

    1. OSM Dataset: Provides real-world city layout geometry by rasterizing OpenStreetMap vector data across 80 cities worldwide into paired semantic maps and height fields covering >6000 km2>6000\text{ km}^2. Rasterization is performed at EPSG:3857 coordinate projection at zoom level 18 (∼0.597 m/pixel\sim 0.597\text{ m/pixel}). Semantic maps delineate 6 classes (roads, buildings, green lands, construction sites, water areas, others). Height fields specify OSM building heights, road height is fixed to 4, water height to 0, and tree heights are sampled from Perlin noise in the range [8,16][8, 16].
    2. GoogleEarth Dataset: Captures photorealistic multi-view urban imagery via Google Earth Studio. It contains 400 orbit trajectories across Manhattan and Brooklyn (New York City), totaling 24,000 images (960×540960 \times 540 resolution, 60 frames per trajectory, orbit radii 125–813 m125\text{--}813\text{ m}, altitudes 112–884 m112\text{--}884\text{ m}, camera elevations 15∘–70∘15^\circ\text{--}70^\circ) covering 25 km225\text{ km}^2. Automated semantic and building instance segmentation masks are computed by projecting the 3D layout volume (extruded from OSM data) onto camera views using calibrated intrinsic and extrinsic camera matrices.
  6. Knowl 6 — Quantitative Evaluation of 3D City Synthesis Performance

    empirical result

    CityDreamer was benchmarked on the GoogleEarth test split against state-of-the-art 3D-aware scene generation methods: SGAM, PersistentNature, and SceneDreamer (adapted with CityDreamer's layout generator). Evaluation evaluated visual quality (FID, KID), 3D geometry accuracy (Depth Error, DE), and multi-view consistency (Camera Error, CE).

    Methods FID ↓\downarrow KID ↓\downarrow DE ↓\downarrow CE ↓\downarrow
    SGAM 277.64 0.358 0.575 239.291
    PersistentNature 123.83 0.109 0.326 86.371
    SceneDreamer 213.56 0.216 0.152 0.186
    CityDreamer 97.38 0.096 0.147 0.060
    • FID / KID: Computed between 15,000 generated frames and 15,000 real GoogleEarth frames.
    • Depth Error (DE): L2 distance between normalized predicted NeRF depth and pseudo-ground-truth depth from DPT on 100 frames.
    • Camera Error (CE): Scale-invariant normalized L2 distance between COLMAP-reconstructed camera trajectories and ground-truth camera trajectories.

    CityDreamer achieved substantially superior visual fidelity (FID 97.3897.38 vs. second-best 123.83123.83) and multi-view geometric consistency (CE 0.0600.060 vs. second-best 0.1860.186).

  7. Knowl 7 — Ablation Study on Generative Scene Parameterization Strategies

    empirical result

    An ablation study evaluated combinations of encoder architectures (Global vs. Local) and positional encodings (Generative HashGrid vs. SinCos) for the City Background Generator (CBG) and Building Instance Generator (BIG).

    CBG BIG FID ↓\downarrow KID ↓\downarrow DE ↓\downarrow CE ↓\downarrow
    Enc. P.E. Enc. P.E.
    Local SinCos Global Hash 219.30 0.233 0.154 0.452
    Local SinCos Local SinCos 107.63 0.125 0.149 0.078
    Global Hash Global Hash 213.56 0.216 0.153 0.186
    Global Hash Local SinCos 97.38 0.096 0.147 0.060

    The results demonstrate that:

    1. A Global Encoder with HashGrid is optimal for background terrain generation, where broad spatial context and cross-scene feature sharing preserve naturalness and multi-view consistency across irregular surfaces.
    2. A Local Encoder with SinCos Positional Encoding is optimal for building instances, as SinCos's inherent periodicity matches the regular, repeating structural patterns found on architectural façades.
  8. Knowl 8 — Ablation on Building Instance Disentanglement and Layout Generation

    empirical result

    Ablations assessed the necessity of separating building instances from background generation, as well as the quality of the VQ-VAE + MaskGIT unbounded layout generator compared to existing layout generation approaches.

    Impact of Building Instance Disentanglement:

    Methods FID ↓\downarrow KID ↓\downarrow DE ↓\downarrow CE ↓\downarrow
    w/o BIG 213.56 0.216 0.152 0.186
    w/o Ins 117.75 0.124 0.148 0.098
    Ours (Full CityDreamer) 97.38 0.096 0.147 0.060
    • Removing the Building Instance Generator (w/o BIG, reverting to a unified SceneDreamer-style network) causes severe performance drop (FID 213.56213.56).
    • Generating all buildings simultaneously without instance labels (w/o Ins) improves over w/o BIG but underperforms full instance disentanglement (FID 117.75117.75 vs. 97.3897.38).

    Layout Generation Quality (Evaluated on 4096×40964096 \times 4096 cropped maps):

    Methods FID ↓\downarrow KID ↓\downarrow
    IPSM 321.47 0.502
    InfinityGAN 183.14 0.288
    Ours (Layout Generator) 124.45 0.123

    The proposed VQ-VAE + MaskGIT approach generates higher-fidelity 2D layouts than procedural street modeling (IPSM) and patch-extrapolation GANs (InfinityGAN).

  9. Knowl 9 — Geometric and Computational Limitations of CityDreamer

    limitation

    CityDreamer is subject to two main technical limitations:

    1. Inability to Model Concave Geometries: Because the 3D scene representation is constructed by vertically extruding a 2.5D height field H\mathbf{H} and semantic map S\mathbf{S}, the representation is strictly elevation-based. Consequently, overhanging structures and concave vertical geometries—such as tunnels, underground passages, bridges, and caves—cannot be represented or generated.
    2. Inference Compute Scaling with Instance Count: Because building instances are extracted, encoded, and volumetrically rendered individually before composition, overall inference computational cost and rendering latency scale with the total number of building instances present in the camera view frustum.

Coverage note — None was omitted; all key architectural components, dataset constructions, mathematical formulations, loss functions, empirical results, ablations, and stated limitations are fully covered.

References

  1. 1.https://openstreetmap.org. 2, 5
  2. 2.https://earth.google.com/studio. 2, 5
  3. 3.Miguel Angel Bautista, Pengsheng Guo, Samira Abnar, Walter Talbott, Alexander Toshev, Zhuoyuan Chen, Laurent Dinh, Shuangfei Zhai, Hanlin Goh, Daniel Ulbricht, Afshin Dehghan, and Joshua M. Susskind. GAUDI: A neural architect for immersive 3d scene generation. In NeurIPS, 2022. 2
  4. 4.Mikolaj Binkowski, Danica J. Sutherland, Michael Arbel, and Arthur Gretton. Demystifying MMD GANs. In ICLR, 2018. 6
  5. 5.Sam Bond-Taylor, Peter Hessey, Hiroshi Sasaki, Toby P. Breckon, and Chris G. Willcocks. Unleashing transformers: Parallel token prediction with discrete absorbing diffusion for fast high-resolution image generation from vectorquantized codes. In ECCV, 2022. 3
  6. 6.Zoya Bylinskii, Laura Mariah Herman, Aaron Hertzmann, Stefanie Hutka, and Yile Zhang. Towards better user studies in computer graphics and vision. Foundations and Trends in Computer Graphics and Vision, 15(3):201–252, 2023. 7
  7. 7.Lucy Chai, Richard Tucker, Zhengqi Li, Phillip Isola, and Noah Snavely. Persistent Nature: A generative model of unbounded 3D worlds. In CVPR, 2023. 1, 2, 6, 8
  8. 8.Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J. Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3D generative adversarial networks. In CVPR, 2022. 2, 6
  9. 9.Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman. MaskGIT: Masked generative image transformer. In CVPR, 2022. 3
  10. 10.Guoning Chen, Gregory Esch, Peter Wonka, Pascal M¨uller, and Eugene Zhang. Interactive procedural street modeling. ACM TOG, 27(3):103, 2008. 7, 8
  11. 11.Zhaoxi Chen, Guangcong Wang, and Ziwei Liu. SceneDreamer: Unbounded 3D scene generation from 2D image collections. TPAMI, 45(12):15562–15576, 2023. 1, 2, 3, 4, 6
  12. 12.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016. 5
  13. 13.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas A. Funkhouser, and Matthias Nießner. ScanNet: Richly-annotated 3D reconstructions of indoor scenes. In CVPR, 2017. 2
  14. 14.Terrance DeVries, Miguel Angel Bautista, Nitish Srivastava, Graham W. Taylor, and Joshua M. Susskind. Unconstrained scene generation with locally conditioned radiance fields. In ICCV, 2021. 2
  15. 15.Patrick Esser, Robin Rombach, and Bj¨orn Ommer. Taming transformers for high-resolution image synthesis. In CVPR, 2021. 2
  16. 16.Huan Fu, Bowen Cai, Lin Gao, Lingxiao Zhang, Jiaming Wang, Cao Li, Qixun Zeng, Chengyue Sun, Rongfei Jia, Binqiang Zhao, and Hao Zhang. 3D-FRONT: 3D furnished rooms with layouts and semantics. In ICCV, 2021. 2
  17. 17.Matheus Gadelha, Subhransu Maji, and Rui Wang. 3d shape induction from 2D views of multiple objects. In 3DV, 2017. 2
  18. 18.Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, and Sanja Fidler. GET3D: A generative model of high quality 3D textured shapes learned from images. In NeurIPS, 2022. 2
  19. 19.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the KITTI vision benchmark suite. In CVPR, 2012. 5
  20. 20.Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial nets. In NIPS, 2014. 2
  21. 21.Jiatao Gu, Lingjie Liu, Peng Wang, and Christian Theobalt. StyleNeRF: A style-based 3D aware generator for highresolution image synthesis. In ICLR, 2022. 2
  22. 22.Zekun Hao, Arun Mallya, Serge J. Belongie, and Ming-Yu Liu. GANCraft: Unsupervised 3D neural rendering of minecraft worlds. In ICCV, 2021. 1, 2, 3
  23. 23.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. GANs trained by a two time-scale update rule converge to a local nash equilibrium. In NIPS, 2017. 6
  24. 24.Fangzhou Hong, Zhaoxi Chen, Yushi Lan, Liang Pan, and Ziwei Liu. EVA3D: compositional 3D human generation from 2D image collections. In ICLR, 2023. 1
  25. 25.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments. IEEE TPAMI, 36(7):1325–1339, 2014. 2
  26. 26.Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In ECCV, 2016. 4
  27. 27.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. IEEE TPAMI, 43(12):4217–4228, 2021. 2
  28. 28.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In CVPR, 2020. 2
  29. 29.Nikos Kolotouros, Thiemo Alldieck, Andrei Zanfir, Eduard Gabriel Bazavan, Mihai Fieraru, and Cristian Sminchisescu. DreamHuman: Animatable 3D avatars from text. arXiv, 2306.09329, 2023. 1
  30. 30.Weijia Li, Yawen Lai, Linning Xu, Yuanbo Xiangli, Jinhua Yu, Conghui He, Gui-Song Xia, and Dahua Lin. OmniCity: Omnipotent city understanding with multi-level and multiview images. In CVPR, 2023. 5
  31. 31.Zhengqi Li, Qianqian Wang, Noah Snavely, and Angjoo Kanazawa. InfiniteNature-Zero: Learning perpetual view generation of natural scenes from single images. In ECCV, 2022. 2
  32. 32.Jae Hyun Lim and Jong Chul Ye. Geometric GAN. arXiv, 1705.02894, 2017. 4
  33. 33.Chieh Hubert Lin, Hsin-Ying Lee, Yen-Chi Cheng, Sergey Tulyakov, and Ming-Hsuan Yang. InfinityGan: Towards infinite-pixel image synthesis. In ICLR, 2022. 7, 8
  34. 34.Chieh Hubert Lin, Hsin-Ying Lee, Willi Menapace, Menglei Chai, Aliaksandr Siarohin, Ming-Hsuan Yang, and Sergey Tulyakov. InfiniCity: Infinite-scale city synthesis. In ICCV, 2023. 1, 2, 3, 6, 7
  35. 35.Liqiang Lin, Yilin Liu, Yue Hu, Xingguang Yan, Ke Xie, and Hui Huang. Capturing, reconstructing, and simulating: The urbanscene3d dataset. In ECCV, 2022. 5
  36. 36.Andrew Liu, Ameesh Makadia, Richard Tucker, Noah Snavely, Varun Jampani, and Angjoo Kanazawa. Infinite Nature: Perpetual view generation of natural scenes from a single image. In ICCV, 2021. 2
  37. 37.Arun Mallya, Ting-Chun Wang, Karan Sapra, and Ming-Yu Liu. World-consistent video-to-video synthesis. In ECCV, 2020. 2
  38. 38.Simon Meister, Junhwa Hur, and Stefan Roth. UnFlow: Unsupervised learning of optical flow with a bidirectional census loss. In AAAI, 2018. 3
  39. 39.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 2, 4
  40. 40.Thu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt, and Yong-Liang Yang. HoloGAN: Unsupervised learning of 3D representations from natural images. In CVPR, 2019. 2
  41. 41.Roy Or-El, Xuan Luo, Mengyi Shan, Eli Shechtman, Jeong Joon Park, and Ira Kemelmacher-Shlizerman. StyleSDF: High-resolution 3D-consistent image and geometry generation. In CVPR, 2022. 2
  42. 42.Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. Semantic image synthesis with spatially-adaptive normalization. In CVPR, 2019. 1, 2
  43. 43.Despoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis, Andreas Geiger, and Sanja Fidler. ATISS: autoregressive transformers for indoor scene synthesis. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, editors, NeurIPS, 2021. 2
  44. 44.Ken Perlin. An image synthesizer. In SIGGRAPH, 1985. 5
  45. 45.Amit Raj, Srinivas Kaza, Ben Poole, Michael Niemeyer, Nataniel Ruiz, Ben Mildenhall, Shiran Zada, Kfir Aberman, Michael Rubinstein, Jonathan T. Barron, Yuanzhen Li, and Varun Jampani. DreamBooth3D: Subject-driven text-to-3D generation. arXiv, 2303.13508, 2023. 1
  46. 46.Ren´e Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE TPAMI, 44(3):1623–1637, 2022. 6
  47. 47.Ali Razavi, Aaron van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with VQ-VAE-2. In NeurIPS, 2019. 3
  48. 48.Johannes L. Sch¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. In CVPR, 2016. 6
  49. 49.Yuan Shen, Wei-Chiu Ma, and Shenlong Wang. SGAM: building a virtual 3D world through simultaneous generation and mapping. In NeurIPS, 2022. 6
  50. 50.Zifan Shi, Yujun Shen, Jiapeng Zhu, Dit-Yan Yeung, and Qifeng Chen. 3D-aware indoor scene synthesis with depth priors. In ECCV, 2022. 2
  51. 51.Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, Raul Mur-Artal, Carl Yuheng Ren, Shobhit Verma, Anton Clarkson, Mingfei Yan, Brian Budge, Yajie Yan, Xiaqing Pan, June Yon, Yuyang Zou, Kimberly Leon, Nigel Carter, Jesus Briales, Tyler Gillingham, Elias Mueggler, Luis Pesqueira, Manolis Savva, Dhruv Batra, Hauke M. Strasdat, Renzo De Nardi, Michael Goesele, Steven Lovegrove, and Richard A. Newcombe. The Replica Dataset: A digital replica of indoor spaces. arXiv, 1906.05797, 2019. 2
  52. 52.Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. In NIPS, 2017. 3
  53. 53.Xinpeng Wang, Chandan Yeshwanth, and Matthias Nießner. Sceneformer: Indoor scene generation with transformers. In 3DV, 2021. 2
  54. 54.Nicholas Weir, David Lindenbaum, Alexei Bastidas, Adam Van Etten, Varun Kumar Vijay, Sean McPherson, Jacob Shermeyer, and Hanlin Tang. Spacenet MVOI: A multi-view overhead imagery dataset. In ICCV, 2019. 5
  55. 55.Jiajun Wu, Chengkai Zhang, Tianfan Xue, Bill Freeman, and Josh Tenenbaum. Learning a probabilistic latent space of object shapes via 3D generative-adversarial modeling. In NIPS, 2016. 2
  56. 56.Tong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang, Jiawei Ren, Liang Pan, Wayne Wu, Lei Yang, Jiaqi Wang, Chen Qian, Dahua Lin, and Ziwei Liu. OmniObject3D: Large-vocabulary 3D object dataset for realistic perception, reconstruction and generation. In CVPR, 2023. 2
  57. 57.Haozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, and Shengping Zhang. Pix2Vox: Context-aware 3D reconstruction from single and multi-view images. In ICCV, 2019. 1
  58. 58.Haozhe Xie, Hongxun Yao, Shengping Zhang, Shangchen Zhou, and Wenxiu Sun. Pix2Vox++: Multi-scale context-aware 3D object reconstruction from single and multiple images. IJCV, 128(12):2919–2935, 2020. 1
  59. 59.Yang Xue, Yuheng Li, Krishna Kumar Singh, and Yong Jae Lee. GIRAFFE HD: A high-resolution 3D-aware generative model. In CVPR, 2022. 2
  60. 60.Chi Zhang, Yiwen Chen, Yijun Fu, Zhenglin Zhou, Gang Yu, Billzb Wang, Bin Fu, Tao Chen, Guosheng Lin, and Chunhua Shen. StyleAvatar3D: Leveraging image-text diffusion models for high-fidelity 3D avatar generation. arXiv, 2305.19012, 2023. 1
  61. 61.Yichao Zhou, Jingwei Huang, Xili Dai, Linjie Luo, Zhili Chen, and Yi Ma. HoliCity: A city-scale data platform for learning holistic 3D structures. arXiv, 2008.03286, 2020. 5

Citation

MLA
Xie, H., et al. “CityDreamer: Compositional Generative Model of Unbounded 3D Cities”. arXiv, 2023, http://arxiv.org/abs/2309.00610v3.
APA
Xie, H., Chen, Z., Hong, F., & Liu, Z. (2023). CityDreamer: Compositional Generative Model of Unbounded 3D Cities. arXiv. http://arxiv.org/abs/2309.00610v3
Chicago
Xie, H., Z. Chen, F. Hong, and Z. Liu. 2023. “CityDreamer: Compositional Generative Model of Unbounded 3D Cities”. arXiv. http://arxiv.org/abs/2309.00610v3.
Harvard
Xie, H. et al. (2023) “CityDreamer: Compositional Generative Model of Unbounded 3D Cities”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2309.00610v3.
Vancouver
1. Xie H, Chen Z, Hong F, Liu Z (2023) CityDreamer: Compositional Generative Model of Unbounded 3D Cities. arXiv

BibTeX

@article{xie2023citydreamer,
  title = {CityDreamer: Compositional Generative Model of Unbounded 3D Cities},
  author = {Xie, Haozhe and Chen, Zhaoxi and Hong, Fangzhou and Liu, Ziwei},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2309.00610v3},
  eprint = {2309.00610}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE