Map2World: Segment Map Conditioned Text to 3D World Generation
Jaeyoung ChungSuyoung LeeJianfeng XiangJiaolong YangKyoung Mu Lee
Presents Map2World, a framework that synthesizes coherent, large-scale 3D environments conditioned on arbitrary user-defined segment maps to overcome scale inconsistencies and rigid grid constraints in 3D world generation.
Generating expansive, high-fidelity 3D virtual environments is critical for industries such as autonomous vehicle simulation, video game development, and immersive virtual reality. While recent generative artificial intelligence models can create individual 3D objects with remarkable fidelity, scaling these models to build entire worlds remains difficult due to a severe shortage of comprehensive, large-scale 3D scene datasets. Existing approaches typically rely on rigid square grid layouts, synthesize narrow single-domain environments like simple driving tracks, or exhibit scale inconsistencies and disconnected seams between adjacent objects.
The article demonstrates Map2World, a generative framework that synthesizes large-scale, highly coherent 3D worlds directly from text prompts and user-defined 2D segmentation maps of arbitrary shapes and sizes. The system aims to provide flexible layout control while maintaining visual fidelity, geometric harmony, and scale consistency across broad environments without requiring costly ground-up retraining.
To overcome the lack of expansive training datasets, the authors built their approach atop a pre-trained state-of-the-art 3D asset generator known as TRELLIS. The methodology introduces three core components: multi-window latent fusion across overlapping 3D volumes to seamlessly blend regions defined by arbitrary boundaries; a spectral-domain noise optimization technique that aligns initial structural generation to intended world scales; and an efficient detail enhancer network. This detail enhancer was trained on 16,000 paired cube subdivisions extracted from 35 high-quality 3D scenes, tuning only a small neural module (approximately 4% of total network parameters) to super-resolve local geometries while referencing adjacent cube context and global structure.
The evaluation yielded several key findings regarding world generation performance and layout control. First, Map2World achieved a superior overall World Quality score of 7.76 out of 10, outperforming baseline models such as SynCity (7.25) and GaussianCube (5.08), with pronounced advantages in world completeness (7.8 versus 6.8 for SynCity) and structural coherence (7.9 versus 7.6). Second, region-wise semantic alignment metrics demonstrated that the generated 3D environments faithfully adhered to irregular, free-form district shapes where prior tile-based baselines failed. Third, spectral-domain noise optimization allowed the framework to achieve spatial scale convergence within five optimization steps at a high learning rate, whereas unparameterized optimization diverged. Finally, the proposed detail enhancement module delivered superior perceptual quality and visual sharpness compared to alternative adapter architectures and guidance strategies.
These findings indicate that industrial 3D generation can move away from rigid grid-based tiling toward flexible, bespoke spatial planning without requiring massive, newly captured world datasets. For technical leaders and simulation developers, this approach substantially lowers computational and data acquisition costs by repurposing pre-trained object-level generative models for macro-scale world synthesis. It also reduces development timelines by allowing human designers or automated pipelines to direct large 3D scene layouts using intuitive 2D semantic maps.
Organizations developing virtual environments should consider adopting latent fusion and parameter-efficient enhancement pipelines for scalable 3D asset production. To further advance production readiness, future development should incorporate richer textured datasets beyond basic meshes and transition toward relative positional encoding schemes. While confidence in the framework's layout flexibility and coherence is high, current limitations stem from inherited absolute positional encodings, which can introduce minor structural shifts when merging smaller sub-cubes into large unified scenes.
- Paper: CityDreamer: Compositional Generative Model of Unbounded 3D Cities, Haozhe Xie et al. (2024). CityDreamer introduces compositionally generating unbounded 3D scenes from layout maps and neural rendering pathways, establishing the foundational paradigm that Map2World generalizes beyond rigid grids.
- Paper: DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis, Yinghao Xu et al. (2023). DisCoScene develops layout-guided disentanglement of object-centric radiance fields and background generators, providing essential conceptual groundwork for controllable and modular 3D scene synthesis.
- Paper: XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies, Xuanchi Ren et al. (2024). XCube demonstrates scalable 3D generative modeling of large outdoor scenes through hierarchical structures and progressive detail refinement, directly addressing the scale and resolution challenges Map2World tackles.
- Paper: Magic3D: High-Resolution Text-to-3D Content Creation, Chen-Hsuan Lin et al. (2022). Magic3D establishes effective coarse-to-fine 3D generation using 2D diffusion priors, underlying the asset generation and detail enhancement techniques leveraged in Map2World.
- Paper: High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs, Ting-Chun Wang et al. (2018). This paper presents foundational methods for conditioning generative pipelines on semantic segment maps and synthesizing high-resolution details via multi-scale architectures.
No sufficiently relevant recommendations were found.
