Map2World: Segment Map Conditioned Text to 3D World Generation

Jaeyoung ChungSuyoung LeeJianfeng XiangJiaolong YangKyoung Mu Lee

article2026arXiv0 citations

Presents Map2World, a framework that synthesizes coherent, large-scale 3D environments conditioned on arbitrary user-defined segment maps to overcome scale inconsistencies and rigid grid constraints in 3D world generation.

Listen

Generating expansive, high-fidelity 3D virtual environments is critical for industries such as autonomous vehicle simulation, video game development, and immersive virtual reality. While recent generative artificial intelligence models can create individual 3D objects with remarkable fidelity, scaling these models to build entire worlds remains difficult due to a severe shortage of comprehensive, large-scale 3D scene datasets. Existing approaches typically rely on rigid square grid layouts, synthesize narrow single-domain environments like simple driving tracks, or exhibit scale inconsistencies and disconnected seams between adjacent objects.

The article demonstrates Map2World, a generative framework that synthesizes large-scale, highly coherent 3D worlds directly from text prompts and user-defined 2D segmentation maps of arbitrary shapes and sizes. The system aims to provide flexible layout control while maintaining visual fidelity, geometric harmony, and scale consistency across broad environments without requiring costly ground-up retraining.

To overcome the lack of expansive training datasets, the authors built their approach atop a pre-trained state-of-the-art 3D asset generator known as TRELLIS. The methodology introduces three core components: multi-window latent fusion across overlapping 3D volumes to seamlessly blend regions defined by arbitrary boundaries; a spectral-domain noise optimization technique that aligns initial structural generation to intended world scales; and an efficient detail enhancer network. This detail enhancer was trained on 16,000 paired cube subdivisions extracted from 35 high-quality 3D scenes, tuning only a small neural module (approximately 4% of total network parameters) to super-resolve local geometries while referencing adjacent cube context and global structure.

The evaluation yielded several key findings regarding world generation performance and layout control. First, Map2World achieved a superior overall World Quality score of 7.76 out of 10, outperforming baseline models such as SynCity (7.25) and GaussianCube (5.08), with pronounced advantages in world completeness (7.8 versus 6.8 for SynCity) and structural coherence (7.9 versus 7.6). Second, region-wise semantic alignment metrics demonstrated that the generated 3D environments faithfully adhered to irregular, free-form district shapes where prior tile-based baselines failed. Third, spectral-domain noise optimization allowed the framework to achieve spatial scale convergence within five optimization steps at a high learning rate, whereas unparameterized optimization diverged. Finally, the proposed detail enhancement module delivered superior perceptual quality and visual sharpness compared to alternative adapter architectures and guidance strategies.

These findings indicate that industrial 3D generation can move away from rigid grid-based tiling toward flexible, bespoke spatial planning without requiring massive, newly captured world datasets. For technical leaders and simulation developers, this approach substantially lowers computational and data acquisition costs by repurposing pre-trained object-level generative models for macro-scale world synthesis. It also reduces development timelines by allowing human designers or automated pipelines to direct large 3D scene layouts using intuitive 2D semantic maps.

Organizations developing virtual environments should consider adopting latent fusion and parameter-efficient enhancement pipelines for scalable 3D asset production. To further advance production readiness, future development should incorporate richer textured datasets beyond basic meshes and transition toward relative positional encoding schemes. While confidence in the framework's layout flexibility and coherence is high, current limitations stem from inherited absolute positional encodings, which can introduce minor structural shifts when merging smaller sub-cubes into large unified scenes.

No sufficiently relevant recommendations were found.

Cover for Map2World: Segment Map Conditioned Text to 3D World Generation

Abstract

3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from inconsistencies in object scale throughout the entire world. In this work, we introduce a novel framework, Map2World, that first enables 3D world generation conditioned on user-defined segment maps of arbitrary shapes and scales, ensuring global-scale consistency and flexibility across expansive environments. To further enhance the quality, we propose a detail enhancer network that generates fine details of the world. The detail enhancer enables the addition of fine-grained details without compromising overall scene coherence by incorporating global structure information. We design the entire pipeline to leverage strong priors from asset generators, achieving robust generalization across diverse domains, even under limited training data for scene generation. Extensive experiments demonstrate that our method significantly outperforms existing approaches in user-controllability, scale consistency, and content coherence, enabling users to generate 3D worlds under more complex conditions.

Citation

MLA
Chung, J., et al. “Map2World: Segment Map Conditioned Text to 3D World Generation”. Lecture Notes in Computer Science, Springer Nature Switzerland, 2026, pp. 337–54, https://doi.org/10.1007/978-3-032-37225-3_19.
APA
Chung, J., Lee, S., Xiang, J., Yang, J., & Lee, K. M. (2026). Map2World: Segment Map Conditioned Text to 3D World Generation. In Lecture Notes in Computer Science (pp. 337–354). Springer Nature Switzerland. https://doi.org/10.1007/978-3-032-37225-3_19
Chicago
Chung, J., S. Lee, J. Xiang, J. Yang, and K. M. Lee. 2026. “Map2World: Segment Map Conditioned Text to 3D World Generation”. In Lecture Notes in Computer Science. Springer Nature Switzerland. https://doi.org/10.1007/978-3-032-37225-3_19.
Harvard
Chung, J. et al. (2026) “Map2World: Segment Map Conditioned Text to 3D World Generation”, Lecture Notes in Computer Science. Springer Nature Switzerland, pp. 337–354. Available at: https://doi.org/10.1007/978-3-032-37225-3_19.
Vancouver
1. Chung J, Lee S, Xiang J, Yang J, Lee KM (2026) Map2World: Segment Map Conditioned Text to 3D World Generation. In: Lecture Notes in Computer Science. Springer Nature Switzerland, pp 337–354

BibTeX

@inbook{Chung_2026, title={Map2World: Segment Map Conditioned Text to 3D World Generation}, ISBN={9783032372253}, ISSN={1611-3349}, url={http://dx.doi.org/10.1007/978-3-032-37225-3_19}, DOI={10.1007/978-3-032-37225-3_19}, booktitle={Computer Vision – ECCV 2026}, publisher={Springer Nature Switzerland}, author={Chung, Jaeyoung and Lee, Suyoung and Xiang, Jianfeng and Yang, Jiaolong and Lee, Kyoung Mu}, year={2026}, pages={337–354} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/