keyword
scene generation
Scene generation is a computer vision and computer graphics process that utilizes generative artificial intelligence models to computationally synthesize complete visual environments containing multiple objects and their background surroundings. Unlike single-object synthesis, it emphasizes structural composition, realistic spatial layouts, geometry, and visual coherence across expansive two-dimensional images or three-dimensional virtual worlds. These systems often operate under user-defined conditions, converting inputs such as text descriptions, semantic layout maps, sketches, or bounding boxes into complex scenes while maintaining consistent lighting, perspective, and object scales. The synthesized outputs can be represented as explicit geometric structures, neural feature fields, or rendered images, enabling applications in virtual reality, video game development, visual effects, and simulation environments for autonomous systems.
3 items

Map2World: Segment Map Conditioned Text to 3D World Generation
Jaeyoung Chung, Suyoung Lee, Jianfeng Xiang, Jiaolong Yang, Kyoung Mu Lee
Why you should read this
Presents Map2World, a framework that synthesizes coherent, large-scale 3D environments conditioned on arbitrary user-defined segment maps to overcome scale inconsistencies and rigid grid constraints in 3D world generation.
3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from inconsistencies in object scale throughout the entire world. In this work, we introduce a novel framework, Map2World, that first enables 3D world generation conditioned on user-defined segment maps of arbitrary shapes and scales, ensuring global-scale consistency and flexibility across expansive environments. To further enhance the quality, we propose a detail enhancer network that generates fine details of the world. The detail enhancer enables the addition of fine-grained details without compromising overall scene coherence by incorporating global structure information. We design the entire pipeline to leverage strong priors from asset generators, achieving robust generalization across diverse domains, even under limited training data for scene generation. Extensive experiments demonstrate that our method significantly outperforms existing approaches in user-controllability, scale consistency, and content coherence, enabling users to generate 3D worlds under more complex conditions.
Added
2026-09-29

XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies
Xuanchi Ren, Jiahui Huang, Xiaohui Zeng, Ken Museth, Sanja Fidler, Francis Williams
Why you should read this
Presents a hierarchical sparse voxel latent diffusion framework built on VDB data structures that rapidly generates high-resolution 3D objects and large-scale outdoor scenes with rich geometric and semantic attributes without test-time optimization.
We present XCube (abbreviated as X3), a novel generative model for high-resolution sparse 3D voxel grids with arbitrary attributes. Our model can generate millions of voxels with a finest effective resolution of up to 10243 in a feed-forward fashion without time-consuming test-time optimization. To achieve this, we employ a hierarchical voxel latent diffusion model which generates progressively higher resolution grids in a coarse-to-fine manner using a custom framework built on the highly efficient VDB data structure. Apart from generating high-resolution objects, we demonstrate the effectiveness of XCube on large outdoor scenes at scales of 100 m×100 m with a voxel size as small as 10 cm. We observe clear qualitative and quantitative improvements over past approaches. In addition to unconditional generation, we show that our model can be used to solve a variety of tasks such as user-guided editing, scene completion from a single scan, and text-to-3D. More results and details can be found on our project webpage.
Added
2026-09-26

GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields
Michael Niemeyer, Andreas Geiger
Why you should read this
Combines volumetric scene representations with Generative Adversarial Networks using compositional localized feature fields to enable 3D-aware image synthesis.
Deep generative models allow for photorealistic image synthesis at high resolutions. But for many applications, this is not enough: content creation also needs to be controllable. While several recent works investigate how to disentangle underlying factors of variation in the data, most of them operate in 2D and hence ignore that our world is three-dimensional. Further, only few works consider the compositional nature of scenes. Our key hypothesis is that incorporating a compositional 3D scene representation into the generative model leads to more controllable image synthesis. Representing scenes as compositional generative neural feature fields allows us to disentangle one or multiple objects from the background as well as individual objects’ shapes and appearances while learning from unstructured and unposed image collections without any additional supervision. Combining this scene representation with a neural rendering pipeline yields a fast and realistic image synthesis model. As evidenced by our experiments, our model is able to disentangle individual objects and allows for translating and rotating them in the scene as well as changing the camera pose.
Added
2026-03-14
