XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies
Xuanchi RenJiahui HuangXiaohui ZengKen MusethSanja FidlerFrancis Williams
Presents a hierarchical sparse voxel latent diffusion framework built on VDB data structures that rapidly generates high-resolution 3D objects and large-scale outdoor scenes with rich geometric and semantic attributes without test-time optimization.
Generating realistic three-dimensional assets and large-scale environments is critical for fields such as robotics, autonomous driving, gaming, and digital simulation. However, existing 3D generative techniques struggle to scale. Methods relying on 2D image models require lengthy per-asset optimization and frequently introduce spatial errors, while direct 3D models are typically restricted to low resolutions or simple single objects due to heavy memory and computational constraints.
The article introduces and evaluates XCube, a 3D generative system designed to rapidly produce high-resolution 3D objects and expansive outdoor driving scenes enriched with structural and semantic details.
To achieve this, the approach generates 3D content across a hierarchy of sparse voxel grids—3D pixel representations that only store data where geometry actually exists. The method uses a cascaded latent diffusion process, which first generates a rough coarse shape and progressively synthesizes finer details at higher resolutions. The system is built on an optimized sparse computing framework that processes high-resolution volumetric grids directly on modern graphics hardware without requiring lengthy test-time optimization. Evaluation was conducted across standard object benchmarks (ShapeNet and Objaverse) and large real-world and synthetic driving scenes (Waymo Open Dataset and Karton City).
The evaluation yielded several key findings. First, the method scales to an effective resolution of 1024-cubed, generating millions of voxels for complex 3D objects in under 30 seconds and full 100-meter by 100-meter outdoor environments with 10-centimeter detail. Second, the framework runs approximately three times faster while consuming half the memory of leading sparse 3D processing alternatives. Third, on standard object generation benchmarks, the system consistently outperformed existing point, mesh, and dense-voxel baselines across shape similarity metrics. Fourth, user studies confirmed substantial quality improvements: 79.2% of evaluators preferred its text-to-3D object generations over a leading baseline, and 66.3% rated its synthetic urban driving scenes as more realistic than actual validation sensor data. Finally, the model proved flexible across practical tasks, including text-to-3D creation, intuitive multi-scale voxel editing, and complete 3D scene reconstruction from a single sparse lidar scan.
These findings demonstrate that direct 3D generative modeling can achieve high spatial resolution and geometric complexity without prohibitive computation times. For industrial applications such as autonomous driving simulation and 3D asset creation pipelines, this significantly lowers the computational cost and turnaround time required to build realistic virtual environments, reducing operational bottlenecks while enhancing consistency.
Organizations developing simulation pipelines or 3D digital content should consider adopting hierarchical sparse representations to scale their generative workflows. Technical teams can integrate coarse-to-fine editing tools into artist workflows and test single-scan completion models to accelerate autonomous vehicle simulation environments. Future initiatives should explore conditioning the architecture on standard 2D reference images and deploying the underlying generative prior across broader downstream perception tasks.
The primary limitation noted in the article is the comparative scarcity of comprehensive 3D training data relative to massive 2D web datasets, which restricts the model's ability to interpret complex, highly nuanced text prompts. Nonetheless, the findings provide a high degree of confidence that sparse hierarchical structures offer an efficient, scalable foundation for direct 3D content generation.
- Paper: 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks, Benjamin Graham et al. (2017). It introduces submanifold sparse convolutional networks, providing the core foundational mechanics for computing efficiently on sparse 3D structures without dilating active voxels.
- Paper: OctNet: Learning Deep 3D Representations at High Resolutions, Gernot Riegler et al. (2016). It introduces octree-based hierarchical voxel grids for scaling deep 3D representations, laying the direct conceptual groundwork for sparse voxel hierarchies in 3D generation.
- Paper: Cascaded Diffusion Models for High Fidelity Image Generation, Jonathan Ho et al. (2021). It establishes the coarse-to-fine cascaded diffusion framework that XCube adapts to generate low-resolution 3D shapes before progressively synthesizing fine voxel details.
- Paper: Objaverse: A Universe of Annotated 3D Objects, Matt Deitke et al. (2022). It supplies the large-scale 3D object dataset Objaverse, which serves as one of the primary training and evaluation benchmarks utilized in XCube.
- Paper: DreamFusion: Text-to-3D using 2D Diffusion, Ben Poole et al. (2023). It pioneered text-to-3D synthesis via Score Distillation Sampling, defining the 2D-prior optimization paradigm and its computational bottlenecks that direct 3D sparse generative models like XCube overcome.
- Paper: CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language, Aditya Sanghi et al. (2023). It establishes multi-resolution coarse-to-fine latent 3D shape generation conditioned on text, providing a direct antecedent to multi-scale hierarchical 3D generative architectures.
- Paper: Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling, Jiajun Wu et al. (2016). It is a seminal work on native 3D volumetric generative modeling, establishing the baseline concepts and resolution bottlenecks of direct 3D grid synthesis.
- Paper: Structured 3D Latents for Scalable and Versatile 3D Generation, Jianfeng Xiang et al. (2025). It generalizes structured 3D representations by combining sparse occupied voxel latents with rectified-flow transformers to decode into diverse 3D formats like meshes and Gaussians.
- Paper: CityDreamer: Compositional Generative Model of Unbounded 3D Cities, Haozhe Xie et al. (2024). It explores an alternative compositional neural rendering paradigm to scale 3D generative models to expansive, unbounded urban environments.
- Paper: SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction, Pin Tang et al. (2024). It applies sparse 3D latent representations and sparse diffusion modules directly to the downstream perception task of autonomous driving semantic occupancy prediction.
- Paper: VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction, Jiaqi Lin et al. (2024). It addresses large-scale scene scalability from a point-based 3D Gaussian Splatting perspective via spatial partitioning rather than hierarchical volumetric grids.
