XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies

Xuanchi RenJiahui HuangXiaohui ZengKen MusethSanja FidlerFrancis Williams

article2024CVPR157 citations

Presents a hierarchical sparse voxel latent diffusion framework built on VDB data structures that rapidly generates high-resolution 3D objects and large-scale outdoor scenes with rich geometric and semantic attributes without test-time optimization.

Listen

Generating realistic three-dimensional assets and large-scale environments is critical for fields such as robotics, autonomous driving, gaming, and digital simulation. However, existing 3D generative techniques struggle to scale. Methods relying on 2D image models require lengthy per-asset optimization and frequently introduce spatial errors, while direct 3D models are typically restricted to low resolutions or simple single objects due to heavy memory and computational constraints.

The article introduces and evaluates XCube, a 3D generative system designed to rapidly produce high-resolution 3D objects and expansive outdoor driving scenes enriched with structural and semantic details.

To achieve this, the approach generates 3D content across a hierarchy of sparse voxel grids—3D pixel representations that only store data where geometry actually exists. The method uses a cascaded latent diffusion process, which first generates a rough coarse shape and progressively synthesizes finer details at higher resolutions. The system is built on an optimized sparse computing framework that processes high-resolution volumetric grids directly on modern graphics hardware without requiring lengthy test-time optimization. Evaluation was conducted across standard object benchmarks (ShapeNet and Objaverse) and large real-world and synthetic driving scenes (Waymo Open Dataset and Karton City).

The evaluation yielded several key findings. First, the method scales to an effective resolution of 1024-cubed, generating millions of voxels for complex 3D objects in under 30 seconds and full 100-meter by 100-meter outdoor environments with 10-centimeter detail. Second, the framework runs approximately three times faster while consuming half the memory of leading sparse 3D processing alternatives. Third, on standard object generation benchmarks, the system consistently outperformed existing point, mesh, and dense-voxel baselines across shape similarity metrics. Fourth, user studies confirmed substantial quality improvements: 79.2% of evaluators preferred its text-to-3D object generations over a leading baseline, and 66.3% rated its synthetic urban driving scenes as more realistic than actual validation sensor data. Finally, the model proved flexible across practical tasks, including text-to-3D creation, intuitive multi-scale voxel editing, and complete 3D scene reconstruction from a single sparse lidar scan.

These findings demonstrate that direct 3D generative modeling can achieve high spatial resolution and geometric complexity without prohibitive computation times. For industrial applications such as autonomous driving simulation and 3D asset creation pipelines, this significantly lowers the computational cost and turnaround time required to build realistic virtual environments, reducing operational bottlenecks while enhancing consistency.

Organizations developing simulation pipelines or 3D digital content should consider adopting hierarchical sparse representations to scale their generative workflows. Technical teams can integrate coarse-to-fine editing tools into artist workflows and test single-scan completion models to accelerate autonomous vehicle simulation environments. Future initiatives should explore conditioning the architecture on standard 2D reference images and deploying the underlying generative prior across broader downstream perception tasks.

The primary limitation noted in the article is the comparative scarcity of comprehensive 3D training data relative to massive 2D web datasets, which restricts the model's ability to interpret complex, highly nuanced text prompts. Nonetheless, the findings provide a high degree of confidence that sparse hierarchical structures offer an efficient, scalable foundation for direct 3D content generation.

arXiv: 2312.03806
Cover for XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies

Abstract

We present XCube (abbreviated as X3), a novel generative model for high-resolution sparse 3D voxel grids with arbitrary attributes. Our model can generate millions of voxels with a finest effective resolution of up to 10243 in a feed-forward fashion without time-consuming test-time optimization. To achieve this, we employ a hierarchical voxel latent diffusion model which generates progressively higher resolution grids in a coarse-to-fine manner using a custom framework built on the highly efficient VDB data structure. Apart from generating high-resolution objects, we demonstrate the effectiveness of XCube on large outdoor scenes at scales of 100 m×100 m with a voxel size as small as 10 cm. We observe clear qualitative and quantitative improvements over past approaches. In addition to unconditional generation, we show that our model can be used to solve a variety of tasks such as user-guided editing, scene completion from a single scan, and text-to-3D. More results and details can be found on our project webpage.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Sparse Structure VAE
  • 3.2. Hierarchical Voxel Latent Diffusion
  • 3.3. Training and Sampling
  • 3.4. Implementation Details
  • 4. Experiments
  • 4.1. Object-level 3D Generation on ShapeNet
  • 4.2. Object-level 3D Generation on Objaverse
  • 4.3. Large-scale Scene-level 3D Generation
  • 4.4. Ablation Study
  • 5. Discussion
  • References

Knowls

  1. Knowl 1 — Hierarchical Voxel Latent Diffusion Probabilistic Model

    model/method

    XCube models high-resolution 3D shapes and large-scale scenes as a sparse voxel hierarchy comprising LL levels of coarse-to-fine sparse voxel grids G={G1,…,GL}\mathcal{G} = \{G_1, \dots, G_L\} and corresponding per-voxel attributes A={A1,…,AL}\mathcal{A} = \{A_1, \dots, A_L\} (such as surface normals, semantic class labels, or truncated signed distance fields (TSDF)). Each finer grid Gl+1G_{l+1} is spatially contained within the coarser grid GlG_l.

    The joint distribution of the hierarchy (G,A)(\mathcal{G}, \mathcal{A}) and its latent representation X={X1,…,XL}\mathcal{X} = \{X_1, \dots, X_L\} is factorized under a Markovian assumption across hierarchical levels:

    p(G,A,X)=∏l=1Lpψl(Gl,Al∣Xl) pθl(Xl∣Cl−1)p(\mathcal{G}, \mathcal{A}, \mathcal{X}) = \prod_{l=1}^L p_{\psi_l}(G_l, A_l \mid X_l) \, p_{\theta_l}(X_l \mid C_{l-1})

    where pψlp_{\psi_l} is a sparse structure Variational Autoencoder (VAE) decoder parameterized by ψl\psi_l, and pθlp_{\theta_l} is a latent diffusion model parameterized by θl\theta_l. The conditioning context Cl−1C_{l-1} from the coarser level is defined as:

    Cl={c,l=0{Gl,Al,c},l>0C_l = \begin{cases} c, & l = 0 \\ \{G_l, A_l, c\}, & l > 0 \end{cases}

    where cc is an optional global condition (e.g., text prompt or category label). Each latent grid XlX_l is defined on the coarser spatial resolution matching Gl−1G_{l-1}.

  2. Knowl 2 — Sparse Structure Variational Autoencoder

    model/method

    For each hierarchical level ll, a sparse structure Variational Autoencoder (VAE) maps a sparse voxel grid GlG_l and its attributes AlA_l into a continuous latent grid representation XlX_l at the spatial resolution of the preceding coarser level Gl−1G_{l-1}.

    The encoder models the posterior qϕ(Xl∣Gl,Al)q_\phi(X_l \mid G_l, A_l) using alternating 3D sparse convolutions and max pooling operations. The decoder models the likelihood pψ(Gl,Al∣Xl)p_\psi(G_l, A_l \mid X_l) by progressively upsampling from XlX_l to the resolution of GlG_l. In each upsampling stage, existing sparse voxels are subdivided into octants and excessive voxels are pruned based on predicted subdivision occupancy masks.

    The training loss for the level-ll VAE is:

    LVAEl=E{Gl,Al}[EXl∼qϕ[BCE(Gl,G~l)+LAttrl(Al,A~l)]+λDKL(qϕ(Xl∣Gl,Al)∥p(Xl))]\mathcal{L}_{\text{VAE}}^l = \mathbb{E}_{\{G_l, A_l\}}\left[\mathbb{E}_{X_l \sim q_\phi}\left[\text{BCE}(G_l, \tilde{G}_l) + \mathcal{L}_{\text{Attr}}^l(A_l, \tilde{A}_l)\right] + \lambda D_{\text{KL}}\left(q_\phi(X_l \mid G_l, A_l) \parallel p(X_l)\right)\right]

    where G~l,A~l\tilde{G}_l, \tilde{A}_l are the decoded grid structure and predicted attributes, BCE(⋅)\text{BCE}(\cdot) denotes binary cross-entropy on sparse voxel occupancy, LAttrl\mathcal{L}_{\text{Attr}}^l supervises attribute reconstruction, p(Xl)=N(0,I)p(X_l) = \mathcal{N}(0, I) is a standard Gaussian prior, and λ\lambda is the KL divergence regularization weight.

  3. Knowl 3 — Voxel Latent Diffusion Objective with v-Parameterization

    equation

    The latent diffusion model pθl(Xl∣Cl−1)p_{\theta_l}(X_l \mid C_{l-1}) at level ll is trained to predict the velocity vector vv following the vv-parameterization scheme:

    LDMl=Et,Xl,ϵ∼N(0,I)[∥vθl(Xl,t,t)−vref∥22]\mathcal{L}_{\text{DM}}^l = \mathbb{E}_{t, X_l, \epsilon \sim \mathcal{N}(0, I)}\left[\left\| v_{\theta_l}(X_{l,t}, t) - v_{\text{ref}} \right\|_2^2\right]

    where:

    vref=αˉtϵ−1−αˉtXlv_{\text{ref}} = \sqrt{\bar{\alpha}_t}\epsilon - \sqrt{1 - \bar{\alpha}_t}X_l

    with discrete timestep t∼[1,T]t \sim [1, T], Gaussian noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I), latent sample Xl∼qϕ(Xl∣Gl,Al)X_l \sim q_\phi(X_l \mid G_l, A_l), noisy latent Xl,t=αˉtXl+1−αˉtϵX_{l,t} = \sqrt{\bar{\alpha}_t}X_l + \sqrt{1 - \bar{\alpha}_t}\epsilon, noise variance schedule αt=1−βt\alpha_t = 1 - \beta_t, and cumulative product αˉt=∏s=0tαs\bar{\alpha}_t = \prod_{s=0}^t \alpha_s.

    The network vθlv_{\theta_l} is implemented as a 3D sparse convolutional network matching the input grid structure. Conditioning on coarser features Al−1A_{l-1} is performed by direct feature concatenation (since XlX_l shares the grid structure of Gl−1G_{l-1}). Timestep conditioning tt is injected via Adaptive Group Normalization (AdaGN), and textual condition cc is encoded via CLIP and injected via cross-attention.

  4. Knowl 4 — Hierarchical Coarse-to-Fine Sparse Voxel Sampling

    algorithm

    Sampling generates a full sparse voxel hierarchy progressively from the coarsest level l=1l=1 to the finest level l=Ll=L using Denoising Diffusion Implicit Models (DDIM).

    Input: Number of hierarchy levels LL, global conditioning cc, diffusion models {pθl}l=1L\{p_{\theta_l}\}_{l=1}^L, VAE decoders {pψl}l=1L\{p_{\psi_l}\}_{l=1}^L, optional refinement networks {Rl}l=1L\{R_l\}_{l=1}^L.
    Output: Final fine-scale sparse voxel grid GLG_L and associated attributes ALA_L.
    Set conditioning context C0=cC_0 = c
    for l=1l = 1 to LL do
        Sample latent Xl∼pθl(Xl∣Cl−1)X_l \sim p_{\theta_l}(X_l \mid C_{l-1}) using DDIM reverse diffusion
        Decode geometry and attributes (G~l,A~l)=pψl(Xl)(\tilde{G}_l, \tilde{A}_l) = p_{\psi_l}(X_l)
        if refinement network RlR_l is enabled then
            (Gl,Al)=Rl(G~l,A~l)(G_l, A_l) = R_l(\tilde{G}_l, \tilde{A}_l)
        else
            (Gl,Al)=(G~l,A~l)(G_l, A_l) = (\tilde{G}_l, \tilde{A}_l)
        end if
        Update condition Cl={Gl,Al,c}C_l = \{G_l, A_l, c\}
    end for
    Extract final surface mesh from TSDF in ALA_L
    return GL,ALG_L, A_L
  5. Knowl 5 — VDB Deep Learning Engine and Error Mitigation in Sparse Hierarchies

    model/method

    XCube incorporates specific architectural and data structure implementations to enable high-resolution sparse voxel generative modeling:

    1. GPU-Accelerated VDB Framework: Sparse voxel grids are stored using the VDB tree structure, requiring 11 MB of memory for 3.4 million active voxels. Custom GPU sparse 3D convolution and pooling operators operate directly on VDB grids, processing a 102431024^3 resolution scene in milliseconds with ∼3×\sim 3\times faster execution and ∼0.5×\sim 0.5\times the memory footprint of TorchSparse.
    2. Early Dilation: In network layers operating at larger voxel sizes, active sparse voxels are dilated by 1 voxel. This populates halo regions with non-zero features, providing surrounding spatial context to subsequent layers.
    3. Refinement Networks: To prevent cascading error accumulation across hierarchical levels, a lightweight refinement network is applied to the output of each VAE decoder (G~l,A~l)(\tilde{G}_l, \tilde{A}_l). Refinement networks are trained using input data generated by decoding noise-perturbed posterior latent samples from the VAE.
  6. Knowl 6 — ShapeNet 3D Generation Performance Benchmark

    data/table

    XCube was evaluated on unconditional 3D shape generation on the ShapeNet benchmark across three standard categories: Airplane (4,145 shapes), Chair (6,778 shapes), and Car (7,496 shapes) voxelized at 5123512^3 resolution. Evaluation uses the 1-Nearest Neighbor Accuracy (1-NNA) metric computed with Chamfer Distance (CD) and Earth Mover's Distance (EMD), where lower percentages denote closer distributional alignment between generated and validation shapes (ideal is 50%).

    Airplane Chair Car
    Method CD (%) EMD (%) CD (%) EMD (%) CD (%) EMD (%)
    Point-based
    PVD 69.55 60.89 57.68 54.95 64.89 54.61
    LION 65.10 60.15 56.72 54.28 60.61 54.94
    Triplane-based
    NFD 57.55 53.47 54.87 54.06 69.49 71.96
    Dense voxel-based
    NWD 59.78 53.84 56.35 57.98 61.75 58.54
    LAS-Diffusion 71.29 56.93 55.17 55.02 75.03 72.10
    3DShape2VecSet 62.75 61.01 54.06 56.79 86.85 80.91
    XCube (Ours) 52.85 49.75 53.99 48.60 57.96 54.43

    XCube achieves state-of-the-art 1-NNA across all evaluated categories, outperforming point-based, triplane-based, and dense voxel-based baselines.

  7. Knowl 7 — Large-Scale Driving Scene Synthesis and Single-Scan Completion

    empirical result

    XCube scales to large outdoor driving environments by modeling 102.4 m×102.4 m102.4\text{ m} \times 102.4\text{ m} spatial chunks at an effective voxel resolution of 102431024^3 (10 cm10\text{ cm} per voxel) on the Waymo Open Dataset (1,000 LiDAR sequence driving scenarios) and Karton City (20 synthetic city blocks cropped into 900 training and 100 validation scenes).

    In an Amazon Mechanical Turk user study presenting 30 human evaluators with 30 pairwise comparisons (900 total comparisons) between validation set scenes and XCube unconditional generations, evaluators ranked XCube outputs as more realistic than the ground-truth data in 66.3% of comparisons.

    When conditioned on a single, unannotated sparse LiDAR scan, XCube performs 3D scene completion, reconstructing dense geometry alongside full semantic annotations and surface normals.

  8. Knowl 8 — Text-to-3D and Category-Conditional Generation on Objaverse

    empirical result

    XCube was trained on Objaverse (approximately 800,000 3D models) using text captions from Cap3D and on the LVIS subset (approximately 40,000 objects) for category-conditional generation, voxelized at 5123512^3 resolution.

    In a blinded user study on Amazon Mechanical Turk comparing text-to-3D geometry against Shap-E across 30 text prompts (30 users per prompt, 900 total pairwise comparisons), participants voted for XCube's untextured geometry over Shap-E's untextured geometry in 79.2% of comparisons due to higher geometric fidelity and detail alignment with the text prompts.

    Feed-forward shape synthesis with XCube requires 30 seconds for geometry generation; combined with off-the-shelf texture synthesis (30 seconds), complete textured 3D assets are produced in approximately 1 minute.

  9. Knowl 9 — Ablation of Progressive Pruning and Hierarchy Configurations

    data/table

    Ablation studies on ShapeNet Chairs evaluate the impact of progressive subdivision/pruning and varying hierarchy depths/resolutions.

    1. Progressive Pruning: Replacing progressive multi-stage pruning with a single-step pruning stage for a 163→128316^3 \to 128^3 VAE reduces reconstruction accuracy (grid Intersection over Union) from 92.88%92.88\% to 89.68%89.68\% and increases GPU memory consumption by a factor of 3×3\times.
    2. Hierarchy Depth and Resolution: Evaluating different hierarchy depth configurations via 1-NNA on ShapeNet Chairs shows that hierarchical models consistently outperform single-level models (163→512316^3 \to 512^3), while two-level and three-level hierarchies achieve comparable generation quality:
    Model Hierarchy CD (%) EMD (%)
    163→512316^3 \to 512^3 (Single-level) 59.31 57.46
    163→1283→512316^3 \to 128^3 \to 512^3 53.99 48.60
    323→1283→512332^3 \to 128^3 \to 512^3 55.39 51.40
    43→163→1283→51234^3 \to 16^3 \to 128^3 \to 512^3 52.88 53.62
  10. Knowl 10 — Limitations of XCube

    limitation

    XCube exhibits two primary limitations:

    1. Dataset Constraints in Text-to-3D: Because existing 3D datasets are substantially smaller in scale and conceptual coverage than web-scale 2D image corpora (e.g., LAION-5B), the text-conditional diffusion models struggle to synthesize accurate shapes when provided with complex, highly compositional text prompts.
    2. Lack of Direct Image Conditioning: The model framework is currently developed for unconditional, text-conditioned, category-conditioned, and single-scan LiDAR conditional generation, and lacks native conditioning mechanisms for single-view or multi-view 2D images.

Coverage note — None was omitted; all key architectural components, mathematical models, algorithms, benchmarks, ablations, applications, and limitations are fully covered.

References

  1. 1.3d karton city model. https://www.turbosquid.com / 3d - models / 3d - karton - city - 2 - model - 1196110, 2023. Accessed: 2023-08-01. 2, 5, 7, 8
  2. 2.Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas J. Guibas. Learning representations and generative models for 3d point clouds. In International Conference on Machine Learning (ICML), pages 40–49, 2018. 2
  3. 3.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 22563–22575, 2023. 2, 4
  4. 4.Tianshi Cao, Karsten Kreis, Sanja Fidler, Nicholas Sharp, and Kangxue Yin. Texfusion: Synthesizing 3d textures with text-guided image diffusion models. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2023. 6
  5. 5.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015. 1, 2, 5
  6. 6.Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. In European Conference on Computer Vision (ECCV), pages 333–350, 2022. 2
  7. 7.Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5939–5948, 2019. 2
  8. 8.Zhaoxi Chen, Guangcong Wang, and Ziwei Liu. Scenedreamer: Unbounded 3d scene generation from 2d image collections. arXiv preprint arXiv:2302.01330, 2023. 2
  9. 9.Guillaume Chereau. Goxel: 3d voxel editor, 2023. Accessed: 2023-11-16. 6
  10. 10.Christopher B. Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction. In European Conference on Computer Vision (ECCV), pages 628–644, 2016. 5
  11. 11.Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, et al. Objaverse-xl: A universe of 10m+ 3d objects. arXiv preprint arXiv:2307.05663, 2023. 2
  12. 12.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 13142–13153, 2023. 1, 2, 5, 6, 7
  13. 13.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. In Advances in Neural Information Processing Systems, pages 8780–8794, 2021. 4
  14. 14.Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, and Sanja Fidler. Get3d: A generative model of high quality 3d textured shapes learned from images. In Advances in Neural Information Processing Systems, 2022. 2
  15. 15.Lin Gao, Jie Yang, Tong Wu, Yu-Jie Yuan, Hongbo Fu, Yu-Kun Lai, and Hao Zhang. SDM-NET: deep generative network for structured deformable mesh. ACM Transactions on Graphics (TOG), 38(6):243:1–243:15, 2019. 2
  16. 16.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in neural information processing systems, 27, 2014. 2
  17. 17.Anchit Gupta, Wenhan Xiong, Yixin Nie, Ian Jones, and Barlas Oguz. 3dgen: Triplane latent diffusion for textured mesh generation. arXiv preprint arXiv:2303.05371, 2023. 1, 3
  18. 18.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, 2020. 2, 4
  19. 19.Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research, 2022. 4, 5
  20. 20.Jiahui Huang, Zan Gojcic, Matan Atzmon, Or Litany, Sanja Fidler, and Francis Williams. Neural kernel surface reconstruction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4369–4379, 2023. 3, 5
  21. 21.Ka-Hei Hui, Ruihui Li, Jingyu Hu, and Chi-Wing Fu. Neural wavelet-domain diffusion for 3d shape generation. In SIGGRAPH Asia, 2022. 1, 2, 5
  22. 22.Moritz Ibing, Gregor Kobsik, and Leif Kobbelt. Octree transformer: Autoregressive 3d shape generation on hierarchically structured sequences. arXiv preprint arXiv:2111.12480, 2021. 2
  23. 23.Heewoo Jun and Alex Nichol. Shap-e: Generating conditional 3d implicit functions. arXiv preprint arXiv:2305.02463, 2023. 2, 6, 7
  24. 24.Seung Wook Kim, Bradley Brown, Kangxue Yin, Karsten Kreis, Katja Schwarz, Daiqing Li, Robin Rombach, Antonio Torralba, and Sanja Fidler. Neuralfield-ldm: Scene generation with hierarchical latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 8496–8506, 2023. 2
  25. 25.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. In ICLR, 2014. 2
  26. 26.Jumin Lee, Woobin Im, Sebin Lee, and Sung-Eui Yoon. Diffusion probabilistic models for scene-scale 3d categorical data. arXiv preprint arXiv:2301.00527, 2023. 1
  27. 27.Muheng Li, Yueqi Duan, Jie Zhou, and Jiwen Lu. Diffusion-sdf: Text-to-shape via voxelized diffusion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 12642–12651, 2023. 2
  28. 28.Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 300–309, 2023. 2
  29. 29.Chieh Hubert Lin, Hsin-Ying Lee, Willi Menapace, Menglei Chai, Aliaksandr Siarohin, Ming-Hsuan Yang, and Sergey Tulyakov. Infinicity: Infinite-scale city synthesis. arXiv preprint arXiv:2301.09637, 2023. 2
  30. 30.Minghua Liu, Ruoxi Shi, Linghao Chen, Zhuoyang Zhang, Chao Xu, Xinyue Wei, Hansheng Chen, Chong Zeng, Jiayuan Gu, and Hao Su. One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion. arXiv preprint arXiv:2311.07885, 2023. 1
  31. 31.Minghua Liu, Chao Xu, Haian Jin, Linghao Chen, Zexiang Xu, Hao Su, et al. One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization. arXiv preprint arXiv:2306.16928, 2023. 2
  32. 32.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. arXiv preprint arXiv:2303.11328, 2023. 1, 2
  33. 33.Andrew Luo, Tianqin Li, Wen-Hao Zhang, and Tai Sing Lee. Surfgen: Adversarial 3d shape synthesis with explicit surface discriminators. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 16218–16228, 2021. 2
  34. 34.Shitong Luo and Wei Hu. Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2837–2845, 2021. 2, 5
  35. 35.Tiange Luo, Chris Rockwell, Honglak Lee, and Justin Johnson. Scalable 3d captioning with pretrained models. In Advances in Neural Information Processing Systems, 2023. 6
  36. 36.Ken Museth. VDB: high-resolution sparse volumes with dynamic topology. ACM Transactions on Graphics (TOG), 32(3):27:1–27:22, 2013. 2, 5
  37. 37.George Kiyohiro Nakayama, Mikaela Angelina Uy, Jiahui Huang, Shi-Min Hu, Ke Li, and Leonidas Guibas. Difffacto: Controllable part-based 3d point cloud generation with cross diffusion. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2023. 2
  38. 38.Gimin Nam, Mariem Khlifi, Andrew Rodriguez, Alberto Tono, Linqi Zhou, and Paul Guerrero. 3d-ldm: Neural implicit 3d shape generation with latent diffusion models. arXiv preprint arXiv:2212.00842, 2022. 2, 3
  39. 39.Charlie Nash, Yaroslav Ganin, S. M. Ali Eslami, and Peter W. Battaglia. Polygen: An autoregressive generative model of 3d meshes. In International Conference on Machine Learning (ICML), pages 7220–7229, 2020. 2
  40. 40.Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generating 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 2, 6
  41. 41.Evangelos Ntavelis, Aliaksandr Siarohin, Kyle Olszewski, Chaoyang Wang, Luc Van Gool, and Sergey Tulyakov. Autodecoding latent 3d diffusion models. In Advances in Neural Information Processing Systems, 2023. 2, 6
  42. 42.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 2
  43. 43.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In ICLR, 2023. 1
  44. 44.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), pages 8748–8763, 2021. 4
  45. 45.Alexander Raistrick, Lahav Lipson, Zeyu Ma, Lingjie Mei, Mingzhe Wang, Yiming Zuo, Karhan Kayan, Hongyu Wen, Beining Han, Yihan Wang, Alejandro Newell, Hei Law, Ankit Goyal, Kaiyu Yang, and Jia Deng. Infinite photorealistic worlds using procedural generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 12630–12641, 2023. 2
  46. 46.Danilo Rezende and Shakir Mohamed. Variational inference with normalizing flows. In International Conference on Machine Learning (ICML), pages 1530–1538, 2015. 2
  47. 47.Elad Richardson, Gal Metzer, Yuval Alaluf, Raja Giryes, and Daniel Cohen-Or. Texture: Text-guided texturing of 3d shapes. In SIGGRAPH, 2023. 6, 7
  48. 48.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 10674–10685, 2022. 2, 3, 4
  49. 49.Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In ICLR, 2022. 4
  50. 50.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. LAION-5B: an open large-scale dataset for training next generation image-text models. In Advances in Neural Information Processing Systems, 2022. 8
  51. 51.Etai Sella, Gal Fiebelman, Peter Hedman, and Hadar Averbuch-Elor. Vox-e: Text-guided voxel editing of 3d objects. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2023. 2
  52. 52.J. Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner, Jiajun Wu, and Gordon Wetzstein. 3d neural field generation using triplane diffusion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 20875–20886, 2023. 2, 5
  53. 53.Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning (ICML), pages 2256–2265, 2015. 2
  54. 54.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2021. 4
  55. 55.Chunyi Sun, Junlin Han, Weijian Deng, Xinlong Wang, Zishan Qin, and Stephen Gould. 3d-gpt: Procedural 3d modeling with large language models. arXiv preprint arXiv:2310.12945, 2023. 2
  56. 56.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han, Jiquan Ngiam, Hang Zhao, Aleksei Timofeev, Scott Ettinger, Maxim Krivokon, Amy Gao, Aditya Joshi, Yu Zhang, Jonathon Shlens, Zhifeng Chen, and Dragomir Anguelov. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2443–2451, 2020. 2, 5, 6, 8
  57. 57.Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Viewset diffusion:(0-) image-conditioned 3d generative models from 2d data. arXiv preprint arXiv:2306.07881, 2023. 2
  58. 58.Haotian Tang, Zhijian Liu, Xiuyu Li, Yujun Lin, and Song Han. Torchsparse: Efficient point cloud inference engine. MLSys, 2022. 2, 5
  59. 59.Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. Dreamgaussian: Generative gaussian splatting for efficient 3d content creation. arXiv preprint arXiv:2309.16653, 2023. 6
  60. 60.Jia-Heng Tang, Weikai Chen, Jie Yang, Bo Wang, Songrun Liu, Bo Yang, and Lin Gao. Octfield: Hierarchical implicit functions for 3d modeling. arXiv preprint arXiv:2111.01067, 2021. 2
  61. 61.Maxim Tatarchenko, Alexey Dosovitskiy, and Thomas Brox. Octree generating networks: Efficient convolutional architectures for high-resolution 3d outputs. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 2107–2115, 2017. 2
  62. 62.Arash Vahdat, Karsten Kreis, and Jan Kautz. Score-based generative modeling in latent space. In Advances in Neural Information Processing Systems, pages 11287–11302, 2021. 2, 3
  63. 63.Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al. Conditional image generation with pixelcnn decoders. In Advances in Neural Information Processing Systems, pages 4790–4798, 2016. 2
  64. 64.Peng-Shuai Wang, Yang Liu, and Xin Tong. Dual octree graph networks for learning adaptive volumetric shape representations. ACM Transactions on Graphics (TOG), 41(4): 1–15, 2022. 2
  65. 65.Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation. arXiv preprint arXiv:2305.16213, 2023. 1, 2
  66. 66.Jiajun Wu, Chengkai Zhang, Tianfan Xue, William T. Freeman, and Joshua B. Tenenbaum. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. In Advances in Neural Information Processing Systems, pages 82–90, 2016. 2
  67. 67.Haozhe Xie, Zhaoxi Chen, Fangzhou Hong, and Ziwei Liu. Citydreamer: Compositional generative model of unbounded 3d cities. arXiv preprint arXiv:2309.00610, 2023. 2
  68. 68.Guandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu, Serge J. Belongie, and Bharath Hariharan. Pointflow: 3d point cloud generation with continuous normalizing flows. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 4540–4549, 2019. 5
  69. 69.Xiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, and Karsten Kreis. LION: latent point diffusion models for 3d shape generation. In Advances in Neural Information Processing Systems, 2022. 2, 3, 5
  70. 70.Biao Zhang, Jiapeng Tang, Matthias Niessner, and Peter Wonka. 3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models. ACM Transactions on Graphics (TOG), 42(4):92:1–92:16, 2023. 5
  71. 71.Dongsu Zhang, Changwoon Choi, Jeonghwan Kim, and Young Min Kim. Learning to generate 3d shapes with generative cellular automata. In ICLR, 2021. 2
  72. 72.Xin-Yang Zheng, Hao Pan, Peng-Shuai Wang, Xin Tong, Yang Liu, and Heung-Yeung Shum. Locally attentional SDF diffusion for controllable 3d shape generation. ACM Transactions on Graphics (TOG), 42(4):91:1–91:13, 2023. 2, 5
  73. 73.Linqi Zhou, Yilun Du, and Jiajun Wu. 3d shape generation and completion through point-voxel diffusion. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 5806–5815, 2021. 5

Citation

MLA
Ren, X., et al. “XCube: Large-Scale 3D Generative Modeling Using Sparse Voxel Hierarchies”. arXiv, 2023, http://arxiv.org/abs/2312.03806v2.
APA
Ren, X., Huang, J., Zeng, X., Museth, K., Fidler, S., & Williams, F. (2023). XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies. arXiv. http://arxiv.org/abs/2312.03806v2
Chicago
Ren, X., J. Huang, X. Zeng, K. Museth, S. Fidler, and F. Williams. 2023. “XCube: Large-Scale 3D Generative Modeling Using Sparse Voxel Hierarchies”. arXiv. http://arxiv.org/abs/2312.03806v2.
Harvard
Ren, X. et al. (2023) “XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.03806v2.
Vancouver
1. Ren X, Huang J, Zeng X, Museth K, Fidler S, Williams F (2023) XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies. arXiv

BibTeX

@article{ren2023xcube,
  title = {XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies},
  author = {Ren, Xuanchi and Huang, Jiahui and Zeng, Xiaohui and Museth, Ken and Fidler, Sanja and Williams, Francis},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.03806v2},
  eprint = {2312.03806}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE