Diffusion-SDF: Text-to-Shape via Voxelized Diffusion
Muheng LiYueqi DuanJie ZhouJiwen Lu
Proposes a two-stage 3D generative framework that combines a patch-based signed distance field autoencoder with a voxelized diffusion model to synthesize, complete, and manipulate detailed 3D shapes from text descriptions.
Generating high-quality 3D digital content from natural language descriptions is an increasingly important capability for virtual modeling, design, and simulation. However, existing automated methods struggle to produce diverse 3D shapes that accurately match text prompts while maintaining watertight and precise structural geometries.
The article develops and evaluates Diffusion-SDF, a generative modeling framework designed for text-to-shape synthesis. The main objective is to demonstrate that combining implicit 3D representations with diffusion models can produce higher-quality and more varied 3D shapes from text descriptions than existing approaches.
The authors implemented a two-stage generative pipeline. In the first stage, a patch-wise autoencoder compresses 3D shapes—represented as voxelized truncated signed distance fields—into localized, independent latent representations across 13 object categories from the ShapeNet repository. In the second stage, a voxelized diffusion model equipped with a customized dual-network architecture generates shape representations guided by text embeddings. The model was trained and evaluated using the Text2Shape benchmark, which contains approximately 75,000 text-shape pairs.
The findings show that Diffusion-SDF substantially outperforms prior state-of-the-art methods across multiple performance dimensions. First, the framework achieved a classification accuracy of 88.56% on generated shapes, outperforming the closest baseline by about 5 percentage points and more than doubling older methods. Second, it delivered an intersection-over-union fidelity score of 0.194 and the highest text-alignment similarity score among evaluated models. Third, the system drastically improved generation diversity, reducing the total mutual difference score to 0.169 compared to 0.581 for the leading baseline. Finally, qualitative assessments confirmed the method successfully performs complex downstream tasks, including text-guided shape completion of missing parts and localized shape manipulation.
These results indicate that combining patch-independent implicit representations with diffusion models resolves major quality and diversity bottlenecks in automated 3D asset generation. For commercial workflows, this capability can reduce the time, manual labor, and production costs required for early-stage 3D content creation, making text-driven design accessible to non-specialists.
Organizations exploring automated 3D modeling should consider piloting patch-based diffusion frameworks for iterative asset drafting and editing. Before deploying these models into broader production pipelines, developers should expand training datasets to cover a wider range of object categories beyond common furniture types and investigate zero-shot synthesis leveraging pre-trained vision-language models.
Confidence in the reported benchmarks is supported by quantitative comparisons and ablation studies. However, practical application is currently bounded by the limited scope of available paired text-shape datasets, which were restricted primarily to chairs and tables, as well as resolution constraints imposed by voxel patch boundaries.
- Paper: AutoSDF: Shape Priors for 3D Completion, Reconstruction and Generation, Paritosh Mittal et al. (2022). AutoSDF introduced patch-wise encoding of truncated signed distance fields for generic 3D shape priors and completion, establishing the localized implicit representation adapted by Diffusion-SDF.
- Paper: LION: Latent Point Diffusion Models for 3D Shape Generation, Xiaohui Zeng et al. (2022). LION establishes hierarchical latent diffusion modeling for generative 3D shape synthesis on ShapeNet, providing foundational methodology for generative 3D architectures.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This seminal work establishes the mathematical formulation and denoising objectives of diffusion probabilistic models that underpin the voxelized generative pipeline of Diffusion-SDF.
- Paper: Magic3D: High-Resolution Text-to-3D Content Creation, Chen-Hsuan Lin et al. (2022). Magic3D details high-resolution text-to-3D asset creation and multi-stage volumetric optimization, serving as critical context and a baseline for text-guided 3D modeling.
- Paper: GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models, Alexander Quinn Nichol et al. (2022). GLIDE develops text-guided diffusion models with classifier-free guidance, establishing core conditioning principles adapted for text-to-3D generation.
- Paper: XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies, Xuanchi Ren et al. (2024). XCube scales direct 3D volumetric generative diffusion beyond dense localized voxel grids by employing sparse voxel hierarchies to synthesize full scenes and objects up to 1024-cubed resolution.
- Paper: Text-to-3D using Gaussian Splatting, Zilong Chen et al. (2024). GSGEN advances text-to-3D generation by replacing voxelized implicit fields with explicit 3D Gaussian Splatting and joint 2D-3D score distillation to overcome structural artifacts.
- Paper: Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors, Guocheng Qian et al. (2024). Magic123 builds on single-modality 3D generative priors by combining both 2D and 3D diffusion models to generate high-resolution textured 3D meshes.
- Paper: MatFuse: Controllable Material Generation with Diffusion Models, Giuseppe Vecchio et al. (2024). MatFuse extends controllable 3D diffusion workflows from raw geometric shape synthesis to photorealistic, multi-attribute material map generation and editing.
- Paper: Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based Rendering, Kim Youwang et al. (2024). Paint-it complements text-driven 3D geometric synthesis by leveraging pre-trained diffusion models to synthesize high-fidelity physically-based rendering texture maps for untextured meshes.
