Diffusion-SDF: Text-to-Shape via Voxelized Diffusion

Muheng LiYueqi DuanJie ZhouJiwen Lu

article2023CVPR168 citations

Proposes a two-stage 3D generative framework that combines a patch-based signed distance field autoencoder with a voxelized diffusion model to synthesize, complete, and manipulate detailed 3D shapes from text descriptions.

Listen

Generating high-quality 3D digital content from natural language descriptions is an increasingly important capability for virtual modeling, design, and simulation. However, existing automated methods struggle to produce diverse 3D shapes that accurately match text prompts while maintaining watertight and precise structural geometries.

The article develops and evaluates Diffusion-SDF, a generative modeling framework designed for text-to-shape synthesis. The main objective is to demonstrate that combining implicit 3D representations with diffusion models can produce higher-quality and more varied 3D shapes from text descriptions than existing approaches.

The authors implemented a two-stage generative pipeline. In the first stage, a patch-wise autoencoder compresses 3D shapes—represented as voxelized truncated signed distance fields—into localized, independent latent representations across 13 object categories from the ShapeNet repository. In the second stage, a voxelized diffusion model equipped with a customized dual-network architecture generates shape representations guided by text embeddings. The model was trained and evaluated using the Text2Shape benchmark, which contains approximately 75,000 text-shape pairs.

The findings show that Diffusion-SDF substantially outperforms prior state-of-the-art methods across multiple performance dimensions. First, the framework achieved a classification accuracy of 88.56% on generated shapes, outperforming the closest baseline by about 5 percentage points and more than doubling older methods. Second, it delivered an intersection-over-union fidelity score of 0.194 and the highest text-alignment similarity score among evaluated models. Third, the system drastically improved generation diversity, reducing the total mutual difference score to 0.169 compared to 0.581 for the leading baseline. Finally, qualitative assessments confirmed the method successfully performs complex downstream tasks, including text-guided shape completion of missing parts and localized shape manipulation.

These results indicate that combining patch-independent implicit representations with diffusion models resolves major quality and diversity bottlenecks in automated 3D asset generation. For commercial workflows, this capability can reduce the time, manual labor, and production costs required for early-stage 3D content creation, making text-driven design accessible to non-specialists.

Organizations exploring automated 3D modeling should consider piloting patch-based diffusion frameworks for iterative asset drafting and editing. Before deploying these models into broader production pipelines, developers should expand training datasets to cover a wider range of object categories beyond common furniture types and investigate zero-shot synthesis leveraging pre-trained vision-language models.

Confidence in the reported benchmarks is supported by quantitative comparisons and ablation studies. However, practical application is currently bounded by the limited scope of available paired text-shape datasets, which were restricted primarily to chairs and tables, as well as resolution constraints imposed by voxel patch boundaries.

Cover for Diffusion-SDF: Text-to-Shape via Voxelized Diffusion

Abstract

With the rising industrial attention to 3D virtual modeling technology, generating novel 3D content based on specified conditions (e.g. text) has become a hot issue. In this paper, we propose a new generative 3D modeling framework called Diffusion-SDF for the challenging task of text-to-shape synthesis. Previous approaches lack flexibility in both 3D data representation and shape generation, thereby failing to generate highly diversified 3D shapes conforming to the given text descriptions. To address this, we propose a SDF autoencoder together with the Voxelized Diffusion model to learn and generate representations for voxelized signed distance fields (SDFs) of 3D shapes. Specifically, we design a novel UinU-Net architecture that implants a local-focused inner network inside the standard U-Net architecture, which enables better reconstruction of patch-independent SDF representations. We extend our approach to further text-to-shape tasks including text-conditioned shape completion and manipulation. Experimental results show that Diffusion-SDF generates both higher quality and more diversified 3D shapes that conform well to given text descriptions when compared to previous approaches. Code is available at: https://github.com/ttilm/Diffusion-SDF.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Autoencoding Signed Distance Fields
  • 3.2. Voxelized Diffusion Models (VDMs)
  • 4. Experiments
  • 4.1. Text-Conditioned Shape Generation
  • 4.2. Ablation Studies
  • 4.3. Text-Guided Shape Completion
  • 4.4. Text-Guided Shape Manipulation
  • 5. Limitations and Conclusion
  • References

Knowls

  1. Knowl 1 — Two-Stage Diffusion-SDF Generative Framework

    model/method

    Diffusion-SDF is a text-conditioned 3D shape synthesis framework that represents shapes via voxelized truncated signed distance fields (TSDFs) through a two-stage decoupled generation pipeline:

    1. Stage I (Patch-Wise SDF Compression): A 3D shape truncated signed distance field x∈RD×D×Dx \in \mathbb{R}^{D \times D \times D} is split into N=(D/P)3N = (D/P)^3 local volumetric patches Xp=[xp1,…,xpN]X_p = [x_{p1}, \dots, x_{pN}] of size P×P×PP \times P \times P. A local shape encoder Eloc\mathcal{E}_{loc} maps each patch independently into a Gaussian latent vector zpn=Eloc(xpn)∈Rcz_{pn} = \mathcal{E}_{loc}(x_{pn}) \in \mathbb{R}^c. The latent representations are repositioned into a 3D latent grid z∈R(D/P)×(D/P)×(D/P)×cz \in \mathbb{R}^{(D/P) \times (D/P) \times (D/P) \times c}. A patch-joint decoder D\mathcal{D} reconstructs the full continuous TSDF grid x~=D(z)\tilde{x} = \mathcal{D}(z). The isosurface of x~\tilde{x} is subsequently extracted as a 3D surface mesh using Marching Cubes.

    2. Stage II (Latent Voxelized Diffusion): A denoising diffusion probabilistic model operates over the compressed 3D latent grid representations zz. An autoencoder-based neural score estimator ϵθ(zt,t,τ(c))\epsilon_\theta(z_t, t, \tau(c))—termed UinU-Net—is conditioned on text embeddings τ(c)\tau(c) extracted from a frozen CLIP text encoder τ\tau, gradually denoising a pure Gaussian noise grid into a valid shape latent grid z0z_0 matching input text cc.

  2. Knowl 2 — Patch-Wise SDF Autoencoder

    model/method

    The patch-wise SDF autoencoder compresses a 3D voxelized truncated signed distance field (TSDF) while enforcing a patch-independent Gaussian latent distribution.

    Given a TSDF grid x∈RD3x \in \mathbb{R}^{D^3} (with grid resolution D=64D=64), the volume is divided into N=(D/P)3N = (D/P)^3 non-overlapping patches Xp=[xp1,…,xpN]X_p = [x_{p1}, \dots, x_{pN}], each of size P3P^3. The local encoder Eloc\mathcal{E}_{loc} encodes each patch independently:

    zpn=Eloc(xpn)∈Rcz_{pn} = \mathcal{E}_{loc}(x_{pn}) \in \mathbb{R}^c

    where cc is the latent channel dimension. The patch embeddings Zp=[zp1,…,zpN]Z_p = [z_{p1}, \dots, z_{pN}] are arranged into a regular 3D latent grid z∈R(D/P)×(D/P)×(D/P)×cz \in \mathbb{R}^{(D/P) \times (D/P) \times (D/P) \times c}. The patch-joint decoder D\mathcal{D} processes this full grid to reconstruct the complete SDF field:

    x~=D(z)\tilde{x} = \mathcal{D}(z)

    The autoencoder is trained using a composite objective consisting of an L1\mathcal{L}_1 reconstruction loss and a Kullback-Leibler (KL) divergence regularization loss:

    LAE=∥x−x~∥1+λKLDKL(q(z∣x)∥N(0,I))\mathcal{L}_{AE} = \|x - \tilde{x}\|_1 + \lambda_{\text{KL}} D_{\text{KL}}(q(z|x) \parallel \mathcal{N}(0, \mathbf{I}))

    The KL penalty enforces that each patch latent embedding is approximately distributed according to an independent standard normal distribution N(0,I)\mathcal{N}(0, \mathbf{I}), providing an isotropic Gaussian prior for subsequent diffusion modeling while preserving patch-to-global spatial continuity in the decoder.

  3. Knowl 3 — UinU-Net Score Estimator Architecture

    model/method

    The UinU-Net is a specialized 3D neural score estimator ϵθ(zt,t,τ(c))\epsilon_\theta(z_t, t, \tau(c)) designed for denoising voxelized latent TSDF grids ztz_t. Standard 3D U-Nets apply hierarchical 3×3×33 \times 3 \times 3 convolutions, which mix global and neighborhood features but struggle to reconstruct patch-independent local Gaussian representations. UinU-Net embeds an inner local network inside an outer 3D U-Net backbone:

    • Outer Network: A standard 3D U-Net backbone featuring 3×3×33 \times 3 \times 3 convolutional residual blocks, cross-attention layers conditioned on text embeddings τ(c)\tau(c), downsampling/upsampling layers, and skip connections across hierarchical resolutions (e.g., 83×1928^3 \times 192, 43×3844^3 \times 384, 23×7682^3 \times 768).
    • Inner Network: Operating directly at the native latent resolution D/PD/P (e.g., 83×88^3 \times 8), it utilizes 1×1×11 \times 1 \times 1 convolutional ResNet blocks to process patch-wise features independently without spatial mixing.
    • Spatial Transformer: A self-attention transformer module within the inner path takes positionally embedded patch representations as tokens to explicitly model non-local patch-to-patch dependencies.
    • Inner-Outer Skip Connections: Synthesis pathways of the inner network are concatenated directly to corresponding stages in the outer U-Net, allowing independent local patch representations to guide global shape formation.
  4. Knowl 4 — Latent Voxelized Diffusion and Classifier-Free Text Guidance

    model/method

    The Voxelized Diffusion Model (VDM) models the distribution of latent TSDF grids z0∼q(z0)z_0 \sim q(z_0) over a Markov chain of TT timesteps. The forward diffusion transition is:

    q(zt∣z0)=N(zt;αˉtz0,(1−αˉt)I)q(z_t | z_0) = \mathcal{N}\left(z_t; \sqrt{\bar{\alpha}_t} z_0, (1 - \bar{\alpha}_t) \mathbf{I}\right)

    where αt=1−βt\alpha_t = 1 - \beta_t and αˉt=∏i=1tαi\bar{\alpha}_t = \prod_{i=1}^t \alpha_i. The training objective for the score estimator ϵθ\epsilon_\theta optimizes the MSE of predicted noise on latent samples:

    LDiffusion-SDF=Ez0,t,ϵ∼N(0,I)[∥ϵ−ϵθ(zt,t,τ(c))∥2]\mathcal{L}_{\text{Diffusion-SDF}} = \mathbb{E}_{z_0, t, \boldsymbol{\epsilon} \sim \mathcal{N}(0, \mathbf{I})}\left[\left\|\boldsymbol{\epsilon} - \boldsymbol{\epsilon}_\theta(z_t, t, \tau(c))\right\|^2\right]

    where t∼U(1,T)t \sim \mathcal{U}(1, T), zt=αˉtz0+1−αˉtϵz_t = \sqrt{\bar{\alpha}_t} z_0 + \sqrt{1 - \bar{\alpha}_t} \boldsymbol{\epsilon}, and τ(c)\tau(c) is the text embedding of prompt cc obtained from a frozen CLIP text encoder.

    During reverse diffusion sampling, classifier-free guidance is applied with guidance scale s≥1s \ge 1:

    ϵ^θ(zt∣τ(c))=ϵθ(zt∣∅)+s⋅(ϵθ(zt∣τ(c))−ϵθ(zt∣∅))\hat{\boldsymbol{\epsilon}}_\theta(z_t | \tau(c)) = \boldsymbol{\epsilon}_\theta(z_t | \emptyset) + s \cdot \left(\boldsymbol{\epsilon}_\theta(z_t | \tau(c)) - \boldsymbol{\epsilon}_\theta(z_t | \emptyset)\right)

    where ϵθ(zt∣∅)\boldsymbol{\epsilon}_\theta(z_t | \emptyset) denotes the unconditional score computed with an empty text prompt. Sampling is accelerated using DDIM, reducing inference from T=1000T=1000 to 50 steps.

  5. Knowl 5 — Text-Guided Shape Completion via Masked Latent Diffusion

    algorithm

    Text-guided shape completion synthesizes missing 3D shape regions from a partial shape input under the guidance of a text prompt cc. The known partial shape is encoded into latent grid z0z_0, and missing regions are represented by a binary mask m∈{0,1}(D/P)3m \in \{0, 1\}^{(D/P)^3}, where m=1m=1 indicates known patches and m=0m=0 indicates unobserved patches.

    Input: Partial shape latent z0z_0, binary patch mask mm, text condition cc, total DDIM steps KK, noise schedule αˉt,βt\bar{\alpha}_t, \beta_t
    Output: Completed shape latent z0z_0
    Initialize ztK∼N(0,I)z_{t_K} \sim \mathcal{N}(0, \mathbf{I})
    for step k=Kk = K down to 11 do
        t=tkt = t_k
        tprev=tk−1t_{prev} = t_{k-1}
        Estimate denoised state z~tprev∼N(μθ(zt,t,τ(c)),βtI)\tilde{z}_{t_{prev}} \sim \mathcal{N}(\boldsymbol{\mu}_\theta(z_t, t, \tau(c)), \beta_t \mathbf{I}) using classifier-free guidance
        Sample forward diffused known patches z^tprev=αˉtprevz0+1−αˉtprevϵ\hat{z}_{t_{prev}} = \sqrt{\bar{\alpha}_{t_{prev}}} z_0 + \sqrt{1 - \bar{\alpha}_{t_{prev}}} \boldsymbol{\epsilon}, where ϵ∼N(0,I)\boldsymbol{\epsilon} \sim \mathcal{N}(0, \mathbf{I})
        Combine known and generated patches: ztprev=(1−m)⊙z~tprev+m⊙z^tprevz_{t_{prev}} = (1 - m) \odot \tilde{z}_{t_{prev}} + m \odot \hat{z}_{t_{prev}}
    end
    return z0z_0

    The final latent grid z0z_0 is decoded into a complete TSDF grid x~=D(z0)\tilde{x} = \mathcal{D}(z_0), from which a mesh is reconstructed via Marching Cubes.

  6. Knowl 6 — Text-Guided Shape Manipulation via Latent Cycle-Sampling

    algorithm

    Text-guided shape manipulation modifies structural or semantic attributes of an existing 3D shape zinitz_{\text{init}} according to an instructive text prompt cc using a forward-reverse cycle-sampling strategy.

    Input: Initial shape latent zinitz_{\text{init}}, target text prompt cc, intermediate step tmid∈(0,T)t_{\text{mid}} \in (0, T), noise schedule αˉ\bar{\alpha}
    Output: Manipulated shape latent zgoalz_{\text{goal}}
    Sample random noise ϵ∼N(0,I)\boldsymbol{\epsilon} \sim \mathcal{N}(0, \mathbf{I})
    Compute intermediate noisy latent: zmid=αˉtmidzinit+1−αˉtmidϵz_{\text{mid}} = \sqrt{\bar{\alpha}_{t_{\text{mid}}}} z_{\text{init}} + \sqrt{1 - \bar{\alpha}_{t_{\text{mid}}}} \boldsymbol{\epsilon}
    Initialize reverse trajectory at step tmidt_{\text{mid}} with ztmid=zmidz_{t_{\text{mid}}} = z_{\text{mid}}
    for step t=tmidt = t_{\text{mid}} down to 11 do
        Compute conditioned score ϵ^θ(zt∣τ(c))\hat{\boldsymbol{\epsilon}}_\theta(z_t | \tau(c)) with classifier-free guidance
        Update zt−1z_{t-1} from ztz_t following standard reverse diffusion step
    end
    zgoal=z0z_{\text{goal}} = z_0
    return zgoalz_{\text{goal}}

    The choice of tmidt_{\text{mid}} governs the trade-off between preservation of original characteristics and fidelity to the prompt: setting tmidt_{\text{mid}} to 600∼800600\sim800 out of 1000 DDPM steps alters high-level attributes (such as leg style or sofa vs. chair) while preserving base shape identity. Excessively small tmidt_{\text{mid}} leaves the shape unchanged, whereas overly large tmidt_{\text{mid}} destructs original features.

  7. Knowl 7 — Quantitative Evaluation on Text-to-Shape Generation Benchmark

    data/table

    Diffusion-SDF was quantitatively compared against state-of-the-art implicit text-to-shape methods on the chair category of the Text2Shape dataset using four metrics:

    • IoU ↑\uparrow: Intersection over Union of occupancy voxel grids downsampled to 32332^3, measuring geometric fidelity to ground truth.
    • Acc ↑\uparrow: Classification accuracy of generated shapes across 13 ShapeNet categories using a pre-trained voxel classifier.
    • CLIP-S ↑\uparrow: Maximum cosine similarity between the text description and 5 rendered 2D views of the generated 3D shape using pre-trained CLIP.
    • TMD ↓\downarrow: Total Mutual Difference, computed as the average IoU among k=10k = 10 samples generated per text query (lower indicates higher generative diversity).
    Methods IoU↑\uparrow Acc↑\uparrow CLIP-S↑\uparrow TMD↓\downarrow
    Liu et al. 0.160 34.79 29.94 0.891
    AutoSDF 0.187 83.88 29.10 0.581
    Diffusion-SDF (Ours) 0.194 88.56 30.88 0.169

    Diffusion-SDF outperforms prior methods across reconstruction fidelity (IoU), category classification accuracy (Acc), and text-shape semantic alignment (CLIP-S), while achieving a significantly lower TMD score (0.169 vs. 0.581 for AutoSDF), indicating substantial improvements in sample diversity.

  8. Knowl 8 — Ablation Studies on Denoiser Architecture Components

    data/table

    Ablations evaluate the contribution of the UinU-Net denoiser design over a standard U-Net backbone and dissect its internal components on the Text2Shape chair split:

    Architecture Variant IoU↑\uparrow Acc↑\uparrow CLIP-S↑\uparrow TMD↓\downarrow
    Standard U-Net 0.187 80.98 29.35 0.171
    UinU-Net (Full) 0.194 88.56 30.88 0.169
    – w/o in-out concat 0.198 83.91 28.68 0.181
    – w/o spatial attention 0.194 85.56 30.82 0.161

    Key findings:

    1. Incorporating the inner network (UinU-Net) improves classification accuracy from 80.98% to 88.56% and CLIP-S from 29.35 to 30.88, demonstrating that local patch modeling helps recover independent patch distributions.
    2. Removing inner-outer skip concatenation causes a major decline in semantic conformance (CLIP-S drops from 30.88 to 28.68; Acc drops from 88.56% to 83.91%).
    3. Removing the spatial self-attention module drops accuracy from 88.56% to 85.56%, indicating that capturing patch-to-patch relational context is critical for preserving global shape semantics.
  9. Knowl 9 — Limitations of Diffusion-SDF

    limitation

    Diffusion-SDF has several documented limitations:

    1. Dataset Scope: Training and evaluation are constrained by existing shape-text paired datasets, which primarily encompass chairs and tables from ShapeNet in Text2Shape, restricting generalized synthesis across diverse open-domain categories.
    2. Patch Boundary Artifacts: Partitioning the TSDF volume into regular patch grids can cause visible discontinuities or cut region artifacts across adjacent patch boundaries in complex geometry. Mitigating this requires higher grid resolutions DD and smaller patch sizes PP, which increases computation.
    3. Intermediate Step Sensitivity in Manipulation: Text-guided manipulation is sensitive to the chosen cutoff step tmidt_{\text{mid}}; improperly calibrated values either fail to alter the initial shape or completely erase its geometric identity.

Coverage note — None omitted; all primary architectural components, formulation steps, downstream task algorithms, experimental results, ablations, and limitations are covered.

References

  1. 1.Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas J. Guibas. Learning representations and generative models for 3D point clouds. In ICML, pages 40–49, 2018. 1, 2, 5
  2. 2.Panos Achlioptas, Leonidas J. Guibas, Noah D. Goodman, Judy Fan, and Robert X. D. Hawkins. Shapeglot: Learning language for shape differentiation. In ICCV, pages 8937–8946, 2019. 6
  3. 3.Rohan Chabra, Jan Eric Lenssen, Eddy Ilg, Tanner Schmidt, Julian Straub, Steven Lovegrove, and Richard A. Newcombe. Deep local shapes: Learning local SDF priors for detailed 3D reconstruction. In ECCV, pages 608–625, 2020. 1, 3
  4. 4.Paramanand Chandramouli and Kanchana Vaishnavi Gandikota. Ldedit: Towards generalized text guided image manipulation via latent diffusion models. abs/2210.02249, 2022. 5
  5. 5.Angel X. Chang, Thomas A. Funkhouser, Leonidas J. Guibas, Pat Hanrahan, Qi-Xing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. ShapeNet: An information-rich 3D model repository. abs/1512.03012, 2015. 5, 6
  6. 6.Kevin Chen, Christopher B. Choy, Manolis Savva, Angel X. Chang, Thomas A. Funkhouser, and Silvio Savarese. Text2Shape: Generating shapes from natural language by learning joint embeddings. In ACCV, pages 100–116, 2018. 2, 5, 6, 8
  7. 7.Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss, Mohammad Norouzi, and William Chan. Wavegrad: Estimating gradients for waveform generation. In ICLR, 2021. 2
  8. 8.Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In CVPR, pages 5939–5948, 2019. 1, 2
  9. 9.Christopher B. Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3D-R2^2N2^2: A unified approach for single and multi-view 3D object reconstruction. In ECCV, pages 628–644, 2016. 1, 2
  10. 10.Özgün Çiçek, Ahmed Abdulkadir, Soeren S. Lienkamp, Thomas Brox, and Olaf Ronneberger. 3D U-Net: Learning dense volumetric segmentation from sparse annotation. In MICCAI, pages 424–432, 2016. 4
  11. 11.Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah. Diffusion models in vision: A survey. 2022. 2
  12. 12.Angela Dai, Charles Ruizhongtai Qi, and Matthias Nießner. Shape completion using 3D-encoder-predictor CNNs and shape synthesis. In CVPR, pages 6545–6554, 2017. 1
  13. 13.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat gans on image synthesis. In NeurIPS, pages 8780–8794, 2021. 2
  14. 14.Yueqi Duan, Haidong Zhu, He Wang, Li Yi, Ram Nevatia, and Leonidas J. Guibas. Curriculum deepsdf. In ECCV, pages 51–67, 2020. 1
  15. 15.Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, and Sanja Fidler. GET3D: A generative model of high quality 3D textured shapes learned from images. 2022. 1, 5
  16. 16.Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial networks. Commun. ACM, 63(11):139–144, 2020. 2
  17. 17.Thibault Groueix, Matthew Fisher, Vladimir G. Kim, Bryan C. Russell, and Mathieu Aubry. A papier-mâché approach to learning 3D surface generation. In CVPR, pages 216–224, 2018. 1
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 4
  19. 19.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020. 2, 3, 4
  20. 20.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 5
  21. 21.Moritz Ibing, Isaak Lim, and Leif Kobbelt. 3D shape generation with grid-based implicit functions. In CVPR, pages 13559–13568, 2021. 1
  22. 22.Chiyu "Max" Jiang, Avneesh Sud, Ameesh Makadia, Jingwei Huang, Matthias Nießner, and Thomas A. Funkhouser. Local implicit grid representations for 3D scenes. In CVPR, pages 6000–6009, 2020. 3
  23. 23.Yue Jiang, Dantong Ji, Zhizhong Han, and Matthias Zwicker. SDFDiff: Differentiable rendering of signed distance fields for 3D shape optimization. In CVPR, pages 1248–1258, 2020. 1
  24. 24.Hiroharu Kato, Yoshitaka Ushiku, and Tatsuya Harada. Neural 3D mesh renderer. In CVPR, pages 3907–3916, 2018. 1
  25. 25.Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. Diffusionclip: Text-guided diffusion models for robust image manipulation. In CVPR, pages 2416–2425, 2022. 5
  26. 26.Diederik P. Kingma and Max Welling. Auto-encoding variational bayes. In ICLR, 2014. 2, 3
  27. 27.Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. Diffwave: A versatile diffusion model for audio synthesis. In ICLR, 2021. 2
  28. 28.Ke Li and Jitendra Malik. Implicit maximum likelihood estimation. 2018. 2
  29. 29.Zhengzhe Liu, Yi Wang, Xiaojuan Qi, and Chi-Wing Fu. Towards implicit text-guided 3D shape generation. In CVPR, pages 17875–17885, 2022. 2, 5, 6
  30. 30.Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. In CVPR, pages 11451–11461, 2022. 5
  31. 31.Shitong Luo and Wei Hu. Diffusion probabilistic models for 3D point cloud generation. In CVPR, pages 2837–2845, 2021. 2
  32. 32.Lars M. Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3D reconstruction in function space. In CVPR, pages 4460–4470, 2019. 1, 2
  33. 33.Oscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim, and Rana Hanocka. Text2mesh: Text-driven neural stylization for meshes. In CVPR, pages 13482–13492, 2022. 2, 5
  34. 34.Paritosh Mittal, Yen-Chi Cheng, Maneesh Singh, and Shubham Tulsiani. Autosdf: Shape priors for 3D completion, reconstruction and generation. In CVPR, pages 306–315, 2022. 1, 2, 5, 6
  35. 35.Charlie Nash, Yaroslav Ganin, S. M. Ali Eslami, and Peter W. Battaglia. Polygen: An autoregressive generative model of 3D meshes. In ICML, pages 7220–7229, 2020. 1
  36. 36.Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In ICML, pages 16784–16804, 2022. 2
  37. 37.Jeong Joon Park, Peter Florence, Julian Straub, Richard A. Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In CVPR, pages 165–174, 2019. 1
  38. 38.Charles Ruizhongtai Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. Pointnet: Deep learning on point sets for 3D classification and segmentation. In CVPR, pages 77–85, 2017. 1
  39. 39.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J. Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In NeurIPS, pages 5099–5108, 2017. 1
  40. 40.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763, 2021. 5, 6
  41. 41.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with CLIP latents. 2022. 2
  42. 42.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, pages 8821–8831, 2021. 2
  43. 43.Danilo Jimenez Rezende and Shakir Mohamed. Variational inference with normalizing flows. In ICML, pages 1530–1538, 2015. 2
  44. 44.Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. Stochastic backpropagation and approximate inference in deep generative models. In ICML, pages 1278–1286, 2014. 3
  45. 45.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, pages 10674–10685, 2022. 2, 5, 7
  46. 46.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional networks for biomedical image segmentation. In MICCAI, pages 234–241, 2015. 2
  47. 47.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. 2022. 2
  48. 48.Aditya Sanghi, Hang Chu, Joseph G. Lambourne, Ye Wang, Chin-Yi Cheng, Marco Fumero, and Kamal Rahimi Malekshan. Clip-forge: Towards zero-shot text-to-shape generation. In CVPR, pages 18582–18592, 2022. 2, 5, 6, 8
  49. 49.Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, pages 2256–2265, 2015. 2, 3
  50. 50.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2021. 5
  51. 51.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, pages 5998–6008, 2017. 4
  52. 52.Jiajun Wu, Chengkai Zhang, Tianfan Xue, Bill Freeman, and Josh Tenenbaum. Learning a probabilistic latent space of object shapes via 3D generative-adversarial modeling. In NeurIPS, pages 82–90, 2016. 2
  53. 53.Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 3D ShapeNets: A deep representation for volumetric shapes. In CVPR, pages 1912–1920, 2015. 1
  54. 54.Qiangeng Xu, Weiyue Wang, Duygu Ceylan, Radomír Mech, and Ulrich Neumann. DISN: deep implicit surface network for high-quality single-view 3D reconstruction. In NeurIPS, pages 490–500, 2019. 1, 2
  55. 55.Xinchen Yan, Jimei Yang, Ersin Yumer, Yijie Guo, and Honglak Lee. Perspective transformer nets: Learning single-view 3D object reconstruction without 3D supervision. In NeurIPS, pages 1696–1704, 2016. 2
  56. 56.Guandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu, Serge J. Belongie, and Bharath Hariharan. Pointflow: 3D point cloud generation with continuous normalizing flows. In ICCV, pages 4540–4549, 2019. 1, 2
  57. 57.Xumin Yu, Yongming Rao, Ziyi Wang, Zuyan Liu, Jiwen Lu, and Jie Zhou. Pointr: Diverse point cloud completion with geometry-aware transformers. In ICCV, pages 12478–12487, 2021. 5
  58. 58.Xiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, and Karsten Kreis. LION: latent point diffusion models for 3D shape generation. 2022. 2
  59. 59.Linqi Zhou, Yilun Du, and Jiajun Wu. 3D shape generation and completion through point-voxel diffusion. In ICCV, pages 5806–5815, 2021. 2, 5

Citation

MLA
Li, M., et al. “Diffusion-SDF: Text-to-Shape via Voxelized Diffusion”. arXiv, 2022, http://arxiv.org/abs/2212.03293v2.
APA
Li, M., Duan, Y., Zhou, J., & Lu, J. (2022). Diffusion-SDF: Text-to-Shape via Voxelized Diffusion. arXiv. http://arxiv.org/abs/2212.03293v2
Chicago
Li, M., Y. Duan, J. Zhou, and J. Lu. 2022. “Diffusion-SDF: Text-to-Shape via Voxelized Diffusion”. arXiv. http://arxiv.org/abs/2212.03293v2.
Harvard
Li, M. et al. (2022) “Diffusion-SDF: Text-to-Shape via Voxelized Diffusion”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.03293v2.
Vancouver
1. Li M, Duan Y, Zhou J, Lu J (2022) Diffusion-SDF: Text-to-Shape via Voxelized Diffusion. arXiv

BibTeX

@article{li2022diffusion,
  title = {Diffusion-SDF: Text-to-Shape via Voxelized Diffusion},
  author = {Li, Muheng and Duan, Yueqi and Zhou, Jie and Lu, Jiwen},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.03293v2},
  eprint = {2212.03293}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE