CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language

Aditya SanghiRao FuVivian LiuKarl D. D. WillisHooman ShayaniAmir Hosein KhasahmadiSrinath SridharDaniel Ritchie

article2023CVPR73 citations

Proposes CLIP-Sculptor, a multi-resolution generative framework that synthesizes diverse, high-fidelity 3D shapes from natural language prompts without requiring paired text-shape training data by combining discrete latent transformers with annealed classifier-free guidance.

Listen

Generating 3D shapes directly from natural language prompts has significant potential across digital content creation, robotics, and industrial design. However, developing robust text-to-3D systems is hindered by the scarcity of large datasets containing paired text descriptions and 3D shapes. While existing methods attempt to bypass this limitation using pretrained vision-language models, they typically suffer from slow optimization times, low geometric fidelity, and limited shape diversity.

The article introduces and evaluates CLIP-Sculptor, a multi-resolution generative framework designed to produce high-fidelity, diverse 3D shapes from text queries without requiring paired text and 3D data during training. The primary objective is to demonstrate that combining discrete latent representations with hierarchical transformer models and an adaptive guidance schedule significantly improves generation accuracy and diversity while maintaining rapid inference speeds.

To achieve zero-shot text-to-shape synthesis, the approach uses a three-stage training pipeline. First, it encodes 3D voxel grids into discrete representations at both low and high resolutions using vector-quantized autoencoders. Second, it trains a coarse transformer model conditioned on image embeddings from a pretrained vision-language model (CLIP) using rendered 2D views of 3D shapes, perturbing embeddings with noise to bridge the cross-modal gap. Third, a fine transformer performs latent super-resolution to upscale the coarse representations to high-resolution geometry. Generation quality is controlled at inference time using an annealed classifier-free guidance schedule, which dynamically adjusts guidance intensity across iterative decoding steps. The framework was evaluated on the ShapeNet benchmark across 13 and 55 object categories.

The experimental findings demonstrate substantial performance advantages over existing baselines. Quantitative evaluations show that CLIP-Sculptor achieves superior semantic accuracy and diversity, reaching a classifier accuracy of approximately 87.5% and reducing the Fréchet Inception Distance (FID) to 1480.11 on ShapeNet13, outperforming competing zero-shot baselines by a wide margin. In addition, the method executes inference in roughly 0.91 seconds per shape, representing a dramatic speed advantage over optimization-based alternatives that require between 30 minutes and 24 hours. The findings also confirm that annealed guidance scheduling consistently yields a superior quality-diversity trade-off compared to conventional constant-scale guidance, and that latent-based super-resolution outperforms direct 3D synthesis baselines.

These results show that automated 3D shape synthesis can be achieved efficiently without costly manual text-shape annotation. By generating editable voxel-based geometries rapidly, the framework reduces production timelines and computational overhead, functioning as a practical aid to enhance designer workflows rather than replace them. The findings also establish that dynamic guidance schedules can improve outputs in masked generative models.

For future development, the article recommends investigating implicit shape representations to capture finer geometric details, incorporating architectures capable of numerical counting (such as specific numbers of components), and expanding models to synthesize surface textures. While the current results are highly reliable for common categories represented in vision-language models, stakeholders should note limitations when processing out-of-distribution prompts, fine geometric structures, or multi-attribute counting tasks.

arXiv: 2211.01427
Cover for CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language

Abstract

Recent works have demonstrated that natural language can be used to generate and edit 3D shapes. However, these methods generate shapes with limited fidelity and diversity. We introduce CLIP-Sculptor, a method to address these constraints by producing high-fidelity and diverse 3D shapes without the need for (text, shape) pairs during training. CLIP-Sculptor achieves this in a multi-resolution approach that first generates in a low-dimensional latent space and then upscales to a higher resolution for improved shape fidelity. For improved shape diversity, we use a discrete latent space which is modeled using a transformer conditioned on CLIP’s image-text embedding space. We also present a novel variant of classifier-free guidance, which improves the accuracy-diversity trade-off. Finally, we perform extensive experiments demonstrating that CLIP-Sculptor outperforms state-of-the-art baselines.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Multi-Resolution Voxel VQ-VAEs
  • 3.2. CLIP-Conditioned Coarse Transformer
  • 3.3. Super-Resolution with Fine Transformer
  • 3.4. Annealed Classifier-Free Guidance
  • 4. Experiments
  • 4.1. Evaluating Shape Diversity and Accuracy
  • 4.2. Major Components for Conditional Generation
  • 4.3. Annealing Strategy for Scale Parameter
  • 4.4. Super-Resolution
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — CLIP-Sculptor Multi-Stage Architecture for Zero-Shot Text-to-3D Generation

    model/method

    CLIP-Sculptor is a generative framework that synthesizes diverse, high-fidelity 3D shapes from natural language text prompts without requiring paired (text,shape)(text, shape) training data. The method is trained in three stages:

    1. Stage 1 (Discrete Latent Autoencoding): Two separate 3D Vector-Quantized Variational Autoencoders (VQ-VAEs) are trained on voxel grids of resolution 32332^3 and 64364^3, compressing them into discrete latent grid codebooks E32\mathbf{E}_{32} of resolution 434^3 and E64\mathbf{E}_{64} of resolution 838^3, respectively.

    2. Stage 2 (Coarse Shape Generation): A coarse masked transformer Tc(⋅)T_c(\cdot) is trained on the low-resolution latent grids E32\mathbf{E}_{32}. To enable zero-shot text-to-shape conditioning without text data during training, the transformer is conditioned on perturbed image embeddings produced by the pre-trained CLIP image encoder from rendered multi-view 2D images {Ir}\{\mathbf{I}_r\} of the 3D shapes.

    3. Stage 3 (Latent Super-Resolution): A fine transformer Tf(⋅)T_f(\cdot) is trained to perform super-resolution by unmasking the high-resolution latent grid E64\mathbf{E}_{64} conditioned on the predicted coarse latent grids E~32\tilde{\mathbf{E}}_{32} from the coarse transformer via cross-attention.

    At inference time, an input text prompt is encoded by the pre-trained CLIP text encoder fT(⋅)f_T(\cdot) and mapped into conditioning vectors for the coarse transformer. Iterative unmasking with an annealed classifier-free guidance schedule produces a low-resolution discrete representation E~32\tilde{\mathbf{E}}_{32}, which the fine transformer upsamples into a high-resolution discrete grid E64\mathbf{E}_{64}. The pre-trained Stage 1 64364^3 VQ-VAE decoder converts E64\mathbf{E}_{64} into the final output voxel shape V64\mathbf{V}_{64}.

  2. Knowl 2 — Multi-Resolution Voxel VQ-VAE Representation

    model/method

    To capture discrete representations of 3D geometry and avoid posterior collapse, CLIP-Sculptor trains two hierarchical 3D Vector-Quantized Variational Autoencoders (VQ-VAEs) using ResNet-based volumetric 3D convolutional neural network (CNN) encoders and decoders:

    • A low-resolution VQ-VAE compresses a 32332^3 voxel grid V32\mathbf{V}_{32} into a 434^3 discrete latent token grid E32\mathbf{E}_{32}, which captures global semantic shape geometry for conditional generation.
    • A high-resolution VQ-VAE compresses a 64364^3 voxel grid V64\mathbf{V}_{64} into an 838^3 discrete latent token grid E64\mathbf{E}_{64}, which encodes fine local details.

    Both VQ-VAEs are trained using mean squared error (MSE) reconstruction loss, codebook commitment loss, and an exponential moving average (EMA) codebook update rule.

  3. Knowl 3 — Modality-Gap Perturbed CLIP Conditioning Vector

    equation

    To train the coarse shape generator without paired text-shape annotations, CLIP-Sculptor leverages the joint embedding space of a pre-trained CLIP model. During training, conditioning vectors are derived from rendered images Ir\mathbf{I}_r of the 3D shape from random viewpoint r∈Rr \in R. To bridge the empirical modality gap between CLIP image embeddings and inference-time CLIP text embeddings, isotropic Gaussian noise is added to the normalized image embedding:

    c^=fI(Ir)+γ⋅ϵ⋅∥fI(Ir)∥2∥ϵ∥2\hat{\mathbf{c}} = f_{I}(\mathbf{I}_r) + \frac{\gamma \cdot \boldsymbol{\epsilon} \cdot \|f_{I}(\mathbf{I}_r)\|_2}{\|\boldsymbol{\epsilon}\|_2}

    where fI(⋅)f_I(\cdot) denotes the frozen CLIP image encoder, ϵ∼N(0,I)\boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}) is standard Gaussian noise, and γ∈R+\gamma \in \mathbb{R}^+ is a scaling hyperparameter controlling the magnitude of noise perturbation (empirically optimal around γ=1.2\gamma = 1.2).

    The perturbed vector c^\hat{\mathbf{c}} is normalized and processed through a multi-layer perceptron (MLP) mapping network with LL layers to yield the conditioning vector c~\tilde{\mathbf{c}}, which modulates each transformer block by predicting adaptive layer-normalization affine parameters.

  4. Knowl 4 — Step-Unrolled Training Objective for Coarse Transformer

    equation

    To counteract autoregressive drift caused by compounding errors during iterative masked decoding at inference time, the coarse masked transformer Tc(⋅)T_c(\cdot) is trained with Step-Unrolled Training (SUT). The objective maximizes the log-likelihood of ground truth tokens given both randomly masked inputs and the network's own intermediate predictions:

    L=1N∑n=1N(log⁡p(E32,n∣E32,nmsk,c)+log⁡p(E32,n∣E~32,n,c))\mathcal{L} = \frac{1}{N} \sum_{n=1}^N \left( \log p(\mathbf{E}_{32,n} \mid \mathbf{E}^{\text{msk}}_{32,n}, \mathbf{c}) + \log p(\mathbf{E}_{32,n} \mid \tilde{\mathbf{E}}_{32,n}, \mathbf{c}) \right)

    where NN is the total number of training samples, E32,n\mathbf{E}_{32,n} is the ground-truth discrete low-resolution latent grid of sample nn, E32,nmsk\mathbf{E}^{\text{msk}}_{32,n} is the latent grid masked with a masking ratio sampled uniformly from [0,1][0, 1], c\mathbf{c} is the conditioning vector, and E~32,n\tilde{\mathbf{E}}_{32,n} is the coarse transformer's own unmasked prediction from a prior forward pass.

  5. Knowl 5 — Discrete Latent Super-Resolution via Fine Transformer

    model/method

    Super-resolution is performed directly in discrete latent space using a fine masked transformer Tf(⋅)T_f(\cdot). The network takes as input a coarse latent grid E32\mathbf{E}_{32} (size 434^3) and a masked fine latent grid E64msk\mathbf{E}^{\text{msk}}_{64} (size 838^3), using cross-attention layers to condition the generation of high-resolution tokens on the coarse representation.

    To align the training distribution with inference conditions, the fine transformer is trained using the coarse transformer's predicted output E~32\tilde{\mathbf{E}}_{32} rather than ground-truth low-resolution tokens E32\mathbf{E}_{32}. At inference, an initially fully masked grid E64msk\mathbf{E}^{\text{msk}}_{64} is iteratively unmasked over TT steps, and the resulting unmasked grid E64\mathbf{E}_{64} is decoded by the Stage 1 64364^3 VQ-VAE decoder into the final 64364^3 voxel grid V64\mathbf{V}_{64}.

  6. Knowl 6 — Annealed Classifier-Free Guidance for Masked Transformers

    equation

    During inference iterative decoding over TT steps, classifier-free guidance extrapolates the conditional token probability distribution away from the unconditional distribution using a time-dependent guidance scale schedule a(t)a(t):

    P^t(c)=Pt(0)+a(t)(Pt(c)−Pt(0))\hat{P}_{t}(\mathbf{c}) = P_{t}(\mathbf{0}) + a(t) \Big( P_{t}(\mathbf{c}) - P_{t}(\mathbf{0}) \Big)

    where:

    • Pt(c)=p(E32,t∣E32,t−1msk,c)P_{t}(\mathbf{c}) = p(\mathbf{E}_{32,t} \mid \mathbf{E}^{\text{msk}}_{32,t-1}, \mathbf{c}) is the conditional distribution given text conditioning vector c\mathbf{c},
    • Pt(0)=p(E32,t∣E32,t−1msk,0)P_{t}(\mathbf{0}) = p(\mathbf{E}_{32,t} \mid \mathbf{E}^{\text{msk}}_{32,t-1}, \mathbf{0}) is the unconditional distribution obtained with a null embedding (trained by dropping out conditioning with probability ρ%\rho \%),
    • a(t)a(t) is a continuous, monotonically decreasing guidance schedule (such as square root, linear, or cosine decay functions).

    A higher guidance scale a(t)a(t) in early decoding steps enforces strong prompt fidelity when most tokens are masked, while decreasing a(t)a(t) in later steps encourages shape diversity and prevents visual artifacts.

  7. Knowl 7 — Quantitative Text-to-Shape Generation Performance on ShapeNet13

    data/table

    Quantitative comparison of text-to-shape generation evaluated on 234 predetermined text queries on the ShapeNet13 benchmark using Classifier Accuracy (Acc %, measuring semantic alignment) and Fréchet Inception Distance (FID, measuring diversity and shape fidelity, where lower is better). Baselines include CLIP-Forge (CF) with various latent sampling techniques (mean shape CF-MS, Gaussian CF-G, truncated Gaussian CF-TG, clipped Gaussian CF-CG) and AutoSDF augmented with CLIP embeddings (ZS-ASDF).

    Method CF-MS CF-G CF-TG CF-CG ZS-ASDF CS-const. CS-sqrt CS-linear CS-cosine
    FID ↓\downarrow 2425.25 2233.48 2141.61 2100.67 7332.93 1821.78 1480.11 1629.51 1725.63
    ACC ↑\uparrow 83.33 62.81 68.71 71.11 39.14 86.59 87.08 87.50 87.27

    All CLIP-Sculptor (CS) variants significantly outperform CLIP-Forge and ZS-ASDF in both accuracy and FID. Annealed guidance schedules (square root, linear, cosine) outperform constant scale guidance (CS-const.), achieving improved accuracy-diversity Pareto frontiers (e.g., CS-sqrt reduces FID from 1821.78 to 1480.11 while improving classification accuracy from 86.59% to 87.08%).

  8. Knowl 8 — Ablations on Noise Perturbation, Mapping Layers, and Super-Resolution Baselines

    data/table

    Ablation experiments evaluate the impact of conditioning noise scale γ\gamma, mapping network depth LL, and super-resolution design on ShapeNet13 generation fidelity (FID ↓\downarrow) and semantic accuracy (Acc ↑\uparrow):

    (a) Effect of varying the Noise parameter γ\gamma
    γ\gamma ×\times (no noise) 0.5 0.8 1.0 1.2 1.5
    FID ↓\downarrow 1720.02 1764.98 1484.61 1703.38 1447.91 1478.17
    ACC ↑\uparrow 64.87 75.73 77.41 79.09 79.63 78.47
    (b) Effect of varying number of mapping network layers LL
    LL 0 1 2 3 4 5
    FID ↓\downarrow 2874.87 1716.73 1518.97 1447.91 1532.50 1424.46
    ACC ↑\uparrow 62.72 78.70 79.17 79.63 77.16 76.09
    (c) Comparison with super-resolution baselines
    Method 3D-UNet 64364^3-DS (Direct Synth.) CLIP-Sculptor
    FID ↓\downarrow 2056.92 2196.96 1910.28
    ACC ↑\uparrow 86.65 77.92 86.85

    The optimal noise scale γ=1.2\gamma = 1.2 boosts accuracy by 14.76%14.76\% compared to training without noise. Using L=3L=3 mapping layers provides the best accuracy-FID balance. The discrete latent super-resolution transformer in CLIP-Sculptor outperforms both a voxel-domain 3D-UNet super-resolution baseline and a direct 64364^3 synthesis transformer without multi-resolution staging.

  9. Knowl 9 — Effect of Conditioning Dropout, Guidance Scale, and Step-Unrolled Training

    data/table

    Evaluation of conditioning dropout rate ρ%\rho \%, constant classifier-free guidance scale a(t)=ka(t) = k, and Step-Unrolled Training (SUT) on ShapeNet13 text-to-shape synthesis:

    a(t)=3a(t) = 3 a(t)=2.5a(t) = 2.5 a(t)=2a(t) = 2 a(t)=1.5a(t) = 1.5 a(t)=1a(t) = 1
    ρ%\rho\% SUT FID ↓\downarrow Acc ↑\uparrow FID ↓\downarrow Acc ↑\uparrow FID ↓\downarrow Acc ↑\uparrow FID ↓\downarrow Acc ↑\uparrow FID ↓\downarrow Acc ↑\uparrow
    5 ×\times 2059.0 85.73 1970.1 85.47 1790.9 85.43 1536.5 83.98 1227.5 79.17
    10 ×\times 1893.2 84.46 1821.2 84.50 1684.3 83.41 1522.4 82.34 1348.4 77.76
    15 ×\times 2086.9 85.03 1964.7 84.99 1851.9 84.36 1660.9 83.14 1485.5 78.19
    20 ×\times 2062.5 83.39 1972.3 82.95 1892.9 82.74 1733.3 81.13 1566.6 75.94
    5 ✓\checkmark 2039.8 87.69 2011.1 87.39 1811.8 87.40 1678.6 86.24 1517.9 82.18

    Increasing the guidance scale a(t)a(t) increases classification accuracy at the expense of sample diversity (higher FID). A dropout rate of ρ=5%\rho = 5\% to 15%15\% establishes an effective balance. Incorporating SUT consistently increases classification accuracy across all guidance scale settings (e.g., improving accuracy from 85.73%85.73\% to 87.69%87.69\% at a(t)=3a(t)=3 with ρ=5%\rho=5\%).

  10. Knowl 10 — Limitations of CLIP-Sculptor

    limitation

    CLIP-Sculptor is subject to four primary limitations:

    1. Fine Structural Details: The model struggles to synthesize small-scale geometric topological features, such as a chair with a void or hole in its backrest.
    2. Numerical Counting in Prompts: The system cannot reliably generate shapes that depend on explicit counting concepts specified in natural language (e.g., "chair with four slats").
    3. CLIP Distributional Bias: Text prompts describing object semantics absent from or poorly represented in pre-trained CLIP feature representations fail to generate accurate geometries.
    4. Lack of Surface Appearance: The model is restricted to voxel geometry synthesis and lacks the capability to generate surface textures or colors.

Coverage note — None was omitted; all substantive contributions, including the multi-stage discrete voxel architecture, training losses, guidance scheduling formulations, quantitative benchmarks, ablation studies, and limitations, are represented.

References

  1. 1.Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas Guibas. Learning representations and generative models for 3d point clouds. In International conference on machine learning, pages 40–49. PMLR, 2018. 2
  2. 2.Panos Achlioptas, Judy Fan, X.D. Robert Hawkins, D. Noah Goodman, and J. Leonidas Guibas. ShapeGlot: Learning language for shape differentiation. CoRR, abs/1905.02925, 2019. 1
  3. 3.Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. Shapenet: An information-rich 3d model repository, 2015. cite arxiv:1512.03012. 5
  4. 4.Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman. Maskgit: Masked generative image transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11315–11325, 2022. 2, 3, 4, 5
  5. 5.Kevin Chen, Christopher B Choy, Manolis Savva, Angel X Chang, Thomas Funkhouser, and Silvio Savarese. Text2shape: Generating shapes from natural language by learning joint embeddings. In Asian Conference on Computer Vision, pages 100–116. Springer, 2018. 1, 2
  6. 6.Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5939–5948, 2019. 2
  7. 7.Christopher B Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction. In European conference on computer vision, pages 628–644. Springer, 2016. 5
  8. 8.Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341, 2020. 2
  9. 9.Rao Fu, Xiao Zhan, Yiwen Chen, Daniel Ritchie, and Srinath Sridhar. Shapecrafter: A recursive text-conditioned 3d shape generation model, 2022. 1, 2
  10. 10.Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. Make-a-scene: Scenebased text-to-image generation with human priors, 2022. 2
  11. 11.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H. Bermano, Gal Chechik, and Daniel Cohen-Or. An image is worth one word: Personalizing text-to-image generation using textual inversion, 2022. 2
  12. 12.Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10696–10706, 2022. 2, 4
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 4
  14. 14.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017. 5
  15. 15.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. Imagen video: High definition video generation with diffusion models, 2022. 2
  16. 16.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 2, 4, 5
  17. 17.Moritz Ibing, Isaak Lim, and Leif Kobbelt. 3d shape generation with grid-based implicit functions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13559–13568, 2021. 2
  18. 18.Ajay Jain, Ben Mildenhall, Jonathan T Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object generation with dream fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 867–876, 2022. 1, 2, 5, 6
  19. 19.Nasir Khalid, Tianhao Xie, Eugene Belilovsky, and Tiberiu Popa. Clip-mesh: Generating textured meshes from text using pretrained image-text models. ACM Transactions on Graphics (TOG), Proc. SIGGRAPH Asia, 2022. 1, 2, 5, 6
  20. 20.Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Zou. Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning. arXiv preprint arXiv:2203.02053, 2022. 4
  21. 21.Vivian Liu, Han Qiao, and Lydia Chilton. Opal: Multimodal image generation for news illustration, 2022. 1
  22. 22.Vivian Liu, Jo Vermeulen, George Fitzmaurice, and Justin Matejka. 3dall-e: Integrating text-to-image ai in 3d design workflows, 2022. 1
  23. 23.Zhengzhe Liu, Yi Wang, Xiaojuan Qi, and Chi-Wing Fu. Towards implicit text-guided 3d shape generation. arXiv preprint arXiv:2203.14622, 2022. 1, 2
  24. 24.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2019. 5
  25. 25.Oscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim, and Rana Hanocka. Text2mesh: Text-driven neural stylization for meshes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13492–13502, 2022. 1, 2
  26. 26.Paritosh Mittal, Yen-Chi Cheng, Maneesh Singh, and Shubham Tulsiani. Autosdf: Shape priors for 3d completion, reconstruction and generation. arXiv preprint arXiv:2203.09516, 2022. 1, 2, 5
  27. 27.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2021. 4
  28. 28.Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. arXiv preprint arXiv:1711.00937, 2017. 2, 3, 4
  29. 29.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv, 2022. 2
  30. 30.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021. 1, 2
  31. 31.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 1
  32. 32.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021. 1
  33. 33.Ali Razavi, Aaron Van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems, 32, 2019. 2, 4
  34. 34.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2021. 2
  35. 35.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, June 2022. 1
  36. 36.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022. 2, 4
  37. 37.Aditya Sanghi, Hang Chu, Joseph G Lambourne, Ye Wang, Chin-Yi Cheng, Marco Fumero, and Kamal Rahimi Malekshan. Clip-forge: Towards zero-shot text-to-shape generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18603–18613, 2022. 1, 2, 5, 6
  38. 38.Nikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Erich Elsen, and Aaron van den Oord. Step-unrolled denoising autoencoders for text generation. arXiv preprint arXiv:2112.06749, 2021. 4
  39. 39.Nikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Erich Elsen, and Aaron van den Oord. Step-unrolled denoising autoencoders for text generation, 2021. 7
  40. 40.Mohit Shridhar, Lucas Manuelli, and Dieter Fox. Cliport: What and where pathways for robotic manipulation. In Proceedings of the 5th Conference on Robot Learning (CoRL), 2021. 1
  41. 41.Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. Clip-nerf: Text-and-image driven manipulation of neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3835–3844, 2022. 1, 2
  42. 42.Xingguang Yan, Liqiang Lin, Niloy J Mitra, Dani Lischinski, Daniel Cohen-Or, and Hui Huang. Shapeformer: Transformer-based shape completion via sparse representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6239–6249, 2022. 2
  43. 43.Guandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu, Serge Belongie, and Bharath Hariharan. Pointflow: 3d point cloud generation with continuous normalizing flows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4541–4550, 2019. 2
  44. 44.Yufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu, Jiuxiang Gu, Jinhui Xu, and Tong Sun. Lafite: Towards language-free training for text-toimage generation. arXiv preprint arXiv:2111.13792, 2021. 4
  45. 45.Adrian Łańcucki, Jan Chorowski, Guillaume Sanchez, Ricard Marxer, Nanxin Chen, Hans J. G. A. Dolfing, Sameer Khurana, Tanel Alumäe, and Antoine Laurent. Robust training of vector quantized bottleneck models, 2020. 4

Citation

MLA
Sanghi, A., et al. “CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language”. arXiv, 2022, http://arxiv.org/abs/2211.01427v4.
APA
Sanghi, A., Fu, R., Liu, V., Willis, K., Shayani, H., Khasahmadi, A. H., Sridhar, S., & Ritchie, D. (2022). CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language. arXiv. http://arxiv.org/abs/2211.01427v4
Chicago
Sanghi, A., R. Fu, V. Liu, et al. 2022. “CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language”. arXiv. http://arxiv.org/abs/2211.01427v4.
Harvard
Sanghi, A. et al. (2022) “CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.01427v4.
Vancouver
1. Sanghi A, Fu R, Liu V, Willis K, Shayani H, Khasahmadi AH, Sridhar S, Ritchie D (2022) CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language. arXiv

BibTeX

@article{sanghi2022clip,
  title = {CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language},
  author = {Sanghi, Aditya and Fu, Rao and Liu, Vivian and Willis, Karl and Shayani, Hooman and Khasahmadi, Amir Hosein and Sridhar, Srinath and Ritchie, Daniel},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.01427v4},
  eprint = {2211.01427}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE