CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language
Aditya SanghiRao FuVivian LiuKarl D. D. WillisHooman ShayaniAmir Hosein KhasahmadiSrinath SridharDaniel Ritchie
Proposes CLIP-Sculptor, a multi-resolution generative framework that synthesizes diverse, high-fidelity 3D shapes from natural language prompts without requiring paired text-shape training data by combining discrete latent transformers with annealed classifier-free guidance.
Generating 3D shapes directly from natural language prompts has significant potential across digital content creation, robotics, and industrial design. However, developing robust text-to-3D systems is hindered by the scarcity of large datasets containing paired text descriptions and 3D shapes. While existing methods attempt to bypass this limitation using pretrained vision-language models, they typically suffer from slow optimization times, low geometric fidelity, and limited shape diversity.
The article introduces and evaluates CLIP-Sculptor, a multi-resolution generative framework designed to produce high-fidelity, diverse 3D shapes from text queries without requiring paired text and 3D data during training. The primary objective is to demonstrate that combining discrete latent representations with hierarchical transformer models and an adaptive guidance schedule significantly improves generation accuracy and diversity while maintaining rapid inference speeds.
To achieve zero-shot text-to-shape synthesis, the approach uses a three-stage training pipeline. First, it encodes 3D voxel grids into discrete representations at both low and high resolutions using vector-quantized autoencoders. Second, it trains a coarse transformer model conditioned on image embeddings from a pretrained vision-language model (CLIP) using rendered 2D views of 3D shapes, perturbing embeddings with noise to bridge the cross-modal gap. Third, a fine transformer performs latent super-resolution to upscale the coarse representations to high-resolution geometry. Generation quality is controlled at inference time using an annealed classifier-free guidance schedule, which dynamically adjusts guidance intensity across iterative decoding steps. The framework was evaluated on the ShapeNet benchmark across 13 and 55 object categories.
The experimental findings demonstrate substantial performance advantages over existing baselines. Quantitative evaluations show that CLIP-Sculptor achieves superior semantic accuracy and diversity, reaching a classifier accuracy of approximately 87.5% and reducing the Fréchet Inception Distance (FID) to 1480.11 on ShapeNet13, outperforming competing zero-shot baselines by a wide margin. In addition, the method executes inference in roughly 0.91 seconds per shape, representing a dramatic speed advantage over optimization-based alternatives that require between 30 minutes and 24 hours. The findings also confirm that annealed guidance scheduling consistently yields a superior quality-diversity trade-off compared to conventional constant-scale guidance, and that latent-based super-resolution outperforms direct 3D synthesis baselines.
These results show that automated 3D shape synthesis can be achieved efficiently without costly manual text-shape annotation. By generating editable voxel-based geometries rapidly, the framework reduces production timelines and computational overhead, functioning as a practical aid to enhance designer workflows rather than replace them. The findings also establish that dynamic guidance schedules can improve outputs in masked generative models.
For future development, the article recommends investigating implicit shape representations to capture finer geometric details, incorporating architectures capable of numerical counting (such as specific numbers of components), and expanding models to synthesize surface textures. While the current results are highly reliable for common categories represented in vision-language models, stakeholders should note limitations when processing out-of-distribution prompts, fine geometric structures, or multi-attribute counting tasks.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. Introduces the joint vision-language contrastive embedding space (CLIP) that CLIP-Sculptor conditions its shape generation transformer upon.
- Paper: StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery, Or Patashnik et al. (2021). Pioneered zero-shot generative manipulation guided by pretrained CLIP representations in latent spaces, laying the foundation for zero-shot text-driven synthesis without paired training data.
- Paper: Magic3D: High-Resolution Text-to-3D Content Creation, Chen-Hsuan Lin et al. (2022). Presents coarse-to-fine multi-resolution optimization for text-driven 3D generation, establishing baseline methodologies and motivation for multi-stage shape fidelity refinement.
- Paper: DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Jeong Joon Park et al. (2019). Establishes continuous signed distance function representations conditioned on latent codes for 3D shape generation, which underpins modern latent-space 3D modeling.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). Provides the foundational two-stage paradigm of training discrete latent tokenizers and autoregressive transformers conditioned on text embeddings for diverse generation.
- Paper: Hierarchical Text-Conditional Image Generation with CLIP Latents, Aditya Ramesh et al. (2022). Demonstrates hierarchical generation conditioned on CLIP latent spaces with guidance techniques to significantly improve generative diversity over realism trade-offs.
- Paper: DreamFusion: Text-to-3D using 2D Diffusion, Ben Poole et al. (2023). Extends text-to-3D synthesis beyond latent shape models by optimizing neural radiance fields directly through 2D diffusion priors using Score Distillation Sampling.
- Paper: ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation, Zhengyi Wang et al. (2023). Generalizes text-to-3D optimization by introducing Variational Score Distillation to achieve higher fidelity and diversity in generated 3D representations.
- Paper: Text-to-3D using Gaussian Splatting, Zilong Chen et al. (2024). Advances zero-shot text-to-3D asset creation by replacing implicit volumetric representations with 3D Gaussian Splatting for faster rendering and finer geometric detail.
- Paper: LRM: Large Reconstruction Model for Single Image to 3D, Yicong Hong et al. (2024). Scales feed-forward 3D asset generation directly via large transformer reconstruction models, bypassing iterative per-shape optimization.
- Paper: Structured 3D Latents for Scalable and Versatile 3D Generation, Jianfeng Xiang et al. (2025). Unifies structured 3D latents across multiple representations using flow transformers to achieve state-of-the-art fidelity and output versatility.
