Text-to-3D using Gaussian Splatting

Zilong ChenFeng WangYikai WangHuaping Liu

article2024CVPR397 citations

Presents GSGEN, a text-to-3D generation framework that pairs 3D Gaussian Splatting with joint 2D and 3D diffusion priors to resolve the multi-face Janus problem and synthesize detailed, geometrically consistent 3D assets.

Listen

Automated generation of three-dimensional digital assets directly from text prompts has become a critical capability for interactive media, gaming, simulation, and design. However, existing methods relying on two-dimensional image diffusion models and implicit neural volumetric representations frequently suffer from severe structural defects, such as the multi-face "Janus problem" where objects generate multiple front sides or collapsed geometry. Additionally, these approaches struggle to capture intricate, high-frequency surface details and typically require extensive rendering and optimization time due to their implicit mathematical formulations.

The article introduces and evaluates GSGEN, a novel framework designed to synthesize high-fidelity, geometrically accurate 3D objects from text prompts. It demonstrates how leveraging 3D Gaussian Splatting—an explicit point-like representation—enables direct geometric guidance and superior rendering quality compared to conventional implicit techniques.

To achieve this, the approach employs a two-stage optimization process combined with geometric initialization. The system initializes Gaussian positions using a 3D point cloud diffusion model (Point-E) to establish a rough anisotropic structure. In the first stage, geometry optimization jointly applies 2D image score distillation and 3D point cloud score distillation to enforce 3D structural consistency. In the second stage, appearance refinement optimizes fine-grained textures using only 2D image guidance while applying a custom "compactness-based densification" strategy. This technique inserts new Gaussians between existing neighboring points to close structural gaps and enhance surface continuity without destabilizing optimization.

The findings demonstrate four main outcomes. First, integrating explicit 3D point cloud priors with 2D image priors successfully mitigates the Janus problem, establishing coherent geometry even on complex asymmetric prompts like animals and vehicles. Second, the explicit Gaussian Splatting representation significantly outperforms mesh- and implicit-based baselines in rendering fine, high-frequency textures such as fur, feathers, and patterned surfaces. Third, the compactness-based densification resolves optimization instability under score distillation sampling, preventing the over-smoothing caused by high gradient thresholds and the uncontrolled point explosion caused by low thresholds. Fourth, the framework achieves this high visual fidelity in approximately 40 minutes per asset, matching the runtime of standard mesh-based methods while delivering noticeably sharper geometric and textural quality.

These results show that transitioning from implicit coordinate networks to explicit Gaussian representations unlocks substantial performance and visual quality improvements in automated asset creation. For organizations developing 3D pipelines, this approach reduces the risk of generating physically implausible models, shortens design iterations, and eliminates costly manual touch-ups needed to repair multi-face artifacts.

Teams exploring generative 3D workflows should consider adopting explicit 3D Gaussian representations and dual 2D/3D prior guidance for asset synthesis pipelines. Decision-makers should evaluate pilot implementations where high surface detail is needed, such as character design or detailed props. Future technical exploration should evaluate pairing the framework with more advanced multi-view diffusion models and specialized large language models to broaden prompt complexity.

Confidence in these findings is high for individual object synthesis, as qualitative and ablation experiments consistently validate the architectural choices. However, limitations remain: the framework struggles with prompts requiring complex logic or dense multi-object scene descriptions, bounded by the language comprehension limits of the underlying text encoders. Furthermore, if a text prompt triggers extreme bias in the guidance diffusion models, geometric artifacts can still occasionally occur.

Cover for Text-to-3D using Gaussian Splatting

Abstract

Automatic text-to-3D generation that combines Score Distillation Sampling (SDS) with the optimization of volume rendering has achieved remarkable progress in synthesizing realistic 3D objects. Yet most existing text-to-3D methods by SDS and volume rendering suffer from inaccurate geometry, e.g., the Janus issue, since it is hard to explicitly integrate 3D priors into implicit 3D representations. Besides, it is usually time-consuming for them to generate elaborate 3D models with rich colors. In response, this paper proposes GSGEN, a novel method that adopts Gaussian Splatting, a recent state-of-the-art representation, to text-to-3D generation. GSGEN aims at generating high-quality 3D objects and addressing existing shortcomings by exploiting the explicit nature of Gaussian Splatting that enables the incorporation of 3D prior. Specifically, our method adopts a progressive optimization strategy, which includes a geometry optimization stage and an appearance refinement stage. In geometry optimization, a coarse representation is established under 3D point cloud diffusion prior along with the ordinary 2D SDS optimization, ensuring a sensible and 3D-consistent rough shape. Subsequently, the obtained Gaussians undergo an iterative appearance refinement to enrich texture details. In this stage, we increase the number of Gaussians by compactness-based densification to enhance continuity and improve fidelity. With these designs, our approach can generate 3D assets with delicate details and accurate geometry. Extensive evaluations demonstrate the effectiveness of our method, especially for capturing high-frequency components. Our code is available at https://github.com/gsgen3d/gsgen.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. 3D Scene Representations
  • 2.2. Diffusion Models
  • 2.3. Text-to-3D Generation
  • 3. Preliminary
  • 3.1. Score Distillation Sampling
  • 3.2. 3D Gaussian Splatting
  • 4. Approach
  • 4.1. Geometry Optimization
  • 4.2. Appearance Refinement
  • 4.3. Initialization with Geometry Prior
  • 5. Experiments
  • 5.1. Implementation Details
  • 5.2. Text-to-3D Generation
  • 5.3. Ablation Study
  • 6. Limitations and Conclusion
  • 7. Acknowledgement
  • References

Knowls

  1. Knowl 1 — GSGEN Framework for Text-to-3D Generation

    model/method

    GSGEN is a text-to-3D generative framework that models 3D content using explicit 3D Gaussian Splatting (3DGS) optimized through Score Distillation Sampling (SDS). Unlike implicit neural representations (such as Neural Radiance Fields or DMTet) that struggle to integrate direct geometric guidance and often suffer from the multi-face "Janus" problem or over-smoothed surfaces, GSGEN leverages the explicit point-like nature of 3D Gaussians to incorporate both 3D geometric priors and 2D text-to-image diffusion priors.

    The framework operates in three sequential phases:

    1. Initialization with Geometry Prior: 3D Gaussian positions are initialized using a text-to-point-cloud diffusion model (Point-E) conditioned on the input text prompt, breaking spatial symmetry from the start.
    2. Geometry Optimization: The Gaussians are optimized under joint guidance from a 3D point cloud diffusion model (Point-E) via a 3D SDS loss and a 2D image diffusion model (Stable Diffusion) via a 2D SDS loss to establish a coherent, 3D-consistent coarse structure.
    3. Appearance Refinement: The 3D SDS loss is removed to avoid disturbing fine texture synthesis. The Gaussians are iteratively refined using 2D SDS alongside spatial regularization and a compactness-based densification mechanism to capture high-frequency details (such as animal fur, feathers, and intricate textures).
  2. Knowl 2 — Geometry Optimization Stage with Joint 2D and 3D Score Distillation Sampling

    model/method

    In the geometry optimization stage of GSGEN, the goal is to construct a coarse 3D shape that avoids multi-face Janus artifacts and geometric collapse. Because a 3D point cloud can be interpreted as a set of isotropic Gaussians, Gaussian Splatting enables direct application of geometric supervision to Gaussian center positions.

    Rather than performing direct point-cloud registration or rigid alignment (which introduces scaling and degeneration challenges), GSGEN optimizes Gaussian parameters under a combined 2D and 3D Score Distillation Sampling (SDS) objective. The 3D SDS loss leverages a pre-trained text-to-point-cloud diffusion model (Point-E) to guide the 3D positions of the Gaussians directly in space, while the 2D SDS loss leverages a text-to-image diffusion model (Stable Diffusion) to supervise the 2D differentiable splatted renderings across sampled camera viewpoints.

  3. Knowl 3 — Geometry Optimization Gradient Formulation

    equation

    During the geometry optimization stage of GSGEN, the gradient update for the Gaussian Splatting parameters θ\theta is defined by:

    ∇θLgeometry=EϵI,t[wI(t)(ϵϕ(xt;y,t)−ϵI)∂x∂θ]+λ3D⋅EϵP,t[wP(t)(ϵψ(pt;y,t)−ϵP)]\nabla_\theta \mathcal{L}_\text{geometry} = \mathbb{E}_{\epsilon_I, t} \left[ w_I(t) \left( \epsilon_\phi(x_t; y, t) - \epsilon_I \right) \frac{\partial x}{\partial \theta} \right] + \lambda_\text{3D} \cdot \mathbb{E}_{\epsilon_P, t} \left[ w_P(t) \left( \epsilon_\psi(p_t; y, t) - \epsilon_P \right) \right]

    where:

    • θ\theta denotes the trainable 3D Gaussian parameters (positions, covariances, colors, and opacities).
    • x=g(θ)x = g(\theta) is the 2D image rendered from the Gaussians under camera parameters through splatting function gg.
    • xtx_t is the noisy 2D rendered image at diffusion timestep tt, ϵI∼N(0,I)\epsilon_I \sim \mathcal{N}(0, \mathbf{I}) is 2D Gaussian noise, and wI(t)w_I(t) is a weighting function for 2D SDS.
    • ϵϕ(xt;y,t)\epsilon_\phi(x_t; y, t) is the predicted noise score from a 2D text-to-image diffusion model (Stable Diffusion) conditioned on text prompt embedding yy.
    • ptp_t represents the noisy 3D Gaussian positions at diffusion timestep tt, ϵP∼N(0,I)\epsilon_P \sim \mathcal{N}(0, \mathbf{I}) is 3D Gaussian noise, and wP(t)w_P(t) is a weighting function for 3D SDS.
    • ϵψ(pt;y,t)\epsilon_\psi(p_t; y, t) is the predicted score from a 3D text-to-point-cloud diffusion model (Point-E).
    • λ3D\lambda_\text{3D} is the loss weight balancing the 3D point cloud prior with the 2D image prior.
  4. Knowl 4 — Appearance Refinement Stage and Spatial Regularization

    model/method

    While 3D point cloud diffusion priors establish consistent coarse geometry, retaining the 3D SDS loss during later stages can constrain high-frequency texture formation and lead to blurred or under-detailed appearances. GSGEN addresses this by transitioning to an appearance refinement stage that uses only 2D image diffusion SDS.

    To ensure that the Gaussians do not drift away from the geometry established in the first stage or generate floating background artifacts, GSGEN introduces two spatial regularizers:

    1. Mean Position Regularization: Penalizes the L2L_2 norm of Gaussian positions ∑i∥pi∥\sum_i \|p_i\| to maintain global structural coherence and penalize significant drift from the origin.
    2. Distance-Weighted Opacity Regularization: Penalizes Gaussian opacities proportionally to their distance from the center, formulated as ∑isg(∥pi∥)⋅oi\sum_i \text{sg}(\|p_i\|) \cdot o_i, where sg(⋅)\text{sg}(\cdot) denotes the stop-gradient operator. This penalizes distant or detached "floater" Gaussians while leaving central geometry intact.
  5. Knowl 5 — Appearance Refinement Gradient Formulation

    equation

    In the appearance refinement stage of GSGEN, the parameter gradient update ∇θLrefine\nabla_\theta \mathcal{L}_\text{refine} is given by:

    ∇θLrefine=λSDS EϵI,t[wI(t)(ϵϕ(xt;y,t)−ϵI)∂x∂θ]+λmean ∇θ∑i∥pi∥+λopacity ∇θ∑isg(∥pi∥)⋅oi\nabla_\theta \mathcal{L}_\text{refine} = \lambda_\text{SDS} \, \mathbb{E}_{\epsilon_I, t} \left[ w_I(t) \left( \epsilon_\phi(x_t; y, t) - \epsilon_I \right) \frac{\partial x}{\partial \theta} \right] + \lambda_\text{mean} \, \nabla_\theta \sum_i \|p_i\| + \lambda_\text{opacity} \, \nabla_\theta \sum_i \text{sg}(\|p_i\|) \cdot o_i

    where:

    • θ\theta represents the set of 3D Gaussian parameters.
    • xx is the rendered 2D view generated by tile-based rasterization.
    • ϵϕ(xt;y,t)\epsilon_\phi(x_t; y, t) is the noise prediction from the 2D image diffusion model for noisy image xtx_t, text embedding yy, and timestep tt, with noise ϵI∼N(0,I)\epsilon_I \sim \mathcal{N}(0, \mathbf{I}) and weighting wI(t)w_I(t).
    • pi∈R3p_i \in \mathbb{R}^3 and oi∈[0,1]o_i \in [0, 1] represent the 3D spatial center position and opacity of the ii-th Gaussian, respectively.
    • sg(⋅)\text{sg}(\cdot) is the stop-gradient operation, ensuring that distance ∥pi∥\|p_i\| acts as a fixed coefficient for the opacity gradient without backpropagating into positions.
    • λSDS\lambda_\text{SDS}, λmean\lambda_\text{mean}, and λopacity\lambda_\text{opacity} are scalar loss weighting hyperparameters.
  6. Knowl 6 — Compactness-Based Densification and Pruning Algorithm

    algorithm

    Standard Gaussian Splatting densifies primitives using view-space positional gradient splitting. Under stochastic Score Distillation Sampling (SDS) gradients, a small threshold creates excessive spurious Gaussians from noise spikes, whereas a large threshold causes under-densification and blurred geometry. GSGEN resolves this dilemma by using a conservative gradient split threshold (Tpos=0.02T_\text{pos} = 0.02) complemented by a compactness-based densification routine to fill structural voids, along with periodic opacity and radius pruning.

    Input: Gaussian set G={(pi,ri,ci,oi,Σi)}i=1NG = \{(p_i, r_i, c_i, o_i, \Sigma_i)\}_{i=1}^N, nearest neighbor count KK, opacity threshold αmin=0.05\alpha_\text{min} = 0.05
    Output: Densified and pruned Gaussian set GG
    Construct KD-Tree from Gaussian center positions {pi}i=1N\{p_i\}_{i=1}^N
    for each Gaussian i∈{1,…,N}i \in \{1, \dots, N\} do
        Find the KK nearest neighbors of pip_i using the KD-Tree
        for each neighbor j∈KNN(i)j \in \text{KNN}(i) do
            dij←∥pi−pj∥2d_{ij} \leftarrow \|p_i - p_j\|_2
            if dij<ri+rjd_{ij} < r_i + r_j then
                rnew←(ri+rj)−dijr_\text{new} \leftarrow (r_i + r_j) - d_{ij}
                pnew←pi+pj2p_\text{new} \leftarrow \frac{p_i + p_j}{2}
                Initialize new Gaussian at pnewp_\text{new} with radius rnewr_\text{new}
                Insert new Gaussian into GG
            end if
        end for
    end for
    for each Gaussian k∈Gk \in G do
        if ok<αmino_k < \alpha_\text{min} or radius of Gaussian kk exceeds maximum world-space or view-space limit then
            Remove Gaussian kk from GG
        end if
    end for
    return GG
  7. Knowl 7 — Geometry Prior Initialization for 3D Gaussians

    model/method

    Initializing 3D Gaussians from a standard origin-centered isotropic distribution frequently leads to severe shape collapse or symmetrical multi-face Janus artifacts in SDS optimization. GSGEN initializes Gaussian positions using geometric priors:

    1. Text-Conditioned Generation: For general text-to-3D generation, initial positions pip_i are sampled from a point cloud generated by Point-E conditioned on the text prompt. While Point-E produces colored point clouds, Gaussian colors are initialized randomly because directly using Point-E colors degrades downstream texture optimization under 2D SDS. Scales and opacities are initialized with fixed constant values, and rotation matrices are initialized to the identity matrix I\mathbf{I}.
    2. User-Guided Shape Initialization: When a user provides a reference shape (mesh or point cloud), initial points are extracted by applying Farthest Point Sampling (FPS) to point clouds or uniform surface sampling to meshes to obtain a compact initial subset.
  8. Knowl 8 — Experimental Setup and Training Hyperparameters of GSGEN

    experimental setup

    GSGEN is evaluated on text-to-3D generation using the following implementation parameters:

    • Diffusion Guidance Models: Stable Diffusion v1.5 (runwayml/stable-diffusion-v1-5) with classifier-free guidance scale 100100 and view-dependent prompt engineering. Point-E is used for 3D text-to-point-cloud guidance.
    • Camera Sampling: Focal length, elevation, and azimuth ranges follow DreamFusion conventions, with stratified sampling applied over azimuth angles to ensure uniform viewpoint coverage.
    • Loss Hyperparameters:
      • Geometry Optimization stage: λSDS=0.1\lambda_\text{SDS} = 0.1, λ3D=0.01\lambda_\text{3D} = 0.01.
      • Appearance Refinement stage: λSDS=0.1\lambda_\text{SDS} = 0.1, λmean=1.0\lambda_\text{mean} = 1.0, λopacity=100.0\lambda_\text{opacity} = 100.0.
    • Optimization and Pruning Schedule:
      • Positional gradient-based Gaussian splitting is executed every 500 iterations with threshold Tpos=0.02T_\text{pos} = 0.02.
      • Compactness-based densification is executed every 1000 iterations.
      • Opacity pruning (removing Gaussians with opacity oi<αmin=0.05o_i < \alpha_\text{min} = 0.05 or excessively large world/view-space radius) is executed every 200 iterations.
    • Runtime: Synthesizing a single 3D asset requires approximately 40 minutes.
  9. Knowl 9 — Empirical Quality and Geometric Consistency vs. Text-to-3D Baselines

    empirical result

    GSGEN is qualitatively compared against existing text-to-3D methods, including DreamFusion, Magic3D, Fantasia3D, and ProlificDreamer:

    • Janus Issue Mitigation: Under identical 2D SDS guidance and text prompts describing asymmetric objects (such as a corgi, panda, or bulldozer), baseline methods (DreamFusion, Magic3D, ProlificDreamer) frequently suffer from multi-face Janus artifacts and collapsed geometry. GSGEN consistently generates 3D-consistent, single-headed geometries.
    • High-Frequency Detail Preservation: Compared to mesh-based methods (Magic3D, Fantasia3D) that yield over-smoothed surfaces due to DMTet discretization constraints, the explicit 3D Gaussian representation in GSGEN faithfully recovers fine, high-frequency textural structures, such as peacock feathers, sushi toppings, thatched cottage roofs, and animal fur.
    • Computational Efficiency: GSGEN synthesizes full 3D assets in ~40 minutes, matching the runtime of Magic3D and Fantasia3D while providing higher textural fidelity and avoiding the multi-face degradation observed in ProlificDreamer.
  10. Knowl 10 — Ablation on Initialization, 3D SDS Guidance, and Densification Strategy

    empirical result

    Ablation experiments evaluate the contribution of individual components in GSGEN:

    • Initialization Impact: Replacing Point-E point cloud initialization with an origin-centered Gaussian distribution (DreamFusion-style) leads to severe structural degeneration and collapsed geometry on asymmetric prompts (e.g., streaming engine train, corgi with top hat), confirming that anisotropic point cloud initialization is necessary to break early geometric symmetry.
    • 3D SDS Guidance Impact: Removing 3D SDS guidance (setting λ3D=0\lambda_\text{3D} = 0) while retaining Point-E initialization reintroduces the Janus problem on asymmetric subjects (e.g., dogs, pandas). The coarse 3D SDS guidance rectifies major spatial deviations early in optimization, even when Point-E point clouds are imperfect.
    • Densification Strategy Impact: Using only gradient-based splitting with a low threshold (Tpos=0.0002T_\text{pos} = 0.0002) results in noisy Gaussian explosion from stochastic SDS gradients. Using a high threshold (Tpos=0.02T_\text{pos} = 0.02) without compactness densification produces over-smoothed, blurry surfaces with missing structural regions. Combining Tpos=0.02T_\text{pos} = 0.02 with compactness-based densification achieves complete geometric coverage and crisp surface sharpness.
  11. Knowl 11 — Limitations on Semantic Complexity and Diffusion Prior Bias

    limitation

    GSGEN exhibits two main limitations:

    1. Complex Language and Compositional Prompts: When input prompts involve complex scene logic or multi-object compositions, generation quality deteriorates due to the limited linguistic reasoning capacities of Point-E and the CLIP text encoder in Stable Diffusion.
    2. Extreme 2D Diffusion Viewpoint Bias: While 3D prior guidance mitigates geometric collapse, it does not fully eliminate Janus artifacts when the text prompt triggers extreme canonical viewpoint bias in the pre-trained 2D diffusion model.

Coverage note — No substantial contributed material was omitted. The knowls cover the full GSGEN framework including initialization, two-stage optimization, loss formulations, compactness-based densification, experimental configuration, baseline comparisons, ablations, and stated limitations.

References

  1. 1.Alex, Misha Konstantinov, apolinário, Daria Bakshandaeva, Ksenia Ivanova, Sayak Paul, Will Berman, and Emad. deepfloyd/if, 2023. 1, 3, 8
  2. 2.Mohammadreza Armandpour, Huangjie Zheng, Ali Sadeghian, Amir Sadeghian, and Mingyuan Zhou. Re-imagine the negative prompt algorithm: Transform 2d diffusion into 3d, alleviate janus problem and beyond. arXiv preprint arXiv:2304.04968, 2023. 2, 5
  3. 3.Benjamin Attal, Jia-Bin Huang, Christian Richardt, Michael Zollhoefer, Johannes Kopf, Matthew O’Toole, and Changil Kim. Hyperreel: High-fidelity 6-dof video with rayconditioned sampling. arXiv preprint arXiv:2301.02238, 2023. 3
  4. 4.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, Tero Karras, and Ming-Yu Liu. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. CoRR, abs/2211.01324, 2022. 3
  5. 5.Fan Bao, Chongxuan Li, Jun Zhu, and Bo Zhang. Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. 3
  6. 6.Fan Bao, Shen Nie, Kaiwen Xue, Yue Cao, Chongxuan Li, Hang Su, and Jun Zhu. All are worth words: A vit backbone for diffusion models. In CVPR, 2023. 3
  7. 7.Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5855–5864, 2021. 3
  8. 8.Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman. Zip-nerf: Anti-aliased gridbased neural radiance fields. ICCV, 2023. 3
  9. 9.Angel X. Chang, Thomas A. Funkhouser, Leonidas J. Guibas, Pat Hanrahan, Qi-Xing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. Shapenet: An information-rich 3d model repository. CoRR, abs/1512.03012, 2015. 3
  10. 10.Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. arXiv preprint arXiv:2203.09517, 2022. 3
  11. 11.Hansheng Chen, Jiatao Gu, Anpei Chen, Wei Tian, Zhuowen Tu, Lingjie Liu, and Hao Su. Single-stage diffusion nerf: A unified approach to 3d generation and reconstruction. In ICCV, 2023. 3
  12. 12.Rui Chen, Yongwei Chen, Ningxin Jiao, and Kui Jia. Fantasia3d: Disentangling geometry and appearance for highquality text-to-3d content creation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 2, 3, 5, 6, 7
  13. 13.Xingyu Chen, Qi Zhang, Xiaoyu Li, Yue Chen, Ying Feng, Xuan Wang, and Jue Wang. Hallucinated neural radiance fields in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12943–12952, 2022. 3
  14. 14.Yongwei Chen, Rui Chen, Jiabao Lei, Yabin Zhang, and Kui Jia. TANGO: text-driven photorealistic and robust 3d stylization via lighting decomposition. In NeurIPS, 2022. 3
  15. 15.Yiwen Chen, Chi Zhang, Xiaofeng Yang, Zhongang Cai, Gang Yu, Lei Yang, and Guosheng Lin. It3d: Improved text-to-3d generation with explicit view synthesis, 2023. 3
  16. 16.Yen-Chi Cheng, Hsin-Ying Lee, Sergey Tulyakov, Alexander Schwing, and Liangyan Gui. SDFusion: Multimodal 3d shape completion, reconstruction, and generation. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  17. 17.Qiyu Dai, Yan Zhu, Yiran Geng, Ciyu Ruan, Jiazhao Zhang, and He Wang. Graspnerf: Multiview-based 6-dof grasp detection for transparent and specular objects using generalizable nerf. In IEEE International Conference on Robotics and Automation, ICRA 2023, London, UK, May 29 - June 2, 2023, pages 1757–1763. IEEE, 2023. 3
  18. 18.Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Objaverse-xl: A universe of 10m+ 3d objects. CoRR, abs/2307.05663, 2023. 3
  19. 19.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 13142–13153. IEEE, 2023. 3
  20. 20.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat gans on image synthesis. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 8780–8794, 2021. 3
  21. 21.Yuval Eldar, Michael Lindenbaum, Moshe Porat, and Yehoshua Y. Zeevi. The farthest point strategy for progressive image sampling. IEEE Trans. Image Process., 6(9):1305–1315, 1997. 7
  22. 22.Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quantized diffusion model for text-to-image synthesis. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 10686–10696. IEEE, 2022. 3
  23. 23.Yuan-Chen Guo, Ying-Tian Liu, Ruizhi Shao, Christian Laforte, Vikram Voleti, Guan Luo, Chia-Hao Chen, Zi-Xin Zou, Chen Wang, Yan-Pei Cao, and Song-Hai Zhang. threestudio: A unified framework for 3d content generation. https://github.com/threestudio-project/threestudio, 2023. 2, 7
  24. 24.Anchit Gupta, Wenhan Xiong, Yixin Nie, Ian Jones, and Barlas Oguz. 3dgen: Triplane latent diffusion for textured mesh generation. CoRR, abs/2303.05371, 2023. 3
  25. 25.Peter Hedman, Pratul P. Srinivasan, Ben Mildenhall, Jonathan T. Barron, and Paul Debevec. Baking neural radiance fields for real-time view synthesis. ICCV, 2021. 3
  26. 26.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. CoRR, abs/2207.12598, 2022. 3
  27. 27.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. 3
  28. 28.Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. simple diffusion: End-to-end diffusion for high resolution images. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, pages 13213–13232. PMLR, 2023. 3
  29. 29.Ajay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object generation with dream fields. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 857–866. IEEE, 2022. 3
  30. 30.Heewoo Jun and Alex Nichol. Shap-e: Generating conditional 3d implicit functions. CoRR, abs/2305.02463, 2023. 3
  31. 31.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4), 2023. 2, 3, 4, 5
  32. 32.Nasir Mohammad Khalid, Tianhao Xie, Eugene Belilovsky, and Popa Tiberiu. Clip-mesh: Generating textured meshes from text using pretrained image-text models. SIGGRAPH Asia 2022 Conference Papers, 2022. 3
  33. 33.Weiyu Li, Rui Chen, Xuelin Chen, and Ping Tan. Sweetdreamer: Aligning geometric priors in 2d diffusion for consistent text-to-3d. arxiv:2310.02596, 2023. 3
  34. 34.Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3, 5, 6, 7
  35. 35.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. https://arxiv.org/abs/2303.11328, 2023. 3
  36. 36.Xingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng, and Qiang Liu. Instaflow: One step is enough for high-quality diffusion-based text-to-image generation. arXiv preprint arXiv:2309.06380, 2023. 3
  37. 37.Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Learning to generate multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023. 3
  38. 38.Jonathan Lorraine, Kevin Xie, Xiaohui Zeng, Chen-Hsuan Lin, Towaki Takikawa, Nicholas Sharp, Tsung-Yi Lin, Ming-Yu Liu, Sanja Fidler, and James Lucas. Att3d: Amortized text-to-3d object synthesis. In International Conference on Computer Vision ICCV, 2023. 3
  39. 39.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. arXiv preprint arXiv:2206.00927, 2022. 3
  40. 40.Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duckworth. NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections. In CVPR, 2021. 3
  41. 41.Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or. Latent-nerf for shape-guided generation of 3d shapes and textures. arXiv preprint arXiv:2211.07600, 2022. 2, 3, 6
  42. 42.Oscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim, and Rana Hanocka. Text2mesh: Text-driven neural stylization for meshes. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 13482–13492. IEEE, 2022. 3
  43. 43.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 2, 3
  44. 44.Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph., 41(4):102:1–102:15, 2022. 3
  45. 45.Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generating 3d point clouds from complex prompts. CoRR, abs/2212.08751, 2022. 3, 5, 6
  46. 46.Tolga Özer and Ömer Türkmen. Low-cost ai-based solar panel detection drone design and implementation for solar power systems. Robotic Intelligence and Automation, 43(6):605–624, 2023. 3
  47. 47.Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M Seitz. Hypernerf: A higher-dimensional representation for topologically varying neural radiance fields. arXiv preprint arXiv:2106.13228, 2021. 3
  48. 48.William Peebles and Saining Xie. Scalable diffusion models with transformers. arXiv preprint arXiv:2212.09748, 2022. 3
  49. 49.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. SDXL: improving latent diffusion models for high-resolution image synthesis. CoRR, abs/2307.01952, 2023. 3
  50. 50.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. 2, 3, 4, 5, 6, 7, 8
  51. 51.Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10318–10327, 2021. 3
  52. 52.Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Skorokhodov, Peter Wonka, Sergey Tulyakov, and Bernard Ghanem. Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors. arXiv preprint arXiv:2306.17843, 2023. 3
  53. 53.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, pages 8748–8763. PMLR, 2021. 3
  54. 54.Amit Raj, Srinivas Kaza, Ben Poole, Michael Niemeyer, Nataniel Ruiz, Ben Mildenhall, Shiran Zada, Kfir Aberman, Michael Rubinstein, Jonathan Barron, Yuanzhen Li, and Varun Jampan. DreamBooth3D: Subject-driven text-to-3d generation. In International Conference on Computer Vision ICCV, 2023. 3
  55. 55.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with CLIP latents. CoRR, abs/2204.06125, 2022. 1, 3
  56. 56.Christian Reiser, Richard Szeliski, Dor Verbin, Pratul P. Srinivasan, Ben Mildenhall, Andreas Geiger, Jonathan T. Barron, and Peter Hedman. Merf: Memory-efficient radiance fields for real-time view synthesis in unbounded scenes. SIGGRAPH, 2023. 3
  57. 57.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 10674–10685. IEEE, 2022. 1, 3, 7
  58. 58.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. In NeurIPS, 2022. 1, 3, 4
  59. 59.Aditya Sanghi, Hang Chu, Joseph G Lambourne, Ye Wang, Chin-Yi Cheng, and Marco Fumero. Clip-forge: Towards zero-shot text-to-shape generation. arXiv preprint arXiv:2110.02624, 2021. 3
  60. 60.Sara Fridovich-Keil and Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. In CVPR, 2023. 3
  61. 61.Junyoung Seo, Wooseok Jang, Min-Seop Kwak, Jaehoon Ko, Hyeonsu Kim, Junho Kim, Jin-Hwa Kim, Jiyoung Lee, and Seungryong Kim. Let 2d diffusion model know 3dconsistency for robust text-to-3d generation. arXiv preprint arXiv:2303.07937, 2023. 2, 5
  62. 62.Tianchang Shen, Jun Gao, Kangxue Yin, Ming-Yu Liu, and Sanja Fidler. Deep marching tetrahedra: a hybrid representation for high-resolution 3d shape synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2021. 2
  63. 63.Yichun Shi, Peng Wang, Jianglong Ye, Long Mai, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv:2308.16512, 2023. 3, 8
  64. 64.Jaehyeok Shim, Changwoo Kang, and Kyungdon Joo. Diffusion-based signed distance fields for 3d shape generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 20887–20897. IEEE, 2023. 3
  65. 65.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. 3
  66. 66.Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. 3
  67. 67.Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5459–5469, 2022. 3
  68. 68.Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul P Srinivasan, Jonathan T Barron, and Henrik Kretzschmar. Block-nerf: Scalable large scene neural view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8248–8258, 2022. 3
  69. 69.Jiaxiang Tang. Stable-dreamfusion: Text-to-3d with stable-diffusion, 2022. https://github.com/ashawkey/stable-dreamfusion. 2, 7
  70. 70.Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. Dreamgaussian: Generative gaussian splatting for efficient 3d content creation. arXiv preprint arXiv:2309.16653, 2023. 3
  71. 71.Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, and Thomas Wolf. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers, 2022. 7
  72. 72.Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. Clip-nerf: Text-and-image driven manipulation of neural radiance fields. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 3825–3834. IEEE, 2022. 3
  73. 73.Feng Wang, Sinan Tan, Xinghang Li, Zeyue Tian, and Huaping Liu. Mixed neural voxels for fast multi-view video synthesis. arXiv preprint arXiv:2212.00190, 2022. 3
  74. 74.Haochen Wang, Xiaodan Du, Jiahao Li, Raymond A. Yeh, and Greg Shakhnarovich. Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 12619–12629. IEEE, 2023. 3
  75. 75.Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, and Baining Guo. RODIN: A generative model for sculpting 3d digital avatars using diffusion. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 4563–4573. IEEE, 2023. 3
  76. 76.Zhongshu Wang, Lingzhi Li, Zhen Shen, Li Shen, and Liefeng Bo. 4k-nerf: High fidelity neural radiance fields at ultra high resolutions. arXiv preprint arXiv:2212.04701, 2022. 3
  77. 77.Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation. arXiv preprint arXiv:2305.16213, 2023. 2, 3, 6, 7
  78. 78.Jianglong Ye, Jiashun Wang, Binghao Huang, Yuzhe Qin, and Xiaolong Wang. Learning continuous grasping function with a dexterous hand from human demonstrations. IEEE Robotics and Automation Letters, 8(5):2882–2889, 2023. 3
  79. 79.Alex Yu, Sara Fridovich-Keil, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. arXiv preprint arXiv:2112.05131, 2021. 3
  80. 80.Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. Plenoctrees for real-time rendering of neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5752–5761, 2021. 3
  81. 81.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020. 3
  82. 82.Xin-Yang Zheng, Hao Pan, Peng-Shuai Wang, Xin Tong, Yang Liu, and Heung-Yeung Shum. Locally attentional sdf diffusion for controllable 3d shape generation. ACM Transactions on Graphics (SIGGRAPH), 42(4), 2023. 3
  83. 83.Shaohong Zhong, Alessandro Albini, Oiwi Parker Jones, Perla Maiolino, and Ingmar Posner. Touching a nerf: Leveraging neural radiance fields for tactile sensory data generation. In Conference on Robot Learning, pages 1618–1628. PMLR, 2023. 3
  84. 84.Yang Zhou, Long Wang, Yongbin Lai, and Xiaolong Wang. The general method of the tanker car mouth pose measurement. Robotic Intelligence and Automation, 43(6):625–636, 2023. 3
  85. 85.Junzhe Zhu and Peiye Zhuang. Hifa: High-fidelity text-to-3d with advanced diffusion guidance. CoRR, abs/2305.18766, 2023. 2, 3
  86. 86.Matthias Zwicker, Hanspeter Pfister, Jeroen van Baar, and Markus H. Gross. EWA volume splatting. In 12th IEEE Visualization Conference, IEEE Vis 2001, San Diego, CA, USA, October 24-26, 2001, Proceedings, pages 29–36. IEEE Computer Society, 2001. 4

Citation

MLA
Chen, Z., et al. “Text-to-3D Using Gaussian Splatting”. arXiv, 2023, http://arxiv.org/abs/2309.16585v4.
APA
Chen, Z., Wang, F., Wang, Y., & Liu, H. (2023). Text-to-3D using Gaussian Splatting. arXiv. http://arxiv.org/abs/2309.16585v4
Chicago
Chen, Z., F. Wang, Y. Wang, and H. Liu. 2023. “Text-to-3D Using Gaussian Splatting”. arXiv. http://arxiv.org/abs/2309.16585v4.
Harvard
Chen, Z. et al. (2023) “Text-to-3D using Gaussian Splatting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2309.16585v4.
Vancouver
1. Chen Z, Wang F, Wang Y, Liu H (2023) Text-to-3D using Gaussian Splatting. arXiv

BibTeX

@article{chen2023text,
  title = {Text-to-3D using Gaussian Splatting},
  author = {Chen, Zilong and Wang, Feng and Wang, Yikai and Liu, Huaping},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2309.16585v4},
  eprint = {2309.16585}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE