Text-to-3D using Gaussian Splatting
Zilong ChenFeng WangYikai WangHuaping Liu
Presents GSGEN, a text-to-3D generation framework that pairs 3D Gaussian Splatting with joint 2D and 3D diffusion priors to resolve the multi-face Janus problem and synthesize detailed, geometrically consistent 3D assets.
Automated generation of three-dimensional digital assets directly from text prompts has become a critical capability for interactive media, gaming, simulation, and design. However, existing methods relying on two-dimensional image diffusion models and implicit neural volumetric representations frequently suffer from severe structural defects, such as the multi-face "Janus problem" where objects generate multiple front sides or collapsed geometry. Additionally, these approaches struggle to capture intricate, high-frequency surface details and typically require extensive rendering and optimization time due to their implicit mathematical formulations.
The article introduces and evaluates GSGEN, a novel framework designed to synthesize high-fidelity, geometrically accurate 3D objects from text prompts. It demonstrates how leveraging 3D Gaussian Splatting—an explicit point-like representation—enables direct geometric guidance and superior rendering quality compared to conventional implicit techniques.
To achieve this, the approach employs a two-stage optimization process combined with geometric initialization. The system initializes Gaussian positions using a 3D point cloud diffusion model (Point-E) to establish a rough anisotropic structure. In the first stage, geometry optimization jointly applies 2D image score distillation and 3D point cloud score distillation to enforce 3D structural consistency. In the second stage, appearance refinement optimizes fine-grained textures using only 2D image guidance while applying a custom "compactness-based densification" strategy. This technique inserts new Gaussians between existing neighboring points to close structural gaps and enhance surface continuity without destabilizing optimization.
The findings demonstrate four main outcomes. First, integrating explicit 3D point cloud priors with 2D image priors successfully mitigates the Janus problem, establishing coherent geometry even on complex asymmetric prompts like animals and vehicles. Second, the explicit Gaussian Splatting representation significantly outperforms mesh- and implicit-based baselines in rendering fine, high-frequency textures such as fur, feathers, and patterned surfaces. Third, the compactness-based densification resolves optimization instability under score distillation sampling, preventing the over-smoothing caused by high gradient thresholds and the uncontrolled point explosion caused by low thresholds. Fourth, the framework achieves this high visual fidelity in approximately 40 minutes per asset, matching the runtime of standard mesh-based methods while delivering noticeably sharper geometric and textural quality.
These results show that transitioning from implicit coordinate networks to explicit Gaussian representations unlocks substantial performance and visual quality improvements in automated asset creation. For organizations developing 3D pipelines, this approach reduces the risk of generating physically implausible models, shortens design iterations, and eliminates costly manual touch-ups needed to repair multi-face artifacts.
Teams exploring generative 3D workflows should consider adopting explicit 3D Gaussian representations and dual 2D/3D prior guidance for asset synthesis pipelines. Decision-makers should evaluate pilot implementations where high surface detail is needed, such as character design or detailed props. Future technical exploration should evaluate pairing the framework with more advanced multi-view diffusion models and specialized large language models to broaden prompt complexity.
Confidence in these findings is high for individual object synthesis, as qualitative and ablation experiments consistently validate the architectural choices. However, limitations remain: the framework struggles with prompts requiring complex logic or dense multi-object scene descriptions, bounded by the language comprehension limits of the underlying text encoders. Furthermore, if a text prompt triggers extreme bias in the guidance diffusion models, geometric artifacts can still occasionally occur.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). Introduces the 3D Gaussian Splatting representation and differentiable rasterization pipeline that GSGEN directly adopts and adapts for text-guided 3D generation.
- Paper: DreamFusion: Text-to-3D using 2D Diffusion, Ben Poole et al. (2023). Establishes the Score Distillation Sampling (SDS) objective using 2D diffusion models, which forms the core text-guided optimization loss used by GSGEN.
- Paper: ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation, Zhengyi Wang et al. (2023). Analyzes and improves upon score distillation techniques for text-to-3D synthesis, providing direct context for addressing geometric issues and high-fidelity rendering.
- Paper: Magic3D: High-Resolution Text-to-3D Content Creation, Chen-Hsuan Lin et al. (2022). Pioneers the two-stage coarse-to-fine optimization framework for text-to-3D generation that motivates GSGEN's progressive geometry and appearance refinement strategy.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). Provides foundational methods for leveraging diffusion models with 3D viewpoint control to guide 3D generation and alleviate multi-view inconsistency.
- Paper: Learning Representations and Generative Models for 3D Point Clouds, Panos Achlioptas et al. (2017). Establishes fundamental 3D point cloud generative modeling techniques that underlie point-based 3D priors incorporated into GSGEN's coarse geometry stage.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). Advances explicit Gaussian rendering from 3D volumetric ellipsoids to geometrically aligned 2D surfels, directly refining the geometric surface accuracy challenges tackled by 3D Gaussian generation methods.
- Paper: Structured 3D Latents for Scalable and Versatile 3D Generation, Jianfeng Xiang et al. (2025). Extends Gaussian- and radiance-based 3D generation to a unified structured latent framework capable of directly decoding into multiple formats including 3D Gaussians.
