GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting
Xiaoyu ZhouXingjian RanYajiao XiongJinlin HeZhiwei LinYongtao WangDeqing SunMing-Hsuan Yang
Proposes a framework that combines large language model layout priors with 3D Gaussian Splatting to generate high-fidelity multi-object 3D scenes from text while supporting conversational interactive editing.
Creating complex 3D virtual environments has traditionally required labor-intensive manual work by specialized artists, limiting scalability and making fast iteration difficult. While recent artificial intelligence systems can generate individual 3D objects from text descriptions, expanding this capability to multi-object scenes remains challenging. Existing methods often suffer from severe geometric distortions, blurry textures, spatial inconsistencies across viewpoints, or a requirement for time-consuming manual 3D layout design.
The article demonstrates and evaluates GALA3D, an end-to-end generative framework designed to produce high-fidelity, complex 3D scenes containing multiple interacting objects directly from natural language prompts, while also supporting conversational scene editing.
To solve spatial and geometric ambiguities, the framework first uses a large language model to automatically extract object relationships from text prompts and propose an initial coarse spatial layout. It then translates these layouts into a layout-guided 3D Gaussian Splatting representation—a method that models 3D space using collections of parameterized 3D ellipsoids. The authors introduce adaptive geometric controls to constrain the shape and density distribution of these Gaussians, coupled with a dual-level optimization strategy that uses multi-view 2D image diffusion models. This optimization simultaneously refines individual objects, enhances global scene interactions, and iteratively adjusts the language model's coarse layout bounding boxes to correct spatial misalignments.
The evaluation demonstrates clear advantages over state-of-the-art alternative approaches. In automated benchmark evaluations across 22 multi-object scenes ranging from 1 to 10 instances, GALA3D achieved an average text-alignment score of 34.57, outperforming leading neural radiance field and Gaussian-based alternatives which averaged between 25.12 and 31.17. In a human evaluation study of 125 participants—nearly 40% of whom were professional artists and 3D modelers—GALA3D secured the highest ratings across all categories, including scene quality (8.42 out of 10), geometric fidelity (8.37), text alignment (8.55), and scene consistency (9.68), compared to competitor averages ranging from roughly 4.18 to 7.12. Ablation studies confirmed that removing adaptive geometry controls or global compositional optimization caused substantial drops in scene realism and visual coherence.
These findings indicate that combining automated layout interpretation with layout-guided Gaussian splatting provides a viable path toward automating end-to-end 3D content creation. For design studios, game developers, and interactive media pipelines, this approach substantially reduces manual 3D asset layout time and production costs while preserving high visual quality. The framework also enables conversational editing, allowing users to move, add, delete, or re-style individual objects in localized areas without corrupting the broader scene geometry.
Organizations evaluating generative 3D tools should consider adopting layout-guided Gaussian representations over traditional monolithic representations for multi-object scenes. Technical teams planning implementation should begin with pilot deployments targeting iterative prototyping and conversational scene customization. Additionally, stakeholders must establish content governance policies to manage risks associated with automated asset generation and malicious synthetic media creation.
Confidence in the reported improvements is high based on consistent quantitative scores, expert user evaluations, and structural ablation tests. However, users should note that the system requires substantial computational resources (evaluations were conducted on an 80 GB GPU) and still relies on upstream 2D diffusion priors, meaning visual quality remains bounded by the capabilities of underlying foundation models.
- Paper: DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis, Yinghao Xu et al. (2023). DisCoScene establishes how layout priors disentangle objects in generated 3D scenes, making its object-centric scene design a direct precursor to GALA3D’s layout-guided composition.
- Paper: Text-to-3D using Gaussian Splatting, Zilong Chen et al. (2024). GSGEN introduces text-conditioned 3D Gaussian asset generation, providing the representation and generation context that GALA3D adapts from objects to laid-out complex scenes.
- Paper: Magic3D: High-Resolution Text-to-3D Content Creation, Chen-Hsuan Lin et al. (2022). Magic3D develops text-to-3D generation through pretrained image diffusion guidance, clarifying the text-conditioned optimization paradigm that GALA3D extends with scene layouts and Gaussians.
- Paper: Diffusion-SDF: Text-to-Shape via Voxelized Diffusion, Muheng Li et al. (2023). Diffusion-SDF shows how text conditioning can guide 3D shape generation, supplying foundational context for GALA3D’s text-driven 3D synthesis.
- Paper: GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields, Michael Niemeyer et al. (2021). GIRAFFE introduces compositional, object-level 3D scene generation, helping explain the scene decomposition and controllability that GALA3D brings to text-to-3D generation.
- Paper: Beyond Voxel 3D Editing: Learning from 3D Masks and Self-Constructed Data, Yizhao Xu et al. (2026). Beyond Voxel extends 3D generation toward instruction-driven editing with spatial masks, carrying layout-aware control into localized modification of generated assets.
- Paper: Map2World: Segment Map Conditioned Text to 3D World Generation, Jaeyoung Chung et al. (2026). Map2World scales layout-conditioned generation from scenes to coherent large environments, continuing GALA3D’s focus on spatial control across 3D compositions.
