GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting

Xiaoyu ZhouXingjian RanYajiao XiongJinlin HeZhiwei LinYongtao WangDeqing SunMing-Hsuan Yang

article2024ICML110 citations

Proposes a framework that combines large language model layout priors with 3D Gaussian Splatting to generate high-fidelity multi-object 3D scenes from text while supporting conversational interactive editing.

Listen

Creating complex 3D virtual environments has traditionally required labor-intensive manual work by specialized artists, limiting scalability and making fast iteration difficult. While recent artificial intelligence systems can generate individual 3D objects from text descriptions, expanding this capability to multi-object scenes remains challenging. Existing methods often suffer from severe geometric distortions, blurry textures, spatial inconsistencies across viewpoints, or a requirement for time-consuming manual 3D layout design.

The article demonstrates and evaluates GALA3D, an end-to-end generative framework designed to produce high-fidelity, complex 3D scenes containing multiple interacting objects directly from natural language prompts, while also supporting conversational scene editing.

To solve spatial and geometric ambiguities, the framework first uses a large language model to automatically extract object relationships from text prompts and propose an initial coarse spatial layout. It then translates these layouts into a layout-guided 3D Gaussian Splatting representation—a method that models 3D space using collections of parameterized 3D ellipsoids. The authors introduce adaptive geometric controls to constrain the shape and density distribution of these Gaussians, coupled with a dual-level optimization strategy that uses multi-view 2D image diffusion models. This optimization simultaneously refines individual objects, enhances global scene interactions, and iteratively adjusts the language model's coarse layout bounding boxes to correct spatial misalignments.

The evaluation demonstrates clear advantages over state-of-the-art alternative approaches. In automated benchmark evaluations across 22 multi-object scenes ranging from 1 to 10 instances, GALA3D achieved an average text-alignment score of 34.57, outperforming leading neural radiance field and Gaussian-based alternatives which averaged between 25.12 and 31.17. In a human evaluation study of 125 participants—nearly 40% of whom were professional artists and 3D modelers—GALA3D secured the highest ratings across all categories, including scene quality (8.42 out of 10), geometric fidelity (8.37), text alignment (8.55), and scene consistency (9.68), compared to competitor averages ranging from roughly 4.18 to 7.12. Ablation studies confirmed that removing adaptive geometry controls or global compositional optimization caused substantial drops in scene realism and visual coherence.

These findings indicate that combining automated layout interpretation with layout-guided Gaussian splatting provides a viable path toward automating end-to-end 3D content creation. For design studios, game developers, and interactive media pipelines, this approach substantially reduces manual 3D asset layout time and production costs while preserving high visual quality. The framework also enables conversational editing, allowing users to move, add, delete, or re-style individual objects in localized areas without corrupting the broader scene geometry.

Organizations evaluating generative 3D tools should consider adopting layout-guided Gaussian representations over traditional monolithic representations for multi-object scenes. Technical teams planning implementation should begin with pilot deployments targeting iterative prototyping and conversational scene customization. Additionally, stakeholders must establish content governance policies to manage risks associated with automated asset generation and malicious synthetic media creation.

Confidence in the reported improvements is high based on consistent quantitative scores, expert user evaluations, and structural ablation tests. However, users should note that the system requires substantial computational resources (evaluations were conducted on an 80 GB GPU) and still relies on upstream 2D diffusion priors, meaning visual quality remains bounded by the capabilities of underlying foundation models.

Cover for GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting

Abstract

We present GALA3D, generative 3D Gaussians with LAyout-guided control, for effective compositional text-to-3D generation. We first utilize large language models (LLMs) to generate the initial layout and introduce a layout-guided 3D Gaussian representation for 3D content generation with adaptive geometric constraints. We

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Layout-guided Gaussian Representation
  • 3.2. Adaptive Geometry Control for Gaussians
  • 3.3. Compositional Optimization with Diffusion Priors
  • 3.4. Layout Refinement
  • 3.5. Total Loss
  • 4. Experimental Results
  • 4.1. Quantitative Comparison
  • 4.2. Qualitative Comparison
  • 4.3. User Study
  • 4.4. Conversational Interactive Editing
  • 4.5. Ablation Studies
  • 5. Conclusion
  • Acknowledgment
  • Impact Statement
  • References

Knowls

  1. Knowl 1 — End-to-end layout-guided scene optimization

    model/method

    GALA3D generates a 3D scene from a text description by asking a large language model (LLM) to extract object instances and their coarse spatial layouts, then optimizing a separate layout-guided Gaussian representation for each instance together with the scene as a whole. Its objective combines instance-level score distillation, layout constraints, layout refinement, scene-level conditioned diffusion, and Gaussian-shape regularization:

    L=∑i=1N(β1LSDS(i)+β2Llayout(i)+β3LDef(i))+β4Lglobal+β5Lreg.\mathcal{L}=\sum_{i=1}^{N}\left(\beta_1\mathcal{L}_{\mathrm{SDS}}^{(i)}+\beta_2\mathcal{L}_{\mathrm{layout}}^{(i)}+\beta_3\mathcal{L}_{\mathrm{Def}}^{(i)}\right)+\beta_4\mathcal{L}_{\mathrm{global}}+\beta_5\mathcal{L}_{\mathrm{reg}}.

    Here, NN is the number of scene instances; the five losses respectively supervise each instance’s diffusion-based appearance, its spatial relationship to its layout, refinement of that layout, diffusion-based optimization of the complete scene, and Gaussian geometry. The reported weights are β1=1\beta_1=1, β2=103\beta_2=10^3, β3=10−1\beta_3=10^{-1}, β4=10−1\beta_4=10^{-1}, and β5=103\beta_5=10^3. Joint optimization lets the scene depart from inaccurate initial LLM layouts while retaining instance-level generation and global interaction constraints.

  2. Knowl 2 — Layout-guided Gaussian scene representation

    model/method

    GALA3D represents a scene as a collection of per-instance Gaussian sets, each attached to a layout defined by a center (xi,yi,zi)(x_i,y_i,z_i), extents (hi,wi,li)(h_i,w_i,l_i) along the x,y,zx,y,z axes, a scale factor kik_i, and a rotation angle ϕi\phi_i about the vertical zz axis. The iith layout and its instance Gaussians are denoted LiL_i and GiG_i, respectively; NN is the number of instances. Each instance Gaussian has a 3D center pip_i, color cic_i, opacity αi\alpha_i, and anisotropic covariance Σobj=RSS⊤R⊤\Sigma_{\mathrm{obj}}=R S S^{\top}R^{\top}, where RR and SS are its rotation and scale matrices.

    To place an instance in the common scene coordinate frame, GALA3D transforms a Gaussian center and covariance as

    pscene=kiRz(ϕi)pi+(xi,yi,zi)⊤,Σscene=ki2Rz(ϕi)ΣobjRz(ϕi)⊤,p_{\mathrm{scene}}=k_iR_z(\phi_i)p_i+(x_i,y_i,z_i)^{\top},\qquad \Sigma_{\mathrm{scene}}=k_i^2R_z(\phi_i)\Sigma_{\mathrm{obj}}R_z(\phi_i)^{\top},

    where Rz(ϕi)R_z(\phi_i) is the rotation matrix for the layout angle. This representation makes object-level Gaussian geometry independently optimizable while allowing layout scale, position, and orientation to control its placement in the composed scene.

  3. Knowl 3 — Adaptive geometric control of Gaussians

    model/method

    GALA3D’s Adaptive Geometry Control (AGC) regulates both where instance Gaussians occur and the shapes of their ellipsoids, instead of relying on standard 3D Gaussian Splatting densification alone. For a Gaussian center pip_i and its layout center ζi=(xi,yi,zi)\zeta_i=(x_i,y_i,z_i), AGC models the reciprocal distance 1/∥pi−ζi∥1/\lVert p_i-\zeta_i\rVert with a truncated folded-normal distribution. The truncation is bounded by the layout center and its boundary; the distribution mean and standard deviation are adjustable, and samples are used to place Gaussians near the layout surface.

    AGC also regularizes Gaussian shape and scale so that overly elongated ellipsoids are compressed and the resulting Gaussian geometry is more regular. The paper reports that this adaptive control produces better-aligned surface Gaussians and improves geometric detail and texture. In the reported implementation, each instance starts with 100,000 Gaussian particles and the standard adaptive density-control procedure is disabled.

  4. Knowl 4 — Two-scale diffusion optimization for objects and scenes

    model/method

    GALA3D combines instance-level and scene-level diffusion supervision. First, it optimizes each instance’s layout-guided Gaussians using MVDream as a multi-view diffusion prior with score distillation sampling (SDS); virtual camera poses are sampled across a full 360∘360^\circ horizontal range, with camera radius set to 34∥(hi,wi,li)∥2\tfrac{3}{4}\lVert(h_i,w_i,l_i)\rVert_2 for an instance whose layout extents are (hi,wi,li)(h_i,w_i,l_i). Second, it renders the full multi-object scene and optimizes its Gaussians using a scene-conditioned diffusion prior: a fine-tuned ControlNet receives rendered 2D layouts from multiple viewpoints and provides text-conditioned scene supervision. The instance and global stages share the same diffusion timestep during optimization, so object-level appearance and scene-level interactions are updated synchronously.

    The reported guidance scales are 50 for MVDream and 100 for ControlNet. The global stage is intended to improve scene-wide coherence and object interactions beyond what independently generated instances and their layouts provide.

  5. Knowl 5 — Layout-based spatial constraint

    equation

    For the iith instance, GALA3D defines a layout loss on Gaussian centers p=(px,py,pz)p=(p_x,p_y,p_z) relative to the instance’s axis-aligned box. The box has center (xi,yi,zi)(x_i,y_i,z_i) and extents (hi,wi,li)(h_i,w_i,l_i) along the x,y,zx,y,z axes. The loss is

    Llayout(i)=1bbox(p)[dx(px,xi,hi)+dy(py,yi,wi)+dz(pz,zi,li)],\mathcal{L}_{\mathrm{layout}}^{(i)}=\mathbf{1}_{\mathrm{bbox}}(p)\left[d^x(p_x,x_i,h_i)+d^y(p_y,y_i,w_i)+d^z(p_z,z_i,l_i)\right],

    where 1bbox(p)\mathbf{1}_{\mathrm{bbox}}(p) is 1 when pp is inside the box and 0 otherwise. For an axis with coordinate pxp_x, center xix_i, and extent hih_i, the distance term is

    dx(px,xi,hi)=min⁡(∣px−(xi+hi/2)∣,∣px−(xi−hi/2)∣),d^x(p_x,x_i,h_i)=\min\left(\left|p_x-(x_i+h_i/2)\right|,\left|p_x-(x_i-h_i/2)\right|\right),

    with the yy and zz terms defined analogously using (yi,wi)(y_i,w_i) and (zi,li)(z_i,l_i). The paper uses this loss to supervise the generated instances’ spatial placement, scale, and geometric consistency with their layout priors.

  6. Knowl 6 — Diffusion-driven refinement of LLM layouts

    model/method

    Because an LLM’s initial scene layout can misplace objects or assign implausible sizes, GALA3D treats layout parameters as learnable during scene generation. For instance ii, the learnable parameters include its layout center ζi=(xi,yi,zi)\zeta_i=(x_i,y_i,z_i), opacity αi\alpha_i, scale factor kik_i, and rotation ϕi\phi_i. The method updates these parameters using a score-distillation gradient from the scene-conditioned diffusion prior, which is conditioned on the full text prompt and rendered layout images. Refinement therefore adjusts the initial LLM interpretation in response to the generated scene, rather than treating the coarse layout as fixed.

  7. Knowl 7 — Conversational editing of generated scenes

    algorithm

    GALA3D supports text-based editing of an existing scene. An LLM translates a user instruction into layout operations, such as adding or deleting an object, changing an object’s position or rotation, or replacing an object. The system then optimizes the layout-guided Gaussian representation in the affected local layout regions while seeking to preserve the rest of the scene. The paper demonstrates this workflow for object addition and removal, spatial rearrangement and rotation, object replacement, style changes, and restoration of an earlier scene state.

  8. Knowl 8 — CLIP comparison across text-to-3D methods

    data/table

    The paper compares zero-shot text-to-3D scene generations using CLIP Score, because these tasks lack ground-truth 3D scenes. The reported average is over 22 scenes containing between 1 and 10 objects; the four individual cases contain 1, 3, 7, and 10 instances, respectively. Higher scores indicate stronger text-image alignment. The comparison shows that GALA3D has the highest reported average and score in each of the four cases. The Set-the-scene baseline uses text plus layout input; the other listed methods use text input.

    Method Average Case 1 Case 2 Case 3 Case 4
    Latent-NeRF 27.772 22.135 27.482 22.203 19.606
    ProlificDreamer 28.401 30.237 21.913 19.219 25.587
    MVDream 30.856 28.756 32.636 26.015 27.417
    SJC 28.775 29.100 31.764 21.154 26.352
    DreamGaussian 25.117 26.281 23.051 18.595 25.739
    GaussianDreamer 28.351 29.469 31.237 25.727 24.143
    GSGEN 30.293 28.932 29.578 29.959 23.927
    LucidDreamer 31.174 28.720 26.533 27.768 26.895
    Set-the-scene 29.628 28.129 19.135 29.003 25.899
    GALA3D 34.573 31.637 37.658 31.459 35.052
  9. Knowl 9 — Human preference evaluation

    empirical result

    A user study compared GALA3D with six text-to-3D methods using generations for 8 descriptions. The 125 participants rated scene quality, geometric fidelity, text alignment, and scene consistency on a 1–10 scale, where higher scores indicate stronger preference; 39.2% of participants were professionals in art design or 3D modeling. GALA3D received the highest mean rating on all four dimensions, with its largest score on scene consistency.

    Method Scene Quality Geometric Fidelity Text Alignment Scene Consistency
    SJC 5.98 5.04 6.76 4.61
    DreamGaussian 5.22 4.18 4.30 5.46
    GaussianDreamer 6.09 5.71 5.23 4.37
    GSGEN 6.54 4.23 5.41 6.25
    LucidDreamer 4.78 5.62 5.03 4.77
    Set-the-scene 6.36 5.03 7.12 6.12
    GALA3D 8.42 8.37 8.55 9.68
  10. Knowl 10 — Ablation evidence for GALA3D components

    empirical result

    The paper reports CLIP Score ablations to assess Adaptive Geometry Control (AGC), the compositional optimization scheme (COS), the layout loss, the Layout Refinement Module (LRM), and global scene diffusion. Removing AGC or COS causes the largest reported declines relative to the full model, while removing each layout or global supervision also lowers the score. The full ablation configuration scores 34.885.

    Model variant CLIP Score
    w/o AGC 32.198
    w/o COS 32.213
    w/o Llayout\mathcal{L}_{\mathrm{layout}} 33.297
    w/o LRM 34.293
    w/o Lglobal\mathcal{L}_{\mathrm{global}} 34.342
    Ours-Full 34.885

    The visual ablations also report that omitting AGC leads to artifacts and blur, omitting layout refinement leaves objects misaligned with the scene, and omitting global scene optimization produces weaker texture and coherence and can create over-constrained boundaries.

Coverage note — No substantial contributed component was omitted; routine Gaussian projection/compositing equations and qualitative image-by-image comparisons were excluded because they add no separate method or quantitative result beyond the knowls above.

References

  1. 1.Chang, A., Monroe, W., Savva, M., Potts, C., and Manning, C. D. Text to 3d scene generation with rich lexical grounding. arXiv preprint arXiv:1505.06289, 2015.
  2. 2.Chen, R., Chen, Y., Jiao, N., and Jia, K. Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation. In ICCV, 2023a.
  3. 3.Chen, Z., Wang, F., and Liu, H. Text-to-3d using gaussian splatting. arXiv preprint arXiv:2309.16585, 2023b.
  4. 4.Cohen-Bar, D., Richardson, E., Metzer, G., Giryes, R., and Cohen-Or, D. Set-the-scene: Global-local training for generating controllable nerf scenes. arXiv preprint arXiv:2303.13450, 2023.
  5. 5.de Queiroz, R. L. and Chou, P. A. Compression of 3d point clouds using a region-adaptive hierarchical transform. IEEE Transactions on Image Processing, 25(8):3947–3956, 2016.
  6. 6.Fang, J., Wang, J., Zhang, X., Xie, L., and Tian, Q. Gaussianeditor: Editing 3d gaussians delicately with text instructions. arXiv preprint arXiv:2311.16037, 2023.
  7. 7.Feng, W., Zhu, W., Fu, T.-j., Jampani, V., Akula, A., He, X., Basu, S., Wang, X. E., and Wang, W. Y. Layoutgpt: Compositional visual planning and generation with large language models. NIPS, 36, 2024.
  8. 8.He, Y., Bai, Y., Lin, M., Zhao, W., Hu, Y., Sheng, J., Yi, R., Li, J., and Liu, Y.-J. T3bench: Benchmarking current progress in text-to-3d generation. arXiv preprint arXiv:2310.02977, 2023.
  9. 9.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  10. 10.Huang, Y., Wang, J., Shi, Y., Qi, X., Zha, Z.-J., and Zhang, L. Dreamtime: An improved optimization strategy for text-to-3d content creation. arXiv preprint arXiv:2306.12422, 2023.
  11. 11.Jain, A., Mildenhall, B., Barron, J. T., Abbeel, P., and Poole, B. Zero-shot text-guided object generation with dream fields. In CVPR, pp. 867–876, 2022.
  12. 12.Kerbl, B., Kopanas, G., Leimkuhler, T., and Drettakis, G. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4), 2023.
  13. 13.Li, X., Wang, H., and Tseng, K.-K. Gaussiandiffusion: 3d gaussian splatting for denoising diffusion probabilistic models with structured noise. arXiv preprint arXiv:2311.11221, 2023.
  14. 14.Liang, Y., Yang, X., Lin, J., Li, H., Xu, X., and Chen, Y. Luciddreamer: Towards high-fidelity text-to-3d generation via interval score matching. arXiv preprint arXiv:2311.11284, 2023.
  15. 15.Lin, C.-H., Gao, J., Tang, L., Takikawa, T., Zeng, X., Huang, X., Kreis, K., Fidler, S., Liu, M.-Y., and Lin, T.-Y. Magic3d: High-resolution text-to-3d content creation. In CVPR, 2023a.
  16. 16.Lin, Y., Bai, H., Li, S., Lu, H., Lin, X., Xiong, H., and Wang, L. Componerf: Text-guided multi-object compositional nerf with editable 3d scene layout. arXiv preprint arXiv:2303.13843, 2023b.
  17. 17.Lin, Y., Wu, H., Wang, R., Lu, H., Lin, X., Xiong, H., and Wang, L. Towards language-guided interactive 3d generation: Llms as layout interpreter with generative feedback. arXiv preprint arXiv:2305.15808, 2023c.
  18. 18.Liu, X., Zhan, X., Tang, J., Shan, Y., Zeng, G., Lin, D., Liu, X., and Liu, Z. Humangaussian: Text-driven 3d human generation with gaussian splatting. arXiv preprint arXiv:2311.17061, 2023.
  19. 19.Low, W. F. and Lee, G. H. Robust e-nerf: Nerf from sparse & noisy events under non-uniform motion. In ICCV, pp. 18335–18346, 2023.
  20. 20.Metzer, G., Richardson, E., Patashnik, O., Giryes, R., and Cohen-Or, D. Latent-nerf for shape-guided generation of 3d shapes and textures. In CVPR, pp. 12663–12673, 2023.
  21. 21.Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
  22. 22.Po, R. and Wetzstein, G. Compositional 3d scene generation using locally conditioned diffusion. arXiv preprint arXiv:2303.12218, 2023.
  23. 23.Poole, B., Jain, A., Barron, J. T., and Mildenhall, B. Dreamfusion: Text-to-3d using 2d diffusion. arXiv, 2022.
  24. 24.Raj, A., Kaza, S., Poole, B., Niemeyer, M., Ruiz, N., Mildenhall, B., Zada, S., Aberman, K., Rubinstein, M., Barron, J., et al. Dreambooth3d: Subject-driven text-to-3d generation. arXiv preprint arXiv:2303.13508, 2023.
  25. 25.Ren, J., He, C., Liu, L., Chen, J., Wang, Y., Song, Y., Li, J., Xue, T., Hu, S., Chen, T., et al. Make-a-character: High quality text-to-3d character generation within minutes. arXiv preprint arXiv:2312.15430, 2023.
  26. 26.Shi, Y., Wang, P., Ye, J., Long, M., Li, K., and Yang, X. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023.
  27. 27.Sun, C., Han, J., Deng, W., Wang, X., Qin, Z., and Gould, S. 3d-gpt: Procedural 3d modeling with large language models. arXiv preprint arXiv:2310.12945, 2023.
  28. 28.Tang, J., Ren, J., Zhou, H., Liu, Z., and Zeng, G. Dreamgaussian: Generative gaussian splatting for efficient 3d content creation. arXiv preprint arXiv:2309.16653, 2023.
  29. 29.Vilesov, A., Chari, P., and Kadambi, A. Cg3d: Compositional generation for text-to-3d via gaussian splatting. arXiv preprint arXiv:2311.17907, 2023.
  30. 30.Wang, H., Du, X., Li, J., Yeh, R. A., and Shakhnarovich, G. Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation. In CVPR, pp. 12619–12629, 2023a.
  31. 31.Wang, Z., Lu, C., Wang, Y., Bao, F., Li, C., Su, H., and Zhu, J. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation. arXiv preprint arXiv:2305.16213, 2023b.
  32. 32.Wen, Z., Liu, Z., Sridhar, S., and Fu, R. Anyhome: Open-vocabulary generation of structured and textured 3d homes. arXiv preprint arXiv:2312.06644, 2023.
  33. 33.Xu, J., Wang, X., Cheng, W., Cao, Y.-P., Shan, Y., Qie, X., and Gao, S. Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20908–20918, 2023.
  34. 34.Yang, Y., Sun, F.-Y., Weihs, L., VanderBilt, E., Herrasti, A., Han, W., Wu, J., Haber, N., Krishna, R., Liu, L., et al. Holodeck: Language guided generation of 3d embodied ai environments. In CVPR, volume 30, pp. 20–25, 2024.
  35. 35.Yi, T., Fang, J., Wu, G., Xie, L., Zhang, X., Liu, W., Tian, Q., and Wang, X. Gaussiandreamer: Fast generation from text to 3d gaussian splatting with point cloud priors. arXiv preprint arXiv:2310.08529, 2023.
  36. 36.Zhang, L., Rao, A., and Agrawala, M. Adding conditional control to text-to-image diffusion models. In ICCV, pp. 3836–3847, 2023a.
  37. 37.Zhang, Q., Wang, C., Siarohin, A., Zhuang, P., Xu, Y., Yang, C., Lin, D., Zhou, B., Tulyakov, S., and Lee, H.-Y. Scenewiz3d: Towards text-guided 3d scene composition. arXiv preprint arXiv:2312.08885, 2023b.

Citation

MLA
Zhou, X., et al. “GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting”. arXiv, 2024, http://arxiv.org/abs/2402.07207v2.
APA
Zhou, X., Ran, X., Xiong, Y., He, J., Lin, Z., Wang, Y., Sun, D., & Yang, M.-H. (2024). GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting. arXiv. http://arxiv.org/abs/2402.07207v2
Chicago
Zhou, X., X. Ran, Y. Xiong, et al. 2024. “GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting”. arXiv. http://arxiv.org/abs/2402.07207v2.
Harvard
Zhou, X. et al. (2024) “GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.07207v2.
Vancouver
1. Zhou X, Ran X, Xiong Y, He J, Lin Z, Wang Y, Sun D, Yang M-H (2024) GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting. arXiv

BibTeX

@article{zhou2024gala3d,
  title = {GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting},
  author = {Zhou, Xiaoyu and Ran, Xingjian and Xiong, Yajiao and He, Jinlin and Lin, Zhiwei and Wang, Yongtao and Sun, Deqing and Yang, Ming-Hsuan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.07207v2},
  eprint = {2402.07207}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/