SceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender Code
Ziniu HuAhmet IscenAashi JainThomas KipfYisong YueDavid A. RossCordelia SchmidAlireza Fathi
Introduces SceneCraft, a dual-loop language model agent that converts text prompts into executable Blender Python scripts to arrange complex multi-asset 3D scenes by combining scene-graph spatial planning, multimodal visual feedback refinement, and autonomous library learning.
Converting natural language descriptions into complex 3D environments is a critical capability for gaming, cinematic production, virtual reality, and architectural design. While modern generative tools can create individual 3D objects, automatically assembling complete scenes with dozens or hundreds of items remains difficult due to the intricate spatial, rotational, and semantic relationships required. The article introduces SceneCraft, an autonomous artificial intelligence agent powered by large language models that translates natural language prompts into executable Blender scripts. The primary objective is to demonstrate that this agent can reliably plan spatial layouts, iteratively critique its own visual outputs, and accumulate reusable design functions to generate rich 3D scenes without expensive model fine-tuning.
The authors evaluated SceneCraft using a dual-loop framework that mirrors human studio workflows. When given a text description, the system decomposes the query into manageable sub-scenes, retrieves appropriate 3D assets, and constructs a relational scene graph detailing spatial rules such as proximity, symmetry, and alignment. In an inner optimization loop, the model converts these relationships into numerical constraints, computes optimal object placements, and renders the scene in Blender. A multimodal vision-language model inspects the rendered image against the original description, identifying layout errors and revising the Python script over several cycles. In the outer loop, SceneCraft identifies recurring code solutions across multiple tasks and compiles them into an expanding spatial skill library. The framework was tested on synthetic spatial benchmarks with defined constraints and on cinematic reconstruction using scenes from the open-source movie Sintel.
The findings show that SceneCraft significantly outperforms existing tools across objective and subjective metrics. On synthetic benchmarks, SceneCraft achieved a constraint passing score of 88.9 out of 100, compared to just 5.6 for the baseline BlenderGPT, while improving image-text alignment scores by over 45%. In human preference studies, evaluators favored SceneCraft over the baseline by wide margins, choosing its outputs 83.6% of the time for spatial composition and more than 74% of the time for visual aesthetics and text fidelity. Ablation experiments confirmed that visual self-critique and the learned skill library are critical; removing the visual feedback loop reduced constraint compliance scores by 38.4 points. In cinematic applications, using SceneCraft layouts as intermediate structural guidance for video generation models cut visual distortion scores from 574 to 317 compared to baseline scene guidance.
These results demonstrate that combining high-level symbolic code generation with multimodal visual feedback offers a practical, low-cost solution for complex 3D layout automation. Because SceneCraft updates an external library of Python code rather than updating underlying neural network weights, it achieves continuous self-improvement at an average cost of only a few cents per scene. This hybrid approach reduces manual labor and production timelines in creative industries while providing fine-grained control over individual asset coordinates. Organizations involved in digital content creation should consider piloting code-generating layout agents to accelerate pre-visualization, virtual staging, and video generation workflows. Future development should focus on extending the agent's capabilities to dynamic camera trajectories, variable lighting, and richer material rendering.
While the findings provide high confidence in the agent's spatial reasoning, certain operational boundaries apply. SceneCraft currently uses fixed ground textures and simplified lighting presets, and it can occasionally produce incorrect asset dimensions if initial object files lack standard scale proportions. Additionally, performance declines when attempting to arrange more than twenty assets in a single step, making structured hierarchical scene decomposition essential for very large projects. Stakeholders should treat the system as a robust layout automation and prototyping tool that benefits from human review during final production stages.
- Paper: GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting, Xiaoyu Zhou et al. (2024). GALA3D likewise uses an LLM to derive scene layouts from text, making its layout-guided approach a direct precursor for understanding SceneCraft’s scene-graph planning.
- Paper: DreamFusion: Text-to-3D using 2D Diffusion, Ben Poole et al. (2023). DreamFusion established text-conditioned 3D synthesis through iterative rendering and pretrained diffusion guidance, grounding SceneCraft’s use of rendered feedback in generative 3D workflows.
- Paper: Magic3D: High-Resolution Text-to-3D Content Creation, Chen-Hsuan Lin et al. (2022). Magic3D develops coarse-to-fine text-to-3D generation with differentiable rendering, useful context for SceneCraft’s script-driven scene creation and visual refinement.
- Paper: Scene Graph Generation by Iterative Message Passing, Danfei Xu et al. (2017). This work formalizes scene graphs as object-and-relation structures, clarifying the structured representation SceneCraft uses as a blueprint for spatially arranging assets.
- Paper: Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models, Ruiyu Wang et al. (2025). CADFusion carries SceneCraft’s code-and-render feedback loop into parametric design, using visual evaluation to improve generated executable geometry.
- Paper: CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates, Shresth Grover et al. (2025). CoSPlan extends scene-graph-based planning with incremental corrections, offering a later example of structured visual plans that are revised as actions unfold.
- Paper: Structured 3D Latents for Scalable and Versatile 3D Generation, Jianfeng Xiang et al. (2025). TRELLIS broadens text-driven 3D generation beyond SceneCraft’s Blender scripts by using structured latents to generate assets across multiple 3D representations.
