SceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender Code

Ziniu HuAhmet IscenAashi JainThomas KipfYisong YueDavid A. RossCordelia SchmidAlireza Fathi

article2024ICML152 citations

Introduces SceneCraft, a dual-loop language model agent that converts text prompts into executable Blender Python scripts to arrange complex multi-asset 3D scenes by combining scene-graph spatial planning, multimodal visual feedback refinement, and autonomous library learning.

Listen

Converting natural language descriptions into complex 3D environments is a critical capability for gaming, cinematic production, virtual reality, and architectural design. While modern generative tools can create individual 3D objects, automatically assembling complete scenes with dozens or hundreds of items remains difficult due to the intricate spatial, rotational, and semantic relationships required. The article introduces SceneCraft, an autonomous artificial intelligence agent powered by large language models that translates natural language prompts into executable Blender scripts. The primary objective is to demonstrate that this agent can reliably plan spatial layouts, iteratively critique its own visual outputs, and accumulate reusable design functions to generate rich 3D scenes without expensive model fine-tuning.

The authors evaluated SceneCraft using a dual-loop framework that mirrors human studio workflows. When given a text description, the system decomposes the query into manageable sub-scenes, retrieves appropriate 3D assets, and constructs a relational scene graph detailing spatial rules such as proximity, symmetry, and alignment. In an inner optimization loop, the model converts these relationships into numerical constraints, computes optimal object placements, and renders the scene in Blender. A multimodal vision-language model inspects the rendered image against the original description, identifying layout errors and revising the Python script over several cycles. In the outer loop, SceneCraft identifies recurring code solutions across multiple tasks and compiles them into an expanding spatial skill library. The framework was tested on synthetic spatial benchmarks with defined constraints and on cinematic reconstruction using scenes from the open-source movie Sintel.

The findings show that SceneCraft significantly outperforms existing tools across objective and subjective metrics. On synthetic benchmarks, SceneCraft achieved a constraint passing score of 88.9 out of 100, compared to just 5.6 for the baseline BlenderGPT, while improving image-text alignment scores by over 45%. In human preference studies, evaluators favored SceneCraft over the baseline by wide margins, choosing its outputs 83.6% of the time for spatial composition and more than 74% of the time for visual aesthetics and text fidelity. Ablation experiments confirmed that visual self-critique and the learned skill library are critical; removing the visual feedback loop reduced constraint compliance scores by 38.4 points. In cinematic applications, using SceneCraft layouts as intermediate structural guidance for video generation models cut visual distortion scores from 574 to 317 compared to baseline scene guidance.

These results demonstrate that combining high-level symbolic code generation with multimodal visual feedback offers a practical, low-cost solution for complex 3D layout automation. Because SceneCraft updates an external library of Python code rather than updating underlying neural network weights, it achieves continuous self-improvement at an average cost of only a few cents per scene. This hybrid approach reduces manual labor and production timelines in creative industries while providing fine-grained control over individual asset coordinates. Organizations involved in digital content creation should consider piloting code-generating layout agents to accelerate pre-visualization, virtual staging, and video generation workflows. Future development should focus on extending the agent's capabilities to dynamic camera trajectories, variable lighting, and richer material rendering.

While the findings provide high confidence in the agent's spatial reasoning, certain operational boundaries apply. SceneCraft currently uses fixed ground textures and simplified lighting presets, and it can occasionally produce incorrect asset dimensions if initial object files lack standard scale proportions. Additionally, performance declines when attempting to arrange more than twenty assets in a single step, making structured hierarchical scene decomposition essential for very large projects. Stakeholders should treat the system as a robust layout automation and prototyping tool that benefits from human review during final production stages.

arXiv: 2403.01248
Cover for SceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender Code

Abstract

This paper introduces SceneCraft, a Large Language Model (LLM) Agent converting text descriptions into Blender-executable Python scripts which render complex scenes with up to a hundred 3D assets. This process requires complex spatial planning and arrangement. We tackle these challenges through a combination of advanced abstraction, strategic planning, and library learning. SceneCraft first models a scene graph as a blueprint, detailing the spatial relationships among assets in the scene. SceneCraft then writes Python scripts based on this graph, translating relationships into numerical constraints for asset layout. Next, SceneCraft leverages the perceptual strengths of vision-language foundation models like GPT-V to analyze rendered images and iteratively refine the scene. On top of this process, SceneCraft features a library learning mechanism that compiles common script functions into a reusable library, facilitating continuous self-improvement without expensive LLM parameter tuning. Our evaluation demonstrates that SceneCraft surpasses existing LLM-based agents in rendering complex scenes, as shown by its adherence to constraints and favorable human assessments. We also showcase the broader application potential of SceneCraft by reconstructing detailed 3D scenes from the Sintel movie and guiding a video generative model with generated scenes as intermediary control signal.

Table of Contents

  • 1. Introduction
  • 2. Approach
  • 2.1. Asset Retrieval and Scene Decomposition
  • 2.2. Scene Graph Construction
  • 2.3. Scene Layout Optimization in a Feedback Loop
  • 2.4. Library Learning
  • 3. Experiments
  • 3.1. Evaluate Scene Synthesis with Given Constraints
  • 3.2. Scene-Guided Video Generation over Sintel Movie
  • 4. Related Works
  • 5. Conclusion
  • Impact Statement
  • References
  • Supplementary Material for SCENECRAFT
  • A. Discussion with other Related Works
  • B. Examples of SceneCraft's Generated Scripts and Rendered Scenes
  • C. List of relationships
  • D. Spatial Skill Library
  • E. Implementation Details
  • F. Examples of annotated queries
  • G. Prompt Used at each stage

Knowls

  1. Knowl 1 — SceneCraft uses coupled per-scene refinement and cross-scene skill learning

    model/method

    SceneCraft turns a natural-language scene query into Blender-executable Python code through two coupled learning loops. In the inner loop, the agent plans asset relationships, generates code and layout constraints, renders the scene in Blender, and uses a vision-language model to critique and revise the result. In the outer loop, it collects code changes made during inner-loop refinement across a batch of queries and consolidates recurring changes into a reusable spatial-skill library. The library is non-parametric Python code, so the system can update its reusable design knowledge without fine-tuning LLM parameters. The intended system can compose scenes containing as many as about 100 assets by decomposing the planning task.

  2. Knowl 2 — Asset retrieval and scene decomposition prepare queries for layout planning

    model/method

    For an input query, an LLM first proposes asset names with detailed visual descriptions. A CLIP-based retrieval system finds the top 10 candidate 3D assets for each description, renders the candidates, and selects the asset with the highest text-to-image score. To make large scenes more manageable, an LLM decomposer then breaks the query into an ordered sequence of sub-scenes, each with a title, a nonempty subset of assets, and a description used to guide later planning. The paper reports that this is intended to keep individual planning problems small; in practice, sub-scenes typically contain 10–20 assets.

  3. Knowl 3 — A relational bipartite graph represents scene plans

    definition

    SceneCraft represents a scene as a relational bipartite graph G(s)=(A,R,E)G(s)=(A,R,E), where ss is the scene, AA is its set of 3D assets, RR is a set of relation nodes, and EE connects each relation node to the subset of assets involved in that relation. Multiple nodes of the same relation type can connect different asset subsets. An LLM planner creates the graph from a sub-scene description and its asset list; the graph serves as an intermediate plan that specifies which spatial or contextual requirements the layout must satisfy. The relation types include proximity, direction, alignment, symmetry, overlap, parallelism, perpendicularity, hierarchy, rotation, repetition, and scaling.

  4. Knowl 4 — Layout is optimized against relation-specific satisfaction scores

    model/method

    Each asset aia_i has a layout matrix L(ai)L(a_i) encoding its position, scale, and orientation in the scene coordinate frame. For each relation node rr, a scoring function FrF_r receives the layout matrices of the connected assets and relation-specific arguments argr\mathrm{arg}_r (for example, a target distance); it returns a real-valued satisfaction score between 0 and 1. SceneCraft’s constraint-based search seeks a collection of layouts LL maximizing the sum of the relation scores:

    L^=arg⁡max⁡L∑r∈RFr({L(ai):ai∈E(r)},argr).\hat L=\arg\max_L\sum_{r\in R}F_r\left(\{L(a_i):a_i\in E(r)\},\mathrm{arg}_r\right).

    Here, RR is the set of relation nodes, E(r)E(r) is the asset subset connected to relation node rr, and L^\hat L is the selected layout. The LLM coder generates Blender code using available library functions and predicts their arguments; Blender renders the resulting scene.

  5. Knowl 5 — Rendered scenes drive per-scene critique and revision

    model/method

    SceneCraft’s inner loop uses rendered images as feedback on the generated layout. At each iteration, a multimodal LLM reviewer receives the Blender-rendered image and the relevant scene description, identifies missing or incorrectly satisfied constraints, and revises the scene representation and code. Depending on the diagnosed error, the reviewer can change the scene-graph edges, modify or add scoring functions, or adjust function arguments. The revised code is used in the next layout search and render. This process is intended to correct both errors in the planned relationships and errors in how those relationships are translated into numerical constraints.

  6. Knowl 6 — The outer loop consolidates reusable constraint-function updates

    model/method

    After inner-loop refinement on each query in a batch QQ, SceneCraft collects the final version of each relation’s scoring function as changed during that query’s refinement. A library learner reviews the per-query versions, looks for recurring modifications or common patterns, and merges a consensus update into the shared spatial-skill library. For example, feedback on a parallelism function that considered only asset locations led to adding orientation similarity to the function. The updates are represented as Python code; the library-learning signal comes from the multimodal reviewer’s query-alignment critiques, not ground-truth scenes, an explicit reward function, human intervention, or LLM parameter tuning.

  7. Knowl 7 — Synthetic-query evaluation separates library learning from testing

    experimental setup

    The authors manually created 40 synthetic queries with ground-truth spatial constraints, using sampled relation constraints to define the queries and assets retrieved from TurboSquid. For evaluation, human annotators also wrote query-specific scoring functions: each returns a score no greater than 1 and reaches 1 only when the constraints are strictly satisfied. SceneCraft uses 20 queries to build its spatial-skill library through dual-loop optimization, while seeing the queries but not the ground-truth constraint scores; performance is then assessed on the other 20 queries. The reported metrics are text-to-image CLIP similarity and the annotator-written constraint score.

  8. Knowl 8 — SceneCraft outperforms BlenderGPT on synthetic-query scores

    empirical result

    On the 20 held-out synthetic queries, SceneCraft scored 69.8 CLIP similarity and 88.9 constraint score, compared with 24.7 and 5.6, respectively, for the modified BlenderGPT baseline. The reported ablations scored 48.3 CLIP similarity and 64.5 constraint score without the learned library, 32.8 and 26.1 without the inner loop, and 19.4 and 3.2 without the relation graph. Thus, each removed component was associated with lower scores in this evaluation; removing the inner loop produced the largest constraint-score drop among the first two ablations, while removing the relation graph produced the lowest scores overall.

  9. Knowl 9 — Human raters prefer SceneCraft across fidelity, composition, and aesthetics

    empirical result

    In a qualitative comparison, 22 responses evaluated 10 randomly selected SceneCraft–BlenderGPT output pairs, with the ordering of the systems randomized. Raters judged text fidelity, composition and constraint agreement (with the ground-truth relations provided), and overall aesthetics. SceneCraft’s reported win rates were 76.8% for text fidelity, 83.6% for composition, and 74.5% for aesthetics; BlenderGPT’s corresponding rates were 12.7%, 11.4%, and 14.5%. The paper reports that SceneCraft was preferred on all three dimensions, with its largest reported advantage in composition.

  10. Knowl 10 — Sintel study tests scene planning as conditioning for video generation

    empirical result

    For a real-scene case study, the authors used the animated Sintel movie, treating its first half as training data and its remaining half as test data. The system was given fixed ground-truth assets and evaluated on layout planning. VideoPoet was fine-tuned on the training set using one ground-truth scene image frame as a conditional input; at test time it generated two-second videos conditioned on scenes produced by BlenderGPT or SceneCraft. SceneCraft achieved layout-matrix similarity 69.3 and scene CLIP similarity 82.7, compared with BlenderGPT’s 27.5 and 41.8. For video comparison, the reported CLIP-based relative-matching scores were 46.2 for SceneCraft and 69.1 for BlenderGPT; FVD, for which lower is better, was 317 and 574, respectively. The text-to-video baselines had no scene metrics: without fine-tuning they scored 56.8 CLIP-based relative matching and 846 FVD, and with fine-tuning they scored 64.2 and 531. The paper presents the generated scenes as a promising means of video control; the reported relative-matching value for SceneCraft is lower than BlenderGPT’s, whereas its FVD is lower.

  11. Knowl 11 — SceneCraft has practical limits in asset scale and rendering control

    limitation

    The authors report that LLM refinement becomes less effective when a sub-scene contains more than about 20 assets, motivating decomposition into smaller planning problems. Their refinement procedure uses a fixed four iterations without early stopping; they chose this as a balance between quality and efficiency after observing that more iterations could continue to improve results and that generation without a useful library could require more than 10. Assets are canonicalized to a default height of 1 and assigned predicted heights, but height is not adjusted during refinement, so some dimensions may remain unrealistic. Lighting and ground texture are fixed rather than controlled by the LLM. The implementation allows at most seven sub-problems and reports an average of 6,000 tokens per query, a maximum of about 15,000, and estimated costs of about 0.12perquerywithGPT−4−turboor0.12 per query with GPT-4-turbo or 0.006 with GPT-3.5.

Coverage note — The paper’s illustrative generated scripts, individual learned helper functions, and prompt templates are omitted because they are examples or implementation detail rather than standalone results needed to reconstruct the central method and evaluations.

Citation

MLA
Hu, Z., et al. “SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code”. arXiv, 2024, http://arxiv.org/abs/2403.01248v1.
APA
Hu, Z., Iscen, A., Jain, A., Kipf, T., Yue, Y., Ross, D. A., Schmid, C., & Fathi, A. (2024). SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code. arXiv. http://arxiv.org/abs/2403.01248v1
Chicago
Hu, Z., A. Iscen, A. Jain, et al. 2024. “SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code”. arXiv. http://arxiv.org/abs/2403.01248v1.
Harvard
Hu, Z. et al. (2024) “SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.01248v1.
Vancouver
1. Hu Z, Iscen A, Jain A, Kipf T, Yue Y, Ross DA, Schmid C, Fathi A (2024) SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code. arXiv

BibTeX

@article{hu2024scenecraft,
  title = {SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code},
  author = {Hu, Ziniu and Iscen, Ahmet and Jain, Aashi and Kipf, Thomas and Yue, Yisong and Ross, David A. and Schmid, Cordelia and Fathi, Alireza},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.01248v1},
  eprint = {2403.01248}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/