A Few Words Go a Long Way: Language Guided Robot Policy Synthesis
Daphne ChenArchit Ritesh JainEric GoossenEmma RomigMichael MurrayNick WalkerMaya Cakmak
Presents ARCHITECT, an interactive framework that uses large language models to synthesize modular, interpretable robot code from human natural language corrections, establishing a reusable skill library that outperforms black-box vision-language-action models on long-horizon manipulation tasks.
Modern generalist robot policies, such as end-to-end vision-language-action (VLA) models, struggle with real-world deployment due to their black-box nature. Minor changes in environment, lighting, or object orientation cause cascading failures that are difficult to correct without gathering costly robot demonstrations and undergoing extensive retraining. Furthermore, one-shot code-generation methods frequently fail due to underspecified instructions and visual ambiguity. To address this adaptability bottleneck, the article evaluates ARCHITECT, an agentic framework that approaches robot manipulation through interactive program synthesis steered by plain-language human corrections and execution tracing.
ARCHITECT uses a large language model to orchestrate modular tools for perception, robot control, and state monitoring without requiring robot-specific training data. When an execution fails or proves suboptimal, an observing human provides direct natural language feedback (for instance, specifying that a grasp should be lower). The system diagnoses the failure through execution traces, re-synthesizes the policy, and distills the correction into a persistent skill library for long-term reuse. Evaluated on a Franka Panda robotic arm across eight manipulation tasks—spanning articulated mechanisms, deformable cloth folding, and long-horizon clutter retrieval—ARCHITECT demonstrated substantial performance advantages. Across complex tasks, it achieved success rates ranging from 70% to 100%, whereas leading VLA models (such as π0 and π0.5) and baseline program synthesis frameworks consistently failed, often scoring 0% success on multi-stage, cloth, and drawer operations. In a six-participant human study, reusing the accumulated skill library reduced required supervisor interventions from an average of 4.67 queries per trial down to 0.83 (an 82% reduction) and enabled zero-shot transfer to novel tasks with a 67% success rate compared to 0% without prior skills.
These findings demonstrate that modular, code-centric architectures combined with structured human feedback offer a practical, interpretable, and data-efficient alternative to black-box robotic models. Retaining corrections in an explicit skill library amortizes human effort, dramatically lowering operating overhead and eliminating the continuous data-collection cycle. Additionally, the analysis indicates that human visual and physical intuition remains essential, as human corrections successfully resolved physical and depth-related errors that automated vision-language feedback could not detect.
Organizations developing or deploying robotic manipulation systems should prioritize modular, steerable architectures that decouple high-level planning from low-level execution primitives. However, decision-makers should note that ARCHITECT remains bounded by the physical precision of its underlying perception, grasp sampling, and motion-planning tools. In addition, the richness of the synthesized skill library depends on encountering sufficient task variation and diverse corrections. Before full-scale deployment in production environments, teams should conduct expanded pilot studies examining diverse user cohorts, complex industrial geometries, and the long-term governance of growing skill libraries.
- Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). Read this foundational Code-as-Policies work first to understand how language models generate executable robot-control programs, the central policy representation ARCHITECT extends.
- Paper: Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, Michael Ahn et al. (2022). Its SayCan framework grounds language-model plans in executable robot skills, preparing you for ARCHITECT’s programmatic use of perception and control tools.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). Inner Monologue establishes how execution feedback can drive language-model replanning, a key precursor to ARCHITECT’s human-guided correction loop.
- Paper: SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills, Zhongxin Guo et al. (2026). SKILL-DISCO carries ARCHITECT’s trace-to-reusable-skill idea further by distilling execution traces into structured, compiled skills and verifying them across tasks.
