A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

Daphne ChenArchit Ritesh JainEric GoossenEmma RomigMichael MurrayNick WalkerMaya Cakmak

article2026arXiv1 citations

Presents ARCHITECT, an interactive framework that uses large language models to synthesize modular, interpretable robot code from human natural language corrections, establishing a reusable skill library that outperforms black-box vision-language-action models on long-horizon manipulation tasks.

Listen

Modern generalist robot policies, such as end-to-end vision-language-action (VLA) models, struggle with real-world deployment due to their black-box nature. Minor changes in environment, lighting, or object orientation cause cascading failures that are difficult to correct without gathering costly robot demonstrations and undergoing extensive retraining. Furthermore, one-shot code-generation methods frequently fail due to underspecified instructions and visual ambiguity. To address this adaptability bottleneck, the article evaluates ARCHITECT, an agentic framework that approaches robot manipulation through interactive program synthesis steered by plain-language human corrections and execution tracing.

ARCHITECT uses a large language model to orchestrate modular tools for perception, robot control, and state monitoring without requiring robot-specific training data. When an execution fails or proves suboptimal, an observing human provides direct natural language feedback (for instance, specifying that a grasp should be lower). The system diagnoses the failure through execution traces, re-synthesizes the policy, and distills the correction into a persistent skill library for long-term reuse. Evaluated on a Franka Panda robotic arm across eight manipulation tasks—spanning articulated mechanisms, deformable cloth folding, and long-horizon clutter retrieval—ARCHITECT demonstrated substantial performance advantages. Across complex tasks, it achieved success rates ranging from 70% to 100%, whereas leading VLA models (such as π0 and π0.5) and baseline program synthesis frameworks consistently failed, often scoring 0% success on multi-stage, cloth, and drawer operations. In a six-participant human study, reusing the accumulated skill library reduced required supervisor interventions from an average of 4.67 queries per trial down to 0.83 (an 82% reduction) and enabled zero-shot transfer to novel tasks with a 67% success rate compared to 0% without prior skills.

These findings demonstrate that modular, code-centric architectures combined with structured human feedback offer a practical, interpretable, and data-efficient alternative to black-box robotic models. Retaining corrections in an explicit skill library amortizes human effort, dramatically lowering operating overhead and eliminating the continuous data-collection cycle. Additionally, the analysis indicates that human visual and physical intuition remains essential, as human corrections successfully resolved physical and depth-related errors that automated vision-language feedback could not detect.

Organizations developing or deploying robotic manipulation systems should prioritize modular, steerable architectures that decouple high-level planning from low-level execution primitives. However, decision-makers should note that ARCHITECT remains bounded by the physical precision of its underlying perception, grasp sampling, and motion-planning tools. In addition, the richness of the synthesized skill library depends on encountering sufficient task variation and diverse corrections. Before full-scale deployment in production environments, teams should conduct expanded pilot studies examining diverse user cohorts, complex industrial geometries, and the long-term governance of growing skill libraries.

Cover for A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

Abstract

While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail. In this work, we propose ARCHITECT, a framework that treats robot policy acquisition as an interactive program synthesis task. ARCHITECT leverages the reasoning capabilities of LLM coding agents to synthesize modular robot programs that utilize a suite of perception and control tools. Unlike end-to-end models where distribution shift leads to unpredictable, cascading failures, our modular architecture allows users to isolate failures and localize feedback at the level of abstraction required. We introduce an iterative process where a human supervisor provides natural language corrections to steer the policy. These corrections are grounded in the policy code by program execution traces and distilled into a persistent skill library, a form of long-term in-context learning which enables the agent to accumulate a repertoire of reusable, interpretable behaviors. In a benchmark evaluation on a Franka Panda robot, ARCHITECT outperforms state-of-the-art VLA models and program synthesis baselines on complex, long-horizon tasks, including articulated object manipulation and cloth folding. Our results demonstrate that the synthesized skill library enables the system to transfer to novel tasks with decreasing human intervention, providing a steerable and data-efficient alternative to black-box robot learning. Website: this https URL

Citation

MLA
Chen, D., et al. “A Few Words Go a Long Way: Language Guided Robot Policy Synthesis”. arXiv, 2026, http://arxiv.org/abs/2607.23784v1.
APA
Chen, D., Jain, A. R., Goossen, E., Romig, E., Murray, M., Walker, N., & Cakmak, M. (2026). A Few Words Go a Long Way: Language Guided Robot Policy Synthesis. arXiv. http://arxiv.org/abs/2607.23784v1
Chicago
Chen, D., A. R. Jain, E. Goossen, et al. 2026. “A Few Words Go a Long Way: Language Guided Robot Policy Synthesis”. arXiv. http://arxiv.org/abs/2607.23784v1.
Harvard
Chen, D. et al. (2026) “A Few Words Go a Long Way: Language Guided Robot Policy Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2607.23784v1.
Vancouver
1. Chen D, Jain AR, Goossen E, Romig E, Murray M, Walker N, Cakmak M (2026) A Few Words Go a Long Way: Language Guided Robot Policy Synthesis. arXiv

BibTeX

@article{chen2026few,
  title = {A Few Words Go a Long Way: Language Guided Robot Policy Synthesis},
  author = {Chen, Daphne and Jain, Archit Ritesh and Goossen, Eric and Romig, Emma and Murray, Michael and Walker, Nick and Cakmak, Maya},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2607.23784v1},
  eprint = {2607.23784}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/