Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
Chengshu LiJacky LiangAndy ZengXinyun ChenKarol HausmanDorsa SadighSergey LevineLi Fei-FeiFei XiaBrian Ichter
Proposes Chain of Code, a reasoning framework that interweaves Python execution with language model emulation for non-executable pseudocode, boosting reasoning accuracy on BIG-Bench Hard to 84% across both algorithmic and semantic tasks.
Large language models frequently struggle with complex reasoning problems that blend exact computation with subjective semantic understanding. Natural language reasoning techniques like Chain of Thought excel at linguistic nuances but fail at precise arithmetic or symbolic manipulation, whereas standard code execution methods fail when encountering non-algorithmic semantic subtasks, such as determining sarcasm or categorizing real-world concepts.
The article demonstrates and evaluates Chain of Code, a framework designed to enhance language model reasoning by combining the algorithmic precision of a code interpreter with the semantic commonsense of a language model. The core objective is to allow models to format complex reasoning problems as programs that seamlessly hand off semantic steps to the language model itself when code execution is not possible.
The researchers evaluated this framework across multiple benchmark datasets, primarily the 23 challenging tasks in BIG-Bench Hard and the GSM8K grade-school math suite. They tested several model families, including OpenAI completion models (text-davinci-003 and smaller variants), instruction-tuned models (GPT-3.5 and GPT-4), and PaLM-2. The approach operates in two stages: first, the model generates reasoning steps formatted as structured code or pseudocode; second, a Python interpreter runs the executable lines, and whenever an unexecutable semantic function is caught, a language model emulator (termed an LMulator) simulates the expected output and updates the shared program state.
The findings establish that Chain of Code significantly improves multi-step reasoning performance. On BIG-Bench Hard, the method achieved an overall accuracy of 84% using text-davinci-003, representing a 12% gain over Chain of Thought (72%) and substantially outperforming the average human baseline of 68%. When paired with GPT-4, accuracy reached 91%. The gains were especially pronounced on algorithmic tasks, where Chain of Code achieved 95% accuracy compared to 71% for Chain of Thought. Ablation studies confirmed that both interpreter execution and language model simulation are essential, as relying solely on Python dropped overall accuracy to 48%. Furthermore, the framework scaled effectively to smaller models and demonstrated strong generalization in cross-task prompting and physical robotics tasks.
These results demonstrate that expressing complex problems through program structure reduces calculation errors without sacrificing natural language understanding. For organizations deploying automated decision systems, customer-facing agents, or robotic workflows, this hybrid structure increases reliability, auditability, and execution accuracy across mixed semantic and numerical operations.
Organizations developing reasoning pipelines should consider adopting hybrid code-execution architectures for complex multi-step tasks rather than relying purely on natural language prompting. For production deployments, teams should implement sandboxing and security safeguards, as executing dynamically generated code introduces vulnerabilities if prompts are maliciously constructed.
Key limitations include increased computation time and context length requirements due to multi-turn execution and state tracking. The current implementation tracks state via string parsing into basic Python types, preventing the modification of custom, non-serialized objects. Further work is required to optimize inference latency, evaluate fine-tuned dedicated emulator models, and expand the framework to richer external tool ecosystems.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). Program of Thoughts establishes the key precedent of separating language-model reasoning from computation by generating code for an interpreter to execute.
- Paper: PAL: Program-aided Language Models, Luyu Gao et al. (2023). PAL shows how executable programs can carry a model’s intermediate reasoning, clarifying the code-driven approach that Chain of Code broadens to semantic subtasks.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Chain-of-Thought prompting provides the step-by-step reasoning baseline that Chain of Code extends by expressing parts of a reasoning trace as code.
- Paper: Steering Large Language Models between Code Execution and Textual Reasoning, Yongchao Chen et al. (2025). This study probes when models should reason in text versus execute code, extending Chain of Code’s central concern with coordinating those two modes.
