Reasoning Like Program Executors
Xinyu PiQian LiuBei ChenMorteza ZiyadiZeqi LinQiang FuYan GaoJian-Guang LouWeizhu Chen
Proposes a pre-training paradigm that teaches language models to predict program execution outputs, transferring formal symbolic reasoning capabilities directly into neural models for downstream natural language tasks.
Existing pre-trained language models achieve near-human scores on standard language understanding benchmarks but struggle with complex reasoning, including numerical calculations, formal logic, and multi-hop deductions. While standard language pre-training captures broad linguistic patterns, it rarely encounters clean, structured reasoning steps at scale. Hybrid systems that connect external symbolic engines to neural networks can execute precise rules, but their reasoning mechanisms remain external and fail to generalize across diverse, unseen language tasks.
The article evaluates a new pre-training paradigm called POET (Program Executor). The primary objective is to demonstrate that language models can internalize formal reasoning principles by learning to mimic deterministic program executors and subsequently transfer these skills to natural language reasoning tasks.
To evaluate this concept, the authors generated synthetic pre-training datasets pairing executable programs and structured contexts with exact execution outputs. They tested three implementations: POET-Math (arithmetic calculators), POET-Logic (first-order logic solvers), and POET-SQL (relational database query engines). These pre-trained models were applied across distinct model architectures, ranging from medium-sized encoders and decoders (BART, RoBERTa) to giant models (T5-11B), and evaluated on six benchmark datasets covering numerical, logical, hybrid table-text, and quantitative inference tasks.
The study established several key findings. First, pre-training on program execution substantially improved downstream reasoning across all evaluated models. On the numerical benchmark DROP, POET-SQL improved exact match accuracy on BART by 11.5 percentage points (from 66.2% to 77.7%) and on the math diagnostic SVAMP by 21.1 percentage points (from 12.4% to 33.5%). Second, integrated executors proved versatile: POET-SQL provided consistent cross-domain gains, outperforming previous specialized reasoning architectures and language-only pre-training methods. Third, the pre-training showed high data efficiency; models trained on only 10% of the synthetic corpus (500,000 examples) achieved performance comparable to those trained on the full 5-million example dataset. Finally, standard natural language understanding was preserved with minimal degradation on standard benchmarks, and the reasoning transfer proved abstract rather than surface-level, as modifying the naturalness of query syntax caused little variation in downstream gains.
These results demonstrate that formal reasoning can be internalized into neural parameters via synthetic data rather than expensive, noisy natural language curation. This approach significantly lowers the data collection barrier for specialized reasoning applications in domains such as finance, analytics, and business intelligence, reducing the risk of calculation errors while avoiding the operational complexity of integrating external symbolic modules.
Organizations developing reasoning-intensive language applications should consider incorporating synthetic execution pre-training rather than relying solely on larger text scraping or few-shot prompting. Future engineering efforts should focus on combining program-execution pre-training with step-by-step prompting methods, expanding program generation grammars beyond fixed templates, and joint training to mitigate minor regressions on general language inference tasks.
Confidence in these findings is high for structured numerical, logical, and database-style tasks within the tested benchmarks. However, leaders should note two main limitations: the model transfers reasoning effectively only when the operations in pre-training closely overlap with the target downstream requirements, and the synthetic data generation in this study relied on template-driven designs rather than fully generalized context-free grammars.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). Program-of-Thoughts establishes the external-execution approach to numerical reasoning that POET recasts as a pre-training signal for internalizing computation.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Chain-of-Thought prompting provides the step-by-step reasoning baseline whose calculation failures motivate POET’s program-execution training.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). Program Synthesis with Large Language Models introduces the code-generation and execution setting that underlies POET’s synthetic program-output training examples.
- Paper: Chain of Code: Reasoning with a Language Model-Augmented Code Emulator, Chengshu Li et al. (2024). Chain of Code extends program-based reasoning beyond deterministic execution by handing semantic steps back to a language model emulator.
- Paper: Steering Large Language Models between Code Execution and Textual Reasoning, Yongchao Chen et al. (2025). This study continues the code-execution line by testing when models should execute code rather than reason in text, exposing a key decision problem for POET-style capabilities.
- Paper: Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective, Yiyao Yu et al. (2025). Chain-of-Reasoning extends execution-based reasoning by unifying Python computation with natural-language and formal-proof reasoning across mathematical tasks.
