Reasoning Like Program Executors

Xinyu PiQian LiuBei ChenMorteza ZiyadiZeqi LinQiang FuYan GaoJian-Guang LouWeizhu Chen

article2022EMNLP61 citations

Proposes a pre-training paradigm that teaches language models to predict program execution outputs, transferring formal symbolic reasoning capabilities directly into neural models for downstream natural language tasks.

Listen

Existing pre-trained language models achieve near-human scores on standard language understanding benchmarks but struggle with complex reasoning, including numerical calculations, formal logic, and multi-hop deductions. While standard language pre-training captures broad linguistic patterns, it rarely encounters clean, structured reasoning steps at scale. Hybrid systems that connect external symbolic engines to neural networks can execute precise rules, but their reasoning mechanisms remain external and fail to generalize across diverse, unseen language tasks.

The article evaluates a new pre-training paradigm called POET (Program Executor). The primary objective is to demonstrate that language models can internalize formal reasoning principles by learning to mimic deterministic program executors and subsequently transfer these skills to natural language reasoning tasks.

To evaluate this concept, the authors generated synthetic pre-training datasets pairing executable programs and structured contexts with exact execution outputs. They tested three implementations: POET-Math (arithmetic calculators), POET-Logic (first-order logic solvers), and POET-SQL (relational database query engines). These pre-trained models were applied across distinct model architectures, ranging from medium-sized encoders and decoders (BART, RoBERTa) to giant models (T5-11B), and evaluated on six benchmark datasets covering numerical, logical, hybrid table-text, and quantitative inference tasks.

The study established several key findings. First, pre-training on program execution substantially improved downstream reasoning across all evaluated models. On the numerical benchmark DROP, POET-SQL improved exact match accuracy on BART by 11.5 percentage points (from 66.2% to 77.7%) and on the math diagnostic SVAMP by 21.1 percentage points (from 12.4% to 33.5%). Second, integrated executors proved versatile: POET-SQL provided consistent cross-domain gains, outperforming previous specialized reasoning architectures and language-only pre-training methods. Third, the pre-training showed high data efficiency; models trained on only 10% of the synthetic corpus (500,000 examples) achieved performance comparable to those trained on the full 5-million example dataset. Finally, standard natural language understanding was preserved with minimal degradation on standard benchmarks, and the reasoning transfer proved abstract rather than surface-level, as modifying the naturalness of query syntax caused little variation in downstream gains.

These results demonstrate that formal reasoning can be internalized into neural parameters via synthetic data rather than expensive, noisy natural language curation. This approach significantly lowers the data collection barrier for specialized reasoning applications in domains such as finance, analytics, and business intelligence, reducing the risk of calculation errors while avoiding the operational complexity of integrating external symbolic modules.

Organizations developing reasoning-intensive language applications should consider incorporating synthetic execution pre-training rather than relying solely on larger text scraping or few-shot prompting. Future engineering efforts should focus on combining program-execution pre-training with step-by-step prompting methods, expanding program generation grammars beyond fixed templates, and joint training to mitigate minor regressions on general language inference tasks.

Confidence in these findings is high for structured numerical, logical, and database-style tasks within the tested benchmarks. However, leaders should note two main limitations: the model transfers reasoning effectively only when the operations in pre-training closely overlap with the target downstream requirements, and the synthetic data generation in this study relied on template-driven designs rather than fully generalized context-free grammars.

Cover for Reasoning Like Program Executors

Abstract

Reasoning over natural language is a long-standing goal for the research community. However, studies have shown that existing language models are inadequate in reasoning. To address the issue, we present PoET, a novel reasoning pre-training paradigm. Through pre-training language models with programs and their execution results, PoET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach. PoET is conceptually simple and can be instantiated by different kinds of program executors. In this paper, we showcase two simple instances PoET-Math and PoET-Logic, in addition to a complex instance, PoET-SQL. Experimental results on six benchmarks demonstrate that PoET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning. PoET opens a new gate on reasoning-enhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 3 Reasoning Like Program Executors
  • 4 POET with Singleton Executors
  • 4.1 Learning from Math Calculators
  • 4.2 Learning from Logic Solvers
  • 4.3 Preliminary Observation
  • 5 POET with Integrated Executors
  • 6 Experiments and Analysis
  • 6.1 Experimental Results
  • 6.2 Pre-training Analysis
  • 7 Discussion and Open Questions
  • 8 Conclusion & Future Work
  • Limitations
  • Acknowledgement
  • References
  • A Program Context Analysis
  • A.1 The Necessity of Program Context
  • A.2 The Variables Design in Program Context
  • B Experimental Setup
  • B.1 Dataset Setup
  • B.2 Baseline Setup
  • C Implementation Details
  • C.2 Pre-training Details
  • C.3 Fine-tuning Details
  • C.4 Fine-tuning Hyperparameters
  • D Fine-grained Analysis
  • E NL Understanding Performance

Knowls

  1. Knowl 1 — POET transfers executor supervision into language-model pre-training

    model/method

    POET trains a language model to reproduce the result of executing a program in its program context. A training example pairs a program (such as a SQL query or arithmetic expression) with the context it runs against (such as a database or variable assignments); a deterministic program executor supplies the execution result as the target. The model is then fine-tuned on natural-language tasks, where a sentence and its natural-language context play roles analogous to the program and program context, and the answer corresponds to the execution result. POET’s motivating hypothesis is that program executors embody formal reasoning knowledge and that a language model can acquire and transfer some of that knowledge through execution-result prediction; the experiments test this transfer rather than establish that hypothesis as a general mechanism.

  2. Knowl 2 — POET-SQL uses database execution to teach multiple reasoning operations

    model/method

    POET-SQL pairs an SQL query with its database and trains a language model to predict the query result. Its SQL templates cover arithmetic, superlative selection, comparisons, aggregation, union, and nested queries. Databases are flattened into input sequences; rows are randomly removed from large databases until the flattened input is under 450 tokens. SQL templates from SQUALL are instantiated over WIKISQL databases to create 5 million examples for encoder–decoder result generation. Because encoder-only models cannot freely generate results absent from the input, their task is instead to tag result tokens with I and other tokens with O, retaining only examples whose results occur in the database; this produces nearly 2 million examples. The result-generation examples are suitable for encoder–decoder models such as BART, while the result-selection task lets encoder-only models such as RoBERTa learn from SQL execution.

  3. Knowl 3 — POET-Math trains arithmetic over relevant and irrelevant variables

    model/method

    POET-Math teaches addition and subtraction by pairing a math expression with a context containing floating-point variable assignments and asking an encoder–decoder model to output the calculated number. Expressions contain at most two operators and use at most three variables; contexts contain no more than 30 variables, including variables not used by the expression. Variable values range from 0.0 to 1000.0. Random generation produced a 4-million-example pre-training corpus. The irrelevant assignments are intended to make the program context resemble natural-language passages that include information unrelated to the question. POET-Math targets arithmetic reasoning needed for questions involving addition or subtraction.

  4. Knowl 4 — POET-Logic learns implication judgments from a solver

    model/method

    POET-Logic trains an encoder-only model to determine whether a conclusion follows necessarily from a set of first-order-logic premises. To synthesize an example, the authors allocate five Boolean variables and sample at most eight variable pairs from the pairwise combinations. For each selected pair (p,q)(p,q), they add one randomly selected statement from {p→q,p→¬q,¬p→¬q,¬p→q}\{p \to q, p \to \neg q, \neg p \to \neg q, \neg p \to q\}. One statement is randomly designated as the conclusion and the remaining statements as premises; Z3 supplies the True/False implication label. The resulting corpus contains 1 million examples, of which nearly 16% have a True label. This instantiation targets logical reasoning, including necessary conditional reasoning.

  5. Knowl 5 — POET improves its backbone models on six reasoning benchmarks

    data/table

    The reported comparisons show gains over the corresponding unmodified backbone models. Exact match (EM) and F1 are percentages; DROP, HotpotQA, and TAT-QA results below are on development sets, while SVAMP and EQUATE results are on test sets. The RoBERTa DROP comparison uses only the span-answer subset. The authors report the POET gains as statistically significant at p<0.05p<0.05.

    • On DROP, BART-Large scores 66.2 EM and 69.2 F1; POET-SQL-BART scores 77.7 EM (+11.5) and 80.6 F1 (+11.4). RoBERTa-Large scores 78.1 EM and 85.3 F1; POET-SQL-RoBERTa scores 79.8 EM (+1.7) and 87.4 F1 (+2.1).
    • On HotpotQA, BART-Large scores 65.6 EM and 78.9 F1; POET-SQL-BART scores 66.5 EM (+0.9) and 79.7 F1 (+0.8). RoBERTa-Large scores 67.6 EM and 81.1 F1; POET-SQL-RoBERTa scores 68.7 EM (+1.1) and 81.6 F1 (+0.5).
    • On TAT-QA, BART-Large scores 38.8 EM and 46.7 F1; POET-SQL-BART scores 41.5 EM (+2.7) and 49.6 F1 (+2.9). RoBERTa-Large scores 55.2 EM and 62.7 F1; POET-SQL-RoBERTa scores 59.1 EM (+3.9) and 65.9 F1 (+3.2).
    • On SVAMP, BART-Large scores 12.4 EM and POET-SQL-BART scores 33.5 EM (+21.1); encoder-only results are not reported.
    • On EQUATE, BART-Large scores 62.6 EM and POET-SQL-BART scores 66.5 EM (+3.9); RoBERTa-Large scores 64.2 EM and POET-SQL-RoBERTa scores 67.5 EM (+3.3).
    • In a separate preliminary comparison, BART-Large obtains 66.2 EM on DROP and POET-Math obtains 75.2; RoBERTa-Large obtains 36.7 EM on LogiQA and POET-Logic obtains 38.9.

    The improvements span numerical, multi-hop, hybrid table-and-text, and quantitative inference settings, even though POET pre-training uses programs and program contexts rather than the natural-language inputs of these downstream tasks.

  6. Knowl 6 — POET models are competitive with specialized reasoning systems

    empirical result

    Across comparisons with prior systems, POET is competitive but does not outperform every specialized model. On DROP, POET-SQL-BART obtains 77.7 EM and 80.6 F1, and POET-Math+SQL-BART obtains 78.0 EM and 80.9 F1; the specialized QDGAT system reports 84.1 EM and 87.1 F1. Among language-model approaches listed for DROP, POET-SQL-BART exceeds PReasM’s 69.4 EM and 72.3 F1. On HotpotQA, POET-SQL-RoBERTa reaches 68.7 EM and 81.6 F1, above SpanBERT’s 67.4 EM and 81.2 F1 but below HGN’s 69.2 EM and 82.2 F1. On TAT-QA, adding POET-SQL-RoBERTa to TAGOP raises results from 55.2 EM and 62.7 F1 to 59.1 EM and 65.9 F1. On EQUATE, POET-SQL-RoBERTa obtains 67.5 EM, compared with 60.7 for Q-REAS.

    POET also improves a very large model: T5-11B rises from 83.5 EM and 85.9 F1 on DROP to 85.2 EM (+1.7) and 87.6 F1 (+1.7) with POET-SQL-T5; on SVAMP, EM rises from 52.9 to 57.4 (+4.5). These results indicate that execution pre-training can benefit both smaller backbones and a model with 11 billion parameters.

  7. Knowl 7 — Execution supervision and pre-training scale affect transfer

    empirical result

    An ablation replacing POET-SQL execution pre-training with SQL language modeling masks SQL queries and trains BART to reconstruct them from the masked query and database. The authors report only trivial downstream performance variance relative to POET-SQL, supporting the importance of predicting execution results rather than merely modeling SQL strings. In scale experiments, pre-training and DROP development performance rise with additional training steps and tend toward an asymptote; adding pre-training examples speeds convergence. A corpus containing 10% of the full POET-SQL data (500,000 examples) reaches approximately the same asymptotic performance as pre-training on the full 5-million-example corpus, indicating data efficiency in these experiments. The authors also observe the reverse direction of transfer: BART pre-trained on DROP has lower train and development perplexity while learning SQL execution than vanilla BART, suggesting that language-model pre-training on natural-language reasoning can assist subsequent learning of program execution.

  8. Knowl 8 — Irrelevant program-context variables improve arithmetic transfer

    empirical result

    In DROP evaluation, adding irrelevant variable assignments to POET-Math’s program contexts improves BART-Large performance. BART-Large without POET-Math scores 66.2 EM and 69.2 F1. POET-Math with no program context scores 67.4 EM and 70.5 F1; with a context containing the required variables and zero irrelevant variables it scores 71.5 EM and 74.5 F1; with 10 irrelevant variables it scores 74.6 EM and 77.5 F1; and with 30 irrelevant variables it scores 75.2 EM and 78.1 F1. Thus, arithmetic skill is learned even without irrelevant context, but including it helps, with the largest tested amount producing the strongest result. The authors interpret the irrelevant assignments as a way to approximate the distracting information common in natural-language passages.

  9. Knowl 9 — Transfer is not dependent on natural-looking inputs and largely preserves language understanding

    empirical result

    Two analyses examine whether POET’s transfer depends on surface naturalness or harms existing natural-language understanding. For POET-SQL-BART on DROP development data, the standard SQL setting scores 77.7 EM and 80.6 F1. Translating the programs into more natural sentences yields 77.2 EM and 79.9 F1; making programs less natural by replacing SQL keywords with low-frequency tokens yields 76.9 EM and 79.7 F1; converting database contexts into natural-language sentences yields 76.5 EM and 79.0 F1. These modest differences do not support the expectation that more natural-looking programs or contexts are necessary for effective transfer.

    For POET-SQL-RoBERTa on natural-language understanding tasks, vanilla RoBERTa-Large versus POET-SQL-RoBERTa scores are 93.3 versus 93.2 F1 on SQuAD, 90.8 versus 89.6 accuracy on MNLI, and 84.9 versus 84.7 F1 on QuoRef. The results are nearly unchanged on SQuAD and QuoRef, while MNLI shows a 1.2-percentage-point decrease. The experiments therefore suggest that POET can add reasoning performance without broadly erasing the tested language-understanding abilities, although the MNLI reduction is a measurable exception.

  10. Knowl 10 — POET’s transfer depends on overlap with downstream reasoning and on program diversity

    limitation

    POET relies on overlap between the reasoning operations represented in pre-training and those needed downstream. The limitation applies even to POET-SQL, which covers multiple kinds of computation: removing all pre-training programs containing mathematical operations leads to poor DROP performance. A second limitation is that the reported program corpora are synthesized from instantiated templates rather than probabilistic context-free grammars. The authors note that grammars could provide more diverse programs and potentially support better generalization, but are more complex to use.

Coverage note — Omitted the detailed answer-type and dataset-subset breakdowns, fine-tuning hyperparameters, and additional implementation specifics because they support rather than define the central method and transfer findings.

References

  1. 1.Daniel Andor, Luheng He, Kenton Lee, and Emily Pitler. 2019. Giving BERT a calculator: Finding operations and arguments with reading comprehension. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5947–5952, Hong Kong, China. Association for Computational Linguistics.
  2. 2.Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016. Neural module networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 39–48.
  3. 3.Akari Asai and Hannaneh Hajishirzi. 2020. Logic-guided data augmentation and regularization for consistent question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5642–5650, Online. Association for Computational Linguistics.
  4. 4.Tarek R. Besold, Artur S. d’Avila Garcez, Sebastian Bader, Howard Bowman, Pedro M. Domingos, Pascal Hitzler, Kai-Uwe Kühnberger, Luís C. Lamb, Daniel Lowd, Priscila Machado Vieira Lima, Leo de Penning, Gadi Pinkas, Hoifung Poon, and Gerson Zaverucha. 2017. Neural-symbolic learning and reasoning: A survey and interpretation. CoRR, abs/1711.03902.
  5. 5.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen tau Yih, and Yejin Choi. 2020. Abductive commonsense reasoning. In International Conference on Learning Representations.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  7. 7.Giovanni Campagna, Agata Foryciarz, Mehrad Moradshahi, and Monica Lam. 2020. Zero-shot transfer learning with synthesized data for multi-domain dialogue state tracking. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 122–132, Online. Association for Computational Linguistics.
  8. 8.Kunlong Chen, Weidi Xu, Xingyi Cheng, Zou Xiaochuan, Yuyu Zhang, Le Song, Taifeng Wang, Yuan Qi, and Wei Chu. 2020a. Question directed graph attention network for numerical reasoning over text. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6759–6768, Online. Association for Computational Linguistics.
  9. 9.Shuang Chen, Qian Liu, Zhiwei Yu, Chin-Yew Lin, Jian-Guang Lou, and Feng Jiang. 2021. ReTraCk: A flexible and efficient framework for knowledge base question answering. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 325–336, Online. Association for Computational Linguistics.
  10. 10.Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2019a. Tabfact: A large-scale dataset for table-based fact verification. In International Conference on Learning Representations.
  11. 11.Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. 2020b. HybridQA: A dataset of multi-hop question answering over tabular and textual data. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1026–1036, Online. Association for Computational Linguistics.
  12. 12.Xinyun Chen, Chen Liang, Adams Wei Yu, Denny Zhou, Dawn Song, and Quoc V. Le. 2020c. Neural symbolic reader: Scalable integration of distributed and symbolic representations for reading comprehension. In International Conference on Learning Representations.
  13. 13.Xinyun Chen, Chang Liu, and Dawn Song. 2019b. Execution-guided neural program synthesis. In International Conference on Learning Representations.
  14. 14.Pradeep Dasigi, Nelson F. Liu, Ana Marasovic, Noah A. ´ Smith, and Matt Gardner. 2019. Quoref: A reading comprehension dataset with questions requiring coreferential reasoning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5925–5932, Hong Kong, China. Association for Computational Linguistics.
  15. 15.Leonardo De Moura and Nikolaj Bjørner. 2008. Z3: An efficient smt solver. In International conference on Tools and Algorithms for the Construction and Analysis of Systems, pages 337–340. Springer.
  16. 16.Xiang Deng, Yu Su, Alyssa Lees, You Wu, Cong Yu, and Huan Sun. 2021. ReasonBERT: Pre-trained to reason with distant supervision. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6112–6127, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  17. 17.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  18. 18.Ming Ding, Chang Zhou, Qibin Chen, Hongxia Yang, and Jie Tang. 2019. Cognitive graph for multi-hop reading comprehension at scale. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2694–2703, Florence, Italy. Association for Computational Linguistics.
  19. 19.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2368–2378, Minneapolis, Minnesota. Association for Computational Linguistics.
  20. 20.Kevin Ellis, Maxwell I. Nye, Yewen Pu, Felix Sosa, Josh Tenenbaum, and Armando Solar-Lezama. 2019. Write, execute, assess: Program synthesis with a REPL. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 9165–9174.
  21. 21.Yuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai, Shuohang Wang, and Jingjing Liu. 2020. Hierarchical graph network for multi-hop question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8823–8838, Online. Association for Computational Linguistics.
  22. 22.Mor Geva, Ankit Gupta, and Jonathan Berant. 2020. Injecting numerical reasoning skills into language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 946–958, Online. Association for Computational Linguistics.
  23. 23.Nitish Gupta, Kevin Lin, Dan Roth, Sameer Singh, and Matt Gardner. 2019. Neural module networks for reasoning over text. CoRR, abs/1912.04971.
  24. 24.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations.
  25. 25.Chadi Helwe, Chloé Clavel, and Fabian M. Suchanek. 2021. Reasoning with transformer-based models: Deep learning, but shallow reasoning. In 3rd Conference on Automated Knowledge Base Construction.
  26. 26.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Eisenschlos. 2020. TaPas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4320–4333, Online. Association for Computational Linguistics.
  27. 27.Minghao Hu, Yuxing Peng, Zhen Huang, and Dongsheng Li. 2019. A multi-type multi-span network for reading comprehension that requires discrete reasoning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1596–1606, Hong Kong, China. Association for Computational Linguistics.
  28. 28.Yinya Huang, Meng Fang, Yu Cao, Liwei Wang, and Xiaodan Liang. 2021. DAGN: Discourse-aware graph network for logical reasoning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5848–5855, Online. Association for Computational Linguistics.
  29. 29.Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2020. SpanBERT: Improving pre-training by representing and predicting spans. Transactions of the Association for Computational Linguistics, 8:64–77.
  30. 30.Tushar Khot, Daniel Khashabi, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2021. Text modular networks: Learning to decompose tasks in the language of existing models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1264–1279, Online. Association for Computational Linguistics.
  31. 31.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.
  32. 32.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  33. 33.Moxin Li, Fuli Feng, Hanwang Zhang, Xiangnan He, Fengbin Zhu, and Tat-Seng Chua. 2022a. Learning to imagine: Integrating counterfactual thinking in neural discrete reasoning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 57–69, Dublin, Ireland. Association for Computational Linguistics.
  34. 34.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022b. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336.
  35. 35.Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020. Logiqa: A challenge dataset for machine reading comprehension with logical reasoning. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20, pages 3622–3628. International Joint Conferences on Artificial Intelligence Organization. Main track.
  36. 36.Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022. TAPEX: Table pre-training via learning a neural SQL executor. In International Conference on Learning Representations.
  37. 37.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  38. 38.Augustus Odena, Kensen Shi, David Bieber, Rishabh Singh, Charles Sutton, and Hanjun Dai. 2020. Bustle: Bottom-up program synthesis through learning-guided exploration.
  39. 39.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53, Minneapolis, Minnesota. Association for Computational Linguistics.
  40. 40.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2080–2094, Online. Association for Computational Linguistics.
  41. 41.Lin Qiu, Yunxuan Xiao, Yanru Qu, Hao Zhou, Lei Li, Weinan Zhang, and Yong Yu. 2019. Dynamically fused graph network for multi-hop reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6140–6150, Florence, Italy. Association for Computational Linguistics.
  42. 42.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  43. 43.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susan-nah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant M. Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake A. Hechtman, Laura Weidinger, Iason Gabriel, William S. Isaac, Edward Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training gopher. CoRR, abs/2112.11446.
  44. 44.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  45. 45.Qiu Ran, Yankai Lin, Peng Li, Jie Zhou, and Zhiyuan Liu. 2019. NumNet: Machine reading comprehension with numerical reasoning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2474–2484, Hong Kong, China. Association for Computational Linguistics.
  46. 46.Abhilasha Ravichander, Aakanksha Naik, Carolyn Rose, and Eduard Hovy. 2019. EQUATE: A benchmark evaluation framework for quantitative reasoning in natural language inference. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 349–361, Hong Kong, China. Association for Computational Linguistics.
  47. 47.Hongyu Ren, Hanjun Dai, Bo Dai, Xinyun Chen, Michihiro Yasunaga, Haitian Sun, Dale Schuurmans, Jure Leskovec, and Denny Zhou. 2021. Lego: Latent execution-guided reasoning for multi-hop question answering on knowledge graphs. In ICML.
  48. 48.Michael Scriven. 1976. Reasoning. New York: McGraw-Hill.
  49. 49.Elad Segal, Avia Efrat, Mor Shoham, Amir Globerson, and Jonathan Berant. 2020. A simple and effective model for answering multi-span questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3074–3080, Online. Association for Computational Linguistics.
  50. 50.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015. Neural machine translation of rare words with subword units. CoRR, abs/1508.07909.
  51. 51.Nan Shao, Yiming Cui, Ting Liu, Shijin Wang, and Guoping Hu. 2020. Is Graph Structure Necessary for Multi-hop Question Answering? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7187–7192, Online. Association for Computational Linguistics.
  52. 52.Tianze Shi, Chen Zhao, Jordan Boyd-Graber, Hal Daumé III, and Lillian Lee. 2020. On the potential of lexico-logical alignments for semantic parsing to SQL queries. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1849–1864, Online. Association for Computational Linguistics.
  53. 53.Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams. 2021. UnNatural Language Inference. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7329–7346, Online. Association for Computational Linguistics.
  54. 54.Shao-Hua Sun, Hyeonwoo Noh, Sriram Somasundaram, and Joseph Lim. 2018. Neural program synthesis from diverse demonstration videos. In International Conference on Machine Learning, pages 4790–4799. PMLR.
  55. 55.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota. Association for Computational Linguistics.
  56. 56.Yonglong Tian, Andrew Luo, Xingyuan Sun, Kevin Ellis, William T. Freeman, Joshua B. Tenenbaum, and Jiajun Wu. 2019. Learning to infer and execute 3d shape programs.
  57. 57.Ming Tu, Kevin Huang, Guangtao Wang, Jing Huang, Xiaodong He, and Bowen Zhou. 2020. Select, answer and explain: Interpretable multi-hop reading comprehension over multiple documents. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 9073–9080. AAAI Press.
  58. 58.Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, and Matt Gardner. 2019. Do NLP models know numbers? probing numeracy in embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5307–5315, Hong Kong, China. Association for Computational Linguistics.
  59. 59.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018a. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
  60. 60.Chenglong Wang, Kedar Tatwawadi, Marc Brockschmidt, Po-Sen Huang, Yi Xin Mao, Oleksandr Polozov, and Rishabh Singh. 2018b. Robust text-to-sql generation with execution-guided decoding. ArXiv, abs/1807.03100.
  61. 61.Shuohang Wang, Mo Yu, Jing Jiang, and Shiyu Chang. 2018c. A co-matching model for multi-choice reading comprehension. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 746–751, Melbourne, Australia. Association for Computational Linguistics.
  62. 62.Siyuan Wang, Wanjun Zhong, Duyu Tang, Zhongyu Wei, Zhihao Fan, Daxin Jiang, Ming Zhou, and Nan Duan. 2022. Logic-driven context extension and data augmentation for logical reasoning of text. In Findings of the Association for Computational Linguistics: ACL 2022, pages 1619–1629, Dublin, Ireland. Association for Computational Linguistics.
  63. 63.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  64. 64.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  65. 65.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  66. 66.Tomer Wolfson, Mor Geva, Ankit Gupta, Matt Gardner, Yoav Goldberg, Daniel Deutch, and Jonathan Berant. 2020. Break it down: A question understanding benchmark. Transactions of the Association for Computational Linguistics, 8:183–198.
  67. 67.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  68. 68.Ori Yoran, Alon Talmor, and Jonathan Berant. 2022. Turning tables: Generating examples from semi-structured tables for endowing language models with reasoning skills. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6016–6031, Dublin, Ireland. Association for Computational Linguistics.
  69. 69.Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2020. Reclor: A reading comprehension dataset requiring logical reasoning. In International Conference on Learning Representations (ICLR).
  70. 70.Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. SWAG: A large-scale adversarial dataset for grounded commonsense inference. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 93–104, Brussels, Belgium. Association for Computational Linguistics.
  71. 71.Jing Zhang, Bo Chen, Lingxi Zhang, Xirui Ke, and Haipeng Ding. 2021. Neural, symbolic and neural-symbolic reasoning on knowledge graphs. AI Open, 2:14–35.
  72. 72.Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2SQL: Generating structured queries from natural language using reinforcement learning. arXiv, abs/1709.00103.
  73. 73.Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3277–3287, Online. Association for Computational Linguistics.
  74. 74.Amit Zohar and Lior Wolf. 2018. Automatic program synthesis of long programs with a learned garbage collector. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pages 2098–2107.

Citation

MLA
Pi, X., et al. “Reasoning Like Program Executors”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 761–79, https://doi.org/10.18653/v1/2022.emnlp-main.48.
APA
Pi, X., Liu, Q., Chen, B., Ziyadi, M., Lin, Z., Fu, Q., Gao, Y., Lou, J.-G., & Chen, W. (2022). Reasoning Like Program Executors. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 761–779. https://doi.org/10.18653/v1/2022.emnlp-main.48
Chicago
Pi, X., Q. Liu, B. Chen, et al. 2022. “Reasoning Like Program Executors”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 761–79. https://doi.org/10.18653/v1/2022.emnlp-main.48.
Harvard
Pi, X. et al. (2022) “Reasoning Like Program Executors”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 761–779. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.48.
Vancouver
1. Pi X, Liu Q, Chen B, Ziyadi M, Lin Z, Fu Q, Gao Y, Lou J-G, Chen W (2022) Reasoning Like Program Executors. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 761–779

BibTeX

@inproceedings{pi-etal-2022-reasoning,
    title = "Reasoning Like Program Executors",
    author = "Pi, Xinyu  and
      Liu, Qian  and
      Chen, Bei  and
      Ziyadi, Morteza  and
      Lin, Zeqi  and
      Fu, Qiang  and
      Gao, Yan  and
      Lou, Jian-Guang  and
      Chen, Weizhu",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.48/",
    doi = "10.18653/v1/2022.emnlp-main.48",
    pages = "761--779"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/