A Survey of Deep Learning for Mathematical Reasoning

Pan LuLiang QiuWenhao YuSean WelleckKai-Wei Chang

article2023ACL209 citations

Provides a systematic taxonomy and evaluation of over 180 studies covering neural architectures, large language models, and benchmarks for automated mathematical problem solving and theorem proving.

Listen

Mathematical reasoning is a fundamental pillar of human intelligence and decision-making across engineering, science, and finance. While artificial intelligence systems have achieved notable success in general language tasks, enabling machines to reliably solve math word problems, prove formal theorems, and interpret multi-modal geometric figures remains a critical bottleneck. The article aims to establish a clear taxonomy of mathematical reasoning tasks, evaluate the performance and limitations of deep learning methods over the past decade, and provide actionable future directions for the field.

The authors conducted a comprehensive review of over 180 studies published between 2013 and 2022 across the machine learning and natural language processing communities. The analysis categorizes benchmarks into key areas such as math word problems, formal and informal theorem proving, geometry problem solving, and general quantitative question answering. It systematically examines three generations of methods: specialized neural networks (sequence-to-sequence, graph-based, and attention architectures), pre-trained language models, and recent large language models utilizing in-context learning and chain-of-thought prompting.

The review identifies critical breakthroughs alongside significant technical vulnerabilities. First, while scaling models enhances performance—exemplified by Google's Minerva achieving 75.0% on advanced STEM benchmarks and GPT-3 reaching 93.0% on MultiArith—state-of-the-art systems remain brittle. When tested on SVAMP, an elementary benchmark with slight phrasing variations, top systems degrade sharply, with Graph2Tree achieving only 43.8% and GPT-3 reaching 63.7%. Second, current tokenization methods inherently fail at numerical representation; models break multi-digit numbers into arbitrary sub-word fragments, leading to systematic calculation errors and poor generalization on large numbers. Third, outcome-based approaches (such as self-consistency sampling) and process-based frameworks (like decomposing problems into sub-tasks or delegating math execution to external computer programs) significantly outperform standard single-pass prompting. Finally, the research landscape remains heavily biased toward text-only English datasets, leaving multi-modal contexts (diagrams and tables) and low-resource domains severely underexplored.

These findings indicate that high benchmark scores can mask a lack of true reasoning capability. The presence of ungrounded outputs and hallucinated logical steps creates major operational and safety risks for organizations attempting to deploy AI for automated finance, engineering, or scientific analytics without oversight. Because systems are highly sensitive to superficial prompt changes and struggle with basic number representations, standard language models cannot be trusted as standalone mathematical solvers.

To address these limitations, stakeholders and practitioners should avoid relying on single-pass model outputs and instead implement hybrid workflows, such as program-aided execution that routes deterministic calculations to verified computing engines. Decision-makers should support research into better numerical encoding formats (such as scientific notation), expand multi-modal benchmark datasets, and integrate reinforcement learning from human or formal theorem feedback. Because this survey focuses on published literature through 2022, continued caution and empirical pilots are necessary as model capabilities rapidly evolve.

Cover for A Survey of Deep Learning for Mathematical Reasoning

Abstract

Mathematical reasoning is a fundamental aspect of human intelligence and is applicable in various fields, including science, engineering, finance, and everyday life. The development of artificial intelligence (AI) systems capable of solving math problems and proving theorems in language has garnered significant interest in the fields of machine learning and natural language processing. For example, mathematics serves as a testbed for aspects of reasoning that are challenging for powerful deep learning models, driving new algorithmic and modeling advances. On the other hand, recent advances in large-scale neural language models have opened up new benchmarks and opportunities to use deep learning for mathematical reasoning. In this survey paper, we review the key tasks, datasets, and methods at the intersection of mathematical reasoning and deep learning over the past decade. We also evaluate existing benchmarks and methods, and discuss future research directions in this domain.

Table of Contents

  • 1 Introduction
  • 2 Mathematical Reasoning Tasks
  • 3 Neural Networks for Mathematical Reasoning
  • 3.1 Seq2Seq-based Networks for Math
  • 3.2 Graph-based Networks for Math
  • 3.3 Attention-based Networks for Math
  • 3.4 Other Neural Networks for Math
  • 4 Pre-trained Language Models for Mathematical Reasoning
  • 4.1 Self-Supervised Learning for Math
  • 4.2 Task-specific Fine-tuning for Math
  • 5 In-context Learning for Mathematical Reasoning
  • 5.1 In-context Example Selection
  • 5.2 High-quality Reasoning Chains
  • 6 Discussion and Findings
  • 6.1 Analysis of Benchmarks
  • 6.2 Analysis of Deep Learning Methods
  • 7 Future Work
  • 7.1 Generalization and Robustness
  • 7.2 Trustworthy Reasoning
  • 7.3 Learning from Feedback
  • 7.4 Multi-modal Mathematical Reasoning
  • 8 Conclusion
  • Limitations
  • Broader Impact
  • References
  • A Mathematical Reasoning Datasets
  • A.1 Math Word Problem Solving
  • A.2 Theorem Proving
  • A.3 Geometry Problem Solving
  • A.4 Math Question Answering
  • A.5 Other Quantitative Problems

Knowls

  1. Knowl 1 — Taxonomy of Mathematical Reasoning Tasks and Benchmark Categories

    definition

    Mathematical reasoning tasks within deep learning research are categorized into five core areas based on problem structure and input-output modalities:

    1. Math Word Problem (MWP) Solving: Systems solve narrative word problems involving characters, entities, and numerical quantities that require one or multiple arithmetic operations. Sub-types include purely textual MWPs (e.g., MathQA, SVAMP, GSM8K, Ape210K) and multimodal MWPs grounded in diagrams or tables (e.g., IconQA, TabMWP).
    2. Theorem Proving (TP): Systems prove mathematical conjectures through deductive reasoning steps. Formulations include formal theorem proving inside Interactive Theorem Provers (ITPs such as Lean, Isabelle, Coq, and Metamath; e.g., CoqGym, PISA, miniF2F), informal theorem proving written in natural language and LaTeX\LaTeX (e.g., NaturalProofs), and hybrid formal/informal proving focused on autoformalization or informal-to-formal proof translation.
    3. Geometry Problem Solving (GPS): Systems process multimodal inputs comprising geometric diagrams alongside natural language descriptions to deduce numeric values, angles, or geometric relationships (e.g., Geometry3K, GeoQA, UniGeo).
    4. Math Question Answering (MathQA): Single-task and unified multi-task reading comprehension and question answering benchmarks that require discrete operations, quantitative commonsense, or mathematical domain knowledge (e.g., DROP, Mathematics, MATH, NumGLUE, Lila, TheoremQA).
    5. Other Quantitative Problems: Specialized quantitative reasoning tasks, such as visual diagram reasoning (FigureQA, DVQA), conversational financial reporting (ConvFinQA, FinQA, TAT-QA), scientific question answering (ScienceQA), and algorithmic program synthesis puzzles (P3).
  2. Knowl 2 — Neural Network Architectural Paradigms for Mathematical Reasoning

    model/method

    Deep neural network architectures for mathematical reasoning are organized across several core structural paradigms:

    • Sequence-to-Sequence (Seq2Seq): Models encode mathematical problem text using recurrent neural networks (e.g., LSTM, GRU, BiLSTM, BiGRU) and generate output token sequences (equations, programs, or proofs) sequentially.
    • Graph- and Tree-Structured Networks:
      • Sequence-to-Tree (Seq2Tree): Models use recurrent encoders to process input text while employing tree-structured decoders to generate expression syntax trees (ASTs) in goal-driven or prefix order.
      • Sequence-to-Directed Acyclic Graph (Seq2DAG): Models decode mathematical equations into DAG structures to capture complex multi-variable relationships and shared intermediate sub-expressions.
      • Graph-to-Tree (Graph2Tree): Models represent input relationships as semantic or dependency graphs via graph neural networks (GNNs) and decode target solutions as expression trees.
    • Attention-Enhanced Networks: Models incorporate self-attention or multi-head attention (e.g., Group-ATT) to resolve long-range token dependencies, align mathematical quantities with entities, and fuse knowledge graphs.
    • Specialized Architectures:
      • Convolutional Neural Networks (CNNs) and WaveNets applied to premise selection and proof search in theorem proving.
      • Visual Question Answering (VQA) fusion backbones (combining ResNet/Faster R-CNN visual representations with language embeddings) for geometry and chart parsing.
      • Deep Q-Networks (DQN) for navigating arithmetic search spaces.
  3. Knowl 3 — Pre-training Data Sources and Objectives for Mathematical Language Models

    model/method

    Language models adapted for mathematical reasoning rely on specialized pre-training configurations:

    • Scale Thresholding: Empirical analysis indicates model scale correlates with mathematical competence, with models exceeding 50 billion parameters demonstrating distinct thresholding capabilities over complex multi-step reasoning (e.g., Minerva 540B based on PaLM).
    • Pre-training Corpora:
      1. Curated Natural Corpora: Webpages filtered for mathematical content, arXiv preprints, and step-by-step educational datasets containing natural text and LaTeX\LaTeX formulations (such as AMPS, comprising Khan Academy and Mathematica solutions).
      2. Synthetic and Engine-Generated Corpora: Synthesized theorem libraries generated via interactive provers (e.g., Isabelle in LISA), executable SQL query-result pairs (TaPEx), and template-generated numerical problems (GenBERT).
    • Pre-training Objectives:
      • Standard Self-Supervised Tasks: Masked Language Modeling (MLM) and Causal Language Modeling (CLM) executed over scientific and formal corpora.
      • Customized Numeracy/Reasoning Objectives: Specialized multi-task suites incorporating numerical properties and logic operations (MWP-BERT with 8 numeracy-augmented tasks), as well as curriculum learning of foundational reasoning primitives including deduction, induction, and abduction (e.g., LIME).
  4. Knowl 4 — Prompting and In-Context Learning Strategies for Large Language Models in Mathematics

    model/method

    In-context learning (ICL) prompts large language models (LLMs) with few-shot demonstrations at inference time without modifying model weights. Strategies for mathematical reasoning fall into two complementary components:

    • In-Context Example Selection:
      • Semantic Retrieval: Dynamically retrieving candidate exemplars that are semantically nearest to the input query.
      • Complexity-Based Prompting: Selecting demonstrations containing high reasoning complexity (e.g., chains with greater numbers of intermediate steps).
      • Policy-Guided Selection: Optimizing prompt selection from candidate pools via reinforcement learning policy gradients (e.g., PromptPG).
    • Reasoning Chain Generation:
      • Process-Based Approaches: Structuring demonstrations into multi-stage decompositions, such as least-to-most prompting (breaking a problem into sequential sub-problems where subsequent queries consume prior solutions) or Program-of-Thoughts (PoT) / Program-Aided Language Models (PAL) (generating executable code whose calculation is delegated to external interpreters).
      • Outcome-Based Approaches: Mitigating single-path decoding errors through multi-path sampling and consistency aggregation (e.g., Self-Consistency, which samples diverse reasoning paths and selects the marginal consensus answer via majority vote) or prompt diversification through self-teaching.
  5. Knowl 5 — In-Context Learning Paradigms for Mathematical Reasoning with Large Language Models

    data/table

    The operational landscape of chain-of-thought (CoT) and programmatic in-context learning frameworks for mathematical reasoning is summarized below across underlying model engines, demonstration selection sources, rationale representations, and post-processing methods:

    Model Framework Engine ICL Source Rationale Type Rationale Source Post Method
    Few-shot-CoT PaLM (540B) Random Language Hand-crafted –
    Self-Consistency-CoT Codex (175B) Random Language Hand-crafted Self-consistency
    Least-to-most CoT Codex (175B) Random Language Hand-crafted –
    PromptPG-CoT GPT-3 (175B) RL Language Hand-crafted –
    Retrieval-CoT GPT-3 (175B) Retrieval Language Auto-generated –
    Auto-CoT Codex (175B) Clustering Language Auto-generated –
    Complexity-CoT GPT-3 (175B) Complexity Language Hand-crafted Self-consistency
    Few-shot-PoT GPT-3 (175B) Random Code Hand-crafted –

    All evaluated GPT-3 instances utilize text-davinci-002, and Codex instances utilize code-davinci-002. Hand-crafted rationales denote human-written step-by-step explanations, whereas code rationales delegate arithmetic execution directly to an external Python interpreter.

  6. Knowl 6 — Subword Tokenization Pathologies and Multi-Digit Arithmetic Degradation

    empirical result

    Standard deep learning language models represent numbers via standard subword tokenization schemes (such as Byte-Pair Encoding or WordPiece), creating inconsistent, non-continuous token boundaries for numerically adjacent values. For example, the integer 15981598 is tokenized as ["15", "98"], whereas its comma-formatted equivalent 1,5981,598 is split into three distinct tokens ["1", ",", "598"].

    This lack of structural and spatial numeracy representation causes severe out-of-distribution performance degradation as numerical magnitude increases, as demonstrated across multiple model architectures on identical arithmetic structures:

    Arithmetic Input T5 (Large) UnifiedQA (Large) GPT-3 (text-davinci-002) GPT-3 (text-davinci-003)
    3 balls + 5 balls = Incorrect 5 balls (Incorrect) 8 balls (Correct) 8 balls (Correct)
    23 balls + 145 balls = Incorrect Incorrect 58 balls (Incorrect) 168 balls (Correct)
    23 balls + 1,855 balls = Incorrect Incorrect 2,878 balls (Incorrect) 2,988 balls (Incorrect)

    Even scaled models like GPT-3 (text-davinci-003) fail when arithmetic arguments reach 4-digit numbers (23+1,855=1,878≠2,98823 + 1,855 = 1,878 \neq 2,988).

  7. Knowl 7 — Semantic Inconsistency and Adversarial Sensitivity in Large Language Model Reasoning

    empirical result

    While scaled models achieve high accuracy on standard benchmarks (e.g., zero-shot CoT Minerva 540B achieves 75.0% on MMLU-STEM and few-shot CoT GPT-3 175B reaches 93.0% on MultiArith), their reasoning remains fragile when subjected to minor semantic perturbations or adversarial rephrasings.

    On the SVAMP benchmark—which introduces simple word variations to elementary arithmetic problems—state-of-the-art neural methods exhibit drastic accuracy drops: Graph2Tree drops to 43.8% and zero-shot CoT GPT-3 (175B) reaches only 63.7%.

    Furthermore, zero-shot GPT-3 (text-davinci-002) outputs logically contradictory answers under minor variations of problem phrasing:

    Problem Prompt GPT-3 Prediction
    John had 8 balls and he gave 3 to Mary. How many balls does John have now? John has 5 balls. (Correct)
    John had 3 apples. John had 8 balls and he gave 3 to Mary. How many balls does Mary have now? Mary has 5 balls. (Incorrect)
    John had 8 balls and he gave 3 to Mary. Who has more balls now? John has more balls. (Correct)
    John had 8 balls and he gave 3 to Mary. Does John have more balls now? No, John has 5 balls now. (Inconsistent)
    John had 8 balls and he gave 4 to Mary. Does John have more balls now? No, John has 4 balls now. (Correct)
    John had 8 balls and he gave 4 to Mary. Who has more balls now? John has more balls. (Inconsistent)

    These inconsistencies demonstrate that correct predictions often stem from spurious surface patterns rather than invariant deductive reasoning.

  8. Knowl 8 — Taxonomy of Rationale Formalisms in Mathematical Reasoning Benchmarks

    model/method

    Benchmark annotations for intermediate reasoning steps have evolved across four primary rationale representations to improve solver interpretability and multi-step deduction:

    1. Mathematical Equations / Expression Trees: Numerical and algebraic equation templates providing the direct formula required to calculate the answer (e.g., Math23K, Ape210K, ASDiv, SVAMP).
    2. Domain-Specific Programs and Logic Forms: Executable domain-specific operations or formal logic representations that map structured linguistic entities into deterministic symbolic operations (e.g., MathQA, Geometry3K, QuaRel).
    3. General-Purpose Programming Scripts: Python code snippets that encode reasoning logic programmatically, offloading calculation to Python interpreters and enhancing readability (e.g., MathQA-Python, Lila).
    4. Natural Language Chains-of-Thought: Multi-step, human-written explanatory sequences in natural language and LaTeX\LaTeX describing the step-by-step logic and arithmetic deductions (e.g., GSM8K, MATH, ScienceQA, TabMWP).
  9. Knowl 9 — Future Research Directions for Deep Learning in Mathematical Reasoning

    model/method

    Four key open frontiers are identified for advancing deep learning in mathematical reasoning:

    1. Generalization, Robustness, and Disentangling Memorization: Developing inference-time search algorithms and training strategies that generalize to longer reasoning chains and higher numerical values, while disentangling training corpus term frequencies and memorization from true multi-step deductive problem solving.
    2. Trustworthy and Verifiable Reasoning: Mitigating ungrounded hallucinations by grounding model outputs in formal theorems, incorporating calibrated uncertainty mechanisms for unknown states, and developing dedicated verification modules that locate intermediate deduction errors.
    3. Learning from Multi-Source Feedback: Extending reinforcement learning beyond human preference ranking (RLHF) to incorporate deterministic automated feedback from formal interactive theorem provers (e.g., Lean, Isabelle, Coq REPLs) and code execution environments.
    4. Multi-modal Mathematical Integration: Bridging the semantic gap between natural-image visual encoders and abstract technical diagrams, plots, and hierarchical tables through unified multimodal reasoning architectures.
  10. Knowl 10 — Scope and Systematic Limitations of the Mathematical Reasoning Survey

    limitation

    The survey outlines three explicit systematic limitations:

    1. Temporal and Scope Focus: The review focuses primarily on literature at the intersection of deep learning and mathematical reasoning from the decade spanning 2013 to 2022, omitting extensive historical analysis of pre-deep learning classical and symbolic AI systems.
    2. Curated Selection: Analysis is based on a selected set of over 180 representative papers and benchmarks, which may not encompass every niche implementation or dataset.
    3. Pace of Model Evolution: Due to the rapid pace of large language model development, emerging foundation models and concurrent prompting techniques may surpass specific quantitative benchmarks established during the survey window.

Coverage note — None was omitted; the extracted knowls cover the complete taxonomy of mathematical tasks, neural network architectures, pre-training methodologies, in-context learning paradigms, benchmark rationale formalisms, empirical analyses of numeracy and consistency, future research directions, and stated limitations.

References

  1. 1.Alexander A. Alemi, François Chollet, Niklas Een, Geoffrey Irving, Christian Szegedy, and Josef Urban. 2016. Deepmath - deep sequence models for premise selection. Advances in neural information processing systems (NeurIPS), 29.
  2. 2.Reem Alghamdi, Zhenwen Liang, and Xiangliang Zhang. 2022. Armath: a dataset for solving arabic math word problems. In Proceedings of the Thirteenth Language Resources and Evaluation Conference (LREC), pages 351–362.
  3. 3.Chris Alvin, Sumit Gulwani, Rupak Majumdar, and Supratik Mukhopadhyay. 2017. Synthesis of solutions for shaded area geometry problems. In The Thirtieth International Flairs Conference.
  4. 4.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 2357–2367.
  5. 5.Connor Anderson and Ryan Farrell. 2022. Improving fractal pre-training. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1300–1309.
  6. 6.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), pages 6077–6086.
  7. 7.Cem Anil, Yuhuai Wu, Anders Johan Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Venkatesh Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. 2022. Exploring length generalization in large language models. In Advances in Neural Information Processing Systems (NeurIPS).
  8. 8.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732.
  9. 9.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations (ICLR).
  10. 10.Kshitij Bansal, Sarah Loos, Markus Rabe, Christian Szegedy, and Stewart Wilcox. 2019. Holist: An environment for machine learning of higher order logic theorem proving. In International Conference on Machine Learning (ICML), pages 454–463. PMLR.
  11. 11.Bruno Barras, Samuel Boutin, Cristina Cornes, Judicaël Courant, Yann Coscoy, David Delahaye, Daniel de Rauglaudre, Jean-Christophe Filliâtre, Eduardo Giménez, Hugo Herbelin, et al. 1999. The coq proof assistant reference manual. INRIA, version, 6(11).
  12. 12.Taylor Berg-Kirkpatrick and Daniel Spokoyny. 2020. An empirical investigation of contextualized number prediction. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4754–4764.
  13. 13.Arindam Bhattacharya. 2017. A survey of question answering for math and science problem. arXiv preprint arXiv:1705.04530.
  14. 14.Daniel G Bobrow. 1964. Natural language input for a computer problem solving system. AI Technical Reports.
  15. 15.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS), 33:1877–1901.
  16. 16.Jie Cao and Jing Xiao. 2022. An augmented benchmark dataset for geometric question answering through dual parallel text encoding. In Proceedings of the 29th International Conference on Computational Linguistics (COLING), pages 1511–1520.
  17. 17.Yixuan Cao, Feng Hong, Hongwei Li, and Ping Luo. 2021. A bottom-up dag structure extraction model for math word problems. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), pages 39–46.
  18. 18.François Charton. 2022. Linear algebra with transformers. Transactions on Machine Learning Research.
  19. 19.Jiaqi Chen, Tong Li, Jinghui Qin, Pan Lu, Liang Lin, Chongyu Chen, and Xiaodan Liang. 2022a. Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression. In The 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  20. 20.Jiaqi Chen, Jianheng Tang, Jinghui Qin, Xiaodan Liang, Lingbo Liu, Eric Xing, and Liang Lin. 2021a. Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning. In Findings of the Association for Computational Linguistics (ACL), pages 513–523.
  21. 21.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021b. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  22. 22.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022b. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588.
  23. 23.Wenhu Chen, Ming Yin, Max Ku, Elaine Wan, Xueguang Ma, Jianyu Xu, Tony Xia, Xinyi Wang, and Pan Lu. 2023. Theoremqa: A theorem-driven question answering dataset. arXiv preprint arXiv:2305.12524.
  24. 24.Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan R Routledge, et al. 2021c. Finqa: A dataset of numerical reasoning over financial data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3697–3711.
  25. 25.Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. 2022c. Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering. arXiv preprint arXiv:2210.03849.
  26. 26.Ting-Rui Chiang and Yun-Nung Chen. 2019. Semantically-aligned equation generation for solving and reasoning math word problems. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 2656–2668.
  27. 27.Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021. Unifying vision-and-language tasks via text generation. In Proceedings of the 38th International Conference on Machine Learning (ICML), pages 1931–1942.
  28. 28.Kyunghyun Cho, Bart van Merrienboer Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using rnn encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1724–1734.
  29. 29.Shang-Ching Chou, Xiao-Shan Gao, and Jing-Zhong Zhang. 1996. Automated generation of readable proofs with geometric invariants. Journal of Automated Reasoning, 17(3):325–347.
  30. 30.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  31. 31.Peter Clark, Oren Etzioni, Tushar Khot, Daniel Khashabi, Bhavana Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, Niket Tandon, et al. 2020. From ‘f’to ‘a’on the ny regents science exams: An overview of the aristo project. AI Magazine, 41(4):39–53.
  32. 32.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  33. 33.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 4171–4186.
  34. 34.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 2368–2378.
  35. 35.Edward A Feigenbaum et al. 1963. Computers and thought. McGraw-Hill.
  36. 36.Yu Feng, Jing Zhang, Xiaokang Zhang, Lemao Liu, Cuiping Li, and Hong Chen. 2021. Injecting numerical reasoning skills into knowledge base question answering models. arXiv preprint arXiv:2112.06109.
  37. 37.Deborah Ferreira and André Freitas. 2020a. Natural language premise selection: Finding supporting statements for mathematical text. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 2175–2182.
  38. 38.Deborah Ferreira and André Freitas. 2020b. Premise selection in natural language mathematical texts. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 7365–7374.
  39. 39.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2023. Complexity-based prompting for multi-step reasoning. In International Conference on Learning Representations (ICLR).
  40. 40.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
  41. 41.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2022. Pal: Program-aided language models. arXiv preprint arXiv:2211.10435.
  42. 42.Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven CH Hoi, Xiaogang Wang, and Hongsheng Li. 2019. Dynamic fusion with intra-and inter-modality attention flow for visual question answering. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 6639–6648.
  43. 43.Thibault Gauthier, Cezary Kaliszyk, Josef Urban, Ramana Kumar, and Michael Norrish. 2021. TacticToe: Learning to Prove with Tactics. Journal of Automated Reasoning.
  44. 44.Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. 2017. Convolutional sequence to sequence learning. In International conference on machine learning (ICML), pages 1243–1252. PMLR.
  45. 45.Herbert Gelernter, James R Hansen, and Donald W Loveland. 1960. Empirical explorations of the geometry theorem machine. In Papers presented at the May 3-5, 1960, western joint IRE-AIEE-ACM computer conference, pages 143–149.
  46. 46.Mor Geva, Ankit Gupta, and Jonathan Berant. 2020. Injecting numerical reasoning skills into language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 946–958.
  47. 47.Kevin Gimpel, Dipanjan Das, and Noah A Smith. 2010. Distributed asynchronous online learning for natural language processing. In Proceedings of the Fourteenth Conference on Computational Natural Language Learning, pages 213–222.
  48. 48.Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al. 2022. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375.
  49. 49.Adam Grabowski, Artur Korniłowicz, and Adam Naumowicz. 2015. Four decades of mizar. Journal of Automated Reasoning, 55(3):191–198.
  50. 50.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International Conference on Machine Learning (ICML), pages 3929–3938. PMLR.
  51. 51.Jesse Michael Han, Jason Rute, Yuhuai Wu, Edward W Ayers, and Stanislas Polu. 2022. Proof artifact co-training for theorem proving with language models. In International Conference on Learning Representations (ICLR).
  52. 52.Yihan Hao, Mingliang Zhang, Fei Yin, and Linlin Huang. 2022. Pgdp5k: A diagram parsing dataset for plane geometry problems. In 26th International Conference on Pattern Recognition (ICPR).
  53. 53.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), pages 770–778.
  54. 54.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021a. Measuring massive multitask language understanding. In International Conference on Learning Representations (ICLR).
  55. 55.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021b. Measuring mathematical problem solving with the math dataset. In 35th Conference on Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks.
  56. 56.Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song. 2020. Pretrained transformers improve out-of-distribution robustness. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 2744–2751.
  57. 57.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Mueller, Francesco Piccinno, and Julian Eisenschlos. 2020. Tapas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4320–4333.
  58. 58.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation, 9(8):1735–1780.
  59. 59.Yining Hong, Qing Li, Daniel Ciao, Siyuan Huang, and Song-Chun Zhu. 2021a. Learning by fixing: Solving math word problems with weak supervision. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), pages 4959–4967.
  60. 60.Yining Hong, Qing Li, Ran Gong, Daniel Ciao, Siyuan Huang, and Song-Chun Zhu. 2021b. Smart: A situation model for algebra story problems via attributed grammar. In AAAI, pages 13009–13017.
  61. 61.Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014. Learning to solve arithmetic word problems with verb categorization. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  62. 62.Daniel Huang, Prafulla Dhariwal, Dawn Song, and Ilya Sutskever. 2019. Gamepad: A learning environment for theorem proving. In International Conference on Learning Representations (ICLR).
  63. 63.Danqing Huang, Jing Liu, Chin-Yew Lin, and Jian Yin. 2018. Neural math word problem solver with reinforcement learning. In Proceedings of the 27th International Conference on Computational Linguistics (COLING), pages 213–223.
  64. 64.Danqing Huang, Shuming Shi, Chin-Yew Lin, and Jian Yin. 2017. Learning fine-grained expressions to solve math word problems. In Proceedings of Empirical Methods in Natural Language Processing (EMNLP), pages 805–814.
  65. 65.Danqing Huang, Shuming Shi, Chin-Yew Lin, Jian Yin, and Wei-Ying Ma. 2016. How well do computers solve math word problems? large-scale dataset construction and evaluation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL), pages 887–896.
  66. 66.Albert Q. Jiang, Sean Welleck, Jin Peng Zhou, Wenda Li, Jiacheng Liu, Mateja Jamnik, Timothée Lacroix, Yuhuai Wu, and Guillaume Lample. 2022a. Draft, sketch, and prove: Guiding formal theorem provers with informal proofs. In Submitted to The Eleventh International Conference on Learning Representations.
  67. 67.Albert Qiaochu Jiang, Wenda Li, Jesse Michael Han, and Yuhuai Wu. 2021. Lisa: Language models of isabelle proofs. In 6th Conference on Artificial Intelligence and Theorem Proving (AITP).
  68. 68.Albert Qiaochu Jiang, Wenda Li, Szymon Tworkowski, Konrad Czechowski, Tomasz Odrzygóźć, Piotr Miłoś, Yuhuai Wu, and Mateja Jamnik. 2022b. Thor: Wielding hammers to integrate language models and automated theorem provers. Advances in Neural Information Processing Systems (NeurIPS), 35:8360–8373.
  69. 69.Zhanming Jie, Jierui Li, and Wei Lu. 2022. Learning to reason deductively: Math word problem solving as complex relation extraction. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pages 5944–5955.
  70. 70.Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi. 2022. Maieutic prompting: Logically consistent reasoning with recursive explanations. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1266–1279.
  71. 71.Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan. 2018. Dvqa: Understanding data visualizations via question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5648–5656.
  72. 72.Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Ákos Kádár, Adam Trischler, and Yoshua Bengio. 2018. Figureqa: An annotated figure dataset for visual reasoning. In International Conference on Learning Representations (ICLR).
  73. 73.Cezary Kaliszyk, François Chollet, and Christian Szegedy. 2017. Holstep: A machine learning dataset for higher-order logic theorem proving. In International Conference on Learning Representations (ICLR).
  74. 74.Ashwin Kalyan, Abhinav Kumar, Arjun Chandrasekaran, Ashish Sabharwal, and Peter Clark. 2021. How much coffee was consumed during emnlp 2019? fermi problems: A new reasoning challenge for ai. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7318–7328.
  75. 75.Nikhil Kandpal, H. Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2022. Large language models struggle to learn long-tail knowledge. ArXiv, abs/2211.08411.
  76. 76.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. Unifiedqa: Crossing format boundaries with a single qa system. In Findings of the Association for Computational Linguistics (EMNLP), pages 1896–1907.
  77. 77.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2022. Decomposed prompting: A modular approach for solving complex tasks. arXiv preprint arXiv:2210.02406.
  78. 78.Bugeun Kim, Kyung Seo Ki, Donggeon Lee, and Gahgene Gweon. 2020. Point to the expression: Solving algebraic word problems using the expression-pointer transformer model. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3768–3779.
  79. 79.Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. 2018. Bilinear attention networks. In Advances in Neural Information Processing Systems (NeurIPS), pages 1571–1581.
  80. 80.Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML), pages 5583–5594.
  81. 81.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In 36th Conference on Neural Information Processing Systems (NeurIPS).
  82. 82.Rik Koncel-K., Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. Mawps: A math word problem repository. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL), pages 1152–1157.
  83. 83.Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. 2015. Parsing algebraic word problems into equations. Transactions of the Association for Computational Linguistics (TACL), 3:585–597.
  84. 84.Kundan Krishna, Jeffrey Bigham, and Zachary C Lipton. 2021. Does pretraining for summarization require knowledge transfer? In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3178–3189.
  85. 85.Nate Kushman, Yoav Artzi, Luke Zettlemoyer, and Regina Barzilay. 2014. Learning to automatically solve algebra word problems. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (ACL), pages 271–281.
  86. 86.Guillaume Lample and François Charton. 2020. Deep learning for symbolic mathematics. In International Conference on Learning Representations (ICLR).
  87. 87.Guillaume Lample, Timothee Lacroix, Marie-Anne Lachaux, Aurelien Rodriguez, Amaury Hayat, Thibaut Lavril, Gabriel Ebner, and Xavier Martinet. 2022. Hypertree proof search for neural theorem proving. Advances in Neural Information Processing Systems (NeurIPS), 35:26337–26349.
  88. 88.Yihuai Lan, Lei Wang, Qiyuan Zhang, Yunshi Lan, Bing Tian Dai, Yan Wang, Dongxiang Zhang, and Ee-Peng Lim. 2022. Mwptoolkit: an open-source framework for deep learning-based math word problem solvers. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), pages 13188–13190.
  89. 89.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942.
  90. 90.Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324.
  91. 91.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 7871–7880.
  92. 92.Aitor Lewkowycz, Anders Johan Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Venkatesh Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. 2022. Solving quantitative reasoning problems with language models. In Advances in Neural Information Processing Systems (NeurIPS).
  93. 93.Jierui Li, Lei Wang, Jipeng Zhang, Yan Wang, Bing Tian Dai, and Dongxiang Zhang. 2019. Modeling intra-relation in math word problems with different functional multi-head attentions. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 6162–6167.
  94. 94.Jiwei Li, Alexander H Miller, Sumit Chopra, Marc’Aurelio Ranzato, and Jason Weston. 2017. Dialogue learning with human-in-the-loop. In International Conference on Learning Representations (ICLR).
  95. 95.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2020a. What does bert with vision look at? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 5265–5275.
  96. 96.Shucheng Li, Lingfei Wu, Shiwei Feng, Fangli Xu, Fengyuan Xu, and Sheng Zhong. 2020b. Graph-to-tree neural networks for learning structured input-output translation with applications to semantic parsing and math word problem. In Findings of the Association for Computational Linguistics (EMNLP), pages 2841–2852.
  97. 97.Wenda Li, Lei Yu, Yuhuai Wu, and Lawrence C Paulson. 2021. Isarstep: a benchmark for high-level mathematical reasoning. In International Conference on Learning Representations (ICLR).
  98. 98.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022a. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336.
  99. 99.Zhongli Li, Wenxuan Zhang, Chao Yan, Qingyu Zhou, Chao Li, Hongzhi Liu, and Yunbo Cao. 2022b. Seeking patterns, not just memorizing procedures: Contrastive learning for solving math word problems. In Findings of the Association for Computational Linguistics (ACL), pages 2486–2496.
  100. 100.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022a. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
  101. 101.Percy Liang and Dan Klein. 2009. Online em for unsupervised models. In Proceedings of human language technologies: The 2009 annual conference of the North American chapter of the association for computational linguistics (NAACL), pages 611–619.
  102. 102.Zhenwen Liang, Jipeng Zhang, Lei Wang, Wei Qin, Yunshi Lan, Jie Shao, and Xiangliang Zhang. 2022b. Mwp-bert: Numeracy-augmented pre-training for math word problem solving. In Findings of the Association for Computational Linguistics (NAACL), pages 997–1009.
  103. 103.Bill Yuchen Lin, Seyeon Lee, Rahul Khanna, and Xiang Ren. 2020. Birds have four legs?! numersense: Probing numerical commonsense knowledge of pre-trained language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6862–6868.
  104. 104.Xin Lin, Zhenya Huang, Hongke Zhao, Enhong Chen, Qi Liu, Hao Wang, and Shijin Wang. 2021. Hms: A hierarchical solver with dependency-enhanced understanding for math word problem. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), pages 4232–4240.
  105. 105.Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL), pages 158–167.
  106. 106.Jiachang Liu, Dinghan Shen, Yizhe Zhang, William B Dolan, Lawrence Carin, and Weizhu Chen. 2022a. What makes good in-context examples for gpt-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114.
  107. 107.Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022b. TAPEX: Table pre-training via learning a neural SQL executor. In International Conference on Learning Representations.
  108. 108.Qianying Liu, Wenyu Guan, Sujian Li, Fei Cheng, Daisuke Kawahara, and Sadao Kurohashi. 2020. Reverse operation based data augmentation for solving math word problems. IEEE Transactions on Audio, Speech and Language Processing.
  109. 109.Qianying Liu, Wenyv Guan, Sujian Li, and Daisuke Kawahara. 2019a. Tree-structured decoding for solving math word problems. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pages 2370–2379.
  110. 110.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b. Roberta: A robustly optimized bert pretraining approach. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT).
  111. 111.Sarah Loos, Geoffrey Irving, Christian Szegedy, and Cezary Kaliszyk. 2017. Deep network guided proof search. arXiv preprint arXiv:1701.06972.
  112. 112.Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-Chun Zhu. 2021a. Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning. In The 59th Annual Meeting of the Association for Computational Linguistics (ACL).
  113. 113.Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022a. Learn to explain: Multimodal reasoning via thought chains for science question answering. In The 36th Conference on Neural Information Processing Systems (NeurIPS).
  114. 114.Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023. Chameleon: Plug-and-play compositional reasoning with large language models. arXiv preprint arXiv:2304.09842.
  115. 115.Pan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, and Ashwin Kalyan. 2022b. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. In International Conference on Learning Representations (ICLR).
  116. 116.Pan Lu, Liang Qiu, Jiaqi Chen, Tony Xia, Yizhou Zhao, Wei Zhang, Zhou Yu, Xiaodan Liang, and Song-Chun Zhu. 2021b. Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning. In The 35th Conference on Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks.
  117. 117.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022c. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pages 8086–8098.
  118. 118.The mathlib Community. 2020. The lean mathematical library. In CPP 2020 - Proceedings of the 9th ACM SIGPLAN International Conference on Certified Programs and Proofs, co-located with POPL 2020.
  119. 119.Jordan Meadows and Andre Freitas. 2022. A survey in mathematical language processing. arXiv preprint arXiv:2205.15231.
  120. 120.Norman D. Megill and David A. Wheeler. 2019. Metamath: A Computer Language for Mathematical Proofs. Lulu Press, Morrisville, North Carolina. http://us.metamath.org/downloads/metamath.pdf.
  121. 121.Yuanliang Meng and Anna Rumshisky. 2019. Solving math word problems with double-decoder transformer. arXiv preprint arXiv:1908.10924.
  122. 122.Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2020. A diverse corpus for evaluating and developing english math word problem solvers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 975–984.
  123. 123.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? Proceedings of Empirical Methods in Natural Language Processing (EMNLP).
  124. 124.Shervin Minaee, Nal Kalchbrenner, Erik Cambria, Narjes Nikzad, Meysam Chenaghlu, and Jianfeng Gao. 2021. Deep learning based text classification: a comprehensive review. ACM Computing Surveys (CSUR), 54(3):1–40.
  125. 125.Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, and Ashwin Kalyan. 2022a. Lila: A unified benchmark for mathematical reasoning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  126. 126.Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Sachdeva, Peter Clark, Chitta Baral, and Ashwin Kalyan. 2022b. Numglue: A suite of fundamental yet challenging mathematical reasoning tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pages 3505–3523.
  127. 127.Eric Mitchell, Joseph J. Noh, Siyan Li, William S. Armstrong, Ananth Agarwal, Patrick Liu, Chelsea Finn, and Christopher D. Manning. 2022. Enhancing self-consistency and performance of pretrained language models with nli. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics.
  128. 128.Leonardo de Moura, Soonho Kong, Jeremy Avigad, Floris van Doorn, and Jakob von Raumer. 2015. The lean theorem prover (system description). In International Conference on Automated Deduction, pages 378–388. Springer.
  129. 129.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.
  130. 130.Allen Newell, John Clifford Shaw, and Herbert A Simon. 1957. Empirical explorations of the logic theory machine: A case study in heuristic. In Proceedings of the Western Joint Computer Conference, IRE-AIEE-ACM 1957.
  131. 131.Ansong Ni, Jeevana Priya Inala, Chenglong Wang, Oleksandr Polozov, Christopher Meek, Dragomir Radev, and Jianfeng Gao. 2023. Learning from self-sampled correct and partially-correct programs. In International Conference on Learning Representations (ICLR).
  132. 132.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. 2021. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114.
  133. 133.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (NeurIPS).
  134. 134.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are nlp models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HIT), pages 2080–2094.
  135. 135.Lawrence C. Paulson. 1994. Isabelle - A Generic Theorem Prover (with a contribution by T. Nipkow), volume 828 of Lecture Notes in Computer Science. Springer.
  136. 136.Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI).
  137. 137.Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, and Ilya Sutskever. 2023. Formal mathematics statement curriculum learning. In International Conference on Learning Representations (ICLR), volume abs/2202.01344.
  138. 138.Stanislas Polu and Ilya Sutskever. 2020. Generative language modeling for automated theorem proving. arXiv preprint arXiv:2009.03393.
  139. 139.Jinghui Qin, Xiaodan Liang, Yining Hong, Jianheng Tang, and Liang Lin. 2021. Neural-symbolic solver for math word problems with auxiliary tasks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL), pages 5870–5881.
  140. 140.Jinghui Qin, Lihui Lin, Xiaodan Liang, Rumin Zhang, and Liang Lin. 2020. Semantically-aligned universal tree-structured solver for math word problems. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3780–3789.
  141. 141.Liang Qiu, Yizhou Zhao, Jinchao Li, Pan Lu, Baolin Peng, Jianfeng Gao, and Song-Chun Zhu. 2022a. Valuenet: A new dataset for human value driven dialogue system. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), pages 2468–2484.
  142. 142.Liang Qiu, Yizhou Zhao, Yuan Liang, Pan Lu, Weiyan Shi, Zhou Yu, and Song-chun Zhu. 2022b. Towards socially intelligent agents with mental state transition and human value. In Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 146–158.
  143. 143.Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. 2020. Pre-trained models for natural language processing: A survey. Science China Technological Sciences, 63(10):1872–1897.
  144. 144.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2020. Language models are unsupervised multitask learners. OpenAI Blog.
  145. 145.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  146. 146.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR), 21:1–67.
  147. 147.Abhilasha Ravichander, Aakanksha Naik, Carolyn Rose, and Eduard Hovy. 2019. Equate: A benchmark evaluation framework for quantitative reasoning in natural language inference. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 349–361.
  148. 148.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022. Impact of pretraining term frequencies on few-shot numerical reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 840–854.
  149. 149.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems (NeurIPS), 28.
  150. 150.Ryokan Ri and Yoshimasa Tsuruoka. 2022. Pretraining with artificial language: Studying transferable knowledge in language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pages 7302–7315.
  151. 151.Benjamin Robaidek, Rik Koncel-Kedziorski, and Hannaneh Hajishirzi. 2018. Data-driven methods for solving algebra word problems. arXiv preprint arXiv:1804.10718.
  152. 152.Subhro Roy and Dan Roth. 2015. Solving general arithmetic word problems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1743–1752.
  153. 153.Subhro Roy and Dan Roth. 2017. Unit dependency graph and its application to arithmetic word problem solving. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI).
  154. 154.Subhro Roy and Dan Roth. 2018. Mapping to declarative knowledge for word problem solving. Transactions of the Association for Computational Linguistics (TACL), 6:159–172.
  155. 155.Subhro Roy, Tim Vieira, and Dan Roth. 2015. Reasoning about quantities in natural language. Transactions of the Association for Computational Linguistics (TACL), 3:1–13.
  156. 156.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. North American Chapter of the Association for Computational Linguistics (NAACL).
  157. 157.Mrinmaya Sachan, Kumar Dubey, and Eric Xing. 2017. From textbooks to knowledge: A case study in harvesting axiomatic knowledge from textbooks to solve geometry problems. In Proceedings of Empirical Methods in Natural Language Processing (EMNLP), pages 773–784.
  158. 158.Mrinmaya Sachan and Eric Xing. 2017. Learning to solve geometry problems from natural language demonstrations in textbooks. In Proceedings of the 6th Joint Conference on Lexical and Computational Semantics, pages 251–261.
  159. 159.David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. 2020. Analysing mathematical reasoning abilities of neural models. In International Conference on Learning Representations (ICLR).
  160. 160.Tal Schuster, Ashwin Kalyan, Alex Polozov, and Adam Tauman Kalai. 2021. Programming puzzles. In Thirty-fifth Conference on Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track.
  161. 161.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL), pages 1715–1725.
  162. 162.Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni, and Clint Malcolm. 2015. Solving geometry problems: Combining text and diagram interpretation. In Proceedings of Empirical Methods in Natural Language Processing (EMNLP), pages 1466–1476.
  163. 163.Jianhao Shen, Yichun Yin, Lin Li, Lifeng Shang, Xin Jiang, Ming Zhang, and Qun Liu. 2021. Generate & rank: A multi-task framework for math word problems. In Findings of the Association for Computational Linguistics (EMNLP), pages 2269–2279.
  164. 164.Yibin Shen and Cheqing Jin. 2020. Solving math word problems with multi-encoders and multi-decoders. In Proceedings of the 28th International Conference on Computational Linguistics (COLING), pages 2924–2934.
  165. 165.Shuming Shi, Yuehui Wang, Chin-Yew Lin, Xiaojiang Liu, and Yong Rui. 2015. Automatically solving number word problems by semantic parsing and reasoning. In Proceedings of the 2015 conference on empirical methods in natural language processing (EMNLP), pages 1132–1142.
  166. 166.Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019. Mass: Masked sequence to sequence pre-training for language generation. In 36th International Conference on Machine Learning (ICML).
  167. 167.Kai Sun, Dian Yu, Jianshu Chen, Dong Yu, Yejin Choi, and Claire Cardie. 2019. Dream: A challenge data set and models for dialogue-based reading comprehension. Transactions of the Association for Computational Linguistics (TACL), 7:217–231.
  168. 168.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. Advances in neural information processing systems (NeurIPS), 27.
  169. 169.Oyvind Tafjord, Peter Clark, Matt Gardner, Wen-tau Yih, and Ashish Sabharwal. 2019. Quarel: A dataset and models for answering questions about qualitative relationships. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), pages 7063–7071.
  170. 170.Kai Sheng Tai, Richard Socher, and Christopher D Manning. 2015. Improved semantic representations from tree-structured long short-term memory networks. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (ACL), pages 1556–1566.
  171. 171.Avijit Thawani, Jay Pujara, and Ashwin Kalyan. 2022. Estimating numbers without regression. In 36th Conference on Neural Information Processing Systems (NeurIPS 2022) Workshop on MATH-AI.
  172. 172.Avijit Thawani, Jay Pujara, Pedro A Szekely, and Filip Ilievski. 2021. Representing numbers in nlp: a survey and a vision. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HIT), pages 644–656.
  173. 173.Shounaak Ughade and Satish Kumbhar. 2019. Survey on mathematical word problem solving using natural language processing. In 2019 1st International Conference on Innovations in Information and Communication Technology (ICIICT), pages 1–5. IEEE.
  174. 174.Shyam Upadhyay and Ming-Wei Chang. 2015. Draw: A challenging and diverse algebra word problem set. Technical report, Citeseer.
  175. 175.Shyam Upadhyay and Ming-Wei Chang. 2017. Annotating derivations: A new evaluation strategy and dataset for algebra word problems. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (ACL), pages 494–504.
  176. 176.Josef Urban. 2006. Mptp 0.2: Design, implementation, and initial experiments. Journal of Automated Reasoning, 37(1):21–43.
  177. 177.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS), pages 5998–6008.
  178. 178.Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, and Matt Gardner. 2019. Do nlp models know numbers? probing numeracy in embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5307–5315.
  179. 179.Lei Wang, Yan Wang, Deng Cai, Dongxiang Zhang, and Xiaojiang Liu. 2018a. Translating a math word problem to a expression tree. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1064–1069.
  180. 180.Lei Wang, Dongxiang Zhang, Lianli Gao, Jingkuan Song, Long Guo, and Heng Tao Shen. 2018b. Mathdqn: Solving arithmetic word problems via deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI).
  181. 181.Lei Wang, Dongxiang Zhang, Jipeng Zhang, Xing Xu, Lianli Gao, Bing Tian Dai, and Heng Tao Shen. 2019. Template-based math word problem solvers with recursive neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), pages 7144–7151.
  182. 182.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations (ICLR).
  183. 183.Yan Wang, Xiaojiang Liu, and Shuming Shi. 2017. Deep neural solver for math word problems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 845–854.
  184. 184.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems (NeurIPS).
  185. 185.Sean Welleck, Jiacheng Liu, Ronan Le Bras, Hannaneh Hajishirzi, Yejin Choi, and Kyunghyun Cho. 2021. Naturalproofs: Mathematical theorem proving in natural language. In Thirty-fifth Conference on Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track.
  186. 186.Sean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi, and Yejin Choi. 2022a. Naturalprover: Grounded mathematical proof generation with language models. In Advances in Neural Information Processing Systems (NeurIPS).
  187. 187.Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel Khashabi, and Yejin Choi. 2023. Generating sequences by learning to self-correct. In International Conference on Learning Representations (ICLR).
  188. 188.Sean Welleck, Peter West, Jize Cao, and Yejin Choi. 2022b. Symbolic brittleness in sequence models: on systematic generalization in symbolic mathematics. In AAAI.
  189. 189.Wu Wen-Tsun. 1986. Basic principles of mechanical theorem proving in elementary geometries. Journal of automated Reasoning, 2(3):221–252.
  190. 190.Daniel Whalen. 2016. Holophrasm: a neural automated theorem prover for higher-order logic. arXiv preprint arXiv:1608.02644.
  191. 191.Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. 2018. Cbam: Convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV), pages 3–19.
  192. 192.Qinzhuo Wu, Qi Zhang, Jinlan Fu, and Xuan-Jing Huang. 2020. A knowledge-aware sequence-to-tree network for math word problem solving. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7137–7146.
  193. 193.Qinzhuo Wu, Qi Zhang, and Zhongyu Wei. 2021a. An edge-enhanced hierarchical graph-to-tree network for math word problem solving. In Findings of the Association for Computational Linguistics (EMNLP), pages 1473–1482.
  194. 194.Qinzhuo Wu, Qi Zhang, Zhongyu Wei, and Xuan-Jing Huang. 2021b. Math word problem solving with explicit numerical values. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL), pages 5859–5869.
  195. 195.Xingjiao Wu, Luwei Xiao, Yixuan Sun, Junhang Zhang, Tianlong Ma, and Liang He. 2022a. A survey of human-in-the-loop for machine learning. Future Generation Computer Systems.
  196. 196.Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144.
  197. 197.Yuhuai Wu, Albert Jiang, Jimmy Ba, and Roger Baker Grosse. 2021c. Int: An inequality benchmark for evaluating generalization in theorem proving. In International Conference on Learning Representations (ICLR).
  198. 198.Yuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus Norman Rabe, Charles E Staats, Mateja Jamnik, and Christian Szegedy. 2022b. Autoformalization with large language models. In Advances in Neural Information Processing Systems (NeurIPS).
  199. 199.Yuhuai Wu, Felix Li, and Percy Liang. 2022c. Insights into pre-training via simpler synthetic tasks. arXiv preprint arXiv:2206.10139.
  200. 200.Yuhuai Wu, Markus N Rabe, Wenda Li, Jimmy Ba, Roger B Grosse, and Christian Szegedy. 2021d. Lime: Learning inductive bias for primitives of mathematical reasoning. In International Conference on Machine Learning (ICML), pages 11251–11262. PMLR.
  201. 201.Zhipeng Xie and Shichao Sun. 2019. A goal-driven tree-structured neural model for math word problems. In International Joint Conference on Artificial Intelligence (IJCAI), pages 5299–5305.
  202. 202.Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In International conference on machine learning (ICML), pages 2048–2057. PMLR.
  203. 203.Kaiyu Yang and Jia Deng. 2019. Learning to prove theorems via interacting with proof assistants. In International Conference on Machine Learning (ICML), pages 6984–6994. PMLR.
  204. 204.Zheng Ye, Shang-Ching Chou, and Xiao-Shan Gao. 2008. An introduction to java geometry expert. In International workshop on automated deduction in geometry, pages 189–195. Springer.
  205. 205.Wei Yu, Mengzhu Wang, Xiaodong Wang, Xun Zhou, Yongfu Zha, Yongjian Zhang, Shuyu Miao, and Jingdong Liu. 2021a. Geore: A relation extraction dataset for chinese geometry problems. In 35th Conference on Neural Information Processing Systems (NeurIPS) Workshop on Math AI for Education (MATHAI4ED).
  206. 206.Weijiang Yu, Yingpeng Wen, Fudan Zheng, and Nong Xiao. 2021b. Improving math word problems with pre-trained knowledge and hierarchical reasoning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3384–3394.
  207. 207.Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2023. Generate rather than retrieve: Large language models are strong context generators. In International Conference on Learning Representations (ICLR).
  208. 208.Klim Zaporojets, Giannis Bekoulis, Johannes Deleu, Thomas Demeester, and Chris Develder. 2021. Solving arithmetic word problems by scoring equations with recursive neural networks. Expert Systems with Applications, 174:114704.
  209. 209.Dongxiang Zhang, Lei Wang, Luming Zhang, Bing Tian Dai, and Heng Tao Shen. 2019. The gap of semantic parsing: A survey on automatic math word problem solvers. IEEE transactions on pattern analysis and machine intelligence, 42(9):2287–2305.
  210. 210.Jipeng Zhang, Roy Ka-Wei Lee, Ee-Peng Lim, Wei Qin, Lei Wang, Jie Shao, and Qianru Sun. 2020a. Teacher-student networks with multiple decoders for solving math word problem. In International Joint Conference on Artificial Intelligence (IJCAI).
  211. 211.Jipeng Zhang, Lei Wang, Roy Ka-Wei Lee, Yi Bin, Yan Wang, Jie Shao, and Ee-Peng Lim. 2020b. Graph-to-tree learning for solving math word problems. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 3928–3937.
  212. 212.Ming-Liang Zhang, Fei Yin, Yi-Han Hao, and Cheng-Lin Liu. 2022. Learning to understand plane geometry diagram. In 36th Conference on Neural Information Processing Systems (NeurIPS) Workshop on MATH-AI.
  213. 213.Qiyuan Zhang, Lei Wang, Sicheng Yu, Shuohang Wang, Yang Wang, Jing Jiang, and Ee-Peng Lim. 2021. Noahqa: Numerical reasoning with interpretable graph question answering dataset. In Findings of the Association for Computational Linguistics (EMNLP), pages 4147–4161.
  214. 214.Wenhe Zhang, Chi Zhang, Yixin Zhu, and Song-Chun Zhu. 2020c. Machine number sense: A dataset of visual arithmetic problems for abstract and relational reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), pages 1332–1340.
  215. 215.Xikun Zhang, Deepak Ramachandran, Ian Tenney, Yanai Elazar, and Dan Roth. 2020d. Do language embeddings capture scales? In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 292–299.
  216. 216.Xikun Zhang, Deepak Ramachandran, Ian Tenney, Yanai Elazar, and Dan Roth. 2020e. Do language embeddings capture scales? In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 292–299.
  217. 217.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020f. Dialogpt: Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations.
  218. 218.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2023. Automatic chain of thought prompting in large language models. In International Conference on Learning Representations (ICLR).
  219. 219.Wei Zhao, Mingyue Shang, Yang Liu, Liang Wang, and Jingming Liu. 2020. Ape210k: A large-scale and template-rich dataset of math word problems. arXiv preprint arXiv:2009.11506.
  220. 220.Yilun Zhao, Yunxiang Li, Chenying Li, and Rui Zhang. 2022. Multihiertt: Numerical reasoning over multi hierarchical tabular and textual data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pages 6588–6600.
  221. 221.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning (ICML), pages 12697–12706. PMLR.
  222. 222.Kunhao Zheng, Jesse Michael Han, and Stanislas Polu. 2022. Minif2f: a cross-system benchmark for formal olympiad-level mathematics. In International Conference on Learning Representations (ICLR).
  223. 223.Ben Zhou, Daniel Khashabi, Qiang Ning, and Dan Roth. 2019. "Going on a vacation" takes longer than "Going for a walk": A Study of Temporal Commonsense Understanding. In Proc. of the Conference on Empirical Methods in Natural Language Processing (EMNLP).
  224. 224.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2023. Least-to-most prompting enables complex reasoning in large language models. In International Conference on Learning Representations (ICLR).
  225. 225.Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-JCNLP), pages 3277–3287.

Citation

MLA
Lu, P., et al. “A Survey of Deep Learning for Mathematical Reasoning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 14605–31, https://doi.org/10.18653/v1/2023.acl-long.817.
APA
Lu, P., Qiu, L., Yu, W., Welleck, S., & Chang, K.-W. (2023). A Survey of Deep Learning for Mathematical Reasoning. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14605–14631. https://doi.org/10.18653/v1/2023.acl-long.817
Chicago
Lu, P., L. Qiu, W. Yu, S. Welleck, and K.-W. Chang. 2023. “A Survey of Deep Learning for Mathematical Reasoning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14605–31. https://doi.org/10.18653/v1/2023.acl-long.817.
Harvard
Lu, P. et al. (2023) “A Survey of Deep Learning for Mathematical Reasoning”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 14605–14631. Available at: https://doi.org/10.18653/v1/2023.acl-long.817.
Vancouver
1. Lu P, Qiu L, Yu W, Welleck S, Chang K-W (2023) A Survey of Deep Learning for Mathematical Reasoning. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 14605–14631

BibTeX

@inproceedings{lu-etal-2023-survey,
    title = "A Survey of Deep Learning for Mathematical Reasoning",
    author = "Lu, Pan  and
      Qiu, Liang  and
      Yu, Wenhao  and
      Welleck, Sean  and
      Chang, Kai-Wei",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.817/",
    doi = "10.18653/v1/2023.acl-long.817",
    pages = "14605--14631"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/