Built independently by an author, for readers. Read the story and support ChapterPal

keyword

quantitative reasoning

Quantitative reasoning is the cognitive and computational ability to understand, analyze, and apply numerical and mathematical concepts to solve problems and draw logical conclusions. It encompasses the interpretation of information expressed in numerical, symbolic, or natural language formats, requiring more than basic arithmetic calculation to involve multi-step problem formulation, relational analysis, and logical deduction. Practitioners and automated systems utilizing quantitative reasoning evaluate evidence, identify relevant mathematical structures, model real-world scenarios, and verify the validity of derived results. This capability is fundamental across scientific, technical, and everyday decision-making contexts, representing a critical dimension of both human intelligence and the evaluation of advanced computational problem-solving systems.

6 items

MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning

MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning

Chengpeng Li, Zheng Yuan, Hongyi Yuan, Guanting Dong, Keming Lu, Jiancan Wu, Chuanqi Tan, Xiang Wang, Chang Zhou

OrganizationsAlibaba GroupInstitute of DataspaceUniversity of Science and Technology of China

Why you should read this

Presents a systematic investigation into query and response augmentation for mathematical reasoning, establishing state-of-the-art open-source models while identifying quantitative scaling laws and critical out-of-domain generalization limits between GSM8K and MATH.

In math reasoning with large language models (LLMs), fine-tuning data augmentation by query evolution and diverse reasoning paths is empirically verified effective, profoundly narrowing the gap between open-sourced LLMs and cutting-edge proprietary LLMs. In this paper, we conduct an investigation for such data augmentation in math reasoning and are intended to answer: (1) What strategies of data augmentation are more effective; (2) What is the scaling relationship between the amount of augmented data and model performance; and (3) Can data augmentation incentivize generalization to out-of-domain mathematical reasoning tasks? To this end, we create two new dataset AugGSM8K and AugMATH, by complicating and diversifying the queries and sampling multiple reasoning paths from GSM8K and MATH. We obtained a series of LLMs called MuggleMath by fine-tuning LLaMA models on AugGSM8K and AugMATH. MuggleMath substantially achieves new state-of-the-art on GSM8K and MATH. A log-linear relationship and a segmented log-linear are presented between MuggleMath’s performance and the amount of augmented data on GSM8K and MATH, respectively. We also find that it is weak in out-of-domain math reasoning generalization from AugGSM8K to MATH and from AugMATH to GSM8K, which suggests that augmenting queries that cover a broader range of subjects is more beneficial for generalization.

Added

2026-10-03

Reasoning Like Program Executors

Reasoning Like Program Executors

Xinyu Pi, Qian Liu, Bei Chen, Morteza Ziyadi, Zeqi Lin, Qiang Fu, Yan Gao, Jian-Guang Lou, Weizhu Chen

OrganizationsMicrosoftSea AI LabUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes a pre-training paradigm that teaches language models to predict program execution outputs, transferring formal symbolic reasoning capabilities directly into neural models for downstream natural language tasks.

Reasoning over natural language is a long-standing goal for the research community. However, studies have shown that existing language models are inadequate in reasoning. To address the issue, we present PoET, a novel reasoning pre-training paradigm. Through pre-training language models with programs and their execution results, PoET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach. PoET is conceptually simple and can be instantiated by different kinds of program executors. In this paper, we showcase two simple instances PoET-Math and PoET-Logic, in addition to a complex instance, PoET-SQL. Experimental results on six benchmarks demonstrate that PoET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning. PoET opens a new gate on reasoning-enhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors.

Added

2026-10-03

A Survey of Deep Learning for Mathematical Reasoning

A Survey of Deep Learning for Mathematical Reasoning

Pan Lu, Liang Qiu, Wenhao Yu, Sean Welleck, Kai-Wei Chang

OrganizationsUniversity of California, Los AngelesUniversity of Notre DameUniversity of Washington

Why you should read this

Provides a systematic taxonomy and evaluation of over 180 studies covering neural architectures, large language models, and benchmarks for automated mathematical problem solving and theorem proving.

Mathematical reasoning is a fundamental aspect of human intelligence and is applicable in various fields, including science, engineering, finance, and everyday life. The development of artificial intelligence (AI) systems capable of solving math problems and proving theorems in language has garnered significant interest in the fields of machine learning and natural language processing. For example, mathematics serves as a testbed for aspects of reasoning that are challenging for powerful deep learning models, driving new algorithmic and modeling advances. On the other hand, recent advances in large-scale neural language models have opened up new benchmarks and opportunities to use deep learning for mathematical reasoning. In this survey paper, we review the key tasks, datasets, and methods at the intersection of mathematical reasoning and deep learning over the past decade. We also evaluate existing benchmarks and methods, and discuss future research directions in this domain.

Added

2026-09-28

NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks

NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks

Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva, Peter Clark, Chitta Baral, Ashwin Kalyan

OrganizationsAllen Institute for AIArizona State UniversityMicrosoft

Why you should read this

Presents NUMGLUE, an eight-task benchmark spanning roughly 100,000 problems that exposes large language models' severe arithmetic brittleness compared to human reasoning while showing that joint multi-task training significantly boosts numerical performance across diverse question formats.

Given the ubiquitous nature of numbers in text, reasoning with numbers to perform simple calculations is an important skill of AI systems. While many datasets and models have been developed to this end, state-of-the-art AI systems are brittle; failing to perform the underlying mathematical reasoning when they appear in a slightly different scenario. Drawing inspiration from GLUE (Wang et al., 2018) that was proposed in the context of natural language understanding, we propose NUMGLUE, a multi-task benchmark that evaluates the performance of AI systems on eight different tasks, that at their core require simple arithmetic understanding. We show that this benchmark is far from being solved with neural models including state-of-the-art large-scale language models performing significantly worse than humans (lower by 46.4%). Further, NUMGLUE promotes sharing knowledge across tasks, especially those with limited training data as evidenced by the superior performance (average gain of 3.4% on each task) when a model is jointly trained on all the tasks as opposed to task-specific modeling. Finally, we hope that NUMGLUE will encourage systems that perform robust and general arithmetic reasoning within language, a first step towards being able to perform more complex mathematical reasoning.

Added

2026-09-26

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Zhihong Shao, Peiyi Wang, Qihao Zhu, R. Xu, Jun-Mei Song, Mingchuan Zhang, Y. K. Li, Yu Wu, Daya Guo

OrganizationsDeepSeek AIPeking UniversityTsinghua University

Why you should read this

Introduces DeepSeekMath 7B and Group Relative Policy Optimization (GRPO), demonstrating that an open 7B-parameter model can approach closed frontier performance on the competition-level MATH benchmark through large-scale pre-training and memory-efficient reinforcement learning.

Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark without relying on external toolkits and voting techniques, approaching the performance level of Gemini-Ultra and GPT-4. Self-consistency over 64 samples from DeepSeekMath 7B achieves 60.9% on MATH. The mathematical reasoning capability of DeepSeekMath is attributed to two key factors: First, we harness the significant potential of publicly available web data through a meticulously engineered data selection pipeline. Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.

Added

2026-09-14

WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, Yansong Tang, Dongmei Zhang

OrganizationsMicrosoftShenzhen Institute of Advanced Technology, Chinese Academy of SciencesTsinghua University

Why you should read this

Introduces WizardMath, a novel approach using Reinforcement Learning from Evol-Instruct Feedback (RLEIF) that significantly elevates the mathematical reasoning capabilities of open-source LLMs to surpass even proprietary models like GPT-3.5-Turbo and Claude 2 without external tools.

Large language models (LLMs), such as GPT-4, have shown remarkable performance in natural language processing (NLP) tasks, including challenging mathematical reasoning. However, most existing open-source models are only pre-trained on large-scale internet data and without math-related optimization. In this paper, we present WizardMath, which enhances the mathematical CoT reasoning abilities of LLMs without using external python tools, by applying our proposed Reinforcement Learning from Evol-Instruct Feedback (RLEIF) method to the domain of math. Through extensive experiments on two mathematical reasoning benchmarks, namely GSM8k and MATH, we reveal the extraordinary capabilities of our model. Remarkably, WizardMath-Mistral 7B surpasses top-tier open-source LLMs by a substantial margin with higher data efficiency. Furthermore, WizardMath 70B even outperforms GPT-3.5-Turbo, Claude 2, Gemini Pro and GPT-4-early-version. Additionally, our preliminary exploration highlights the pivotal role of instruction evolution and process supervision in achieving exceptional math performance. For more details refer to this https URL

Added

2026-05-14

Creative Commons License