Built independently by an author, for readers. Read the story and support ChapterPal

keyword

numerical commonsense

Numerical commonsense refers to the intuitive, everyday understanding of numbers, quantities, and basic numerical relationships that humans routinely use to interpret the physical world. In artificial intelligence and natural language processing, it denotes the capability of a system to recognize and reason over implicit numerical properties and quantitative facts associated with real-world entities, such as the standard count of limbs on an animal, typical ranges of physical dimensions, or reasonable durations of everyday events. This form of commonsense bridges quantitative logic and language comprehension, allowing models to evaluate whether numerical statements are plausible, estimate approximate magnitudes, and perform basic arithmetic reasoning within everyday contexts without requiring explicit formal calculation.

1 item

NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks

NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks

Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva, Peter Clark, Chitta Baral, Ashwin Kalyan

OrganizationsAllen Institute for AIArizona State UniversityMicrosoft

Why you should read this

Presents NUMGLUE, an eight-task benchmark spanning roughly 100,000 problems that exposes large language models' severe arithmetic brittleness compared to human reasoning while showing that joint multi-task training significantly boosts numerical performance across diverse question formats.

Given the ubiquitous nature of numbers in text, reasoning with numbers to perform simple calculations is an important skill of AI systems. While many datasets and models have been developed to this end, state-of-the-art AI systems are brittle; failing to perform the underlying mathematical reasoning when they appear in a slightly different scenario. Drawing inspiration from GLUE (Wang et al., 2018) that was proposed in the context of natural language understanding, we propose NUMGLUE, a multi-task benchmark that evaluates the performance of AI systems on eight different tasks, that at their core require simple arithmetic understanding. We show that this benchmark is far from being solved with neural models including state-of-the-art large-scale language models performing significantly worse than humans (lower by 46.4%). Further, NUMGLUE promotes sharing knowledge across tasks, especially those with limited training data as evidenced by the superior performance (average gain of 3.4% on each task) when a model is jointly trained on all the tasks as opposed to task-specific modeling. Finally, we hope that NUMGLUE will encourage systems that perform robust and general arithmetic reasoning within language, a first step towards being able to perform more complex mathematical reasoning.

Added

2026-09-26