Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Pan LuLiang QiuKai-Wei ChangYing Nian WuSong-Chun ZhuTanmay RajpurohitPeter ClarkAshwin Kalyan
Introduces TabMWP, a 38,431-question benchmark for mathematical reasoning over semi-structured tables, and develops PromptPG, a policy gradient technique that dynamically selects optimal in-context examples to improve large language model reasoning accuracy and stability.
Mathematical reasoning over structured and semi-structured data is an essential capability for artificial intelligence, yet existing benchmarks focus almost exclusively on pure text. Real-world documents—such as financial statements, medical records, and invoices—combine text with tabular structures, requiring systems to cross-reference table cells and perform arithmetic steps. While large language models like GPT-3 show promise on text-based math word problems, their performance in few-shot settings is notoriously sensitive to how demonstration examples are selected, often causing severe performance fluctuations when handling complex, heterogeneous tabular inputs.
The article addresses this limitation with two main objectives: introducing Tabular Math Word Problems (TABMWP), a large-scale open-domain benchmark combining tabular data and mathematical reasoning, and presenting PROMPTPG, a dynamic prompt-learning framework that uses reinforcement learning to select effective demonstration examples for large language models.
To establish the benchmark, the researchers collected 38,431 grade-level math problems across grades 1 through 8, each paired with a table presented in image, semi-structured, and structured formats. Each problem includes detailed multi-step natural language solutions. The team then formulated PROMPTPG, which deploys a lightweight policy network on top of a fixed language model to learn which candidate demonstration examples maximize answer accuracy when querying GPT-3. The system was evaluated across multiple baselines, including fine-tuned tabular and general question-answering models (TAPEX and UnifiedQA) as well as zero-shot and few-shot GPT-3 variants.
The article reports several key findings. First, PROMPTPG achieved an overall accuracy of 68.23%, outperforming the strongest baseline (few-shot chain-of-thought GPT-3 with random selection at 62.92%) by 5.31 percentage points. Second, PROMPTPG substantially reduced prediction variance compared to random demonstration selection, demonstrating consistent stability. Third, dynamic reinforcement learning-based selection outperformed heuristic strategies, including semantic similarity and complexity matching. Fourth, an input ablation study confirmed that both the table and the question text are strictly indispensable; removing either degraded model accuracy to near-zero levels. Finally, a substantial human evaluation benchmark of 90.22% accuracy revealed a 21.99 percentage point performance gap between human intelligence and the best model.
These findings indicate that learning-to-prompt frameworks provide a cost-effective, high-performing alternative to fine-tuning massive models or relying on unstable heuristic prompts. Dynamic example selection mitigates operational risks associated with unpredictable language model outputs in numerical tasks. However, error analyses show that models still struggle with complex tabular layouts (such as stem-and-leaf plots), intricate multi-step arithmetic, and rigid output formatting during automated parsing.
For practical application, organizations implementing language models for tabular and quantitative reasoning should adopt learned prompt-selection strategies rather than static or random demonstrations to maximize accuracy and minimize variance. Future work should focus on closing the 22% gap with human performance by improving logical reasoning over intricate tables, handling complex arithmetic sequences, and building more robust answer extraction pipelines.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). It introduces chain-of-thought prompting for multi-step reasoning in large language models, providing the essential few-shot prompting baseline that PROMPTPG aims to dynamically optimize.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). It establishes the foundational paradigm of in-context few-shot learning with GPT-3, which the source relies upon and seeks to stabilize through policy-gradient prompt selection.
- Paper: Making Pre-trained Language Models Better Few-shot Learners, Tianyu Gao et al. (2021). It demonstrates dynamic demonstration sampling and automated prompt design for few-shot learning, laying early groundwork for learning-to-prompt frameworks.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). It addresses prompt sensitivity and discrete prompt brittleness by introducing prompt tuning with continuous parameters, motivating learned prompt optimization.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). It formulates automated prompt generation and selection as an optimization problem, serving as key conceptual prior work for dynamic prompt learning.
- Paper: Are NLP Models really able to Solve Simple Math Word Problems?, Arkil Patel et al. (2021). It highlights how NLP models rely on superficial cues rather than genuine reasoning in math word problems, motivating the need for robust benchmarks like TABMWP.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). It provides a comprehensive taxonomy of prompt-based learning and demonstration ensembling strategies in NLP that contextualizes learned prompt selection.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). It demonstrates step-by-step reasoning prompts in zero-shot settings, contextualizing the few-shot versus zero-shot baselines evaluated in the source.
- Paper: MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts, Pan Lu et al. (2023). It extends mathematical reasoning evaluation beyond semi-structured text and tables into diverse visual contexts and multimodal foundation models.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). It advances reasoning over tables and structured knowledge by introducing an iterative reading-then-reasoning interface for large language models.
- Paper: Rethinking Tabular Data Understanding with Large Language Models, Tianyang Liu et al. (2024). It critically examines language model robustness to structural variations in tables, directly continuing the source's investigation into tabular reasoning capabilities.
- Paper: DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, Zhihong Shao et al. (2024). It advances reinforcement learning for mathematical reasoning by introducing group relative policy optimization to train language models directly.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). It generalizes structured prompt-based reasoning beyond linear chains into arbitrary graph architectures for complex problem-solving.
- Paper: Large Language Models Can Be Easily Distracted by Irrelevant Context, Freda Shi et al. (2023). It evaluates the distractibility of prompting methods when math problems contain irrelevant context, testing the robustness of prompt-driven arithmetic reasoning.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). It benchmarks integrated multimodal capabilities, combining tabular interpretation, OCR, and arithmetic calculation in complex real-world tasks.
