Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Pan LuLiang QiuKai-Wei ChangYing Nian WuSong-Chun ZhuTanmay RajpurohitPeter ClarkAshwin Kalyan

article2023ICLR468 citations

Introduces TabMWP, a 38,431-question benchmark for mathematical reasoning over semi-structured tables, and develops PromptPG, a policy gradient technique that dynamically selects optimal in-context examples to improve large language model reasoning accuracy and stability.

Listen

Mathematical reasoning over structured and semi-structured data is an essential capability for artificial intelligence, yet existing benchmarks focus almost exclusively on pure text. Real-world documents—such as financial statements, medical records, and invoices—combine text with tabular structures, requiring systems to cross-reference table cells and perform arithmetic steps. While large language models like GPT-3 show promise on text-based math word problems, their performance in few-shot settings is notoriously sensitive to how demonstration examples are selected, often causing severe performance fluctuations when handling complex, heterogeneous tabular inputs.

The article addresses this limitation with two main objectives: introducing Tabular Math Word Problems (TABMWP), a large-scale open-domain benchmark combining tabular data and mathematical reasoning, and presenting PROMPTPG, a dynamic prompt-learning framework that uses reinforcement learning to select effective demonstration examples for large language models.

To establish the benchmark, the researchers collected 38,431 grade-level math problems across grades 1 through 8, each paired with a table presented in image, semi-structured, and structured formats. Each problem includes detailed multi-step natural language solutions. The team then formulated PROMPTPG, which deploys a lightweight policy network on top of a fixed language model to learn which candidate demonstration examples maximize answer accuracy when querying GPT-3. The system was evaluated across multiple baselines, including fine-tuned tabular and general question-answering models (TAPEX and UnifiedQA) as well as zero-shot and few-shot GPT-3 variants.

The article reports several key findings. First, PROMPTPG achieved an overall accuracy of 68.23%, outperforming the strongest baseline (few-shot chain-of-thought GPT-3 with random selection at 62.92%) by 5.31 percentage points. Second, PROMPTPG substantially reduced prediction variance compared to random demonstration selection, demonstrating consistent stability. Third, dynamic reinforcement learning-based selection outperformed heuristic strategies, including semantic similarity and complexity matching. Fourth, an input ablation study confirmed that both the table and the question text are strictly indispensable; removing either degraded model accuracy to near-zero levels. Finally, a substantial human evaluation benchmark of 90.22% accuracy revealed a 21.99 percentage point performance gap between human intelligence and the best model.

These findings indicate that learning-to-prompt frameworks provide a cost-effective, high-performing alternative to fine-tuning massive models or relying on unstable heuristic prompts. Dynamic example selection mitigates operational risks associated with unpredictable language model outputs in numerical tasks. However, error analyses show that models still struggle with complex tabular layouts (such as stem-and-leaf plots), intricate multi-step arithmetic, and rigid output formatting during automated parsing.

For practical application, organizations implementing language models for tabular and quantitative reasoning should adopt learned prompt-selection strategies rather than static or random demonstrations to maximize accuracy and minimize variance. Future work should focus on closing the 22% gap with human performance by improving logical reasoning over intricate tables, handling complex arithmetic sequences, and building more robust answer extraction pipelines.

Cover for Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Abstract

Mathematical reasoning, a core ability of human intelligence, presents unique challenges for machines in abstract thinking and logical reasoning. Recent large pre-trained language models such as GPT-3 have achieved remarkable progress on mathematical reasoning tasks written in text form, such as math word problems (MWP). However, it is unknown if the models can handle more complex problems that involve math reasoning over heterogeneous information, such as tabular data. To fill the gap, we present Tabular Math Word Problems (TabMWP), a new dataset containing 38,431 open-domain grade-level problems that require mathematical reasoning on both textual and tabular data. Each question in TabMWP is aligned with a tabular context, which is presented as an image, semi-structured text, and a structured table. There are two types of questions: free-text and multi-choice, and each problem is annotated with gold solutions to reveal the multi-step reasoning process. We evaluate different pre-trained models on TabMWP, including the GPT-3 model in a few-shot setting. As earlier studies suggest, since few-shot GPT-3 relies on the selection of in-context examples, its performance is unstable and can degrade to near chance. The unstable issue is more severe when handling complex problems like TabMWP. To mitigate this, we further propose a novel approach, PromptPG, which utilizes policy gradient to learn to select in-context examples from a small amount of training data and then constructs the corresponding prompt for the test example. Experimental results show that our method outperforms the best baseline by 5.31% on the accuracy metric and reduces the prediction variance significantly compared to random selection, which verifies its effectiveness in selecting in-context examples.

Table of Contents

  • 1 Introduction
  • 2 The TabMWP Dataset
  • 2.1 Task Formulation
  • 2.2 Dataset Construction
  • 2.3 Dataset Statistics
  • 3 Methods
  • 3.1 Few-shot GPT-3 for TabMWP
  • 3.2 Dynamic Prompting via Policy Gradient
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Experimental Results
  • 4.3 Ablation Study
  • 4.4 Case Study
  • 5 Related Work
  • 5.1 Math Word Problems
  • 5.2 Table QA Datasets
  • 5.3 Prompt Learning for Language Models
  • 6 Conclusion
  • 7 Acknowledgment
  • References
  • A Appendix
  • A.1 Dataset collection
  • A.2 Human study
  • A.3 The PromptPG Algorithm
  • A.4 Implementation Details
  • A.5 More Experimental Results
  • A.6 Related work of Policy Gradient
  • A.7 Case study examples

Knowls

  1. Knowl 1 — PROMPTPG: Dynamic Prompt Learning via Policy Gradient for In-Context Example Selection

    model/method

    PROMPTPG is a reinforcement learning framework that dynamically selects optimal in-context demonstration examples for few-shot prompting of large language models (such as GPT-3) without relying on brute-force search or manually handcrafted heuristics.

    Given a target math problem pip_i, an agent selects KK in-context demonstration examples ei={ei1,ei2,…,eiK}e_i = \{e_i^1, e_i^2, \dots, e_i^K\} from a candidate pool EcandE_{\text{cand}} according to a parameterized policy πθ(ei∣pi)\pi_\theta(e_i \mid p_i), where each example is sampled independently:

    eik∼πθ(ei∣pi),eik∈Ecand,k∈{1,2,…,K}e_i^k \sim \pi_\theta(e_i \mid p_i), \quad e_i^k \in E_{\text{cand}}, \quad k \in \{1, 2, \dots, K\}

    The selected demonstration examples and the target problem pip_i are formatted as an input prompt and passed to the frozen language model to produce a prediction a^i=GPT-3(ei,pi)\hat{a}_i = \text{GPT-3}(e_i, p_i) containing a generated chain-of-thought solution followed by the final answer. The environment computes a scalar reward rir_i by evaluating the generated answer a^i\hat{a}_i against the ground truth aia_i:

    ri=R(a^i∣pi)=EVAL(a^i,ai)={+1if a^i matches ai,−1otherwise.r_i = R(\hat{a}_i \mid p_i) = \text{EVAL}(\hat{a}_i, a_i) = \begin{cases} +1 & \text{if } \hat{a}_i \text{ matches } a_i, \\ -1 & \text{otherwise.} \end{cases}

    The objective is to maximize the expected reward Eei∼πθ(ei∣pi)[R(GPT-3(ei,pi))]\mathbb{E}_{e_i \sim \pi_\theta(e_i \mid p_i)}[R(\text{GPT-3}(e_i, p_i))]. The policy parameters θ\theta are updated using the REINFORCE policy gradient algorithm, increasing the selection probabilities of candidate sets that lead to correct predictions while penalizing those that produce incorrect predictions.

  2. Knowl 2 — Candidate Selection Policy and Problem Representation in PROMPTPG

    equation

    In PROMPTPG, candidate examples ei∈Ecande_i \in E_{\text{cand}} and the target problem pip_i are encoded using a pre-trained bidirectional transformer (BERT) whose parameters are kept frozen, followed by a single learnable linear layer with weight matrix WW and bias vector bb (learnable parameter set θ={W,b}\theta = \{W, b\} with hidden embedding dimension 768):

    h(ei)=W(BERT(ei))+bh(e_i) = W(\text{BERT}(e_i)) + b

    h(pi)=W(BERT(pi))+bh(p_i) = W(\text{BERT}(p_i)) + b

    where BERT(⋅)\text{BERT}(\cdot) denotes the contextualized embedding from the final pooling layer corresponding to the [CLS]\text{[CLS]} token.

    The conditional probability of selecting candidate example eie_i given problem pip_i is computed via a softmax distribution over inner products:

    πθ(ei∣pi)=exp⁡(h(ei)⋅h(pi))∑ei′∈Ecandexp⁡(h(ei′)⋅h(pi))\pi_\theta(e_i \mid p_i) = \frac{\exp(h(e_i) \cdot h(p_i))}{\sum_{e'_i \in E_{\text{cand}}} \exp(h(e'_i) \cdot h(p_i))}

    During prompting, KK examples are sampled independently from πθ(ei∣pi)\pi_\theta(e_i \mid p_i) to construct the prompt.

  3. Knowl 3 — PROMPTPG Training Algorithm via REINFORCE Policy Gradient

    algorithm

    PROMPTPG optimizes the prompt selection policy network using Monte Carlo sampling and the REINFORCE policy gradient algorithm. The policy network is parameterized by θ={W,b}\theta = \{W, b\} on top of a frozen BERT encoder. The training process uses the Adam optimizer with an initial learning rate of 1×10−31 \times 10^{-3}, a batch size of 20, and a maximum of 30 epochs, with early stopping if the batch loss evaluates to NaN.

    Input: Initial policy parameters θ0\theta_0, training problem set PtrainP_{\text{train}}, candidate example pool EcandE_{\text{cand}}, number of training epochs NN, number of in-context shots KK
    Output: Learned policy parameters θ\theta
    function REINFORCE(πθ0\pi_{\theta_0}, PtrainP_{\text{train}}, EcandE_{\text{cand}}, NN)
        Initialize policy network parameters θ←θ0\theta \leftarrow \theta_0
        for epoch =1,2,…,N= 1, 2, \dots, N do
            for each batch Pbatch⊆PtrainP_{\text{batch}} \subseteq P_{\text{train}} do
                Lbatch←0\mathcal{L}_{\text{batch}} \leftarrow 0
                for each problem pi∈Pbatchp_i \in P_{\text{batch}} do
                    for k=1,…,Kk = 1, \dots, K do
                        Sample in-context example eik∼πθ(ei∣pi)e_i^k \sim \pi_\theta(e_i \mid p_i) from EcandE_{\text{cand}}
                    Construct prompt from selected examples (ei1,…,eiK)(e_i^1, \dots, e_i^K) and problem pip_i
                    Generate answer a^i←GPT-3(ei1,…,eiK,pi)\hat{a}_i \leftarrow \text{GPT-3}(e_i^1, \dots, e_i^K, p_i)
                    Compute reward ri←EVAL(a^i,ai)∈{−1,+1}r_i \leftarrow \text{EVAL}(\hat{a}_i, a_i) \in \{-1, +1\}
                    Lbatch←Lbatch−ri⋅ln⁡πθ(ei∣pi)\mathcal{L}_{\text{batch}} \leftarrow \mathcal{L}_{\text{batch}} - r_i \cdot \ln \pi_\theta(e_i \mid p_i)
                Update θ\theta to minimize Lbatch\mathcal{L}_{\text{batch}} using the policy gradient estimator
                if Lbatch\mathcal{L}_{\text{batch}} contains NaN then
                    Stop training early
        return πθ\pi_\theta
  4. Knowl 4 — The Tabular Math Word Problems (TABMWP) Dataset

    definition

    Tabular Math Word Problems (TABMWP) is an open-domain benchmark consisting of 38,431 grade-level mathematical word problems requiring multi-hop numerical and logical reasoning across heterogeneous textual and tabular data.

    Each problem p=(t,q)p = (t, q) consists of a tabular context tt, a textual question qq, optional multiple-choice options c={c1,…,cn}c = \{c_1, \dots, c_n\} or a target unit uu, a ground truth answer aa, and an annotated multi-step natural language explanation ss. The tabular context tt is provided in three interchangeable formats:

    1. An image screenshot of the table.
    2. A semi-structured text format flattening rows separated by newlines \n and columns separated by vertical bars |.
    3. A structured spreadsheet format executable by database query packages.

    The dataset is partitioned into 23,059 training instances (60%), 7,686 development instances (20%), and 7,686 test instances (20%). It features 37,644 unique tables (average of 5.9 rows, 2.2 columns, and 12.9 cells; 60.5% contain a title). Question formats comprise:

    • Free-text questions (74.7% of total): The answer is purely numerical, consisting of an integer (59.50% of total) or a decimal/fraction (15.23% of total).
    • Multi-choice questions (25.3% of total): The answer is a text span selected from candidate options, categorized as extractive (13.01% of total), Boolean (10.97% of total), or other text (1.29% of total).
  5. Knowl 5 — TABMWP Benchmark Performance Across Baselines and PROMPTPG

    data/table

    The table below evaluates multiple QA and TableQA models alongside zero-shot and few-shot GPT-3 (text-davinci-002) variants on the test split of the TABMWP dataset. Evaluations report accuracy across question types (FREE: free-text, MC: multi-choice), answer types (INT: integer, DEC: decimal, EXTR: extractive text, BOOL: Boolean, OTH: other text), grade levels (grades 1–6 vs. 7–8), and the overall average.

    Method Train Data Strategy FREE MC INT DEC EXTR BOOL OTH 1–6 7–8 Avg.
    Heuristic guess – – 6.71 39.81 8.37 0.26 30.80 51.22 26.67 17.55 12.27 15.29
    Human performance – – 84.61 93.32 84.95 83.29 97.18 88.69 96.20 94.27 81.28 90.22
    Pre-trained Baselines
    UnifiedQASMALL_{\text{SMALL}} – – 1.18 43.62 1.37 0.43 38.70 49.78 37.14 15.57 7.65 12.18
    UnifiedQABASE_{\text{BASE}} – – 4.60 43.02 5.28 1.97 37.08 50.11 38.10 17.14 11.11 14.56
    UnifiedQALARGE_{\text{LARGE}} – – 4.48 48.80 5.19 1.72 48.33 50.33 40.00 19.78 10.87 15.96
    TAPEXBASE_{\text{BASE}} – – 7.32 39.76 8.68 2.06 35.06 47.11 20.95 18.67 11.81 15.73
    TAPEXLARGE_{\text{LARGE}} – – 8.80 46.59 10.62 1.72 46.91 48.11 30.48 22.65 13.18 18.59
    Fine-tuned Baselines
    UnifiedQASMALL_{\text{SMALL}} 23,059 – 22.27 51.31 27.27 2.83 52.28 48.11 69.52 35.85 21.71 29.79
    UnifiedQABASE_{\text{BASE}} 23,059 – 34.02 70.68 40.74 7.90 84.09 55.67 73.33 53.31 30.46 43.52
    UnifiedQALARGE_{\text{LARGE}} 23,059 – 48.67 82.18 55.97 20.26 94.63 68.89 79.05 65.92 45.92 57.35
    TAPEXBASE_{\text{BASE}} 23,059 – 39.59 73.09 46.85 11.33 84.19 61.33 69.52 56.70 37.02 48.27
    TAPEXLARGE_{\text{LARGE}} 23,059 – 51.00 80.02 59.92 16.31 95.34 64.00 73.33 67.11 47.07 58.52
    Prompting Baselines w/ GPT-3
    Zero-shot – – 53.57 66.67 55.55 45.84 78.22 55.44 54.29 63.37 48.41 56.96
    Zero-shot-CoT – – 54.36 66.92 55.82 48.67 78.82 55.67 51.43 63.62 49.59 57.61
    Few-shot (2-shot) 2 Random 54.69 64.11 58.36 40.40 75.95 52.41 53.02 63.10 49.16 57.13
    Few-shot-CoT (2-shot) 2 Random 60.76 69.09 60.04 63.58 76.49 61.19 67.30 68.62 55.31 62.92
    PROMPTPG w/ GPT-3 (Ours)
    Few-shot-CoT (2-shot) 160+20 Dynamic 66.17 74.11 64.12 74.16 76.19 72.81 65.71 71.20 64.27 68.23

    PROMPTPG establishes state-of-the-art accuracy of 68.23% on TABMWP, outperforming 2-shot random Few-shot-CoT GPT-3 by 5.31% and the best fine-tuned specialized tabular model (TAPEXLARGE_{\text{LARGE}}) by 9.71%.

  6. Knowl 6 — Comparison of Prompt Selection Strategies for Few-Shot In-Context Learning

    data/table

    The performance of 2-shot Chain-of-Thought GPT-3 on 1,000 TABMWP development examples was evaluated under different demonstration selection strategies over three repeated trials:

    Selection Strategy Accuracy (%)
    Same question type 66.2±0.6066.2 \pm 0.60
    Same answer type 67.9±0.3867.9 \pm 0.38
    Same grade level 67.9±1.8767.9 \pm 1.87
    Most complex (# of table cells) 64.0±0.4264.0 \pm 0.42
    Most complex (# of question words) 68.2±0.2668.2 \pm 0.26
    Random selection 65.2±4.0165.2 \pm 4.01
    Manual selection (fixed w/ top 2) 66.9±0.0066.9 \pm 0.00
    Nearest neighbor (semantic retrieval) 68.2±0.2968.2 \pm 0.29
    PROMPTPG (Ours) 70.9±1.27\mathbf{70.9 \pm 1.27}

    Random selection suffers from high variance (±4.01%\pm 4.01\%). While semantic retrieval via nearest neighbor reduces variance and improves accuracy by selecting superficially similar text, PROMPTPG achieves the highest overall accuracy (70.9%70.9\%) by learning to select examples sharing underlying multi-step mathematical reasoning structures.

  7. Knowl 7 — Blind Studies on Input Modalities in TABMWP

    data/table

    To measure whether problems in TABMWP strictly require joint textual and tabular reasoning, zero-shot GPT-3 was evaluated on 1,000 development examples under varying input ablations, where TT denotes tabular context, QQ denotes question text, CC denotes candidate choices, and AA denotes answer prediction:

    Model Format FREE MC INT DEC EXTR BOOL OTH 1–6 7–8 Avg.
    Heuristic guess TQ(C)→\rightarrowA 7.31 40.36 9.20 0.00 34.44 47.32 50.00 17.99 13.96 16.40
    Zero-shot GPT-3 T→\rightarrowA 8.28 0.36 10.24 0.67 0.66 0.00 0.00 9.41 1.02 6.10
    Zero-shot GPT-3 Q→\rightarrowA 9.24 1.09 10.94 2.68 1.32 0.89 0.00 10.23 2.03 7.00
    Zero-shot GPT-3 T(C)→\rightarrowA 8.28 41.82 10.24 0.67 36.42 50.89 25.00 23.60 8.12 17.50
    Zero-shot GPT-3 Q(C)→\rightarrowA 9.10 33.09 10.94 2.01 25.17 44.64 25.00 21.29 7.11 15.70
    Zero-shot GPT-3 TQ→\rightarrowA 55.31 68.36 56.60 50.34 79.47 54.46 58.33 66.34 47.46 58.90
    Zero-shot GPT-3 TQ(C)→\rightarrowA 54.76 72.00 56.42 48.32 76.82 66.07 66.67 67.00 47.97 59.50

    Removing either the tabular context (Q→AQ \rightarrow A, 7.00%) or the question text (T→AT \rightarrow A, 6.10%) collapses performance to near-chance levels on multi-choice questions and single digits on free-text questions, demonstrating that both the tabular context and question text are essential.

  8. Knowl 8 — Impact of Training Size, Candidate Pool Size, and Shot Count on PROMPTPG

    empirical result

    Systematic sweeps on 1,000 TABMWP development instances demonstrate the parametric dynamics of PROMPTPG:

    1. Number of Training Examples: With a candidate pool of 20, varying training instances from 20 to 320 shows that accuracy peaks at approximately 160 training examples. Beyond 160 examples, prediction accuracy decreases and variance increases due to inefficient exploitation in larger training spaces.
    2. Candidate Pool Size: When evaluated with 80 or 160 training examples, increasing the candidate pool size from 2 up to 320 causes accuracy to peak at 20 candidates before declining. Small pools restrict the action space for exploring problem types, whereas excessively large candidate sets expand the search space beyond what policy gradients can optimize efficiently with small sample budgets.
    3. Number of In-Context Shots: When evaluating random few-shot-CoT GPT-3 across shot counts, accuracy scales from 65.2±4.01%65.2 \pm 4.01\% (2-shot) to 65.7±1.16%65.7 \pm 1.16\% (3-shot), 67.7±0.78%67.7 \pm 0.78\% (4-shot), and plateaus at 67.5±0.98%67.5 \pm 0.98\% (5-shot). PROMPTPG with only 2 shots attains 70.9±1.27%70.9 \pm 1.27\%, outperforming random selection even when random selection uses 5 demonstration shots.
  9. Knowl 9 — Failure Modes and Reasoning Bottlenecks in PROMPTPG

    limitation

    Qualitative and quantitative error analysis reveals key failure modes of PROMPTPG on semi-structured math reasoning:

    1. Abstract and Domain-Specific Table Layouts: The model struggles to understand specialized table encodings, such as Stem-and-Leaf plots, frequently confusing stem/leaf coordinate indices with raw numbers.
    2. Long Arithmetic Chains: Performance degrades when problems require multi-step arithmetic over many tabular rows (e.g., computing the mean across 8 or more entries), resulting in arithmetic calculation slips.
    3. Indirect Temporal and Schedule Matching: When a question poses an implicit condition that does not match an exact tabular entry (such as finding the next available departure after a query time not listed in the table schedule), the model often fails to identify the correct interval.
    4. Output Formatting Violations: In certain free-text questions, the language model outputs correct intermediate reasoning and the correct numerical answer within the natural language text, but omits the expected final output syntax (e.g., "The answer is X"), causing regex-based evaluators to fail to extract the prediction.

Coverage note — None was omitted.

References

  1. 1.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL), pp. 2357–2367, 2019.
  2. 2.Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva Tb, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap. Distributed distributional deterministic policy gradients. arXiv preprint arXiv:1804.08617, 2018.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS), 33:1877–1901, 2020.
  4. 4.Jiaqi Chen, Tong Li, Jinghui Qin, Pan Lu, Liang Lin, Chongyu Chen, and Xiaodan Liang. Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression. In The 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022.
  5. 5.Wenhu Chen, Ming-Wei Chang, Eva Schlinger, William Yang Wang, and William W Cohen. Open question answering over tables and text. In International Conference on Learning Representations (ICLR), 2020a.
  6. 6.Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. Hybridqa: A dataset of multi-hop question answering over tabular and textual data. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 1026–1036, 2020b.
  7. 7.Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan R Routledge, et al. Finqa: A dataset of numerical reasoning over financial data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 3697–3711, 2021.
  8. 8.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  10. 10.Amelia Glaese, Nat McAleese, Maja Trebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375, 2022.
  11. 11.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning (ICML), pp. 2790–2799. PMLR, 2019.
  12. 12.Danqing Huang, Shuming Shi, Chin-Yew Lin, Jian Yin, and Wei-Ying Ma. How well do computers solve math word problems? large-scale dataset construction and evaluation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 887–896, 2016.
  13. 13.Danqing Huang, Shuming Shi, Chin-Yew Lin, and Jian Yin. Learning fine-grained expressions to solve math word problems. In Proceedings of Empirical Methods in Natural Language Processing (EMNLP), pp. 805–814, 2017.
  14. 14.Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. Search-based neural structured learning for sequential question answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 1821–1831, 2017.
  15. 15.Sujay Kumar Jauhar, Peter Turney, and Eduard Hovy. Tabmcq: A dataset of general knowledge tables and multiple-choice questions. arXiv preprint arXiv:1602.03960, 2016.
  16. 16.Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan. Dvqa: Understanding data visualizations via question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5648–5656, 2018.
  17. 17.Kirthevasan Kandasamy, Yoram Bachrach, Ryota Tomioka, Daniel Tarlow, and David Carter. Batch policy gradient methods for improving neural conversation models. arXiv preprint arXiv:1702.03334, 2017.
  18. 18.Yannis Katsis, Saneem Chemmengath, Vishwajeet Kumar, Samarth Bharadwaj, Mustafa Canim, Michael Glass, Alfio Gliozzo, Feifei Pan, Jaydeep Sen, Karthik Sankaranarayanan, et al. Ait-qa: Question answering dataset over complex tables in the airline industry. arXiv preprint arXiv:2106.12944, 2021.
  19. 19.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. Unifiedqa: Crossing format boundaries with a single qa system. In Findings of the Association for Computational Linguistics (EMNLP), pp. 1896–1907, 2020.
  20. 20.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  21. 21.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916, 2022.
  22. 22.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7871–7880, Online, July 2020. Association for Computational Linguistics (ACL).
  23. 23.Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015.
  24. 24.Jiachang Liu, Dinghan Shen, Yizhe Zhang, William B Dolan, Lawrence Carin, and Weizhu Chen. What makes good in-context examples for gpt-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pp. 100–114, 2022a.
  25. 25.Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. Tapex: Table pre-training via learning a neural sql executor. In International Conference on Learning Representations (ICLR), 2022b.
  26. 26.Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-Chun Zhu. Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning. In The 59th Annual Meeting of the Association for Computational Linguistics (ACL), 2021a.
  27. 27.Pan Lu, Liang Qiu, Jiaqi Chen, Tony Xia, Yizhou Zhao, Wei Zhang, Zhou Yu, Xiaodan Liang, and Song-Chun Zhu. Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning. In The 35th Conference on Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks, 2021b.
  28. 28.Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. In The 36th Conference on Neural Information Processing Systems (NeurIPS 2022), 2022a.
  29. 29.Pan Lu, Liang Qiu, Wenhao Yu, Sean Welleck, and Kai-Wei Chang. A survey of deep learning for mathematical reasoning. arXiv preprint arXiv:2212.10535, 2022b.
  30. 30.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 8086–8098, 2022c.
  31. 31.Xiaojian Ma, Silong Yong, Zilong Zheng, Qing Li, Yitao Liang, Song-Chun Zhu, and Siyuan Huang. Sqa3d: Situated question answering in 3d scenes. arXiv preprint arXiv:2210.07474, 2022.
  32. 32.Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. A diverse corpus for evaluating and developing english math word problem solvers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 975–984, 2020.
  33. 33.Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, and Ashwin Kalyan. Lila: A unified benchmark for mathematical reasoning. In The 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022.
  34. 34.Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In International conference on machine learning, pp. 1928–1937. PMLR, 2016.
  35. 35.Linyong Nan, Chiachun Hsieh, Ziming Mao, Xi Lin, Neha Verma, Rui Zhang, Wojciech Kryściński, Nick Schoelkopf, Riley Kong, Xiangru Tang, et al. Fetaqa: Free-form table question answering. Transactions of the Association for Computational Linguistics (TACL), 10:35–49, 2022.
  36. 36.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.
  37. 37.Panupong Pasupat and Percy Liang. Compositional semantic parsing on semi-structured tables. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (ACL-IJNLP), pp. 1470–1480, 2015.
  38. 38.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. Are nlp models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL), pp. 2080–2094, 2021.
  39. 39.Jan Peters and Stefan Schaal. Policy gradient methods for robotics. In 2006 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 2219–2225. IEEE, 2006.
  40. 40.Liang Qiu, Yizhou Zhao, Jinchao Li, Pan Lu, Baolin Peng, Jianfeng Gao, and Song-Chun Zhu. Valuenet: A new dataset for human value driven dialogue system. IEEE Trans. on Pattern Analysis and Machine Intelligence (TPAMI), 44(5):2468–2484, 2022.
  41. 41.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR), 21:1–67, 2020.
  42. 42.Subhro Roy and Dan Roth. Unit dependency graph and its application to arithmetic word problem solving. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2017.
  43. 43.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  44. 44.David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. Deterministic policy gradient algorithms. In International conference on machine learning, pp. 387–395. PMLR, 2014.
  45. 45.Richard S Sutton, Andrew G Barto, et al. Introduction to reinforcement learning. MIT press Cambridge, 1998.
  46. 46.Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant. Multimodalqa: complex question answering over text, tables and images. In International Conference on Learning Representations (ICLR), 2020.
  47. 47.Shyam Upadhyay and Ming-Wei Chang. Annotating derivations: A new evaluation strategy and dataset for algebra word problems. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (ACL), pp. 494–504, 2017.
  48. 48.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
  49. 49.Yan Wang, Xiaojiang Liu, and Shuming Shi. Deep neural solver for math word problems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 845–854, 2017.
  50. 50.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  51. 51.Ronald J Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3):229–256, 1992.
  52. 52.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 3911–3921, 2018.
  53. 53.Yilun Zhao, Yunxiang Li, Chenying Li, and Rui Zhang. Multihiertt: Numerical reasoning over multi hierarchical tabular and textual data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 6588–6600, 2022.
  54. 54.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning (ICML), pp. 12697–12706. PMLR, 2021.
  55. 55.Victor Zhong, Caiming Xiong, and Richard Socher. Seq2sql: Generating structured queries from natural language using reinforcement learning. arXiv preprint arXiv:1709.00103, 2017.
  56. 56.Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-JCNLP), pp. 3277–3287, 2021.

Citation

MLA
Lu, P., et al. “Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning”. arXiv, 2022, http://arxiv.org/abs/2209.14610v3.
APA
Lu, P., Qiu, L., Chang, K.-W., Wu, Y. N., Zhu, S.-C., Rajpurohit, T., Clark, P., & Kalyan, A. (2022). Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning. arXiv. http://arxiv.org/abs/2209.14610v3
Chicago
Lu, P., L. Qiu, K.-W. Chang, et al. 2022. “Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning”. arXiv. http://arxiv.org/abs/2209.14610v3.
Harvard
Lu, P. et al. (2022) “Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2209.14610v3.
Vancouver
1. Lu P, Qiu L, Chang K-W, Wu YN, Zhu S-C, Rajpurohit T, Clark P, Kalyan A (2022) Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning. arXiv

BibTeX

@article{lu2022dynamic,
  title = {Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning},
  author = {Lu, Pan and Qiu, Liang and Chang, Kai-Wei and Wu, Ying Nian and Zhu, Song-Chun and Rajpurohit, Tanmay and Clark, Peter and Kalyan, Ashwin},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2209.14610v3},
  eprint = {2209.14610}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors