Are NLP Models really able to Solve Simple Math Word Problems?

Arkil PatelSatwik BhattamishraNavin Goyal

article2021NAACL1,346 citations

Demonstrates that current math word problem solvers rely on shallow heuristics rather than mathematical reasoning, and introduces the SVAMP challenge benchmark to accurately evaluate arithmetic comprehension in language models.

Listen

Recent advances in natural language processing have led to high test scores on standard arithmetic word problem benchmarks, creating a common belief that elementary-level math problems are effectively solved. Consequently, research and development resources have shifted toward more complex challenges. However, high accuracy on existing benchmarks may reflect reliance on superficial cues rather than genuine mathematical reasoning, creating performance risks for real-world automated educational tools.

The article evaluates whether current state-of-the-art models truly solve simple, elementary-school math word problems or merely exploit statistical shortcuts in standard benchmarks. To test this, the authors introduce a new challenge dataset to reliably measure model robustness against basic reasoning and structural variations.

The authors conducted an empirical investigation using leading model architectures (Sequence-to-Sequence, Goal-Driven Tree-Structured models, and Graph-to-Tree models) across two standard benchmarks: MAWPS (2,373 problems) and ASDiv-A (1,218 problems). They first stripped the question sentences from the problems to test if models could solve them using only the descriptive narrative. Next, they evaluated a constrained baseline stripped of word-order awareness to see if bag-of-words keyword associations sufficed. Finally, the authors constructed SVAMP, a challenge dataset of 1,000 elementary-level problems created by applying controlled variations—altering question focus, modifying reasoning steps, or inserting irrelevant details—to seed problems.

The investigation produced four critical findings. First, existing benchmarks are heavily compromised by superficial artifacts: top models solved 64.4% of ASDiv-A problems and 77.7% of MAWPS problems with the question entirely removed. Second, a constrained model lacking any word-order understanding achieved 77.9% accuracy on MAWPS and 51.2% on ASDiv-A by latching onto single trigger words. Third, when evaluated on SVAMP, top-performing models experienced severe performance drops, with the best model achieving only 43.8% accuracy despite the dataset containing no higher-level operations. Fourth, model accuracy degraded substantially as problem complexity grew slightly, dropping from 78.3% on two-number problems in SVAMP to just 25.4% on problems containing three or four numbers.

These findings demonstrate that high benchmark accuracies are misleading; current models rely heavily on shallow heuristics rather than actual language understanding or numerical reasoning. Deploying these systems into real-world settings, such as automated tutoring, carries substantial performance and user-trust risks because slight changes in phrasing cause models to fail unpredictably. The assumption that basic arithmetic word problems are solved is incorrect.

Stakeholders and developers should halt the premature shift away from elementary reasoning tasks and incorporate challenge sets like SVAMP into evaluation pipelines to ensure robust testing. Training regimes must combine standard datasets with structural variations to mitigate shallow heuristic exploitation. Moving forward, research should focus on developing model architectures capable of genuine contextual reasoning and number binding before deploying automated solvers in high-stakes educational applications.

The scope of the article is limited to English-language, one-unknown arithmetic problems up to a fourth-grade level. While the results clearly demonstrate the brittleness of current language models on these benchmarks, further research is required to determine whether these vulnerabilities generalize across other languages and more complex multi-variable domains.

Cover for Are NLP Models really able to Solve Simple Math Word Problems?

Abstract

The problem of designing NLP solvers for math word problems (MWP) has seen sustained research activity and steady gains in the test accuracy. Since existing solvers achieve high performance on the benchmark datasets for elementary level MWPs containing one-unknown arithmetic word problems, such problems are often considered "solved" with the bulk of research attention moving to more complex MWPs. In this paper, we restrict our attention to English MWPs taught in grades four and lower. We provide strong evidence that the existing MWP solvers rely on shallow heuristics to achieve high performance on the benchmark datasets. To this end, we show that MWP solvers that do not have access to the question asked in the MWP can still solve a large fraction of MWPs. Similarly, models that treat MWPs as bag-of-words can also achieve surprisingly high accuracy. Further, we introduce a challenge dataset, SVAMP, created by applying carefully chosen variations over examples sampled from existing datasets. The best accuracy achieved by state-of-the-art models is substantially lower on SVAMP, thus showing that much remains to be done even for the simplest of the MWPs.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background
  • 3.1 Problem Formulation
  • 3.2 Datasets and Methods
  • 4 Deficiencies in existing datasets
  • 4.1 Evaluation on Question-removed MWPs
  • 4.2 Performance of a constrained model
  • 4.3 Analyzing the attention weights
  • 5 SVAMP
  • 5.1 Creating SVAMP
  • 5.1.1 Variations
  • 5.1.2 Protocol for creating variations
  • 5.2 Dataset Properties
  • 5.3 Experiments on SVAMP
  • 6 Final Remarks
  • References
  • A Experiments with Transformer
  • B Implementation Details
  • C Creation Protocol
  • D Analyzing Attention Weights
  • E Examples of Simple Problems
  • F Ethical Considerations
  • Appendix: Are NLP Models really able to Solve Simple Math Word Problems?
  • A Implementation Details
  • B Creation Protocol
  • C Analyzing Attention Weights
  • References

Knowls

  1. Knowl 1 — SVAMP Challenge Dataset for Elementary Math Word Problems

    definition

    The SVAMP (Simple Variations on Arithmetic Math word Problems) dataset is a challenge evaluation benchmark consisting of 1,000 one-unknown English arithmetic word problems taught in grade 4 and below. Every problem can be solved by a mathematical expression using numbers from the narrative and arithmetic operators from the set {+,−,∗,/}\{+, -, *, /\} with at most two operators.

    SVAMP is constructed by applying controlled variations to 100 seed problems selected from the ASDiv-A dataset. The 100 seed examples are chosen by partitioning single-operator problems in ASDiv-A across arithmetic types (28 Addition, 33 Subtraction, 19 Multiplication, 20 Division) and selecting the centroid problem from KK-Means clustering over RoBERTa sentence embeddings within each type.

    Dataset # Problems # Equation Templates # Avg Operators Corpus Lexicon Diversity (CLD)
    MAWPS 2373 39 1.78 0.26
    ASDiv-A 1218 19 1.23 0.50
    SVAMP 1000 26 1.24 0.22

    An Equation Template is defined by formatting the target equation into prefix notation and masking all numbers with a generic symbol.

  2. Knowl 2 — Performance Collapse of State-of-the-Art Math Word Problem Solvers on SVAMP

    empirical result

    State-of-the-art neural math word problem solvers trained on the combination of standard benchmark datasets MAWPS and ASDiv-A experience a substantial performance drop when evaluated on SVAMP, failing to solve more than half of the elementary-level problems.

    Seq2Seq GTS Graph2Tree
    Subset Scratch RoBERTa Scratch RoBERTa Scratch RoBERTa
    Full Set 24.2% 40.3% 30.8% 41.0% 36.5% 43.8%
    One-Operator 25.4% 42.6% 31.7% 44.6% 42.9% 51.9%
    Two-Operator 20.3% 33.1% 27.9% 29.7% 16.1% 17.8%
    Addition 28.5% 41.9% 35.8% 36.3% 24.9% 36.8%
    Subtraction 22.3% 35.1% 26.7% 36.9% 41.3% 41.3%
    Multiplication 17.9% 38.7% 29.2% 38.7% 27.4% 35.8%
    Division 29.3% 56.3% 39.5% 61.1% 40.7% 65.3%

    On the benchmark datasets, these models achieve execution accuracies between 76.9%76.9\% and 88.7%88.7\%. On SVAMP, the best model (Graph2Tree with RoBERTa pretrained embeddings) achieves an execution accuracy of only 43.8%43.8\%. Accuracy is particularly low on two-operator problems (17.8%17.8\% for Graph2Tree with RoBERTa). When 5-fold cross-validation is performed with SVAMP included in the training data alongside MAWPS and ASDiv-A, the best model achieves only approximately 65%65\% accuracy.

  3. Knowl 3 — Heuristic Reliance of MWP Solvers on Question-Removed Benchmarks

    empirical result

    Math word problem models can correctly solve a large majority of benchmark problems even when the question text is entirely removed from test instances, demonstrating that standard benchmarks contain strong superficial artifacts in their narrative bodies.

    A math word problem P=(w1,…,wn)P = (w_1, \dots, w_n) is composed of a narrative body B=(w1,…,wk)B = (w_1, \dots, w_k) and a question Q=(wk+1,…,wn)Q = (w_{k+1}, \dots, w_n). When baseline models with RoBERTa embeddings are trained normally on MAWPS and ASDiv-A but tested on instances where QQ is deleted:

    Model MAWPS (w/o Question) ASDiv-A (w/o Question)
    Seq2Seq 77.4% 58.7%
    GTS 76.2% 60.7%
    Graph2Tree 77.7% 64.4%

    When test problems are partitioned into an Easy subset (problems the model solves correctly without QQ) and a Hard subset (problems the model fails without QQ), models exhibit a massive performance discrepancy on ASDiv-A (e.g., Graph2Tree achieves 92.8%92.8\% on Easy vs. 63.3%63.3\% on Hard; GTS achieves 91.6%91.6\% on Easy vs. 65.3%65.3\% on Hard), indicating that overall benchmark accuracy is inflated by shallow heuristics.

  4. Knowl 4 — Shallow Word-Association Heuristics in Orderless Constrained MWP Solvers

    empirical result

    A constrained neural architecture that completely lacks word-order information can solve the majority of benchmark math word problems by memorizing associations between individual trigger words and target arithmetic operations.

    The constrained model replaces the sequential encoder with a Feed-Forward Network (FFN) that independently projects non-contextual token embeddings into hidden representations without recurrence or attention. The average of these hidden vectors initializes an LSTM decoder with Luong attention over the unordered representations.

    Constrained Model Variant MAWPS ASDiv-A
    FFN + LSTM Decoder (Trained from Scratch) 75.1% 46.3%
    FFN + LSTM Decoder (Non-contextual RoBERTa Embeddings) 77.9% 51.2%
    Majority Template Baseline 17.7% 21.2%

    Analyzing the attention weights of the trained constrained model shows that it typically assigns an attention weight of nearly 1.01.0 to a single keyword (such as 'every', 'each', 'left', 'sold', or 'cut off') and outputs the corresponding operator regardless of how the surrounding narrative is altered.

  5. Knowl 5 — Taxonomy of MWP Variations for Robust Evaluation

    definition

    A framework of 9 fine-grained variations categorized into three core axes designed to assess model robustness on arithmetic word problems:

    1. Question Sensitivity: Assesses whether model predictions depend on the question asked while keeping the body BB unchanged.
    • Same Object, Different Structure: The target unknown object is retained, but the syntactic structure of the question is modified.
    • Different Object, Same Structure: The question asks about a different quantity/entity, but uses the same syntactic structure.
    • Different Object, Different Structure: Both the target entity and the syntactic structure of the question are altered.
    1. Reasoning Ability: Assesses whether the model updates its mathematical reasoning in response to semantic modifications in the narrative.
    • Add relevant information: Extra facts are introduced into the narrative that change the output equation.
    • Change information: Core factual quantities or relationships in the narrative are altered.
    • Invert operation: The original unknown quantity is provided as a known value, and the question asks for a previously given quantity.
    1. Structural Invariance: Assesses whether the model remains invariant to superficial phrasing and ordering changes that preserve the underlying mathematical equation.
    • Add irrelevant information: Numerically or contextually distractive details are added that do not affect the solution.
    • Change order of objects: The order of entities or nouns in the text is swapped.
    • Change order of phrases: The order of number-containing clauses/sentences is permuted.
  6. Knowl 6 — Diagnostic Evaluation of Question Sensitivity and Orderless Baselines on SVAMP

    empirical result

    Evaluating question-removed and word-orderless models on SVAMP confirms that the dataset successfully eliminates the superficial heuristics present in MAWPS and ASDiv-A.

    Model (RoBERTa) SVAMP (w/o Question) ASDiv-A (w/o Question)
    Seq2Seq 29.2% 58.7%
    GTS 28.6% 60.7%
    Graph2Tree 30.8% 64.4%

    When question text is removed on SVAMP, model accuracy falls to approximately 29%–31%29\%\text{--}31\%, roughly half the accuracy achieved on question-removed ASDiv-A (58.7%–64.4%58.7\%\text{--}64.4\%).

    Furthermore, the bag-of-words constrained model (FFN + LSTM decoder with non-contextual RoBERTa embeddings) trained on MAWPS and ASDiv-A achieves only 18.3%18.3\% accuracy on SVAMP (and 17.5%17.5\% when trained from scratch), which is only marginally higher than the majority template baseline of 11.7%11.7\% and far below its 77.9%77.9\% performance on MAWPS and 51.2%51.2\% on ASDiv-A.

  7. Knowl 7 — Impact of Specific Variation Categories on SVAMP Model Accuracy

    empirical result

    Ablation analysis on SVAMP demonstrates how different variation categories contribute to model failure. The metric Δ\Delta is the change in execution accuracy of the best-performing model (Graph2Tree with RoBERTa) when all examples containing a given category or variation are removed from the test set: Δ=Acc(Full−Category)−Acc(Full)\Delta = \text{Acc}(\text{Full} - \text{Category}) - \text{Acc}(\text{Full})

    Removed Category / Variation # Removed Examples Change in Accuracy (Δ\Delta)
    Question Sensitivity (Category) 462 +13.7%
    – Same Object, Different Structure 325 +7.3%
    – Different Object, Same Structure 69 +1.5%
    – Different Object, Different Structure 74 +1.3%
    Structural Invariance (Category) 467 +4.5%
    – Add Irrelevant Information 281 +6.9%
    – Change Order of Objects 107 +2.3%
    – Change Order of Phrases 152 -3.3%
    Reasoning Ability (Category) 649 -3.3%
    – Add Relevant Information 264 +5.5%
    – Change Information 149 +3.2%
    – Invert Operation 255 -10.2%

    Removing Question Sensitivity variations yields the largest gain (+13.7%+13.7\%), identifying it as the primary challenge for models. Invert Operation is the only variation whose removal sharply decreases accuracy (−10.2%-10.2\%), because inverted problems resemble standard training instances from ASDiv-A.

  8. Knowl 8 — Sensitivity of MWP Solvers to the Number of Quantities in Context

    empirical result

    Current state-of-the-art MWP models degrade sharply in accuracy when an arithmetic problem contains more than two numbers in its text, revealing an inability to bind quantities to their syntactic context.

    Evaluating the Graph2Tree model (trained on MAWPS and ASDiv-A with RoBERTa embeddings) across subsets of problems grouped by the count of numbers present in the narrative:

    Evaluation Setting 2 Numbers 3 Numbers 4 Numbers
    ASDiv-A (5-fold Cross-Validation) 93.3% 59.0% 47.5%
    SVAMP (Zero-shot Transfer) 78.3% 25.4% 25.4%

    While Graph2Tree solves 78.3%78.3\% of 2-number problems in SVAMP, performance plummets to 25.4%25.4\% on 3-number and 4-number problems, showing that adding irrelevant or multi-step numerical values causes catastrophic failure.

  9. Knowl 9 — Stepwise Protocol for Constructing Math Word Problem Challenge Variations

    algorithm

    A systematic 6-step ordering protocol for generating varied math word problems from a base seed template to ensure structural coverage without combinatorial inconsistency:

    Input: Base Example template with tagged numbers (NUM), person names (NAME), objects (OBJs, OBJp), and modifiers (MOD)
    Output: Set of validated Variation Examples with equations and variation labels
    Step 1: Apply Question Sensitivity variations to the Base Example (altering target unknown entity or question syntax while holding body constant)
    Step 2: Apply Invert Operation variation to the Base Example and to all variations generated in Step 1
    Step 3: Apply Add relevant information variation to the Base Example; then treat the resulting examples as Base Examples and apply Question Sensitivity variations
    Step 4: Apply Add irrelevant information variation to the Base Example and to all variations obtained so far
    Step 5: Apply Change information variation to the Base Example and to all variations obtained so far
    Step 6: Apply Change order of Objects and Change order of Phrases variations to the Base Example and to all variations obtained so far
    Discard any resulting problem requiring more than two operators
    Verify grammatical and logical correctness through independent human review
  10. Knowl 10 — Insufficiency of Lexical Diversity for Assessing MWP Dataset Quality

    theoretical result

    Corpus Lexicon Diversity (CLD) is insufficient as a proxy metric for math word problem dataset difficulty or quality.

    Although ASDiv-A has a higher CLD (0.500.50) than MAWPS (0.260.26) and SVAMP (0.220.22), state-of-the-art models achieve 82.2%82.2\% on ASDiv-A but only 43.8%43.8\% on SVAMP. Furthermore, a Graph2Tree model trained on ASDiv-A achieves 82%82\% accuracy when transferred to MAWPS, whereas training on MAWPS and testing on ASDiv-A yields only 73%73\%. Dataset difficulty and solver robustness are governed by structural variation, entity-quantity associations, and sensitivity to the question, rather than vocabulary breadth.

Coverage note — None was omitted. All primary contributions—including the diagnosis of dataset artifacts via question removal and orderless models, the creation protocol and properties of the SVAMP dataset, and extensive evaluation and ablation results—are captured in the knowls.

References

  1. 1.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. MathQA: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2357–2367, Minneapolis, Minnesota. Association for Computational Linguistics.
  2. 2.Yonatan Belinkov and James Glass. 2019. Analysis methods in neural language processing: A survey. Transactions of the Association for Computational Linguistics, 7:49–72.
  3. 3.Zheng Cai, Lifu Tu, and Kevin Gimpel. 2017. Pay attention to the ending:strong neural baselines for the ROC story cloze task. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 616–622, Vancouver, Canada. Association for Computational Linguistics.
  4. 4.Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hanna Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020. Evaluating models’ local decision boundaries via contrast sets.
  5. 5.Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 107–112, New Orleans, Louisiana. Association for Computational Linguistics.
  6. 6.Danqing Huang, Shuming Shi, Chin-Yew Lin, and Jian Yin. 2017. Learning fine-grained expressions to solve math word problems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 805–814, Copenhagen, Denmark. Association for Computational Linguistics.
  7. 7.Danqing Huang, Shuming Shi, Chin-Yew Lin, Jian Yin, and Wei-Ying Ma. 2016a. How well do computers solve math word problems? large-scale dataset construction and evaluation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 887–896, Berlin, Germany. Association for Computational Linguistics.
  8. 8.Danqing Huang, Shuming Shi, Chin-Yew Lin, Jian Yin, and Wei-Ying Ma. 2016b. How well do computers solve math word problems? large-scale dataset construction and evaluation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 887–896, Berlin, Germany. Association for Computational Linguistics.
  9. 9.Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. MAWPS: A math word problem repository. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1152–1157, San Diego, California. Association for Computational Linguistics.
  10. 10.Tal Linzen. 2020. How can we accelerate progress towards human-like linguistic generalization? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5210–5217, Online. Association for Computational Linguistics.
  11. 11.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  12. 12.Thang Luong, Hieu Pham, and Christopher D. Manning. 2015. Effective approaches to attention-based neural machine translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 1412–1421, Lisbon, Portugal. Association for Computational Linguistics.
  13. 13.Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428–3448, Florence, Italy. Association for Computational Linguistics.
  14. 14.Shen-yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2020. A diverse corpus for evaluating and developing English math word problem solvers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 975–984, Online. Association for Computational Linguistics.
  15. 15.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial NLI: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4885–4901, Online. Association for Computational Linguistics.
  16. 16.Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018. Hypothesis only baselines in natural language inference. In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, pages 180–191, New Orleans, Louisiana. Association for Computational Linguistics.
  17. 17.Jinghui Qin, Lihui Lin, Xiaodan Liang, Rumin Zhang, and Liang Lin. 2020. Semantically-aligned universal tree-structured solver for math word problems.
  18. 18.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online. Association for Computational Linguistics.
  19. 19.Shachar Rosenman, Alon Jacovi, and Yoav Goldberg. 2020. Exposing Shallow Heuristics of Relation Extraction Models with Challenge Data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3702–3710, Online. Association for Computational Linguistics.
  20. 20.Subhro Roy and Dan Roth. 2018. Mapping to declarative knowledge for word problem solving. Transactions of the Association for Computational Linguistics, 6:159–172.
  21. 21.Mrinmaya Sachan and Eric Xing. 2017. Learning to solve geometry problems from natural language demonstrations in textbooks. In *Proceedings of the 6th Joint Conference on Lexical and Computational Semantics (SEM 2017), pages 251–261, Vancouver, Canada. Association for Computational Linguistics.
  22. 22.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  23. 23.Yan Wang, Xiaojiang Liu, and Shuming Shi. 2017. Deep neural solver for math word problems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 845–854, Copenhagen, Denmark. Association for Computational Linguistics.
  24. 24.Sarah Wiegreffe and Yuval Pinter. 2019. Attention is not not explanation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 11–20, Hong Kong, China. Association for Computational Linguistics.
  25. 25.Zhipeng Xie and Shichao Sun. 2019. A goal-driven tree-structured neural model for math word problems. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19, pages 5299–5305. International Joint Conferences on Artificial Intelligence Organization.
  26. 26.D. Zhang, L. Wang, L. Zhang, B. T. Dai, and H. T. Shen. 2020. The gap of semantic parsing: A survey on automatic math word problem solvers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(9):2287–2305.
  27. 27.Jipeng Zhang, Lei Wang, Roy Ka-Wei Lee, Yi Bin, Yan Wang, Jie Shao, and Ee-Peng Lim. 2020. Graph-to-tree learning for solving math word problems. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3928–3937, Online. Association for Computational Linguistics.

Citation

MLA
Patel, A., et al. “Are NLP Models Really Able to Solve Simple Math Word Problems?”. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021, pp. 2080–94, https://doi.org/10.18653/v1/2021.naacl-main.168.
APA
Patel, A., Bhattamishra, S., & Goyal, N. (2021). Are NLP Models really able to Solve Simple Math Word Problems?. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2080–2094. https://doi.org/10.18653/v1/2021.naacl-main.168
Chicago
Patel, A., S. Bhattamishra, and N. Goyal. 2021. “Are NLP Models Really Able to Solve Simple Math Word Problems?”. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2080–94. https://doi.org/10.18653/v1/2021.naacl-main.168.
Harvard
Patel, A., Bhattamishra, S. and Goyal, N. (2021) “Are NLP Models really able to Solve Simple Math Word Problems?”, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 2080–2094. Available at: https://doi.org/10.18653/v1/2021.naacl-main.168.
Vancouver
1. Patel A, Bhattamishra S, Goyal N (2021) Are NLP Models really able to Solve Simple Math Word Problems?. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 2080–2094

BibTeX

@inproceedings{patel-etal-2021-nlp,
    title = "Are {NLP} Models really able to Solve Simple Math Word Problems?",
    author = "Patel, Arkil  and
      Bhattamishra, Satwik  and
      Goyal, Navin",
    editor = "Toutanova, Kristina  and
      Rumshisky, Anna  and
      Zettlemoyer, Luke  and
      Hakkani-Tur, Dilek  and
      Beltagy, Iz  and
      Bethard, Steven  and
      Cotterell, Ryan  and
      Chakraborty, Tanmoy  and
      Zhou, Yichao",
    booktitle = "Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jun,
    year = "2021",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.naacl-main.168/",
    doi = "10.18653/v1/2021.naacl-main.168",
    pages = "2080--2094"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/