LAMBADA: Backward Chaining for Automated Reasoning in Natural Language

Mehran KazemiNajoung KimDeepti BhatiaXin XuDeepak Ramachandran

article2023ACL107 citations

Proposes LAMBADA, a modular backward chaining framework that recursively decomposes natural language reasoning goals using few-shot prompted language models to significantly improve proof accuracy and query efficiency over forward reasoning methods.

Listen

Large language models frequently struggle with complex, multi-step logical reasoning and automated knowledge discovery. Existing methods primarily rely on forward reasoning—starting from known facts to deduce conclusions—which triggers an exponential combinatorial explosion of possible search paths as reasoning depth increases. Consequently, state-of-the-art models suffer from high failure rates, hallucinated proofs, and an inability to accurately handle unprovable statements.

The article evaluates whether backward chaining, a classical automated reasoning strategy that works in reverse from the target conclusion back to supporting facts, can provide more accurate, efficient, and robust text-based logical deduction in large language models. To test this, the authors introduce LAMBADA, a hybrid architecture combining backward-chaining search control with four specialized, few-shot prompted language model modules: fact verification, rule selection, goal decomposition, and polarity agreement.

The approach was benchmarked using the PaLM 540B model across challenging deduction datasets—including ProofWriter, PrOntoQA, and the naturalistic ParaRules corpus—featuring reasoning depths up to five steps and three-way classification outcomes: proved, disproved, or unknown. The evaluations measured label accuracy, proof validity, computational query efficiency, and robustness to lexical and syntactic variations.

The results show that backward chaining substantially outperforms standard prompting and forward modular systems. On 5-hop reasoning under open-world assumptions, LAMBADA achieved a 44% relative accuracy improvement over standard chain-of-thought prompting and a 56% improvement over the forward-chaining Selection-Inference baseline. On the naturalistic ParaRules dataset, LAMBADA attained 86% accuracy compared to 60% for chain-of-thought. Furthermore, an audit of 5-hop proofs revealed that while standard chain-of-thought produced structurally valid proofs in only 28% of seemingly correct answers—frequently relying on hallucinations and spurious correlations—LAMBADA generated valid proofs in 94% of cases. LAMBADA was also significantly more computationally efficient, requiring up to 11.8 times fewer model inference queries than modular forward-chaining alternatives at maximum depth.

These findings demonstrate that separating high-level search planning from natural language translation resolves core reliability issues in automated reasoning. For business and technical operations, this modular backward-chaining design drastically lowers operational compute costs and mitigates compliance, safety, and hallucination risks associated with deploying language models in high-stakes reasoning pipelines.

Organizations developing complex reasoning systems should transition from unconstrained forward-prompting techniques to structured, goal-directed modular workflows. Future initiatives should focus on extending this architecture to open-domain environments where rules are not fully provided in advance, supporting batch processing for lower latency, and exploring fine-tuning strategies on smaller, more cost-effective language models using backward reasoning traces.

The current implementation remains limited to deductive reasoning tasks with explicitly provided, prompt-sized rule sets, and its recursive nature demands sequential model queries that prevent out-of-the-box batching. Nevertheless, the experimental results provide high confidence that goal-directed backward chaining provides a superior, robust framework for automated text-based reasoning.

arXiv: 2212.13894

No sufficiently relevant recommendations were found.

Cover for LAMBADA: Backward Chaining for Automated Reasoning in Natural Language

Abstract

Remarkable progress has been made on automated reasoning with natural text, by using Language Models (LMs) and methods such as Chain-of-Thought and Selection-Inference. These techniques search for proofs in the forward direction from axioms to the conclusion, which suffers from a combinatorial explosion of the search space, and thus high failure rates for problems requiring longer chains of reasoning. The classical automated reasoning literature has shown that reasoning in the backward direction (i.e. from the intended conclusion to supporting axioms) is significantly more efficient at proof-finding. Importing this intuition into the LM setting, we develop a Backward Chaining algorithm, called LAMBADA, that decomposes reasoning into four sub-modules. These sub-modules are simply implemented by few-shot prompted LM inference. We show that LAMBADA achieves sizable accuracy boosts over state-of-the-art forward reasoning methods on two challenging logical reasoning datasets, particularly when deep and accurate proof chains are required.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 LAMBADA: Language Model Augmented Backward Chaining
  • 3.1 Backward Chaining
  • 3.2 LMModules in LAMBADA
  • 3.2.1 Fact Check
  • 3.2.2 Rule Selection
  • 3.2.3 Goal Decomposition
  • 3.2.4 Sign Agreement
  • 3.3 The LAMBADA Algorithm
  • 4 Experimental Setup
  • 4.1 Baselines
  • 4.2 Datasets
  • 5 Results
  • 5.1 Label Prediction Accuracy
  • 5.2 Proof Accuracy
  • 5.3 Forward vs. Backward Chaining
  • 5.4 Does Backward CoT Suffice?
  • 5.5 Qualitative Analysis
  • 5.6 Individual Module Analysis
  • 5.7 The Role of Scale
  • 5.8 Number of Inference Calls
  • 5.9 Lexical Robustness
  • 6 Conclusion and Future Directions
  • Limitations
  • References
  • A Caching and Avoiding Loops for LAMBADA
  • B Additional Results and Analyses
  • B.1 Qualitative Analysis
  • B.2 Further Analysis of CoT
  • B.3 Forward Chaining Becomes Progressively More Difficult
  • B.4 Confusion Matrices
  • B.5 Lexical Sensitivity Analysis
  • C Combinatorial Search Issue in Forward Chaining
  • D Implementation Details
  • D.1 Datasets for Individual Module Evaluation
  • D.2 Quality Issues in ParaRules
  • D.3 Prompts

Knowls

  1. Knowl 1 — LAMBADA performs goal-directed backward chaining with language-model modules

    algorithm

    LAMBADA takes a theory consisting of natural-language facts FF and implication rules RR, a goal GG, and a maximum search depth DD. It returns PROVED, DISPROVED, or UNKNOWN, where UNKNOWN means that neither the goal nor its negation was established. At each goal, it first checks whether a fact proves or disproves the goal; if so, it returns that result. If no fact settles the goal and the remaining depth is zero, it returns UNKNOWN. Otherwise, it identifies rules whose consequent unifies with the goal, tries shorter rules first, and decomposes the goal into the antecedent subgoals of each rule. It recursively evaluates every subgoal at depth D−1D-1 and treats the antecedents conjunctively: a rule can support the goal only if all its subgoals are PROVED. When they are, a sign-comparison step determines whether the rule consequent agrees with the goal, yielding PROVED if they agree and DISPROVED otherwise. If no selected rule has all its antecedents proved, the result is UNKNOWN. The search is depth-first, and rule length is used as a heuristic for trying rules with fewer likely subgoals earlier.

    To avoid repeated work, LAMBADA caches results for repeated theory-and-goal queries. A proved or disproved result obtained at depth dd can be reused at depths at least dd; an unknown result at depth dd can be reused at smaller depths. It also stops a branch when the same subgoal recurs on that branch, preventing exact-match cycles.

  2. Knowl 2 — Four prompted modules translate text into backward-chaining operations

    model/method

    LAMBADA uses few-shot prompted language-model inference for four operations. Fact Check selects a fact relevant to the current goal and determines whether that fact entails the goal, entails its negation, or settles neither. If a trial returns UNKNOWN, the selected fact is removed and the process can be retried; the experiments used two trials. Rule Selection first identifies each rule's consequent independently of the current goal, then compares those consequents with the goal and returns the rules whose consequents unify with it. Goal Decomposition takes a selected rule and a goal and produces the subgoals corresponding to the rule's antecedents. Sign Agreement compares the polarity of the goal with the polarity of the selected rule's consequent, determining whether proving the antecedent proves or disproves the goal. These operations let the symbolic search structure handle proof planning while the language model handles interpretation of natural-language facts, rules, and goals. The fact-selection setup is suited to the evaluated datasets, whose goals can be proved or disproved from a single fact; the authors note that selecting multiple facts would be needed in other settings.

  3. Knowl 3 — LAMBADA improves label accuracy on deeper and more naturalistic reasoning tasks

    empirical result

    Across ProofWriter, PrOntoQA, and ParaRules, LAMBADA generally outperformed Chain-of-Thought (CoT) and Selection-Inference (SI), with especially strong gains on examples requiring deeper proofs and on the three-label setting that includes UNKNOWN. On ProofWriter-PUD at depth 5, the plotted accuracies are approximately 0.50 for CoT, 0.46 for SI, and 0.72 for LAMBADA; the paper reports these as relative improvements of 44% over CoT and 56% over SI. On depth-5 PrOntoQA, the plotted accuracies are approximately 0.70 for CoT, 0.45 for SI, and 0.96 for LAMBADA, corresponding to reported relative improvements of 37% and 113%, respectively. On ParaRules, which uses more varied crowdworker-rewritten language, LAMBADA reaches 0.86 accuracy versus 0.60 for CoT, a reported relative improvement of 43%. The comparisons support the benefit of backward, modular reasoning, including in naturalistic text; they do not imply that LAMBADA wins by the same margin on every dataset or depth.

  4. Knowl 4 — Correct labels from CoT often lack valid proof chains

    empirical result

    The authors manually assessed proofs on 50 randomly sampled depth-5 ProofWriter-PD examples for which CoT had predicted the proof label correctly, and performed a comparable assessment of LAMBADA's proofs. Only 28% of the CoT proof chains were judged correct, compared with 94% of LAMBADA's. Hallucination was the main identified CoT error, accounting for 48% of the analyzed examples; other observed failures included mishandling conjunctions and invalid derivations. In the hallucination cases, CoT sometimes introduced unsupported facts or rules that provided shortcuts to the right label. Thus label accuracy alone overstated CoT's ability to produce faithful deductions on this evaluation.

  5. Knowl 5 — Forward-chaining inference quality declines as SI expands its theory

    empirical result

    On depth-5 PrOntoQA, the authors examined SI examples whose correct label was PROVED but for which SI failed to find a proof, and measured whether each successive inference was on the dataset's proof chain. The success rates for SI's first through fifth inferences were 0.53, 0.47, 0.34, 0.31, and 0.31. The theory grows as SI adds conclusions, so later selections must search a larger set of facts and rules; the observed decline is consistent with that search becoming harder. A separate analysis of five SI inferences on depth-5 ProofWriter-PUD examples where SI incorrectly returned UNKNOWN found exactly one unique inference in 7% of cases, two in 10%, three in 20%, four in 34%, and five in 29%. These results document both declining later-step success and repeated, redundant conclusions in the evaluated forward-chaining system.

  6. Knowl 6 — Backward-direction prompting alone does not reproduce LAMBADA's proof quality

    empirical result

    The authors compared ordinary forward CoT with a backward-CoT variant whose proof explanations proceed from the goal toward supporting premises. Their label accuracies were broadly comparable: forward CoT did better on ProofWriter-PUD, while backward CoT did better on ProofWriter-PD. However, backward CoT produced substantially less accurate proofs than forward CoT on the sampled depth-5 ProofWriter-PD examples. The comparison indicates that reversing the direction of a free-form CoT explanation was not sufficient to obtain LAMBADA's proof quality; the paper attributes the stronger result to the modular formulation that constrains individual reasoning operations.

  7. Knowl 7 — LAMBADA uses substantially fewer language-model calls than SI

    empirical result

    On ProofWriter-PUD, the average number of language-model inference calls per example at depths 0, 1, 2, 3, and 5 was 2.98, 7.26, 10.14, 12.77, and 18.53 for LAMBADA, compared with 27.79, 27.30, 57.22, 99.45, and 219.34 for SI. The paper summarizes the difference as 3.8 times fewer calls for LAMBADA at depth 1 and 11.8 times fewer at depth 5. This measures model calls, not total compute or latency; the authors also note that LAMBADA still requires more calls than a single-pass approach such as CoT.

  8. Knowl 8 — LAMBADA is robust to novel lexical items and rule templates

    empirical result

    To test sensitivity to surface form, the authors modified ProofWriter-PUD test examples by replacing entity names, animal names, adjectives, and verbs with items absent from the demonstrations and dataset, and separately replaced rule templates with novel templates. They created two variants for each type of modification and evaluated LAMBADA using the same few-shot examples as on the original test set. Accuracy across the modified sets remained in roughly the same range as on the original set, with some depth-specific fluctuations rather than a consistent deterioration. The authors also report that LAMBADA on these modified sets remained substantially more accurate than the baselines' results on the original test set.

  9. Knowl 9 — Module accuracy depends on the scale of the language model

    empirical result

    On isolated module-evaluation sets drawn from ProofWriter validation data, PaLM 540B achieved 0.94 accuracy on Fact Check with one selected-fact trial, while allowing two trials raised performance to near perfect; Sign Agreement was also near perfect. Rule Selection was the weakest module, followed by Goal Decomposition. With PaLM 62B, Goal Decomposition and Sign Agreement remained comparatively strong, whereas Fact Check and Rule Selection dropped substantially. With PaLM 8B, accuracy fell markedly for all four modules, in some cases approaching chance performance. The authors suggest that smaller language models may need finer-grained decomposition, particularly for the one-to-many comparisons used in selection.

  10. Knowl 10 — Evaluation used few-shot PaLM inference on three deductive benchmarks

    experimental setup

    Unless otherwise specified, the experiments used PaLM 540B with decoding temperature zero and few-shot prompting rather than task-specific fine-tuning. ProofWriter was evaluated on its open-world-assumption subset using the first 1,000 test examples: ProofWriter-PD removes UNKNOWN examples for binary evaluation, while ProofWriter-PUD retains PROVED, DISPROVED, and UNKNOWN; the benchmark includes reasoning depths 0, at most 1, at most 2, at most 3, and at most 5 hops. PrOntoQA used the fictional-character version at depths 1, 3, and 5. ParaRules uses naturalistic rewrites of ProofWriter-style facts and rules and includes examples requiring up to five hops; the authors manually checked and corrected the first 500 test examples and evaluated on that set. For binary evaluations, LAMBADA's UNKNOWN and DISPROVED predictions were combined into one class. The comparisons included CoT and SI, with SI omitted from the ParaRules comparison because of its low performance on the other evaluated datasets and its high number of model calls.

  11. Knowl 11 — The evaluated formulation has explicit scope limitations

    limitation

    The reported LAMBADA method addresses classification-style deductive entailment: deciding whether a goal is proved, disproved, or neither, rather than answering open-ended questions such as identifying an entity's color. It assumes that all rules are supplied and that the full rule set fits in the prompt, and it is limited to deduction using modus ponens rather than rule forms such as proof by contradiction or disjunction elimination. Its recursive module calls depend on earlier call outputs, preventing straightforward batching of those calls. Although it makes fewer language-model calls than SI, it still makes substantially more calls than CoT, increasing inference cost. These are stated limitations of the studied formulation, not claims that the method cannot be extended.

Coverage note — The paper's illustrative search traces and detailed prompt exemplars are omitted because they demonstrate the method rather than add distinct results; auxiliary confusion matrices and the manual ParaRules correction categories are likewise not expanded into separate knowls.

References

  1. 1.Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. 2022. Exploring length generalization in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 38546–38556. Curran Associates, Inc.
  2. 2.Gregor Betz, Christian Voigt, and Kyle Richardson. 2021. Critical thinking for language models. In Proceedings of the 14th International Conference on Computational Semantics (IWCS), pages 63–75, Groningen, The Netherlands (online). Association for Computational Linguistics.
  3. 3.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghe-mawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankara-narayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. arXiv:2204.02311.
  6. 6.Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2021. Transformers as soft reasoners over language. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI’20.
  7. 7.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv:2110.14168.
  8. 8.Antonia Creswell and Murray Shanahan. 2022. Faithful reasoning using large language models. arXiv:2208.14271.
  9. 9.Antonia Creswell, Murray Shanahan, and Irina Higgins. 2023. Selection-inference: Exploiting large language models for interpretable logical reasoning. In The Eleventh International Conference on Learning Representations.
  10. 10.Ido Dagan, Dan Roth, Mark Sammons, and Fabio Massimo Zanzotto. 2013. Recognizing textual entailment: Models and applications. Synthesis Lectures on Human Language Technologies, 6(4):1–220.
  11. 11.Bhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, and Peter Clark. 2021. Explaining answers with entailment trees. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7358–7370, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  12. 12.Dheeru Dua, Shivanshu Gupta, Sameer Singh, and Matt Gardner. 2022. Successive prompting for decomposing complex questions. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1251–1265, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  13. 13.Artur d’Avila Garcez and Luis C Lamb. 2020. Neurosymbolic ai: the 3rd wave. arXiv:2012.05876.
  14. 14.Nicolas Gontier, Koustuv Sinha, Siva Reddy, and Chris Pal. 2020. Measuring systematic generalization in neural proof generation with transformers. In Advances in Neural Information Processing Systems, volume 33, pages 22231–22242. Curran Associates, Inc.
  15. 15.Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Luke Benson, Lucy Sun, Ekaterina Zubova, Yujie Qiao, Matthew Burtell, et al. 2022. FOLIO: Natural language reasoning with first-order logic. arXiv:2209.00840.
  16. 16.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1. Curran.
  17. 17.Carl Hewitt. 1969. Planner: A language for proving theorems in robots. In Proceedings of the 1st International Joint Conference on Artificial Intelligence, IJCAI’69, page 295–301, San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.
  18. 18.Jie Huang and Kevin Chen-Chuan Chang. 2022. Towards reasoning in large language models: A survey. arXiv:2212.10403.
  19. 19.Harsh Jhamtani and Peter Clark. 2020. Learning to explain: Datasets and models for identifying valid reasoning chains in multihop question-answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 137–150, Online. Association for Computational Linguistics.
  20. 20.Nora Kassner, Benno Krojer, and Hinrich Schütze. 2020. Are pretrained language models symbolic reasoners over knowledge? In Proceedings of the 24th Conference on Computational Natural Language Learning, pages 552–564, Online. Association for Computational Linguistics.
  21. 21.Mehran Kazemi, Sid Mittal, and Deepak Ramachandran. 2023. Understanding finetuning for factual knowledge extraction from language models. arXiv:2301.11293.
  22. 22.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2023. Decomposed prompting: A modular approach for solving complex tasks. In The Eleventh International Conference on Learning Representations.
  23. 23.Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023. Transformers learn shortcuts to automata. In The Eleventh International Conference on Learning Representations.
  24. 24.Gary Marcus. 2020. The next decade in AI: four steps towards robust artificial intelligence. arXiv:2002.06177.
  25. 25.John McCarthy. 1959. Programs with common sense. In Proceedings of the Teddington Conference on the Mechanization of Thought Processes, pages 75–91, London. Her Majesty’s Stationary Office.
  26. 26.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2022. Show your work: Scratchpads for intermediate computation with language models. In Deep Learning for Code Workshop.
  27. 27.David L Poole and Alan K Mackworth. 2010. Artificial Intelligence: foundations of computational agents. Cambridge University Press.
  28. 28.Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving. 2021. Scaling language models: Methods, analysis & insights from training Gopher. arXiv:2112.11446.
  29. 29.Mohammed Saeed, Naser Ahmadi, Preslav Nakov, and Paolo Papotti. 2021. RuleBERT: Teaching soft rules to pre-trained language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1460–1476, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  30. 30.Abulhair Saparov and He He. 2023. Language models are greedy reasoners: A systematic formal analysis of chain-of-thought. In The Eleventh International Conference on Learning Representations.
  31. 31.Imanol Schlag, Sainbayar Sukhbaatar, Asli Celikyilmaz, Wen-tau Yih, Jason Weston, Jürgen Schmidhuber, and Xian Li. 2023. Large language model programs. arXiv preprint arXiv:2305.05364.
  32. 32.Viktor Schlegel, Kamen Pavlov, and Ian Pratt-Hartmann. 2022. Can transformers reason in fragments of natural language? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11184–11199, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  33. 33.Jianhao Shen, Yichun Yin, Lin Li, Lifeng Shang, Xin Jiang, Ming Zhang, and Qun Liu. 2021. Generate & rank: A multi-task framework for math word problems. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2269–2279, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  34. 34.Zayne Sprague, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett. 2022. Natural language deduction with incomplete information. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8230–8258, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  35. 35.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. 2022. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv:2210.09261.
  36. 36.Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2021. ProofWriter: Generating implications, proofs, and abductive statements over natural language. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3621–3634, Online. Association for Computational Linguistics.
  37. 37.Karthik Valmeekam, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati. 2022. Large language models still can’t plan (a benchmark for LLMs on planning and reasoning about change). In NeurIPS 2022 Foundation Models for Decision Making Workshop.
  38. 38.Boshi Wang, Xiang Deng, and Huan Sun. 2022. Iteratively prompt pre-trained language models for chain of thought. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2714–2730, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  39. 39.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc.
  40. 40.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics.
  41. 41.Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022. STaR: Bootstrapping reasoning with reasoning. In Advances in Neural Information Processing Systems, volume 35, pages 15476–15488. Curran Associates, Inc.
  42. 42.Hanlin Zhang, Ziyang Li, Jiani Huang, Mayur Naik, and Eric Xing. 2022a. Improved logical reasoning of language models via differentiable symbolic programming. In First Workshop on Pre-training: Perspectives, Pitfalls, and Paths Forward at ICML 2022.
  43. 43.Honghua Zhang, Liunian Harold Li, Tao Meng, Kai-Wei Chang, and Guy Van den Broeck. 2022b. On the paradox of learning to reason from data. arXiv:2205.11502.
  44. 44.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H. Chi. 2023. Least-to-most prompting enables complex reasoning in large language models. In The Eleventh International Conference on Learning Representations.
  45. 45.Hattie Zhou, Azade Nova, Hugo Larochelle, Aaron Courville, Behnam Neyshabur, and Hanie Sedghi. 2022. Teaching algorithmic reasoning via in-context learning. arXiv:2211.09066.

Citation

MLA
Kazemi, M., et al. “LAMBADA: Backward Chaining for Automated Reasoning in Natural Language”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 6547–68, https://doi.org/10.18653/v1/2023.acl-long.361.
APA
Kazemi, M., Kim, N., Bhatia, D., Xu, X., & Ramachandran, D. (2023). LAMBADA: Backward Chaining for Automated Reasoning in Natural Language. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6547–6568. https://doi.org/10.18653/v1/2023.acl-long.361
Chicago
Kazemi, M., N. Kim, D. Bhatia, X. Xu, and D. Ramachandran. 2023. “LAMBADA: Backward Chaining for Automated Reasoning in Natural Language”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6547–68. https://doi.org/10.18653/v1/2023.acl-long.361.
Harvard
Kazemi, M. et al. (2023) “LAMBADA: Backward Chaining for Automated Reasoning in Natural Language”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6547–6568. Available at: https://doi.org/10.18653/v1/2023.acl-long.361.
Vancouver
1. Kazemi M, Kim N, Bhatia D, Xu X, Ramachandran D (2023) LAMBADA: Backward Chaining for Automated Reasoning in Natural Language. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6547–6568

BibTeX

@inproceedings{kazemi-etal-2023-lambada,
    title = "{LAMBADA}: Backward Chaining for Automated Reasoning in Natural Language",
    author = "Kazemi, Mehran  and
      Kim, Najoung  and
      Bhatia, Deepti  and
      Xu, Xin  and
      Ramachandran, Deepak",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.361/",
    doi = "10.18653/v1/2023.acl-long.361",
    pages = "6547--6568"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/