LAMBADA: Backward Chaining for Automated Reasoning in Natural Language
Mehran KazemiNajoung KimDeepti BhatiaXin XuDeepak Ramachandran
Proposes LAMBADA, a modular backward chaining framework that recursively decomposes natural language reasoning goals using few-shot prompted language models to significantly improve proof accuracy and query efficiency over forward reasoning methods.
Large language models frequently struggle with complex, multi-step logical reasoning and automated knowledge discovery. Existing methods primarily rely on forward reasoning—starting from known facts to deduce conclusions—which triggers an exponential combinatorial explosion of possible search paths as reasoning depth increases. Consequently, state-of-the-art models suffer from high failure rates, hallucinated proofs, and an inability to accurately handle unprovable statements.
The article evaluates whether backward chaining, a classical automated reasoning strategy that works in reverse from the target conclusion back to supporting facts, can provide more accurate, efficient, and robust text-based logical deduction in large language models. To test this, the authors introduce LAMBADA, a hybrid architecture combining backward-chaining search control with four specialized, few-shot prompted language model modules: fact verification, rule selection, goal decomposition, and polarity agreement.
The approach was benchmarked using the PaLM 540B model across challenging deduction datasets—including ProofWriter, PrOntoQA, and the naturalistic ParaRules corpus—featuring reasoning depths up to five steps and three-way classification outcomes: proved, disproved, or unknown. The evaluations measured label accuracy, proof validity, computational query efficiency, and robustness to lexical and syntactic variations.
The results show that backward chaining substantially outperforms standard prompting and forward modular systems. On 5-hop reasoning under open-world assumptions, LAMBADA achieved a 44% relative accuracy improvement over standard chain-of-thought prompting and a 56% improvement over the forward-chaining Selection-Inference baseline. On the naturalistic ParaRules dataset, LAMBADA attained 86% accuracy compared to 60% for chain-of-thought. Furthermore, an audit of 5-hop proofs revealed that while standard chain-of-thought produced structurally valid proofs in only 28% of seemingly correct answers—frequently relying on hallucinations and spurious correlations—LAMBADA generated valid proofs in 94% of cases. LAMBADA was also significantly more computationally efficient, requiring up to 11.8 times fewer model inference queries than modular forward-chaining alternatives at maximum depth.
These findings demonstrate that separating high-level search planning from natural language translation resolves core reliability issues in automated reasoning. For business and technical operations, this modular backward-chaining design drastically lowers operational compute costs and mitigates compliance, safety, and hallucination risks associated with deploying language models in high-stakes reasoning pipelines.
Organizations developing complex reasoning systems should transition from unconstrained forward-prompting techniques to structured, goal-directed modular workflows. Future initiatives should focus on extending this architecture to open-domain environments where rules are not fully provided in advance, supporting batch processing for lower latency, and exploring fine-tuning strategies on smaller, more cost-effective language models using backward reasoning traces.
The current implementation remains limited to deductive reasoning tasks with explicitly provided, prompt-sized rule sets, and its recursive nature demands sequential model queries that prevent out-of-the-box batching. Nevertheless, the experimental results provide high confidence that goal-directed backward chaining provides a superior, robust framework for automated text-based reasoning.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Read this foundational account of few-shot Chain-of-Thought first: LAMBADA contrasts its forward reasoning setup with a backward-chaining alternative.
- Paper: Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations, Jaehun Jung et al. (2022). Its recursively generated explanations and logical constraints offer useful precedent for understanding LAMBADA’s decomposition of language-model reasoning into structured proof steps.
No sufficiently relevant recommendations were found.
