Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning
Oyvind TafjordBhavana Dalvi MishraPeter Clark
Proposes a question-answering framework that couples backward-chaining entailment generation with self-query verification to construct multi-step reasoning trees directly reflecting a language model's internal beliefs.
Large pretrained language models excel at answering complex questions, yet they typically function as black boxes. When these models provide explanations or rationale chains, those explanations are frequently unfaithful, meaning the final answer does not strictly follow from the reasoning, or untruthful, meaning the model generates statements it does not internally verify as true. This lack of transparency limits organizational trust, complicates compliance and safety audits, and makes it difficult to diagnose why a model failed. The article addresses this challenge by designing an architecture that answers questions through systematic, verifiable reasoning grounded in the model's own internal beliefs.
The main objective of the article is to demonstrate Entailer, a question-answering framework that produces faithful, multistep entailment proofs by combining backward-chaining generation with self-verification. The approach utilizes an 11-billion parameter language model trained across three core tasks: generating premises that imply a given answer hypothesis, scoring whether the model believes those individual premises are true, and validating whether the logical entailment step is sound. To find the strongest answer, the system searches backward from candidate answers, generates potential premises, and filters them out if self-querying shows low belief scores or flawed logic. The model was trained on the EntailmentBank science dataset augmented by thousands of crowdsourced positive and negative examples, and was subsequently evaluated zero-shot without task-specific fine-tuning on multiple-choice reasoning benchmarks.
The findings show that Entailer generates structured reasoning proofs while preserving strong baseline accuracy. On benchmark evaluations, Entailer achieved a question-answering accuracy of roughly 75%, matching the accuracy of direct, unreasoned answers. In human evaluations, human judges determined that over 70% of Entailer's generated reasoning chains clearly proved the conclusion from the premises, compared to only 34% for explanations from a leading alternative question-answering model. Additionally, annotators confirmed that approximately 90% of the self-verified premises generated by Entailer were factually correct, preferring its structured explanations over the baseline by a margin of 57% to 23%. An analysis of errors revealed that 47% stemmed from reasoning flaws such as near-tautologies, 33% from incorrect factual beliefs, and 20% from dataset ambiguities.
These results demonstrate that language models can expose their latent knowledge as logical proof trees without incurring an accuracy penalty. For operational workflows, this capability enhances transparency and enables targeted debugging: when an answer is wrong, operators can inspect the reasoning tree to identify precisely which belief or deduction failed. This structured visibility lays the foundation for interactive and teachable artificial intelligence systems, where human feedback can correct an isolated erroneous premise rather than requiring opaque, expensive model retraining.
Stakeholders should treat Entailer as an effective framework for high-stakes domains that demand explainable reasoning, such as technical diagnostics or compliance support. Next steps should focus on implementing user-feedback loops where human corrections are stored in a retrieval memory to dynamically override flawed beliefs. However, leaders should note key operational limitations: Entailer is computationally intensive, requiring up to 360 seconds per question due to recursive multi-step search, and it can occasionally produce circular reasoning or verify contradictory statements. Initial deployments should therefore target asynchronous diagnostic tasks rather than real-time consumer applications until faster search mechanisms are implemented.
- Paper: Language Models (Mostly) Know What They Know, Saurav Kadavath et al. (2022). Read this account of how language models assess the correctness of their own answers first; it grounds Entailer’s use of self-reported beliefs to screen premises.
- Paper: LAMBADA: Backward Chaining for Automated Reasoning in Natural Language, Mehran Kazemi et al. (2023). LAMBADA carries backward-chaining language-model reasoning forward with modular verification and explicit handling of conclusions that are unproved or unknown.
- Paper: ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness, Archiki Prasad et al. (2023). ReCEval extends proof-chain work by evaluating each reasoning step for logical correctness and usefulness, including on EntailmentBank.
