Faithfulness Tests for Natural Language Explanations
Pepa AtanasovaOana-Maria CamburuChristina LiomaThomas LukasiewiczJakob Grue SimonsenIsabelle Augenstein
Proposes two novel evaluation tests using counterfactual input editing and input reconstruction to diagnose whether natural language explanations genuinely reflect the decision-making processes of neural models.
Artificial intelligence systems increasingly generate natural language explanations to justify their predictions in human-readable free text. However, if an explanation fails to reflect the model's actual internal decision-making process—a quality known as faithfulness—it can dangerously mislead users, obscure model flaws, and create unjustified trust in high-stakes deployments.
The article establishes a framework to evaluate whether free-text explanations accurately describe why a system arrived at a given decision. Specifically, it introduces two distinct diagnostic tests to quantitatively assess explanation faithfulness across multiple model architectures.
To conduct this evaluation, the authors tested four model configurations based on the T5 language model across three established reasoning datasets: e-SNLI, CoS-E, and ComVE. The first evaluation method, a counterfactual test, inserts words into the input to deliberately flip the model's prediction and checks whether the resulting explanation mentions the critical inserted words. The second method, an input reconstruction test, extracts the stated reasons from an explanation to build a new input prompt and tests whether the model still produces the original prediction based solely on those stated reasons.
The experiments revealed substantial rates of unfaithful explanations across all evaluated systems. In the counterfactual test, the combination of editing techniques uncovered unfaithfulness in 31% to 59% of test cases across the benchmarks. In the input reconstruction test, the stated reasons failed to justify the original prediction in 8% to 14% of e-SNLI cases and up to 40% of ComVE cases. Notably, no single model architecture—whether generating explanations jointly with predictions or conditioned sequentially—consistently produced faithful explanations.
These findings indicate that current language models frequently generate plausible-sounding justifications that do not correspond to their real computational reasoning. For decision-makers, this introduces operational and safety risks: deploying these models with the assumption that their written justifications explain their behavior could mask critical biases or systematic errors, undermining safety, compliance, and governance efforts.
Organizations should treat generated explanations as helpful text rather than certified evidence of model behavior. Before relying on natural language explanations in sensitive applications, technical teams should implement automated diagnostic tests to audit faithfulness. Future work must develop broader test suites and machine-learning-driven input reconstruction tools to evaluate a wider variety of tasks and datasets.
Readers should note that the diagnostic tests provide a lower bound on unfaithfulness rather than a comprehensive guarantee. Passing these tests does not prove that a model is completely faithful, as models may still overlook other contributing factors. Additionally, the input reconstruction test currently relies on dataset-specific heuristic rules that cannot yet be applied across every type of reasoning task.
- Paper: Attention is not Explanation, Sarthak Jain et al. (2019). This seminal paper exposes the fundamental unfaithfulness of attention-based explanations in NLP, providing the foundational motivation for developing rigorous faithfulness tests for language model explanations.
- Paper: Sanity Checks for Saliency Maps, Julius Adebayo et al. (2018). It introduces sanity checks and input/parameter randomization to evaluate explanation fidelity, establishing the testing paradigms that the source adapts for natural language explanations.
- Paper: Counterfactual Explanations without Opening the Black Box: Automated Decisions and the GDPR, Sandra Wachter et al. (2017). This work formulates the principles of counterfactual explanations for black-box systems, directly underpinning the source's counterfactual input editing methodology.
- Paper: Do Feature Attribution Methods Correctly Attribute Features?, Yilun Zhou et al. (2022). It formalizes ground-truth faithfulness evaluation by perturbing and modifying datasets, laying the ground for testing whether explanations genuinely reflect internal decision logic.
- Paper: OpenXAI: Towards a Transparent Evaluation of Model Explanations, Chirag Agarwal et al. (2022). This paper establishes quantitative evaluation metrics and benchmarks for explanation faithfulness, providing essential background for rigorous interpretability assessment.
- Paper: Interpretable Explanations of Black Boxes by Meaningful Perturbation, Ruth Fong et al. (2017). It introduces input perturbation and deletion games to assess explanation fidelity, directly inspiring the source's input reconstruction and editing tests.
- Paper: Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations, Jaehun Jung et al. (2022). This work explores generating recursive natural language explanations for language model inference, providing the model generation setting that faithfulness tests aim to evaluate.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It demonstrates how natural language chain-of-thought explanations systematically rationalize biased inputs rather than reflecting true model reasoning, extending the critical study of natural language explanation faithfulness.
- Paper: RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations, Jing Huang et al. (2024). This work provides an empirical benchmark to evaluate whether interpretability methods accurately disentangle underlying representations in language models, building upon faithfulness evaluation frameworks.
- Paper: Towards Best Practices of Activation Patching in Language Models: Metrics and Methods, Fred Zhang et al. (2024). It investigates causal activation patching to localize internal reasoning components, continuing the pursuit of verified, faithful mechanistic explanations in language models.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). This paper advances the evaluation of multi-step reasoning by measuring intermediate reasoning quality and validity rather than final accuracy alone.
- Paper: Chain-of-Thought Reasoning Without Prompting, Xuezhi Wang et al. (2024). It investigates whether intrinsic reasoning paths in language models can be decoded faithfully without explicit prompt-induced rationalization.
