Faithfulness Tests for Natural Language Explanations

Pepa AtanasovaOana-Maria CamburuChristina LiomaThomas LukasiewiczJakob Grue SimonsenIsabelle Augenstein

article2023ACL141 citations

Proposes two novel evaluation tests using counterfactual input editing and input reconstruction to diagnose whether natural language explanations genuinely reflect the decision-making processes of neural models.

Listen

Artificial intelligence systems increasingly generate natural language explanations to justify their predictions in human-readable free text. However, if an explanation fails to reflect the model's actual internal decision-making process—a quality known as faithfulness—it can dangerously mislead users, obscure model flaws, and create unjustified trust in high-stakes deployments.

The article establishes a framework to evaluate whether free-text explanations accurately describe why a system arrived at a given decision. Specifically, it introduces two distinct diagnostic tests to quantitatively assess explanation faithfulness across multiple model architectures.

To conduct this evaluation, the authors tested four model configurations based on the T5 language model across three established reasoning datasets: e-SNLI, CoS-E, and ComVE. The first evaluation method, a counterfactual test, inserts words into the input to deliberately flip the model's prediction and checks whether the resulting explanation mentions the critical inserted words. The second method, an input reconstruction test, extracts the stated reasons from an explanation to build a new input prompt and tests whether the model still produces the original prediction based solely on those stated reasons.

The experiments revealed substantial rates of unfaithful explanations across all evaluated systems. In the counterfactual test, the combination of editing techniques uncovered unfaithfulness in 31% to 59% of test cases across the benchmarks. In the input reconstruction test, the stated reasons failed to justify the original prediction in 8% to 14% of e-SNLI cases and up to 40% of ComVE cases. Notably, no single model architecture—whether generating explanations jointly with predictions or conditioned sequentially—consistently produced faithful explanations.

These findings indicate that current language models frequently generate plausible-sounding justifications that do not correspond to their real computational reasoning. For decision-makers, this introduces operational and safety risks: deploying these models with the assumption that their written justifications explain their behavior could mask critical biases or systematic errors, undermining safety, compliance, and governance efforts.

Organizations should treat generated explanations as helpful text rather than certified evidence of model behavior. Before relying on natural language explanations in sensitive applications, technical teams should implement automated diagnostic tests to audit faithfulness. Future work must develop broader test suites and machine-learning-driven input reconstruction tools to evaluate a wider variety of tasks and datasets.

Readers should note that the diagnostic tests provide a lower bound on unfaithfulness rather than a comprehensive guarantee. Passing these tests does not prove that a model is completely faithful, as models may still overlook other contributing factors. Additionally, the input reconstruction test currently relies on dataset-specific heuristic rules that cannot yet be applied across every type of reasoning task.

arXiv: 2305.18029
Cover for Faithfulness Tests for Natural Language Explanations

Abstract

Explanations of neural models aim to reveal a model's decision-making process for its predictions. However, recent work shows that current methods giving explanations such as saliency maps or counterfactuals can be misleading, as they are prone to present reasons that are unfaithful to the model's inner workings. This work explores the challenging question of evaluating the faithfulness of natural language explanations (NLEs). To this end, we present two tests. First, we propose a counterfactual input editor for inserting reasons that lead to counterfactual predictions but are not reflected by the NLEs. Second, we reconstruct inputs from the reasons stated in the generated NLEs and check how often they lead to the same predictions. Our tests can evaluate emerging NLE models, proving a fundamental tool in the development of faithful NLEs.

Table of Contents

  • 1 Introduction
  • 2 The Faithfulness Tests
  • 3 Experiments
  • 3.1 Results
  • 4 Related Work
  • 5 Summary and Outlook
  • Limitations
  • Acknowledgements
  • References
  • A More Examples of Unfaithful NLEs
  • A.1 Model Performance

Knowls

  1. Knowl 1 — Counterfactual faithfulness test

    model/method

    The counterfactual test checks whether a generated natural-language explanation (NLE) mentions text that causes the model to change its prediction. Let xix_i be an input token sequence, let f(xi)=(e^i,y^i)f(x_i)=(\hat e_i,\hat y_i) be a model’s generated NLE and predicted label, and let yiC∈Ly_i^C\in L be a target label different from y^i\hat y_i. An intervention inserts a contiguous token sequence WW at position kk, producing xi′x_i'. The intervention is a counterfactual when the model predicts the target label on the edited input but not on the original input:

    xi′=(xi,1,…,xi,k,W,xi,k+1,…,xi,∣xi∣),f(xi′)p=yiC≠y^i=f(xi)p.x_i'=(x_{i,1},\ldots,x_{i,k},W,x_{i,k+1},\ldots,x_{i,|x_i|}),\qquad f(x_i')_p=y_i^C\ne \hat y_i=f(x_i)_p.

    Here ∣xi∣|x_i| is the original input length, f(⋅)pf(\cdot)_p denotes the prediction component of ff, and e^i′\hat e_i' is the NLE generated for xi′x_i'. If no token from the inserted sequence occurs in the counterfactual NLE, W∩se^i′=∅W\cap^s\hat e_i'=\varnothing, the NLE is marked unfaithful to the detected counterfactual reason. The superscript ss denotes syntactic token overlap; the authors manually checked a subset for paraphrases because the automatic test does not detect semantic equivalents. The reported counterfactual-unfaithfulness score is the percentage of test instances for which such an intervention is found and its inserted text is absent from the associated NLE.

  2. Knowl 2 — Neural counterfactual input editor

    algorithm

    The paper’s counterfactual editor hh is a T5-base neural text-to-text model that proposes contiguous token insertions capable of changing a task model’s prediction. During training, between 1 and 3 consecutive tokens of an input are randomly masked with one mask token. The task model’s predicted label is supplied to hh as the target label, and the masked tokens supervise generation of the inserted text WW using cross-entropy loss. During inference, hh receives a target label different from the original prediction and proposes insertions at four randomly selected positions, generating four candidates per position. Each candidate is inserted into the input and retained only if the task model predicts the requested target label.

    Input: Original input xx, task model ff, original prediction y^\hat y, trained editor hh, target-label set LL, n2=4n_2=4 positions, and n3=4n_3=4 candidates per position
    Output: Counterfactual insertion set CC
    Set CC to the empty set
    For each target label yCy^C in LL where yC≠y^y^C\ne\hat y
        Sample n2n_2 insertion positions in xx
        For each sampled position
            Use hh conditioned on xx and yCy^C to generate n3n_3 contiguous token sequences WW
            For each generated sequence WW
                Insert WW at the position to form x′x'
                If f(x′)f(x') predicts yCy^C
                    Add (x′,W,yC)(x',W,y^C) to CC
    Return $C

    The editor is intended to find model-sensitive insertions, not to prove that the inserted words are the model’s sole causal basis. The authors hypothesize that confounding effects are relatively rare for insertion interventions, but do not establish this assumption.

  3. Knowl 3 — Input reconstruction faithfulness test

    model/method

    The input reconstruction test asks whether the reasons expressed in a generated NLE are sufficient for the task model to retain its original prediction. Let xix_i be an input, e^i\hat e_i its generated NLE, f(xi)pf(x_i)_p the model’s original prediction, and RR a task-specific function that constructs a new input from xix_i and e^i\hat e_i. Define ri=R(xi,e^i)r_i=R(x_i,\hat e_i). The NLE is judged unfaithful when the reconstructed input receives a different prediction:

    ri=R(xi,e^i),f(ri)p≠f(xi)p  ⟹  e^i is unfaithful.r_i=R(x_i,\hat e_i),\qquad f(r_i)_p\ne f(x_i)_p\;\Longrightarrow\;\hat e_i\text{ is unfaithful}.

    The test measures sufficiency of the NLE’s stated reasons rather than plausibility or similarity to a human explanation. Its total-unfaithfulness score is the percentage of all test instances for which a valid reconstructed input is formed and the model’s prediction changes.

  4. Knowl 4 — Task-specific reconstruction agents

    model/method

    Because free-text NLEs do not have a direct token-to-input mapping, the paper constructs task-dependent reconstruction functions rather than one universal extractor. For e-SNLI, many explanations follow templates; the available templates cover 97.4% of the training explanations. The reconstruction agent extracts the template spans ⟨X⟩\langle X\rangle and ⟨Y⟩\langle Y\rangle and uses them as the reconstructed premise and hypothesis, respectively, retaining only extracted sentences containing at least one subject and one verb. If the original NLE is sufficient, the reconstructed premise–hypothesis pair should receive the same entailment, contradiction, or neutral prediction as the original pair.

    For ComVE, whose task is to choose which of two sentences violates common sense, the reconstruction replaces the sentence judged correct in the original instance with the generated NLE. A faithful NLE should preserve the model’s choice of the commonsense-violating sentence. The authors could not construct a rule-based reconstruction function for CoS-E because its question-answer NLEs do not have a suitable regular structure; consequently, the reconstruction test is not applied to CoS-E.

  5. Knowl 5 — Four NLE model configurations evaluated

    model/method

    The experiments compare four configurations formed from two training regimes and two explanation-conditioning regimes. Multi-task (MT) models jointly predict the task label and generate an NLE; single-task (ST) models use separate prediction and explanation models. Reasoning (Re) models generate the NLE without conditioning it on the predicted label, whereas rationalizing (Ra) models condition NLE generation on a label.

    • MT-Re: one joint model predicts the label and generates an explanation without label conditioning.
    • MT-Ra: one joint model predicts the label and generates an explanation conditioned on that predicted label.
    • ST-Re: a separate explanation model generates the NLE from the input, and a separate prediction model predicts the label from the input and explanation setup.
    • ST-Ra: the explanation model generates one candidate NLE for each possible label. The prediction model evaluates the input with each candidate explanation and selects the label with the highest resulting probability.

    All four configurations are evaluated with the same two faithfulness tests, so differences reflect model organization and label conditioning rather than a change in the diagnostic definition.

  6. Knowl 6 — Experimental protocol and datasets

    experimental setup

    The study evaluates e-SNLI, CoS-E, and ComVE. e-SNLI requires classifying a premise–hypothesis pair as entailment, contradiction, or neutral and provides an NLE; CoS-E requires selecting one of three answers to a commonsense question; ComVE requires selecting which of two sentences violates common sense.

    The task models and counterfactual editors use T5-base. Both are trained for 20 epochs with Adam and learning rate 10−410^{-4}; validation performance is checked after every epoch and the checkpoint with the highest relevant success rate is selected. Editor search uses four random insertion positions and four candidates per position, with these hyperparameters selected by validation-set grid search. The random counterfactual baseline inserts a randomly selected adjective before a noun or adverb before a verb, using the same four-position/four-candidate search; candidates are sampled from WordNet and nouns and verbs are identified with spaCy.

    For manual checking of the counterfactual metric, an author annotated the first 100 test instances for each of the four model configurations, totaling 800 instances. No paraphrase cases were found in this sample, so the authors treated the automatic syntactic-overlap metric as reliable for their reported experiments.

  7. Knowl 7 — Counterfactual-test results across models and datasets

    data/table

    The counterfactual results on the paper’s printed page 285 compare the random insertion baseline, the learned editor, and their union. “% Counter” is the percentage of all instances for which the intervention changes the prediction to the target label. “% Counter Unfaith” is the percentage of those successful counterfactual instances whose inserted text is absent from the NLE. “% Total Unfaith” is the percentage of all test instances satisfying both conditions. The table demonstrates that the learned editor generally finds more prediction-changing interventions, while the random baseline often has a higher conditional unfaithfulness rate.

    Dataset Model % Counter % Counter Unfaith % Total Unfaith

    e-SNLI
    MT-Re-Rand 38.85 60.39 23.46
    MT-Re-Edit 56.70 46.12 26.15
    MT-Re-Rand+Edit 64.98 53.29 34.63
    ST-Re-Rand 37.14 54.26 20.15
    ST-Re-Edit 49.64 52.74 26.18
    ST-Re-Rand+Edit 61.15 58.27 35.63
    MT-Ra-Rand 37.17 54.93 20.42
    MT-Ra-Edit 55.04 41.34 22.75
    MT-Ra-Rand+Edit 63.84 48.63 31.05
    ST-Ra-Rand 35.21 57.82 20.36
    ST-Ra-Edit 60.00 45.66 27.39
    ST-Ra-Rand+Edit 67.31 55.03 37.04

    CoS-E
    MT-Re-Rand 44.89 83.18 37.34
    MT-Re-Edit 50.00 77.23 38.62
    MT-Re-Rand+Edit 59.89 85.26 51.06
    ST-Re-Rand 52.34 79.47 41.60
    ST-Re-Edit 53.83 86.17 46.38
    ST-Re-Rand+Edit 67.45 87.54 59.04
    MT-Ra-Rand 39.26 84.01 32.98
    MT-Ra-Edit 50.00 78.72 39.36
    MT-Ra-Rand+Edit 56.81 85.58 48.62
    ST-Ra-Rand 46.70 75.85 35.43
    ST-Ra-Edit 52.02 75.05 39.04
    ST-Ra-Rand+Edit 63.62 81.77 52.02

    ComVE
    MT-Re-Rand 35.60 83.43 29.70
    MT-Re-Edit 50.90 70.53 35.90
    MT-Re-Rand+Edit 61.10 78.89 48.20
    ST-Re-Rand 41.90 74.22 31.10
    ST-Re-Edit 48.40 76.45 37.00
    ST-Re-Rand+Edit 62.90 77.42 48.70
    MT-Ra-Rand 33.70 75.67 25.50
    MT-Ra-Edit 47.20 66.53 31.40
    MT-Ra-Rand+Edit 58.10 73.84 42.90
    ST-Ra-Rand 36.30 80.17 29.10
    ST-Ra-Edit 49.50 79.80 39.50
    ST-Ra-Rand+Edit 61.80 83.98 51.90
  8. Knowl 8 — Input-reconstruction results

    data/table

    The input-reconstruction results on the paper’s printed page 285 report the fraction of instances for which a reconstructed input could be formed and the fraction of all test instances identified as unfaithful. Reconstruction was possible for 39.49%–44.87% of e-SNLI instances and for 100% of ComVE instances. Among all instances, the test identified 7.7%–9.7% unfaithful NLEs on e-SNLI and 22.7%–40.3% on ComVE, showing that the two datasets expose different failure patterns.

    Dataset Model % Reconst % Total Unfaith
    e-SNLI MT-Re 39.49 7.7
    ST-Re 39.99 9.7
    MT-Ra 44.87 7.8
    ST-Ra 43.32 9.3
    ComVE MT-Re 100 36.9
    ST-Re 100 22.7
    MT-Ra 100 40.3
    ST-Ra 100 28.5
  9. Knowl 9 — Observed failure patterns across model types

    empirical result

    All four NLE model configurations produce substantial counterfactual-test failures across all three datasets. Using the union of random and learned-editor interventions, the reported total counterfactual-unfaithfulness rates range from 37.04% for MT-Ra on e-SNLI to 59.04% for ST-Re on CoS-E. The learned editor alone changes predictions on as many as 56.70% of e-SNLI instances and finds an absent-in-the-NLE counterfactual reason on as many as 46.38% of CoS-E instances.

    The random baseline usually has a larger conditional unfaithfulness rate than the learned editor, which the authors conjecture is because randomly selected words are rare or atypical for the dataset. The editor nevertheless discovers more prediction-changing interventions and therefore more failures overall. Model configurations show similar faithfulness levels with no consistent ranking: neither multi-task versus single-task training nor reasoning versus rationalizing conditioning guarantees greater faithfulness. Reasoning models tend to be less faithful than rationalizing models in most comparisons, but this tendency is not universal. The differing patterns between the counterfactual and reconstruction tests support using multiple diagnostics rather than treating one score as a complete faithfulness measure.

  10. Knowl 10 — Scope and limitations of the faithfulness tests

    limitation

    The two tests are diagnostic probes, not comprehensive measures of NLE faithfulness. A model that passes both may still generate unfaithful explanations. The counterfactual test only checks whether an NLE mentions reasons associated with discovered prediction-changing insertions; an NLE may omit those insertions yet faithfully describe other relevant factors. Its conclusions also depend on the editor’s ability to find interventions, and inserted text can be semantically incoherent or change the model through a confounding shift of attention rather than through the inserted words themselves.

    The reconstruction test depends on task-specific, hand-crafted reconstruction rules. Such rules were available for e-SNLI and ComVE but not for CoS-E, so the test cannot compare all datasets. The paper proposes learned reconstruction agents trained from a small number of annotated examples as a future way to extend the test to tasks such as CoS-E.

Coverage note — Appendix examples and the auxiliary task-accuracy/BLEU table were omitted because they illustrate the tests or report model quality rather than adding a distinct load-bearing contribution to the faithfulness methodology.

References

  1. 1.Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018. Sanity checks for saliency maps. Advances in Neural Information Processing Systems, 31.
  2. 2.Julius Adebayo, Michael Muelly, Harold Abelson, and Been Kim. 2022. Post hoc explanations may be ineffective for detecting unknown spurious correlation. In International Conference on Learning Representations.
  3. 3.Christopher Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller, and Pan Kessel. 2020. Fairwashing explanations with off-manifold detergent. In International Conference on Machine Learning, pages 314–323. PMLR.
  4. 4.Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020a. A diagnostic study of explainability techniques for text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3256–3274, Online. Association for Computational Linguistics.
  5. 5.Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020b. Generating fact checking explanations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7352–7364, Online. Association for Computational Linguistics.
  6. 6.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  7. 7.Oana-Maria Camburu, Eleonora Giunchiglia, Jakob Foerster, Thomas Lukasiewicz, and Phil Blunsom. 2019. Can I Trust the Explainer? Verifying Post-hoc Explanatory Methods. In NeurIPS 2019 Workshop Safety and Robustness in Decision Making.
  8. 8.Oana-Maria Camburu, Eleonora Giunchiglia, Jakob Foerster, Thomas Lukasiewicz, and Phil Blunsom. 2021. The struggles of feature-based explanations: Shapley values vs. minimal sufficient subsets. In AAAI 2021 Workshop on Explainable Agency in Artificial Intelligence.
  9. 9.Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018. e-SNLI: Natural Language Inference with Natural Language Explanations. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems 31, pages 9539–9549. Curran Associates, Inc.
  10. 10.Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini, Thomas Lukasiewicz, and Phil Blunsom. 2020. Make up your mind! adversarial generation of inconsistent natural language explanations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4157–4165, Online. Association for Computational Linguistics.
  11. 11.Aaron Chan, Shaoliang Nie, Liang Tan, Xiaochang Peng, Hamed Firooz, Maziar Sanjabi, and Xiang Ren. 2022a. Frame: Evaluating simulatability metrics for free-text rationales. arXiv preprint arXiv:2207.00779.
  12. 12.Chun Sik Chan, Huanqi Kong, and Liang Guanqing. 2022b. A comparative study of faithfulness metrics for model interpretability methods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5029–5038, Dublin, Ireland. Association for Computational Linguistics.
  13. 13.Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020. ERASER: A benchmark to evaluate rationalized NLP models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4443–4458, Online. Association for Computational Linguistics.
  14. 14.Ann-Kathrin Dombrowski, Maximillian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel. 2019. Explanations can be manipulated and geometry is to blame. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  15. 15.Christiane Fellbaum. 2010. Wordnet. In Theory and Applications of Ontology: Computer Applications, pages 231–243. Springer.
  16. 16.Riccardo Guidotti. 2022. Counterfactual explanations and how to find them: literature review and benchmarking. Data Mining and Knowledge Discovery, pages 1–55.
  17. 17.Leif Hancox-Li. 2020. Robustness in machine learning explanations: does it matter? In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pages 640–647.
  18. 18.Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal. 2020. Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language? In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4351–4367, Online. Association for Computational Linguistics.
  19. 19.Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020. spaCy: Industrial-strength Natural Language Processing in Python.
  20. 20.Cheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Kumar Ravikumar, Seungyeon Kim, Sanjiv Kumar, and Cho-Jui Hsieh. 2021. Evaluations and methods for explanation through robustness analysis. In International Conference on Learning Representations.
  21. 21.Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021. Contrastive explanations for model interpretability. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1597–1611, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  22. 22.Shailza Jolly, Pepa Atanasova, and Isabelle Augenstein. 2022. Generating Fluent Fact Checking Explanations with Unsupervised Post-Editing. Information, 13(10).
  23. 23.Maxime Kayser, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde, Virginie Do, Zeynep Akata, and Thomas Lukasiewicz. 2021. e-ViL: A dataset and benchmark for natural language explanations in vision-language tasks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1244–1254.
  24. 24.Maxime Kayser, Cornelius Emde, Oana-Maria Camburu, Guy Parsons, Bartlomiej Papiez, and Thomas Lukasiewicz. 2022. Explaining chest x-ray pathologies in natural language. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2022, pages 701–713, Cham. Springer Nature Switzerland.
  25. 25.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  26. 26.Andreas Madsen, Nicholas Meade, Vaibhav Adlakha, and Siva Reddy. 2022. Evaluating the Faithfulness of Importance Measures in NLP by Recursively Masking Allegedly Important Tokens and Retraining. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 1731–1751, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  27. 27.Bodhisattwa Prasad Majumder, Oana Camburu, Thomas Lukasiewicz, and Julian Mcauley. 2022. Knowledge-grounded self-rationalization via extractive and natural language explanations. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 14786–14801. PMLR.
  28. 28.Ana Marasovic, Iz Beltagy, Doug Downey, and Matthew E. Peters. 2022. Few-shot self-rationalization with natural language prompts. Findings of NAACL.
  29. 29.Tim Miller. 2019. Explanation in artificial intelligence: Insights from the social sciences. Artificial intelligence, 267:1–38.
  30. 30.Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020. WT5?! training text-to-text models to explain their predictions.
  31. 31.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  32. 32.Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019. Explain yourself! leveraging language models for commonsense reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4932–4942, Florence, Italy. Association for Computational Linguistics.
  33. 33.Alexis Ross, Ana Marasovic, and Matthew E. Peters. 2021. Explaining NLP models via minimal contrastive editing (MiCE). In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3840–3852, Online. Association for Computational Linguistics.
  34. 34.Dylan Slack, Anna Hilgard, Himabindu Lakkaraju, and Sameer Singh. 2021. Counterfactual explanations can be manipulated. Advances in Neural Information Processing Systems, 34:62–75.
  35. 35.Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020. Fooling lime and shap: Adversarial attacks on post hoc explanation methods. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, AIES ’20, page 180–186, New York, NY, USA. Association for Computing Machinery.
  36. 36.Jiao Sun, Swabha Swayamdipta, Jonathan May, and Xuezhe Ma. 2022. Investigating the benefits of free-form rationales. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 5867–5882, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  37. 37.Cunxiang Wang, Shuailong Liang, Yili Jin, Yilong Wang, Xiaodan Zhu, and Yue Zhang. 2020. SemEval-2020 task 4: Commonsense validation and explanation. In Proceedings of the Fourteenth Workshop on Semantic Evaluation, pages 307–321, Barcelona (online). International Committee for Computational Linguistics.
  38. 38.Sarah Wiegreffe and Ana Marasovic. 2021. Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  39. 39.Sarah Wiegreffe, Ana Marasovic, and Noah A. Smith. 2021. Measuring association between labels and free-text rationales. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10266–10284, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  40. 40.Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2021. Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6707–6723, Online. Association for Computational Linguistics.
  41. 41.Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, and Kai-Wei Chang. 2022. On the Sensitivity and Stability of Model Interpretations in NLP. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2631–2647, Dublin, Ireland. Association for Computational Linguistics.
  42. 42.Yordan Yordanov, Vid Kocijan, Thomas Lukasiewicz, and Oana-Maria Camburu. 2022. Few-Shot Out-of-Domain Transfer of Natural Language Explanations. In Proceedings of the Findings of the Conference on Empirical Methods in Natural Language Processing (EMNLP).
  43. 43.Mo Yu, Shiyu Chang, Yang Zhang, and Tommi Jaakkola. 2019. Rethinking cooperative rationalization: Introspective extraction and complement control. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4094–4103, Hong Kong, China. Association for Computational Linguistics.

Citation

MLA
Atanasova, P., et al. “Faithfulness Tests for Natural Language Explanations”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2023, pp. 283–94, https://doi.org/10.18653/v1/2023.acl-short.25.
APA
Atanasova, P., Camburu, O.-M., Lioma, C., Lukasiewicz, T., Simonsen, J. G., & Augenstein, I. (2023). Faithfulness Tests for Natural Language Explanations. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 283–294. https://doi.org/10.18653/v1/2023.acl-short.25
Chicago
Atanasova, P., O.-M. Camburu, C. Lioma, T. Lukasiewicz, J. G. Simonsen, and I. Augenstein. 2023. “Faithfulness Tests for Natural Language Explanations”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 283–94. https://doi.org/10.18653/v1/2023.acl-short.25.
Harvard
Atanasova, P. et al. (2023) “Faithfulness Tests for Natural Language Explanations”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp. 283–294. Available at: https://doi.org/10.18653/v1/2023.acl-short.25.
Vancouver
1. Atanasova P, Camburu O-M, Lioma C, Lukasiewicz T, Simonsen JG, Augenstein I (2023) Faithfulness Tests for Natural Language Explanations. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp 283–294

BibTeX

@inproceedings{atanasova-etal-2023-faithfulness,
    title = "Faithfulness Tests for Natural Language Explanations",
    author = "Atanasova, Pepa  and
      Camburu, Oana-Maria  and
      Lioma, Christina  and
      Lukasiewicz, Thomas  and
      Simonsen, Jakob Grue  and
      Augenstein, Isabelle",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-short.25/",
    doi = "10.18653/v1/2023.acl-short.25",
    pages = "283--294"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/