XAI for Transformers: Better Explanations through Conservative Propagation
Ameen AliThomas SchnakeOliver EberleGrégoire MontavonKlaus-Robert MüllerLior Wolf
Extends Layer-wise Relevance Propagation to Transformers by identifying and fixing conservation failures in attention heads and LayerNorm layers, producing more reliable feature attributions across text, vision, and graph benchmarks.
Modern artificial intelligence increasingly relies on transformer models across language processing, computer vision, and scientific data analysis. Because these models contain up to billions of parameters, their internal decision-making is opaque and difficult to verify. In sensitive settings like automated hiring or risk assessment, stakeholders require reliable explainable artificial intelligence (XAI) to verify that decisions are trustworthy, fair, and free of bias. However, existing interpretability tools often struggle with the unique internal mechanics of transformers.
The article evaluates why traditional attribution methods fail on transformers and demonstrates a theoretically sound, conservative propagation technique that significantly improves explanation accuracy across diverse datasets and architectures.
To address this challenge, the authors analyzed standard gradient-based interpretability through the framework of Layer-wise Relevance Propagation (LRP), specifically evaluating whether the axiom of conservation—the requirement that input importance scores sum to the model's total output score—holds across internal layers. After identifying where standard gradient calculations fail mathematically, they introduced modified propagation rules for attention heads and normalization layers. The method was benchmarked across nine datasets covering natural language sentiment and emotion detection, graph-based digit recognition, and molecular property prediction, using input perturbation tests that measure output stability when removing or adding key features.
The investigation produced several key findings. First, standard gradient-based attribution severely violates conservation in transformers, primarily due to attention gating and variance rescaling in layer normalization; on image-based graph benchmarks, naive gradient attributions were almost anticorrelated with model outputs. Second, the proposed method, which treats these gating and normalization scaling factors as locally constant during backpropagation, restored conservation and achieved the highest explanation performance across all tested datasets. For example, on the Stanford Sentiment Treebank dataset, the area under the error curve during feature removal dropped from 2.10 with naive gradient attribution to 1.56 with the combined approach. Third, the technique improved computational efficiency, running in 0.012 seconds per sample on benchmark language data compared to 0.017 seconds for standard gradient methods and 0.024 seconds for the original model prediction. Fourth, applied to bias auditing in sentiment analysis models, the method successfully surfaced specific entity biases (such as disparate sentiment shifts tied to particular names) without requiring synthetic or out-of-distribution test inputs.
These findings indicate that organizations relying on standard gradient or attention visualization tools may be making governance and safety decisions based on flawed or misleading explanations. By ensuring mathematical conservation throughout the network, the proposed approach lowers operational risk and compliance overhead in high-stakes deployments, providing leaders with high-fidelity insights into why a model produces a specific prediction without increasing compute costs.
Organizations deploying transformer models should adopt conservative propagation rules in place of raw attention or standard gradient attribution when auditing model decisions and monitoring for demographic bias. Implementation is straightforward, as the adjustment only requires strategically detaching specific internal terms in the code during the backward attribution pass rather than retraining the underlying models.
While confidence in these findings is reinforced by consistent mathematical derivations and comprehensive empirical validation across language, vision, and molecular tasks, minor conservation gaps remain due to unhandled bias terms in linear layers. Stakeholders should consider this method a robust tool for feature-level attribution and bias detection, while continuing to pair it with broader system-level validation practices.
- Paper: Methods for interpreting and understanding deep neural networks, Grégoire Montavon et al. (2018). Its account of LRP and deep Taylor decomposition supplies the propagation framework that this paper extends to Transformer-specific layers.
- Paper: Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models, Wojciech Samek et al. (2017). Its comparison of gradient sensitivity with relevance propagation clarifies the attribution methods and reliability problem that motivate the Transformer adaptation.
- Paper: Evaluating the Visualization of What a Deep Neural Network Has Learned, Wojciech Samek et al. (2015). Its perturbation-based evaluation of LRP against gradient explanations provides useful grounding for the faithfulness criteria applied to Transformer explanations.
- Paper: Value bounds and Convergence Analysis for Averages of LRP attributions, Alexander Binder et al. (2025). It carries LRP attribution into later convergence and stability analysis, including tests on a Swin Transformer, extending the source’s concern with reliable Transformer explanations.
