GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers
Ali ModarressiMohsen FayyazYadollah YaghoobzadehMohammad Taher Pilehvar
Proposes GlobEnc, a layer-aggregation framework that measures global token attribution by accounting for full Transformer encoder components, including layer normalizations and residual connections, to achieve higher correlation with gradient-based saliency scores than attention-only baselines.
Modern natural language processing relies heavily on Transformer models, yet interpreting how these complex architectures arrive at specific decisions remains an active challenge. Standard interpretability techniques often rely on raw attention weights or focus exclusively on the self-attention sub-layer, which produces misleading explanations. While gradient-based methods provide more reliable importance scores, they are computationally intensive and slow. The article introduces and evaluates GlobEnc, a novel method designed to quantify how much each input word contributes to a model's final output by incorporating almost all components within the entire encoder block across all layers.
The authors evaluated GlobEnc using fine-tuned BERT-base, BERT-large, and ELECTRA-base models across three benchmark text classification tasks: sentiment analysis (SST2), natural language inference (MNLI), and hate speech detection (HateXplain). The approach accounts for vector magnitudes (norms), residual connections, and both layer normalizations within the encoder layer, then aggregates these layer-level contributions across the entire network using an attention rollout technique. The resulting global attribution scores were benchmarked against gradient-based saliency scores to measure explanation faithfulness.
The analysis revealed that GlobEnc consistently outperforms existing attention-based attribution methods, achieving high Spearman rank correlations with gradient saliency across all datasets (reaching 0.72 to 0.78 on BERT-base and up to 0.83 on BERT-large), while raw attention weights showed negligible or negative correlations (ranging from -0.11 to 0.12 on BERT-base). The authors found that accounting for vector norms and exact residual connections is essential, as models naturally preserve individual word representations far more than mixing them. Crucially, the study discovered that including only the first layer normalization severely degrades attribution accuracy due to outlier weights, but incorporating both layer normalizations restores high fidelity because their outlier effects counteract one another. Furthermore, multi-layer aggregation proved vital, showing that models identify key decision-driving words within the first few layers before stabilizing.
These findings demonstrate that organizations can achieve highly faithful model explanations without the prohibitive computational costs of gradient-based techniques, as computing the proposed encoder maps takes seconds compared to hours for gradient alternatives. This provides a practical, scalable mechanism for auditing model behavior, ensuring regulatory compliance, and managing performance risks in high-stakes deployments.
Teams seeking to interpret Transformer architectures should transition away from raw attention weights or isolated sub-layer analyses in favor of full encoder layer aggregation. Future efforts should focus on validating GlobEnc across generative language models and broader operational datasets before fully standardizing its use in production monitoring.
The main limitation of this approach is its necessary omission of the non-linear feed-forward sub-layer, which cannot be linearly decomposed into token contributions. Nevertheless, confidence in the reported results is high, as the strong correlation gains remain robust across diverse model architectures, sizes, and evaluation tasks.
- Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). This paper establishes the foundational post hoc aggregation methods (attention rollout and attention flow) across Transformer layers, highlighting the failure of raw attention weights and motivating GlobEnc's whole-encoder attribution approach.
- Paper: Attention is not Explanation, Sarthak Jain et al. (2019). This work demonstrates that raw attention weights do not provide faithful explanations, laying the conceptual groundwork for developing attribution methods that incorporate other layer components.
- Paper: Transformer Feed-Forward Layers Are Key-Value Memories, Mor Geva et al. (2020). This paper reveals the functional importance of feed-forward networks in Transformers, underpinning GlobEnc's premise that non-attention components in the encoder block must be integrated into attribution analysis.
- Paper: What Does BERT Look at? An Analysis of BERT’s Attention, Kevin Clark et al. (2019). This foundational study analyzes individual attention head behaviors in BERT, providing key empirical context on how Transformers route information across layers.
- Paper: Axiomatic Attribution for Deep Networks, Mukund Sundararajan et al. (2017). This work introduces Integrated Gradients and axiomatic feature attribution, which serve as standard gradient-based saliency baselines used to evaluate token attribution faithfuless in GlobEnc.
- Paper: Methods for interpreting and understanding deep neural networks, Grégoire Montavon et al. (2018). This comprehensive overview covers decomposition-based explanation techniques like Layer-wise Relevance Propagation that inspire layer-by-layer attribution aggregation in neural networks.
- Paper: Measuring the Mixing of Contextual Information in the Transformer, Javier Ferrando et al. (2022). Building upon multi-component Transformer interpretation, this work specifically decomposes and measures how contextual information mixes through attention, residual connections, and subsequent transformations.
- Paper: Learning to Estimate Shapley Values with Vision Transformers, Ian Connick Covert et al. (2023). This paper extends the exploration of attribution beyond raw attention by developing efficient Shapley value estimation frameworks for vision transformers.
- Paper: What the DAAM: Interpreting Stable Diffusion Using Cross Attention, Raphael Tang et al. (2023). This work applies layer-aggregated token attribution principles to generative multimodal models by tracing cross-attention maps across diffusion steps.
