Measuring the Mixing of Contextual Information in the Transformer
Javier FerrandoGerard I. GállegoMarta R. Costa-jussà
Proposes ALTI, an interpretability method that tracks information flow across full Transformer attention blocks to generate input attributions that surpass gradient-based techniques in faithfulness and reliability.
Modern natural language processing relies heavily on Transformer architectures, which aggregate contextual information across complex internal layers. However, understanding how these models mix information to reach specific decisions remains an open challenge. Traditional methods that rely solely on attention weights or standard gradient calculations often misrepresent token importance, creating reliability and transparency issues for organizations deploying artificial intelligence in critical workflows.
The article introduces and evaluates ALTI (Aggregation of Layer-wise Token-to-token Interactions), a novel interpretability method designed to provide faithful input attribution scores by tracking how contextual information flows and mixes throughout the attention blocks of Transformer models.
The researchers developed ALTI by decomposing the complete attention block—including multi-head attention, residual connections, and layer normalization—and calculating token interactions using Manhattan distance (the absolute sum of component differences) rather than Euclidean norm metrics. They then aggregated these layer-level contributions across the entire network. The approach was evaluated across three widely used language models (BERT, DistilBERT, and RoBERTa) on text classification and syntactic agreement benchmarks (SST-2, IMDB, Yelp, and a Wikipedia subject-verb agreement dataset), measuring faithfulness through standard erasure metrics and assessing robustness across multiple random model initializations.
The evaluation yielded several key findings. First, ALTI consistently outperformed existing gradient-based and attention-based attribution methods across all models and benchmarks, exceeding the leading gradient baseline by an average of 58% in comprehensiveness and 38% in sufficiency. Second, ALTI demonstrated substantial performance advantages on longer, multi-sentence inputs where gradient techniques degraded. Third, robustness tests across ten differently initialized BERT models showed that ALTI produced significantly higher ranking stability and correlation than alternative explainability methods. Finally, ablation analyses confirmed that using Manhattan distance rather than Euclidean distance better mitigates distortion caused by outlier embedding dimensions.
These findings indicate that ALTI delivers more accurate and consistent explanations of model behavior without requiring costly ad-hoc retraining. For technical leaders and operational teams, adopting this approach reduces the compliance and validation risks associated with deploying opaque machine learning systems, while improving error analysis in language-driven applications.
Organizations implementing Transformer models should consider integrating ALTI into their model auditing and explainability pipelines, especially for long-text classification tasks. Before full deployment, teams should note that ALTI currently tracks representation mixing within the Transformer body and does not account for final external classification layers. Future technical efforts should focus on extending the framework to generate class-specific explanations and evaluating its application across broader generative and multimodal tasks.
- Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). It introduces foundational methods like attention rollout and attention flow to track information mixing across Transformer layers, establishing the exact problem and framework ALTI builds upon.
- Paper: Attention is not Explanation, Sarthak Jain et al. (2019). It demonstrates that raw attention weights fail as faithful explanations of token importance, motivating ALTI's layer-wise decomposition approach.
- Paper: Axiomatic Attribution for Deep Networks, Mukund Sundararajan et al. (2017). It formalizes axiomatic feature attribution via Integrated Gradients, providing the theoretical benchmark against which ALTI measures its improvements in attribution faithfulness.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). It characterizes the geometric anisotropy and context-mixing behaviors of Transformer representations across layers that ALTI explicitly measures.
- Paper: A Primer in BERTology: What We Know About How BERT Works, Anna Rogers et al. (2020). It synthesizes empirical findings on layer-wise representation mixing and attention head behaviors across BERT models, providing essential context for evaluating ALTI's attribution results.
- Paper: What Does BERT Look at? An Analysis of BERT’s Attention, Kevin Clark et al. (2019). It analyzes BERT's internal attention patterns and shows why individual attention heads alone do not fully reflect input importance.
- Paper: GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers, Ali Modarressi et al. (2022). It extends layer-wise attribution in Transformers by incorporating the full encoder block—including vector norms and layer normalizations—to compute global token attributions.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). It advances from measuring low-level token-to-token mixing to a top-down representation engineering framework for monitoring and steering high-level concepts in language models.
- Paper: Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham et al. (2023). It investigates how mixed representation spaces in language models can be disentangled into interpretable, monosemantic features using sparse autoencoders.
- Paper: Faith-Shap: The Faithful Shapley Interaction Index, Che-Ping Tsai et al. (2023). It builds upon feature attribution and interaction metrics by formalizing game-theoretic higher-order feature interactions for faithful model explanations.
- Paper: Supervising Model Attention with Human Explanations for Robust Natural Language Inference, Joe Stacey et al. (2022). It operationalizes insights into attention and token attribution by directly supervising internal attention distributions to align with human explanations.
