Measuring the Mixing of Contextual Information in the Transformer

Javier FerrandoGerard I. GállegoMarta R. Costa-jussà

article2022EMNLP89 citations

Proposes ALTI, an interpretability method that tracks information flow across full Transformer attention blocks to generate input attributions that surpass gradient-based techniques in faithfulness and reliability.

Listen

Modern natural language processing relies heavily on Transformer architectures, which aggregate contextual information across complex internal layers. However, understanding how these models mix information to reach specific decisions remains an open challenge. Traditional methods that rely solely on attention weights or standard gradient calculations often misrepresent token importance, creating reliability and transparency issues for organizations deploying artificial intelligence in critical workflows.

The article introduces and evaluates ALTI (Aggregation of Layer-wise Token-to-token Interactions), a novel interpretability method designed to provide faithful input attribution scores by tracking how contextual information flows and mixes throughout the attention blocks of Transformer models.

The researchers developed ALTI by decomposing the complete attention block—including multi-head attention, residual connections, and layer normalization—and calculating token interactions using Manhattan distance (the absolute sum of component differences) rather than Euclidean norm metrics. They then aggregated these layer-level contributions across the entire network. The approach was evaluated across three widely used language models (BERT, DistilBERT, and RoBERTa) on text classification and syntactic agreement benchmarks (SST-2, IMDB, Yelp, and a Wikipedia subject-verb agreement dataset), measuring faithfulness through standard erasure metrics and assessing robustness across multiple random model initializations.

The evaluation yielded several key findings. First, ALTI consistently outperformed existing gradient-based and attention-based attribution methods across all models and benchmarks, exceeding the leading gradient baseline by an average of 58% in comprehensiveness and 38% in sufficiency. Second, ALTI demonstrated substantial performance advantages on longer, multi-sentence inputs where gradient techniques degraded. Third, robustness tests across ten differently initialized BERT models showed that ALTI produced significantly higher ranking stability and correlation than alternative explainability methods. Finally, ablation analyses confirmed that using Manhattan distance rather than Euclidean distance better mitigates distortion caused by outlier embedding dimensions.

These findings indicate that ALTI delivers more accurate and consistent explanations of model behavior without requiring costly ad-hoc retraining. For technical leaders and operational teams, adopting this approach reduces the compliance and validation risks associated with deploying opaque machine learning systems, while improving error analysis in language-driven applications.

Organizations implementing Transformer models should consider integrating ALTI into their model auditing and explainability pipelines, especially for long-text classification tasks. Before full deployment, teams should note that ALTI currently tracks representation mixing within the Transformer body and does not account for final external classification layers. Future technical efforts should focus on extending the framework to generate class-specific explanations and evaluating its application across broader generative and multimodal tasks.

arXiv: 2203.04212
Cover for Measuring the Mixing of Contextual Information in the Transformer

Abstract

The Transformer architecture aggregates input information through the self-attention mechanism, but there is no clear understanding of how this information is mixed across the entire model. Additionally, recent works have demonstrated that attention weights alone are not enough to describe the flow of information. In this paper, we consider the whole attention block –multi-head attention, residual connection, and layer normalization– and define a metric to measure token-to-token interactions within each layer. Then, we aggregate layer-wise interpretations to provide input attribution scores for model predictions. Experimentally, we show that our method, ALTI (Aggregation of Layer-wise Token-to-token Interactions), provides more faithful explanations and increased robustness than gradient-based methods.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Attention Block Decomposition
  • 2.2 Attention Rollout
  • 3 Proposed Approach
  • 4 Experimental Setup
  • 4.1 Models
  • 4.2 Faithfulness Metrics
  • 4.3 Input Attribution Methods
  • 5 Results
  • 5.1 Faithfulness Results
  • 5.2 Robustness Analysis
  • 5.3 Ablation Study
  • 5.4 Addition of Layer Norm 2
  • 6 Conclusions
  • Limitations
  • Ethical Considerations
  • 7 Acknowledgements
  • References
  • A Layer Normalization decomposition
  • B RoBERTa and DistilBERT Results
  • C Qualitative Examples

Knowls

  1. Knowl 1 — Attention Block Linear Decomposition via Layer Normalization

    theoretical result

    In a Transformer layer with hidden dimension dd and HH attention heads (each of dimension dh=d/Hd_h = d/H), the output vector yi∈Rdy_i \in \mathbb{R}^d of the attention block for token position i∈{1,…,J}i \in \{1, \dots, J\} (comprising multi-head self-attention, residual connection, and layer normalization) can be decomposed into an exact sum of transformed token vectors and bias terms:

    yi=∑j=1JTi(xj)+1σ(x^i+xi)LbO+βy_i = \sum_{j=1}^J T_i(x_j) + \frac{1}{\sigma(\hat{x}_i + x_i)} L b_O + \beta

    where X=(x1,…,xJ)∈Rd×JX = (x_1, \dots, x_J) \in \mathbb{R}^{d \times J} are the layer input token representations, bO∈Rdb_O \in \mathbb{R}^d is the projection bias of the multi-head attention module, and x^i=∑h=1HWOhzih+bO\hat{x}_i = \sum_{h=1}^H W_O^h z_i^h + b_O with head state zih=∑j=1JAi,jhWVhxj∈Rdhz_i^h = \sum_{j=1}^J A_{i,j}^h W_V^h x_j \in \mathbb{R}^{d_h} (Ai,jhA_{i,j}^h being the attention weight from position ii to position jj in head hh, WVh∈Rdh×dW_V^h \in \mathbb{R}^{d_h \times d} the value weight matrix, and WOh∈Rd×dhW_O^h \in \mathbb{R}^{d \times d_h} the partitioned output projection matrix).

    Layer normalization LN(u)=u−μ(u)σ(u)⊙γ+β\text{LN}(u) = \frac{u - \mu(u)}{\sigma(u)} \odot \gamma + \beta is reformulated linearly as LN(u)=1σ(u)Lu+β\text{LN}(u) = \frac{1}{\sigma(u)} L u + \beta, where γ,β∈Rd\gamma, \beta \in \mathbb{R}^d are learned scale and bias vectors, and L∈Rd×dL \in \mathbb{R}^{d \times d} is the linear operator L=diag(γ)(I−1d11⊤)L = \text{diag}(\gamma)\left(I - \frac{1}{d}\mathbf{1}\mathbf{1}^\top\right).

    The transformed vectors Ti(xj)∈RdT_i(x_j) \in \mathbb{R}^d describing the contribution of token representation xjx_j to position ii are defined as:

    Ti(xj)={1σ(x^i+xi)L(∑h=1HWOhAi,jhWVhxj)if j≠i1σ(x^i+xi)L(∑h=1HWOhAi,jhWVhxj+xi)if j=iT_i(x_j) = \begin{cases} \frac{1}{\sigma(\hat{x}_i + x_i)} L \left( \sum_{h=1}^H W_O^h A_{i,j}^h W_V^h x_j \right) & \text{if } j \neq i \\[8pt] \frac{1}{\sigma(\hat{x}_i + x_i)} L \left( \sum_{h=1}^H W_O^h A_{i,j}^h W_V^h x_j + x_i \right) & \text{if } j = i \end{cases}

  2. Knowl 2 — ALTI Layer-Wise Token-to-Token Interaction Metric

    model/method

    In the ALTI (Aggregation of Layer-wise Token-to-token Interactions) framework, the contribution ci,jc_{i,j} of input token representation xjx_j to the attention block output representation yiy_i is defined by measuring the geometric proximity of the transformed vector Ti(xj)T_i(x_j) to the resultant output vector yiy_i.

    First, the Manhattan distance between the output vector and the transformed vector is computed:

    di,j=∥yi−Ti(xj)∥1d_{i,j} = \|y_i - T_i(x_j)\|_1

    The use of the ℓ1\ell_1 norm rather than the ℓ2\ell_2 norm avoids disproportionate weighting of outlier embedding dimensions that occur in contextualized representations. The proximity is then quantified as −di,j-d_{i,j}, truncated at the ℓ1\ell_1 magnitude of yiy_i to neglect vectors pointing too far from yiy_i, and normalized across all tokens j∈{1,…,J}j \in \{1, \dots, J\}:

    ci,j=max⁡(0,−di,j+∥yi∥1)∑k=1Jmax⁡(0,−di,k+∥yi∥1)c_{i,j} = \frac{\max\left(0, -d_{i,j} + \|y_i\|_1\right)}{\sum_{k=1}^J \max\left(0, -d_{i,k} + \|y_i\|_1\right)}

    Collecting all pairwise interactions yields the layer-wise token contribution matrix C∈RJ×JC \in \mathbb{R}^{J \times J}, where each row ii sums to 1.

  3. Knowl 3 — ALTI Multi-Layer Attribution Rollout

    model/method

    To track the global flow of information from the initial input tokens across ll Transformer layers into intermediate or final representations, ALTI aggregates the layer-wise contribution matrices C1,…,ClC^1, \dots, C^l via sequential matrix multiplication (a rollout on the contribution graph):

    Rl=Cl⋅Cl−1⋯C1R^l = C^l \cdot C^{l-1} \cdots C^1

    where Cm∈RJ×JC^m \in \mathbb{R}^{J \times J} is the layer-wise contribution matrix at layer mm, and Rl∈RJ×JR^l \in \mathbb{R}^{J \times J} is the global attribution matrix at layer ll.

    For sequence classification tasks, the input attribution scores corresponding to the full input sequence are extracted from the row of RLR^L associated with the sentence representation token at the final layer LL, such as R[CLS]L∈RJR^L_{[\text{CLS}]} \in \mathbb{R}^J in classification models or R[MASK]L∈RJR^L_{[\text{MASK}]} \in \mathbb{R}^J in masked language modeling tasks.

  4. Knowl 4 — Faithfulness Comparison of ALTI against Gradient and Attention Attribution Baselines

    data/table

    The faithfulness of interpretability methods is evaluated across four datasets: Stanford Sentiment Treebank (SST-2), Yelp Review Polarity, IMDB, and Subject-Verb Agreement (SVA) across BERT, RoBERTa, and DistilBERT. Faithfulness is assessed via Comprehensiveness (Comp., higher is better), which measures the average probability drop when deleting the top-k%k\% most important tokens, and Sufficiency (Suff., lower is better), which measures the probability drop when keeping only the top-k%k\% tokens, with k∈{0,5,10,20,50}k \in \{0, 5, 10, 20, 50\}.

    BERT RoBERTa DistilBERT
    SST-2 Yelp IMDB SVA SST-2 SST-2
    Methods Comp.↑\uparrow Suff.↓\downarrow Comp. Suff. Comp. Suff. Comp. Suff. Comp. Suff. Comp. Suff.
    Gradℓ2\text{Grad}_{\ell2} 0.204 0.076 0.083 0.101 0.192 0.052 0.284 0.145 0.190 0.075 0.230 0.066
    IGℓ2\text{IG}_{\ell2} 0.223 0.084 0.111 0.024 0.214 0.060 0.315 0.171 0.223 0.074 0.295 0.048
    IGμ\text{IG}_\mu 0.211 0.112 0.106 0.022 0.179 0.063 0.317 0.174 0.231 0.067 0.279 0.064
    G×Iℓ2\text{G} \times \text{I}_{\ell2} 0.199 0.080 0.081 0.104 0.197 0.056 0.279 0.149 0.187 0.081 0.235 0.065
    G×Iμ\text{G} \times \text{I}_\mu 0.207 0.073 0.087 0.098 0.213 0.054 0.285 0.145 0.191 0.078 0.237 0.065
    Rollout 0.074 0.270 0.076 0.102 0.090 0.185 0.108 0.292 0.076 0.179 0.152 0.147
    Globenc 0.174 0.125 0.118 0.053 0.207 0.117 0.251 0.178 0.161 0.119 0.240 0.072
    ALTI 0.317 0.044 0.255 0.022 0.308 0.031 0.372 0.088 0.269 0.057 0.332 0.034

    ALTI outperforms all gradient-based methods, Attention Rollout, and Globenc across all models and datasets, beating the strongest gradient baseline (IGℓ2\text{IG}_{\ell2}) by an average of 58% in comprehensiveness and 38% in sufficiency.

  5. Knowl 5 — Attribution Robustness Across Random Weight Initializations

    empirical result

    The robustness of interpretability methods under the implementation invariance criterion is measured using 10 MultiBERTs models that share the identical architecture and training data but differ only in their random weight initializations.

    For each method, the similarity between explanations from all model pairs is evaluated using:

    1. The Jaccard similarity between the top-25% ranked tokens (Jaccard-25%\text{Jaccard-25\%}).
    2. The Spearman's rank correlation coefficient across ranked token attributions.

    ALTI produces significantly more consistent attributions across randomly initialized models than gradient-based methods:

    • ALTI achieves a median Jaccard-25%\text{Jaccard-25\%} score of approximately 0.750.75, whereas gradient methods (Gradℓ2\text{Grad}_{\ell2}, IGℓ2\text{IG}_{\ell2}, IGμ\text{IG}_\mu, G×Iℓ2\text{G} \times \text{I}_{\ell2}, G×Iμ\text{G} \times \text{I}_\mu) achieve medians between 0.380.38 and 0.650.65.
    • ALTI achieves a median Spearman rank correlation of approximately 0.930.93, whereas gradient methods fall between 0.650.65 and 0.800.80.
  6. Knowl 6 — Faithfulness Comparison of Proximity-Based Token Contributions versus Vector Norms

    data/table

    To isolate the effect of computing layer-wise token contributions via vector proximity to the output representation (yiy_i) versus using the vector norm (∥Ti(xj)∥2\|T_i(x_j)\|_2, the Norms method of Kobayashi et al.), both methods are evaluated when aggregated using the rollout matrix multiplication across BERT, RoBERTa, and DistilBERT. For a direct comparison independent of norm selection, the ℓ2\ell_2 version of ALTI (ALTI ℓ2\ell_2) is compared with Norms.

    BERT RoBERTa DistilBERT
    Dataset Metric ALTI ℓ2\ell_2 Norms ALTI ℓ2\ell_2 Norms ALTI ℓ2\ell_2 Norms
    SST-2 Comp.↑\uparrow 0.300 0.134 0.230 0.168 0.292 0.157
    Suff.↓\downarrow 0.045 0.180 0.075 0.135 0.051 0.152
    IMDB Comp. 0.286 0.148 0.203 0.133 0.254 0.175
    Suff. 0.031 0.115 0.086 0.162 0.053 0.104
    Yelp Comp. 0.212 0.082 0.108 0.064 0.200 0.093
    Suff. 0.020 0.057 0.104 0.180 0.031 0.114
    SVA Comp. 0.374 0.273 0.379 0.275 0.357 0.315
    Suff. 0.088 0.152 0.114 0.178 0.096 0.117

    Measuring contributions via proximity to the resultant output vector yiy_i substantially outperforms standalone vector norms in every dataset, architecture, and metric.

  7. Knowl 7 — Ablation of Distance Metric Norms in ALTI

    data/table

    Evaluating ALTI with Manhattan distance (ℓ1\ell_1 norm) versus Euclidean distance (ℓ2\ell_2 norm) demonstrates that the ℓ1\ell_1 norm provides superior or equal faithfulness across nearly all models and datasets.

    BERT RoBERTa DistilBERT
    Dataset Metric ℓ1\ell_1 ℓ2\ell_2 ℓ1\ell_1 ℓ2\ell_2 ℓ1\ell_1 ℓ2\ell_2
    SST-2 Comp.↑\uparrow 0.317 0.300 0.269 0.230 0.332 0.292
    Suff.↓\downarrow 0.044 0.045 0.057 0.075 0.034 0.051
    IMDB Comp. 0.308 0.286 0.266 0.203 0.304 0.254
    Suff. 0.031 0.031 0.050 0.086 0.039 0.053
    Yelp Comp. 0.221 0.212 0.138 0.108 0.237 0.200
    Suff. 0.020 0.020 0.070 0.104 0.017 0.031
    SVA Comp. 0.372 0.374 0.390 0.379 0.382 0.357
    Suff. 0.088 0.088 0.109 0.114 0.084 0.096

    The advantage of ℓ1\ell_1 over ℓ2\ell_2 stems from its robustness to rogue outlier embedding dimensions that distort Euclidean distances. The advantage is less pronounced on BERT due to the lower anisotropy observed in its transformed representations compared to RoBERTa and DistilBERT.

  8. Knowl 8 — Impact of Secondary Feed-Forward Layer Normalization on ALTI

    empirical result

    Incorporating the second layer normalization (LN2) of the Transformer feed-forward network (FFN) module into the ALTI decomposition—analogous to its inclusion in Globenc—produces a negligible difference in attribution faithfulness.

    Across 10 random seeds of BERT on the SST-2 dataset, the probability drop curve when removing important tokens for ALTI with LN2 is virtually identical to standard ALTI without LN2, demonstrating that tracking token-to-token interactions across the attention block alone captures the relevant contextual mixing.

  9. Knowl 9 — Class-Agnostic Representation Attribution Limitation

    limitation

    ALTI measures the aggregation and mixing of contextual information from input tokens into the hidden representations of the final Transformer layer (such as the [CLS] or [MASK] token representation). However, it does not incorporate the final task-specific classification head parameters (linear layer weights, biases, and activation functions) situated on top of the Transformer encoder.

    As a result, ALTI provides attributions reflecting input contribution to the final representation rather than class-specific explanations for individual output class logits, in contrast to gradient-based attribution methods that differentiate specific class logits.

Coverage note — Appendix tables reporting DistilBERT and RoBERTa results on IMDB, Yelp, and SVA (Tables 6 and 7) were omitted as their conclusions and comparative dynamics are already fully captured in the main results table and ablation tables.

References

  1. 1.Samira Abnar and Willem Zuidema. 2020. Quantifying attention flow in transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4190–4197, Online. Association for Computational Linguistics.
  2. 2.Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020. A diagnostic study of explainability techniques for text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3256–3274, Online. Association for Computational Linguistics.
  3. 3.Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, and Katja Filippova. 2021. "will you find these shortcuts?" a protocol for evaluating the faithfulness of input salience methods for text classification.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020. On identifiability in transformers. In International Conference on Learning Representations.
  6. 6.Xingyu Cai, Jiaji Huang, Yuchen Bian, and Kenneth Church. 2021. Isotropy in the contextual embedding space: Clusters and manifolds. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  7. 7.Chun Sik Chan, Huanqi Kong, and Liang Guanqing. 2022. A comparative study of faithfulness metrics for model interpretability methods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5029–5038, Dublin, Ireland. Association for Computational Linguistics.
  8. 8.Hila Chefer, Shir Gur, and Lior Wolf. 2021a. Generic attention-model explainability for interpreting bimodal and encoder-decoder transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 397–406.
  9. 9.Hila Chefer, Shir Gur, and Lior Wolf. 2021b. Transformer interpretability beyond attention visualization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 782–791.
  10. 10.Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. What does BERT look at? an analysis of BERT’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 276–286, Florence, Italy. Association for Computational Linguistics.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  12. 12.Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020. ERASER: A benchmark to evaluate rationalized NLP models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4443–4458, Online. Association for Computational Linguistics.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.
  14. 14.Kawin Ethayarajh. 2019. How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 55–65, Hong Kong, China. Association for Computational Linguistics.
  15. 15.Yoav Goldberg. 2019. Assessing bert’s syntactic abilities. CoRR, abs/1901.05287.
  16. 16.Sarthak Jain and Byron C. Wallace. 2019. Attention is not Explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3543–3556, Minneapolis, Minnesota. Association for Computational Linguistics.
  17. 17.Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020. Attention is not only a weight: Analyzing transformers with vector norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7057–7075, Online. Association for Computational Linguistics.
  18. 18.Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2021. Incorporating Residual and Normalization Layers into Analysis of Masked Language Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4547–4568, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  19. 19.Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson. 2020. Captum: A unified and generic model interpretability library for pytorch.
  20. 20.Olga Kovaleva, Saurabh Kulshreshtha, Anna Rogers, and Anna Rumshisky. 2021. BERT busters: Outlier dimensions that disrupt transformers. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3392–3405, Online. Association for Computational Linguistics.
  21. 21.Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019. Revealing the dark secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4365–4374, Hong Kong, China. Association for Computational Linguistics.
  22. 22.Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016. Visualizing and understanding neural models in NLP. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 681–691, San Diego, California. Association for Computational Linguistics.
  23. 23.Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016. Assessing the ability of lstms to learn syntax-sensitive dependencies. Trans. Assoc. Comput. Linguistics, 4:521–535.
  24. 24.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  25. 25.Kaiji Lu, Zifan Wang, Piotr Mardziel, and Anupam Datta. 2021. Influence patterns for explaining information flow in BERT. In Advances in Neural Information Processing Systems.
  26. 26.Ziyang Luo, Artur Kulmizev, and Xiaoxi Mao. 2021. Positional artefacts propagate through masked language model embeddings. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5312–5327, Online. Association for Computational Linguistics.
  27. 27.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  28. 28.Andreas Madsen, Nicholas Meade, Vaibhav Adlakha, and Siva Reddy. 2021a. Evaluating the faithfulness of importance measures in nlp by recursively masking allegedly important tokens and retraining.
  29. 29.Andreas Madsen, Siva Reddy, and Sarath Chandar. 2021b. Post-hoc interpretability for neural nlp: A survey.
  30. 30.Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, and Mohammad Taher Pilehvar. 2022. Globenc: Quantifying global token attribution by incorporating the whole encoder layer in transformers.
  31. 31.John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020. Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 119–126.
  32. 32.Damian Pascual, Gino Brunner, and Roger Wattenhofer. 2021. Telling BERT’s full story: from local attention to global aggregation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 105–124, Online. Association for Computational Linguistics.
  33. 33.Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C. Lipton. 2020. Learning to deceive with attention-based explanations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4782–4793, Online. Association for Computational Linguistics.
  34. 34.Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. "why should i trust you?": Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, page 1135–1144, New York, NY, USA. Association for Computing Machinery.
  35. 35.Hassan Sajjad, Narine Kokhlikyan, Fahim Dalvi, and Nadir Durrani. 2021. Fine-grained interpretation and causation analysis in deep NLP models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Tutorials, pages 5–10, Online. Association for Computational Linguistics.
  36. 36.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. ArXiv, abs/1910.01108.
  37. 37.Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Raluca Turc, Jacob Eisenstein, Dipanjan Das, and Ellie Pavlick. 2022. The multiBERTs: BERT reproductions for robustness analysis. In International Conference on Learning Representations.
  38. 38.Sofia Serrano and Noah A. Smith. 2019. Is attention interpretable? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2931–2951, Florence, Italy. Association for Computational Linguistics.
  39. 39.Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje. 2016. Not just a black box: Learning important features through propagating activation differences. CoRR, abs/1605.01713.
  40. 40.Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014. Deep inside convolutional networks: Visualising image classification models and saliency maps. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Workshop Track Proceedings.
  41. 41.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
  42. 42.Kaiser Sun and Ana Marasovic´. 2021. Effective attention sheds light on interpretability. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4126–4135, Online. Association for Computational Linguistics.
  43. 43.Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 3319–3328, International Convention Centre, Sydney, Australia. PMLR.
  44. 44.William Timkey and Marten van Schijndel. 2021. All bark and no bite: Rogue dimensions in transformer language models obscure representational quality. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4527–4546, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  45. 45.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  46. 46.Jesse Vig and Yonatan Belinkov. 2019. Analyzing the structure of attention in a transformer language model. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 63–76, Florence, Italy. Association for Computational Linguistics.
  47. 47.Sarah Wiegreffe and Yuval Pinter. 2019. Attention is not not explanation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 11–20, Hong Kong, China. Association for Computational Linguistics.
  48. 48.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  49. 49.Muhammad Bilal Zafar, Michele Donini, Dylan Slack, Cedric Archambeau, Sanjiv Das, and Krishnaram Kenthapadi. 2021. On the lack of robust interpretability of neural text classifiers. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3730–3740, Online. Association for Computational Linguistics.
  50. 50.Kerem Zaman and Yonatan Belinkov. 2022. A multilingual perspective towards the evaluation of attribution methods in natural language inference.
  51. 51.Matthew D. Zeiler and Rob Fergus. 2014. Visualizing and understanding convolutional networks. In Computer Vision – ECCV 2014, pages 818–833, Cham. Springer International Publishing.
  52. 52.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc.

Citation

MLA
Ferrando, J., et al. “Measuring the Mixing of Contextual Information in the Transformer”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 8698–714, https://doi.org/10.18653/v1/2022.emnlp-main.595.
APA
Ferrando, J., Gállego, G. I., & Costa-jussà, M. R. (2022). Measuring the Mixing of Contextual Information in the Transformer. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 8698–8714. https://doi.org/10.18653/v1/2022.emnlp-main.595
Chicago
Ferrando, J., G. I. Gállego, and M. R. Costa-jussà. 2022. “Measuring the Mixing of Contextual Information in the Transformer”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 8698–8714. https://doi.org/10.18653/v1/2022.emnlp-main.595.
Harvard
Ferrando, J., Gállego, G.I. and Costa-jussà, M.R. (2022) “Measuring the Mixing of Contextual Information in the Transformer”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 8698–8714. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.595.
Vancouver
1. Ferrando J, Gállego GI, Costa-jussà MR (2022) Measuring the Mixing of Contextual Information in the Transformer. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 8698–8714

BibTeX

@inproceedings{ferrando-etal-2022-measuring,
    title = "Measuring the Mixing of Contextual Information in the Transformer",
    author = "Ferrando, Javier  and
      G{\'a}llego, Gerard I.  and
      Costa-juss{\`a}, Marta R.",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.595/",
    doi = "10.18653/v1/2022.emnlp-main.595",
    pages = "8698--8714"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/