Information Flow Routes: Automatically Interpreting Language Models at Scale

Javier FerrandoElena Voita

article2024EMNLP111 citations

Proposes an attribution-based method to automatically extract information flow subgraphs for transformer language model predictions in a single forward pass, providing an interpretation approach roughly 100 times faster than activation patching without requiring human-crafted templates.

Listen

Modern artificial intelligence relies heavily on large Transformer-based language models, yet understanding their internal decision-making processes remains a difficult challenge. Existing interpretability techniques largely depend on activation patching, which requires human experts to manually design contrastive prompt templates and execute numerous intervention passes. This manual approach is slow, costly, and limited to narrow, pre-selected test cases.

The article introduces an automated method called information flow routes to efficiently interpret language models at scale. The primary objective is to demonstrate that model predictions can be explained by tracing internal computation subgraphs using an attribution-based approach rather than activation patching.

The researchers formulated the internal representations and operations of Transformer models as directed graphs and applied a top-down attribution algorithm to extract the most relevant pathways for any given prediction. They evaluated the framework on standard mechanistic interpretability benchmarks, such as Indirect Object Identification and Greater-Than tasks, and conducted broad empirical analyses across subsets of the C4 dataset, multilingual text from FLORES-200, code data from CodeParrot, and arithmetic tasks using GPT-2, OPT-125m, and Llama 2 models.

The analysis yielded several key findings. First, the attribution method uncovers internal computational circuits roughly 100 times faster than patching-based methods, reducing execution time from eight minutes to five seconds in benchmark tests. Second, the method evaluates the overall importance of components directly from a single forward pass without suffering from template fragility or self-repair artifacts. Third, across general text, attention heads in the lower layers exhibit universal functional roles, such as previous-token tracking and subword merging. Fourth, component activation strongly aligns with input linguistic properties, notably clustering by parts of speech for function words. Finally, distinct sets of attention heads specialize in specific tasks, such as coding, arithmetic, and non-English languages, directly promoting topic-specific tokens.

These findings indicate that language model computations are modular and can be systematically mapped without burdensome manual intervention. By drastically reducing the computational and human overhead of internal model auditing, this approach enables scalable monitoring of model behaviors, safety mechanisms, and domain-specific operations across large-scale deployments.

The article suggests that organizations should leverage automated attribution graphs to audit model pathways and identify specialized components across various domains. Future work should expand the analysis to fine-grained neuron-level studies in feed-forward layers and investigate whether anomalous internal pathways, such as punctuation tokens inadvertently acting as sentence boundaries, cause generation errors.

The study's primary limitation is that empirical evaluations were conducted on GPT-2, OPT, and Llama 2 model families. While the theoretical framework generalizes to any standard Transformer architecture, stakeholders should maintain moderate caution when applying these observations to models with fundamentally different designs without further validation.

Ferrando et al (2024).pdf
  • Paper: Towards Automated Circuit Discovery for Mechanistic Interpretability, Arthur Conmy et al. (2023). ACDC establishes how mechanistic interpretability extracts computational circuits with activation patching, the baseline that Information Flow Routes automates and seeks to avoid.
  • Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). Its attention rollout and flow methods show how Transformer information can be traced through layered graphs, a useful foundation for understanding the source’s attribution-based routes.
  • Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). This later critique examines the reliability of attribution and circuit-discovery explanations, challenging the confidence readers might place in automated routes and their apparent findings.
Cover for Information Flow Routes: Automatically Interpreting Language Models at Scale

Abstract

Information flows by routes inside the network via mechanisms implemented in the model. These routes can be represented as graphs where nodes correspond to token representations and edges to computations. We automatically build these graphs in a top-down manner, for each prediction leaving only the most important nodes and edges. In contrast to the existing workflows relying on activation patching, we do this through attribution: this allows us to efficiently uncover existing circuits with just a single forward pass. Unlike with patching, we do not need a human to carefully design prediction templates, and we can extract information flow routes for any prediction (not just the ones among the allowed templates). As a result, we can analyze model behavior in general, for specific types of predictions, or different domains. We experiment with Llama 2 and show that some attention head roles are overall important, e.g. previous token heads and subword merging heads. Next, we find similarities in Llama 2 behavior when handling tokens of the same part of speech. Finally, we show that some model components can be specialized on domains such as coding or multilingual texts.

Table of Contents

  • 1 Introduction
  • 2 Extracting Information Flow Routes
  • 2.1 Evaluating Edge Importance
  • 2.1.1 Definition of Importance
  • 2.1.2 Defining FFN Edges
  • 2.1.3 Defining Attention Edges
  • 2.2 Extracting the Important Subgraph
  • 3 Information Flow vs Patching Circuits
  • 3.1 Indirect Object Identification
  • 3.2 Greater-than
  • 3.3 Hundred Times Faster than Patching
  • 4 General Experiments
  • 4.1 Component Importance for POS
  • 4.2 Bottom-to-Top Patterns
  • 4.3 Positional and Subword Merging Heads
  • 5 Domain-Specific Model Components
  • 5.1 Components are Specialized
  • 5.2 Specialized Heads Output Topic-Related Concepts
  • 6 Additional Related Work
  • 7 Conclusions
  • 8 Limitations
  • 9 Acknowledgements
  • References
  • A Background on Activation Patching
  • B Details about the Information Flow Routes
  • B.1 Linearizing the Layer Normalization
  • B.2 Further details about the implementation of Information Flow Routes
  • B.3 Folding the Layernorm
  • C Examples of Routes
  • D Peculiar Information Flow Patterns, or Periods Acting as BOS
  • E Specialization

Knowls

  1. Knowl 1 — Top-down extraction of prediction-specific information-flow routes

    algorithm

    The method represents a Transformer computation as a directed graph: nodes are token representations at residual-stream stages, and edges are operations such as residual connections, attention contributions, and feed-forward network (FFN) updates. To extract a route for one next-token prediction, it runs the model once and caches the activations, then traces backwards from the final residual-stream representation at the predicting position. At each visited node, it retains incoming edges whose local attribution importance exceeds a threshold τ\tau and adds their source nodes to the traversal. The process ends when no newly retained source nodes remain, producing a prediction-specific subgraph without activation interventions.

    For a node representation y=∑j=1mzjy=\sum_{j=1}^{m}z_j, where each zjz_j is an incoming vector in the same representation space as yy, the score used to select an edge is

    importance⁡(zj,y)=proximity⁡(zj,y)∑k=1mproximity⁡(zk,y),proximity⁡(zj,y)=max⁡(∥y∥1−∥zj−y∥1,0).\operatorname{importance}(z_j,y)=\frac{\operatorname{proximity}(z_j,y)}{\sum_{k=1}^{m}\operatorname{proximity}(z_k,y)},\qquad \operatorname{proximity}(z_j,y)=\max\bigl(\|y\|_1-\|z_j-y\|_1,0\bigr).

    In the reported implementation, attention sub-edges below τ\tau are removed and the remaining sub-edge scores are renormalized before scores are aggregated across heads. The traversal uses the aggregated attention edges and the FFN and residual edges. The threshold controls route sparsity; the paper does not prescribe a universal value.

  2. Knowl 2 — Local attribution measures an incoming vector’s contribution to a sum

    equation

    The route method scores each incoming vector by its proximity to the resulting node representation, rather than by an activation-patching intervention. If y=∑j=1mzjy=\sum_{j=1}^{m}z_j, with yy and each zjz_j vectors in the same representation space, then

    proximity⁡(zj,y)=max⁡(−∥zj−y∥1+∥y∥1,0),importance⁡(zj,y)=proximity⁡(zj,y)∑k=1mproximity⁡(zk,y).\operatorname{proximity}(z_j,y)=\max\left(-\|z_j-y\|_1+\|y\|_1,0\right),\qquad \operatorname{importance}(z_j,y)=\frac{\operatorname{proximity}(z_j,y)}{\sum_{k=1}^{m}\operatorname{proximity}(z_k,y)}.

    Here ∥⋅∥1\|\cdot\|_1 is the L1L_1 norm. A vector closer to the total yy receives greater importance, while a vector whose distance from yy exceeds ∥y∥1\|y\|_1 receives zero proximity. The resulting importance scores are normalized across the incoming vectors when their total proximity is nonzero. These scores describe local contributions to a sum; they are not causal effects measured by changing an activation.

  3. Knowl 3 — Attention contributions are decomposed by source token and head

    model/method

    For a destination token at position pp, the attention update is decomposed into separate contributions from each source position j≤pj\leq p and attention head hh. With residual-stream row vector xjx_j at source position jj, attention weight αp,jh\alpha^h_{p,j}, and the linearized layer-normalization map LL, the head-specific output matrix is WOVh=WVhWOhW^h_{OV}=W^h_VW^h_O and the contribution is αp,jhfh(xj)\alpha^h_{p,j}f^h(x_j), where fh(xj)=xjLWOVhf^h(x_j)=x_jLW^h_{OV}. Thus the attention-block update is the sum of these contributions over heads and source positions. The head-specific output matrices are the corresponding portions of the attention output matrix; WVhW^h_V and WOhW^h_O are the value and head-specific output matrices, respectively.

    Each sub-edge is scored against the attention-block output representation xpAx^A_p using the local vector-attribution measure. The scores for a given source position are then summed across heads. For j≠pj\ne p, this sum is the attention-edge importance from position jj to pp; for j=pj=p, the importance of the current-token residual connection is added as well. This decomposition allows the route graph to distinguish which source-token representations contribute through attention, instead of treating the whole attention block as one undivided update.

  4. Knowl 4 — Routes recover task-specific circuits while exposing contrastive-template sensitivity

    empirical result

    On GPT-2 Small, information-flow routes for indirect object identification (IOI) and the greater-than number-prediction task contained attention heads previously identified as task-relevant by activation patching. The routes also retained heads that contributed to the predictions more generally: for IOI, such heads appeared in the original and contrastive predictions, while comparing the two route sets removed these generic heads and highlighted task-specific differences. Previous-token heads, for example, contributed to both IOI versions rather than being specific to IOI.

    The greater-than results depended on the contrastive example. Comparing a prompt ending in year digits YY=01YY=01 with a greater-than prompt did not substantially change overall component contributions: predicting a number greater than 01 still requires a constraint that excludes values such as 00. With a contrastive prompt using YY=00YY=00, the heads associated with greater-than reasoning became less important. The authors use this variation to argue that patching-based conclusions can depend on a human-selected contrastive template, whereas routes can show both overall contributions and differences between chosen predictions.

  5. Knowl 5 — Route extraction was about 100 times faster than automated patching

    empirical result

    For circuit discovery on a batch of 50 IOI examples, the paper compares information-flow route extraction with ACDC, an automated circuit-discovery method based on activation patching. ACDC required 8 minutes on a single GPU, whereas route extraction took 5 seconds, which the authors report as roughly a 100-fold reduction in runtime. Route extraction obtains the routes after a forward pass and caching internal activations, rather than performing the many patched forward passes required by the comparison method.

  6. Knowl 6 — Component-importance patterns vary with token part of speech and subword position

    empirical result

    For Llama 2-7B, the authors formed a component-importance vector for each prediction from attention-head contributions summed over source positions and FFN contributions, then visualized the vectors with t-SNE. The analysis used predictions from 1,000 C4 sentences and colored points by the input token’s part of speech (POS), the next token’s POS, or whether the input was the first or a later subword. POS tags came from the NLTK universal tagset; the next token was taken from the dataset reference, not generated by the model, and its POS was assigned to the token’s subwords.

    Function-word inputs formed clusters associated with POS, whereas content-word inputs were much more mixed, with only some separation such as verbs from nouns. Importance patterns also differed strongly between first and later subwords; the number-token clusters separated in a way consistent with first versus later digits. Comparing input-token and next-token coloring, the patterns tracked the input token’s POS more clearly than the next token’s POS.

  7. Knowl 7 — Attention and FFN contributions are concentrated lower in the network, with a strong final FFN

    empirical result

    In Llama 2-7B, the authors measured attention-head activation frequency and FFN importance across layers, using an edge threshold of τ=0.01\tau=0.01. A head counted as active when at least one of its sub-edges appeared in the extracted route. The results show greater attention and FFN activity in lower layers and declining activity for individual heads and FFNs higher in the network. The final-layer FFN is an exception to the general decline: it has high importance, indicating that this last FFN update substantially changes the representation used for prediction.

  8. Knowl 8 — Previous-token and subword-merging heads are highly important head functions

    empirical result

    The paper classifies attention heads by their token-contribution patterns, rather than by attention weights alone. A head is a previous-token head if, in at least 70% of cases, more than half of its influence is assigned to the immediately preceding token. A head is a subword-merging head if it is not a previous-token head and, in at least 70% of cases, later subwords assign more than half their influence to preceding subwords of the same word; for the first subword, the head’s overall influence must be no more than 0.005% in at least 70% of cases.

    In Llama 2-7B, several previous-token heads were found, and almost all were among the most important heads in their layers. Several subword-merging heads occurred in the lower network and were also among the most important heads in their layers. These heads chiefly mattered for later subwords, consistent with the observed difference in importance patterns between first and later subwords.

  9. Knowl 9 — Llama 2-7B components specialize across domains

    empirical result

    The authors measured average component importance for Llama 2-7B on C4 text, English, Russian, Italian, and Spanish data from FLORES-200, CodeParrot code, and addition/subtraction data. Components with low importance on general C4 text could be highly important in a particular domain. The important heads differed across code, non-English text, and arithmetic, supporting the finding that many components are specialized rather than uniformly important across inputs. The FFN results also varied by domain: the final FFN was less relevant for non-English data than for other domains, and some layers’ FFN importance fell to zero on addition or subtraction.

    The experiments in the paper cover GPT-2, OPT, and Llama 2 model families; whether the findings generalize to other Transformer-based language models was not tested.

  10. Knowl 10 — Domain-specific heads write interpretable topic-related concepts

    empirical result

    To investigate what domain-specific attention heads write to the residual stream, the authors applied singular value decomposition to a head’s combined value-output matrix, WOVh=UΣVTW^h_{OV}=U\Sigma V^T, and projected the right singular vectors through the model’s unembedding matrix to obtain token-level interpretations. In the top singular-value directions they examined, code-specific heads promoted coding- and technology-related tokens, including Apple, iOS, Xcode, and iPhone. Heads most active on non-English inputs promoted tokens from multiple languages while avoiding English tokens; their interpretable outputs also included place names and currency symbols such as Sydney and £. This analysis links the domain specialization observed in component importance to the concepts promoted by those heads.

Coverage note — The supplementary observation that the first period in a text can behave like a BOS token, and the finer addition-versus-subtraction head comparisons, are omitted as exploratory analyses rather than central contributions; the model-family scope limitation is included with the domain results.

References

  1. 1.Jasmijn Bastings and Katja Filippova. 2020. The elephant in the interpretability room: Why use attention as explanation when we have saliency methods? In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 149–155, Online. Association for Computational Linguistics.
  2. 2.Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2019. Identifying and controlling important neurons in neural machine translation. In International Conference on Learning Representations.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  4. 4.Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. 2023. Towards automated circuit discovery for mechanistic interpretability. In Advances in Neural Information Processing Systems, volume 36, pages 16318–16352. Curran Associates, Inc.
  5. 5.Gonçalo M. Correia, Vlad Niculae, and André F. T. Martins. 2019. Adaptively sparse transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2174–2184, Hong Kong, China. Association for Computational Linguistics.
  6. 6.David Dale, Elena Voita, Loic Barrault, and Marta R. Costa-jussà. 2023a. Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity Even better. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 36–50, Toronto, Canada. Association for Computational Linguistics.
  7. 7.David Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loic Barrault, and Marta R. Costa-jussà. 2023b. Halomi: A manually annotated benchmark for multilingual hallucination and omission detection in machine translation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore. Association for Computational Linguistics.
  8. 8.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread.
  9. 9.Javier Ferrando, Gerard I. Gállego, and Marta R. Costa-jussà. 2022. Measuring the mixing of contextual information in the transformer. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8698–8714, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  10. 10.Atticus Geiger, Kyle Richardson, and Christopher Potts. 2020. Neural natural language inference models partially embed theories of lexical entailment and negation. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 163–173, Online. Association for Computational Linguistics.
  11. 11.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022. The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10:522–538.
  12. 12.Nuno M. Guerreiro, Duarte Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, and André F. T. Martins. 2023. Hallucinations in large multilingual translation models. Preprint, arXiv:2303.16104.
  13. 13.Michael Hanna, Ollie Liu, and Alexandre Variengien. 2023. How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model. In Advances in Neural Information Processing Systems, volume 36, pages 76033–76060. Curran Associates, Inc.
  14. 14.Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov. 2024. Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms. Preprint, arXiv:2403.17806.
  15. 15.Roee Hendel, Mor Geva, and Amir Globerson. 2023. In-context learning creates task vectors. Preprint, arXiv:2310.15916.
  16. 16.Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020. Attention is not only a weight: Analyzing transformers with vector norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7057–7075, Online. Association for Computational Linguistics.
  17. 17.János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda. 2024. Atp*: An efficient and scalable method for localizing llm behaviour to components. Preprint, arXiv:2403.00745.
  18. 18.Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg. 2023. The hydra effect: Emergent self-repair in language model computations. Preprint, arXiv:2307.15771.
  19. 19.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. Advances in Neural Information Processing Systems, 36.
  20. 20.Beren Millidge and Sid Black. 2022. The singular value decompositions of transformer weight matrices are highly interpretable.
  21. 21.Raul Molina. 2023. Traveling words: A geometric interpretation of transformers. Preprint, arXiv:2309.07315.
  22. 22.Team NLLB, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No language left behind: Scaling human-centered machine translation. Preprint, arXiv:2207.04672.
  23. 23.Judea Pearl. 2009. Causality, 2 edition. Cambridge University Press.
  24. 24.Slav Petrov, Dipanjan Das, and Ryan McDonald. 2012. A universal part-of-speech tagset. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), pages 2089–2096, Istanbul, Turkey. European Language Resources Association (ELRA).
  25. 25.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(1).
  26. 26.Cody Rushing and Neel Nanda. 2024. Explorations of self-repair in language models. Preprint, arXiv:2402.15390.
  27. 27.Alessandro Stolfo, Yonatan Belinkov, and Mrinmaya Sachan. 2023. Understanding arithmetic reasoning in language models using causal mediation analysis. Preprint, arXiv:2305.15054.
  28. 28.Aaquib Syed, Can Rager, and Arthur Conmy. 2023. Attribution patching outperforms automated circuit discovery. Preprint, arXiv:2310.10348.
  29. 29.Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau. 2023. Function vectors in large language models. Preprint, arXiv:2310.15213.
  30. 30.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. Preprint, arXiv:2302.13971.
  31. 31.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models. Preprint, arXiv:2307.09288.
  32. 32.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  33. 33.Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020. Investigating gender bias in language models using causal mediation analysis. In Advances in Neural Information Processing Systems, volume 33, pages 12388–12401. Curran Associates, Inc.
  34. 34.Elena Voita, Javier Ferrando, and Christoforos Nalmpantis. 2023. Neurons in large language models: Dead, n-gram, positional. Preprint, arXiv:2309.04827.
  35. 35.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5797–5808, Florence, Italy. Association for Computational Linguistics.
  36. 36.Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In The Eleventh International Conference on Learning Representations.

Citation

MLA
Ferrando, J., and E. Voita. “Information Flow Routes: Automatically Interpreting Language Models at Scale”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 17432–45, https://doi.org/10.18653/v1/2024.emnlp-main.965.
APA
Ferrando, J., & Voita, E. (2024). Information Flow Routes: Automatically Interpreting Language Models at Scale. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 17432–17445. https://doi.org/10.18653/v1/2024.emnlp-main.965
Chicago
Ferrando, J., and E. Voita. 2024. “Information Flow Routes: Automatically Interpreting Language Models at Scale”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 17432–45. https://doi.org/10.18653/v1/2024.emnlp-main.965.
Harvard
Ferrando, J. and Voita, E. (2024) “Information Flow Routes: Automatically Interpreting Language Models at Scale”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 17432–17445. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.965.
Vancouver
1. Ferrando J, Voita E (2024) Information Flow Routes: Automatically Interpreting Language Models at Scale. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 17432–17445

BibTeX

@inproceedings{ferrando-voita-2024-information,
    title = "Information Flow Routes: Automatically Interpreting Language Models at Scale",
    author = "Ferrando, Javier  and
      Voita, Elena",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.965/",
    doi = "10.18653/v1/2024.emnlp-main.965",
    pages = "17432--17445"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/