Information Flow Routes: Automatically Interpreting Language Models at Scale
Javier FerrandoElena Voita
Proposes an attribution-based method to automatically extract information flow subgraphs for transformer language model predictions in a single forward pass, providing an interpretation approach roughly 100 times faster than activation patching without requiring human-crafted templates.
Modern artificial intelligence relies heavily on large Transformer-based language models, yet understanding their internal decision-making processes remains a difficult challenge. Existing interpretability techniques largely depend on activation patching, which requires human experts to manually design contrastive prompt templates and execute numerous intervention passes. This manual approach is slow, costly, and limited to narrow, pre-selected test cases.
The article introduces an automated method called information flow routes to efficiently interpret language models at scale. The primary objective is to demonstrate that model predictions can be explained by tracing internal computation subgraphs using an attribution-based approach rather than activation patching.
The researchers formulated the internal representations and operations of Transformer models as directed graphs and applied a top-down attribution algorithm to extract the most relevant pathways for any given prediction. They evaluated the framework on standard mechanistic interpretability benchmarks, such as Indirect Object Identification and Greater-Than tasks, and conducted broad empirical analyses across subsets of the C4 dataset, multilingual text from FLORES-200, code data from CodeParrot, and arithmetic tasks using GPT-2, OPT-125m, and Llama 2 models.
The analysis yielded several key findings. First, the attribution method uncovers internal computational circuits roughly 100 times faster than patching-based methods, reducing execution time from eight minutes to five seconds in benchmark tests. Second, the method evaluates the overall importance of components directly from a single forward pass without suffering from template fragility or self-repair artifacts. Third, across general text, attention heads in the lower layers exhibit universal functional roles, such as previous-token tracking and subword merging. Fourth, component activation strongly aligns with input linguistic properties, notably clustering by parts of speech for function words. Finally, distinct sets of attention heads specialize in specific tasks, such as coding, arithmetic, and non-English languages, directly promoting topic-specific tokens.
These findings indicate that language model computations are modular and can be systematically mapped without burdensome manual intervention. By drastically reducing the computational and human overhead of internal model auditing, this approach enables scalable monitoring of model behaviors, safety mechanisms, and domain-specific operations across large-scale deployments.
The article suggests that organizations should leverage automated attribution graphs to audit model pathways and identify specialized components across various domains. Future work should expand the analysis to fine-grained neuron-level studies in feed-forward layers and investigate whether anomalous internal pathways, such as punctuation tokens inadvertently acting as sentence boundaries, cause generation errors.
The study's primary limitation is that empirical evaluations were conducted on GPT-2, OPT, and Llama 2 model families. While the theoretical framework generalizes to any standard Transformer architecture, stakeholders should maintain moderate caution when applying these observations to models with fundamentally different designs without further validation.
- Paper: Towards Automated Circuit Discovery for Mechanistic Interpretability, Arthur Conmy et al. (2023). ACDC establishes how mechanistic interpretability extracts computational circuits with activation patching, the baseline that Information Flow Routes automates and seeks to avoid.
- Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). Its attention rollout and flow methods show how Transformer information can be traced through layered graphs, a useful foundation for understanding the source’s attribution-based routes.
- Paper: The Dead Salmons of AI Interpretability, Maxime Méloux et al. (2025). This later critique examines the reliability of attribution and circuit-discovery explanations, challenging the confidence readers might place in automated routes and their apparent findings.
