The Linear Representation Hypothesis and the Geometry of Large Language Models
Kiho ParkYo Joong ChoeVictor Veitch
Formalizes the linear representation hypothesis using counterfactual pairs to unify linear probing and steering under a causally grounded inner product for large language model representations.
As large language models become central to enterprise applications, organizations require reliable methods to understand model reasoning and control output behavior. A long-standing assumption in artificial intelligence is the linear representation hypothesis, which suggests that models represent high-level concepts (such as language, tense, or gender) as linear directions in their internal mathematical spaces. However, the precise definition of linear representation has remained ambiguous, spanning distinct concepts such as geometric subspaces, linear measurement probes, and model-steering interventions. Furthermore, standard geometric operations like similarity and projection are mathematically unidentifiable during training, leaving practitioners unsure of how to correctly measure or manipulate these internal vectors.
The article aims to formalize the linear representation hypothesis through a causal framework, resolve geometric ambiguity by introducing a principled inner product, and demonstrate how this structure unifies model interpretation and targeted behavioral control.
To achieve this, the authors develop mathematical proofs using counterfactual concept pairs (pairs of words or phrases that vary only in one specific attribute). They introduce the concept of a "causal inner product," an algebraic metric designed to ensure that causally independent concepts remain mathematically orthogonal (perpendicular). The authors demonstrate that this metric can be computed directly from the inverse covariance matrix of the model's output vocabulary. They validate their theoretical framework empirically using the 7-billion-parameter LLaMA-2 model across 27 distinct linguistic, semantic, and morphological concepts, while also testing comparisons on the Gemma-2B model.
The findings provide strong evidence that high-level concepts are represented linearly and predictably within language models. Across 26 of the 27 evaluated concepts, differences between counterfactual word pairs align consistently along a common directional vector compared to random pairs, confirming the subspace hypothesis. In addition, the estimated causal inner product effectively renders causally separable concepts orthogonal, revealing clear semantic clusters. The mathematical bridge established between input and output representations enables concept directions to function successfully as linear measurement probes without absorbing spurious background correlations. Finally, intervening on the input representation by adding scaled concept vectors systematically alters target outputs (such as promoting the completion "queen" over "king") while leaving unrelated concepts completely untouched.
These results establish that organizations can reliably interpret and steer generative AI models using simple, computationally efficient linear algebra rather than expensive retraining or prompt tuning. By utilizing the causal inner product rather than standard Euclidean metrics—which fail to capture true semantic geometry in models with tied embeddings like Gemma—practitioners gain precise control mechanisms that mitigate safety and bias risks without causing unintended side effects in off-target model outputs.
Leaders and technical teams should consider adopting this causal inner product framework to build lightweight monitoring probes and steering vectors for critical safety and domain-specific tasks. Before wide deployment, organizations should conduct pilot testing on domain-relevant concepts to verify linearity, using simple counterfactual pairs to construct custom intervention vectors.
The findings are supported by consistent mathematical proofs and clear empirical validation, giving high confidence in the core theory. However, minor limitations exist: certain complex concepts (such as whole-to-part relationships) do not exhibit linear representations, multi-token words introduce tokenization noise, and the current study does not evaluate intermediate layer activations or parameter-level interpretability.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). Introduces representation engineering and empirical methods for linear probing and representation steering in large language models that the source formalizes geometrically.
- Paper: A Structural Probe for Finding Syntax in Word Representations, John Hewitt et al. (2019). Establishes structural probing techniques that evaluate how linguistic relations and geometric distances are embedded within neural representations.
- Paper: Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham et al. (2023). Demonstrates the extraction of interpretable linear directions and superposition dynamics that underpin linear feature representations in language models.
- Paper: Direct and Indirect Effects, Judea Pearl (2001). Provides the foundational structural counterfactual framework used by the source to define causal notions of linear representation.
- Paper: Linguistic Regularities in Continuous Space Word Representations, Tomáš Mikolov et al. (2013). Pioneered the empirical finding that semantic and syntactic concepts correspond to linear directions in continuous representation spaces.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). Analyzes the non-uniform geometric properties and anisotropy of contextualized word representation spaces across transformer layers.
- Paper: On the Origins of Linear Representations in Large Language Models, Yibo Jiang et al. (2024). Extends the study of linear representation geometry by theoretically analyzing why language model pretraining objectives and optimization dynamics naturally generate linear concept directions.
- Paper: Language Models Are Implicitly Continuous, Samuele Marro et al. (2025). Explores how continuous geometric interpolations in embedding space elicit graded conceptual behaviors in large language models.
- Paper: The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook, Xinlei Yu et al. (2026). Synthesizes modern paradigms of latent-space computation and representation geometry in large language models.
- Paper: Language Models are Injective and Hence Invertible, Giorgos Nikolaou et al. (2025). Investigates the mathematical injectivity and invertibility of transformer hidden representations.
- Paper: Evaluating the World Model Implicit in a Generative Model, Keyon Vafa et al. (2024). Critiques the fidelity of linear probing and investigates whether internal representations constitute coherent world models.
