On the Origins of Linear Representations in Large Language Models
Yibo JiangGoutham RajendranPradeep Kumar RavikumarBryon AragamVictor Veitch
Proves through a latent variable framework that the next-token prediction objective combined with gradient descent's implicit bias mathematically drives large language models to linearly encode high-level semantic concepts.
Large language models often organize high-level semantic concepts into straight, directional lines within their internal mathematical spaces. While this "linear representation" allows practitioners to identify or edit concepts such as gender or tense, it has remained an empirical observation without a clear mathematical foundation. Consequently, researchers and decision-makers could not determine whether these linear representations reflect a genuine structural property of the learning process or merely a coincidental illusion.
The article aims to explain the theoretical origins of linear representations in language models by analyzing how standard training objectives and optimization methods naturally produce linear and orthogonal concept geometry.
To achieve this, the authors introduce a theoretical framework called the latent conditional model, which represents human concepts as underlying binary variables governed by a network of probabilistic dependencies. Under this framework, generating sentences and predicting the next word are modeled as estimating conditional probabilities across these latent concepts. The authors evaluate this setup through mathematical proofs, numerical simulations across varying concept dimensions, and empirical validations using the open-source LLaMA-2 model.
The analysis yields several key findings. First, standard language model training objectives inherently promote linearity via two mechanisms: statistical matching of relative probabilities (log-odds) and the implicit bias of gradient descent optimization, which acts like a maximum-margin classifier driving directional alignment. Second, numerical simulations demonstrate near-perfect linearity and embedding alignment under full context conditions, achieving average cosine similarities between 0.97 and 0.99 across varying variable counts. Third, the framework proves that unrelated concepts naturally arrange themselves at right angles (orthogonally) to one another. Fourth, the core linear behavior persists robustly even when models are trained with missing context combinations, incomplete concept vocabularies, reduced vector dimensions, or alternative optimizers like Adam. Finally, empirical tests on LLaMA-2 confirm theoretical predictions: matching sentence and word concepts exhibited systematic directional alignment (averaging 0.042 with peaks at 0.161) compared to non-matching pairs (averaging 0.011).
These findings provide rigorous theoretical justification for concept-based model steering, safety interventions, and interpretability techniques. They show that linear representations are not arbitrary artifacts of specific neural architectures, but inevitable geometric consequences of training language models via next-token prediction and gradient descent. This lowers the perceived risk of relying on linear steering vectors for alignment, model editing, and auditing.
Organizations developing or deploying language models should confidently leverage linear representation methods for mechanistic interpretability and model safety controls. However, because real-world natural language contains non-linear nuances, teams should avoid treating linear steering as a complete solution and conduct pilot evaluations for complex, multi-dimensional concepts before relying on them in high-stakes governance environments.
Confidence in the mathematical proofs and simulation results is high. Nevertheless, readers should exercise measured caution when applying these insights directly to massive production models, as the theoretical framework relies on simplifying assumptions—such as discrete binary concepts, perfect token-to-concept mappings, and simplified classification subproblems—which result in lower observed cosine similarities in natural text compared to controlled simulations.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). It introduces the empirical practice of representation engineering and linear concept steering in LLMs, providing the foundational phenomenon that the source paper sets out to explain mathematically.
- Paper: Sparse Autoencoders Find Highly Interpretable Features in Language Models, Hoagy Cunningham et al. (2023). It demonstrates how linear feature directions emerge and represent distinct concepts under superposition, directly motivating the source's investigation into why linear concept representations form.
- Paper: Neural Word Embedding as Implicit Matrix Factorization, Omer Levy et al. (2014). It establishes the theoretical link between word prediction objectives and implicit factorization of PMI statistics, laying the mathematical groundwork for explaining linear representations via statistical log-odds matching.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). It provides foundational empirical analyses on the geometric anisotropy and directional properties of contextual embeddings across transformer layers.
- Paper: A Structural Probe for Finding Syntax in Word Representations, John Hewitt et al. (2019). It pioneers linear structural probing in language models, establishing that high-level structures are organized along linear subspaces in representation space.
- Paper: Linguistic Regularities in Continuous Space Word Representations, Tomáš Mikolov et al. (2013). It serves as the original benchmark discovering linear vector offsets and semantic arithmetic in continuous word representation spaces.
- Paper: Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations, Francesco Locatello et al. (2018). It provides critical theoretical background on the impossibility and necessary inductive biases of learning disentangled representations in unsupervised latent spaces.
- Paper: Superposition Yields Robust Neural Scaling, Yizhou Liu et al. (2025). It explores how superposition and geometric overlaps among concept representation vectors mathematically govern scaling laws in large language models.
- Paper: Language Models Are Implicitly Continuous, Samuele Marro et al. (2025). It extends the study of internal representation geometry by analyzing continuous linear interpolations and continuous dynamics in pretrained LLM embedding spaces.
- Paper: The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook, Xinlei Yu et al. (2026). It provides a comprehensive survey synthesizing the mechanisms, capabilities, and continuous representation properties of modern latent spaces in generative AI.
- Paper: Low-dimensional topology of deep neural networks, Junyu Ren et al. (2026). It investigates the underlying topological mechanisms required for neural architectures to transform entangled representations into linearly separable spaces.
