On the Origins of Linear Representations in Large Language Models

Yibo JiangGoutham RajendranPradeep Kumar RavikumarBryon AragamVictor Veitch

article2024ICML61 citations

Proves through a latent variable framework that the next-token prediction objective combined with gradient descent's implicit bias mathematically drives large language models to linearly encode high-level semantic concepts.

Listen

Large language models often organize high-level semantic concepts into straight, directional lines within their internal mathematical spaces. While this "linear representation" allows practitioners to identify or edit concepts such as gender or tense, it has remained an empirical observation without a clear mathematical foundation. Consequently, researchers and decision-makers could not determine whether these linear representations reflect a genuine structural property of the learning process or merely a coincidental illusion.

The article aims to explain the theoretical origins of linear representations in language models by analyzing how standard training objectives and optimization methods naturally produce linear and orthogonal concept geometry.

To achieve this, the authors introduce a theoretical framework called the latent conditional model, which represents human concepts as underlying binary variables governed by a network of probabilistic dependencies. Under this framework, generating sentences and predicting the next word are modeled as estimating conditional probabilities across these latent concepts. The authors evaluate this setup through mathematical proofs, numerical simulations across varying concept dimensions, and empirical validations using the open-source LLaMA-2 model.

The analysis yields several key findings. First, standard language model training objectives inherently promote linearity via two mechanisms: statistical matching of relative probabilities (log-odds) and the implicit bias of gradient descent optimization, which acts like a maximum-margin classifier driving directional alignment. Second, numerical simulations demonstrate near-perfect linearity and embedding alignment under full context conditions, achieving average cosine similarities between 0.97 and 0.99 across varying variable counts. Third, the framework proves that unrelated concepts naturally arrange themselves at right angles (orthogonally) to one another. Fourth, the core linear behavior persists robustly even when models are trained with missing context combinations, incomplete concept vocabularies, reduced vector dimensions, or alternative optimizers like Adam. Finally, empirical tests on LLaMA-2 confirm theoretical predictions: matching sentence and word concepts exhibited systematic directional alignment (averaging 0.042 with peaks at 0.161) compared to non-matching pairs (averaging 0.011).

These findings provide rigorous theoretical justification for concept-based model steering, safety interventions, and interpretability techniques. They show that linear representations are not arbitrary artifacts of specific neural architectures, but inevitable geometric consequences of training language models via next-token prediction and gradient descent. This lowers the perceived risk of relying on linear steering vectors for alignment, model editing, and auditing.

Organizations developing or deploying language models should confidently leverage linear representation methods for mechanistic interpretability and model safety controls. However, because real-world natural language contains non-linear nuances, teams should avoid treating linear steering as a complete solution and conduct pilot evaluations for complex, multi-dimensional concepts before relying on them in high-stakes governance environments.

Confidence in the mathematical proofs and simulation results is high. Nevertheless, readers should exercise measured caution when applying these insights directly to massive production models, as the theoretical framework relies on simplifying assumptions—such as discrete binary concepts, perfect token-to-concept mappings, and simplified classification subproblems—which result in lower observed cosine similarities in natural text compared to controlled simulations.

arXiv: 2403.03867
Cover for On the Origins of Linear Representations in Large Language Models

Abstract

Recent works have argued that high-level semantic concepts are encoded “linearly” in the representation space of large language models. In this work, we study the origins of such linear representations. To that end, we introduce a simple latent variable model to abstract and formalize the concept dynamics of the next token prediction. We use this formalism to show that the next token prediction objective (softmax with cross-entropy) and the implicit bias of gradient descent together promote the linear representation of concepts. Experiments show that linear representations emerge when learning from data matching the latent variable model, confirming that this simple structure already suffices to yield linear representations. We additionally confirm some predictions of the theory using the LLaMA-2 large language model, giving evidence that the simplified model yields generalizable insights.

Table of Contents

  • 1. Introduction
  • 2. Problem setting
  • 2.1. Latent conditional model
  • 2.2. Next token prediction
  • 3. Linearity
  • 3.1. Linearity from log-odds
  • 3.2. Linearity from implicit bias of gradient descent
  • 4. Orthogonality
  • 5. Experiments
  • 5.1. Simulated experiments
  • 5.2. Experiments with large language models
  • 6. Literature review
  • 7. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Linearity from log-odds: Proof of Theorem 3.2
  • B. Linearity from log-odds for general MRFs
  • C. Linearity from the implicit bias of gradient descent
  • D. Orthogonality
  • E. Simulated Experiments
  • F. Experiments with large language models

Knowls

  1. Knowl 1 — Latent Conditional Model for Next-Token Prediction in Language Models

    model/method

    The latent conditional framework models the next-token prediction task by assuming context sentences and target tokens share an underlying discrete concept space.

    Let VC={C1,…,Cm}V_C = \{C_1, \dots, C_m\} denote a set of mm binary latent concept random variables, with joint realization c=(c1,…,cm)∈C={0,1}mc = (c_1, \dots, c_m) \in \mathcal{C} = \{0, 1\}^m. Dependencies among concepts are structured as a Markov Random Field over an undirected graph GC=(VC,EC)G_C = (V_C, E_C), satisfying p(Ci∣C[m]∖{i})=p(Ci∣Cne(i))p(C_i \mid C_{[m]\setminus \{i\}}) = p(C_i \mid C_{\text{ne}(i)}) where ne(i)\text{ne}(i) denotes the neighbors of node ii.

    The next token Y∈YY \in \mathcal{Y} is generated via an injective map from concept space C\mathcal{C} to token space Y\mathcal{Y}, with inverse map hy:Y→Ch_y: \mathcal{Y} \to \mathcal{C}.

    A context sentence X∈XX \in \mathcal{X} maps to a context vector in D={⋄,0,1}m\mathcal{D} = \{\diamond, 0, 1\}^m via hx:X→Dh_x: \mathcal{X} \to \mathcal{D}, where ⋄\diamond denotes an unobserved/unknown concept. For a prompt xx, the set of compatible concepts is: Cx={c∈{0,1}m:P(x∣C=c)≠0}C^x = \{c \in \{0, 1\}^m : P(x \mid C = c) \neq 0\} The core concepts fixed by xx are defined as: core(x)={i∈[m]:(ci=1,∀c∈Cx) or (ci=0,∀c∈Cx)}\text{core}(x) = \{i \in [m] : (c_i = 1, \forall c \in C^x) \text{ or } (c_i = 0, \forall c \in C^x)\} The mapping hx(x)h_x(x) assigns the fixed value cic_i for i∈core(x)i \in \text{core}(x) and ⋄\diamond otherwise. The next-token distribution factorizes through the core concepts: p(y∣x)=p(c∣ccore(x))p(y \mid x) = p(c \mid c_{\text{core}(x)}) where c=hy(y)c = h_y(y). A language model estimates this distribution using an embedding function f:D→Rdf: \mathcal{D} \to \mathbb{R}^d and an unembedding function g:C→Rdg: \mathcal{C} \to \mathbb{R}^d via softmax parameterization: p^(c∣d)=exp⁡(f(d)Tg(c))∑c′∈Cexp⁡(f(d)Tg(c′))\hat{p}(c \mid d) = \frac{\exp(f(d)^T g(c))}{\sum_{c' \in \mathcal{C}} \exp(f(d)^T g(c'))}

  2. Knowl 2 — Linearly Encoded and Matched Representations

    definition

    Let C={0,1}m\mathcal{C} = \{0, 1\}^m denote the set of binary latent concept vectors and D={⋄,0,1}m\mathcal{D} = \{\diamond, 0, 1\}^m denote the set of conditioning contexts. For a vector c∈Cc \in \mathcal{C} and t∈{0,1}t \in \{0, 1\}, let c(i→t)c^{(i \to t)} denote the vector identical to cc except that its ii-th entry is set to tt. Pairs (c(i→1),c(i→0))(c^{(i \to 1)}, c^{(i \to 0)}) that differ in only one concept are termed counterfactual pairs, and their difference vector in representation space is called a steering vector. For any vector vv, let Cone(v)={αv:α>0}\text{Cone}(v) = \{\alpha v : \alpha > 0\}.

    Given embedding function f:D^→Rdf: \hat{\mathcal{D}} \to \mathbb{R}^d and unembedding function g:C^→Rdg: \hat{\mathcal{C}} \to \mathbb{R}^d defined on subsets C^⊆C\hat{\mathcal{C}} \subseteq \mathcal{C} and D^⊆D\hat{\mathcal{D}} \subseteq \mathcal{D}:

    1. A latent concept CiC_i has a linearly encoded representation in the unembedding space if there exists a unit vector u∈Rdu \in \mathbb{R}^d such that: g(c(i→1))−g(c(i→0))∈Cone(u)∀c∈C^g(c^{(i \to 1)}) - g(c^{(i \to 0)}) \in \text{Cone}(u) \quad \forall c \in \hat{\mathcal{C}}

    2. A latent concept CiC_i has a linearly encoded representation in the embedding space if there exists a unit vector v∈Rdv \in \mathbb{R}^d such that: f(d(i→1))−f(d(i→0))∈Cone(v)∀d∈D^f(d^{(i \to 1)}) - f(d^{(i \to 0)}) \in \text{Cone}(v) \quad \forall d \in \hat{\mathcal{D}}

    3. The concept CiC_i has a matched-representation if it is linearly encoded in both spaces with identical direction, i.e., u=vu = v.

  3. Knowl 3 — Linearity of Concept Steering Vectors from Matching Log-Odds

    theoretical result

    Let C={0,1}m\mathcal{C} = \{0, 1\}^m and D={⋄,0,1}m\mathcal{D} = \{\diamond, 0, 1\}^m. For any concept index i∈[m]i \in [m], define the steering vector Δc,i=g(c(i→1))−g(c(i→0))\Delta_{c, i} = g(c^{(i \to 1)}) - g(c^{(i \to 0)}) and let Δc,i‾=ΠiΔc,i\overline{\Delta_{c, i}} = \Pi_i \Delta_{c, i}, where Πi\Pi_i is the orthogonal projection operator onto the subspace span{f(d):d∈D,di=⋄}\text{span}\{f(d) : d \in \mathcal{D}, d_i = \diamond\}.

    If the predicted conditional distribution p^(c∣d)=softmax(f(d)Tg(c))\hat{p}(c \mid d) = \text{softmax}(f(d)^T g(c)) satisfies the log-odds matching condition: ln⁡p^(c(i→0)∣d)p^(c(i→1)∣d)=ln⁡p(Ci=0)p(Ci=1)\ln \frac{\hat{p}(c^{(i \to 0)} \mid d)}{\hat{p}(c^{(i \to 1)} \mid d)} = \ln \frac{p(C_i = 0)}{p(C_i = 1)} for all concept vectors c∈Cc \in \mathcal{C} and all contexts d∈Dd \in \mathcal{D} with di=⋄d_i = \diamond, then all projected steering vectors Δc,i‾\overline{\Delta_{c, i}} over all c∈Cc \in \mathcal{C} are parallel to each other. In particular, this condition holds identically when all latent concept variables C1,…,CmC_1, \dots, C_m are jointly independent under the true distribution pp.

  4. Knowl 4 — Subspace Dimension Bound for Markov Random Field Latent Concepts

    theoretical result

    Let concepts C1,…,CmC_1, \dots, C_m follow a Markov Random Field with graph GC=(VC,EC)G_C = (V_C, E_C) and neighborhood set ne(i)\text{ne}(i) for each concept i∈[m]i \in [m]. Let Πi\Pi_i denote the orthogonal projection operator onto span{f(d):d∈D,di=⋄,dj≠⋄  ∀j∈ne(i)}\text{span}\{f(d) : d \in \mathcal{D}, d_i = \diamond, d_j \neq \diamond \; \forall j \in \text{ne}(i)\}, and define the projected steering vector Δc,i‾=Πi(g(c(i→1))−g(c(i→0)))\overline{\Delta_{c, i}} = \Pi_i (g(c^{(i \to 1)}) - g(c^{(i \to 0)})).

    If for a fixed concept i∈[m]i \in [m], any c∈Cc \in \mathcal{C}, and any context d∈Dd \in \mathcal{D} with di=⋄d_i = \diamond and dj≠⋄d_j \neq \diamond for all j∈ne(i)j \in \text{ne}(i), the model satisfies: ln⁡p^(c(i→0)∣d)p^(c(i→1)∣d)=ln⁡p(Ci=0∣Cne(i)=dne(i))p(Ci=1∣Cne(i)=dne(i))\ln \frac{\hat{p}(c^{(i \to 0)} \mid d)}{\hat{p}(c^{(i \to 1)} \mid d)} = \ln \frac{p(C_i = 0 \mid C_{\text{ne}(i)} = d_{\text{ne}(i)})}{p(C_i = 1 \mid C_{\text{ne}(i)} = d_{\text{ne}(i)})} then the set of all projected steering vectors {Δc,i‾:c∈C}\{\overline{\Delta_{c, i}} : c \in \mathcal{C}\} spans a linear subspace of dimension at most 2∣ne(i)∣2^{|\text{ne}(i)|}.

  5. Knowl 5 — Emergence of Unembedding Linearity via Gradient Descent with Fixed Embeddings

    theoretical result

    Let concept index i∈[m]i \in [m], d⋄=[⋄,…,⋄]d^\diamond = [\diamond, \dots, \diamond], and conditioning set D^={d(i→1)⋄,d(i→0)⋄}\hat{\mathcal{D}} = \{d^\diamond_{(i \to 1)}, d^\diamond_{(i \to 0)}\}. Let Δc,i=g(c(i→1))−g(c(i→0))\Delta_{c, i} = g(c^{(i \to 1)}) - g(c^{(i \to 0)}). Consider training the unembeddings gg under fixed embeddings ff (with f(d(i→1)⋄)≠f(d(i→0)⋄)f(d^\diamond_{(i \to 1)}) \neq f(d^\diamond_{(i \to 0)})) by minimizing the subproblem loss: L({Δc,i}c∈C,f(d(i→1)⋄),f(d(i→0)⋄))=∑c∈C(exp⁡(−Δc,iTf(d(i→1)⋄))+exp⁡(Δc,iTf(d(i→0)⋄)))L(\{\Delta_{c, i}\}_{c \in \mathcal{C}}, f(d^\diamond_{(i \to 1)}), f(d^\diamond_{(i \to 0)})) = \sum_{c \in \mathcal{C}} \left(\exp(-\Delta_{c, i}^T f(d^\diamond_{(i \to 1)})) + \exp(\Delta_{c, i}^T f(d^\diamond_{(i \to 0)}))\right) where ci=1c_i = 1 for all c∈Cc \in \mathcal{C}.

    When optimizing gg via gradient descent with an appropriate step size, the steering vectors for the same concept become asymptotically parallel: lim⁡t→∞⟨Δc1,it,Δc2,it⟩∥Δc1,it∥∥Δc2,it∥=1∀c1,c2∈C\lim_{t \to \infty} \frac{\langle \Delta_{c^1, i}^t, \Delta_{c^2, i}^t \rangle}{\|\Delta_{c^1, i}^t\| \|\Delta_{c^2, i}^t\|} = 1 \quad \forall c^1, c^2 \in \mathcal{C} where the superscript tt denotes parameter values after tt iterations.

  6. Knowl 6 — Asymptotic Collinearity of Two Vectors Under Exponential Loss Minimization

    theoretical result

    Consider the exponential loss function L(u,v)=exp⁡(−uTv)L(u, v) = \exp(-u^T v) optimized over vectors u,v∈Rdu, v \in \mathbb{R}^d using gradient descent with step size η<1L(u0,v0)\eta < \frac{1}{L(u_0, v_0)}.

    If the initialization satisfies u0≠−αv0u_0 \neq -\alpha v_0 for all α>0\alpha > 0, the gradient descent iterates (ut,vt)(u_t, v_t) satisfy:

    1. lim⁡t→∞L(ut,vt)=0\lim_{t \to \infty} L(u_t, v_t) = 0
    2. lim⁡t→∞∥ut∥=∞\lim_{t \to \infty} \|u_t\| = \infty and lim⁡t→∞∥vt∥=∞\lim_{t \to \infty} \|v_t\| = \infty
    3. The cosine similarity cos⁡(ut,vt)=⟨ut,vt⟩∥ut∥∥vt∥\cos(u_t, v_t) = \frac{\langle u_t, v_t \rangle}{\|u_t\| \|v_t\|} increases monotonically with iteration tt
    4. lim⁡t→∞cos⁡(ut,vt)=1\lim_{t \to \infty} \cos(u_t, v_t) = 1
  7. Knowl 7 — Joint Alignment of Embedding and Unembedding Steering Vectors by Gradient Descent

    theoretical result

    Let i∈[m]i \in [m], d⋄=[⋄,…,⋄]d^\diamond = [\diamond, \dots, \diamond], Δc,i=g(c(i→1))−g(c(i→0))\Delta_{c, i} = g(c^{(i \to 1)}) - g(c^{(i \to 0)}), and consider the exponential subproblem objective: L({Δc,i}c∈C,f(d(i→1)⋄),f(d(i→0)⋄))=∑c∈C(exp⁡(−Δc,iTf(d(i→1)⋄))+exp⁡(Δc,iTf(d(i→0)⋄)))L(\{\Delta_{c, i}\}_{c \in \mathcal{C}}, f(d^\diamond_{(i \to 1)}), f(d^\diamond_{(i \to 0)})) = \sum_{c \in \mathcal{C}} \left(\exp(-\Delta_{c, i}^T f(d^\diamond_{(i \to 1)})) + \exp(\Delta_{c, i}^T f(d^\diamond_{(i \to 0)}))\right) Suppose that at initialization (t=0t = 0), all elements in the set {Δc,i}c∈C∪{f(d(i→1)⋄),f(d(i→0)⋄)}\{\Delta_{c, i}\}_{c \in \mathcal{C}} \cup \{f(d^\diamond_{(i \to 1)}), f(d^\diamond_{(i \to 0)})\} are mutually orthogonal, with ∥f(d(i→1)⋄)∥=∥f(d(i→0)⋄)∥\|f(d^\diamond_{(i \to 1)})\| = \|f(d^\diamond_{(i \to 0)})\| and ∥Δc,i∥\|\Delta_{c, i}\| equal across all cc.

    Then optimizing both ff and gg jointly via gradient descent yields asymptotic alignment within the unembedding space and across the embedding and unembedding spaces: lim⁡t→∞cos⁡(Δc1,it,Δc2,it)=1∀c1,c2∈C\lim_{t \to \infty} \cos(\Delta_{c^1, i}^t, \Delta_{c^2, i}^t) = 1 \quad \forall c^1, c^2 \in \mathcal{C} lim⁡t→∞cos⁡(Δc1,it,ft(d(i→1)⋄)−ft(d(i→0)⋄))=1∀c1∈C\lim_{t \to \infty} \cos(\Delta_{c^1, i}^t, f^t(d^\diamond_{(i \to 1)}) - f^t(d^\diamond_{(i \to 0)})) = 1 \quad \forall c^1 \in \mathcal{C}

  8. Knowl 8 — Orthogonality of Representations for Graph-Separated Concepts

    theoretical result

    Let C={0,1}m\mathcal{C} = \{0, 1\}^m, D={⋄,0,1}m\mathcal{D} = \{\diamond, 0, 1\}^m, and assume true distribution p(c)>0p(c) > 0 for all c∈Cc \in \mathcal{C}. Let CiC_i and CjC_j be two latent variables separated in the Markov Random Field graph GCG_C.

    For any binary vector c∈Cc \in \mathcal{C}, define Dc={d∈D:di=⋄,p(c∣d)>0}D_c = \{d \in \mathcal{D} : d_i = \diamond, p(c \mid d) > 0\}. If the model perfectly estimates the conditional distribution such that p^(c∣d)=p(c∣d)\hat{p}(c \mid d) = p(c \mid d) for all d∈Dcd \in D_c, then: (g(c(i→1))−g(c(i→0)))⊥(f(d(j→cj))−f(d(j→⋄)))∀d∈Dc(g(c^{(i \to 1)}) - g(c^{(i \to 0)})) \perp (f(d^{(j \to c_j)}) - f(d^{(j \to \diamond)})) \quad \forall d \in D_c Furthermore, if CiC_i and CjC_j admit matched representations (i.e., embedding and unembedding steering directions coincide), their unembedding steering vectors are orthogonal: (g(c(i→1))−g(c(i→0)))⊥(g(c(j→1))−g(c(j→0)))∀c∈C(g(c^{(i \to 1)}) - g(c^{(i \to 0)})) \perp (g(c^{(j \to 1)}) - g(c^{(j \to 0)})) \quad \forall c \in \mathcal{C}

  9. Knowl 9 — Emergence of Linearity and Matched Representations Across Latent Dimensions in Simulation

    data/table

    To evaluate the emergence of linear representations, models were trained on data sampled from random directed acyclic graphs (DAGs) with m∈{3,4,5,6,7}m \in \{3, 4, 5, 6, 7\} variables. Conditional distributions were parameterized by Bernoulli distributions with parameters sampled uniformly from [0.3,0.7][0.3, 0.7]. Models parameterized ff and gg as linear lookup tables mapping one-hot encodings to Rd\mathbb{R}^d with representation dimension d=md = m, trained with SGD (learning rate 0.10.1, batch size 100100) using cross-entropy loss on randomly masked vectors.

    The table below reports the average cosine similarities among steering vectors in the unembedding space, among steering vectors in the embedding space, and between unembedding and embedding steering vectors for the same concept:

    mm UNEMBEDDING EMBEDDING UNEMBEDDING AND EMBEDDING
    3 0.972±0.0060.972 \pm 0.006 0.982±0.0050.982 \pm 0.005 0.980±0.0050.980 \pm 0.005
    4 0.975±0.0050.975 \pm 0.005 0.971±0.0050.971 \pm 0.005 0.973±0.0050.973 \pm 0.005
    5 0.988±0.0040.988 \pm 0.004 0.981±0.0040.981 \pm 0.004 0.984±0.0040.984 \pm 0.004
    6 0.997±0.0000.997 \pm 0.000 0.985±0.0020.985 \pm 0.002 0.990±0.0010.990 \pm 0.001
    7 0.995±0.0010.995 \pm 0.001 0.972±0.0040.972 \pm 0.004 0.981±0.0030.981 \pm 0.003

    Standard errors are computed over 100 runs for m=3,4m=3, 4, 50 runs for m=5m=5, 20 runs for m=6m=6, and 10 runs for m=7m=7. The near-unity values confirm that the next-token prediction objective induces both linear encoding and matching alignment between embedding and unembedding spaces.

  10. Knowl 10 — Alignment Between Context Differences and Token Steering Vectors in LLaMA-2

    empirical result

    Representational alignment was evaluated on the pre-trained LLaMA-2-7B model across two natural language benchmarks:

    1. Multilingual Context Pairs: Using sentence pairs from OPUS Books across 4 language pairs (French--Spanish, French--German, English--French, German--Spanish), embedding steering vectors were computed as the average difference between representations of equivalent sentences in two languages. Unembedding steering vectors were computed across 27 grammatical and semantic token concepts. The cosine similarity between embedding differences for a language translation concept and its matching unembedding concept vector was highest compared to non-matching concepts (e.g., French--Spanish context embedding differences had absolute cosine similarity ≈0.34\approx 0.34 with the French ⇒\Rightarrow Spanish token steering vector, significantly higher than with unrelated grammatical concepts).

    2. Winograd Schema Pairs: Sentence pairs differing in 1--2 words that resolve pronoun ambiguity were evaluated. Embedding steering vectors were formed from the first differing token between pairs. Across pairs, matching context-concept pairs achieved an average cosine similarity of 0.0420.042 (maximum 0.1610.161), compared to 0.0110.011 (maximum 0.0810.081) for non-matching pairs, confirming that context embedding differences align preferentially with unembedding steering vectors of the corresponding concept.

  11. Knowl 11 — Robustness of Linear Representations to Incomplete Data and Optimizers

    empirical result

    In synthetic latent graphical model experiments, linear and matched representations proved robust under several non-ideal training conditions:

    1. Choice of Optimizer: Training with Adam (learning rate 0.0010.001) on the complete conditional set produced cosine similarities ranging from 0.9100.910 to 0.9960.996 across m=3m=3 to 77 latent variables.

    2. Incomplete Conditioning Contexts (∣D^∣<∣D∣|\hat{\mathcal{D}}| < |\mathcal{D}|): Restricting the model to subsets of conditioning masks (e.g., maximum 50 or 100 masks for m=10m=10 variables) maintained unembedding cosine similarities of 0.974±0.0090.974 \pm 0.009 (50 masks) and 0.957±0.0090.957 \pm 0.009 (100 masks).

    3. Incomplete Concept Vectors (∣C^∣<∣C∣|\hat{\mathcal{C}}| < |\mathcal{C}|): Introducing zero-probability concept configurations via rejection sampling yielded unembedding cosine similarities of 0.951±0.0110.951 \pm 0.011 for 10 variables and 0.896±0.0110.896 \pm 0.011 for 12 variables.

    4. Dimension Reduction: Decreasing the representation dimension dd below the number of latent variables mm (e.g., evaluating d<7d < 7 for 7 variables, and d<10d < 10 for 10 variables) resulted in only minor reductions in steering vector cosine similarity.

Coverage note — None was omitted; all key theoretical definitions, log-odds theorems, gradient descent alignment proofs, orthogonality results, and empirical simulations / LLM validations were distilled into self-contained knowls.

References

  1. 1.Allen, C. and Hospedales, T. Analogies explained: Towards understanding word embeddings. In International Conference on Machine Learning, pp. 223–231. PMLR, 2019.
  2. 2.Allen, C., Balazevic, I., and Hospedales, T. What the vec? towards probabilistically grounded embeddings. Advances in neural information processing systems, 32, 2019.
  3. 3.Arora, S., Ge, R., Halpern, Y., Mimno, D. M., Moitra, A., Sontag, D. A., Wu, Y., and Zhu, M. A practical algorithm for topic modeling with provable guarantees. In Proceedings of the 30th International Conference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16-21 June 2013, volume 28 of JMLR Workshop and Conference Proceedings, pp. 280–288. JMLR.org, 2013.
  4. 4.Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A. Random walks on context spaces: Towards an explanation of the mysteries of semantic word embeddings. arXiv preprint arXiv:1502.03520, pp. 385–399, 2015.
  5. 5.Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A. A latent variable model approach to pmi-based word embeddings. Transactions of the Association for Computational Linguistics, 4:385–399, 2016.
  6. 6.Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A. Linear algebraic structure of word senses, with applications to polysemy. Transactions of the Association for Computational Linguistics, 6:483–495, 2018.
  7. 7.Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A. Network dissection: Quantifying interpretability of deep visual representations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6541–6549, 2017.
  8. 8.Blei, D. M. and Lafferty, J. D. Dynamic topic models. In Proceedings of the 23rd international conference on Machine learning, pp. 113–120, 2006.
  9. 9.Buchholz, S., Rajendran, G., Rosenfeld, E., Aragam, B., Schölkopf, B., and Ravikumar, P. Learning linear causal representations from interventions under general nonlinear mixing. arXiv preprint arXiv:2306.02235, 2023.
  10. 10.Burns, C., Ye, H., Klein, D., and Steinhardt, J. Discovering latent knowledge in language models without supervision. arXiv preprint arXiv:2212.03827, 2022.
  11. 11.Chang, T. A., Tu, Z., and Bergen, B. K. The geometry of multilingual language model representations. arXiv preprint arXiv:2205.10964, 2022.
  12. 12.Chen, B., Fu, Y., Xu, G., Xie, P., Tan, C., Chen, M., and Jing, L. Probing bert in hyperbolic spaces. arXiv preprint arXiv:2104.03869, 2021.
  13. 13.Dasgupta, S. Learning mixtures of gaussians. In 40th Annual Symposium on Foundations of Computer Science (Cat. No. 99CB37039), pp. 634–644. IEEE, 1999.
  14. 14.Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., et al. Toy models of superposition. arXiv preprint arXiv:2209.10652, 2022.
  15. 15.Engel, J., Hoffman, M., and Roberts, A. Latent constraints: Learning to generate conditionally from unconditional generative models. arXiv preprint arXiv:1711.05772, 2017.
  16. 16.Ethayarajh, K., Duvenaud, D., and Hirst, G. Towards understanding linear word analogies. arXiv preprint arXiv:1810.04882, 2018.
  17. 17.Falck, F., Zhang, H., Willetts, M., Nicholson, G., Yau, C., and Holmes, C. C. Multi-facet clustering variational autoencoders. Advances in Neural Information Processing Systems, 34, 2021.
  18. 18.Frandsen, A. and Ge, R. Understanding composition of word embeddings via tensor decomposition. arXiv preprint arXiv:1902.00613, 2019.
  19. 19.Gittens, A., Achlioptas, D., and Mahoney, M. W. Skip-gram - zipf + uniform = vector additivity. In Annual Meeting of the Association for Computational Linguistics, 2017a.
  20. 20.Gittens, A., Achlioptas, D., and Mahoney, M. W. Skip-gram- zipf+ uniform= vector additivity. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 69–76, 2017b.
  21. 21.Gladkova, A., Drozd, A., and Matsuoka, S. Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn’t. In Proceedings of the NAACL Student Research Workshop, pp. 8–15, 2016.
  22. 22.Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D. Finding neurons in a haystack: Case studies with sparse probing. arXiv preprint arXiv:2305.01610, 2023.
  23. 23.Hyvärinen, A., Khemakhem, I., and Monti, R. Identifiability of latent-variable and structural-equation models: from linear to nonlinear. arXiv preprint arXiv:2302.02672, 2023.
  24. 24.Jiang, Y. and Aragam, B. Learning latent causal graphs with unknown interventions. In Advances in Neural Information Processing Systems, 2023.
  25. 25.Jiang, Y., Aragam, B., and Veitch, V. Uncovering meanings of embeddings via partial orthogonality. arXiv preprint arXiv:2310.17611, 2023.
  26. 26.Khemakhem, I., Kingma, D., Monti, R., and Hyvärinen, A. Variational autoencoders and nonlinear ica: A unifying framework. In International Conference on Artificial Intelligence and Statistics, pp. 2207–2217. PMLR, 2020.
  27. 27.Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International conference on machine learning, pp. 2668–2677. PMLR, 2018.
  28. 28.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In Bengio, Y. and LeCun, Y. (eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015.
  29. 29.Kivva, B., Rajendran, G., Ravikumar, P., and Aragam, B. Learning latent causal graphs via mixture oracles. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 18087–18101, 2021.
  30. 30.Kivva, B., Rajendran, G., Ravikumar, P., and Aragam, B. Identifiability of deep generative models without auxiliary information. Advances in Neural Information Processing Systems, 35:15687–15701, 2022.
  31. 31.Koller, D. and Friedman, N. Probabilistic graphical models: principles and techniques. MIT press, 2009.
  32. 32.Kushner, H. and Yin, G. G. Stochastic approximation and recursive algorithms and applications, volume 35. Springer Science & Business Media, 2003.
  33. 33.Lachapelle, S., Rodríguez, P., Sharma, Y., Everett, K., Priol, R. L., Lacoste, A., and Lacoste-Julien, S. Disentanglement via mechanism sparsity regularization: A new principle for nonlinear ICA. In Schölkopf, B., Uhler, C., and Zhang, K. (eds.), 1st Conference on Causal Learning and Reasoning, CLeaR 2022, Sequoia Conference Center, Eureka, CA, USA, 11-13 April, 2022, volume 177 of Proceedings of Machine Learning Research, pp. 428–484. PMLR, 2022.
  34. 34.Levesque, H., Davis, E., and Morgenstern, L. The winograd schema challenge. In Thirteenth International Conference on the Principles of Knowledge Representation and Reasoning. Citeseer, 2012.
  35. 35.Li, B., Zhou, H., He, J., Wang, M., Yang, Y., and Li, L. On the sentence embeddings from pre-trained language models. arXiv preprint arXiv:2011.05864, 2020.
  36. 36.Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M. Inference-time intervention: Eliciting truthful answers from a language model. arXiv preprint arXiv:2306.03341, 2023.
  37. 37.Liang, V. W., Zhang, Y., Kwon, Y., Yeung, S., and Zou, J. Y. Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning. Advances in Neural Information Processing Systems, 35: 17612–17625, 2022.
  38. 38.McGrath, T., Kapishnikov, A., Tomašev, N., Pearce, A., Wattenberg, M., Hassabis, D., Kim, B., Paquet, U., and Kramnik, V. Acquisition of chess knowledge in alphazero. Proceedings of the National Academy of Sciences, 119 (47):e2206625119, 2022.
  39. 39.Mikolov, T., Yih, W.-t., and Zweig, G. Linguistic regularities in continuous space word representations. In Proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: Human language technologies, pp. 746–751, 2013.
  40. 40.Mimno, D. and Thompson, L. The strange geometry of skipgram with negative sampling. In Conference on Empirical Methods in Natural Language Processing, 2017.
  41. 41.Moschella, L., Maiorca, V., Fumero, M., Norelli, A., Locatello, F., and Rodolà, E. Relative representations enable zero-shot latent space communication. arXiv preprint arXiv:2209.15430, 2022.
  42. 42.Nanda, N., Lee, A., and Wattenberg, M. Emergent linear representations in world models of self-supervised sequence models. arXiv preprint arXiv:2309.00941, 2023.
  43. 43.OpenAI. GPT-4 technical report, 2023.
  44. 44.Park, K., Choe, Y. J., and Veitch, V. The linear representation hypothesis and the geometry of large language models, 2023.
  45. 45.Pearl, J. Causality. Cambridge university press, 2009.
  46. 46.Pennington, J., Socher, R., and Manning, C. D. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp. 1532–1543, 2014.
  47. 47.Radford, A., Metz, L., and Chintala, S. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434, 2015.
  48. 48.Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J. Svcca: Singular vector canonical correlation analysis for deep understanding and improvement. stat, 1050:19, 2017.
  49. 49.Rajendran, G., Kivva, B., Gao, M., and Aragam, B. Structure learning in polynomial time: Greedy algorithms, bregman information, and exponential families. Advances in Neural Information Processing Systems, 34:18660–18672, 2021.
  50. 50.Rajendran, G., Reizinger, P., Brendel, W., and Ravikumar, P. An interventional perspective on identifiability in gaussian lti systems with independent component analysis. arXiv preprint arXiv:2311.18048, 2023.
  51. 51.Rajendran, G., Buchholz, S., Aragam, B., Schölkopf, B., and Ravikumar, P. Learning interpretable concepts: Unifying causal representation learning and foundation models. arXiv preprint, 2024.
  52. 52.Reif, E., Yuan, A., Wattenberg, M., Viegas, F. B., Coenen, A., Pearce, A., and Kim, B. Visualizing and measuring the geometry of bert. Advances in Neural Information Processing Systems, 32, 2019.
  53. 53.Ri, N., Lee, F.-T., and Verma, N. Contrastive loss is all you need to recover analogies as parallel lines. arXiv preprint arXiv:2306.08221, 2023.
  54. 54.Rudolph, M. and Blei, D. Dynamic bernoulli embeddings for language evolution. arXiv preprint arXiv:1703.08052, 2017.
  55. 55.Rudolph, M., Ruiz, F., Mandt, S., and Blei, D. Exponential family embeddings. Advances in Neural Information Processing Systems, 29, 2016.
  56. 56.Schölkopf, B. and von Kügelgen, J. From statistical to causal learning. arXiv preprint arXiv:2204.00607, 2022.
  57. 57.Schölkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y. Toward causal representation learning. Proceedings of the IEEE, 109(5): 612–634, 2021. arXiv:2102.11107.
  58. 58.Schut, L., Tomasev, N., McGrath, T., Hassabis, D., Paquet, U., and Kim, B. Bridging the human-ai knowledge gap: Concept discovery and transfer in alphazero. arXiv preprint arXiv:2310.16410, 2023.
  59. 59.Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N. The implicit bias of gradient descent on separable data. The Journal of Machine Learning Research, 19(1):2822–2878, 2018.
  60. 60.Spirtes, P., Glymour, C. N., and Scheines, R. Causation, prediction, and search. MIT press, 2000.
  61. 61.Squires, C. and Uhler, C. Causal structure learning: a combinatorial perspective. Foundations of Computational Mathematics, pp. 1–35, 2022.
  62. 62.Tiedemann, J. Parallel data, tools and interfaces in opus. In Chair), N. C. C., Choukri, K., Declerck, T., Dogan, M. U., Maegaard, B., Mariani, J., Odijk, J., and Piperidis, S. (eds.), Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC’12), Istanbul, Turkey, may 2012. European Language Resources Association (ELRA). ISBN 978-2-9517408-7-7.
  63. 63.Tigges, C., Hollinsworth, O. J., Geiger, A., and Nanda, N. Linear representations of sentiment in large language models. arXiv preprint arXiv:2310.15154, 2023.
  64. 64.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  65. 65.Trager, M., Perera, P., Zancato, L., Achille, A., Bhatia, P., and Soatto, S. Linear spaces of meanings: compositional structures in vision-language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15395–15404, 2023.
  66. 66.Varici, B., Acarturk, E., Shanmugam, K., Kumar, A., and Tajer, A. Score-based causal representation learning with interventions. arXiv preprint arXiv:2301.08230, 2023.
  67. 67.Volpi, R. and Malagò, L. Evaluating natural alpha embeddings on intrinsic and extrinsic tasks. In Workshop on Representation Learning for NLP, 2020.
  68. 68.Volpi, R. and Malagò, L. Natural alpha embeddings. Information Geometry, 4(1):3–29, 2021.
  69. 69.Wang, Z., Gui, L., Negrea, J., and Veitch, V. Concept algebra for score-based conditional model. In ICML 2023 Workshop on Structured Probabilistic Inference {&} Generative Modeling, 2023.
  70. 70.Wu, J., Zou, D., Braverman, V., and Gu, Q. Direction matters: On the implicit bias of stochastic gradient descent with moderate learning rate. arXiv preprint arXiv:2011.02538, 2020.
  71. 71.Zimmermann, R. S., Sharma, Y., Schneider, S., Bethge, M., and Brendel, W. Contrastive learning inverts the data generating process. In International Conference on Machine Learning, pp. 12979–12990. PMLR, 2021.

Citation

MLA
Jiang, Y., et al. “On the Origins of Linear Representations in Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2403.03867v1.
APA
Jiang, Y., Rajendran, G., Ravikumar, P., Aragam, B., & Veitch, V. (2024). On the Origins of Linear Representations in Large Language Models. arXiv. http://arxiv.org/abs/2403.03867v1
Chicago
Jiang, Y., G. Rajendran, P. Ravikumar, B. Aragam, and V. Veitch. 2024. “On the Origins of Linear Representations in Large Language Models”. arXiv. http://arxiv.org/abs/2403.03867v1.
Harvard
Jiang, Y. et al. (2024) “On the Origins of Linear Representations in Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.03867v1.
Vancouver
1. Jiang Y, Rajendran G, Ravikumar P, Aragam B, Veitch V (2024) On the Origins of Linear Representations in Large Language Models. arXiv

BibTeX

@article{jiang2024the,
  title = {On the Origins of Linear Representations in Large Language Models},
  author = {Jiang, Yibo and Rajendran, Goutham and Ravikumar, Pradeep and Aragam, Bryon and Veitch, Victor},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.03867v1},
  eprint = {2403.03867}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/