The Linear Representation Hypothesis and the Geometry of Large Language Models

Kiho ParkYo Joong ChoeVictor Veitch

article2024ICML746 citationsBest Paper Award (ICML 2024 Workshop on Mechanistic Interpretability)

Formalizes the linear representation hypothesis using counterfactual pairs to unify linear probing and steering under a causally grounded inner product for large language model representations.

Listen

As large language models become central to enterprise applications, organizations require reliable methods to understand model reasoning and control output behavior. A long-standing assumption in artificial intelligence is the linear representation hypothesis, which suggests that models represent high-level concepts (such as language, tense, or gender) as linear directions in their internal mathematical spaces. However, the precise definition of linear representation has remained ambiguous, spanning distinct concepts such as geometric subspaces, linear measurement probes, and model-steering interventions. Furthermore, standard geometric operations like similarity and projection are mathematically unidentifiable during training, leaving practitioners unsure of how to correctly measure or manipulate these internal vectors.

The article aims to formalize the linear representation hypothesis through a causal framework, resolve geometric ambiguity by introducing a principled inner product, and demonstrate how this structure unifies model interpretation and targeted behavioral control.

To achieve this, the authors develop mathematical proofs using counterfactual concept pairs (pairs of words or phrases that vary only in one specific attribute). They introduce the concept of a "causal inner product," an algebraic metric designed to ensure that causally independent concepts remain mathematically orthogonal (perpendicular). The authors demonstrate that this metric can be computed directly from the inverse covariance matrix of the model's output vocabulary. They validate their theoretical framework empirically using the 7-billion-parameter LLaMA-2 model across 27 distinct linguistic, semantic, and morphological concepts, while also testing comparisons on the Gemma-2B model.

The findings provide strong evidence that high-level concepts are represented linearly and predictably within language models. Across 26 of the 27 evaluated concepts, differences between counterfactual word pairs align consistently along a common directional vector compared to random pairs, confirming the subspace hypothesis. In addition, the estimated causal inner product effectively renders causally separable concepts orthogonal, revealing clear semantic clusters. The mathematical bridge established between input and output representations enables concept directions to function successfully as linear measurement probes without absorbing spurious background correlations. Finally, intervening on the input representation by adding scaled concept vectors systematically alters target outputs (such as promoting the completion "queen" over "king") while leaving unrelated concepts completely untouched.

These results establish that organizations can reliably interpret and steer generative AI models using simple, computationally efficient linear algebra rather than expensive retraining or prompt tuning. By utilizing the causal inner product rather than standard Euclidean metrics—which fail to capture true semantic geometry in models with tied embeddings like Gemma—practitioners gain precise control mechanisms that mitigate safety and bias risks without causing unintended side effects in off-target model outputs.

Leaders and technical teams should consider adopting this causal inner product framework to build lightweight monitoring probes and steering vectors for critical safety and domain-specific tasks. Before wide deployment, organizations should conduct pilot testing on domain-relevant concepts to verify linearity, using simple counterfactual pairs to construct custom intervention vectors.

The findings are supported by consistent mathematical proofs and clear empirical validation, giving high confidence in the core theory. However, minor limitations exist: certain complex concepts (such as whole-to-part relationships) do not exhibit linear representations, multi-token words introduce tokenization noise, and the current study does not evaluate intermediate layer activations or parameter-level interpretability.

Cover for The Linear Representation Hypothesis and the Geometry of Large Language Models

Abstract

Informally, the "linear representation hypothesis" is the idea that high-level concepts are represented linearly as directions in some representation space. In this paper, we address two closely related questions: What does "linear representation" actually mean? And, how do we make sense of geometric notions (e.g., cosine similarity and projection) in the representation space? To answer these, we use the language of counterfactuals to give two formalizations of linear representation, one in the output (word) representation space, and one in the input (context) space. We then prove that these connect to linear probing and model steering, respectively. To make sense of geometric notions, we use the formalization to identify a particular (non-Euclidean) inner product that respects language structure in a sense we make precise. Using this causal inner product, we show how to unify all notions of linear representation. In particular, this allows the construction of probes and steering vectors using counterfactual pairs. Experiments with LLaMA-2 demonstrate the existence of linear representations of concepts, the connection to interpretation and control, and the fundamental role of the choice of inner product. Code is available at github.com/KihoPark/linear_rep_geometry.

Table of Contents

  • 1. Introduction
  • 2. The Linear Representation Hypothesis
  • 2.1. Concepts
  • 2.2. Unembedding Representations and Measurement
  • 2.3. Embedding Representations and Intervention
  • 3. Inner Product for Language Model Representations
  • 3.1. Causal Inner Products
  • 3.2. An Explicit Form for Causal Inner Product
  • 4. Experiments
  • 5. Discussion and Related Work
  • Acknowledgements
  • Impact Statement
  • References
  • A. Summary of Main Results
  • B. Proofs
  • B.1. Proof of Theorem 2.2
  • B.2. Proof of Lemma 2.4
  • B.3. Proof of Theorem 2.5
  • B.4. Proof of Theorem 3.2
  • B.5. Proof of Theorem 3.4
  • C. Experiment Details
  • D. Additional Results
  • D.1. Histograms of random and counterfactual pairs for all concepts
  • D.2. Comparison with the Euclidean inner products
  • D.3. Additional results from the measurement experiment
  • D.4. Additional results from the intervention experiment
  • D.5. Additional tables of top-5 words after intervention
  • D.6. A sanity check for the estimated causal inner product

Knowls

  1. Knowl 1 — Causal concepts and causal separability

    definition

    A language model maps a context xx to an embedding vector λ(x)∈Λ≃Rd\lambda(x)\in\Lambda\simeq\mathbb{R}^d and a possible output token yy to an unembedding vector γ(y)\gamma(y) in an affine space Γ\Gamma whose difference space has dimension dd, with

    P(Y=y∣x)∝exp⁡ ⁣(λ(x)⊤γ(y)).P(Y=y\mid x)\propto \exp\!\big(\lambda(x)^\top\gamma(y)\big).

    A binary concept WW is a latent variable caused by the context and causing the output token. Its values are ordered as 0⇒10\Rightarrow1, and the corresponding counterfactual outputs are denoted Y(W=0)Y(W=0) and Y(W=1)Y(W=1). The paper assumes that the concept value can be read deterministically from the sampled output, so a concept can be specified by its counterfactual output pairs.

    Two concepts WW and ZZ are causally separable when the joint counterfactual output Y(W=w,Z=z)Y(W=w,Z=z) is well-defined for every w,z∈{0,1}w,z\in\{0,1\}; equivalently, the two concepts can be varied freely and in isolation. For example, language and gender are treated as causally separable, whereas two mutually exclusive language changes such as English ⇒\Rightarrow French and English ⇒\Rightarrow Russian are not.

  2. Knowl 2 — Unembedding directions are measurement representations

    theoretical result

    For a binary concept WW with counterfactual outputs Y(0)Y(0) and Y(1)Y(1), define Cone⁡(v)={av:a>0}\operatorname{Cone}(v)=\{av:a>0\} for a nonzero vector vv. An unembedding representation γˉW\bar\gamma_W is a direction satisfying

    γ ⁣(Y(1))−γ ⁣(Y(0))∈Cone⁡(γˉW)almost surely,\gamma\!\big(Y(1)\big)-\gamma\!\big(Y(0)\big)\in\operatorname{Cone}(\bar\gamma_W)\quad\text{almost surely},

    where γ(y)\gamma(y) is the model's unembedding vector for token yy. The positive cone, rather than an unoriented subspace, preserves the ordering 0⇒10\Rightarrow1; when the representation exists, its direction is unique up to positive scaling.

    This subspace definition implies a linear measurement rule. For every context embedding λ∈Λ\lambda\in\Lambda, there is a scalar α>0\alpha>0, depending only on the counterfactual pair {Y(0),Y(1)}\{Y(0),Y(1)\}, such that

    logit⁡P ⁣(Y=Y(1)∣Y∈{Y(0),Y(1)},λ)=α λ⊤γˉW.\operatorname{logit}P\!\left(Y=Y(1)\mid Y\in\{Y(0),Y(1)\},\lambda\right) =\alpha\,\lambda^\top\bar\gamma_W.

    Thus, the unembedding direction of a concept is a logit-linear predictor for that concept across all counterfactual pairs expressing it. Unlike a probe trained on observational data, it does not incorporate correlations with unrelated concepts.

  3. Knowl 3 — Embedding directions are intervention representations

    theoretical result

    Let WW be a binary concept and let ZZ range over concepts causally separable from WW. An embedding representation λˉW∈Λ\bar\lambda_W\in\Lambda is a direction such that, for context embeddings λ0,λ1\lambda_0,\lambda_1, the change from λ0\lambda_0 to λ1\lambda_1 increases the target concept while preserving the relative distribution of every causally separable off-target concept:

    P(W=1∣λ1)P(W=1∣λ0)>1,P(W,Z∣λ1)P(W,Z∣λ0)=P(W∣λ1)P(W∣λ0).\frac{P(W=1\mid\lambda_1)}{P(W=1\mid\lambda_0)}>1, \qquad \frac{P(W,Z\mid\lambda_1)}{P(W,Z\mid\lambda_0)} = \frac{P(W\mid\lambda_1)}{P(W\mid\lambda_0)}.

    The corresponding embedding difference satisfies λ1−λ0∈Cone⁡(λˉW)\lambda_1-\lambda_0\in\operatorname{Cone}(\bar\lambda_W). If γˉW\bar\gamma_W and γˉZ\bar\gamma_Z are unembedding representations, then

    λˉW⊤γˉW>0,λˉW⊤γˉZ=0.\bar\lambda_W^\top\bar\gamma_W>0, \qquad \bar\lambda_W^\top\bar\gamma_Z=0.

    Conversely, under the condition that γˉW\bar\gamma_W together with d−1d-1 causally separable concept directions forms a basis of Rd\mathbb{R}^d, these two orthogonality and positivity conditions characterize the embedding representation up to positive scaling.

    Adding the embedding direction to a context produces a targeted intervention: for any base embedding λ\lambda and scalar c∈Rc\in\mathbb{R},

    P ⁣(Y=Y(W,1)∣Y∈{Y(W,0),Y(W,1)},λ+cλˉW)P\!\left(Y=Y(W,1)\mid Y\in\{Y(W,0),Y(W,1)\},\lambda+c\bar\lambda_W\right)

    is constant in cc, whereas

    P ⁣(Y=Y(1,Z)∣Y∈{Y(0,Z),Y(1,Z)},λ+cλˉW)P\!\left(Y=Y(1,Z)\mid Y\in\{Y(0,Z),Y(1,Z)\},\lambda+c\bar\lambda_W\right)

    increases with cc. Therefore, an embedding representation changes the target concept without changing causally separable concepts.

  4. Knowl 4 — Language-model geometry is not identified by training

    theoretical result

    The language-model softmax distribution is invariant to arbitrary invertible affine changes of the unembedding coordinates. For any invertible matrix A∈Rd×dA\in\mathbb{R}^{d\times d} and constant vector β∈Rd\beta\in\mathbb{R}^d, transform

    γ(y)↦g(y)=Aγ(y)+β,λ(x)↦ℓ(x)=A−⊤λ(x).\gamma(y)\mapsto g(y)=A\gamma(y)+\beta, \qquad \lambda(x)\mapsto \ell(x)=A^{-\top}\lambda(x).

    Because ℓ(x)⊤g(y)=λ(x)⊤γ(y)+λ(x)⊤A−1β\ell(x)^\top g(y)=\lambda(x)^\top\gamma(y)+\lambda(x)^\top A^{-1}\beta, the additional term is independent of yy and cancels in the softmax normalization. Consequently, maximum-likelihood training based only on output probabilities cannot identify the representation coordinates beyond an invertible affine transformation.

    Concept directions transform by γˉW↦AγˉW\bar\gamma_W\mapsto A\bar\gamma_W, so a fixed geometric operation such as a Euclidean inner product is generally not invariant: γˉW⊤γˉZ\bar\gamma_W^\top\bar\gamma_Z need not equal (AγˉW)⊤(AγˉZ)(A\bar\gamma_W)^\top(A\bar\gamma_Z). Semantic claims based on angles, lengths, or projections therefore require an additional choice of inner product.

  5. Knowl 5 — The causal inner product unifies input and output representations

    theoretical result

    Let Γˉ\bar\Gamma be the dd-dimensional vector space of differences between unembedding vectors. A causal inner product is a positive-definite inner product ⟨⋅,⋅⟩C\langle\cdot,\cdot\rangle_C on Γˉ\bar\Gamma satisfying

    ⟨γˉW,γˉZ⟩C=0\langle\bar\gamma_W,\bar\gamma_Z\rangle_C=0

    whenever concepts WW and ZZ are causally separable.

    Suppose that for every concept WW, there are d−1d-1 concepts causally separable from WW whose unembedding directions, together with γˉW\bar\gamma_W, form a basis of Rd\mathbb{R}^d. The Riesz map induced by the causal inner product sends the unembedding representation of WW to its embedding representation:

    ⟨γˉW,γˉ⟩C=λˉW⊤γˉfor every γˉ∈Γˉ.\langle\bar\gamma_W,\bar\gamma\rangle_C =\bar\lambda_W^\top\bar\gamma \qquad\text{for every }\bar\gamma\in\bar\Gamma.

    Thus, after choosing the causal geometry, the input-context and output-token representations of each concept become the same linear object under the induced identification of the two spaces. In coordinates transformed by the positive-definite square root of the inner-product matrix, ordinary Euclidean operations can be used while preserving this causal semantics.

  6. Knowl 6 — The covariance inverse gives an estimable causal geometry

    theoretical result

    Assume that WW and ZZ are causally separable and that an unembedding vector γ\gamma is sampled uniformly from the model vocabulary. For embedding representations λˉW\bar\lambda_W and λˉZ\bar\lambda_Z, assume that the scalar projections λˉW⊤γ\bar\lambda_W^\top\gamma and λˉZ⊤γ\bar\lambda_Z^\top\gamma are independent; the paper notes that uncorrelatedness is sufficient for the result.

    Suppose a causal inner product has matrix form

    ⟨γˉ,γˉ′⟩C=γˉ⊤Mγˉ′,\langle\bar\gamma,\bar\gamma'\rangle_C =\bar\gamma^\top M\bar\gamma',

    where MM is symmetric positive definite. If W1,…,WdW_1,\ldots,W_d are mutually causally separable and their canonical unembedding directions form the basis

    G=[γˉW1,…,γˉWd],G=[\bar\gamma_{W_1},\ldots,\bar\gamma_{W_d}],

    then

    M−1=GG⊤,G⊤Cov⁡(γ)−1G=D,M^{-1}=GG^\top, \qquad G^\top\operatorname{Cov}(\gamma)^{-1}G=D,

    where DD is a diagonal matrix with positive entries and Cov⁡(γ)\operatorname{Cov}(\gamma) is the covariance of uniformly sampled vocabulary unembeddings. The causal-orthogonality constraints leave the diagonal matrix DD undetermined, so the causal inner product is not unique.

    The paper uses the choice D=IdD=I_d, producing the practical estimator

    ⟨γˉ,γˉ′⟩C:=γˉ⊤Cov⁡(γ)−1γˉ′.\langle\bar\gamma,\bar\gamma'\rangle_C :=\bar\gamma^\top\operatorname{Cov}(\gamma)^{-1}\bar\gamma'.

    This covariance-whitened metric can reject most geometries that are compatible with the softmax but fail to respect causal separability.

  7. Knowl 7 — Counterfactual token pairs estimate concept directions

    model/method

    The empirical concept direction for a binary concept WW is estimated from nWn_W counterfactual token pairs (yi(0),yi(1))(y_i(0),y_i(1)) that differ in the intended concept value. Using the covariance-based causal inner product

    ⟨u,v⟩C=u⊤Cov⁡(γ)−1v,\langle u,v\rangle_C=u^\top\operatorname{Cov}(\gamma)^{-1}v,

    the unnormalized direction is the mean token-vector difference

    γ~W=1nW∑i=1nW[γ(yi(1))−γ(yi(0))],\widetilde\gamma_W =\frac{1}{n_W}\sum_{i=1}^{n_W} \left[\gamma\big(y_i(1)\big)-\gamma\big(y_i(0)\big)\right],

    and its canonical normalized form is

    γˉW=γ~W⟨γ~W,γ~W⟩C.\bar\gamma_W =\frac{\widetilde\gamma_W} {\sqrt{\langle\widetilde\gamma_W,\widetilde\gamma_W\rangle_C}}.

    The study uses LLaMA-2-7B, with a 32,000-token vocabulary and 4,096-dimensional token embeddings. It tests 27 concepts: 22 morphological or semantic relations from BATS 3.0, four translation relations (English ⇒\Rightarrow French, French ⇒\Rightarrow German, French ⇒\Rightarrow Spanish, and German ⇒\Rightarrow Spanish), and frequent ⇒\Rightarrow infrequent. Only words represented by single LLaMA-2 tokens are retained. To evaluate alignment, each pair is projected onto a leave-one-out estimate of its concept direction and compared with projections of 100,000 randomly sampled word pairs.

  8. Knowl 8 — Most tested concepts have approximately linear unembedding representations

    empirical result

    In LLaMA-2-7B, the causal projections of counterfactual token differences are substantially more aligned with their corresponding estimated concept directions than are differences between randomly sampled word pairs. This pattern holds across nearly all of the 27 tested morphological, semantic, frequency, and language concepts, supporting the claim that their values are encoded as approximately one-dimensional directions in the unembedding space.

    The main exception is thing ⇒\Rightarrow part: its counterfactual-pair projections do not show the expected separation from random-pair projections, so this concept does not appear to have a reliable linear representation in the unembedding space. The result also supports using the mean of counterfactual differences as an estimator of a concept direction despite tokenization noise and other imperfections in the word pairs.

  9. Knowl 9 — The covariance metric tracks semantic separability better than Euclidean geometry

    empirical result

    For the 27 estimated LLaMA-2 concept directions, most pairs that are causally separable have inner products close to zero under the covariance-based causal metric. The near-zero pattern is not structureless: related concepts form blocks, such as groups of verbal or morphological relations and groups of language relations. Nonzero entries also reflect semantic structure; for example, English ⇒\Rightarrow French is more related to French ⇒\Rightarrow German and French ⇒\Rightarrow Spanish than to German ⇒\Rightarrow Spanish.

    The Euclidean metric happens to make many separable concepts approximately orthogonal in LLaMA-2, but the causal metric gives clearer semantic behavior. In particular, frequent ⇒\Rightarrow infrequent has substantial spurious Euclidean similarity with several separable concepts that becomes much smaller under the causal metric, while the language relations sharing French become more appropriately related.

    The same comparison on Gemma-2B is more discriminating: causally separable concepts remain approximately orthogonal under the covariance-based metric, whereas Euclidean inner products exhibit substantial semantic confounding. This demonstrates that Euclidean geometry is model-dependent and is not a generally reliable substitute for the causal inner product.

  10. Knowl 10 — Concept directions support probing and controlled steering

    empirical result

    For the measurement test, contexts sampled from French and Spanish Wikipedia are scored with the estimated direction for French ⇒\Rightarrow Spanish. The scores γˉW⊤λ(x)\bar\gamma_W^\top\lambda(x) are lower for French contexts and higher for Spanish contexts, as predicted by the logit-linear measurement result. The direction for the unrelated concept male ⇒\Rightarrow female does not separate the French and Spanish contexts, indicating that the effect is not explained by arbitrary concept directions.

    For steering, the paper constructs an input-space direction directly from the output-space estimate:

    λˉW=Cov⁡(γ)−1γˉW,λC,α(x)=λ(x)+αλˉC,\bar\lambda_W=\operatorname{Cov}(\gamma)^{-1}\bar\gamma_W, \qquad \lambda_{C,\alpha}(x)=\lambda(x)+\alpha\bar\lambda_C,

    where CC is the concept being intervened on and α>0\alpha>0 is the intervention strength. Across paired target concepts, increasing α\alpha in the direction of WW raises the target logit while leaving the logit of a causally separable off-target concept approximately unchanged. Intervening in a direction separable from both tested concepts leaves both logits approximately unchanged.

    For the context fragment “Long live the ”, whose original likely completion is “king”, steering in the male ⇒\Rightarrow female direction with α∈{0,0.1,0.2,0.3,0.4}\alpha\in\{0,0.1,0.2,0.3,0.4\} makes the female completion “queen” the top-ranked token by intermediate intervention strength; at α=0.4\alpha=0.4, “queen” is top-1 and the original “king” has fallen below the top five. The other highly ranked completions also increasingly reflect female or royalty-related meanings.

Coverage note — Detailed proof derivations, the full counterfactual-pair count table, auxiliary histograms, and the paper's discussion of extending the framework to parameters or intermediate-layer activations were omitted because they do not add independent load-bearing results beyond the stated theorems and main experiments.

References

  1. 1.Alain, G. and Bengio, Y. Understanding intermediate layers using linear classifier probes. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=ryF7rTqgl.
  2. 2.Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A. A latent variable model approach to PMI-based word embeddings. Transactions of the Association for Computational Linguistics, 4:385–399, 2016.
  3. 3.Belinkov, Y. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219, 2022.
  4. 4.Bowman, S. R., Vilnis, L., Vinyals, O., Dai, A., Jozefowicz, R., and Bengio, S. Generating sentences from a continuous space. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, pp. 10–21, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/K16-1002. URL https://aclanthology.org/K16-1002.
  5. 5.Chang, T., Tu, Z., and Bergen, B. The geometry of multilingual language model representations. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 119–136, 2022.
  6. 6.Chen, B., Fu, Y., Xu, G., Xie, P., Tan, C., Chen, M., and Jing, L. Probing BERT in hyperbolic spaces. In International Conference on Learning Representations, 2021.
  7. 7.Chiang, H.-Y., Camacho-Collados, J., and Pardos, Z. Understanding the source of semantic regularities in word embeddings. In Proceedings of the 24th Conference on Computational Natural Language Learning, pp. 119–131, 2020.
  8. 8.Choe, Y. J., Park, K., and Kim, D. word2word: A collection of bilingual lexicons for 3,564 language pairs. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pp. 3036–3045, 2020.
  9. 9.Drozd, A., Gladkova, A., and Matsuoka, S. Word embeddings, analogies, and machine learning: Beyond king - man + woman = queen. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical papers, pp. 3519–3530, 2016.
  10. 10.Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1, 2021.
  11. 11.Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., et al. Toy models of superposition. arXiv preprint arXiv:2209.10652, 2022.
  12. 12.Ethayarajh, K. How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 55–65, 2019.
  13. 13.Fournier, L., Dupoux, E., and Dunbar, E. Analogies minus analogy test: measuring regularities in word embeddings. In Proceedings of the 24th Conference on Computational Natural Language Learning, pp. 365–375, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.conll-1.29. URL https://aclanthology.org/2020.conll-1.29.
  14. 14.Geva, M., Caciularu, A., Wang, K., and Goldberg, Y. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pp. 30–45, 2022.
  15. 15.Gladkova, A., Drozd, A., and Matsuoka, S. Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn’t. In Proceedings of the NAACL Student Research Workshop, pp. 8–15, 2016.
  16. 16.Goldberg, Y. and Levy, O. word2vec explained: deriving Mikolov et al.’s negative-sampling word-embedding method. arXiv preprint arXiv:1402.3722, 2014.
  17. 17.Gurnee, W. and Tegmark, M. Language models represent space and time. arXiv preprint arXiv:2310.02207, art. arXiv:2310.02207, October 2023. doi: 10.48550/arXiv.2310.02207.
  18. 18.Hendel, R., Geva, M., and Globerson, A. In-context learning creates task vectors. arXiv preprint arXiv:2310.15916, 2023.
  19. 19.Hernandez, E., Sharma, A. S., Haklay, T., Meng, K., Wattenberg, M., Andreas, J., Belinkov, Y., and Bau, D. Linearity of relation decoding in transformer language models. arXiv preprint arXiv:2308.09124, 2023.
  20. 20.Hewitt, J. and Manning, C. D. A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4129–4138, 2019.
  21. 21.Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A. beta-VAE: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations, 2016.
  22. 22.Higgins, I., Amos, D., Pfau, D., Racaniere, S., Matthey, L., Rezende, D., and Lerchner, A. Towards a definition of disentangled representations. arXiv preprint arXiv:1812.02230, 2018.
  23. 23.Hyvarinen, A. and Morioka, H. Unsupervised feature extraction by time-contrastive learning and nonlinear ICA. Advances in Neural Information Processing Systems, 29, 2016.
  24. 24.Jiang, Y., Aragam, B., and Veitch, V. Uncovering meanings of embeddings via partial orthogonality. arXiv preprint arXiv:2310.17611, 2023.
  25. 25.Khemakhem, I., Kingma, D., Monti, R., and Hyvarinen, A. Variational autoencoders and nonlinear ICA: A unifying framework. In International Conference on Artificial Intelligence and Statistics, pp. 2207–2217. PMLR, 2020.
  26. 26.Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV). In International Conference on Machine Learning, pp. 2668–2677. PMLR, 2018.
  27. 27.Kudo, T. and Richardson, J. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 66–71, 2018.
  28. 28.Lample, G., Conneau, A., Ranzato, M., Denoyer, L., and Jegou, H. Word translation without parallel data. In International Conference on Learning Representations, 2018.
  29. 29.Levy, O. and Goldberg, Y. Linguistic regularities in sparse and explicit word representations. In Proceedings of the Eighteenth Conference on Computational Natural Language Learning, pp. 171–180, 2014.
  30. 30.Li, B., Zhou, H., He, J., Wang, M., Yang, Y., and Li, L. On the sentence embeddings from pre-trained language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 9119–9130, 2020.
  31. 31.Li, K., Hopkins, A. K., Bau, D., Viegas, F., Pfister, H., and Wattenberg, M. Emergent world representations: Exploring a sequence model trained on a synthetic task. In International Conference on Learning Representations, 2022.
  32. 32.Meng, K., Bau, D., Andonian, A., and Belinkov, Y. Locating and editing factual associations in GPT. Advances in Neural Information Processing Systems, 35:17359–17372, 2022.
  33. 33.Merullo, J., Eickhoff, C., and Pavlick, E. Language models implement simple word2vec-style vector arithmetic. arXiv preprint arXiv:2305.16130, 2023.
  34. 34.Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Riviere, M., Kale, M. S., Love, J., et al. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024.
  35. 35.Mikolov, T., Le, Q. V., and Sutskever, I. Exploiting similarities among languages for machine translation. arXiv preprint arXiv:1309.4168, 2013a.
  36. 36.Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. Distributed representations of words and phrases and their compositionality. Advances in Neural Information Processing Systems, 26, 2013b.
  37. 37.Mikolov, T., Yih, W.-T., and Zweig, G. Linguistic regularities in continuous space word representations. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 746–751, 2013c.
  38. 38.Mimno, D. and Thompson, L. The strange geometry of skip-gram with negative sampling. In Palmer, M., Hwa, R., and Riedel, S. (eds.), Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pp. 2873–2878, Copenhagen, Denmark, 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1308. URL https://aclanthology.org/D17-1308.
  39. 39.Moran, G. E., Sridhar, D., Wang, Y., and Blei, D. M. Identifiable deep generative models via sparse decoding. arXiv preprint arXiv:2110.10804, art. arXiv:2110.10804, October 2021. doi: 10.48550/arXiv.2110.10804.
  40. 40.Nanda, N., Lee, A., and Wattenberg, M. Emergent linear representations in world models of self-supervised sequence models. arXiv preprint arXiv:2309.00941, 2023.
  41. 41.nostalgebraist. Interpreting GPT: the logit lens, 2020. URL https://www.alignmentforum.org/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens.
  42. 42.OpenAI. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  43. 43.Peng, X., Stevenson, M., Lin, C., and Li, C. Understanding linearity of cross-lingual word embedding mappings. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=8HuyXvbvqX.
  44. 44.Pennington, J., Socher, R., and Manning, C. D. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1532–1543, 2014.
  45. 45.Perera, P., Trager, M., Zancato, L., Achille, A., and Soatto, S. Prompt algebra for task composition. arXiv preprint arXiv:2306.00310, 2023.
  46. 46.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training. 2018.
  47. 47.Reif, E., Yuan, A., Wattenberg, M., Viegas, F. B., Coenen, A., Pearce, A., and Kim, B. Visualizing and measuring the geometry of BERT. Advances in Neural Information Processing Systems, 32, 2019.
  48. 48.Rogers, A., Kovaleva, O., and Rumshisky, A. A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8:842–866, 2021.
  49. 49.Ruder, S., Vulic, I., and Søgaard, A. A survey of cross-lingual word embedding models. Journal of Artificial Intelligence Research, 65:569–631, 2019.
  50. 50.Scholkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y. Toward causal representation learning. Proceedings of the IEEE, 109(5):612–634, 2021.
  51. 51.Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D. Function vectors in large language models. arXiv preprint arXiv:2310.15213, 2023.
  52. 52.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  53. 53.Trager, M., Perera, P., Zancato, L., Achille, A., Bhatia, P., and Soatto, S. Linear spaces of meanings: Compositional structures in vision-language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15395–15404, 2023.
  54. 54.Turner, A. M., Thiergart, L., Udell, D., Leech, G., Mini, U., and MacDiarmid, M. Activation addition: Steering language models without optimization. arXiv preprint arXiv:2308.10248, art. arXiv:2308.10248, August 2023. doi: 10.48550/arXiv.2308.10248.
  55. 55.Ushio, A., Anke, L. E., Schockaert, S., and Camacho-Collados, J. BERT is to NLP what AlexNet is to CV: Can pre-trained language models identify analogies? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 3609–3624, 2021.
  56. 56.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in Neural Information Processing Systems, 30, 2017.
  57. 57.Vylomova, E., Rimell, L., Cohn, T., and Baldwin, T. Take and took, gaggle and goose, book and read: Evaluating the utility of vector differences for lexical relation learning. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1671–1682, 2016.
  58. 58.Wang, Z., Gui, L., Negrea, J., and Veitch, V. Concept algebra for score-based conditional models. arXiv preprint arXiv:2302.03693, 2023.
  59. 59.Zhu, X. and de Melo, G. Sentence analogies: Linguistic regularities in sentence embeddings. In Proceedings of the 28th International Conference on Computational Linguistics, pp. 3389–3400, 2020.
  60. 60.Zimmermann, R. S., Sharma, Y., Schneider, S., Bethge, M., and Brendel, W. Contrastive learning inverts the data generating process. In International Conference on Machine Learning, pp. 12979–12990. PMLR, 2021.
  61. 61.Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, Z., and Hendrycks, D. Representation engineering: A top-down approach to AI transparency. arXiv preprint arXiv:2310.01405, 2023.

Citation

MLA
Park, K., et al. “The Linear Representation Hypothesis and the Geometry of Large Language Models”. In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024, 2023, http://arxiv.org/abs/2311.03658v2.
APA
Park, K., Choe, Y. J., & Veitch, V. (2023). The Linear Representation Hypothesis and the Geometry of Large Language Models. In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024. http://arxiv.org/abs/2311.03658v2
Chicago
Park, K., Y. J. Choe, and V. Veitch. 2023. “The Linear Representation Hypothesis and the Geometry of Large Language Models”. In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024. http://arxiv.org/abs/2311.03658v2.
Harvard
Park, K., Choe, Y.J. and Veitch, V. (2023) “The Linear Representation Hypothesis and the Geometry of Large Language Models”, In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024 [Preprint]. Available at: http://arxiv.org/abs/2311.03658v2.
Vancouver
1. Park K, Choe YJ, Veitch V (2023) The Linear Representation Hypothesis and the Geometry of Large Language Models. In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024

BibTeX

@article{park2023the,
  title = {The Linear Representation Hypothesis and the Geometry of Large Language Models},
  author = {Park, Kiho and Choe, Yo Joong and Veitch, Victor},
  year = {2023},
  journal = {In Proceedings of the 41st International Conference on Machine Learning (ICML), 2024},
  url = {http://arxiv.org/abs/2311.03658v2},
  eprint = {2311.03658}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/