Built independently by an author, for readers. Read the story and support ChapterPal

keyword

linguistic regularities

Linguistic regularities are systematic, predictable patterns in natural language that govern the structural, syntactic, and semantic relationships between words and phrases. In computational linguistics and machine learning, these regularities encompass both grammatical formations, such as tense inflections and pluralization, and conceptual associations, such as gender pairs, spatial relationships, and topical analogies. Within continuous vector representations and language models, these linguistic relationships often manifest as consistent geometric patterns, where pairs of related words exhibit similar vector offsets or linear directions in representation space, enabling relational reasoning and analogy completion through basic vector arithmetic.

5 items

The Linear Representation Hypothesis and the Geometry of Large Language Models

The Linear Representation Hypothesis and the Geometry of Large Language Models

Kiho Park, Yo Joong Choe, Victor Veitch

OrganizationsUniversity of Chicago

Why you should read this

Formalizes the linear representation hypothesis using counterfactual pairs to unify linear probing and steering under a causally grounded inner product for large language model representations.

Informally, the "linear representation hypothesis" is the idea that high-level concepts are represented linearly as directions in some representation space. In this paper, we address two closely related questions: What does "linear representation" actually mean? And, how do we make sense of geometric notions (e.g., cosine similarity and projection) in the representation space? To answer these, we use the language of counterfactuals to give two formalizations of linear representation, one in the output (word) representation space, and one in the input (context) space. We then prove that these connect to linear probing and model steering, respectively. To make sense of geometric notions, we use the formalization to identify a particular (non-Euclidean) inner product that respects language structure in a sense we make precise. Using this causal inner product, we show how to unify all notions of linear representation. In particular, this allows the construction of probes and steering vectors using counterfactual pairs. Experiments with LLaMA-2 demonstrate the existence of linear representations of concepts, the connection to interpretation and control, and the fundamental role of the choice of inner product. Code is available at github.com/KihoPark/linear_rep_geometry.

Added

2026-09-26

Linguistic Regularities in Continuous Space Word Representations

Linguistic Regularities in Continuous Space Word Representations

Tomáš Mikolov, Wen-tau Yih, Geoffrey Zweig

OrganizationsGoogleMicrosoft

Why you should read this

Demonstrates that neural language models encode syntactic and semantic regularities as constant vector offsets, enabling simple vector arithmetic to solve analogy questions and outperform prior methods on relation similarity benchmarks.

Continuous space language models have recently demonstrated outstanding results across a variety of tasks. In this paper, we examine the vector-space word representations that are implicitly learned by the input-layer weights. We find that these representations are surprisingly good at capturing syntactic and semantic regularities in language, and that each relationship is characterized by a relation-specific vector offset. This allows vector-oriented reasoning based on the offsets between words. For example, the male/female relationship is automatically learned, and with the induced vector representations, “King - Man + Woman” results in a vector very close to “Queen.” We demonstrate that the word vectors capture syntactic regularities by means of syntactic analogy questions (provided with this paper), and are able to correctly answer almost 40% of the questions. We demonstrate that the word vectors capture semantic regularities by using the vector offset method to answer SemEval-2012 Task 2 questions. Remarkably, this method outperforms the best previous systems.

Added

2026-09-11

GloVe: Global Vectors for Word Representation

GloVe: Global Vectors for Word Representation

Jeffrey Pennington, Richard Socher, Christopher D. Manning

OrganizationsStanford University

Why you should read this

Demonstrates how to unify the complementary strengths of global statistical methods and local context window approaches into a single model that efficiently captures meaningful semantic relationships in word vectors through weighted co-occurrence statistics rather than sparse matrix factorization or individual context windows.

Recent methods for learning vector space representations of words have succeeded in capturing fine-grained semantic and syntactic regularities using vector arithmetic, but the origin of these regularities has remained opaque. We analyze and make explicit the model properties needed for such regularities to emerge in word vectors. The result is a new global logbilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods. Our model efficiently leverages statistical information by training only on the nonzero elements in a word-word cooccurrence matrix, rather than on the entire sparse matrix or on individual context windowsinalargecorpus. Themodelproduces a vector space with meaningful substructure, as evidenced by its performance of 75% on a recent word analogy task. It also outperforms related models on similarity tasks and named entity recognition.

Added

2026-02-21

Distributed Representations of Words and Phrases and their Compositionality

Distributed Representations of Words and Phrases and their Compositionality

Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, Jeffrey Dean

OrganizationsGoogle

Why you should read this

Demonstrates how to create simple yet powerful word vector representations that capture meaningful relationships between words and phrases, enabling basic arithmetic operations to produce linguistically sensible results while training significantly faster than previous methods.

The recently introduced continuous Skip-gram model is an efficient method for learning high-quality distributed vector representations that capture a large number of precise syntactic and semantic word relationships. In this paper we present several extensions that improve both the quality of the vectors and the training speed. By subsampling of the frequent words we obtain significant speedup and also learn more regular word representations. We also describe a simple alternative to the hierarchical softmax called negative sampling. An inherent limitation of word representations is their indifference to word order and their inability to represent idiomatic phrases. For example, the meanings of "Canada" and "Air" cannot be easily combined to obtain "Air Canada". Motivated by this example, we present a simple method for finding phrases in text, and show that learning good vector representations for millions of phrases is possible.

Added

2026-02-21