Built independently by an author, for readers. Read the story and support ChapterPal

keyword

word embeddings

Word embeddings are continuous, dense vector representations of individual words in natural language processing, designed to map vocabulary terms into a shared geometric space where semantically and syntactically related words reside close to each other. Learned from large text corpora using statistical co-occurrence patterns or neural network architectures, these distributed representations encode linguistic nuances and conceptual relationships across a fixed number of dimensions. By replacing high-dimensional, sparse representations such as one-hot vectors, word embeddings enable computational models to quantify semantic similarity, compute mathematical analogies, and process human language effectively across a wide range of machine learning and text retrieval tasks.

51 items

Language-driven Semantic Segmentation

Language-driven Semantic Segmentation

Boyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun, René Ranftl

OrganizationsAppleCornell UniversityIntelUniversity of Copenhagen

Why you should read this

Proposes LSeg, a model that aligns per-pixel visual embeddings with text representations through contrastive learning, enabling zero-shot semantic segmentation of arbitrary unseen categories without additional training.

We present LSeg, a novel model for language-driven semantic image segmentation. LSeg uses a text encoder to compute embeddings of descriptive input labels (e.g., "grass" or "building") together with a transformer-based image encoder that computes dense per-pixel embeddings of the input image. The image encoder is trained with a contrastive objective to align pixel embeddings to the text embedding of the corresponding semantic class. The text embeddings provide a flexible label representation in which semantically similar labels map to similar regions in the embedding space (e.g., "cat" and "furry"). This allows LSeg to generalize to previously unseen categories at test time, without retraining or even requiring a single additional training sample. We demonstrate that our approach achieves highly competitive zero-shot performance compared to existing zero- and few-shot semantic segmentation methods, and even matches the accuracy of traditional segmentation algorithms when a fixed label set is provided. Code and demo are available at this https URL.

Added

2026-10-05

An Open Dataset and Model for Language Identification

An Open Dataset and Model for Language Identification

Laurie Burchell, Alexandra Birch, Nikolay Bogoychev, Kenneth Heafield

OrganizationsSchool of InformaticsUniversity of Edinburgh

Why you should read this

Presents an open, manually verified dataset of 121 million lines alongside a fastText language identification model that outperforms existing systems across 201 languages.

Language identification (LID) is a fundamental step in many natural language processing pipelines. However, current LID systems are far from perfect, particularly on lower-resource languages. We present a fasttext LID model which achieves a macro-average F1 score of 0.93 and a false positive rate of 0.033 across 201 languages, outperforming previous work. We achieve this by training on a curated dataset of monolingual data, which we audit manually to ensure reliability. We make both the model and the dataset available to the research community. Finally, we carry out detailed analysis into our model’s performance, both in comparison to existing open models and by language class. For applications such as corpus filtering, LID systems need to be fast, reliable, and cover as many languages as possible. There are several open LID models offering quick classification and high language coverage, such as CLD3 or the work of Costa-jussà et al. (2022). However, to the best of our knowledge, none of the commonly-used scalable LID systems make their training data public. We address this gap by releasing an open and curated dataset for LID, which we audit by hand to assure quality. We also train a high-performing LID model on this dataset to show its utility.

Added

2026-10-03

Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP

Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP

Lukas Galke, Ansgar Scherp

OrganizationsMax Planck Institute for PsycholinguisticsUlm UniversityUniversity of Kiel

Why you should read this

Demonstrates that a simple, wide bag-of-words multi-layer perceptron outperforms complex graph neural networks like TextGCN in inductive text classification while offering significantly faster training and inference than Transformer models on long sequences.

Graph neural networks have triggered a resurgence of graph-based text classification methods, defining today’s state of the art. We show that a wide multi-layer perceptron (MLP) using a Bag-of-Words (BoW) outperforms the recent graph-based models TextGCN and HeteGCN in an inductive text classification setting and is comparable with HyperGAT. Moreover, we fine-tune a sequence-based BERT and a lightweight DistilBERT model, which both outperform all state-of-the-art models. These results question the importance of synthetic graphs used in modern text classifiers. In terms of efficiency, DistilBERT is still twice as large as our BoW-based wide MLP, while graph-based models like TextGCN require setting up an O(N²) graph, where N is the vocabulary plus corpus size. Finally, since Transformers need to compute O(L²) attention weights with sequence length L, the MLP models show higher training and inference speeds on datasets with long sequences.

Added

2026-10-03

A Statutory Article Retrieval Dataset in French

A Statutory Article Retrieval Dataset in French

Antoine Louis, Gerasimos Spanakis

OrganizationsLaw & Tech LabMaastricht University

Why you should read this

Introduces the first native French statutory article retrieval benchmark containing expert-annotated legal questions paired with Belgian law statutes to evaluate dense and lexical information retrieval models on complex legal text.

Statutory article retrieval is the task of automatically retrieving law articles relevant to a legal question. While recent advances in natural language processing have sparked considerable interest in many legal tasks, statutory article retrieval remains primarily untouched due to the scarcity of large-scale and high-quality annotated datasets. To address this bottleneck, we introduce the Belgian Statutory Article Retrieval Dataset (BSARD), which consists of 1,100+ French native legal questions labeled by experienced jurists with relevant articles from a corpus of 22,600+ Belgian law articles. Using BSARD, we benchmark several state-of-the-art retrieval approaches, including lexical and dense architectures, both in zero-shot and supervised setups. We find that fine-tuned dense retrieval models significantly outperform other systems. Our best performing baseline achieves 74.8% R@100, which is promising for the feasibility of the task and indicates there is still room for improvement. By the specificity of the domain and addressed task, BSARD presents a unique challenge problem for future research on legal information retrieval. Our dataset and source code are publicly available.

Added

2026-10-03

Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts

Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts

Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, Yejin Choi

Why you should read this

Reveals that continuous prompts optimized for language models can be projected onto completely arbitrary or contradictory discrete text without losing task performance, exposing critical flaws in standard prompt interpretability methods.

Fine-tuning continuous prompts for target tasks has recently emerged as a compact alternative to full model fine-tuning. Motivated by these promising results, we investigate the feasibility of extracting a discrete (textual) interpretation of continuous prompts that is faithful to the problem they solve. In practice, we observe a “wayward” behavior between the task solved by continuous prompts and the nearest neighbor discrete projections of these prompts: One can find continuous prompts that solve a task while being projected to an arbitrary text (e.g., definition of a different or even a contradictory task) and simultaneously being within a very small (2%) margin of the best continuous prompt of the same size for the task. We provide intuitions behind this odd and surprising behavior, as well as extensive empirical analyses quantifying the effect of design choices. For instance, larger models exhibit higher waywardness, i.e, we can find prompts that more closely map to any arbitrary text with a smaller drop of accuracy. These findings have important implications relating to the difficulty of faithfully interpreting continuous prompts and their generalization across models and tasks, providing guidance for future progress in prompting language models.

Added

2026-10-02

The Linear Representation Hypothesis and the Geometry of Large Language Models

The Linear Representation Hypothesis and the Geometry of Large Language Models

Kiho Park, Yo Joong Choe, Victor Veitch

OrganizationsUniversity of Chicago

Why you should read this

Formalizes the linear representation hypothesis using counterfactual pairs to unify linear probing and steering under a causally grounded inner product for large language model representations.

Informally, the "linear representation hypothesis" is the idea that high-level concepts are represented linearly as directions in some representation space. In this paper, we address two closely related questions: What does "linear representation" actually mean? And, how do we make sense of geometric notions (e.g., cosine similarity and projection) in the representation space? To answer these, we use the language of counterfactuals to give two formalizations of linear representation, one in the output (word) representation space, and one in the input (context) space. We then prove that these connect to linear probing and model steering, respectively. To make sense of geometric notions, we use the formalization to identify a particular (non-Euclidean) inner product that respects language structure in a sense we make precise. Using this causal inner product, we show how to unify all notions of linear representation. In particular, this allows the construction of probes and steering vectors using counterfactual pairs. Experiments with LLaMA-2 demonstrate the existence of linear representations of concepts, the connection to interpretation and control, and the fundamental role of the choice of inner product. Code is available at github.com/KihoPark/linear_rep_geometry.

Added

2026-09-26

How Do Transformers Learn Topic Structure: Towards a Mechanistic Understanding

How Do Transformers Learn Topic Structure: Towards a Mechanistic Understanding

Yuchen Li, Yuanzhi Li, Andrej Risteski

OrganizationsCarnegie Mellon UniversityMicrosoft

Why you should read this

Proves mathematically and verifies empirically on topic-modeled data how transformer embeddings and self-attention mechanisms independently learn word co-occurrence structures through a distinct two-stage training dynamic.

While the successes of transformers across many domains are indisputable, accurate understanding of the learning mechanics is still largely lacking. Their capabilities have been probed on benchmarks which include a variety of structured and reasoning tasks—but mathematical understanding is lagging substantially behind. Recent lines of work have begun studying representational aspects of this question: that is, the size/depth/complexity of attention-based networks to perform certain tasks. However, there is no guarantee the learning dynamics will converge to the constructions proposed. In our paper, we provide fine-grained mechanistic understanding of how transformers learn “semantic structure”, understood as capturing co-occurrence structure of words. Precisely, we show, through a combination of mathematical analysis and experiments on Wikipedia data and synthetic data modeled by Latent Dirichlet Allocation (LDA), that the embedding layer and the self-attention layer encode the topical structure. In the former case, this manifests as higher average inner product of embeddings between same-topic words. In the latter, it manifests as higher average pairwise attention between same-topic words. The mathematical results involve several assumptions to make the analysis tractable, which we verify on data, and might be of independent interest as well.

Added

2026-09-26

A Sensitivity Analysis of (and Practitioners’ Guide to) Convolutional Neural Networks for Sentence Classification

A Sensitivity Analysis of (and Practitioners’ Guide to) Convolutional Neural Networks for Sentence Classification

Ye Zhang, Byron C. Wallace

OrganizationsUniversity of Texas at Austin

Why you should read this

Establishes practical guidelines for configuring convolutional neural networks in sentence classification by isolating which hyperparameter choices critically influence model accuracy and which can be safely neglected.

Convolutional Neural Networks (CNNs) have recently achieved remarkably strong performance on the practically important task of sentence classification (kim 2014, kalchbrenner 2014, johnson 2014). However, these models require practitioners to specify an exact model architecture and set accompanying hyperparameters, including the filter region size, regularization parameters, and so on. It is currently unknown how sensitive model performance is to changes in these configurations for the task of sentence classification. We thus conduct a sensitivity analysis of one-layer CNNs to explore the effect of architecture components on model performance; our aim is to distinguish between important and comparatively inconsequential design decisions for sentence classification. We focus on one-layer CNNs (to the exclusion of more complex models) due to their comparative simplicity and strong empirical performance, which makes it a modern standard baseline method akin to Support Vector Machine (SVMs) and logistic regression. We derive practical advice from our extensive empirical results for those interested in getting the most out of CNNs for sentence classification in real world settings.

Added

2026-09-25

SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation

SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation

Felix Hill, Roi Reichart, Anna Korhonen

OrganizationsTechnion – Israel Institute of TechnologyUniversity of Cambridge

Why you should read this

Introduces SimLex-999, a gold-standard evaluation benchmark that isolates true semantic similarity from conceptual association across diverse parts of speech and concreteness levels, exposing performance gaps in vector space models.

We present SimLex-999, a gold standard resource for evaluating distributional semantic models that improves on existing resources in several important ways. First, in contrast to gold standards such as WordSim-353 and MEN, it explicitly quantifies similarity rather than association or relatedness, so that pairs of entities that are associated but not actually similar [Freud, psychology] have a low rating. We show that, via this focus on similarity, SimLex-999 incentivizes the development of models with a different, and arguably wider range of applications than those which reflect conceptual association. Second, SimLex-999 contains a range of concrete and abstract adjective, noun and verb pairs, together with an independent rating of concreteness and (free) association strength for each pair. This diversity enables fine-grained analyses of the performance of models on concepts of different types, and consequently greater insight into how architectures can be improved. Further, unlike existing gold standard evaluations, for which automatic approaches have reached or surpassed the inter-annotator agreement ceiling, state-of-the-art models perform well below this ceiling on SimLex-999. There is therefore plenty of scope for SimLex-999 to quantify future improvements to distributional semantic models, guiding the development of the next generation of representation-learning architectures.

Added

2026-09-25

Out of One, Many: Using Language Models to Simulate Human Samples

Out of One, Many: Using Language Models to Simulate Human Samples

Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, David Wingate

OrganizationsBrigham Young University

Why you should read this

Demonstrates that large language models conditioned on detailed demographic profiles can accurately replicate human survey response distributions across diverse subgroups, establishing artificial intelligence as a viable tool for computational social science research.

We propose and explore the possibility that language models can be studied as effective proxies for specific human sub-populations in social science research. Practical and research applications of artificial intelligence tools have sometimes been limited by problematic biases (such as racism or sexism), which are often treated as uniform properties of the models. We show that the "algorithmic bias" within one such tool -- the GPT-3 language model -- is instead both fine-grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property "algorithmic fidelity" and explore its extent in GPT-3. We create "silicon samples" by conditioning the model on thousands of socio-demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT-3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio-cultural context that characterize human attitudes. We suggest that language models with sufficient algorithmic fidelity thus constitute a novel and powerful tool to advance understanding of humans and society across a variety of disciplines.

Added

2026-09-25

A Simple but Tough-to-Beat Baseline for Sentence Embeddings

A Simple but Tough-to-Beat Baseline for Sentence Embeddings

Sanjeev Arora, Yingyu Liang, Tengyu Ma

OrganizationsPrinceton University

Why you should read this

Proposes an unsupervised sentence embedding baseline combining smooth inverse frequency weighting with principal component removal that consistently outperforms complex neural network models on semantic similarity benchmarks.

The success of neural network methods for computing word embeddings has motivated methods for generating semantic embeddings of longer pieces of text, such as sentences and paragraphs. Surprisingly, Wieting et al (ICLR'16) showed that such complicated methods are outperformed, especially in out-of-domain (transfer learning) settings, by simpler methods involving mild retraining of word embeddings and basic linear regression. The method of Wieting et al. requires retraining with a substantial labeled dataset such as Paraphrase Database (Ganitkevitch et al., 2013). The current paper goes further, showing that the following completely unsupervised sentence embedding is a formidable baseline: Use word embeddings computed using one of the popular methods on unlabeled corpus like Wikipedia, represent the sentence by a weighted average of the word vectors, and then modify them a bit using PCA/SVD. This weighting improves performance by about 10% to 30% in textual similarity tasks, and beats sophisticated supervised methods including RNN's and LSTM's. It even improves Wieting et al.'s embeddings. This simple method should be used as the baseline to beat in future, especially when labeled training data is scarce or nonexistent. The paper also gives a theoretical explanation of the success of the above unsupervised method using a latent variable generative model for sentences, which is a simple extension of the model in Arora et al. (TACL'16) with new "smoothing" terms that allow for words occuring out of context, as well as high probabilities for words like and, not in all contexts.

Added

2026-09-25