keyword
word embeddings
Word embeddings are continuous, dense vector representations of individual words in natural language processing, designed to map vocabulary terms into a shared geometric space where semantically and syntactically related words reside close to each other. Learned from large text corpora using statistical co-occurrence patterns or neural network architectures, these distributed representations encode linguistic nuances and conceptual relationships across a fixed number of dimensions. By replacing high-dimensional, sparse representations such as one-hot vectors, word embeddings enable computational models to quantify semantic similarity, compute mathematical analogies, and process human language effectively across a wide range of machine learning and text retrieval tasks.
51 items

Language-driven Semantic Segmentation
Boyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun, René Ranftl
Why you should read this
Proposes LSeg, a model that aligns per-pixel visual embeddings with text representations through contrastive learning, enabling zero-shot semantic segmentation of arbitrary unseen categories without additional training.
We present LSeg, a novel model for language-driven semantic image segmentation. LSeg uses a text encoder to compute embeddings of descriptive input labels (e.g., "grass" or "building") together with a transformer-based image encoder that computes dense per-pixel embeddings of the input image. The image encoder is trained with a contrastive objective to align pixel embeddings to the text embedding of the corresponding semantic class. The text embeddings provide a flexible label representation in which semantically similar labels map to similar regions in the embedding space (e.g., "cat" and "furry"). This allows LSeg to generalize to previously unseen categories at test time, without retraining or even requiring a single additional training sample. We demonstrate that our approach achieves highly competitive zero-shot performance compared to existing zero- and few-shot semantic segmentation methods, and even matches the accuracy of traditional segmentation algorithms when a fixed label set is provided. Code and demo are available at this https URL.
Added
2026-10-05

An Open Dataset and Model for Language Identification
Laurie Burchell, Alexandra Birch, Nikolay Bogoychev, Kenneth Heafield
Why you should read this
Presents an open, manually verified dataset of 121 million lines alongside a fastText language identification model that outperforms existing systems across 201 languages.
Language identification (LID) is a fundamental step in many natural language processing pipelines. However, current LID systems are far from perfect, particularly on lower-resource languages. We present a fasttext LID model which achieves a macro-average F1 score of 0.93 and a false positive rate of 0.033 across 201 languages, outperforming previous work. We achieve this by training on a curated dataset of monolingual data, which we audit manually to ensure reliability. We make both the model and the dataset available to the research community. Finally, we carry out detailed analysis into our model’s performance, both in comparison to existing open models and by language class. For applications such as corpus filtering, LID systems need to be fast, reliable, and cover as many languages as possible. There are several open LID models offering quick classification and high language coverage, such as CLD3 or the work of Costa-jussà et al. (2022). However, to the best of our knowledge, none of the commonly-used scalable LID systems make their training data public. We address this gap by releasing an open and curated dataset for LID, which we audit by hand to assure quality. We also train a high-performing LID model on this dataset to show its utility.
Added
2026-10-03

Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP
Lukas Galke, Ansgar Scherp
Why you should read this
Demonstrates that a simple, wide bag-of-words multi-layer perceptron outperforms complex graph neural networks like TextGCN in inductive text classification while offering significantly faster training and inference than Transformer models on long sequences.
Graph neural networks have triggered a resurgence of graph-based text classification methods, defining today’s state of the art. We show that a wide multi-layer perceptron (MLP) using a Bag-of-Words (BoW) outperforms the recent graph-based models TextGCN and HeteGCN in an inductive text classification setting and is comparable with HyperGAT. Moreover, we fine-tune a sequence-based BERT and a lightweight DistilBERT model, which both outperform all state-of-the-art models. These results question the importance of synthetic graphs used in modern text classifiers. In terms of efficiency, DistilBERT is still twice as large as our BoW-based wide MLP, while graph-based models like TextGCN require setting up an O(N²) graph, where N is the vocabulary plus corpus size. Finally, since Transformers need to compute O(L²) attention weights with sequence length L, the MLP models show higher training and inference speeds on datasets with long sequences.
Added
2026-10-03

Benchmarking Intersectional Biases in NLP
John Lalor, Yi Yang, Kendall Smith, Nicole Forsgren, Ahmed Abbasi
Why you should read this
Reveals that standard debiasing techniques fail to prevent amplified allocational harms across intersecting demographic dimensions in downstream natural language processing tasks despite preserving predictive accuracy.
Added
2026-10-03

A Statutory Article Retrieval Dataset in French
Antoine Louis, Gerasimos Spanakis
Why you should read this
Introduces the first native French statutory article retrieval benchmark containing expert-annotated legal questions paired with Belgian law statutes to evaluate dense and lexical information retrieval models on complex legal text.
Statutory article retrieval is the task of automatically retrieving law articles relevant to a legal question. While recent advances in natural language processing have sparked considerable interest in many legal tasks, statutory article retrieval remains primarily untouched due to the scarcity of large-scale and high-quality annotated datasets. To address this bottleneck, we introduce the Belgian Statutory Article Retrieval Dataset (BSARD), which consists of 1,100+ French native legal questions labeled by experienced jurists with relevant articles from a corpus of 22,600+ Belgian law articles. Using BSARD, we benchmark several state-of-the-art retrieval approaches, including lexical and dense architectures, both in zero-shot and supervised setups. We find that fine-tuned dense retrieval models significantly outperform other systems. Our best performing baseline achieves 74.8% R@100, which is promising for the feasibility of the task and indicates there is still room for improvement. By the specificity of the domain and addressed task, BSARD presents a unique challenge problem for future research on legal information retrieval. Our dataset and source code are publicly available.
Added
2026-10-03

Word Embeddings Are Steers for Language Models
Chi Han, Jialiang Xu, Manling Li, Yi Fung, Chenkai Sun, Nan Jiang, Tarek F. Abdelzaher, Heng Ji
Why you should read this
Proposes LM-Steer, a lightweight approach that applies linear transformations to output word embeddings to achieve parameter-efficient, transferable, and continuous style control across language models without retraining internal weights.
Recent studies have shown that large language models (LLMs), when augmented with additional weights through adapters—for instance, LoRA—can learn specialized tasks without updating model parameters—a paradigm known as either Instruction Tuning (IT) for individual instructions or Task Arithmetic (TA) for task-specific adaptation. In such scenarios, different instruction/task datasets induce distinct steering vectors within LLMs' embedding space, shaping behavior according to downstream objectives. Inspired by recent advances in model merging or parameter composition, we investigate whether multiple steering directions (each optimized toward generating responses on one dataset while discouraging outputs outside its distribution via unlearning-like objectives) can be combined into a single vector enabling zero-shot generalization across diverse tasks.
Added
2026-10-02

Linear Adversarial Concept Erasure
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan Cotterell
Why you should read this
Develops a constrained minimax framework and convex relaxation method, R-LACE, to identify and remove linear concept subspaces from neural representations, providing closed-form and tractable solutions for post-hoc bias mitigation without sacrificing downstream performance.
Machine learning models can leak sensitive information about their training data. Concept erasure aims to remove undesirable concepts—such as protected attributes—from representation spaces while preserving performance downstream. We study concept erasure through the lens of adversarial robustness, introducing LEACE, a closed-form method achieving linear-concept erasure provably...
Added
2026-10-02

Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts
Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, Yejin Choi
Why you should read this
Reveals that continuous prompts optimized for language models can be projected onto completely arbitrary or contradictory discrete text without losing task performance, exposing critical flaws in standard prompt interpretability methods.
Fine-tuning continuous prompts for target tasks has recently emerged as a compact alternative to full model fine-tuning. Motivated by these promising results, we investigate the feasibility of extracting a discrete (textual) interpretation of continuous prompts that is faithful to the problem they solve. In practice, we observe a “wayward” behavior between the task solved by continuous prompts and the nearest neighbor discrete projections of these prompts: One can find continuous prompts that solve a task while being projected to an arbitrary text (e.g., definition of a different or even a contradictory task) and simultaneously being within a very small (2%) margin of the best continuous prompt of the same size for the task. We provide intuitions behind this odd and surprising behavior, as well as extensive empirical analyses quantifying the effect of design choices. For instance, larger models exhibit higher waywardness, i.e, we can find prompts that more closely map to any arbitrary text with a smaller drop of accuracy. These findings have important implications relating to the difficulty of faithfully interpreting continuous prompts and their generalization across models and tasks, providing guidance for future progress in prompting language models.
Added
2026-10-02

Text Embeddings Reveal (Almost) As Much As Text
John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. Rush
Why you should read this
Demonstrates that dense text embeddings fail to preserve privacy by introducing an iterative decoding method that reconstructs exact text inputs and extracts sensitive personal information from clinical records.
How much private information do text embeddings reveal about the original text? We investigate the problem of embedding \textit{inversion}, reconstructing the full text represented in dense text embeddings. We frame the problem as controlled generation: generating text that, when reembedded, is close to a fixed point in latent space. We find that although a naïve model conditioned on the embedding performs poorly, a multi-step method that iteratively corrects and re-embeds text is able to recover of text inputs exactly. We train our model to decode text embeddings from two state-of-the-art embedding models, and also show that our model can recover important personal information (full names) from a dataset of clinical notes. Our code is available on Github: \href{this https URL}{this http URL}.
Added
2026-09-28

On the Origins of Linear Representations in Large Language Models
Yibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam, Victor Veitch
Why you should read this
Proves through a latent variable framework that the next-token prediction objective combined with gradient descent's implicit bias mathematically drives large language models to linearly encode high-level semantic concepts.
Recent works have argued that high-level semantic concepts are encoded “linearly” in the representation space of large language models. In this work, we study the origins of such linear representations. To that end, we introduce a simple latent variable model to abstract and formalize the concept dynamics of the next token prediction. We use this formalism to show that the next token prediction objective (softmax with cross-entropy) and the implicit bias of gradient descent together promote the linear representation of concepts. Experiments show that linear representations emerge when learning from data matching the latent variable model, confirming that this simple structure already suffices to yield linear representations. We additionally confirm some predictions of the theory using the LLaMA-2 large language model, giving evidence that the simplified model yields generalizable insights.
Added
2026-09-26

The Linear Representation Hypothesis and the Geometry of Large Language Models
Kiho Park, Yo Joong Choe, Victor Veitch
Why you should read this
Formalizes the linear representation hypothesis using counterfactual pairs to unify linear probing and steering under a causally grounded inner product for large language model representations.
Informally, the "linear representation hypothesis" is the idea that high-level concepts are represented linearly as directions in some representation space. In this paper, we address two closely related questions: What does "linear representation" actually mean? And, how do we make sense of geometric notions (e.g., cosine similarity and projection) in the representation space? To answer these, we use the language of counterfactuals to give two formalizations of linear representation, one in the output (word) representation space, and one in the input (context) space. We then prove that these connect to linear probing and model steering, respectively. To make sense of geometric notions, we use the formalization to identify a particular (non-Euclidean) inner product that respects language structure in a sense we make precise. Using this causal inner product, we show how to unify all notions of linear representation. In particular, this allows the construction of probes and steering vectors using counterfactual pairs. Experiments with LLaMA-2 demonstrate the existence of linear representations of concepts, the connection to interpretation and control, and the fundamental role of the choice of inner product. Code is available at github.com/KihoPark/linear_rep_geometry.
Added
2026-09-26

How Do Transformers Learn Topic Structure: Towards a Mechanistic Understanding
Yuchen Li, Yuanzhi Li, Andrej Risteski
Why you should read this
Proves mathematically and verifies empirically on topic-modeled data how transformer embeddings and self-attention mechanisms independently learn word co-occurrence structures through a distinct two-stage training dynamic.
While the successes of transformers across many domains are indisputable, accurate understanding of the learning mechanics is still largely lacking. Their capabilities have been probed on benchmarks which include a variety of structured and reasoning tasks—but mathematical understanding is lagging substantially behind. Recent lines of work have begun studying representational aspects of this question: that is, the size/depth/complexity of attention-based networks to perform certain tasks. However, there is no guarantee the learning dynamics will converge to the constructions proposed. In our paper, we provide fine-grained mechanistic understanding of how transformers learn “semantic structure”, understood as capturing co-occurrence structure of words. Precisely, we show, through a combination of mathematical analysis and experiments on Wikipedia data and synthetic data modeled by Latent Dirichlet Allocation (LDA), that the embedding layer and the self-attention layer encode the topical structure. In the former case, this manifests as higher average inner product of embeddings between same-topic words. In the latter, it manifests as higher average pairwise attention between same-topic words. The mathematical results involve several assumptions to make the analysis tractable, which we verify on data, and might be of independent interest as well.
Added
2026-09-26

A Sensitivity Analysis of (and Practitioners’ Guide to) Convolutional Neural Networks for Sentence Classification
Ye Zhang, Byron C. Wallace
Why you should read this
Establishes practical guidelines for configuring convolutional neural networks in sentence classification by isolating which hyperparameter choices critically influence model accuracy and which can be safely neglected.
Convolutional Neural Networks (CNNs) have recently achieved remarkably strong performance on the practically important task of sentence classification (kim 2014, kalchbrenner 2014, johnson 2014). However, these models require practitioners to specify an exact model architecture and set accompanying hyperparameters, including the filter region size, regularization parameters, and so on. It is currently unknown how sensitive model performance is to changes in these configurations for the task of sentence classification. We thus conduct a sensitivity analysis of one-layer CNNs to explore the effect of architecture components on model performance; our aim is to distinguish between important and comparatively inconsequential design decisions for sentence classification. We focus on one-layer CNNs (to the exclusion of more complex models) due to their comparative simplicity and strong empirical performance, which makes it a modern standard baseline method akin to Support Vector Machine (SVMs) and logistic regression. We derive practical advice from our extensive empirical results for those interested in getting the most out of CNNs for sentence classification in real world settings.
Added
2026-09-25

SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation
Felix Hill, Roi Reichart, Anna Korhonen
Why you should read this
Introduces SimLex-999, a gold-standard evaluation benchmark that isolates true semantic similarity from conceptual association across diverse parts of speech and concreteness levels, exposing performance gaps in vector space models.
We present SimLex-999, a gold standard resource for evaluating distributional semantic models that improves on existing resources in several important ways. First, in contrast to gold standards such as WordSim-353 and MEN, it explicitly quantifies similarity rather than association or relatedness, so that pairs of entities that are associated but not actually similar [Freud, psychology] have a low rating. We show that, via this focus on similarity, SimLex-999 incentivizes the development of models with a different, and arguably wider range of applications than those which reflect conceptual association. Second, SimLex-999 contains a range of concrete and abstract adjective, noun and verb pairs, together with an independent rating of concreteness and (free) association strength for each pair. This diversity enables fine-grained analyses of the performance of models on concepts of different types, and consequently greater insight into how architectures can be improved. Further, unlike existing gold standard evaluations, for which automatic approaches have reached or surpassed the inter-annotator agreement ceiling, state-of-the-art models perform well below this ceiling on SimLex-999. There is therefore plenty of scope for SimLex-999 to quantify future improvements to distributional semantic models, guiding the development of the next generation of representation-learning architectures.
Added
2026-09-25

Out of One, Many: Using Language Models to Simulate Human Samples
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, David Wingate
Why you should read this
Demonstrates that large language models conditioned on detailed demographic profiles can accurately replicate human survey response distributions across diverse subgroups, establishing artificial intelligence as a viable tool for computational social science research.
We propose and explore the possibility that language models can be studied as effective proxies for specific human sub-populations in social science research. Practical and research applications of artificial intelligence tools have sometimes been limited by problematic biases (such as racism or sexism), which are often treated as uniform properties of the models. We show that the "algorithmic bias" within one such tool -- the GPT-3 language model -- is instead both fine-grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property "algorithmic fidelity" and explore its extent in GPT-3. We create "silicon samples" by conditioning the model on thousands of socio-demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT-3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio-cultural context that characterize human attitudes. We suggest that language models with sufficient algorithmic fidelity thus constitute a novel and powerful tool to advance understanding of humans and society across a variety of disciplines.
Added
2026-09-25

Recurrent Neural Network for Text Classification with Multi-Task Learning
Pengfei Liu, Xipeng Qiu, Xuanjing Huang
Why you should read this
Proposes three multi-task recurrent neural network architectures that share text representations across related tasks to improve classification accuracy when training data is limited.
Neural network based methods have obtained great progress on a variety of natural language processing tasks. However, in most previous works, the models are learned based on single-task supervised objectives, which often suffer from insufficient training data. In this paper, we use the multi-task learning framework to jointly learn across multiple related tasks. Based on recurrent neural network, we propose three different mechanisms of sharing information to model text with task-specific and shared layers. The entire network is trained jointly on all these tasks. Experiments on four benchmark text classification tasks show that our proposed models can improve the performance of a task with the help of other related tasks.
Added
2026-09-25

Improving Distributional Similarity with Lessons Learned from Word Embeddings
Omer Levy, Yoav Goldberg, Ido Dagan
Why you should read this
Demonstrates that the superior performance of neural word embeddings over traditional count-based models stems from hyperparameter optimizations rather than algorithmic differences, proving that applying these same tuning strategies to count-based methods eliminates the performance gap across semantic benchmarks.
Recent trends suggest that neural-network-inspired word embedding models outperform traditional count-based distributional models on word similarity and analogy detection tasks. We reveal that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves. Furthermore, we show that these modifications can be transferred to traditional distributional models, yielding similar gains. In contrast to prior reports, we observe mostly local or insignificant performance differences between the methods, with no global advantage to any single approach over the others.
Added
2026-09-25

Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data
Emily M. Bender, Alexander Koller
Why you should read this
Establishes foundational theoretical limits for large language models by demonstrating that systems trained exclusively on raw text form cannot, in principle, acquire semantic meaning or achieve genuine natural language understanding without grounding in external intent.
The success of the large neural language models on many NLP tasks is exciting. However, we find that these successes sometimes lead to hype in which these models are being described as "understanding" language or capturing "meaning". In this position paper, we argue that a system trained only on form has a priori no way to learn meaning. In keeping with the ACL 2020 theme of "Taking Stock of Where We've Been and Where We're Going", we argue that a clear understanding of the distinction between form and meaning will help guide the field towards better science around natural language understanding.
Added
2026-09-25

A Simple but Tough-to-Beat Baseline for Sentence Embeddings
Sanjeev Arora, Yingyu Liang, Tengyu Ma
Why you should read this
Proposes an unsupervised sentence embedding baseline combining smooth inverse frequency weighting with principal component removal that consistently outperforms complex neural network models on semantic similarity benchmarks.
The success of neural network methods for computing word embeddings has motivated methods for generating semantic embeddings of longer pieces of text, such as sentences and paragraphs. Surprisingly, Wieting et al (ICLR'16) showed that such complicated methods are outperformed, especially in out-of-domain (transfer learning) settings, by simpler methods involving mild retraining of word embeddings and basic linear regression. The method of Wieting et al. requires retraining with a substantial labeled dataset such as Paraphrase Database (Ganitkevitch et al., 2013). The current paper goes further, showing that the following completely unsupervised sentence embedding is a formidable baseline: Use word embeddings computed using one of the popular methods on unlabeled corpus like Wikipedia, represent the sentence by a weighted average of the word vectors, and then modify them a bit using PCA/SVD. This weighting improves performance by about 10% to 30% in textual similarity tasks, and beats sophisticated supervised methods including RNN's and LSTM's. It even improves Wieting et al.'s embeddings. This simple method should be used as the baseline to beat in future, especially when labeled training data is scarce or nonexistent. The paper also gives a theoretical explanation of the success of the above unsupervised method using a latent variable generative model for sentences, which is a simple extension of the model in Arora et al. (TACL'16) with new "smoothing" terms that allow for words occuring out of context, as well as high probabilities for words like and, not in all contexts.
Added
2026-09-25

Deep Graph Kernels
Pinar Yanardag, S.V.N. Vishwanathan
Why you should read this
Introduces a unified framework that adapts deep language modeling techniques to learn continuous sub-structure embeddings, overcoming diagonal dominance and feature sparsity to boost graph classification accuracy across multiple standard kernels.
In this paper, we present Deep Graph Kernels, a unified frame-work to learn latent representations of sub-structures for graphs, inspired by latest advancements in language modeling and deep learning. Our framework leverages the dependency information between sub-structures by learning their latent representations. We demonstrate instances of our framework on three popular graph kernels, namely Graphlet kernels, Weisfeiler-Lehman subtree kernels, and Shortest-Path graph kernels. Our experiments on several benchmark datasets show that Deep Graph Kernels achieve significant improvements in classification accuracy over state-of-the-art graph kernels.
Added
2026-09-25
