Built independently by an author, for readers. Read the story and support ChapterPal

keyword

sentence embeddings

Sentence embeddings are dense numerical vector representations of entire sentences that capture their underlying semantic and syntactic meaning in a continuous vector space. Unlike word-level representations that model individual words in isolation, sentence embeddings map complete sentences to fixed-length vectors that preserve overall context, structure, and intent. These vectors are typically generated using pretrained neural networks, contrastive learning objectives, or pooling and weighting mechanisms over lexical representations. By placing sentences with similar meanings close to one another in a shared geometric space, sentence embeddings allow machines to efficiently compute semantic similarity using vector distance metrics and support a wide range of natural language processing tasks, such as dense retrieval, clustering, text classification, and semantic textual similarity.

25 items

The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models

The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language Models

Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, Shauli Ravfogel

OrganizationsBar-Ilan UniversityGoogle

Why you should read this

Reveals that large language models internally encode whether a question is answerable in their hidden states even while generating hallucinatory answers, showing that this linearly separable answerability signal can be extracted to curb overconfident errors.

Large language models (LLMs) have been shown to possess impressive capabilities, while also raising crucial concerns about the faithfulness of their responses. A primary issue arising in this context is the management of (un)answerable queries by LLMs, which often results in hallucinatory behavior due to overconfidence. In this paper, we explore the behavior of LLMs when presented with (un)answerable queries. We ask: do models represent the fact that the question is (un)answerable when generating a hallucinatory answer? Our results show strong indications that such models encode the answerability of an input query, with the representation of the first decoded token often being a strong indicator. These findings shed new light on the spatial organization within the latent representations of LLMs, unveiling previously unexplored facets of these models. Moreover, they pave the way for the development of improved decoding techniques with better adherence to factual generation, particularly in scenarios where query (un)answerability is a concern.

Added

2026-10-03

Data Curation Alone Can Stabilize In-context Learning

Data Curation Alone Can Stabilize In-context Learning

Ting-Yun Chang, Robin Jia

OrganizationsUniversity of Southern California

Why you should read this

Demonstrates that selecting high-value training subsets through individual example scoring significantly reduces performance variance and increases accuracy in in-context learning without requiring dynamic prompt retrieval or model calibration.

In-context learning (ICL) enables large language models (LLMs) to perform new tasks by prompting them with a sequence of training examples. However, it is known that ICL is very sensitive to the choice of training examples: randomly sampling examples from a training set leads to high variance in performance. In this paper, we show that carefully curating a subset of training data greatly stabilizes ICL performance without any other changes to the ICL algorithm (e.g., prompt retrieval or calibration). We introduce two methods to choose training subsets—both score training examples individually, then select the highest-scoring ones. CONDACC scores a training example by its average dev-set ICL accuracy when combined with random training examples, while DATAMODELS learns linear regressors that estimate how the presence of each training example influences LLM outputs. Across five tasks and two LLMs, sampling from stable subsets selected by CONDACC and DATAMODELS improves average accuracy over sampling from the entire training set by 7.7% and 6.3%, respectively. Surprisingly, the stable subset examples are not especially diverse in content or low in perplexity, in contrast with other work suggesting that diversity and perplexity are important when prompting LLMs.

Added

2026-10-02

A Closer Look at How Fine-tuning Changes BERT

A Closer Look at How Fine-tuning Changes BERT

Yichu Zhou, Vivek Srikumar

OrganizationsSchool of ComputingUniversity of Utah

Why you should read this

Reveals how fine-tuning improves BERT representations by widening separation distances between distinct label clusters while preserving the model's underlying geometric structure across downstream tasks.

Given the prevalence of pre-trained contextualized representations in today's NLP, there have been many efforts to understand what information they contain, and why they seem to be universally successful. The most common approach to use these representations involves fine-tuning them for an end task. Yet, how fine-tuning changes the underlying embedding space is less studied. In this work, we study the English BERT family and use two probing techniques to analyze how fine-tuning changes the space. We hypothesize that fine-tuning affects classification performance by increasing the distances between examples associated with different labels. We confirm this hypothesis with carefully designed experiments on five different NLP tasks. Via these experiments, we also discover an exception to the prevailing wisdom that "fine-tuning always improves performance". Finally, by comparing the representations before and after fine-tuning, we discover that fine-tuning does not introduce arbitrary changes to representations; instead, it adjusts the representations to downstream tasks while largely preserving the original spatial structure of the data points.

Added

2026-10-02

GPT-RE: In-context Learning for Relation Extraction using Large Language Models

GPT-RE: In-context Learning for Relation Extraction using Large Language Models

Zhen Wan, Fei Cheng, Zhuoyuan Mao, Qianying Liu, Haiyue Song, Jiwei Li, Sadao Kurohashi

OrganizationsKyoto UniversityZhejiang University

Why you should read this

Proposes GPT-RE, a framework that bridges the performance gap between large language models and fully supervised baselines on relation extraction benchmarks by combining task-aware demonstration retrieval with gold-label-induced reasoning.

In spite of the potential for ground-breaking achievements offered by large language models (LLMs) (e.g., GPT-3) via in-context learning (ICL), they still lag significantly behind fully-supervised baselines (e.g., fine-tuned BERT) in relation extraction (RE). This is due to the two major shortcomings of ICL for RE: (1) low relevance regarding entity and relation in existing sentence-level demonstration retrieval approaches for ICL; and (2) the lack of explaining input-label mappings of demonstrations leading to poor ICL effectiveness. In this paper, we propose GPT-RE to successfully address the aforementioned issues by (1) incorporating task-aware representations in demonstration retrieval; and (2) enriching the demonstrations with gold label-induced reasoning logic. We evaluate GPT-RE on four widely-used RE datasets and observe that GPT-RE achieves improvements over not only existing GPT-3 baselines, but also fully-supervised baselines as in Figure 1. Specifically, GPT-RE achieves SOTA performances on the SemEval and SciERC datasets, and competitive performances on the TACRED and ACE05 datasets. Additionally, a critical issue of LLMs revealed by previous work, the strong inclination to wrongly classify NULL examples into other pre-defined labels, is substantially alleviated by our method. We show an empirical analysis.1

Added

2026-09-26

Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives

Unsupervised Sentence Representation via Contrastive Learning with Mixing Negatives

Yanzhao Zhang, Richong Zhang, Samuel Mensah, Xudong Liu, Yongyi Mao

OrganizationsAdvanced Innovation Center for Big Data and Brain ComputingBeihang UniversitySKLSDEUniversity of OttawaUniversity of Sheffield

Why you should read this

Proposes MixCSE, a contrastive sentence representation framework that overcomes vanishing gradient signals by continually generating artificial hard negative features through mixing positive and negative samples, achieving state-of-the-art results on semantic textual similarity and transfer tasks.

Unsupervised sentence representation learning is a fundamental problem in natural language processing. Recently, contrastive learning has made great success on this task. Existing contrastive learning based models usually apply random sampling to select negative examples for training. Previous work in computer vision has shown that hard negative examples help contrastive learning to achieve faster convergence and better optimization for representation learning. However, the importance of hard negatives in contrastive learning for sentence representation is yet to be explored. In this study, we prove that hard negatives are essential for maintaining strong gradient signals in the training process while random sampling negative examples is ineffective for sentence representation. Accordingly, we present a contrastive model, MixCSE, that extends the current state-of-the-art SimCSE by continually constructing hard negatives via mixing both positive and negative features. The superior performance of the proposed approach is demonstrated via empirical studies on Semantic Textual Similarity datasets and Transfer task datasets.

Added

2026-09-26

The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative

The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative

Leonie Weissweiler, Valentin Hofmann, Abdullatif Köksal, Hinrich Schütze

OrganizationsLMU MunichMunich Center for Machine LearningUniversity of Oxford

Why you should read this

Reveals a critical disconnect in pretrained language models by showing that while models like BERT and DeBERTa reliably identify the syntactic structure of comparative correlative constructions, they consistently fail to understand and apply their underlying semantic meaning.

Construction Grammar (CxG) is a paradigm from cognitive linguistics emphasising the connection between syntax and semantics. Rather than rules that operate on lexical items, it posits constructions as the central building blocks of language, i.e., linguistic units of different granularity that combine syntax and semantics. As a first step towards assessing the compatibility of CxG with the syntactic and semantic knowledge demonstrated by state-of-the-art pretrained language models (PLMs), we present an investigation of their capability to classify and understand one of the most commonly studied constructions, the English comparative correlative (CC). We conduct experiments examining the classification accuracy of a syntactic probe on the one hand and the models’ behaviour in a semantic application task on the other, with BERT, RoBERTa, and DeBERTa as the example PLMs. Our results show that all three investigated PLMs are able to recognise the structure of the CC but fail to use its meaning. While human-like performance of PLMs on many NLP tasks has been alleged, this indicates that PLMs still suffer from substantial shortcomings in central domains of linguistic knowledge.

Added

2026-09-26

DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings

DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings

Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljacic, Shang-Wen Li, Scott Yih, Yoon Kim, James R. Glass

OrganizationsMassachusetts Institute of TechnologyMetaMIT-IBM Watson AI LabUniversity of California, Santa Barbara

Why you should read this

Proposes DiffCSE, an unsupervised sentence embedding framework that improves semantic similarity performance by pairing standard contrastive learning with an auxiliary difference prediction task to make representations sensitive to meaning-altering word edits.

We propose DiffCSE, an unsupervised contrastive learning framework for learning sentence embeddings. DiffCSE learns sentence embeddings that are sensitive to the difference between the original sentence and an edited sentence, where the edited sentence is obtained by stochastically masking out the original sentence and then sampling from a masked language model. We show that DiffCSE is an instance of equivariant contrastive learning (Dangovski et al., 2021), which generalizes contrastive learning and learns representations that are insensitive to certain types of augmentations and sensitive to other “harmful” types of augmentations. Our experiments show that DiffCSE achieves state-of-the-art results among unsupervised sentence representation learning methods, outperforming unsupervised SimCSE1 by 2.3 absolute points on semantic textual similarity tasks. 2

Added

2026-09-26

The Linear Representation Hypothesis and the Geometry of Large Language Models

The Linear Representation Hypothesis and the Geometry of Large Language Models

Kiho Park, Yo Joong Choe, Victor Veitch

OrganizationsUniversity of Chicago

Why you should read this

Formalizes the linear representation hypothesis using counterfactual pairs to unify linear probing and steering under a causally grounded inner product for large language model representations.

Informally, the "linear representation hypothesis" is the idea that high-level concepts are represented linearly as directions in some representation space. In this paper, we address two closely related questions: What does "linear representation" actually mean? And, how do we make sense of geometric notions (e.g., cosine similarity and projection) in the representation space? To answer these, we use the language of counterfactuals to give two formalizations of linear representation, one in the output (word) representation space, and one in the input (context) space. We then prove that these connect to linear probing and model steering, respectively. To make sense of geometric notions, we use the formalization to identify a particular (non-Euclidean) inner product that respects language structure in a sense we make precise. Using this causal inner product, we show how to unify all notions of linear representation. In particular, this allows the construction of probes and steering vectors using counterfactual pairs. Experiments with LLaMA-2 demonstrate the existence of linear representations of concepts, the connection to interpretation and control, and the fundamental role of the choice of inner product. Code is available at github.com/KihoPark/linear_rep_geometry.

Added

2026-09-26

Debiased Contrastive Learning of Unsupervised Sentence Representations

Debiased Contrastive Learning of Unsupervised Sentence Representations

Kun Zhou, Beichen Zhang, Wayne Xin Zhao, Ji-Rong Wen

OrganizationsRenmin University of China

Why you should read this

Proposes a debiased contrastive learning framework that improves unsupervised sentence embeddings by downweighting false negatives and generating optimized noise-based negative samples to overcome representation anisotropy.

Recently, contrastive learning has been shown to be effective in improving pre-trained language models (PLM) to derive high-quality sentence representations. It aims to pull close positive examples to enhance the alignment while push apart irrelevant negatives for the uniformity of the whole representation space. However, previous works mostly adopt in-batch negatives or sample from training data at random. Such a way may cause the sampling bias that improper negatives (e.g., false negatives and anisotropy representations) are used to learn sentence representations, which will hurt the uniformity of the representation space. To address it, we present a new framework DCLR (Debiased Contrastive Learning of unsupervised sentence Representations) to alleviate the influence of these improper negatives. In DCLR, we design an instance weighting method to punish false negatives and generate noise-based negatives to guarantee the uniformity of the representation space. Experiments on seven semantic textual similarity tasks show that our approach is more effective than competitive baselines. Our code and data are publicly available at the link: https://github.com/RUCAIBox/DCLR.

Added

2026-09-26

A Simple but Tough-to-Beat Baseline for Sentence Embeddings

A Simple but Tough-to-Beat Baseline for Sentence Embeddings

Sanjeev Arora, Yingyu Liang, Tengyu Ma

OrganizationsPrinceton University

Why you should read this

Proposes an unsupervised sentence embedding baseline combining smooth inverse frequency weighting with principal component removal that consistently outperforms complex neural network models on semantic similarity benchmarks.

The success of neural network methods for computing word embeddings has motivated methods for generating semantic embeddings of longer pieces of text, such as sentences and paragraphs. Surprisingly, Wieting et al (ICLR'16) showed that such complicated methods are outperformed, especially in out-of-domain (transfer learning) settings, by simpler methods involving mild retraining of word embeddings and basic linear regression. The method of Wieting et al. requires retraining with a substantial labeled dataset such as Paraphrase Database (Ganitkevitch et al., 2013). The current paper goes further, showing that the following completely unsupervised sentence embedding is a formidable baseline: Use word embeddings computed using one of the popular methods on unlabeled corpus like Wikipedia, represent the sentence by a weighted average of the word vectors, and then modify them a bit using PCA/SVD. This weighting improves performance by about 10% to 30% in textual similarity tasks, and beats sophisticated supervised methods including RNN's and LSTM's. It even improves Wieting et al.'s embeddings. This simple method should be used as the baseline to beat in future, especially when labeled training data is scarce or nonexistent. The paper also gives a theoretical explanation of the success of the above unsupervised method using a latent variable generative model for sentences, which is a simple extension of the model in Arora et al. (TACL'16) with new "smoothing" terms that allow for words occuring out of context, as well as high probabilities for words like and, not in all contexts.

Added

2026-09-25

What Does BERT Look at? An Analysis of BERT’s Attention

What Does BERT Look at? An Analysis of BERT’s Attention

Kevin Clark, Urvashi Khandelwal, Omer Levy, Christopher D. Manning

OrganizationsMetaStanford University

Why you should read this

Reveals that individual attention heads in BERT encode distinct grammatical relations and coreference patterns, offering an attention-based analysis framework to interpret linguistic knowledge inside transformer models.

Large pre-trained neural networks such as BERT have had great recent success in NLP, motivating a growing body of research investigating what aspects of language they are able to learn from unlabeled data. Most recent analysis has focused on model outputs (e.g., language model surprisal) or internal vector representations (e.g., probing classifiers). Complementary to these works, we propose methods for analyzing the attention mechanisms of pre-trained models and apply them to BERT. BERT's attention heads exhibit patterns such as attending to delimiter tokens, specific positional offsets, or broadly attending over the whole sentence, with heads in the same layer often exhibiting similar behaviors. We further show that certain attention heads correspond well to linguistic notions of syntax and coreference. For example, we find heads that attend to the direct objects of verbs, determiners of nouns, objects of prepositions, and coreferent mentions with remarkably high accuracy. Lastly, we propose an attention-based probing classifier and use it to further demonstrate that substantial syntactic information is captured in BERT's attention.

Added

2026-09-18

Universal Sentence Encoder

Universal Sentence Encoder

Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, Ray Kurzweil

OrganizationsGoogle

Why you should read this

Introduces the Universal Sentence Encoder, providing two pre-trained embedding architectures that balance computational efficiency with accuracy to deliver strong transfer learning performance across diverse language tasks with minimal training data.

We present models for encoding sentences into embedding vectors that specifically target transfer learning to other NLP tasks. The models are efficient and result in accurate performance on diverse transfer tasks. Two variants of the encoding models allow for trade-offs between accuracy and compute resources. For both variants, we investigate and report the relationship between model complexity, resource consumption, the availability of transfer task training data, and task performance. Comparisons are made with baselines that use word level transfer learning via pretrained word embeddings as well as baselines do not use any transfer learning. We find that transfer learning using sentence embeddings tends to outperform word level transfer. With transfer learning via sentence embeddings, we observe surprisingly good performance with minimal amounts of supervised training data for a transfer task. We obtain encouraging results on Word Embedding Association Tests (WEAT) targeted at detecting model bias. Our pre-trained sentence encoding models are made freely available for download and on TF Hub.

Added

2026-09-16

Supervised Learning of Universal Sentence Representations from Natural Language Inference Data

Supervised Learning of Universal Sentence Representations from Natural Language Inference Data

Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, Antoine Bordes

OrganizationsMetaUniversité du Maine

Why you should read this

Demonstrates that training sentence encoders on supervised natural language inference data yields universal embeddings that consistently outperform unsupervised baselines across diverse transfer tasks, establishing inference as an effective pretraining objective for natural language processing.

Many modern NLP systems rely on word embeddings, previously trained in an unsupervised manner on large corpora, as base features. Efforts to obtain embeddings for larger chunks of text, such as sentences, have however not been so successful. Several attempts at learning unsupervised representations of sentences have not reached satisfactory enough performance to be widely adopted. In this paper, we show how universal sentence representations trained using the supervised data of the Stanford Natural Language Inference datasets can consistently outperform unsupervised methods like SkipThought vectors on a wide range of transfer tasks. Much like how computer vision uses ImageNet to obtain features, which can then be transferred to other tasks, our work tends to indicate the suitability of natural language inference for transfer learning to other NLP tasks. Our encoder is publicly available.

Added

2026-09-16

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut

OrganizationsGoogleToyota Technological Institute at Chicago

Why you should read this

Proposes ALBERT, a lightweight BERT architecture that uses parameter-sharing techniques and a sentence-order prediction objective to surpass BERT-large on GLUE, SQuAD, and RACE while drastically reducing memory footprint and training costs.

Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks. However, at some point further model increases become harder due to GPU/TPU memory limitations and longer training times. To address these problems, we present two parameter-reduction techniques to lower memory consumption and increase the training speed of BERT. Comprehensive empirical evidence shows that our proposed methods lead to models that scale much better compared to the original BERT. We also use a self-supervised loss that focuses on modeling inter-sentence coherence, and show it consistently helps downstream tasks with multi-sentence inputs. As a result, our best model establishes new state-of-the-art results on the GLUE, RACE, and \squad benchmarks while having fewer parameters compared to BERT-large. The code and the pretrained models are available at this https URL.

Added

2026-09-08

License

Published with permission