Built independently by an author, for readers. Read the story and support ChapterPal

keyword

unsupervised sentence representation learning

Unsupervised sentence representation learning is a natural language processing methodology that maps entire sentences into dense, fixed-dimensional semantic vectors without relying on human-annotated labels. In this framework, computational models analyze patterns across vast collections of unlabeled text to capture the underlying meaning, context, and structure of sentences. By employing self-supervised objectives, such as contrastive learning, autoencoding, or predictive language modeling, the system learns to position semantically related sentences close together in a continuous vector space while separating unrelated ones. The resulting embeddings provide transferable numerical representations of text that facilitate various downstream tasks, including semantic textual similarity evaluation, information retrieval, text clustering, and ranking.

1 item

RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

Jiduan Liu, Jiahao Liu, Qifan Wang, Jingang Wang, Wei Wu, Yunsen Xian, Dongyan Zhao, Kai Chen, Rui Yan

OrganizationsBeijing Institute for General Artificial IntelligenceMeituanMetaPeking UniversityRenmin University of China

Why you should read this

Proposes RankCSE, an unsupervised sentence representation framework that integrates ranking consistency and listwise ranking distillation into contrastive learning to capture fine-grained semantic similarities and outperform existing baselines on semantic textual similarity tasks.

Unsupervised sentence representation learning is one of the fundamental problems in natural language processing with various downstream applications. Recently, contrastive learning has been widely adopted which derives high-quality sentence representations by pulling similar semantics closer and pushing dissimilar ones away. However, these methods fail to capture the fine-grained ranking information among the sentences, where each sentence is only treated as either positive or negative. In many real-world scenarios, one needs to distinguish and rank the sentences based on their similarities to a query sentence, e.g., very relevant, moderate relevant, less relevant, irrelevant, etc. In this paper, we propose a novel approach, RankCSE, for unsupervised sentence representation learning, which incorporates ranking consistency and ranking distillation with contrastive learning into a unified framework. In particular, we learn semantically discriminative sentence representations by simultaneously ensuring ranking consistency between two representations with different dropout masks, and distilling listwise ranking knowledge from the teacher. An extensive set of experiments are conducted on both semantic textual similarity (STS) and transfer (TR) tasks. Experimental results demonstrate the superior performance of our approach over several state-of-the-art baselines.

Added

2026-09-26