Built independently by an author, for readers. Read the story and support ChapterPal

keyword

RankCSE

RankCSE is an unsupervised machine learning framework designed to generate high-quality sentence embeddings by integrating contrastive learning with learning-to-rank objectives. While standard contrastive sentence embedding methods typically treat data points in a binary manner as either similar or dissimilar, RankCSE captures more nuanced, fine-grained degrees of semantic relevance among sentences. It operates by enforcing ranking consistency between different representations of an input produced under distinct dropout masks while simultaneously distilling listwise ranking knowledge from a teacher model into the representations. This approach enables natural language processing models to produce semantically discriminative vector representations that effectively capture varying levels of semantic similarity for downstream text retrieval and evaluation tasks.

1 item

RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

Jiduan Liu, Jiahao Liu, Qifan Wang, Jingang Wang, Wei Wu, Yunsen Xian, Dongyan Zhao, Kai Chen, Rui Yan

OrganizationsBeijing Institute for General Artificial IntelligenceMeituanMetaPeking UniversityRenmin University of China

Why you should read this

Proposes RankCSE, an unsupervised sentence representation framework that integrates ranking consistency and listwise ranking distillation into contrastive learning to capture fine-grained semantic similarities and outperform existing baselines on semantic textual similarity tasks.

Unsupervised sentence representation learning is one of the fundamental problems in natural language processing with various downstream applications. Recently, contrastive learning has been widely adopted which derives high-quality sentence representations by pulling similar semantics closer and pushing dissimilar ones away. However, these methods fail to capture the fine-grained ranking information among the sentences, where each sentence is only treated as either positive or negative. In many real-world scenarios, one needs to distinguish and rank the sentences based on their similarities to a query sentence, e.g., very relevant, moderate relevant, less relevant, irrelevant, etc. In this paper, we propose a novel approach, RankCSE, for unsupervised sentence representation learning, which incorporates ranking consistency and ranking distillation with contrastive learning into a unified framework. In particular, we learn semantically discriminative sentence representations by simultaneously ensuring ranking consistency between two representations with different dropout masks, and distilling listwise ranking knowledge from the teacher. An extensive set of experiments are conducted on both semantic textual similarity (STS) and transfer (TR) tasks. Experimental results demonstrate the superior performance of our approach over several state-of-the-art baselines.

Added

2026-09-26