RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank
Jiduan LiuJiahao LiuQifan WangJingang WangWei WuYunsen XianDongyan ZhaoKai ChenRui Yan
Proposes RankCSE, an unsupervised sentence representation framework that integrates ranking consistency and listwise ranking distillation into contrastive learning to capture fine-grained semantic similarities and outperform existing baselines on semantic textual similarity tasks.
Modern natural language processing systems often rely on unsupervised sentence representations to convert text into numerical vectors for tasks such as information retrieval, text clustering, and semantic search. While recent contrastive learning techniques have advanced this area by pulling similar sentences together and pushing dissimilar ones apart, they treat non-matching sentences equally as generic negative samples. This binary distinction prevents models from learning fine-grained ranking relationships—such as distinguishing highly relevant sentences from moderately relevant ones—which limits search precision and recommendation quality.
The article introduces and evaluates RankCSE, a framework designed to learn semantically discriminative sentence embeddings by integrating listwise ranking principles directly into unsupervised contrastive learning. The main objective was to demonstrate that incorporating ranking consistency and ranking distillation allows models to generalize nuanced, fine-grained semantic hierarchies without requiring expensive human-annotated data.
The researchers trained their models using one million unlabeled sentences randomly sampled from English Wikipedia across standard pre-trained architectures, including BERT and RoBERTa. The approach combines standard contrastive loss with two new components: a ranking consistency loss that enforces stable similarity orderings across different network augmentations, and a ranking distillation loss that transfers listwise ranking signals from teacher models. The evaluation spanned seven semantic textual similarity benchmarks and seven downstream transfer classification tasks, measuring performance against established baselines using correlation and accuracy metrics.
The key findings show consistent performance advantages across all benchmarks. First, RankCSE outperformed existing unsupervised baselines across all evaluated architectures, achieving average semantic similarity scores of 80.05% to 80.60% on BERT and 79.73% to 80.60% on RoBERTa. Second, the base version of RankCSE surpassed the much larger baseline model (SimCSE-BERTlarge) on semantic similarity tasks by roughly 2%, illustrating substantial representational efficiency. Third, downstream classification tasks improved to an average accuracy of 87.33% on standard benchmarks. Finally, geometric analysis confirmed that RankCSE establishes a superior balance between bringing related sentences close together and maintaining a uniform spread across the representation space, leading to more robust and stable embeddings.
These results demonstrate that incorporating learning-to-rank methods into unsupervised embedding models enables significant quality gains without labeled training data. For organizations deploying search, matching, or recommendation pipelines, this framework delivers higher precision and relevance while avoiding data labeling costs. Moreover, because smaller base models equipped with RankCSE can outperform standard large models, organizations can lower inference latency and cloud compute expenditures.
Engineering and research teams seeking to improve search or text-similarity pipelines should evaluate RankCSE-style ranking objectives as drop-in upgrades over standard contrastive learning methods. Multi-teacher distillation configurations—such as combining predictions from complementary models—are recommended for optimal performance. Before full production deployment, teams should conduct pilot tests to evaluate the added training-time cost, as calculating teacher ranking signals increases training duration to roughly 2 to 3.7 hours on an A100 GPU compared to simpler single-stage contrastive approaches. Future work should focus on automating teacher model selection and extending listwise ranking objectives to domain-specific corpora.
- Paper: SimCSE: Simple Contrastive Learning of Sentence Embeddings, Tianyu Gao et al. (2021). This paper establishes the foundational unsupervised dropout-as-augmentation contrastive sentence embedding framework that RankCSE directly builds upon and enhances.
- Paper: Learning to rank: from pairwise approach to listwise approach, Zhe Cao et al. (2007). This work introduces listwise learning-to-rank probability formulations that provide the core theoretical and mathematical foundation for RankCSE's listwise ranking distillation.
- Paper: DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings, Yung-Sung Chuang et al. (2022). This study advances contrastive sentence embeddings by making representations sensitive to subtle textual differences, motivating RankCSE's shift toward fine-grained ranking objectives.
- Paper: Debiased Contrastive Learning of Unsupervised Sentence Representations, Kun Zhou et al. (2022). This paper analyzes the limitations and biases of standard binary positive-negative sampling in sentence contrastive learning, which RankCSE addresses via continuous ranking consistency.
- Paper: Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, Nils Reimers et al. (2019). This foundational work establishes Siamese BERT sentence embeddings and the standard semantic textual similarity evaluation benchmark utilized by RankCSE.
- Paper: Learning to rank using gradient descent, Chris Burges et al. (2005). This classic paper introduces learning-to-rank optimization principles using gradient descent that underpin ranking-based loss formulations in neural models.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). This paper extends context ranking principles directly into large language models to unify retrieval ranking with generation in RAG pipelines.
- Paper: Learning to Rank in Generative Retrieval, Yongqi Li et al. (2024). This work explores integrating learning-to-rank objectives into generative retrieval architectures to bridge the gap between generation loss and listwise document ranking.
- Paper: Making Text Embedders Few-Shot Learners, Chaofan Li et al. (2025). This work investigates advanced text embedding generation by extending dense representations through in-context learning and instruction tuning in LLMs.
