Built independently by an author, for readers. Read the story and support ChapterPal

keyword

semantic textual similarity

Semantic textual similarity is a natural language processing task and evaluation metric that measures the degree of meaning equivalence between two pieces of text, such as sentences or phrases. Unlike binary classification frameworks such as textual entailment or paraphrase detection, semantic textual similarity assesses semantic overlap on a graded or continuous scale ranging from complete independence to full semantic equivalence. In computational linguistics and machine learning, it serves as a standard benchmark for evaluating the quality of text representations and sentence embeddings, supporting downstream applications that include information retrieval, semantic search, question answering, automated summarization, and machine translation evaluation.

10 items

PromptBERT: Improving BERT Sentence Embeddings with Prompts

PromptBERT: Improving BERT Sentence Embeddings with Prompts

Ting Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang, Deqing Wang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, Qi Zhang

OrganizationsBeihang UniversityMicrosoftZhongguancun Laboratory

Why you should read this

Introduces a prompt-based contrastive learning framework with template denoising that overcomes token embedding biases in pretrained transformers to significantly outperform SimCSE on unsupervised sentence representation benchmarks.

We propose PromptBERT, a novel contrastive learning method for learning better sentence representation. We firstly analyze the drawback of current sentence embedding from original BERT and find that it is mainly due to the static token embedding bias and ineffective BERT layers. Then we propose the first prompt-based sentence embeddings method and discuss two prompt representing methods and three prompt searching methods to make BERT achieve better sentence embeddings. Moreover, we propose a novel unsupervised training objective by the technology of template denoising, which substantially shortens the performance gap between the supervised and unsupervised settings. Extensive experiments show the effectiveness of our method. Compared to SimCSE, PromptBert achieves 2.29 and 2.58 points of improvement based on BERT and RoBERTa in the unsupervised setting. Our code is available at https://github.com/kongds/Prompt-BERT.

Added

2026-09-26

RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

Jiduan Liu, Jiahao Liu, Qifan Wang, Jingang Wang, Wei Wu, Yunsen Xian, Dongyan Zhao, Kai Chen, Rui Yan

OrganizationsBeijing Institute for General Artificial IntelligenceMeituanMetaPeking UniversityRenmin University of China

Why you should read this

Proposes RankCSE, an unsupervised sentence representation framework that integrates ranking consistency and listwise ranking distillation into contrastive learning to capture fine-grained semantic similarities and outperform existing baselines on semantic textual similarity tasks.

Unsupervised sentence representation learning is one of the fundamental problems in natural language processing with various downstream applications. Recently, contrastive learning has been widely adopted which derives high-quality sentence representations by pulling similar semantics closer and pushing dissimilar ones away. However, these methods fail to capture the fine-grained ranking information among the sentences, where each sentence is only treated as either positive or negative. In many real-world scenarios, one needs to distinguish and rank the sentences based on their similarities to a query sentence, e.g., very relevant, moderate relevant, less relevant, irrelevant, etc. In this paper, we propose a novel approach, RankCSE, for unsupervised sentence representation learning, which incorporates ranking consistency and ranking distillation with contrastive learning into a unified framework. In particular, we learn semantically discriminative sentence representations by simultaneously ensuring ranking consistency between two representations with different dropout masks, and distilling listwise ranking knowledge from the teacher. An extensive set of experiments are conducted on both semantic textual similarity (STS) and transfer (TR) tasks. Experimental results demonstrate the superior performance of our approach over several state-of-the-art baselines.

Added

2026-09-26

XNLI: Evaluating Cross-lingual Sentence Representations

XNLI: Evaluating Cross-lingual Sentence Representations

Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R. Bowman, Holger Schwenk, Veselin Stoyanov

OrganizationsMetaNew York University

Why you should read this

Introduces the Cross-lingual Natural Language Inference (XNLI) benchmark across 15 diverse languages, establishing a standard evaluation suite and competitive baselines for cross-lingual sentence representations and zero-shot transfer.

State-of-the-art natural language processing systems rely on supervision in the form of annotated data to learn competent models. These models are generally trained on data in a single language (usually English), and cannot be directly used beyond that language. Since collecting data in every language is not realistic, there has been a growing interest in cross-lingual language understanding (XLU) and low-resource cross-language transfer. In this work, we construct an evaluation set for XLU by extending the development and test sets of the Multi-Genre Natural Language Inference Corpus (MultiNLI) to 15 languages, including low-resource languages such as Swahili and Urdu. We hope that our dataset, dubbed XNLI, will catalyze research in cross-lingual sentence understanding by providing an informative standard evaluation task. In addition, we provide several baselines for multilingual sentence understanding, including two based on machine translation systems, and two that use parallel data to train aligned multilingual bag-of-words and LSTM encoders. We find that XNLI represents a practical and challenging evaluation suite, and that directly translating the test data yields the best performance among available baselines.

Added

2026-09-24

Supervised Learning of Universal Sentence Representations from Natural Language Inference Data

Supervised Learning of Universal Sentence Representations from Natural Language Inference Data

Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, Antoine Bordes

OrganizationsMetaUniversité du Maine

Why you should read this

Demonstrates that training sentence encoders on supervised natural language inference data yields universal embeddings that consistently outperform unsupervised baselines across diverse transfer tasks, establishing inference as an effective pretraining objective for natural language processing.

Many modern NLP systems rely on word embeddings, previously trained in an unsupervised manner on large corpora, as base features. Efforts to obtain embeddings for larger chunks of text, such as sentences, have however not been so successful. Several attempts at learning unsupervised representations of sentences have not reached satisfactory enough performance to be widely adopted. In this paper, we show how universal sentence representations trained using the supervised data of the Stanford Natural Language Inference datasets can consistently outperform unsupervised methods like SkipThought vectors on a wide range of transfer tasks. Much like how computer vision uses ImageNet to obtain features, which can then be transferred to other tasks, our work tends to indicate the suitability of natural language inference for transfer learning to other NLP tasks. Our encoder is publicly available.

Added

2026-09-16

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, Radu Soricut

OrganizationsGoogleToyota Technological Institute at Chicago

Why you should read this

Proposes ALBERT, a lightweight BERT architecture that uses parameter-sharing techniques and a sentence-order prediction objective to surpass BERT-large on GLUE, SQuAD, and RACE while drastically reducing memory footprint and training costs.

Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks. However, at some point further model increases become harder due to GPU/TPU memory limitations and longer training times. To address these problems, we present two parameter-reduction techniques to lower memory consumption and increase the training speed of BERT. Comprehensive empirical evidence shows that our proposed methods lead to models that scale much better compared to the original BERT. We also use a self-supervised loss that focuses on modeling inter-sentence coherence, and show it consistently helps downstream tasks with multi-sentence inputs. As a result, our best model establishes new state-of-the-art results on the GLUE, RACE, and \squad benchmarks while having fewer parameters compared to BERT-large. The code and the pretrained models are available at this https URL.

Added

2026-09-08

License

Published with permission

TinyBERT: Distilling BERT for Natural Language Understanding

TinyBERT: Distilling BERT for Natural Language Understanding

Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, Qun Liu

OrganizationsHuaweiHuazhong University of Science and Technology

Why you should read this

Details an indispensable two-stage distillation blueprint that matches transformer embeddings and self-attention matrices during both the pre-training and task-specific fine-tuning phases.

Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually computationally expensive, so it is difficult to efficiently execute them on resourcerestricted devices. To accelerate inference and reduce model size while maintaining accuracy, we first propose a novel Transformer distillation method that is specially designed for knowledge distillation (KD) of the Transformer-based models. By leveraging this new KD method, the plenty of knowledge encoded in a large "teacher" BERT can be effectively transferred to a small "student" Tiny-BERT. Then, we introduce a new two-stage learning framework for TinyBERT, which performs Transformer distillation at both the pretraining and task-specific learning stages. This framework ensures that TinyBERT can capture the general-domain as well as the task-specific knowledge in BERT.

Added

2026-06-20

Creative Commons License
Text Embeddings by Weakly-Supervised Contrastive Pre-training

Text Embeddings by Weakly-Supervised Contrastive Pre-training

Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei

OrganizationsMicrosoft

Why you should read this

Introduces E5, a novel family of text embeddings trained with weakly-supervised contrastive pre-training on a curated large-scale dataset, achieving state-of-the-art performance across diverse NLP tasks, including outperforming BM25 in zero-shot retrieval and setting new records on MTEB with significantly fewer parameters.

This paper presents E5, a family of state-of-the-art text embeddings that transfer well to a wide range of tasks. The model is trained in a contrastive manner with weak supervision signals from our curated large-scale text pair dataset (called CCPairs). E5 can be readily used as a general-purpose embedding model for any tasks requiring a single-vector representation of texts such as retrieval, clustering, and classification, achieving strong performance in both zero-shot and fine-tuned settings. We conduct extensive evaluations on 56 datasets from the BEIR and MTEB benchmarks. For zero-shot settings, E5 is the first model that outperforms the strong BM25 baseline on the BEIR retrieval benchmark without using any labeled data. When fine-tuned, E5 obtains the best results on the MTEB benchmark, beating existing embedding models with 40x more parameters.

Added

2026-05-03

Creative Commons License