Built independently by an author, for readers. Read the story and support ChapterPal

keyword

semantic similarity

Semantic similarity is a computational and linguistic measure of how closely two units of language, such as words, sentences, or entire texts, correspond in their underlying meaning. Unlike broader semantic relatedness, which encompasses topical associations and functional connections between distinct concepts, semantic similarity specifically evaluates likeness, equivalence, or shared taxonomic characteristics, such as synonymy or shared hypernyms. In natural language processing, it is typically quantified using structural and information-theoretic calculations over lexical taxonomies, statistical analyses of corpus co-occurrences, or geometric distance metrics like cosine similarity applied to dense vector embeddings. Measuring semantic similarity plays a fundamental role in understanding and processing human language, supporting applications such as paraphrase identification, information retrieval, question answering, and automated text evaluation.

25 items

AnyEdit: Edit Any Knowledge Encoded in Language Models

AnyEdit: Edit Any Knowledge Encoded in Language Models

Houcheng Jiang, Junfeng Fang, Ningyu Zhang, Mingyang Wan, Guojun Ma, Xiang Wang, Xiangnan He, Tat-Seng Chua

OrganizationsByteDanceNational University of SingaporeUniversity of Science and Technology of ChinaZhejiang University

Why you should read this

Proposes AnyEdit, an autoregressive framework that breaks down complex, long-form knowledge into sequential chunks to iteratively edit key tokens, enabling existing model editing techniques to update diverse formats like code and mathematics without degrading accuracy over long outputs.

Large language models (LLMs) often produce incorrect or outdated information, necessitating efficient and precise knowledge updates. Current model editing methods, however, struggle with long-form knowledge in diverse formats, such as poetry, code snippets, and mathematical derivations. These limitations arise from their reliance on editing a single token’s hidden state, a limitation we term as “efficacy barrier”. To solve this, we propose AnyEdit, a new autoregressive editing paradigm. It decomposes long-form knowledge into sequential chunks and iteratively edits the key token in each chunk, ensuring consistent and accurate outputs. Theoretically, we ground AnyEdit in the Chain Rule of Mutual Information, showing its ability to update any knowledge within LLMs. Empirically, it outperforms strong baselines by 21.5% on benchmarks including UnKEBench, AKEW, and our new EditEverything dataset for long-form diverse-formatted knowledge. Additionally, AnyEdit serves as a plug-and-play framework, enabling current editing methods to update knowledge with arbitrary length and format, significantly advancing the scope and practicality of LLM knowledge editing. Our code is available at: https://github.com/jianghoucheng/AnyEdit.

Added

2026-10-03

Evaluating Open-Domain Question Answering in the Era of Large Language Models

Evaluating Open-Domain Question Answering in the Era of Large Language Models

Ehsan Kamalloo, Nouha Dziri, Charles L. A. Clarke, Davood Rafiei

OrganizationsAllen Institute for AIUniversity of AlbertaUniversity of Waterloo

Why you should read this

Reveals that standard lexical metrics drastically underestimate generative LLM performance on open-domain question answering benchmarks by missing semantically equivalent answers and failing to handle hallucinations, demonstrating why human evaluation remains indispensable.

Lexical matching remains the de facto evaluation method for open-domain question answering (QA). Unfortunately, lexical matching fails completely when a plausible candidate answer does not appear in the list of gold answers, which is increasingly the case as we shift from extractive to generative models. The recent success of large language models (LLMs) for QA aggravates lexical matching failures since candidate answers become longer, thereby making matching with the gold answers even more challenging. Without accurate evaluation, the true progress in open-domain QA remains unknown. In this paper, we conduct a thorough analysis of various open-domain QA models, including LLMs, by manually evaluating their answers on a subset of NQ-OPEN, a popular benchmark. Our assessments reveal that while the true performance of all models is significantly underestimated, the performance of the InstructGPT (zero-shot) LLM increases by nearly +60%, making it on par with existing top models, and the InstructGPT (few-shot) model actually achieves a new state-of-the-art on NQ-OPEN. We also find that more than 50% of lexical matching failures are attributed to semantically equivalent answers. We further demonstrate that regex matching ranks QA models consistent with human judgments, although still suffering from unnecessary strictness. Finally, we demonstrate that automated evaluation models are a reasonable surrogate for lexical matching in some circumstances, but not for long-form answers generated by LLMs. The automated models struggle in detecting hallucinations in LLM answers and are thus unable to evaluate LLMs. At this time, there appears to be no substitute for human evaluation.

Added

2026-09-30

Language-agnostic BERT Sentence Embedding

Language-agnostic BERT Sentence Embedding

Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, Wei Wang

OrganizationsGoogle

Why you should read this

Presents a multilingual sentence embedding framework combining pre-trained BERT models with dual-encoder translation ranking to reduce parallel data requirements by 80% while achieving state-of-the-art bitext retrieval performance across more than 100 languages.

While BERT is an effective method for learning monolingual sentence embeddings for semantic similarity and embedding based transfer learning (Reimers and Gurevych, 2019), BERT based cross-lingual sentence embeddings have yet to be explored. We systematically investigate methods for learning multilingual sentence embeddings by combining the best methods for learning monolingual and cross-lingual representations including: masked language modeling (MLM), translation language modeling (TLM) (Conneau and Lample, 2019), dual encoder translation ranking (Guo et al., 2018), and additive margin softmax (Yang et al., 2019a). We show that introducing a pre-trained multilingual language model dramatically reduces the amount of parallel training data required to achieve good performance by 80%. Composing the best of these methods produces a model that achieves 83.7% bi-text retrieval accuracy over 112 languages on Tatoeba, well above the 65.5% achieved by Artetxe and Schwenk (2019b), while still performing competitively on monolingual transfer learning benchmarks (Conneau and Kiela, 2018). Parallel data mined from CommonCrawl using our best model is shown to train competitive NMT models for en-zh and en-de. We publicly release our best multilingual sentence embedding model for 109+ languages at this https URL.

Added

2026-09-26

Creative Commons License
PCL: Peer-Contrastive Learning with Diverse Augmentations for Unsupervised Sentence Embeddings

PCL: Peer-Contrastive Learning with Diverse Augmentations for Unsupervised Sentence Embeddings

Qiyu Wu, Chongyang Tao, Tao Shen, Can Xu, Xiubo Geng, Daxin Jiang

OrganizationsMicrosoftUniversity of Tokyo

Why you should read this

Proposes a peer-contrastive learning framework that uses dual cooperating networks and multi-augmentation strategies to prevent shortcut biases and improve unsupervised sentence embeddings across semantic textual similarity benchmarks.

Learning sentence embeddings in an unsupervised manner is fundamental in natural language processing. Recent common practice is to couple pre-trained language models with unsupervised contrastive learning, whose success relies on augmenting a sentence with a semantically-close positive instance to construct contrastive pairs. Nonetheless, existing approaches usually depend on a mono-augmenting strategy, which causes learning shortcuts towards the augmenting biases and thus corrupts the quality of sentence embeddings. A straightforward solution is resorting to more diverse positives from a multi-augmenting strategy, while an open question remains about how to unsupervisedly learn from the diverse positives but with uneven augmenting qualities in the text field. As one answer, we propose a novel Peer-Contrastive Learning (PCL) with diverse augmentations. PCL constructs diverse contrastive positives and negatives at the group level for unsupervised sentence embeddings. PCL performs peer-positive contrast as well as peer-network cooperation, which offers an inherent anti-bias ability and an effective way to learn from diverse augmentations. Experiments on STS benchmarks verify the effectiveness of PCL against its competitors in unsupervised sentence embeddings.^1

Added

2026-09-26

SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation

SimLex-999: Evaluating Semantic Models With (Genuine) Similarity Estimation

Felix Hill, Roi Reichart, Anna Korhonen

OrganizationsTechnion – Israel Institute of TechnologyUniversity of Cambridge

Why you should read this

Introduces SimLex-999, a gold-standard evaluation benchmark that isolates true semantic similarity from conceptual association across diverse parts of speech and concreteness levels, exposing performance gaps in vector space models.

We present SimLex-999, a gold standard resource for evaluating distributional semantic models that improves on existing resources in several important ways. First, in contrast to gold standards such as WordSim-353 and MEN, it explicitly quantifies similarity rather than association or relatedness, so that pairs of entities that are associated but not actually similar [Freud, psychology] have a low rating. We show that, via this focus on similarity, SimLex-999 incentivizes the development of models with a different, and arguably wider range of applications than those which reflect conceptual association. Second, SimLex-999 contains a range of concrete and abstract adjective, noun and verb pairs, together with an independent rating of concreteness and (free) association strength for each pair. This diversity enables fine-grained analyses of the performance of models on concepts of different types, and consequently greater insight into how architectures can be improved. Further, unlike existing gold standard evaluations, for which automatic approaches have reached or surpassed the inter-annotator agreement ceiling, state-of-the-art models perform well below this ceiling on SimLex-999. There is therefore plenty of scope for SimLex-999 to quantify future improvements to distributional semantic models, guiding the development of the next generation of representation-learning architectures.

Added

2026-09-25

Corpus-based and Knowledge-based Measures of Text Semantic Similarity

Corpus-based and Knowledge-based Measures of Text Semantic Similarity

Rada Mihalcea, Courtney Corley, Carlo Strapparava

OrganizationsDepartment of Computer ScienceFondazione Bruno KesslerIstituto per la Ricerca Scientifica e TecnologicaUniversity of North Texas

Why you should read this

Proposes a framework for evaluating short text semantic similarity by combining word-level corpus and knowledge-based metrics with inverse document frequency weighting, significantly outperforming standard lexical and vector-space models on paraphrase recognition tasks.

This paper presents a method for measuring the semantic similarity of texts, using corpus-based and knowledge-based measures of similarity. Previous work on this problem has focused mainly on either large documents (e.g. text classification, information retrieval) or individual words (e.g. synonymy tests). Given that a large fraction of the information available today, on the Web and elsewhere, consists of short text snippets (e.g. abstracts of scientific documents, image captions, product descriptions), in this paper we focus on measuring the semantic similarity of short texts. Through experiments performed on a paraphrase data set, we show that the semantic similarity method outperforms methods based on simple lexical matching, resulting in up to 13% error rate reduction with respect to the traditional vector-based similarity metric.

Added

2026-09-25

Measuring praise and criticism: Inference of semantic orientation from association

Measuring praise and criticism: Inference of semantic orientation from association

Peter D. Turney, Michael L. Littman

OrganizationsNational Research Council CanadaRutgers University

Why you should read this

Proposes a method for automatically determining the positive or negative semantic orientation of words across diverse parts of speech by measuring their statistical association with paradigm seed words using pointwise mutual information and latent semantic analysis.

The evaluative character of a word is called its semantic orientation. Positive semantic orientation indicates praise (e.g., "honest", "intrepid") and negative semantic orientation indicates criticism (e.g., "disturbing", "superfluous"). Semantic orientation varies in both direction (positive or negative) and degree (mild to strong). An automated system for measuring semantic orientation would have application in text classification, text filtering, tracking opinions in online discussions, analysis of survey responses, and automated chat systems (chatbots). This paper introduces a method for inferring the semantic orientation of a word from its statistical association with a set of positive and negative paradigm words. Two instances of this approach are evaluated, based on two different statistical measures of word association: pointwise mutual information (PMI) and latent semantic analysis (LSA). The method is experimentally tested with 3,596 words (including adjectives, adverbs, nouns, and verbs) that have been manually labeled positive (1,614 words) and negative (1,982 words). The method attains an accuracy of 82.8% on the full test set, but the accuracy rises above 95% when the algorithm is allowed to abstain from classifying mild words.

Added

2026-09-24

Universal Sentence Encoder

Universal Sentence Encoder

Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, Ray Kurzweil

OrganizationsGoogle

Why you should read this

Introduces the Universal Sentence Encoder, providing two pre-trained embedding architectures that balance computational efficiency with accuracy to deliver strong transfer learning performance across diverse language tasks with minimal training data.

We present models for encoding sentences into embedding vectors that specifically target transfer learning to other NLP tasks. The models are efficient and result in accurate performance on diverse transfer tasks. Two variants of the encoding models allow for trade-offs between accuracy and compute resources. For both variants, we investigate and report the relationship between model complexity, resource consumption, the availability of transfer task training data, and task performance. Comparisons are made with baselines that use word level transfer learning via pretrained word embeddings as well as baselines do not use any transfer learning. We find that transfer learning using sentence embeddings tends to outperform word level transfer. With transfer learning via sentence embeddings, we observe surprisingly good performance with minimal amounts of supervised training data for a transfer task. We obtain encouraging results on Word Embedding Association Tests (WEAT) targeted at detecting model bias. Our pre-trained sentence encoding models are made freely available for download and on TF Hub.

Added

2026-09-16