Built independently by an author, for readers. Read the story and support ChapterPal

keyword

cosine similarity

Cosine similarity is a metric that measures the similarity between two non-zero vectors by calculating the cosine of the angle between them in a multi-dimensional space. Rather than measuring the difference in magnitude or length between vectors, it evaluates their directional alignment. The resulting value ranges from -1 to 1, where 1 indicates that the vectors point in the exact same direction, 0 indicates that they are orthogonal, and -1 indicates that they point in diametrically opposite directions. Because it is invariant to vector scaling, cosine similarity is extensively used in data science, natural language processing, and machine learning to evaluate the semantic or feature-level relatedness between high-dimensional representations, such as text and image embeddings.

5 items

RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

Jiduan Liu, Jiahao Liu, Qifan Wang, Jingang Wang, Wei Wu, Yunsen Xian, Dongyan Zhao, Kai Chen, Rui Yan

OrganizationsBeijing Institute for General Artificial IntelligenceMeituanMetaPeking UniversityRenmin University of China

Why you should read this

Proposes RankCSE, an unsupervised sentence representation framework that integrates ranking consistency and listwise ranking distillation into contrastive learning to capture fine-grained semantic similarities and outperform existing baselines on semantic textual similarity tasks.

Unsupervised sentence representation learning is one of the fundamental problems in natural language processing with various downstream applications. Recently, contrastive learning has been widely adopted which derives high-quality sentence representations by pulling similar semantics closer and pushing dissimilar ones away. However, these methods fail to capture the fine-grained ranking information among the sentences, where each sentence is only treated as either positive or negative. In many real-world scenarios, one needs to distinguish and rank the sentences based on their similarities to a query sentence, e.g., very relevant, moderate relevant, less relevant, irrelevant, etc. In this paper, we propose a novel approach, RankCSE, for unsupervised sentence representation learning, which incorporates ranking consistency and ranking distillation with contrastive learning into a unified framework. In particular, we learn semantically discriminative sentence representations by simultaneously ensuring ranking consistency between two representations with different dropout masks, and distilling listwise ranking knowledge from the teacher. An extensive set of experiments are conducted on both semantic textual similarity (STS) and transfer (TR) tasks. Experimental results demonstrate the superior performance of our approach over several state-of-the-art baselines.

Added

2026-09-26

CosFace: Large Margin Cosine Loss for Deep Face Recognition

CosFace: Large Margin Cosine Loss for Deep Face Recognition

Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, Wei Liu

OrganizationsColumbia UniversityTencent

Why you should read this

Introduces the Large Margin Cosine Loss (CosFace), which reformulates traditional softmax loss through vector normalization and an angular margin penalty to maximize feature discrimination for deep face recognition.

Face recognition has made extraordinary progress owing to the advancement of deep convolutional neural networks (CNNs). The central task of face recognition, including face verification and identification, involves face feature discrimination. However, the traditional softmax loss of deep CNNs usually lacks the power of discrimination. To address this problem, recently several loss functions such as center loss, large margin softmax loss, and angular softmax loss have been proposed. All these improved losses share the same idea: maximizing inter-class variance and minimizing intra-class variance. In this paper, we propose a novel loss function, namely large margin cosine loss (LMCL), to realize this idea from a different perspective. More specifically, we reformulate the softmax loss as a cosine loss by L2L_2 normalizing both features and weight vectors to remove radial variations, based on which a cosine margin term is introduced to further maximize the decision margin in the angular space. As a result, minimum intra-class variance and maximum inter-class variance are achieved by virtue of normalization and cosine decision margin maximization. We refer to our model trained with LMCL as CosFace. Extensive experimental evaluations are conducted on the most popular public-domain face recognition datasets such as MegaFace Challenge, Youtube Faces (YTF) and Labeled Face in the Wild (LFW). We achieve the state-of-the-art performance on these benchmarks, which confirms the effectiveness of our proposed approach.

Added

2026-09-13