Built independently by an author, for readers. Read the story and support ChapterPal

keyword

soft labels

Soft labels are continuous probability distributions or real-valued vectors assigned across potential target categories, in contrast to hard labels that assign a data point strictly to a single category using discrete binary values. By allocating fractional weights to multiple classes, soft labels represent uncertainty, label noise, partial membership, or genuine variation and disagreement among human annotators. In machine learning, training with soft labels instead of rigid one-hot targets prevents overconfidence, improves model generalization and calibration, and allows models to capture nuanced, many-to-many relationships among data categories or modalities.

3 items

The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

The "Problem" of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

Barbara Plank

OrganizationsLMU MunichLudwig Maximilian University of MunichMaiNLP LabMunich Center for Machine Learning

Why you should read this

Argues that human annotation variation represents meaningful subjectivity rather than noise, synthesizing its impact across data collection, modeling, and evaluation while compiling a unified repository of un-aggregated datasets to guide future machine learning research.

Human variation in labeling is often considered noise. Annotation projects for machine learning (ML) aim at minimizing human label variation, with the assumption to maximize data quality and in turn optimize and maximize machine learning metrics. However, this conventional practice assumes that there exists a ground truth, and neglects that there exists genuine human variation in labeling due to disagreement, subjectivity in annotation or multiple plausible answers. In this position paper, we argue that this big open problem of human label variation persists and critically needs more attention to move our field forward. This is because human label variation impacts all stages of the ML pipeline: data, modeling and evaluation. However, few works consider all of these dimensions jointly; and existing research is fragmented. We reconcile different previously proposed notions of human label variation, provide a repository of publicly-available datasets with un-aggregated labels, depict approaches proposed so far, identify gaps and suggest ways forward. As datasets are becoming increasingly available, we hope that this synthesized view on the “problem” will lead to an open discussion on possible strategies to devise fundamentally new directions.

Added

2026-10-01

SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger

SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger

Yuting Gao, Jinfeng Liu, Zihan Xu, Tong Wu, Enwei Zhang, Ke Li, Jie Yang, Wei Liu, Xing Sun

OrganizationsShanghai Jiao Tong UniversityTencent

Why you should read this

Proposes a relaxed contrastive learning framework that uses fine-grained intra-modal self-similarity and negative-pair disentanglement as soft alignment targets, significantly improving CLIP's zero-shot classification performance on noisy web-scale image-text datasets.

During the preceding biennium, vision-language pre-training has achieved noteworthy success on several downstream tasks. Nevertheless, acquiring high-quality image-text pairs, where the pairs are entirely exclusive of each other, remains a challenging task, and noise exists in the commonly used datasets. To address this issue, we propose SoftCLIP, a novel approach that relaxes the strict one-to-one constraint and achieves a soft cross-modal alignment by introducing a softened target, which is generated from the fine-grained intra-modal self-similarity. The intra-modal guidance is indicative to enable two pairs have some local similarities and model many-to-many relationships between the two modalities. Besides, since the positive still dominates in the softened target distribution, we disentangle the negatives in the distribution to further boost the relation alignment with the negatives in the cross-modal learning. Extensive experiments demonstrate the effectiveness of SoftCLIP. In particular, on ImageNet zero-shot classification task, using CC3M/CC12M as pre-training dataset, SoftCLIP brings a top-1 accuracy improvement of 6.8%/7.2% over the CLIP baseline.

Added

2026-09-26

Learning From Noisy Labels With Deep Neural Networks: A Survey

Learning From Noisy Labels With Deep Neural Networks: A Survey

Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, Jae-Gil Lee

OrganizationsKorea Advanced Institute of Science and TechnologyNaver AI Lab

Why you should read this

Categorizes 62 methods for training deep neural networks with label noise across five distinct paradigms, providing a comparative analysis of evaluation properties, noise estimation techniques, and standard benchmarks.

Deep learning has achieved remarkable success in numerous domains with help from large amounts of big data. However, the quality of data labels is a concern because of the lack of high-quality labels in many real-world scenarios. As noisy labels severely degrade the generalization performance of deep neural networks, learning from noisy labels (robust training) is becoming an important task in modern deep learning applications. In this survey, we first describe the problem of learning with label noise from a supervised learning perspective. Next, we provide a comprehensive review of 62 state-of-the-art robust training methods, all of which are categorized into five groups according to their methodological difference, followed by a systematic comparison of six properties used to evaluate their superiority. Subsequently, we perform an in-depth analysis of noise rate estimation and summarize the typically used evaluation methodology, including public noisy datasets and evaluation metrics. Finally, we present several promising research directions that can serve as a guideline for future studies. All the contents will be available at this https URL.

Added

2026-09-25