keyword
label smoothing
Label smoothing is a regularization technique in machine learning that replaces hard, one-hot encoded target labels with smoothed probability distributions across all possible classes. In standard classification tasks, target labels assign full probability to the ground-truth class and zero to all others, which can cause neural networks to become overconfident and produce excessively large logit outputs. Label smoothing mitigates this issue by computing a weighted average between the true label vector and a uniform distribution, thereby distributing a small fraction of probability mass to non-target classes. This adjustment prevents the model from fitting too tightly to the training labels, reduces overfitting, and improves both prediction calibration and general performance across unseen data.
6 items

On Mixup Regularization
Luigi Carratino, Moustapha Cissé, Rodolphe Jenatton, Jean-Philippe Vert
Why you should read this
Explains the theoretical mechanisms behind Mixup by formalizing it as empirical risk minimization with data transformation and random perturbations, leading to a simple test-time adjustment that improves prediction accuracy and calibration.
Mixup is a data augmentation technique that creates new examples as convex combinations of training points and labels. This simple technique has empirically shown to improve the accuracy of many state-of-the-art models in different settings and applications, but the reasons behind this empirical success remain poorly understood. In this paper we take a substantial step in explaining the theoretical foundations of Mixup, by clarifying its regularization effects. We show that Mixup can be interpreted as standard empirical risk minimization estimator subject to a combination of data transformation and random perturbation of the transformed data. We gain two core insights from this new interpretation. First, the data transformation suggests that, at test time, a model trained with Mixup should also be applied to transformed data, a one-line change in code that we show empirically to improve both accuracy and calibration of the prediction. Second, we show how the random perturbation of the new interpretation of Mixup induces multiple known regularization schemes, including label smoothing and reduction of the Lipschitz constant of the estimator. These schemes interact synergistically with each other, resulting in a self calibrated and effective regularization effect that prevents overfitting and overconfident predictions. We corroborate our theoretical analysis with experiments that support our conclusions.
Added
2026-10-01

SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger
Yuting Gao, Jinfeng Liu, Zihan Xu, Tong Wu, Enwei Zhang, Ke Li, Jie Yang, Wei Liu, Xing Sun
Why you should read this
Proposes a relaxed contrastive learning framework that uses fine-grained intra-modal self-similarity and negative-pair disentanglement as soft alignment targets, significantly improving CLIP's zero-shot classification performance on noisy web-scale image-text datasets.
During the preceding biennium, vision-language pre-training has achieved noteworthy success on several downstream tasks. Nevertheless, acquiring high-quality image-text pairs, where the pairs are entirely exclusive of each other, remains a challenging task, and noise exists in the commonly used datasets. To address this issue, we propose SoftCLIP, a novel approach that relaxes the strict one-to-one constraint and achieves a soft cross-modal alignment by introducing a softened target, which is generated from the fine-grained intra-modal self-similarity. The intra-modal guidance is indicative to enable two pairs have some local similarities and model many-to-many relationships between the two modalities. Besides, since the positive still dominates in the softened target distribution, we disentangle the negatives in the distribution to further boost the relation alignment with the negatives in the cross-modal learning. Extensive experiments demonstrate the effectiveness of SoftCLIP. In particular, on ImageNet zero-shot classification task, using CC3M/CC12M as pre-training dataset, SoftCLIP brings a top-1 accuracy improvement of 6.8%/7.2% over the CLIP baseline.
Added
2026-09-26

Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?
Keshigeyan Chandrasegaran, Ngoc-Trung Tran, Yunqing Zhao, Ngai-Man Cheung
Why you should read this
Resolves the contradiction over whether label smoothing aids or hurts knowledge distillation by identifying systematic representation diffusion at high temperatures, offering a practical guideline to distill from label-smoothed teachers using low temperatures.
This work investigates the compatibility between label smoothing (LS) and knowledge distillation (KD). Contemporary findings addressing this thesis statement take dichotomous standpoints: Müller et al. (2019); Shen et al. (2021b). Critically, there is no effort to understand and resolve these contradictory findings, leaving the primal question — to smooth or not to smooth a teacher network? — unanswered. The main contributions of our work are the discovery, analysis and validation of systematic diffusion as the missing concept which is instrumental in understanding and resolving these contradictory findings. This systematic diffusion essentially curtails the benefits of distilling from an LS-trained teacher, thereby rendering KD at increased temperatures ineffective. Our discovery is comprehensively supported by large-scale experiments, analyses and case studies including image classification, neural machine translation and compact student distillation tasks spanning across multiple datasets and teacher-student architectures. Based on our analysis, we suggest practitioners to use an LS-trained teacher with a low-temperature transfer to achieve high performance students. Code and models are available at https://keshik6.github.io/revisiting-ls-kd-compatibility/
Added
2026-09-26

When Does Label Smoothing Help?
Rafael Müller, Simon Kornblith, Geoffrey E. Hinton
Why you should read this
Demonstrates that label smoothing improves network calibration by forming tight representation clusters, while explaining why this structural change erases inter-class information and degrades knowledge distillation.
The generalization and learning speed of a multi-class neural network can often be significantly improved by using soft targets that are a weighted average of the hard targets and the uniform distribution over labels. Smoothing the labels in this way prevents the network from becoming over-confident and label smoothing has been used in many state-of-the-art models, including image classification, language translation and speech recognition. Despite its widespread use, label smoothing is still poorly understood. Here we show empirically that in addition to improving generalization, label smoothing improves model calibration which can significantly improve beam-search. However, we also observe that if a teacher network is trained with label smoothing, knowledge distillation into a student network is much less effective. To explain these observations, we visualize how label smoothing changes the representations learned by the penultimate layer of the network. We show that label smoothing encourages the representations of training examples from the same class to group in tight clusters. This results in loss of information in the logits about resemblances between instances of different classes, which is necessary for distillation, but does not hurt generalization or calibration of the model's predictions.
Added
2026-09-14

Scalable Person Re-identification: A Benchmark
Liang Zheng, Liyue Shen, Lu Tian, Shengjin Wang, Jingdong Wang, Qi Tian
Why you should read this
Presents a scalable global-feature baseline for person re-identification that matches state-of-the-art performance without relying on complex multi-branch designs prone to overfitting on large datasets.
Person re-identification(Re-ID) using deep learning has made great progress in the past few years, but there is one problem that many state-of-the-art Re-ID methods all use a complex network most of which use the structure of multi-branch and multi-loss function. At present, the database used for Person re-identification is relatively small. This complex network structure may bring a problem that although current methods may perform well in the small databases, but there may be some problems of overfitting problem, once applied in the bigger dataset or real scene these complex methods may perform not well. So this paper mainly proposes a new powerful baseline network. This end-to-end network only uses a global feature and does not use multi-branch structure, but achieves state-of-the-art level. The key point is that this network has good improvement potential to adapt to larger datasets and even practical application scenarios.
Added
2026-09-10

The Annotated Transformer
Alexander M. Rush
Why you should read this
A complete, hands-on tutorial for understanding and building the Transformer - the foundational technology behind modern AI systems like ChatGPT and Google Translate - by walking through the actual code line by line rather than just abstract theory.
An annotated version of the paper "Attention is All You Need" in the form of a line-by-line implementation.
Added
2026-02-21
