keyword
noisy supervision
Noisy supervision is a machine learning training setting in which the supervisory signals provided to guide model optimization, such as target labels or reward feedback, contain errors, inaccuracies, or corrupted annotations. This form of imperfect supervision commonly arises from crowdsourced labeling, automated heuristics, sensor noise, or approximate reward functions rather than verified ground-truth annotations. Because expressive models like deep neural networks can easily memorize incorrect targets and suffer degraded generalization, learning under noisy supervision typically relies on robust loss functions, sample reweighting or filtering, noise transition modeling, and specialized regularization techniques designed to mitigate the influence of incorrect feedback and uncover true underlying patterns.
3 items

Identifiability of Label Noise Transition Matrix
Yang Liu, Hao Cheng, Kun Zhang
Why you should read this
Establishes the theoretical conditions required to identify instance-dependent label noise transition matrices using Kruskal's identifiability theorem, proving why multiple noisy labels are necessary and demonstrating how disentangled representations improve matrix estimation without clean labels.
The noise transition matrix plays a central role in the problem of learning with noisy labels. Among many other reasons, a large number of existing solutions rely on the knowledge of it. Identifying and estimating the transition matrix without ground truth labels is a critical and challenging task. When label noise transition depends on each instance, the problem of identifying the instance-dependent noise transition matrix becomes substantially more challenging. Despite recently proposed solutions for learning from instance-dependent noisy labels, the literature lacks a unified understanding of when such a problem remains identifiable. The goal of this paper is to characterize the identifiability of the label noise transition matrix. Building on Kruskal’s identifiability results, we are able to show the necessity of multiple noisy labels in identifying the noise transition matrix at the instance level. We further instantiate the results to explain the successes of the state-of-the-art solutions and how additional assumptions alleviated the requirement of multiple noisy labels. Our result reveals that disentangled features improve identification. This discovery led us to an approach that improves the estimation of the transition matrix using properly disentangled features. Code is available at https://github.com/UCSC-REAL/Identifiability.
Added
2026-10-03

Co-teaching: Robust training of deep neural networks with extremely noisy labels
Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, Masashi Sugiyama
Why you should read this
Proposes Co-teaching, a training framework where two peer neural networks iteratively select and exchange likely clean samples to prevent deep models from memorizing heavy label noise.
Deep learning with noisy labels is practically challenging, as the capacity of deep models is so high that they can totally memorize these noisy labels sooner or later during training. Nonetheless, recent studies on the memorization effects of deep neural networks show that they would first memorize training data of clean labels and then those of noisy labels. Therefore in this paper, we propose a new deep learning paradigm called Co-teaching for combating with noisy labels. Namely, we train two deep neural networks simultaneously, and let them teach each other given every mini-batch: firstly, each network feeds forward all data and selects some data of possibly clean labels; secondly, two networks communicate with each other what data in this mini-batch should be used for training; finally, each network back propagates the data selected by its peer network and updates itself. Empirical results on noisy versions of MNIST, CIFAR-10 and CIFAR-100 demonstrate that Co-teaching is much superior to the state-of-the-art methods in the robustness of trained deep models.
Added
2026-09-14

When Can LLMs Learn to Reason with Weak Supervision?
Salman Rahman, Jingyan Shen, Anna Mordvina, Hamid Palangi, Saadia Gabriel, Pavel Izmailov
Why you should read this
Reveals that effective generalization in LLMs trained with weak supervision is governed by prolonged pre-saturation training reward dynamics, which is predicted by reasoning faithfulness and enabled by supervised fine-tuning on explicit reasoning traces plus domain-specific continual pre-training.
Large language models have achieved significant reasoning improvements through reinforcement learning with verifiable rewards (RLVR). Yet as model capabilities grow, constructing high-quality reward signals becomes increasingly difficult, making it essential to understand when RLVR can succeed under weaker forms of supervision. We conduct a systematic empirical study across diverse model families and reasoning domains under three weak supervision settings: scarce data, noisy rewards, and self-supervised proxy rewards. We find that generalization is governed by training reward saturation dynamics: models that generalize exhibit a prolonged pre-saturation phase during which training reward and downstream performance climb together, while models that saturate rapidly memorize rather than learn. We identify reasoning faithfulness, defined as the extent to which intermediate steps logically support the final answer, as the pre-RL property that predicts which regime a model falls into, while output diversity alone is uninformative. Motivated by these findings, we disentangle the contributions of continual pre-training and supervised fine-tuning, finding that SFT on explicit reasoning traces is necessary for generalization under weak supervision, while continual pre-training on domain data amplifies the effect. Applied together to Llama3.2-3B-Base, these interventions enable generalization across all three settings where the base model previously failed.
Added
2026-05-14
License
Published with permission
