Adaptive Early-Learning Correction for Segmentation from Noisy Annotations
Sheng LiuKangning LiuWeicheng ZhuYiqiu ShenCarlos Fernandez-Granda
Proposes an adaptive label correction method that prevents deep segmentation models from memorizing noisy annotations by independently monitoring category-specific early-learning dynamics and enforcing spatial consistency.
Semantic segmentation—the automated assignment of a category label to every pixel in an image—is essential for critical applications such as medical imaging and autonomous driving. However, obtaining high-quality pixel-level annotations requires extensive manual labor and domain expertise. In practice, training datasets often contain substantial annotation noise, whether from human reader variation in medical scans or automated generation in weakly supervised pipelines. While deep learning models are known to fit correct labels first before memorizing incorrect ones in standard image classification, this behavior has remained underexplored and poorly handled in pixel-level segmentation tasks.
The article aims to evaluate the learning dynamics of deep segmentation networks trained on noisy pixel-level annotations and to demonstrate a robust training framework that prevents networks from memorizing false annotations. To achieve this, the authors introduce ADaptive Early-Learning corrEction (ADELE), which monitors the training progression of individual semantic categories and adaptively updates noisy labels using confident model predictions before memorization occurs. The method also incorporates a multi-scale consistency regularization term that enforces prediction agreement across rescaled versions of an image, preventing overfitting to spatial annotation errors.
The experimental evaluation demonstrated three primary findings. First, networks trained on noisy segmentations exhibit distinct early-learning and memorization phases, but unlike image classification, these phases occur at different speeds across different object categories due to pixel imbalances. Second, on a thoracic organ computed tomography dataset with simulated human errors, ADELE improved segmentation accuracy from 59.1% to 70.8% mean Intersection over Union (mIoU) at the final epoch, whereas attempting to correct all categories simultaneously severely degraded accuracy to 40.5%. Third, when integrated into weakly-supervised workflows on the standard PASCAL VOC 2012 benchmark, ADELE consistently boosted existing baseline methods and established state-of-the-art performance, achieving up to 69.3% validation mIoU without requiring extra saliency models, and 71.6% when combined with advanced saliency frameworks.
These findings indicate that annotation noise in complex vision tasks can be managed algorithmically without costly complete manual relabeling. Organizations deploying visual AI can reduce manual annotation overhead and better utilize imperfectly labeled datasets or cheaper weak supervision signals. However, the analysis also reveals that treating all categories uniformly during noise correction is counterproductive, making category-adaptive timing essential for operational success.
For practical implementation, teams should integrate adaptive, class-specific label correction and multi-scale consistency into segmentation training pipelines facing label noise. Decision-makers should note that ADELE's effectiveness relies on reasonable initial annotations and may falter when label errors are heavily structured (such as severe category confusion or complete structural omissions). Further testing across additional clinical and industrial domains is recommended to define operational limits before full deployment.
- Paper: Learning From Noisy Labels With Deep Neural Networks: A Survey, Hwanjun Song et al. (2020). Provides a comprehensive survey of methods for learning from noisy labels with deep neural networks, establishing the core taxonomy of noise-handling strategies that contextualizes ADELE's early-learning correction.
- Paper: DivideMix: Learning with Noisy Labels as Semi-supervised Learning, Junnan Li et al. (2020). Introduces the paradigm of leveraging early learning dynamics to split noisy labels and train semi-supervised models, directly underpinning ADELE's adaptive label correction framework.
- Paper: Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach, Giorgio Patrini et al. (2016). Establishes fundamental loss correction methods for deep networks facing label noise, providing key background for algorithmic correction of corrupted training signals.
- Paper: Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training, Yang Zou et al. (2018). Highlights the necessity of class-balanced pseudo-labeling and confidence thresholding in semantic segmentation, directly motivating ADELE's class-adaptive timing.
- Paper: Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation, Liang-Chieh Chen et al. (2018). Defines the standard DeepLabv3+ semantic segmentation architecture and multi-scale context pooling mechanisms that ADELE regularizes against noisy annotations.
- Paper: Learning to Reweight Examples for Robust Deep Learning, Mengye Ren et al. (2018). Demonstrates online meta-weighting strategies to resist noisy labels and class imbalance during training, providing valuable theoretical grounding for dynamic sample and class handling.
- Paper: Learning with Noisy Labels, Nagarajan Natarajan et al. (2013). Presents theoretical foundations and unbiased estimator techniques for learning classifiers under class-conditional label noise.
- Paper: Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor, Hyeokjun Kweon et al. (2023). Extends weakly supervised segmentation by using adversarial game formulation between classifiers and reconstructors to resolve incomplete and noisy localization masks.
- Paper: Token Contrast for Weakly-Supervised Semantic Segmentation, Lixiang Ru et al. (2023). Applies token and patch contrast mechanisms in Vision Transformers to prevent over-smoothing and generate clean pixel-level pseudo-labels in weakly supervised pipelines.
- Paper: Robust Training under Label Noise by Over-parameterization, Sheng Liu et al. (2022). Introduces sparse over-parameterization to mathematically isolate label corruptions during optimization without explicit early-stopping or heuristic thresholding.
- Paper: Selective-Supervised Contrastive Learning with Noisy Labels, Shikun Li et al. (2022). Explores selective-supervised contrastive learning to filter noisy label pairs and build robust representations without needing prior noise-rate estimates.
- Paper: Generative Semantic Segmentation, Jiaqi Chen et al. (2023). Reformulates semantic segmentation as generative mask modeling, providing an alternative paradigm to per-pixel discriminative classification under domain shift.
- Paper: Segment anything in medical images, Jun Ma et al. (2023). Scales promptable medical image segmentation across diverse clinical datasets and modalities using a foundation model approach.
