Adaptive Early-Learning Correction for Segmentation from Noisy Annotations

Sheng LiuKangning LiuWeicheng ZhuYiqiu ShenCarlos Fernandez-Granda

article2022CVPR123 citations

Proposes an adaptive label correction method that prevents deep segmentation models from memorizing noisy annotations by independently monitoring category-specific early-learning dynamics and enforcing spatial consistency.

Listen

Semantic segmentation—the automated assignment of a category label to every pixel in an image—is essential for critical applications such as medical imaging and autonomous driving. However, obtaining high-quality pixel-level annotations requires extensive manual labor and domain expertise. In practice, training datasets often contain substantial annotation noise, whether from human reader variation in medical scans or automated generation in weakly supervised pipelines. While deep learning models are known to fit correct labels first before memorizing incorrect ones in standard image classification, this behavior has remained underexplored and poorly handled in pixel-level segmentation tasks.

The article aims to evaluate the learning dynamics of deep segmentation networks trained on noisy pixel-level annotations and to demonstrate a robust training framework that prevents networks from memorizing false annotations. To achieve this, the authors introduce ADaptive Early-Learning corrEction (ADELE), which monitors the training progression of individual semantic categories and adaptively updates noisy labels using confident model predictions before memorization occurs. The method also incorporates a multi-scale consistency regularization term that enforces prediction agreement across rescaled versions of an image, preventing overfitting to spatial annotation errors.

The experimental evaluation demonstrated three primary findings. First, networks trained on noisy segmentations exhibit distinct early-learning and memorization phases, but unlike image classification, these phases occur at different speeds across different object categories due to pixel imbalances. Second, on a thoracic organ computed tomography dataset with simulated human errors, ADELE improved segmentation accuracy from 59.1% to 70.8% mean Intersection over Union (mIoU) at the final epoch, whereas attempting to correct all categories simultaneously severely degraded accuracy to 40.5%. Third, when integrated into weakly-supervised workflows on the standard PASCAL VOC 2012 benchmark, ADELE consistently boosted existing baseline methods and established state-of-the-art performance, achieving up to 69.3% validation mIoU without requiring extra saliency models, and 71.6% when combined with advanced saliency frameworks.

These findings indicate that annotation noise in complex vision tasks can be managed algorithmically without costly complete manual relabeling. Organizations deploying visual AI can reduce manual annotation overhead and better utilize imperfectly labeled datasets or cheaper weak supervision signals. However, the analysis also reveals that treating all categories uniformly during noise correction is counterproductive, making category-adaptive timing essential for operational success.

For practical implementation, teams should integrate adaptive, class-specific label correction and multi-scale consistency into segmentation training pipelines facing label noise. Decision-makers should note that ADELE's effectiveness relies on reasonable initial annotations and may falter when label errors are heavily structured (such as severe category confusion or complete structural omissions). Further testing across additional clinical and industrial domains is recommended to define operational limits before full deployment.

arXiv: 2110.03740
Cover for Adaptive Early-Learning Correction for Segmentation from Noisy Annotations

Table of Contents

  • 1. Introduction
  • 2. Methodology
  • 2.1. Early learning and memorization in segmentation from noisy annotations
  • 2.2. Adaptive label correction based on early learning
  • 2.3. Multiscale consistency
  • 3. Related work
  • 4. Segmentation on Medical Images with Annotation Noise
  • 5. Noisy Annotations in Weakly-supervised Semantic Segmentation
  • 6. Limitations
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — ADELE for segmentation with noisy pixel annotations

    model/method

    ADELE (ADaptive Early-Learning corrEction) is a training method for semantic segmentation when pixel-level annotations contain errors throughout the training set. It combines two mechanisms: (1) category-specific correction of noisy labels using the segmentation network’s predictions near the end of each category’s early-learning phase, and (2) a multiscale-consistency regularizer that discourages predictions from changing across rescaled versions of the same image. Unlike approaches that correct all classes simultaneously, ADELE detects a separate correction time for every semantic category, allowing different categories to follow their own early-learning and memorization dynamics. The method does not require any completely clean training examples.

  2. Knowl 2 — Category-dependent early learning and memorization

    empirical result

    Segmentation networks trained on noisy pixel labels first learn the correct structure of incorrectly annotated pixels and later memorize the erroneous annotations. For pixels whose available annotation differs from the ground truth, early learning is measured by the IoU between the model prediction and the ground-truth label, denoted IoUel\mathrm{IoU}_{\mathrm{el}}; memorization is measured by the IoU between the model prediction and the incorrect annotation, denoted IoUm\mathrm{IoU}_{\mathrm{m}}. Across medical-image and weakly supervised segmentation experiments, IoUel\mathrm{IoU}_{\mathrm{el}} initially increases and subsequently decreases, whereas IoUm\mathrm{IoU}_{\mathrm{m}} increases as training proceeds. Both the onset of early learning and the onset of memorization occur at different speeds for different semantic categories, rather than simultaneously for all classes.

  3. Knowl 3 — Adaptive timing rule for category-wise label correction

    equation

    For each semantic category, ADELE monitors the training IoU between the model prediction and the currently available noisy annotation. Let tt denote training time measured in epochs, and fit that category’s training-IoU curve by least squares to

    f(t)=a(1−e−btc),f(t)=a\left(1-e^{-b t^{c}}\right),

    where 0<a≤10<a\leq 1, b≥0b\geq 0, and c≥0c\geq 0. Its fitted slope is f′(t)=abce−btctc−1f'(t)=abc e^{-b t^{c}}t^{c-1}. The category becomes eligible for label correction when the relative slope change from the first epoch to the current epoch exceeds the threshold rr:

    ∣f′(1)−f′(t)∣∣f′(1)∣>r.\frac{|f'(1)-f'(t)|}{|f'(1)|}>r.

    ADELE uses r=0.9r=0.9 and evaluates this condition at every epoch. Because the condition is evaluated independently for every category, different classes receive different correction times.

  4. Knowl 4 — Confidence-filtered iterative label correction

    algorithm

    ADELE’s label-correction procedure takes a segmentation network, noisy pixel annotations, and the hyperparameters r=0.9r=0.9 and confidence threshold τ=0.8\tau=0.8 as input. It outputs a network trained on progressively corrected annotations.

    Input: segmentation network, noisy pixel annotations, correction threshold r = 0.9, confidence threshold tau = 0.8
    Initialize the network and retain the noisy annotations as the current training targets
    For each training epoch t:
        Train the network using the current annotations
        For every semantic category:
            Compute the category-specific training IoU against the current noisy annotations
            Fit the category-specific exponential curve f(t) = a(1 - exp(-b t^c)) by least squares
            If |f'(1) - f'(t)| / |f'(1)| > r, mark that category as ready for correction
        For every category marked ready, and at every later epoch:
            Compute the network’s pixelwise class probabilities
            Replace a noisy category label with the model-predicted label only when its confidence exceeds tau
        Continue training with the updated annotations
    Return the trained segmentation network

    The correction is class-adaptive and is performed without access to ground-truth labels during training. In ADELE, predictions averaged across multiple input scales can be used as the predictions for correction.

  5. Knowl 5 — Multiscale-consistency regularization

    equation

    For an input image xx, ADELE evaluates the same segmentation network on s=3s=3 rescaled versions: downscaling by 0.70.7, no rescaling, and upscaling by 1.51.5. Let pk(x)p_k(x) be the class-probability prediction from scale kk, after rescaling its output back to a common spatial resolution, and define the average prediction

    q(x)=1s∑k=1spk(x).q(x)=\frac{1}{s}\sum_{k=1}^{s}p_k(x).

    The multiscale-consistency penalty is the average Kullback–Leibler divergence between each scale-specific prediction and the average prediction:

    LMultiscale(x)=1s∑k=1sKL(pk(x) ∥ q(x)).\mathcal{L}_{\mathrm{Multiscale}}(x)=\frac{1}{s}\sum_{k=1}^{s}\mathrm{KL}\bigl(p_k(x)\,\|\,q(x)\bigr).

    The penalty is applied only when the maximum class-probability entry of q(x)q(x) exceeds the confidence threshold ρ=0.8\rho=0.8. It is added to the cross-entropy loss on the available annotations with weight λ=1\lambda=1. The averaged prediction q(x)q(x) is also used to obtain more reliable labels during ADELE’s correction step.

  6. Knowl 6 — Medical-image noisy-annotation experiment

    experimental setup

    The medical experiment uses the SegTHOR thoracic-organ CT dataset. Each 3D scan is divided into 2D slices resized to 256×256256\times256 pixels, with five labels: esophagus, heart, trachea, aorta, and background. Patient-level separation produces 3,638 training slices, 570 validation slices, and 580 test slices. Training annotations are corrupted by applying random dilation and erosion to the ground-truth organ masks, modeling common human annotation errors; the main corruption level has noisy-to-clean annotation mIoU of 0.60.6. All training annotations are corrupted, while validation and test annotations remain clean. Performance is measured by mean Intersection over Union (mIoU), using a UNet with multiscale inputs as the baseline.

  7. Knowl 7 — Medical segmentation performance and ablations

    data/table

    On SegTHOR, ADELE substantially improves the baseline and remains better after training has continued to the final epoch. The class-adaptive correction mechanism is essential: correcting every category at the same time is harmful. The method also improves across a broad range of corruption levels, with especially substantial gains for moderate noise.

    The first table reports test mIoU (%) over five realizations of the noisy annotations; values are means with standard deviations. “Best val” selects the epoch with the best validation performance, “Max test” is the highest test performance during training, and “Last epoch” evaluates the final epoch.

    Method Best val Max test Last epoch
    Baseline 62.6 ±\pm 2.3 63.3 ±\pm 2.0 59.1 ±\pm 1.3
    ADELE without class-adaptive correction 40.7 ±\pm 2.5 40.7 ±\pm 2.4 40.5 ±\pm 2.3
    ADELE 71.1 ±\pm 0.7 71.2 ±\pm 0.6 70.8 ±\pm 0.7

    The ablation values below are validation mIoU (%) at the last epoch. Multiscale input augmentation improves over single-scale training, multiscale consistency improves it further, and label correction provides the largest gains when combined with the consistency regularizer.

    Dataset Label correction Single scale Multiscale input augmentation Multiscale consistency regularization
    SegTHOR No 58.8 60.7 62.5
    SegTHOR Yes 63.2 69.8 72.2
    PASCAL VOC 2012 No 64.5 65.5 66.7
    PASCAL VOC 2012 Yes 65.6 67.3 69.3
  8. Knowl 8 — Weakly supervised semantic segmentation evaluation

    experimental setup

    ADELE is also evaluated in the standard weakly supervised semantic segmentation pipeline, where image-level labels are used by a classification model to generate noisy pixel-level annotations and a segmentation model is then trained on those annotations. The dataset is PASCAL VOC 2012 with 21 classes including background, containing 1,464 training images, 1,449 validation images, and 1,456 test images. Following the experimental protocol used by the authors, training uses an augmented set of 10,582 images. ADELE is applied to pixel annotations generated by AffinityNet, SEAM, and ICD, without external datasets or external saliency maps. The segmentation model uses the same inference pipeline as SEAM, including multiscale inference and a fully connected CRF.

  9. Knowl 9 — PASCAL VOC performance and state-of-the-art comparison

    data/table

    ADELE consistently improves the AffinityNet, SEAM, and ICD weakly supervised segmentation baselines on both PASCAL VOC validation and test sets. The combinations ADELE+SEAM and ADELE+ICD achieve the strongest results among the listed methods, despite using only image-level supervision. With the external-saliency-based NSROM method, ADELE reaches 71.6 mIoU on validation and 72.0 mIoU on test, reported as state-of-the-art for the ResNet segmentation-backbone setting.

    The following comparison reports mIoU (%) on PASCAL VOC 2012. The first seven methods are prior systems using a ResNet-101 backbone; the final three are ADELE combinations using a ResNet-38 backbone.

    DSRG ICD SCE AffinityNet SSDD SEAM CONTA ADELE+AffinityNet ADELE+SEAM ADELE+ICD
    Val 61.4 64.1 66.1 61.7 64.9 64.5 66.1 64.8 69.3 68.6
    Test 63.2 64.3 65.9 63.7 65.5 65.7 66.7 65.5 68.8 68.9

    Category-wise analysis shows gains for most classes, but little or no improvement for categories whose initial segmentation is already very poor.

  10. Knowl 10 — Dependence on initial annotation quality and noise structure

    limitation

    ADELE’s success depends partly on the quality of the initial pixel-level annotations. When those annotations are very inaccurate, ADELE may provide only a marginal improvement or may reduce performance. Highly structured errors can be particularly problematic: if the noisy labels contain insufficient information to reveal the correct object structure, the network may not undergo a useful early-learning phase, so prediction-based correction cannot recover the missing or misclassified regions. Examples include an initial mask that completely encompasses a bicycle or a chair systematically labeled as a sofa.

Coverage note — Appendix-only hyperparameter sweeps, additional qualitative examples, and category-wise visualization details were omitted because they support the main method and reported results without adding separate load-bearing contributions.

References

  1. 1.Rolf Adams and Leanne Bischof. Seeded region growing. IEEE Transactions on Pattern Analysis and Machine Intelligence, 16(6):641–647, 1994.
  2. 2.Jiwoon Ahn, Sunghyun Cho, and Suha Kwak. Weakly supervised learning of instance segmentation with inter-pixel relations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  3. 3.Jiwoon Ahn and Suha Kwak. Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  4. 4.Devansh Arpit, Stanisław Jastrz˛ebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien. A closer look at memorization in deep networks. In International Conference on Machine Learning, 2017.
  5. 5.Andrew J Asman and Bennett A Landman. Formulating spatially varying performance in the statistical fusion framework. IEEE Transactions on Medical Imaging, 31(6):1326–1336, 2012.
  6. 6.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A Raffel. Mixmatch: A holistic approach to semi-supervised learning. In Advances in Neural Information Processing Systems, 2019.
  7. 7.Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, and Ming-Hsuan Yang. Weakly-supervised semantic segmentation via sub-category exploration. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  8. 8.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv preprint arXiv:1412.7062, 2014.
  9. 9.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(4):834–848, 2017.
  10. 10.Hao Cheng, Zhaowei Zhu, Xingyu Li, Yifei Gong, Xing Sun, and Yang Liu. Learning with instance-dependent label noise: A sample sieve approach. arXiv preprint arXiv:2010.02347, 2020.
  11. 11.Jifeng Dai, Kaiming He, and Jian Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, 2015.
  12. 12.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Ieee, 2009.
  13. 13.Mark Everingham, SM Ali Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge: A retrospective. International Journal of Computer Vision, 111(1):98–136, 2015.
  14. 14.Junsong Fan, Zhaoxiang Zhang, Chunfeng Song, and Tieniu Tan. Learning integral objects with intra-class discriminator for weakly-supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  15. 15.Geoff French, Timo Aila, Samuli Laine, Michal Mackiewicz, and Graham Finlayson. Semi-supervised semantic segmentation needs strong, high-dimensional perturbations. arXiv preprint arXiv:1906.01916, 2019.
  16. 16.Bharath Hariharan, Pablo Arbeláez, Lubomir Bourdev, Subhransu Maji, and Jitendra Malik. Semantic contours from inverse detectors. In Proceedings of the IEEE International Conference on Computer Vision, 2011.
  17. 17.Mohammad Havaei, Axel Davy, David Warde-Farley, Antoine Biard, Aaron Courville, Yoshua Bengio, Chris Pal, Pierre-Marc Jodoin, and Hugo Larochelle. Brain tumor segmentation with deep neural networks. Medical Image Analysis, 35:18–31, 2017.
  18. 18.Zilong Huang, Xinggang Wang, Jiasi Wang, Wenyu Liu, and Jingdong Wang. Weakly-supervised semantic segmentation network with deep seeded region growing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  19. 19.Huaizu Jiang, Jingdong Wang, Zejian Yuan, Yang Wu, Nanning Zheng, and Shipeng Li. Salient object detection: A discriminative regional feature integration approach. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2013.
  20. 20.Peng-Tao Jiang, Qibin Hou, Yang Cao, Ming-Ming Cheng, Yunchao Wei, and Hong-Kai Xiong. Integral object mining via online attention accumulation. In Proceedings of the IEEE International Conference on Computer Vision, 2019.
  21. 21.Zequn Jie, Yunchao Wei, Xiaojie Jin, Jiashi Feng, and Wei Liu. Deep self-taught learning for weakly supervised object localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  22. 22.Eytan Kats, Jacob Goldberger, and Hayit Greenspan. A soft staple algorithm combined with anatomical knowledge. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 510–517. Springer, 2019.
  23. 23.Jongmok Kim, Jooyoung Jang, and Hyunwoo Park. Structured consistency loss for semi-supervised semantic segmentation. arXiv preprint arXiv:2001.04647, 2020.
  24. 24.Alexander Kolesnikov and Christoph H Lampert. Seed, expand and constrain: Three principles for weakly-supervised image segmentation. In Proceedings of the European Conference on Computer Vision. Springer, 2016.
  25. 25.Philipp Krähenbühl and Vladlen Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. Advances in Neural Information Processing Systems, 24, 2011.
  26. 26.Samuli Laine and Timo Aila. Temporal ensembling for semi-supervised learning. In International Conference on Learning Representations, 2018.
  27. 27.Zoé Lambert, Caroline Petitjean, Bernard Dubray, and Su Kuan. Segthor: Segmentation of thoracic organs at risk in ct images. In 2020 Tenth International Conference on Image Processing Theory, Tools and Applications (IPTA). IEEE, 2020.
  28. 28.Junnan Li, Richard Socher, and Steven C.H. Hoi. Dividemix: Learning with noisy labels as semi-supervised learning. In International Conference on Learning Representations, 2020.
  29. 29.Qizhu Li, Anurag Arnab, and Philip HS Torr. Weakly-and semi-supervised panoptic segmentation. In Proceedings of the European Conference on Computer Vision, 2018.
  30. 30.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  31. 31.Di Lin, Yuanfeng Ji, Dani Lischinski, Daniel Cohen-Or, and Hui Huang. Multi-scale context intertwining for semantic segmentation. In Proceedings of the European Conference on Computer Vision, 2018.
  32. 32.Kangning Liu, Yiqiu Shen, Nan Wu, Jakub Chledowski, Carlos Fernandez-Granda, and Krzysztof J Geras. Weakly-supervised high-resolution segmentation of mammography images for breast cancer diagnosis. In Medical Imaging with Deep Learning, 2021.
  33. 33.Sheng Liu, Jonathan Niles-Weed, Narges Razavian, and Carlos Fernandez-Granda. Early-learning regularization prevents memorization of noisy labels. Advances in Neural Information Processing Systems, 2020.
  34. 34.Yaoru Luo, Guole Liu, Wenjing Li, Yuanhao Guo, and Ge Yang. Deep neural networks learn meta-structures to segment fluorescence microscopy images. arXiv preprint arXiv:2103.11594, 2021.
  35. 35.Shaobo Min, Xuejin Chen, Zheng-Jun Zha, Feng Wu, and Yongdong Zhang. A two-stream mutual attention network for semi-supervised biomedical segmentation with noisy labels. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 4578–4585, 2019.
  36. 36.Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(8):1979–1993, 2018.
  37. 37.Scott E. Reed, Honglak Lee, Dragomir Anguelov, Christian Szegedy, Dumitru Erhan, and Andrew Rabinovich. Training deep neural networks on noisy labels with bootstrapping. CoRR, abs/1412.6596, 2015.
  38. 38.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2015.
  39. 39.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  40. 40.Jo Schlemper, Ozan Oktay, Liang Chen, Jacqueline Matthew, Caroline Knight, Bernhard Kainz, Ben Glocker, and Daniel Rueckert. Attention-gated networks for improving ultrasound scan plane detection. arXiv preprint arXiv:1804.05338, 2018.
  41. 41.Wataru Shimoda and Keiji Yanai. Self-supervised difference detection for weakly-supervised semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, 2019.
  42. 42.Yucheng Shu, Xiao Wu, and Weisheng Li. Lvc-net: Medical image segmentation with noisy label based on local visual cues. In International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2019.
  43. 43.Kihyuk Sohn, David Berthelot, Chun-Liang Li, Zizhao Zhang, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. arXiv preprint arXiv:2001.07685, 2020.
  44. 44.Chunfeng Song, Yan Huang, Wanli Ouyang, and Liang Wang. Box-driven class-wise region masking and filling rate guided loss for weakly supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  45. 45.Guolei Sun, Wenguan Wang, Jifeng Dai, and Luc Van Gool. Mining cross-image semantics for weakly supervised semantic segmentation. In Proceedings of the European Conference on Computer Vision. Springer, 2020.
  46. 46.Daiki Tanaka, Daiki Ikami, Toshihiko Yamasaki, and Kiyoharu Aizawa. Joint optimization framework for learning with noisy labels. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  47. 47.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Advances in Neural Information Processing Systems, 2017.
  48. 48.Michael Treml, José Arjona-Medina, Thomas Unterthiner, Rupesh Durgesh, Felix Friedmann, Peter Schuberth, Andreas Mayr, Martin Heusel, Markus Hofmarcher, Michael Widrich, et al. Speeding up semantic segmentation for autonomous driving. In MLITS, NIPS Workshop, 2016.
  49. 49.Paul Vernaza and Manmohan Chandraker. Learning random-walk label propagation for weakly-supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  50. 50.Guotai Wang, Xinglong Liu, Chaoping Li, Zhiyong Xu, Jiugen Ruan, Haifeng Zhu, Tao Meng, Kang Li, Ning Huang, and Shaoting Zhang. A noise-robust framework for automatic segmentation of covid-19 pneumonia lesions from ct images. IEEE Transactions on Medical Imaging, 39(8):2653–2663, 2020.
  51. 51.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  52. 52.Yude Wang, Jie Zhang, Meina Kan, Shiguang Shan, and Xilin Chen. Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  53. 53.Yunchao Wei, Jiashi Feng, Xiaodan Liang, Ming-Ming Cheng, Yao Zhao, and Shuicheng Yan. Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, July 2017.
  54. 54.Yunchao Wei, Xiaodan Liang, Yunpeng Chen, Xiaohui Shen, Ming-Ming Cheng, Jiashi Feng, Yao Zhao, and Shuicheng Yan. Stc: A simple to complex framework for weakly-supervised semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(11):2314–2320, 2016.
  55. 55.Yunchao Wei, Huaxin Xiao, Honghui Shi, Zequn Jie, Jiashi Feng, and Thomas S Huang. Revisiting dilated convolution: A simple approach for weakly-and semi-supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  56. 56.Zifeng Wu, Chunhua Shen, and Anton Van Den Hengel. Wider or deeper: Revisiting the resnet model for visual recognition. Pattern Recognition, 90:119–133, 2019.
  57. 57.Xiaobo Xia, Tongliang Liu, Bo Han, Chen Gong, Nannan Wang, Zongyuan Ge, and Yi Chang. Robust early-learning: Hindering the memorization of noisy labels. In ICLR, 2021.
  58. 58.Maoke Yang, Kun Yu, Chi Zhang, Zhiwei Li, and Kuiyuan Yang. Denseaspp for semantic segmentation in street scenes. In CVPR, 2018.
  59. 59.Yazhou Yao, Tao Chen, Guosen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu, Zhenmin Tang, and Jian Zhang. Non-salient region object mining for weakly supervised semantic segmentation. arXiv preprint arXiv:2103.14581, 2021.
  60. 60.Kun Yi and Jianxin Wu. Probabilistic end-to-end noise correction for learning with noisy labels. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  61. 61.Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, and Kaizhu Huang. Reliability does matter: An end-to-end weakly supervised semantic segmentation approach. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020.
  62. 62.Dong Zhang, Hanwang Zhang, Jinhui Tang, Xiansheng Hua, and Qianru Sun. Causal intervention for weakly-supervised semantic segmentation. arXiv preprint arXiv:2009.12547, 2020.
  63. 63.Le Zhang, Ryutaro Tanno, Mou-Cheng Xu, Chen Jin, Joseph Jacob, Olga Ciccarelli, Frederik Barkhof, and Daniel C Alexander. Disentangling human error from the ground truth in segmentation of medical images. arXiv preprint arXiv:2007.15963, 2020.
  64. 64.Tianyi Zhang, Guosheng Lin, Weide Liu, Jianfei Cai, and Alex Kot. Splitting vs. merging: Mining object regions with discrepancy and intersection loss for weakly supervised semantic segmentation. In Proceedings of the European Conference on Computer Vision, 2020.
  65. 65.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  66. 66.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  67. 67.Tianyi Zhou, Shengjie Wang, and Jeff Bilmes. Robust curriculum learning: From clean label detection to noisy label self-correction. In International Conference on Learning Representations, 2020.

Citation

MLA
Liu, S., et al. “Adaptive Early-Learning Correction for Segmentation from Noisy Annotations”. arXiv, 2021, http://arxiv.org/abs/2110.03740v2.
APA
Liu, S., Liu, K., Zhu, W., Shen, Y., & Fernandez-Granda, C. (2021). Adaptive Early-Learning Correction for Segmentation from Noisy Annotations. arXiv. http://arxiv.org/abs/2110.03740v2
Chicago
Liu, S., K. Liu, W. Zhu, Y. Shen, and C. Fernandez-Granda. 2021. “Adaptive Early-Learning Correction for Segmentation from Noisy Annotations”. arXiv. http://arxiv.org/abs/2110.03740v2.
Harvard
Liu, S. et al. (2021) “Adaptive Early-Learning Correction for Segmentation from Noisy Annotations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2110.03740v2.
Vancouver
1. Liu S, Liu K, Zhu W, Shen Y, Fernandez-Granda C (2021) Adaptive Early-Learning Correction for Segmentation from Noisy Annotations. arXiv

BibTeX

@article{liu2021adaptive,
  title = {Adaptive Early-Learning Correction for Segmentation from Noisy Annotations},
  author = {Liu, Sheng and Liu, Kangning and Zhu, Weicheng and Shen, Yiqiu and Fernandez-Granda, Carlos},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2110.03740v2},
  eprint = {2110.03740}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE