Estimating Instance-dependent Bayes-label Transition Matrix using a Deep Neural Network

Shuo YangErkun YangBo HanYang LiuMin XuGang NiuTongliang Liu

article2022ICML65 citations

Proposes modeling instance-dependent label noise by estimating transitions from Bayes optimal labels to noisy labels using a deep neural network, shrinking the search space and improving classification accuracy on noisy datasets.

Listen

Modern machine learning applications rely heavily on massive datasets collected through web scraping, online queries, and crowdsourcing platforms. Because manual verification at scale is prohibitively expensive, these real-world datasets inevitably contain substantial amounts of mislabeled data. Deep neural networks tend to memorize these incorrect labels, which severely degrades model performance and reliability in deployment. A particularly challenging and realistic scenario is instance-dependent label noise, where the likelihood of a label error depends directly on the specific features or quality of an individual sample. Addressing this issue is critical for organizations deploying artificial intelligence systems where data quality cannot be manually guaranteed.

The article evaluates a new framework for training accurate deep learning classifiers on datasets corrupted by instance-dependent label noise. Specifically, it demonstrates that modeling the transition probabilities from optimal target categories (termed Bayes optimal labels) to observed noisy labels enables a deep neural network to directly estimate instance-specific noise patterns without relying on clean labels or rigid hand-crafted assumptions.

To achieve this, the approach identifies a subset of high-confidence training examples whose optimal labels can be mathematically inferred from the noisy data. A dedicated deep neural network is then trained on these distilled examples to estimate the transition behavior from optimal to noisy labels for any given input. Finally, this transition network is fixed and used to correct the loss function while training the main classification model across the entire noisy dataset. The authors validated this method using three benchmark image datasets corrupted with synthetic noise levels ranging from 10% to 50% across multiple runs, as well as one large-scale real-world dataset containing over one million noisily labeled clothing images.

The key findings demonstrate significant performance gains across all testing environments. First, the proposed method consistently outperformed leading baseline approaches across all noise levels on the synthetic benchmark datasets. Second, the performance advantage widened substantially as noise severity increased; for instance, on the CIFAR-10 image benchmark, the method achieved a 5.83 percentage point lead over the top-performing baseline at 10% noise, expanding to a 7.01 percentage point lead at 40% noise. Similar gains occurred on the SVHN digit dataset, where the performance margin grew from 2.60 percentage points at 10% noise to 7.14 percentage points at 50% noise. Third, on the real-world clothing dataset, the framework reached 73.39% test accuracy when combined with matrix revision, exceeding all baseline techniques.

These results indicate that estimating noise transitions based on optimal labels reduces mathematical uncertainty and provides better generalization compared to traditional methods that attempt to model clean distributions directly. For organizations, this approach mitigates the risk of model failure and reduces data cleaning costs by allowing systems to train effectively on imperfect, cheaply gathered datasets without requiring manual relabeling.

Organizations handling large, imperfectly labeled data should consider adopting parametric transition models and loss-correction techniques when training deep networks. To improve outcomes, technical teams can combine this method with existing matrix revision techniques to further boost accuracy. Prior to large-scale production deployment, engineering teams should conduct pilot tests on domain-specific data to tune the noise-bound threshold and verify stability across varying operational conditions.

The study operates under the assumption of bounded noise rates, meaning that error probabilities do not exceed a certain maximum threshold. While the method demonstrates high empirical reliability on standard image recognition benchmarks, caution is warranted when applying it to domains with severe label corruption exceeding theoretical bounds or data modalities outside standard image classification.

  • Paper: Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach, Giorgio Patrini et al. (2016). Its forward loss correction uses a label-transition matrix to adjust training, the direct foundation this paper adapts from class-level noise rates to instance-specific transitions.
  • Paper: Learning with Noisy Labels, Nagarajan Natarajan et al. (2013). Its unbiased-loss and label-dependent-cost corrections establish the earlier theory for learning under noisy labels that contextualizes the source’s transition-based loss correction.

No sufficiently relevant recommendations were found.

Cover for Estimating Instance-dependent Bayes-label Transition Matrix using a Deep Neural Network

Abstract

In label-noise learning, estimating the transition matrix is a hot topic as the matrix plays an important role in building statistically consistent classifiers. Traditionally, the transition from clean labels to noisy labels (i.e., clean-label transition matrix (CLTM)) has been widely exploited to learn a clean label classifier by employing the noisy data. Motivated by that classifiers mostly output Bayes optimal labels for prediction, in this paper, we study to directly model the transition from Bayes optimal labels to noisy labels (i.e., Bayes-label transition matrix (BLTM)) and learn a classifier to predict Bayes optimal labels. Note that given only noisy data, it is ill-posed to estimate either the CLTM or the BLTM. But favorably, Bayes optimal labels have less uncertainty compared with the clean labels, i.e., the class posteriors of Bayes optimal labels are one-hot vectors while those of clean labels are not. This enables two advantages to estimate the BLTM, i.e., (a) a set of examples with theoretically guaranteed Bayes optimal labels can be collected out of noisy data; (b) the feasible solution space is much smaller. By exploiting the advantages, we estimate the BLTM parametrically by employing a deep neural network, leading to better generalization and superior classification performance.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 4. Method
  • 4.1. Bayes-label transition matrix
  • 4.2. Collecting Bayes Optimal Labels
  • 4.3. Bayes Label Transition Network
  • 4.4. Classifier Training with Forward Correction
  • 5. Experiments
  • 5.1. Experiment setup
  • 5.1.1. DATASETS
  • 5.1.2. IMPLEMENTATION DETAILS
  • 5.1.3. COMPARISON METHODS
  • 5.2. Comparison with the State-of-the-Arts
  • 6. Conclusion
  • 7. Acknowledgements
  • References

Knowls

  1. Knowl 1 — Bayes-label transition matrix and posterior relation

    definition

    Let XX be an input instance, with clean label YY and noisy label Y~\widetilde Y, and let Y∗=arg⁡max⁡y∈{1,…,C}P(Y=y∣X)Y^* = \arg\max_{y\in\{1,\ldots,C\}} P(Y=y\mid X) be its Bayes-optimal label among CC classes. The Bayes-label transition matrix (BLTM) is the instance-dependent matrix whose entry in row ii, column jj is

    Ti,j∗(x)=P(Y~=j∣Y∗=i,X=x).T^*_{i,j}(x)=P(\widetilde Y=j\mid Y^*=i,X=x).

    Each row is a probability distribution over noisy labels. If q(x)q(x) is the column vector of probabilities for Bayes-optimal labels and s(x)s(x) is the column vector of noisy-label probabilities, then s(x)=T∗(x)⊤q(x)s(x)=T^*(x)^\top q(x). When T∗(x)T^*(x) is invertible, q(x)=(T∗(x)⊤)−1s(x)q(x)=(T^*(x)^\top)^{-1}s(x). The proposal models this transition rather than the conventional clean-label-to-noisy-label transition: Bayes-optimal labels are deterministic given the instance, whereas clean labels can be stochastic under the clean class posterior.

  2. Knowl 2 — Distilling theoretically guaranteed Bayes-optimal labels

    theoretical result

    Under the paper's bounded instance-dependent noise setting, let ρmax⁡<1\rho_{\max}<1 be an upper bound on the noise rates and let ηy(x)=P(Y~=y∣X=x)\eta_{y}(x)=P(\widetilde Y=y\mid X=x) be the noisy class posterior for class yy. An instance whose noisy posterior for a class exceeds the threshold

    ηy(x)>1+ρmax⁡2\eta_y(x)>\frac{1+\rho_{\max}}{2}

    is collected with inferred Bayes-optimal label yy. The resulting distilled example retains the original noisy label as well as the inferred label, which can differ. In practice the method evaluates this criterion using an estimated noisy posterior; the guarantee applies when the threshold condition is correctly identified under the bounded-noise assumptions.

  3. Knowl 3 — Neural estimation of the instance-dependent BLTM

    model/method

    A Bayes-label transition network takes an input instance xx and, with parameters θ\theta, outputs an estimated C×CC\times C transition matrix T^∗(x;θ)\widehat T^*(x;\theta). Its rows represent conditional distributions of noisy labels given each Bayes-optimal class. The network is trained on mm distilled examples (xi,y~i,y^i∗)(x_i,\widetilde y_i,\widehat y_i^*), where the two labels are encoded as one-hot row vectors in R1×C\mathbb R^{1\times C}. It minimizes cross-entropy between the observed noisy label and the noisy-label distribution implied by the inferred Bayes label:

    R^1(θ)=−1m∑i=1my~ilog⁡ ⁣(y^i∗T^∗(xi;θ)).\widehat R_1(\theta)=-\frac{1}{m}\sum_{i=1}^{m}\widetilde y_i\log\!\left(\widehat y_i^*\widehat T^*(x_i;\theta)\right).

    The logarithm and product are interpreted coordinatewise and the resulting classwise terms are summed. An example with inferred Bayes class ii directly trains row ii of its predicted transition matrix; examples from other classes train their corresponding rows, while network parameters can be shared across classes. The paper's rationale is that using deterministic Bayes labels reduces the transition-estimation hypothesis space relative to estimating transitions from stochastic clean labels. The transition network is trained only on distilled examples, so its use on other training examples depends on generalization of noise-causing patterns; the authors report that this generalization works empirically.

  4. Knowl 4 — Forward correction for learning the Bayes classifier

    model/method

    Let f(x;w)∈R1×Cf(x;w)\in\mathbb R^{1\times C} be a classifier's predicted distribution over Bayes-optimal labels, parameterized by ww, and let the fixed, learned BLTM be T^∗(x;θ)\widehat T^*(x;\theta). The implied noisy-label distribution is f(x;w)T^∗(x;θ)f(x;w)\widehat T^*(x;\theta). Training on all nn noisy examples minimizes

    R^2(w)=−1n∑i=1ny~ilog⁡ ⁣(f(xi;w)T^∗(xi;θ)),\widehat R_2(w)=-\frac{1}{n}\sum_{i=1}^{n}\widetilde y_i\log\!\left(f(x_i;w)\widehat T^*(x_i;\theta)\right),

    where y~i\widetilde y_i is the one-hot noisy label. The transition network is held fixed while this loss trains the classifier. The paper invokes the forward-correction consistency result: if the estimated transition matrix is unbiased, the minimizer of the noisy-distribution objective coincides with the minimizer of cross-entropy under the Bayes-optimal-label distribution.

  5. Knowl 5 — Two-stage training procedure and experimental settings

    algorithm

    The method first obtains a set of high-confidence Bayes labels and uses them to train an instance-dependent transition network; it then freezes that network and trains a classifier on the full noisy dataset using forward correction. The inputs are a noisy training set, an upper bound ρmax⁡\rho_{\max} for the distillation threshold, and a classification network; the output is a classifier predicting Bayes-optimal labels.

    Input: Noisy training set, class count CC, and distillation bound ρmax⁡\rho_{\max}
    Output: Classifier trained to predict Bayes-optimal labels
    1. Warm up a probabilistic classifier on the noisy training set for 5 epochs. Use SGD with momentum 0.9, batch size 128, and learning rate 0.01; for Clothing1M use an ImageNet-pretrained ResNet-50 and learning rate 0.001.
    2. Estimate the noisy class posterior η^y(x)\widehat\eta_y(x) with the warmed-up classifier.
    3. Initialize an empty distilled set. For every noisy example (x,y~)(x,\widetilde y) and class yy, if η^y(x)>(1+ρmax⁡)/2\widehat\eta_y(x)>(1+\rho_{\max})/2, add (x,y~,y^∗=y)(x,\widetilde y,\widehat y^*=y) to the set.
    4. Initialize the Bayes-label transition network using the classifier's architecture, with its final layer adapted to output a C×CC\times C matrix. Train it on the distilled set for 5 epochs using the cross-entropy objective for noisy labels, SGD, momentum 0.9, and learning rate 0.01.
    5. Freeze the transition network. Train the classifier on the entire noisy training set with forward-corrected cross-entropy for 50 epochs on F-MNIST, CIFAR-10, and SVHN, or 10 epochs on Clothing1M. Use Adam with learning rate 5×10−75\times10^{-7} and weight decay 10−410^{-4}.
    6. Return the trained classifier.

    In the experiments, the distillation bound was manually set to ρmax⁡=0.3\rho_{\max}=0.3. For synthetic-noise generation, the separate upper bound was 0.60.6.

  6. Knowl 6 — Evaluation datasets and protocol

    experimental setup

    The evaluation used three synthetically corrupted datasets—F-MNIST (60,000 training and 10,000 test images), CIFAR-10 (50,000 training and 10,000 test images), and SVHN (73,257 training and 26,032 test images)—and Clothing1M, which has 1 million noisy training images and 10,000 cleanly labeled test images. Synthetic corruption followed bounded instance-dependent noise settings with noise levels of 10%, 20%, 30%, 40%, and 50%; each synthetic-dataset experiment was repeated five times. Ten percent of noisy training examples were held out for model selection. Clothing1M training and validation used noisy examples only. No data augmentation was used. The reported synthetic-dataset metrics are classification accuracy in percent, given as mean and standard deviation.

  7. Knowl 7 — F-MNIST accuracy under instance-dependent noise

    empirical result

    On F-MNIST, the paper compares the proposed BLTM method, its transition-matrix-revised variant BLTM-V, and PTD, the closely related instance-dependent transition baseline, across increasing synthetic noise rates. Accuracies are percentages, reported as mean ±\pm standard deviation; values below follow noise rates 10%, 20%, 30%, 40%, and 50% in order. PTD achieved 92.03±0.3392.03\pm0.33, 90.78±0.6490.78\pm0.64, 87.86±0.7887.86\pm0.78, 79.46±1.5879.46\pm1.58, and 73.38±2.2573.38\pm2.25. BLTM achieved 96.06±0.7196.06\pm0.71, 94.97±0.3394.97\pm0.33, 91.47±1.3691.47\pm1.36, 82.88±2.7282.88\pm2.72, and 76.35±3.7976.35\pm3.79. BLTM-V achieved 96.93±0.3196.93\pm0.31, 95.55±0.5995.55\pm0.59, 92.24±1.8792.24\pm1.87, 83.43±1.7283.43\pm1.72, and 76.89±4.2676.89\pm4.26. Both proposed variants outperform PTD at every tested noise level, and matrix revision improves BLTM at each level.

  8. Knowl 8 — CIFAR-10 accuracy under instance-dependent noise

    empirical result

    On CIFAR-10, the comparison tests PTD, BLTM, and BLTM-V at synthetic instance-dependent noise levels of 10%, 20%, 30%, 40%, and 50%, in that order. Classification accuracies are percentages reported as mean ±\pm standard deviation. PTD scored 76.33±0.3876.33\pm0.38, 76.05±1.7276.05\pm1.72, 75.42±1.3375.42\pm1.33, 65.92±2.3365.92\pm2.33, and 56.63±1.8856.63\pm1.88. BLTM scored 81.73±0.5681.73\pm0.56, 80.26±0.6380.26\pm0.63, 77.69±1.3777.69\pm1.37, 71.96±2.2771.96\pm2.27, and 59.15±3.1159.15\pm3.11. BLTM-V scored 82.16±1.0182.16\pm1.01, 80.37±1.9880.37\pm1.98, 78.82±1.0778.82\pm1.07, 72.93±4.0072.93\pm4.00, and 60.33±5.2960.33\pm5.29. Both proposed variants exceed PTD at all five levels; the advantage remains pronounced at 40% and 50% noise, the more challenging conditions.

  9. Knowl 9 — SVHN accuracy under instance-dependent noise

    empirical result

    On SVHN, PTD, BLTM, and BLTM-V were evaluated at synthetic instance-dependent noise levels of 10%, 20%, 30%, 40%, and 50%, in that order. Classification accuracies are percentages reported as mean ±\pm standard deviation. PTD achieved 93.77±0.3393.77\pm0.33, 92.59±1.0792.59\pm1.07, 89.64±1.9889.64\pm1.98, 83.56±2.2183.56\pm2.21, and 71.57±3.3271.57\pm3.32. BLTM achieved 96.05±0.3296.05\pm0.32, 94.97±0.5894.97\pm0.58, 93.99±1.2493.99\pm1.24, 87.67±1.2987.67\pm1.29, and 78.13±4.6278.13\pm4.62. BLTM-V achieved 96.37±0.7796.37\pm0.77, 95.12±0.4095.12\pm0.40, 94.69±0.2494.69\pm0.24, 88.13±3.2388.13\pm3.23, and 78.71±4.3778.71\pm4.37. Both proposed variants outperform PTD at each tested level, with the paper reporting that the advantage grows at higher noise rates.

  10. Knowl 10 — Clothing1M accuracy with real-world noisy labels

    empirical result

    On Clothing1M, models were trained and validated using noisy examples, with clean labels used for the 10,000-image test set. Test classification accuracy was 70.07% for PTD, 70.26% for its revised variant PTD-V, 73.33% for BLTM, and 73.39% for BLTM-V. Thus both proposed variants outperform the reported PTD variants on this real-world noisy dataset; matrix revision provides a further, smaller increase for BLTM.

Coverage note — The full baseline-by-baseline accuracy grids are omitted; the empirical knowls retain the proposed variants and the central PTD comparison across every reported noise level, plus the real-world benchmark.

References

  1. 1.Angluin, D. and Laird, P. Learning from noisy examples. Machine Learning, 2(4):343–370, 1988.
  2. 2.Bartlett, P. L., Jordan, M. I., and McAuliffe, J. D. Convexity, classification, and risk bounds. Journal of the American Statistical Association, 101(473):138–156, 2006.
  3. 3.Berthon, A., Han, B., Niu, G., Liu, T., and Sugiyama, M. Confidence scores make instance-dependent label-noise learning possible. arXiv preprint arXiv:2001.03772, 2020.
  4. 4.Biggio, B., Nelson, B., and Laskov, P. Support vector machines under adversarial label noise. In ACML, 2011.
  5. 5.Blum, A., Kalai, A., and Wasserman, H. Noise-tolerant learning, the parity problem, and the statistical query model. Journal of the ACM, 50(4):506–519, 2003.
  6. 6.Cheng, H., Zhu, Z., Li, X., Gong, Y., Sun, X., and Liu, Y. Learning with instance-dependent label noise: A sample sieve approach. In ICLR, 2021.
  7. 7.Cheng, J., Liu, T., Ramamohanarao, K., and Tao, D. Learning with bounded instance-and label-dependent label noise. In ICML, 2020.
  8. 8.Goldberger, J. and Ben-Reuven, E. Training deep neural-networks using a noise adaptation layer. In ICLR, 2017.
  9. 9.Guo, S., Huang, W., Zhang, H., Zhuang, C., Dong, D., Scott, M. R., and Huang, D. Curriculumnet: Weakly supervised learning from large-scale web images. In ECCV, pp. 135–150, 2018.
  10. 10.Han, B., Yao, J., Niu, G., Zhou, M., Tsang, I., Zhang, Y., and Sugiyama, M. Masking: A new perspective of noisy supervision. In NeurIPS, pp. 5836–5846, 2018a.
  11. 11.Han, B., Yao, Q., Yu, X., Niu, G., Xu, M., Hu, W., Tsang, I., and Sugiyama, M. Co-teaching: Robust training of deep neural networks with extremely noisy labels. In NeurIPS, pp. 8527–8537, 2018b.
  12. 12.Han, B., Niu, G., Yu, X., Yao, Q., Xu, M., Tsang, I. W., and Sugiyama, M. Sigua: Forgetting may make learning with noisy labels more robust. In ICML, 2020.
  13. 13.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, pp. 770–778, 2016.
  14. 14.Hendrycks, D., Mazeika, M., Wilson, D., and Gimpel, K. Using trusted data to train deep networks on labels corrupted by severe noise. In NeurIPS, 2018.
  15. 15.Jiang, L., Zhou, Z., Leung, T., Li, L.-J., and Fei-Fei, L. MentorNet: Learning data-driven curriculum for very deep neural networks on corrupted labels. In ICML, pp. 2309–2318, 2018.
  16. 16.Kremer, J., Sha, F., and Igel, C. Robust active label correction. In AISTATS, pp. 308–316, 2018.
  17. 17.Li, J., Socher, R., and Hoi, S. C. Dividemix: Learning with noisy labels as semi-supervised learning. In ICLR, 2020a.
  18. 18.Li, M., Soltanolkotabi, M., and Oymak, S. Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks. In AISTATS, 2020b.
  19. 19.Li, X., Liu, T., Han, B., Niu, G., and Sugiyama, M. Provably end-to-end label-noise learning without anchor points. arXiv preprint arXiv:2102.02400, 2021.
  20. 20.Li, Y., Yang, J., Song, Y., Cao, L., Luo, J., and Li, L.-J. Learning from noisy labels with distillation. In ICCV, pp. 1910–1918, 2017.
  21. 21.Liu, S., Niles-Weed, J., Razavian, N., and Fernandez-Granda, C. Early-learning regularization prevents memorization of noisy labels. In NeurIPS, 2020.
  22. 22.Liu, T. and Tao, D. Classification with noisy labels by importance reweighting. IEEE Transactions on pattern analysis and machine intelligence, 38(3):447–461, 2016.
  23. 23.Liu, T., Lugosi, G., Neu, G., and Tao, D. Algorithmic stability and hypothesis complexity. In ICML, 2017.
  24. 24.Liu, Y. and Guo, H. Peer loss functions: Learning from noisy labels without knowing noise rates. In ICML, 2020.
  25. 25.Lyu, Y. and Tsang, I. W. Curriculum loss: Robust learning and generalization against label corruption. In ICLR, 2020.
  26. 26.Ma, X., Wang, Y., Houle, M. E., Zhou, S., Erfani, S. M., Xia, S.-T., Wijewickrema, S., and Bailey, J. Dimensionality-driven learning with noisy labels. In ICML, pp. 3361–3370, 2018.
  27. 27.Ma, X., Huang, H., Wang, Y., Romano, S., Erfani, S. M., and Bailey, J. Normalized loss functions for deep learning with noisy labels. In ICML, 2020.
  28. 28.Malach, E. and Shalev-Shwartz, S. Decoupling” when to update” from” how to update”. In NeurIPS, pp. 960–970, 2017.
  29. 29.Manwani, N. and Sastry, P. Noise tolerance under risk minimization. IEEE Transactions on Cybernetics, 2013.
  30. 30.Natarajan, N., Dhillon, I. S., Ravikumar, P. K., and Tewari, A. Learning with noisy labels. In NeurIPS, pp. 1196–1204, 2013.
  31. 31.Nguyen, D. T., Mummadi, C. K., Ngo, T. P. N., Nguyen, T. H. P., Beggel, L., and Brox, T. Self: Learning to filter noisy labels with self-ensembling. In ICLR, 2020.
  32. 32.Northcutt, C. G., Wu, T., and Chuang, I. L. Learning with confident examples: Rank pruning for robust classification with noisy labels. In UAI, 2017.
  33. 33.Patrini, G., Rozza, A., Krishna Menon, A., Nock, R., and Qu, L. Making deep neural networks robust to label noise: A loss correction approach. In CVPR, pp. 1944–1952, 2017.
  34. 34.Ren, M., Zeng, W., Yang, B., and Urtasun, R. Learning to reweight examples for robust deep learning. In ICML, pp. 4331–4340, 2018.
  35. 35.Scott, C. A rate of convergence for mixture proportion estimation, with application to learning from noisy labels. In AISTATS, pp. 838–846, 2015.
  36. 36.Shu, J., Zhao, Q., Xu, Z., and Meng, D. Meta transition adaptation for robust deep learning with noisy labels. arXiv preprint arXiv:2006.05697, 2020.
  37. 37.Tanaka, D., Ikami, D., Yamasaki, T., and Aizawa, K. Joint optimization framework for learning with noisy labels. In CVPR, 2018.
  38. 38.Thekumparampil, K. K., Khetan, A., Lin, Z., and Oh, S. Robustness of conditional gans to noisy labels. In NeurIPS, pp. 10271–10282, 2018.
  39. 39.Vahdat, A. Toward robustness against label noise in training deep discriminative neural networks. In NeurIPS, pp. 5596–5605, 2017.
  40. 40.Veit, A., Alldrin, N., Chechik, G., Krasin, I., Gupta, A., and Belongie, S. Learning from noisy large-scale datasets with minimal supervision. In CVPR, pp. 839–847, 2017.
  41. 41.Wang, K., Peng, X., Yang, S., Yang, J., Zhu, Z., Wang, X., and You, Y. Reliable label correction is a good booster when learning with extremely noisy labels. arXiv preprint arXiv:2205.00186, 2022.
  42. 42.Wang, X., Wang, S., Wang, J., Shi, H., and Mei, T. Co-mining: Deep face recognition with noisy labels. In ICCV, pp. 9358–9367, 2019.
  43. 43.Wang, X., Hua, Y., Kodirov, E., Clifton, D. A., and Robertson, N. M. Proselflc: Progressive self label correction for training robust deep neural networks. In CVPR, 2021.
  44. 44.Wu, S., Xia, X., Liu, T., Han, B., Gong, M., Wang, N., Liu, H., and Niu, G. Class2simi: A new perspective on learning with label noise. arXiv preprint arXiv:2006.07831, 2020.
  45. 45.Xia, X., Liu, T., Wang, N., Han, B., Gong, C., Niu, G., and Sugiyama, M. Are anchor points really indispensable in label-noise learning? In NeurIPS, pp. 6838–6849, 2019.
  46. 46.Xia, X., Liu, T., Han, B., Wang, N., Gong, M., Liu, H., Niu, G., Tao, D., and Sugiyama, M. Part-dependent label noise: Towards instance-dependent label noise. In NeurIPS, 2020a.
  47. 47.Xia, X., Liu, T., Han, B., Wang, N., Gong, M., Liu, H., Niu, G., Tao, D., and Sugiyama, M. Parts-dependent label noise: Towards instance-dependent label noise. In NeurIPS, 2020b.
  48. 48.Xia, X., Liu, T., Han, B., Gong, C., Wang, N., Ge, Z., and Chang, Y. Robust early-learning: Hindering the memorization of noisy labels. In ICLR, 2021.
  49. 49.Xiao, T., Xia, T., Yang, Y., Huang, C., and Wang, X. Learning from massive noisy labeled data for image classification. In CVPR, pp. 2691–2699, 2015.
  50. 50.Xu, Y., Cao, P., Kong, Y., and Wang, Y. L dmi: A novel information-theoretic loss function for training deep nets robust to label noise. In NeurIPS, pp. 6222–6233, 2019.
  51. 51.Yan, Y., Rosales, R., Fung, G., Subramanian, R., and Dy, J. Learning from multiple annotators with varying expertise. Machine learning, 95(3):291–327, 2014.
  52. 52.Yang, S., Liu, L., and Xu, M. Free lunch for few-shot learning: Distribution calibration. In ICLR, 2021a.
  53. 53.Yang, S., Wu, S., Liu, T., and Xu, M. Bridging the gap between few-shot and many-shot learning via distribution calibration, 2021b.
  54. 54.Yao, Q., Yang, H., Han, B., Niu, G., and Kwok, J. T. Searching to exploit memorization effect in learning with noisy labels. In ICML, 2020a.
  55. 55.Yao, Y., Liu, T., Han, B., Gong, M., Deng, J., Niu, G., and Sugiyama, M. Dual t: Reducing estimation error for transition matrix in label-noise learning. In NeurIPS, 2020b.
  56. 56.Yu, X., Liu, T., Gong, M., Zhang, K., Batmanghelich, K., and Tao, D. Transfer learning with label noise. arXiv preprint arXiv:1707.09724, 2017.
  57. 57.Yu, X., Liu, T., Gong, M., and Tao, D. Learning with biased complementary labels. In ECCV, pp. 68–83, 2018.
  58. 58.Yu, X., Han, B., Yao, J., Niu, G., Tsang, I. W., and Sugiyama, M. How does disagreement benefit co-teaching? In ICML, 2019.
  59. 59.Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. Understanding deep learning requires rethinking generalization. In ICLR, 2017a.
  60. 60.Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017b.
  61. 61.Zhang, Y., Zheng, S., Wu, P., Goswami, M., and Chen, C. Learning with feature-dependent label noise: A progressive approach. In ICLR, 2021.
  62. 62.Zhang, Z. and Sabuncu, M. Generalized cross entropy loss for training deep neural networks with noisy labels. In NeurIPS, pp. 8778–8788, 2018.
  63. 63.Zheng, G., Awadallah, A. H., and Dumais, S. T. Meta label correction for noisy label learning. In AAAI, 2021.
  64. 64.Zheng, S., Wu, P., Goswami, A., Goswami, M., Metaxas, D., and Chen, C. Error-bounded correction of noisy labels. In ICML, pp. 11447–11457, 2020.
  65. 65.Zhu, Z., Liu, T., and Liu, Y. A second-order approach to learning with instance-dependent label noise. arXiv preprint arXiv:2012.11854, 2020.
  66. 66.Zhu, Z., Song, Y., and Liu, Y. Clusterability as an alternative to anchor points when learning with noisy labels. arXiv preprint arXiv:2102.05291, 2021.

Citation

MLA
Yang, S., et al. “Estimating Instance-dependent Bayes-label Transition Matrix Using a Deep Neural Network”. International Conference on Machine Learning, vol. 162, 2022, pp. 25302–12, https://proceedings.mlr.press/v162/yang22p.html.
APA
Yang, S., Yang, E., Han, B., Liu, Y., Xu, M., Niu, G., & Liu, T. (2022). Estimating Instance-dependent Bayes-label Transition Matrix using a Deep Neural Network. International Conference on Machine Learning, 162, 25302–25312. https://proceedings.mlr.press/v162/yang22p.html
Chicago
Yang, S., E. Yang, B. Han, et al. 2022. “Estimating Instance-dependent Bayes-label Transition Matrix Using a Deep Neural Network”. International Conference on Machine Learning 162: 25302–12. https://proceedings.mlr.press/v162/yang22p.html.
Harvard
Yang, S. et al. (2022) “Estimating Instance-dependent Bayes-label Transition Matrix using a Deep Neural Network”, International Conference on Machine Learning. PMLR, pp. 25302–25312. Available at: https://proceedings.mlr.press/v162/yang22p.html.
Vancouver
1. Yang S, Yang E, Han B, Liu Y, Xu M, Niu G, Liu T (2022) Estimating Instance-dependent Bayes-label Transition Matrix using a Deep Neural Network. In: International Conference on Machine Learning. PMLR, pp 25302–25312

BibTeX

@InProceedings{pmlr-v162-yang22p,
  title = 	 {Estimating Instance-dependent {B}ayes-label Transition Matrix using a Deep Neural Network},
  author =       {Yang, Shuo and Yang, Erkun and Han, Bo and Liu, Yang and Xu, Min and Niu, Gang and Liu, Tongliang},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {25302--25312},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/yang22p/yang22p.pdf},
  url = 	 {https://proceedings.mlr.press/v162/yang22p.html},
  abstract = 	 {In label-noise learning, estimating the transition matrix is a hot topic as the matrix plays an important role in building statistically consistent classifiers. Traditionally, the transition from clean labels to noisy labels (i.e., clean-label transition matrix (CLTM)) has been widely exploited to learn a clean label classifier by employing the noisy data. Motivated by that classifiers mostly output Bayes optimal labels for prediction, in this paper, we study to directly model the transition from Bayes optimal labels to noisy labels (i.e., Bayes-label transition matrix (BLTM)) and learn a classifier to predict Bayes optimal labels. Note that given only noisy data, it is ill-posed to estimate either the CLTM or the BLTM. But favorably, Bayes optimal labels have less uncertainty compared with the clean labels, i.e., the class posteriors of Bayes optimal labels are one-hot vectors while those of clean labels are not. This enables two advantages to estimate the BLTM, i.e., (a) a set of examples with theoretically guaranteed Bayes optimal labels can be collected out of noisy data; (b) the feasible solution space is much smaller. By exploiting the advantages, we estimate the BLTM parametrically by employing a deep neural network, leading to better generalization and superior classification performance.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/