Identifiability of Label Noise Transition Matrix

Yang LiuHao ChengKun Zhang

article2023ICML58 citations

Establishes the theoretical conditions required to identify instance-dependent label noise transition matrices using Kruskal's identifiability theorem, proving why multiple noisy labels are necessary and demonstrating how disentangled representations improve matrix estimation without clean labels.

Listen

Supervised machine learning models rely heavily on large amounts of labeled data, but real-world training datasets frequently contain label errors, known as label noise. Correcting for this noise requires estimating the noise transition matrix—the mathematical relationship describing how true labels are corrupted into observed noisy labels. While classic methods assume that error rates are uniform across a category, real-world errors typically depend on the specific features of each individual data instance. However, without access to clean ground truth data, it is mathematically uncertain whether instance-dependent noise transition matrices can be uniquely identified and correctly estimated. Applying incorrect matrices degrades model accuracy and can create a false sense of algorithmic fairness.

The article establishes a rigorous theoretical foundation for when instance-dependent noise transition matrices can be uniquely identified from data containing label noise. It aims to determine the necessary conditions for identifying these noise patterns, explain why specific empirical techniques succeed, and provide actionable methods to enhance identifiability.

The authors analyze the problem mathematically by establishing a fundamental connection between learning with noisy labels and classical statistical theorems on latent variables and multi-dimensional arrays, specifically Kruskal’s identifiability theorem. They also evaluate their theoretical insights experimentally by training deep neural networks across benchmark image classification datasets, measuring estimation error reductions and classification accuracy under varying noise rates.

The analysis yields several key findings regarding the estimation of label noise. First, the article proves that observing only a single noisy label per instance is mathematically insufficient for identification; exactly three independent and informative noisy labels per instance are both necessary and sufficient to guarantee unique identification. Second, the authors demonstrate how existing single-label methods succeed in practice by effectively bypassing this three-label requirement: algorithms leverage local data smoothness (borrowing labels from nearest neighbors) or group data with shared noise structures. Third, the study reveals that using disentangled feature representations—representations where learned attributes are statistically independent given the true label—substitutes for multiple labels and enables matrix identification even from a single noisy observation. Fourth, empirical experiments confirm that fully disentangled feature encoders reduce matrix estimation error by more than 50% to 75% compared to baseline weakly supervised encoders and achieve substantially higher model test accuracy (e.g., reaching 73.2% test accuracy under 30% noise compared to 66.6% for standard self-supervised methods).

These findings provide clear operational clarity for designing machine learning pipelines on noisy data. They show that simply training models on larger volumes of single-label noisy data without structural constraints cannot resolve underlying noise ambiguities. Instead, practitioners face two viable pathways: collecting multiple independent label annotations per sample or utilizing self-supervised methods to learn disentangled feature representations prior to downstream classification.

Organizations developing machine learning systems on noisy or crowdsourced data should adopt self-supervised and invariant representation pre-training to generate disentangled features before attempting label correction. When annotation budgets allow, soliciting at least three independent annotations per data point is strongly recommended to mathematically ensure noise identifiability.

The theoretical guarantees assume discretized feature representations and conditionally independent observations, while practical continuous deep learning pipelines may deviate slightly from these idealized settings. Nevertheless, the consistent convergence between the mathematical proofs and empirical evaluations provides high confidence that adopting disentangled representations significantly improves robustness against label noise.

No sufficiently relevant recommendations were found.

Cover for Identifiability of Label Noise Transition Matrix

Abstract

The noise transition matrix plays a central role in the problem of learning with noisy labels. Among many other reasons, a large number of existing solutions rely on the knowledge of it. Identifying and estimating the transition matrix without ground truth labels is a critical and challenging task. When label noise transition depends on each instance, the problem of identifying the instance-dependent noise transition matrix becomes substantially more challenging. Despite recently proposed solutions for learning from instance-dependent noisy labels, the literature lacks a unified understanding of when such a problem remains identifiable. The goal of this paper is to characterize the identifiability of the label noise transition matrix. Building on Kruskal’s identifiability results, we are able to show the necessity of multiple noisy labels in identifying the noise transition matrix at the instance level. We further instantiate the results to explain the successes of the state-of-the-art solutions and how additional assumptions alleviated the requirement of multiple noisy labels. Our result reveals that disentangled features improve identification. This discovery led us to an approach that improves the estimation of the transition matrix using properly disentangled features. Code is available at https://github.com/UCSC-REAL/Identifiability.

Table of Contents

  • 1. Introduction
  • 1.1. Related works
  • 2. Formulation
  • 3. Preliminary
  • 3.1. Preliminaries using irreducibility and anchor points
  • 3.2. Kruskal's identifiability result
  • 4. Instance-Level Identifiability
  • 4.1. Single noisy label might not be sufficient
  • 4.2. The necessity of multiple noisy labels
  • 5. Instantiations and Practical Implications of Our Identifiability Results
  • 5.1. Leveraging smoothness and clusterability of X
  • 5.2. Leveraging smoothness and clusterability of T ( X )
  • 5.3. Smoothness and clusterability of T ( X ) with unknown groupings
  • 6. Empirical Evidence: Disentangled Features
  • 7. Concluding Remarks
  • References
  • Appendix: Identifiability of Label Noise Transition Matrix
  • Notation Table
  • A. Omitted Proofs
  • Proof for Lemma 3.3
  • Proof for Theorem 3.4
  • Proof for Theorem 4.2
  • Proof for Theorem 5.2
  • Proof for Theorem 5.5
  • Proof for Theorem 5.6
  • Proof for Theorem 5.7
  • B. Generic identifiability
  • C. More experiments
  • C.1. More training details for Table 1
  • C.2. Training performance using estimated transition matrix
  • C.3. Initializing DNN using disentangled features
  • C.4. Verifying the importance of characterizing the identifiability of the label noise transition matrix
  • C.4.1. CIFAR10 experiment
  • C.4.2. Gaussian experiment

Knowls

  1. Knowl 1 — Three independent noisy labels identify an instance-level transition matrix

    theoretical result

    Fix a feature value xx in a KK-class classification problem. Let KxK_x be the number of clean classes with positive conditional probability at xx. Suppose pp noisy labels are drawn independently and identically conditional on the clean label YY and X=xX=x, all using the same transition matrix T(x)T(x), whose entry Tij(x)T_{ij}(x) is P(Y~=j∣Y=i,X=x)P(\widetilde Y=j\mid Y=i,X=x). A noisy label is informative at xx when the rows of T(x)T(x) corresponding to the KxK_x active clean classes are linearly independent, equivalently when their rank is KxK_x.

    When Kx≥2K_x\ge 2, three informative noisy labels are sufficient, and fewer than three do not suffice to guarantee, identification of T(x)T(x) from their joint conditional distribution. Identification is up to permutation of the clean-class labels. The result gives a condition for general instance-level identification; it does not rule out identification from fewer labels in special parameter settings.

  2. Knowl 2 — Disentangled features identify transition matrices within known groups

    theoretical result

    Suppose the feature values are partitioned into known groups g∈Gg\in G, with a common noise transition matrix T(x)T(x) for all feature values in each group. For an input xx in a group, observe one informative noisy label and a representation R(x)=(R1(x),…,Rd∗(x))R(x)=(R_1(x),\ldots,R_{d^*}(x)) of finite-valued features that are conditionally independent given the clean label YY. For each feature RiR_i, define its class-conditional probability matrix by Mi[j,k]=P(Ri=k∣Y=j)M_i[j,k]=P(R_i=k\mid Y=j); call RiR_i informative when its Kruskal rank (the largest number of rows for which every subset is linearly independent) is at least 22.

    A sufficient condition for identifying the group's transition matrix is d∗≥Kd^*\ge K, where KK is the total number of classes. This result needs neither multiple noisy labels at each input nor an anchor-point condition, but it does require the stated informative, conditionally independent features. Identification is up to clean-label permutation.

  3. Knowl 3 — Identifiability with unobserved group membership

    theoretical result

    Suppose each input belongs to one of ∣G∣|G| groups, and the transition matrix is shared within each group, but group membership is unobserved. Treat the pair of group and clean label, (G,Y)(G,Y), as the latent class; the resulting latent space has size ∣G∣K|G|K. Given one informative noisy label, a sufficient condition for identifying the group-specific transition matrix for an input in group gg is at least d∗≥2∣G∣K−1d^*\ge 2|G|K-1 informative, finite-valued, disentangled features. Here disentanglement requires the feature coordinates to be conditionally independent under the latent group-and-label classes, and each feature's class-conditional probability matrix must have Kruskal rank at least 22.

    The condition identifies the transition matrix while accommodating unknown group membership, up to latent-label permutation. Its required feature count grows linearly with the number of groups.

  4. Knowl 4 — Generic identification requires fewer known-group features

    theoretical result

    For a known group gg, let Kg∗=max⁡x∈gKxK_g^* = \max_{x\in g}K_x be the largest number of active clean classes among its inputs. Under the known-group setup with one informative noisy label and d∗d^* conditionally independent, informative, finite-valued features, the transition matrix is generically identifiable if each feature has at least two possible values and

    d∗≥⌈log⁡2(2Kg∗+12)⌉.d^*\ge \left\lceil\log_2\left(\frac{2K_g^*+1}{2}\right)\right\rceil.

    Generic identifiability means that the condition holds except on a measure-zero set of parameter values, and identification is up to label permutation. This is a weaker guarantee than identification at every parameter setting, and can require substantially fewer features than the sufficient d∗≥Kd^*\ge K condition.

  5. Knowl 5 — Pooling distinct group transitions incurs a minimum estimation error

    theoretical result

    Suppose two groups have transition matrices T1T_1 and T2T_2 of the same dimensions, but an estimator uses a single matrix T∗T^* for both groups. For any such estimate, the sum of its Frobenius-norm errors across the two groups satisfies

    ∥T1−T∗∥F+∥T2−T∗∥F≥12∥T1−T2∥F.\|T_1-T^*\|_F+\|T_2-T^*\|_F\ge \frac{1}{\sqrt{2}}\|T_1-T_2\|_F.

    Thus, when group membership is not modeled and the groups have different transition matrices, no shared estimate can be accurate for both beyond this separation-dependent lower bound.

  6. Knowl 6 — Two-nearest-neighbor clusterability can supply three label observations

    model/method

    The paper's 2-nearest-neighbor clusterability condition requires that an input XX and its two nearest neighbors X1,X2X_1,X_2 have the same clean label and transition matrix: Y=Y1=Y2Y=Y_1=Y_2 and T(X)=T(X1)=T(X2)T(X)=T(X_1)=T(X_2). Under this condition, their three noisy labels can serve as the three independent observations needed for instance-level identification, provided the labels are generated independently conditional on their shared clean label and transition matrix.

    The paper also gives a finite discrete-domain construction supporting this condition. Let each feature value xx have a positive weight qxq_x, sampled independently and uniformly from a finite set of positive weights, and let the feature distribution be D(x)=qx/∑uquD(x)=q_x/\sum_{u}q_u. In the construction, sufficiently close triplets share a clean label and transition matrix, and the triplet's noisy labels are generated from that shared clean label. The stated result is that if the sample size NN satisfies N>4∑xqx/min⁡xqxN>4\sum_x q_x/\min_x q_x, then each observed input and its two nearest neighbors satisfy the 2-nearest-neighbor clusterability condition with probability at least 1−Nexp⁡(−2N)1-N\exp(-2N). This guarantee is specific to the described construction.

  7. Knowl 7 — More disentangled features reduce transition-matrix estimation error

    empirical result

    The paper compared HOC transition-matrix estimates using three CIFAR-10 feature types: weakly supervised features, SimCLR features, and IPIRM features. All encoders used a ResNet-50 backbone and were trained on CIFAR-100; their features were then used to estimate the transition matrix on CIFAR-10. Weak supervision was produced by cross-entropy training with symmetric noise rate 0.10.1; SimCLR provided self-supervised features; and IPIRM was used as the more fully disentangled representation. Each experiment was run three times. The metric is 100∑i=1K∑j=1K∣T^ij−Tij∣/K2100\sum_{i=1}^{K}\sum_{j=1}^{K}|\widehat T_{ij}-T_{ij}|/K^2, reported as mean ±\pm standard deviation. IPIRM has the lowest reported error in every noise condition, and both self-supervised feature types outperform the weakly supervised baseline.

    Feature type asymm. 0.3 asymm. 0.4 inst. 0.4 inst. 0.5 inst. 0.6
    Weakly-Supervised 14.51±0.414.51\pm0.4 15.2±0.0215.2\pm0.02 8.39±0.058.39\pm0.05 6.91±0.066.91\pm0.06 6.18±0.156.18\pm0.15
    SimCLR 4.42±0.014.42\pm0.01 4.41±0.014.41\pm0.01 2.91±0.022.91\pm0.02 2.55±0.042.55\pm0.04 2.64±0.032.64\pm0.03
    IPIRM 3.73±0.023.73\pm0.02 3.74±0.013.74\pm0.01 2.47±0.032.47\pm0.03 2.20±0.022.20\pm0.02 2.37±0.062.37\pm0.06
  8. Knowl 8 — Disentangled-feature estimates and initializations improve noisy-label accuracy

    empirical result

    Two experiments tested whether more disentangled features help downstream classification. In the first, forward loss correction (FW) used HOC-estimated transition matrices on CIFAR-10 with instance-dependent noise. In the second, CIFAR-100 models were trained with ordinary cross-entropy, initialized from either random weights or feature encoders trained using SimCLR or IPIRM. The reported test accuracies show that IPIRM outperforms SimCLR in both settings, and both feature-based approaches outperform their respective baselines.

    FW method on CIFAR-10 inst. 0.3 inst. 0.4 inst. 0.5 inst. 0.6
    SimCLR features 66.61 65.82 64.51 62.81
    IPIRM features 73.24 72.54 71.33 69.42
    CIFAR-100 cross-entropy initialization inst. 0.3 inst. 0.4 inst. 0.5 inst. 0.6
    Random initialization 43.47 35.17 27.07 18.25
    SimCLR initialization 58.95 49.7 36.87 25.07
    IPIRM initialization 64.92 56.18 43.75 30.36

    The CIFAR-100 initialization experiment used cross-entropy training and the same training hyperparameters across initialization methods. These results support the paper's empirical claim that disentangled features are useful both for estimating a transition matrix and for training directly on noisy data.

  9. Knowl 9 — A badly specified transition matrix can erase or reverse correction gains

    empirical result

    The paper evaluated forward loss correction (FW) under two settings to illustrate the consequences of transition-matrix estimation error. In a CIFAR-10 experiment, 50% of examples were corrupted using a uniform-off-diagonal transition matrix whose diagonal entries were evenly spaced from 0.90.9 to 0.20.2; the other examples were uncorrupted. On ResNet-34, FW with the true matrix T1T_1 improved over cross-entropy (CE), while FW with a reversed-diagonal matrix T3T_3 performed worse than CE. A matrix T2T_2 with every diagonal entry set to 0.40.4 gave an intermediate result.

    CIFAR-10 method Test accuracy
    CE 79.34
    FW with T1T_1 (true matrix) 82.62
    FW with T2T_2 (all diagonals 0.40.4) 81.65
    FW with T3T_3 (diagonal from 0.20.2 to 0.90.9) 78.13

    A separate binary Gaussian experiment used X∼N(0,3)X\sim\mathcal N(0,3), P(Y=1∣X)=sigmoid⁡(X)P(Y=1\mid X)=\operatorname{sigmoid}(X), and the true transition matrix [0.90.10.20.8]\begin{bmatrix}0.9&0.1\\0.2&0.8\end{bmatrix}. With 5,000 sampled examples, the reported test accuracies were 83.2283.22 for CE and 83.3183.31 for FW using the estimated matrix. The average estimated matrix was [0.9830.0170.0080.992]\begin{bmatrix}0.983&0.017\\0.008&0.992\end{bmatrix}, close to the identity rather than the true matrix, consistent with the negligible improvement from correction.

  10. Knowl 10 — The identifiability guarantees do not by themselves recover the clean posterior

    limitation

    The paper's feature-based results concern identifying the label-noise transition matrix under assumptions such as finite-valued, informative, conditionally independent features. The existence of such a disentangled representation does not, by itself, imply that the clean-label posterior P(Y∣X)P(Y\mid X) can be directly inferred; identifying that structure is a separate problem on a larger space. The stated guarantees therefore support transition-matrix identification under their assumptions, rather than establishing full recovery of the clean prediction problem in general.

Coverage note — The binary class-dependent mixture-proportion discussion and Kruskal's general criterion are omitted as reproduced preliminary material; proof details and the HOC/noise-generation procedures are omitted because they are not standalone contributions.

References

  1. 1.Allman, E. S., Matias, C., and Rhodes, J. A. Identifiability of parameters in latent structure models with many observed variables. The Annals of Statistics, 37(6A):3099–3132, 2009.
  2. 2.Anandkumar, A., Ge, R., Hsu, D., Kakade, S. M., and Telgarsky, M. Tensor decompositions for learning latent variable models. Journal of machine learning research, 15:2773–2832, 2014.
  3. 3.Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. Invariant risk minimization. arXiv preprint arXiv:1907.02893, 2019.
  4. 4.Bahri, D., Jiang, H., and Gupta, M. Deep k-nn for noisy labels. In International Conference on Machine Learning, pp. 540–550. PMLR, 2020.
  5. 5.Berthon, A., Han, B., Niu, G., Liu, T., and Sugiyama, M. Confidence scores make instance-dependent label-noise learning possible. arXiv preprint arXiv:2001.03772, 2020.
  6. 6.Blanchard, G., Lee, G., and Scott, C. Semi-supervised novelty detection. The Journal of Machine Learning Research, 11:2973–3009, 2010.
  7. 7.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pp. 1597–1607. PMLR, 2020.
  8. 8.Cheng, H., Zhu, Z., Li, X., Gong, Y., Sun, X., and Liu, Y. Learning with instance-dependent label noise: A sample sieve approach. In International Conference on Learning Representations, 2021a.
  9. 9.Cheng, H., Zhu, Z., Sun, X., and Liu, Y. Demystifying how self-supervised features improve training from noisy labels. arXiv preprint arXiv:2110.09022, 2021b.
  10. 10.Cheng, J., Liu, T., Ramamohanarao, K., and Tao, D. Learning with bounded instance-and label-dependent label noise. In Proceedings of the 37th International Conference on Machine Learning, ICML ’20, 2020.
  11. 11.Clogg, C. C. Latent class models. In Handbook of statistical modeling for the social and behavioral sciences, pp. 311–359. Springer, 1995.
  12. 12.Feldman, V. Does learning require memorization? a short tale about a long tail. In Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing, pp. 954–959, 2020.
  13. 13.Ghosh, A. and Lan, A. Contrastive learning improves model robustness under label noise. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2703–2708, 2021.
  14. 14.Han, B., Yao, Q., Yu, X., Niu, G., Xu, M., Hu, W., Tsang, I., and Sugiyama, M. Co-teaching: Robust training of deep neural networks with extremely noisy labels. In Advances in neural information processing systems, pp. 8527–8537, 2018.
  15. 15.Han, B., Niu, G., Yu, X., Yao, Q., Xu, M., Tsang, I., and Sugiyama, M. Sigua: Forgetting may make learning with noisy labels more robust. In International Conference on Machine Learning, pp. 4006–4016. PMLR, 2020.
  16. 16.Higgins, I., Amos, D., Pfau, D., Racaniere, S., Matthey, L., Rezende, D., and Lerchner, A. Towards a definition of disentangled representations. arXiv preprint arXiv:1812.02230, 2018.
  17. 17.Jiang, L., Zhou, Z., Leung, T., Li, L.-J., and Fei-Fei, L. Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels. In International Conference on Machine Learning, pp. 2304–2313. PMLR, 2018.
  18. 18.Karger, D. R., Oh, S., and Shah, D. Iterative learning for reliable crowdsourcing systems. In Advances in neural information processing systems, pp. 1953–1961, 2011.
  19. 19.Kruskal, J. B. More factors than subjects, tests and treatments: an indeterminacy theorem for canonical decomposition and individual differences scaling. Psychometrika, 41(3):281–293, 1976.
  20. 20.Kruskal, J. B. Three-way arrays: rank and uniqueness of trilinear decompositions, with application to arithmetic complexity and statistics. Linear algebra and its applications, 18(2):95–138, 1977.
  21. 21.Li, J., Socher, R., and Hoi, S. C. Dividemix: Learning with noisy labels as semi-supervised learning. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=HJgExaVtwr.
  22. 22.Li, X., Liu, T., Han, B., Niu, G., and Sugiyama, M. Provably end-to-end label-noise learning without anchor points. arXiv preprint arXiv:2102.02400, 2021.
  23. 23.Liu, Q., Peng, J., and Ihler, A. T. Variational inference for crowdsourcing. Advances in neural information processing systems, 25:692–700, 2012.
  24. 24.Liu, T. and Tao, D. Classification with noisy labels by importance reweighting. IEEE Transactions on pattern analysis and machine intelligence, 38(3):447–461, 2015.
  25. 25.Liu, Y. Understanding instance-level label noise: Disparate impacts and treatments, 2021.
  26. 26.Liu, Y. and Liu, M. An online learning approach to improving the quality of crowd-sourcing. In Proceedings of the 2015 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems, SIGMETRICS ’15, pp. 217–230, New York, NY, USA, 2015. ACM. ISBN 978-1-4503-3486-0. doi: 10.1145/2745844.2745874. URL http://doi.acm.org/10.1145/2745844.2745874.
  27. 27.Liu, Y. and Wang, J. Can less be more? when increasing-to-balancing label noise rates considered beneficial. NeurIPS’21.
  28. 28.Liu, Y., Wang, J., and Chen, Y. Surrogate scoring rules and a dominant truth serum. ACM EC, 2020.
  29. 29.Menon, A., Van Rooyen, B., Ong, C. S., and Williamson, B. Learning from corrupted binary labels via class-probability estimation. In International Conference on Machine Learning, pp. 125–134, 2015.
  30. 30.Natarajan, N., Dhillon, I. S., Ravikumar, P. K., and Tewari, A. Learning with noisy labels. In Advances in neural information processing systems, pp. 1196–1204, 2013.
  31. 31.Nguyen, D. T., Mummadi, C. K., Ngo, T. P. N., Nguyen, T. H. P., Beggel, L., and Brox, T. Self: Learning to filter noisy labels with self-ensembling. arXiv preprint arXiv:1910.01842, 2019.
  32. 32.Northcutt, C., Jiang, L., and Chuang, I. Confident learning: Estimating uncertainty in dataset labels. Journal of Artificial Intelligence Research, 70:1373–1411, 2021.
  33. 33.Northcutt, C. G., Wu, T., and Chuang, I. L. Learning with confident examples: Rank pruning for robust classification with noisy labels. UAI, 2017.
  34. 34.Patrini, G., Rozza, A., Krishna Menon, A., Nock, R., and Qu, L. Making deep neural networks robust to label noise: A loss correction approach. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
  35. 35.Scott, C. A rate of convergence for mixture proportion estimation, with application to learning from noisy labels. In AISTATS, 2015.
  36. 36.Sidiropoulos, N. D. and Bro, R. On the uniqueness of multilinear decomposition of n-way arrays. Journal of chemometrics, 14(3):229–239, 2000.
  37. 37.Sidiropoulos, N. D., De Lathauwer, L., Fu, X., Huang, K., Papalexakis, E. E., and Faloutsos, C. Tensor decomposition for signal processing and machine learning. IEEE Transactions on Signal Processing, 65(13):3551–3582, 2017.
  38. 38.Steenbrugge, X., Leroux, S., Verbelen, T., and Dhoedt, B. Improving generalization for abstract reasoning tasks using disentangled feature representations. arXiv preprint arXiv:1811.04784, 2018.
  39. 39.Traganitis, P. A., Pages-Zamora, A., and Giannakis, G. B. Blind multiclass ensemble classification. IEEE Transactions on Signal Processing, 66(18):4737–4752, 2018.
  40. 40.Van den Oord, A., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. arXiv e-prints, pp. arXiv–1807, 2018.
  41. 41.Wang, J., Liu, Y., and Levy, C. Fair classification with group-dependent label noise. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, pp. 526–536, New York, NY, USA, 2021a. Association for Computing Machinery. ISBN 9781450383097. doi: 10.1145/3442188.3445915. URL https://doi.org/10.1145/3442188.3445915.
  42. 42.Wang, Q., Yao, J., Gong, C., Liu, T., Gong, M., Yang, H., and Han, B. Learning with group noise. arXiv preprint arXiv:2103.09468, 2021b.
  43. 43.Wang, T., Yue, Z., Huang, J., Sun, Q., and Zhang, H. Self-supervised learning disentangled group representation as feature. Advances in Neural Information Processing Systems, 34, 2021c.
  44. 44.Wei, H., Feng, L., Chen, X., and An, B. Combating noisy labels by agreement: A joint training method with co-regularization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13726–13735, 2020.
  45. 45.Wei, J., Zhu, Z., Cheng, H., Liu, T., Niu, G., and Liu, Y. Learning with noisy labels revisited: A study using real-world human annotations. arXiv preprint arXiv:2110.12088, 2021.
  46. 46.Wei, J., Zhu, Z., Niu, G., Liu, T., Liu, S., Sugiyama, M., and Liu, Y. Fairness improves learning from noisily labeled long-tailed data. arXiv preprint arXiv:2303.12291, 2023.
  47. 47.Xia, X., Liu, T., Wang, N., Han, B., Gong, C., Niu, G., and Sugiyama, M. Are anchor points really indispensable in label-noise learning? Advances in Neural Information Processing Systems, 32, 2019.
  48. 48.Xia, X., Liu, T., Han, B., Gong, C., Wang, N., Ge, Z., and Chang, Y. Robust early-learning: Hindering the memorization of noisy labels. In International Conference on Learning Representations, 2020a.
  49. 49.Xia, X., Liu, T., Han, B., Wang, N., Gong, M., Liu, H., Niu, G., Tao, D., and Sugiyama, M. Part-dependent label noise: Towards instance-dependent label noise. Advances in Neural Information Processing Systems, 33:7597–7610, 2020b.
  50. 50.Xiao, T., Xia, T., Yang, Y., Huang, C., and Wang, X. Learning from massive noisy labeled data for image classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2691–2699, 2015.
  51. 51.Yang, S., Yang, E., Han, B., Liu, Y., Xu, M., Niu, G., and Liu, T. Estimating instance-dependent label-noise transition matrix using dnns. arXiv preprint arXiv:2105.13001, 2021.
  52. 52.Yao, Y., Liu, T., Han, B., Gong, M., Deng, J., Niu, G., and Sugiyama, M. Dual t: Reducing estimation error for transition matrix in label-noise learning. arXiv preprint arXiv:2006.07805, 2020a.
  53. 53.Yao, Y., Liu, T., Han, B., Gong, M., Niu, G., Sugiyama, M., and Tao, D. Towards mixture proportion estimation without irreducibility. arXiv preprint arXiv:2002.03673, 2020b.
  54. 54.Yao, Y., Liu, T., Gong, M., Han, B., Niu, G., and Zhang, K. Instance-dependent label-noise learning under a structural causal model. Advances in Neural Information Processing Systems, 34, 2021.
  55. 55.Zhang, Y., Chen, X., Zhou, D., and Jordan, M. I. Spectral methods meet em: A provably optimal algorithm for crowdsourcing. Advances in neural information processing systems, 27, 2014.
  56. 56.Zhang, Y., Niu, G., and Sugiyama, M. Learning noise transition matrix from only noisy labels via total variation regularization. arXiv preprint arXiv:2102.02414, 2021.
  57. 57.Zheltonozhskii, E., Baskin, C., Mendelson, A., Bronstein, A. M., and Litany, O. Contrast to divide: Self-supervised pre-training for learning with noisy labels. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 1657–1667, 2022.
  58. 58.Zhu, X. Semi-supervised learning with graphs. Carnegie Mellon University, 2005.
  59. 59.Zhu, X., Ghahramani, Z., and Lafferty, J. D. Semi-supervised learning using gaussian fields and harmonic functions. In Proceedings of the 20th International conference on Machine learning (ICML-03), pp. 912–919, 2003.
  60. 60.Zhu, Z., Dong, Z., Cheng, H., and Liu, Y. A good representation detects noisy labels. arXiv preprint arXiv:2110.06283, 2021a.
  61. 61.Zhu, Z., Liu, T., and Liu, Y. A second-order approach to learning with instance-dependent label noise. CVPR, 2021b.
  62. 62.Zhu, Z., Song, Y., and Liu, Y. Clusterability as an alternative to anchor points when learning with noisy labels. ICML, 2021c.
  63. 63.Zhu, Z., Wang, J., and Liu, Y. Beyond images: Label noise transition matrix estimation for tasks with lower-quality features. arXiv preprint arXiv:2202.01273, 2022a.
  64. 64.Zhu, Z., Yao, Y., Sun, J., Liu, Y., and Li, H. Evaluating fairness without sensitive attributes: A framework using only auxiliary models. arXiv preprint arXiv:2210.03175, 2022b.

Citation

MLA
Liu, Y., et al. “Identifiability of Label Noise Transition Matrix”. International Conference on Machine Learning, vol. 202, 2023, pp. 21475–96, https://proceedings.mlr.press/v202/liu23g.html.
APA
Liu, Y., Cheng, H., & Zhang, K. (2023). Identifiability of Label Noise Transition Matrix. International Conference on Machine Learning, 202, 21475–21496. https://proceedings.mlr.press/v202/liu23g.html
Chicago
Liu, Y., H. Cheng, and K. Zhang. 2023. “Identifiability of Label Noise Transition Matrix”. International Conference on Machine Learning 202: 21475–96. https://proceedings.mlr.press/v202/liu23g.html.
Harvard
Liu, Y., Cheng, H. and Zhang, K. (2023) “Identifiability of Label Noise Transition Matrix”, International Conference on Machine Learning. PMLR, pp. 21475–21496. Available at: https://proceedings.mlr.press/v202/liu23g.html.
Vancouver
1. Liu Y, Cheng H, Zhang K (2023) Identifiability of Label Noise Transition Matrix. In: International Conference on Machine Learning. PMLR, pp 21475–21496

BibTeX

@InProceedings{pmlr-v202-liu23g,
  title = 	 {Identifiability of Label Noise Transition Matrix},
  author =       {Liu, Yang and Cheng, Hao and Zhang, Kun},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {21475--21496},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/liu23g/liu23g.pdf},
  url = 	 {https://proceedings.mlr.press/v202/liu23g.html},
  abstract = 	 {The noise transition matrix plays a central role in the problem of learning with noisy labels. Among many other reasons, a large number of existing solutions rely on the knowledge of it. Identifying and estimating the transition matrix without ground truth labels is a critical and challenging task. When label noise transition depends on each instance, the problem of identifying the instance-dependent noise transition matrix becomes substantially more challenging. Despite recently proposed solutions for learning from instance-dependent noisy labels, the literature lacks a unified understanding of when such a problem remains identifiable. The goal of this paper is to characterize the identifiability of the label noise transition matrix. Building on Kruskal’s identifiability results, we are able to show the necessity of multiple noisy labels in identifying the noise transition matrix at the instance level. We further instantiate the results to explain the successes of the state-of-the-art solutions and how additional assumptions alleviated the requirement of multiple noisy labels. Our result reveals that disentangled features improve identification. This discovery led us to an approach that improves the estimation of the transition matrix using properly disentangled features. Code is available at https://github.com/UCSC-REAL/Identifiability.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/