keyword
misclassification detection
Misclassification detection is the process of identifying when a machine learning classification model has produced or is likely to produce an erroneous prediction on a given input. Typically implemented through uncertainty estimation, confidence scoring, or anomaly detection techniques, it aims to assess the reliability of a model output without relying on immediate access to ground-truth labels. By separating dependable predictions from potential mistakes, misclassification detection allows automated systems to selectively abstain from uncertain decisions, route ambiguous cases to human review, and improve overall system safety and reliability in critical decision-making environments.
2 items

Hybrid Uncertainty Quantification for Selective Text Classification in Ambiguous Tasks
Artem Vazhentsev, Gleb Kuzmin, Akim Tsvigun, Alexander Panchenko, Maxim Panov, Mikhail Burtsev, Artem Shelmanov
Why you should read this
Proposes a hybrid uncertainty quantification method that unites epistemic and aleatoric measures to reliably identify prediction errors and improve selective classification in subjective natural language processing tasks like toxicity detection.
Many text classification tasks are inherently ambiguous, which results in automatic systems having a high risk of making mistakes, in spite of using advanced machine learning models. For example, toxicity detection in user-generated content is a subjective task, and notions of toxicity can be annotated according to a variety of definitions that can be in conflict with one another. Instead of relying solely on automatic solutions, moderation of the most difficult and ambiguous cases can be delegated to human workers. Potential mistakes in automated classification can be identified by using uncertainty estimation (UE) techniques. Although UE is a rapidly growing field within natural language processing, we find that state-of-the-art UE methods estimate only epistemic uncertainty and show poor performance, or under-perform trivial methods for ambiguous tasks such as toxicity detection. We argue that in order to create robust uncertainty estimation methods for ambiguous tasks it is necessary to account also for aleatoric uncertainty. In this paper, we propose a new uncertainty estimation method that combines epistemic and aleatoric UE methods. We show that by using our hybrid method, we can outperform state-of-the-art UE methods for toxicity detection and other ambiguous text classification tasks¹.
Added
2026-10-05


Uncertainty Estimation of Transformer Predictions for Misclassification Detection
Artem Vazhentsev, Gleb Kuzmin, Artem Shelmanov, Akim Tsvigun, Evgenii Tsymbalov, Kirill Fedyanin, Maxim Panov, Alexander Panchenko, Gleb Gusev, Mikhail Burtsev, Manvel Avetisian, Leonid Zhukov
Why you should read this
Develops computationally efficient uncertainty estimation methods for Transformer models in text classification and named entity recognition, demonstrating that a spectral-normalized Mahalanobis distance approach can rival or exceed heavy deep ensembles at detecting misclassifications.
Uncertainty estimation (UE) of model predictions is a crucial step for a variety of tasks such as active learning, misclassification detection, adversarial attack detection, out-of-distribution detection, etc. Most of the works on modeling the uncertainty of deep neural networks evaluate these methods on image classification tasks. Little attention has been paid to UE in natural language processing. To fill this gap, we perform a vast empirical investigation of state-of-the-art UE methods for Transformer models on misclassification detection in named entity recognition and text classification tasks and propose two computationally efficient modifications, one of which approaches or even outperforms computationally intensive methods1.
Added
2026-10-03
