Uncertainty Estimation of Transformer Predictions for Misclassification Detection

Artem VazhentsevGleb KuzminArtem ShelmanovAkim TsvigunEvgenii TsymbalovKirill FedyaninMaxim PanovAlexander PanchenkoGleb GusevMikhail Burtsev

article2022ACL61 citations

Develops computationally efficient uncertainty estimation methods for Transformer models in text classification and named entity recognition, demonstrating that a spectral-normalized Mahalanobis distance approach can rival or exceed heavy deep ensembles at detecting misclassifications.

Listen

Modern language models based on Transformer architectures are widely deployed across high-stakes domains such as clinical medicine, compliance, and automated conversational systems. However, these models frequently produce erroneous or overconfident predictions. Identifying when a model is likely to make an error—known as misclassification detection through uncertainty estimation—is vital for deploying safe systems and determining when human intervention is needed. Existing high-performing uncertainty methods require running multiple models in parallel or performing numerous computational passes for a single prediction, resulting in severe latency, high energy costs, and massive memory footprints that prevent practical real-time deployment.

The article systematically evaluates state-of-the-art uncertainty estimation techniques on text processing tasks and introduces two new, computationally lightweight methods designed to detect model misclassifications without heavy resource overhead. The researchers conducted rigorous empirical experiments across text classification and named entity recognition benchmarks using large models (ELECTRA and DeBERTa). They assessed traditional output probabilities, resource-intensive ensembles and stochastic dropout methods, training loss regularizations, and deterministic distance-based approaches across multiple randomized experimental runs.

The findings establish that computationally cheap methods can match or exceed the performance of heavy, expensive techniques on text classification. The proposed approach combining Mahalanobis distance with spectral normalization achieved the best results among efficient alternatives, outperforming resource-intensive methods on benchmark tasks like sentence acceptability while reducing error risk curves by more than 46% compared to baseline softmax scores on paraphrase tasks. The second proposed modification, which samples diverse dropout masks only in the final classification layer, reduced computational overhead by over 99.5% relative to standard stochastic dropout while still improving upon the baseline. Additionally, simulation of human-in-the-loop workflows revealed that rejecting the most uncertain 40% of predictions for manual review pushed system accuracy above 98.5% using the lightweight distance-based method, outperforming deep ensembles.

These results provide a clear practical path to deploying reliable, uncertainty-aware language models at low operational cost. Organizations can dramatically reduce safety risks and error rates by routing uncertain outputs to human reviewers without purchasing additional hardware or slowing user response times. However, the evaluation also revealed task-dependent differences: while lightweight deterministic methods dominated standard text classification, sequence-level entity recognition still saw computational ensembles retain an advantage, and loss regularization techniques proved counterproductive for entity extraction tasks.

Decision-makers should adopt the spectral-normalized Mahalanobis distance technique for standard text classification pipelines seeking low-latency risk filtering, while reserving ensemble methods for complex sequence tagging until lightweight alternatives improve. Future research should focus on establishing stronger mathematical guarantees for spectral constraints within Transformer self-attention layers to close the remaining performance gap on token- and sequence-level tasks.

Cover for Uncertainty Estimation of Transformer Predictions for Misclassification Detection

Abstract

Uncertainty estimation (UE) of model predictions is a crucial step for a variety of tasks such as active learning, misclassification detection, adversarial attack detection, out-of-distribution detection, etc. Most of the works on modeling the uncertainty of deep neural networks evaluate these methods on image classification tasks. Little attention has been paid to UE in natural language processing. To fill this gap, we perform a vast empirical investigation of state-of-the-art UE methods for Transformer models on misclassification detection in named entity recognition and text classification tasks and propose two computationally efficient modifications, one of which approaches or even outperforms computationally intensive methods1.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background and Methods
  • 3.1 Softmax Response
  • 3.2 Monte Carlo Dropout
  • 3.3 Deterministic Uncertainty Estimation
  • 3.4 Training Loss Regularization
  • 4 Experimental Setup
  • 4.1 Metrics
  • 4.2 Datasets
  • 4.3 Model Choice and Hyperparameter Selection
  • 5 Results and Discussion
  • 5.1 Monte Carlo Dropout and Regularization
  • 5.2 Deterministic Methods
  • 5.3 Best Results
  • 6 Conclusion
  • Acknowledgements
  • References
  • A Dataset Statistics, Hyperparameter Values, and Hardware Configuration
  • B Additional Experimental Results with DeBERTa
  • C Additional Ablation Studies for DDPP

Knowls

  1. Knowl 1 — Mahalanobis uncertainty with spectral normalization

    model/method

    The paper’s MD SN method combines Mahalanobis-distance uncertainty estimation with spectral normalization of the Transformer’s classification-head weight matrix. Let WW be that matrix, ν=∥W∥2\nu=\|W\|_2 its spectral norm estimated by power iteration during training, and W~=W/ν\widetilde W=W/\nu. At inference, a penultimate-layer representation xx is transformed to h~(x)=W~x+b\widetilde h(x)=\widetilde W x+b. For each class cc, let μc\mu_c be the centroid of the transformed training representations and let Σ\Sigma be their covariance matrix. The uncertainty of a test example with representation h~i\widetilde h_i is its minimum class-conditional Mahalanobis distance:

    u_{\mathrm{MD}}(i)=\min_c(\widetilde h_i-\mu_c)^\top\Sigma^{-1}(\widetilde h_i-\mu_c).$$ The intended motivation is that spectral normalization makes representations more distance-preserving, benefiting a method that measures proximity to the training distribution. The normalization is applied to the classification-head matrix, rather than requiring a major change to the Transformer architecture.
  2. Knowl 2 — Diverse DPP Monte Carlo dropout

    model/method

    The proposed DDPP method increases the diversity of dropout masks used by DPP Monte Carlo dropout, which applies stochasticity to the Transformer’s final classification head while computing the rest of the network once. It first generates a pool of masks by sampling from a determinantal point process (DPP) whose kernel is the correlation matrix of neuron outputs on training data. It then selects the inference committee from this pool in one of two ways: (1) DDPP (+DPP) applies a second DPP to the mask pool, using an RBF-similarity kernel between mask vectors, to select masks that activate different neurons; or (2) DDPP (+OOD) selects masks with the highest probability-variance (PV) scores on an out-of-domain dataset, favoring masks that produce diverse predictions. The selected masks produce stochastic predictions from which the chosen Monte Carlo uncertainty score is calculated.

    For the reported experiments, the mask pool contained 100 masks and the selected committee contained 50 masks for text classification and 20 for CoNLL-2003. The maximum active-neuron fraction was selected by validation search: for ELECTRA, DDPP (+DPP) used 0.55, 0.45, 0.40, and 0.60 on MRPC, SST-2, CoLA, and CoNLL-2003, respectively; DDPP (+OOD) used 0.40, 0.35, 0.45, and 0.60. For DeBERTa, those fractions were 0.60, 0.60, 0.60, and 0.30 for DDPP (+DPP), and 0.45 on all four datasets for DDPP (+OOD). The OOD-selection experiments used 5,000 IMDB test examples. Because DPP masks usually activate fewer than half of the neurons and can yield poorly calibrated predictions, the paper temperature-scales predictions using a held-out dataset.

  3. Knowl 3 — Experimental scope and comparison setup

    experimental setup

    The empirical study evaluates uncertainty for misclassification detection in text classification and named entity recognition (NER), using ELECTRA (110 million parameters) and DeBERTa (138 million parameters). Text classification datasets are MRPC, SST-2, and CoLA; NER uses CoNLL-2003. The training data are subsampled to 10% for SST-2 and CoNLL-2003, and the available GLUE validation splits are used as test sets for the GLUE datasets. For DDPP (+OOD), the authors use 5,000 randomly selected examples from the IMDB test set to choose masks. Each experiment is run with six random seeds, and results are reported as means and standard deviations. Hyperparameters for UE methods are selected using a held-out validation set and RCC-AUC; model-training hyperparameters are selected separately using accuracy for classification and span-based F1 for sequence tagging.

  4. Knowl 4 — Uncertainty evaluation and sequence-level NER scoring

    definition

    The study evaluates whether uncertainty scores rank model errors above correct predictions using risk-coverage curve area (RCC-AUC) and reversed pair proportion (RPP); lower values are better for both. RCC-AUC is the area under cumulative misclassification risk as the fraction of retained predictions (coverage) varies, so a lower area indicates that rejecting high-uncertainty predictions removes errors earlier. For nn test instances, uncertainty scores uiu_i and indicator losses lil_i, the paper defines

    RPP=1n2∑i,j=1n1[ui>uj, li<lj].\mathrm{RPP}=\frac{1}{n^2}\sum_{i,j=1}^{n}\mathbf{1}[u_i>u_j,\ l_i<l_j].

    RPP counts pairs for which the more uncertain instance has lower loss than the less uncertain one; reported RPP values are multiplied by 100. For NER, token-level evaluation treats each token as an instance. Sequence-level loss is one if any token in the sequence is misclassified and zero otherwise. Sequence uncertainty is aggregated by taking the maximum token uncertainty for Mahalanobis-distance methods and the sum for the other methods.

  5. Knowl 5 — Training regularizers evaluated with uncertainty methods

    equation

    The experiments combine uncertainty estimators with two auxiliary training-loss regularizers. In either case, the task loss is augmented as L=Ltask+λLregL=L_{\mathrm{task}}+\lambda L_{\mathrm{reg}}, where λ\lambda controls regularization strength.

    For confident error regularization (CER), a batch has mm examples, picp_i^c is example ii’s predicted probability for class cc, and eie_i is defined in the paper as 1 when the prediction matches the true label and 0 otherwise. The loss is

    LCER=∑i,j=1mΔi,j1[ei>ej],Δi,j=max⁡{0,max⁡cpic−max⁡cpjc}2.L_{\mathrm{CER}}=\sum_{i,j=1}^{m}\Delta_{i,j}\mathbf{1}[e_i>e_j],\qquad \Delta_{i,j}=\max\{0,\max_c p_i^c-\max_c p_j^c\}^{2}.

    For metric regularization, hi∈Rdh_i\in\mathbb{R}^d is an example’s penultimate-layer representation, ScS_c is the batch subset with class cc, [z]+=max⁡(0,z)[z]_+=\max(0,z), and ϵ,γ>0\epsilon,\gamma>0 are hyperparameters. With D(hi,hj)=∥hi−hj∥22/dD(h_i,h_j)=\|h_i-h_j\|_2^2/d, the regularizer is

    Lmetric=∑c(Lintra(c)+ϵ∑q≠cLinter(c,q)),L_{\mathrm{metric}}=\sum_c\left(L_{\mathrm{intra}}(c)+\epsilon\sum_{q\ne c}L_{\mathrm{inter}}(c,q)\right),

    Lintra(c)=2∣Sc∣2−∣Sc∣∑i,j∈Sc, i<jD(hi,hj),Linter(c,q)=1∣Sc∣∣Sq∣∑i∈Sc,j∈Sq[γ−D(hi,hj)]+.L_{\mathrm{intra}}(c)=\frac{2}{|S_c|^2-|S_c|}\sum_{i,j\in S_c,\ i<j}D(h_i,h_j),\qquad L_{\mathrm{inter}}(c,q)=\frac{1}{|S_c||S_q|}\sum_{i\in S_c,j\in S_q}[\gamma-D(h_i,h_j)]_+.

    Thus the metric loss penalizes within-class distances and penalizes between-class distances that fall below the margin.

  6. Knowl 6 — MD SN results on ELECTRA

    empirical result

    On ELECTRA, MD SN reduced RCC-AUC relative to the softmax-response (SR) baseline on all three text-classification datasets for the reported configurations, with the strongest selected results varying by dataset. On MRPC, MD SN with metric regularization achieved RCC-AUC 12.04±1.3312.04\pm1.33 and RPP 1.56±0.121.56\pm0.12, compared with SR’s 22.32±8.0822.32\pm8.08 and 2.58±0.652.58\pm0.65. On SST-2, unregularized MD SN achieved 11.77±1.3311.77\pm1.33 and 0.83±0.080.83\pm0.08, versus SR’s 17.93±3.8417.93\pm3.84 and 1.22±0.281.22\pm0.28. On CoLA, MD SN with CER achieved 37.82±2.9137.82\pm2.91 and 1.90±0.121.90\pm0.12, versus SR’s 49.48±3.7149.48\pm3.71 and 2.35±0.252.35\pm0.25. On CoNLL-2003, MD SN did not improve both reported metrics at token level: its metric-regularized result was RCC-AUC 6.90±1.216.90\pm1.21 and RPP 0.11±0.020.11\pm0.02, while SR was 6.08±0.626.08\pm0.62 and 0.10±0.010.10\pm0.01. At sequence level, MD SN with metric regularization achieved 17.02±3.3917.02\pm3.39 and 2.01±0.402.01\pm0.40, compared with SR’s 18.81±3.3518.81\pm3.35 and 2.21±0.292.21\pm0.29.

  7. Knowl 7 — Relative performance of inexpensive and intensive methods

    empirical result

    Across both evaluated Transformer models, computationally inexpensive methods were generally on par with or better than computationally intensive methods on text classification. MD SN was the strongest inexpensive alternative on the text-classification tasks, and on CoLA it substantially outperformed even the intensive methods, including Monte Carlo dropout, deep ensembles, and MSD. The pattern differed for NER: intensive methods, especially deep ensembles, performed better. On token-level CoNLL-2003, only deep ensembles substantially improved on the SR baseline. At sequence level, some inexpensive methods—including MD with CER and the DDPP variants—improved on SR, but generally only approached the intensive methods’ performance. These results therefore support the paper’s efficiency claim for text classification more strongly than for NER.

  8. Knowl 8 — DDPP performance and inference cost

    empirical result

    DDPP did not generally outperform standard Monte Carlo dropout, but its diversity modifications produced useful gains over the SR baseline in selected settings. On ELECTRA sequence-level CoNLL-2003, DDPP (+DPP) with PV achieved RCC-AUC 16.78±2.4416.78\pm2.44 and RPP 1.93±0.201.93\pm0.20, while DDPP (+OOD) with PV achieved 16.75±2.3116.75\pm2.31 and 1.94±0.211.94\pm0.21; SR achieved 18.81±3.3518.81\pm3.35 and 2.21±0.292.21\pm0.29. On token-level CoNLL-2003, DDPP did not substantially outperform SR. The method’s central advantage is inference cost: its overhead is the same as the original DPP Monte Carlo dropout and is reported as less than 0.5% of the overhead introduced by standard Monte Carlo dropout. The paper’s ablation comparison also reports that adding mask diversity and calibration improves on the original DPP Monte Carlo-dropout setup in several evaluated settings, although gains are not uniform across datasets and metrics.

  9. Knowl 9 — Effect of regularization on uncertainty estimates

    empirical result

    CER and metric regularization often improved uncertainty results on text classification but were generally harmful for NER. Both regularizers substantially improved on the SR baseline for MRPC and SST-2 in the reported experiments; CER gave the strongest MRPC result, while metric regularization was especially effective on SST-2 and CoLA. Adding regularization to Monte Carlo dropout usually helped on text classification, and regularization often complemented DDPP there. For Mahalanobis-distance methods, regularization improved results on CoLA and CoNLL-2003, and also on SST-2 for unmodified MD and MRPC for MD SN. The paper reports that regularization reduced the difference between MD and MD SN on text classification and could give MD a slight advantage over MD SN on CoNLL-2003.

  10. Knowl 10 — Accuracy gains from selective rejection on MRPC

    empirical result

    The paper’s accuracy-rejection evaluation treats rejected examples as if a human expert supplies their correct labels. On MRPC with ELECTRA, rejecting the 20% most uncertain examples using Monte Carlo dropout with CER and PV increased accuracy from 88.4% to 96.0%, a 1.3-percentage-point gain over the corresponding SR-baseline result. At a 40% rejection rate, deep ensembles achieved 98.2% accuracy, while the inexpensive MD SN method with metric regularization achieved 98.5%; the latter was reported as 1.7 percentage points above the SR baseline. These results illustrate how uncertainty ranking can reduce the amount of human relabeling needed to reach a target accuracy.

  11. Knowl 11 — Limitation of spectral-normalization guarantees for Transformers

    limitation

    The theoretical motivation for spectral normalization does not provide a general bi-Lipschitz guarantee for the Transformer architecture used in the experiments. Spectral normalization is proven to enforce the relevant constraint for transformations formed by standard residual-network blocks under the stated conditions, but Transformer self-attention blocks have a more complicated structure, so that guarantee does not generally transfer. Consequently, MD SN’s distance-preservation rationale is empirically motivated for these Transformers rather than established by the cited theoretical result; the paper identifies methods for constraining self-attention blocks as future work.

Coverage note — The complete per-configuration score grids, full dataset statistics, and exhaustive tuned hyperparameter tables are not reproduced; they are repetitive supporting details, while the main methods, evaluation design, comparative outcomes, and stated limitation are captured.

References

  1. 1.Arsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, and Dmitry P. Vetrov. 2020. Pitfalls of in-domain uncertainty estimation and ensembling in deep learning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  2. 2.Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra. 2015. Weight uncertainty in neural network. In Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, volume 37 of JMLR Workshop and Conference Proceedings, pages 1613–1622. JMLR.org.
  3. 3.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: pre-training text encoders as discriminators rather than generators. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  4. 4.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Stroudsburg, PA, USA. Association for Computational Linguistics.
  5. 5.William B. Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005).
  6. 6.Ran El-Yaniv and Yair Wiener. 2010. On the foundations of noise-free selective classification. J. Mach. Learn. Res., 11:1605–1641.
  7. 7.Angelos Filos, Sebastian Farquhar, Aidan N. Gomez, Tim G. J. Rudner, Zachary Kenton, Lewis Smith, Milad Alizadeh, Arnoud de Kroon, and Yarin Gal. 2019. A systematic comparison of bayesian deep learning robustness in diabetic retinopathy tasks. CoRR, abs/1912.10481.
  8. 8.Yarin Gal and Zoubin Ghahramani. 2016. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 1050–1059, New York, New York, USA. PMLR.
  9. 9.Yarin Gal, Riashat Islam, and Zoubin Ghahramani. 2017. Deep bayesian active learning with image data. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pages 1183–1192. PMLR.
  10. 10.Yonatan Geifman and Ran El-Yaniv. 2017. Selective classification for deep neural networks. Advances in Neural Information Processing Systems, 30:4878–4887.
  11. 11.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pages 1321–1330. PMLR.
  12. 12.Jianfeng He, Xuchao Zhang, Shuo Lei, Zhiqian Chen, Fanglan Chen, Abdulaziz Alhamadani, Bei Xiao, and Chang-Tien Lu. 2020. Towards more accurate uncertainty estimation in text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 8362–8372. Association for Computational Linguistics.
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 770–778. IEEE Computer Society.
  14. 14.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: decoding-enhanced bert with disentangled attention. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  15. 15.Neil Houlsby, Ferenc Huszar, Zoubin Ghahramani, and Máté Lengyel. 2011. Bayesian active learning for classification and preference learning. CoRR, abs/1112.5745.
  16. 16.Yibo Hu and Latifur Khan. 2021. Uncertainty-aware reliable text classification. In KDD ’21: The 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Virtual Event, Singapore, August 14-18, 2021, pages 628–636. ACM.
  17. 17.Siddhartha Jain, Ge Liu, Jonas Mueller, and David Gifford. 2020. Maximizing overall diversity for improved uncertainty estimates in deep ensembles. Proceedings of the AAAI Conference on Artificial Intelligence, 34(04):4264–4271.
  18. 18.Alex Kulesza and Ben Taskar. 2012. Determinantal point processes for machine learning. Foundations and Trends® in Machine Learning, 5(2–3):123–286.
  19. 19.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017. Simple and scalable predictive uncertainty estimation using deep ensembles. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NeurIPS’17, page 6405–6416, Red Hook, NY, USA. Curran Associates Inc.
  20. 20.Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. 2018. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, volume 31, pages 7167–7177.
  21. 21.Jeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran, Tania Bedrax-Weiss, and Balaji Lakshminarayanan. 2020. Simple and principled uncertainty estimation with deterministic deep learning via distance awareness. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  22. 22.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized bert pretraining approach.
  23. 23.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  24. 24.Andrey Malinin and Mark J. F. Gales. 2018. Predictive uncertainty estimation via prior networks. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pages 7047–7058.
  25. 25.Jishnu Mukhoti, Andreas Kirsch, Joost van Amersfoort, Philip H. S. Torr, and Yarin Gal. 2021. Deterministic neural networks with appropriate inductive biases capture epistemic and aleatoric uncertainty. CoRR, abs/2102.11582.
  26. 26.Alexander Podolskiy, Dmitry Lipin, Andrey Bout, Ekaterina Artemova, and Irina Piontkovskaya. 2021. Revisiting mahalanobis distance for transformer-based out-of-domain detection. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pages 13675–13682. AAAI Press.
  27. 27.Carl Edward Rasmussen. 2003. Gaussian processes in machine learning. In Advanced Lectures on Machine Learning, ML Summer Schools 2003, Canberra, Australia, February 2-14, 2003, Tübingen, Germany, August 4-16, 2003, Revised Lectures, volume 3176 of Lecture Notes in Computer Science, pages 63–71. Springer.
  28. 28.Burr Settles. 2009. Active learning literature survey. Computer Sciences Technical Report 1648, University of Wisconsin–Madison.
  29. 29.Artem Shelmanov, Evgenii Tsymbalov, Dmitri Puzyrev, Kirill Fedyanin, Alexander Panchenko, and Maxim Panov. 2021. How certain is your Transformer? In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1833–1840, Online. Association for Computational Linguistics.
  30. 30.Yilin Shen, Yen-Chang Hsu, Avik Ray, and Hongxia Jin. 2021. Enhancing the generalization for intent classification and out-of-domain detection in SLU. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2443–2453, Online. Association for Computational Linguistics.
  31. 31.Lewis Smith and Yarin Gal. 2018. Understanding measures of uncertainty for adversarial example detection. In Proceedings of the Thirty-Fourth Conference on Uncertainty in Artificial Intelligence, UAI 2018, Monterey, California, USA, August 6-10, 2018, pages 560–569. AUAI Press.
  32. 32.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
  33. 33.Sunil Thulasidasan, Gopinath Chennupati, Jeff A. Bilmes, Tanmoy Bhattacharya, and Sarah Michalak. 2019. On mixup training: Improved calibration and predictive uncertainty for deep neural networks. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13888–13899.
  34. 34.Erik F. Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, pages 142–147.
  35. 35.Joost Van Amersfoort, Lewis Smith, Yee Whye Teh, and Yarin Gal. 2020. Uncertainty estimation using a single deep deterministic neural network. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 9690–9700. PMLR.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, 4-9 December 2017, Long Beach, CA, USA, pages 5998–6008.
  37. 37.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
  38. 38.Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641.
  39. 39.Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. 2021. The art of abstention: Selective prediction and error regularization for natural language processing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1040–1051, Online. Association for Computational Linguistics.
  40. 40.Zhiyuan Zeng, Keqing He, Yuanmeng Yan, Zijun Liu, Yanan Wu, Hong Xu, Huixing Jiang, and Weiran Xu. 2021. Modeling discriminative representations for out-of-domain detection with supervised contrastive learning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 870–878, Online. Association for Computational Linguistics.
  41. 41.Xuchao Zhang, Fanglan Chen, Chang-Tien Lu, and Naren Ramakrishnan. 2019. Mitigating uncertainty in document classification. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3126–3136, Minneapolis, Minnesota. Association for Computational Linguistics.

Citation

MLA
Vazhentsev, A., et al. “Uncertainty Estimation of Transformer Predictions for Misclassification Detection”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 8237–52, https://doi.org/10.18653/v1/2022.acl-long.566.
APA
Vazhentsev, A., Kuzmin, G., Shelmanov, A., Tsvigun, A., Tsymbalov, E., Fedyanin, K., Panov, M., Panchenko, A., Gusev, G., Burtsev, M., Avetisian, M., & Zhukov, L. (2022). Uncertainty Estimation of Transformer Predictions for Misclassification Detection. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8237–8252. https://doi.org/10.18653/v1/2022.acl-long.566
Chicago
Vazhentsev, A., G. Kuzmin, A. Shelmanov, et al. 2022. “Uncertainty Estimation of Transformer Predictions for Misclassification Detection”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8237–52. https://doi.org/10.18653/v1/2022.acl-long.566.
Harvard
Vazhentsev, A. et al. (2022) “Uncertainty Estimation of Transformer Predictions for Misclassification Detection”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8237–8252. Available at: https://doi.org/10.18653/v1/2022.acl-long.566.
Vancouver
1. Vazhentsev A, Kuzmin G, Shelmanov A, et al (2022) Uncertainty Estimation of Transformer Predictions for Misclassification Detection. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8237–8252

BibTeX

@inproceedings{vazhentsev-etal-2022-uncertainty,
    title = "Uncertainty Estimation of Transformer Predictions for Misclassification Detection",
    author = "Vazhentsev, Artem  and
      Kuzmin, Gleb  and
      Shelmanov, Artem  and
      Tsvigun, Akim  and
      Tsymbalov, Evgenii  and
      Fedyanin, Kirill  and
      Panov, Maxim  and
      Panchenko, Alexander  and
      Gusev, Gleb  and
      Burtsev, Mikhail  and
      Avetisian, Manvel  and
      Zhukov, Leonid",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.566/",
    doi = "10.18653/v1/2022.acl-long.566",
    pages = "8237--8252"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/