keyword
ROC analysis
ROC analysis, or Receiver Operating Characteristic analysis, is a statistical and graphical technique used to evaluate, visualize, and compare the performance of binary classification models across all possible decision thresholds. The method plots a model's true positive rate against its false positive rate as the discrimination threshold varies from one extreme to the other. Because these rates are independent of class distributions and misclassification costs, ROC analysis provides a robust framework for assessing a classifier's intrinsic diagnostic ability in uncertain or skewed environments. It also supports model selection and threshold optimization through tools such as the area under the curve metric and the ROC convex hull, which identifies the best attainable performance under varying operational conditions.
4 items

Optimal Strategies for Reject Option Classifiers
Vojtech Franc, Daniel Prusa, Václav Vorácek
Why you should read this
Unifies cost-based, bounded-improvement, and bounded-abstention selective classification models by proving they share the same optimal strategy, while developing two Fisher consistent algorithms to learn optimal rejection functions for arbitrary black-box classifiers across diverse prediction tasks.
In classification with a reject option, the classifier is allowed in uncertain cases to abstain from prediction. The classical cost-based model of a reject option classifier requires the rejection cost to be defined explicitly. The alternative bounded-improvement model and the bounded-abstention model avoid the notion of the reject cost. The bounded-improvement model seeks a classifier with a guaranteed selective risk and maximal cover. The bounded-abstention model seeks a classifier with guaranteed cover and minimal selective risk. We prove that despite their different formulations the three rejection models lead to the same prediction strategy: the Bayes classifier endowed with a randomized Bayes selection function. We define the notion of a proper uncertainty score as a scalar summary of the prediction uncertainty sufficient to construct the randomized Bayes selection function. We propose two algorithms to learn the proper uncertainty score from examples for an arbitrary black-box classifier. We prove that both algorithms provide Fisher consistent estimates of the proper uncertainty score and demonstrate their efficiency in different prediction problems, including classification, ordinal regression, and structured output classification.
Added
2026-09-26

Robust Classification for Imprecise Environments
F. Provost, Tom Fawcett
Why you should read this
Introduces the ROC convex hull method to construct hybrid classifiers that guarantee optimal performance across changing misclassification costs and class distributions without committing to a single model in advance.
In real-world environments it usually is difficult to specify target operating conditions precisely, for example, target misclassification costs. This uncertainty makes building robust classification systems problematic. We show that it is possible to build a hybrid classifier that will perform at least as well as the best available classifier for any target conditions. In some cases, the performance of the hybrid actually can surpass that of the best known classifier. This robust performance extends across a wide variety of comparison frameworks, including the optimization of metrics such as accuracy, expected cost, lift, precision, recall, and workforce utilization. The hybrid also is efficient to build, to store, and to update. The hybrid is based on a method for the comparison of classifier performance that is robust to imprecise class distributions and misclassification costs. The ROC convex hull (ROCCH) method combines techniques from ROC analysis, decision analysis and computational geometry, and adapts them to the particulars of analyzing learned classifiers. The method is efficient and incremental, minimizes the management of classifier performance data, and allows for clear visual comparisons and sensitivity analyses. Finally, we point to empirical evidence that a robust hybrid classifier indeed is needed for many real-world problems.
Added
2026-09-25

Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation
David M. W. Powers
Why you should read this
Exposes critical biases in standard metrics like Precision, Recall, and F-measure that allow inferior models to appear superior, while presenting chance-corrected alternatives such as Informedness and Markedness for rigorous binary and multi-class evaluation.
Commonly used evaluation measures including Recall, Precision, F-Measure and Rand Accuracy are biased and should not be used without clear understanding of the biases, and corresponding identification of chance or base case levels of the statistic. Using these measures a system that performs worse in the objective sense of Informedness, can appear to perform better under any of these commonly used measures. We discuss several concepts and measures that reflect the probability that prediction is informed versus chance. Informedness and introduce Markedness as a dual measure for the probability that prediction is marked versus chance. Finally we demonstrate elegant connections between the concepts of Informedness, Markedness, Correlation and Significance as well as their intuitive relationships with Recall and Precision, and outline the extension from the dichotomous case to the general multi-class case.
Added
2026-09-12

The relationship between Precision-Recall and ROC curves
Jesse Davis, Mark Goadrich
Why you should read this
Demonstrates mathematically why Precision-Recall curves fundamentally outperform ROC curves for imbalanced datasets, establishing core principles of metric selection.
Receiver Operator Characteristic (ROC) curves are commonly used to present results for binary decision problems in machine learning. However, when dealing with highly skewed datasets, Precision-Recall (PR) curves give a more informative picture of an algorithm's performance. We show that a deep connection exists between ROC space and PR space, such that a curve dominates in ROC space if and only if it dominates in PR space. A corollary is the notion of an achievable PR curve, which has properties much like the convex hull in ROC space; we show an efficient algorithm for computing this curve. Finally, we also note differences in the two types of curves are significant for algorithm design. For example, in PR space it is incorrect to linearly interpolate between points. Furthermore, algorithms that optimize the area under the ROC curve are not guaranteed to optimize the area under the PR curve.
Added
2026-03-22
