Metrics for Multi-Class Classification: an Overview
Margherita GrandiniEnrico BagliGiorgio Visani
Analyzes key multi-class classification metrics by detailing their individual strengths, limitations, and practical applications across model selection and hyperparameter tuning.
Modern predictive applications frequently require machine learning models to categorize data across more than two potential outcomes. Evaluating these multi-class models correctly is vital, as selecting an improper performance metric can create misleading impressions of operational readiness and obscure critical errors in automated decision-making. The article evaluates the principal performance indicators used in multi-class classification, assessing their mathematical formulation, behavioral characteristics, operational advantages, and inherent limitations.
To conduct this evaluation, the analysis examines standard cross-tabulation tools (confusion matrices) and probability distributions across varied data environments. The review examines standard accuracy, balanced accuracy variations, harmonic score averages, probabilistic cross-entropy, the Matthews correlation coefficient, and Cohen's Kappa score across both balanced and heavily imbalanced class distributions.
The findings show that standard accuracy treats all data records equally, making it effective for evenly distributed classes but highly vulnerable when categories are imbalanced. For example, in an evaluated imbalanced scenario, overall accuracy registered at 0.689 while balanced accuracy dropped to 0.615, demonstrating how dominant classes can mask severe failure rates in rare categories where individual recall might fall near 0.08. Additionally, the analysis proves that calculating an aggregated overall harmonic mean across individual records (Micro F1-Score) produces the exact same value as standard accuracy, whereas averaging metrics on a per-category basis (Macro F1-Score) grants equal importance to all classes regardless of size. The article also finds that probabilistic cross-entropy assesses predictions without relying on confusion matrices but evaluates only the assigned probability of the correct class, ignoring broader probability distribution shifts across incorrect classes. Finally, correlation and agreement metrics such as the Matthews correlation coefficient and Cohen's Kappa incorporate every cell of the evaluation matrix, effectively identifying models that default completely to a single majority class.
These distinctions carry significant operational and risk implications. Relying solely on overall accuracy in heavily imbalanced business environments exposes organizations to substantial failure risks by concealing underperformance in rare, high-consequence categories. Conversely, selecting class-level measures allows leadership to pinpoint specific functional weaknesses before deploying algorithms into production environments.
Organizations should align metric selection directly with operational priorities: standard accuracy should be reserved for balanced datasets, class-averaged or weighted recall measures should be used to protect minority categories, and Cohen's Kappa should be applied when benchmarking performance across distinct datasets. Analysts must also account for specific metric vulnerabilities, such as instability in Matthews correlation coefficients during training or cross-entropy's blind spots regarding alternative class probabilities.
- Paper: A Simple Generalisation of the Area Under the ROC Curve for Multiple Class Classification Problems, DAVID J. HAND et al. (2001). This seminal paper introduces the pairwise averaging generalization of the AUC metric for multi-class problems, serving as foundational reading for understanding multi-class evaluation criteria.
- Paper: Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation, David M. W. Powers (2011). This work establishes the theoretical limitations of precision, recall, and accuracy while developing chance-corrected evaluation metrics essential for rigorous classification assessment.
- Paper: Using AUC and accuracy in evaluating learning algorithms, Jin Huang et al. (2005). It proves the formal discriminancy and statistical consistency of AUC over standard predictive accuracy in both binary and multi-class settings.
- Paper: Robust Classification for Imprecise Environments, F. Provost et al. (2000). It presents ROC convex hull analysis to evaluate classifier performance under shifting class distributions and unequal misclassification costs.
- Paper: Statistical Comparisons of Classifiers over Multiple Data Sets, Janez Demšar (2006). This text provides standard statistical testing frameworks for reliably comparing classification algorithms and metric outcomes across multiple benchmark datasets.
- Paper: Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms, Thomas G. Dietterich (1998). It offers essential groundwork on the statistical significance and error behavior of metrics used when comparing supervised classification models.
- Paper: On the Evaluation of (Meta-)solver Approaches, Roberto Amadini et al. (2023). This paper extends performance evaluation methodology to meta-solvers, analyzing how metric selection and ranking criteria can alter comparative conclusions.
