Metrics for Multi-Class Classification: an Overview

Margherita GrandiniEnrico BagliGiorgio Visani

article2020arXiv1,327 citations

Analyzes key multi-class classification metrics by detailing their individual strengths, limitations, and practical applications across model selection and hyperparameter tuning.

Listen

Modern predictive applications frequently require machine learning models to categorize data across more than two potential outcomes. Evaluating these multi-class models correctly is vital, as selecting an improper performance metric can create misleading impressions of operational readiness and obscure critical errors in automated decision-making. The article evaluates the principal performance indicators used in multi-class classification, assessing their mathematical formulation, behavioral characteristics, operational advantages, and inherent limitations.

To conduct this evaluation, the analysis examines standard cross-tabulation tools (confusion matrices) and probability distributions across varied data environments. The review examines standard accuracy, balanced accuracy variations, harmonic score averages, probabilistic cross-entropy, the Matthews correlation coefficient, and Cohen's Kappa score across both balanced and heavily imbalanced class distributions.

The findings show that standard accuracy treats all data records equally, making it effective for evenly distributed classes but highly vulnerable when categories are imbalanced. For example, in an evaluated imbalanced scenario, overall accuracy registered at 0.689 while balanced accuracy dropped to 0.615, demonstrating how dominant classes can mask severe failure rates in rare categories where individual recall might fall near 0.08. Additionally, the analysis proves that calculating an aggregated overall harmonic mean across individual records (Micro F1-Score) produces the exact same value as standard accuracy, whereas averaging metrics on a per-category basis (Macro F1-Score) grants equal importance to all classes regardless of size. The article also finds that probabilistic cross-entropy assesses predictions without relying on confusion matrices but evaluates only the assigned probability of the correct class, ignoring broader probability distribution shifts across incorrect classes. Finally, correlation and agreement metrics such as the Matthews correlation coefficient and Cohen's Kappa incorporate every cell of the evaluation matrix, effectively identifying models that default completely to a single majority class.

These distinctions carry significant operational and risk implications. Relying solely on overall accuracy in heavily imbalanced business environments exposes organizations to substantial failure risks by concealing underperformance in rare, high-consequence categories. Conversely, selecting class-level measures allows leadership to pinpoint specific functional weaknesses before deploying algorithms into production environments.

Organizations should align metric selection directly with operational priorities: standard accuracy should be reserved for balanced datasets, class-averaged or weighted recall measures should be used to protect minority categories, and Cohen's Kappa should be applied when benchmarking performance across distinct datasets. Analysts must also account for specific metric vulnerabilities, such as instability in Matthews correlation coefficients during training or cross-entropy's blind spots regarding alternative class probabilities.

arXiv: 2008.05756
  • Paper: On the Evaluation of (Meta-)solver Approaches, Roberto Amadini et al. (2023). This paper extends performance evaluation methodology to meta-solvers, analyzing how metric selection and ranking criteria can alter comparative conclusions.
Cover for Metrics for Multi-Class Classification: an Overview

Abstract

Classification tasks in machine learning involving more than two classes are known by the name of "multi-class classification". Performance indicators are very useful when the aim is to evaluate and compare different classification models or machine learning techniques. Many metrics come in handy to test the ability of a multi-class classifier. Those metrics turn out to be useful at different stage of the development process, e.g. comparing the performance of two different models or analysing the behaviour of the same model by tuning different parameters. In this white paper we review a list of the most promising multi-class metrics, we highlight their advantages and disadvantages and show their possible usages during the development of a classification model.

Table of Contents

  • 1 Introduction
  • 1.1 Confusion Matrix
  • 1.2 Precision & Recall
  • 2 Accuracy
  • 3 Balanced Accuracy
  • 3.1 Balanced Accuracy Weighted
  • 4 F1-Score
  • 4.1 F1-Score Binary case
  • 4.2 F1-Score Multi-class case
  • 4.2.1 Macro F1-Score
  • 4.3 Micro F1-Score
  • 4.4 Cross Entropy
  • 5 Independence between two Random Discrete Variables
  • 5.1 Mattheus Correlation Coefficient
  • 5.1.1 Mattheus Correlation Coefficient for binary classification
  • 5.1.2 Mattheus Correlation Coefficient for Multi-class Classification
  • 5.1.3 Pros and Cons of MCC
  • 5.2 Cohen’s Kappa
  • 5.2.1 Cohen’s Kappa for binary classification
  • 5.2.2 Cohen’s Kappa for multi-class cases
  • 5.2.3 Useful Applications
  • 6 Conclusions
  • References

Knowls

  1. Knowl 1 — Multi-Class Classification Accuracy

    equation

    In a multi-class classification problem with KK classes, let C∈RK×KC \in \mathbb{R}^{K \times K} denote the confusion matrix where entry CijC_{ij} represents the number of samples whose true class is ii and predicted class is jj. The overall classification accuracy is the proportion of correctly classified instances out of all evaluated instances:

    Accuracy=∑k=1KCkk∑i=1K∑j=1KCij\text{Accuracy} = \frac{\sum_{k=1}^K C_{kk}}{\sum_{i=1}^K \sum_{j=1}^K C_{ij}}

    Accuracy measures the empirical probability that an individual chosen uniformly at random from the dataset is correctly classified. Every sample contributes equally to the metric, which means that performance on majority (highly populated) classes dominates the score, while performance on minority classes has minimal impact. Consequently, high accuracy can conceal significant misclassification rates on minority classes in imbalanced datasets.

  2. Knowl 2 — Multi-Class Balanced Accuracy

    equation

    For a multi-class classification problem with KK classes and a confusion matrix C∈RK×KC \in \mathbb{R}^{K \times K}, the Balanced Accuracy is defined as the unweighted arithmetic mean of the recall (sensitivity) achieved across each individual class:

    Balanced Accuracy=1K∑k=1KRecallk=1K∑k=1KCkk∑j=1KCkj\text{Balanced Accuracy} = \frac{1}{K} \sum_{k=1}^K \text{Recall}_k = \frac{1}{K} \sum_{k=1}^K \frac{C_{kk}}{\sum_{j=1}^K C_{kj}}

    where Recallk=CkkTotalrow,k\text{Recall}_k = \frac{C_{kk}}{\text{Total}_{\text{row}, k}} represents the proportion of instances with true label kk that are correctly predicted as class kk.

    By assigning equal weight (1/K1/K) to each class regardless of its size in the dataset, Balanced Accuracy is insensitive to class imbalance and prevents high accuracy on majority classes from masking poor performance on under-represented classes. When the dataset is perfectly balanced across classes, Balanced Accuracy converges to standard classification accuracy.

  3. Knowl 3 — Multi-Class Matthews Correlation Coefficient (MCC)

    equation

    The multi-class extension of the Matthews Correlation Coefficient (MCC), also known as the KK-category correlation coefficient, measures the correlation between true and predicted categorical assignments using all entries of the confusion matrix C∈RK×KC \in \mathbb{R}^{K \times K}.

    Let:

    • c=∑k=1KCkkc = \sum_{k=1}^K C_{kk} (total number of correctly classified instances)
    • s=∑i=1K∑j=1KCijs = \sum_{i=1}^K \sum_{j=1}^K C_{ij} (total number of instances)
    • pk=∑i=1KCikp_k = \sum_{i=1}^K C_{ik} (number of times class kk was predicted, corresponding to column total kk)
    • tk=∑j=1KCkjt_k = \sum_{j=1}^K C_{kj} (number of times class kk truly occurred, corresponding to row total kk)

    The multi-class MCC is defined as:

    MCC=c⋅s−∑k=1Kpktk(s2−∑k=1Kpk2)(s2−∑k=1Ktk2)\text{MCC} = \frac{c \cdot s - \sum_{k=1}^K p_k t_k}{\sqrt{\left(s^2 - \sum_{k=1}^K p_k^2\right)\left(s^2 - \sum_{k=1}^K t_k^2\right)}}

    Properties of multi-class MCC:

    • Range: MCC∈[−1,1]\text{MCC} \in [-1, 1]. An MCC of +1+1 indicates perfect prediction, 00 indicates performance equivalent to random assignment, and negative values indicate inverse prediction.
    • Robustness: Because every cell in the confusion matrix is incorporated in both numerator and denominator, MCC provides a balanced evaluation even under severe class imbalance. If a model assigns all samples to a single class, MCC drops to 00.
    • Lower bound variability: The minimum possible value of MCC is not universally −1-1; depending on KK and the marginal distribution of the ground-truth classes, the minimum attainable value fluctuates between −1-1 and 00.
  4. Knowl 4 — Multi-Class Cohen's Kappa Statistic

    equation

    Cohen's Kappa statistic (κ\kappa) evaluates the concordance between predicted classifications and true ground-truth labels by adjusting observed accuracy for agreement expected purely by chance. For a KK-class confusion matrix C∈RK×KC \in \mathbb{R}^{K \times K} with total samples s=∑i=1K∑j=1KCijs = \sum_{i=1}^K \sum_{j=1}^K C_{ij}, correctly predicted samples c=∑k=1KCkkc = \sum_{k=1}^K C_{kk}, predicted column totals pk=∑i=1KCikp_k = \sum_{i=1}^K C_{ik}, and actual row totals tk=∑j=1KCkjt_k = \sum_{j=1}^K C_{kj}:

    κ=Po−Pe1−Pe=c⋅s−∑k=1Kpktks2−∑k=1Kpktk\kappa = \frac{P_o - P_e}{1 - P_e} = \frac{c \cdot s - \sum_{k=1}^K p_k t_k}{s^2 - \sum_{k=1}^K p_k t_k}

    where:

    • Po=csP_o = \frac{c}{s} is the observed accuracy.
    • Pe=∑k=1Kpktks2P_e = \frac{\sum_{k=1}^K p_k t_k}{s^2} is the expected accuracy by chance under the assumption that predicted and actual classifications are statistically independent random categorical variables.

    Properties and comparison:

    • The statistic takes values in [−1,+1][-1, +1], where κ=1\kappa = 1 denotes complete agreement, κ=0\kappa = 0 denotes agreement equal to random chance, and κ<0\kappa < 0 denotes agreement worse than chance.
    • Multi-class Cohen's Kappa shares an identical numerator with multi-class MCC (c⋅s−∑k=1Kpktkc \cdot s - \sum_{k=1}^K p_k t_k), but its denominator (s2−∑k=1Kpktks^2 - \sum_{k=1}^K p_k t_k) is smaller than the geometric mean used in MCC, yielding slightly higher numerical values for κ\kappa.
    • Because subtracting PeP_e eliminates the baseline agreement attributable to dataset-specific marginal class distributions, κ\kappa allows performance comparison across distinct datasets with differing class distributions.
  5. Knowl 5 — Macro-Averaged F1-Score

    equation

    In multi-class classification across KK classes, the Macro F1F_1-score is the harmonic mean of the unweighted arithmetic means of per-class Precision and per-class Recall. For each class k∈{1,…,K}k \in \{1, \dots, K\}, per-class precision and recall are computed using true positives (TPk=CkkTP_k = C_{kk}), false positives (FPk=∑i≠kCikFP_k = \sum_{i \neq k} C_{ik}), and false negatives (FNk=∑j≠kCkjFN_k = \sum_{j \neq k} C_{kj}):

    Precisionk=TPkTPk+FPk,Recallk=TPkTPk+FNk\text{Precision}_k = \frac{TP_k}{TP_k + FP_k}, \quad \text{Recall}_k = \frac{TP_k}{TP_k + FN_k}

    The macro-averaged precision and recall are:

    Macro-Precision=1K∑k=1KPrecisionk,Macro-Recall=1K∑k=1KRecallk\text{Macro-Precision} = \frac{1}{K} \sum_{k=1}^K \text{Precision}_k, \quad \text{Macro-Recall} = \frac{1}{K} \sum_{k=1}^K \text{Recall}_k

    The Macro F1F_1-score is then defined as:

    Macro F1-Score=2⋅Macro-Precision⋅Macro-RecallMacro-Precision+Macro-Recall\text{Macro } F_1\text{-Score} = 2 \cdot \frac{\text{Macro-Precision} \cdot \text{Macro-Recall}}{\text{Macro-Precision} + \text{Macro-Recall}}

    Because each class contributes equally to the macro averages regardless of class frequency, the Macro F1F_1-score evaluates model effectiveness at the class level and is heavily penalized if any class (including minority classes) suffers from poor precision or recall.

  6. Knowl 6 — Equivalence of Micro-Averaged F1-Score to Accuracy

    theoretical result

    In multi-class classification with KK classes and a dataset of total size N=∑i=1K∑j=1KCijN = \sum_{i=1}^K \sum_{j=1}^K C_{ij}, micro-averaging aggregates true positives, false positives, and false negatives globally across all classes before computing precision and recall.

    The micro-averaged precision and micro-averaged recall are defined as:

    Micro-Precision=∑k=1KTPk∑k=1K(Total Columnk)=∑k=1KTPkN\text{Micro-Precision} = \frac{\sum_{k=1}^K TP_k}{\sum_{k=1}^K (\text{Total Column}_k)} = \frac{\sum_{k=1}^K TP_k}{N}

    Micro-Recall=∑k=1KTPk∑k=1K(Total Rowk)=∑k=1KTPkN\text{Micro-Recall} = \frac{\sum_{k=1}^K TP_k}{\sum_{k=1}^K (\text{Total Row}_k)} = \frac{\sum_{k=1}^K TP_k}{N}

    Because the sum of all column totals and the sum of all row totals both equal the total number of samples NN, Micro-Precision\text{Micro-Precision} and Micro-Recall\text{Micro-Recall} are mathematically identical. Consequently, their harmonic mean, the Micro F1F_1-score, simplifies directly to overall classification accuracy:

    Micro F1-Score=∑k=1KTPkN=Accuracy\text{Micro } F_1\text{-Score} = \frac{\sum_{k=1}^K TP_k}{N} = \text{Accuracy}

    Thus, Micro F1F_1-score evaluates performance at the instance level rather than the class level and shares all properties of standard accuracy, including vulnerability to class imbalance.

  7. Knowl 7 — Multi-Class Categorical Cross-Entropy and Non-Target Class Insensitivity

    equation

    For a dataset of NN observations where each instance ii has ground-truth class yi∈{1,…,K}y_i \in \{1, \dots, K\} and predicted class probabilities p(Y^i=k∣Xi)p(\hat{Y}_i = k \mid X_i) for k∈{1,…,K}k \in \{1, \dots, K\}, the average categorical cross-entropy is:

    H(p,p^)=−1N∑i=1N∑k=1Kp(Yi=k∣Xi)log⁡p(Y^i=k∣Xi)=−1N∑i=1Nlog⁡p(Y^i=yi∣Xi)H(p, \hat{p}) = -\frac{1}{N} \sum_{i=1}^N \sum_{k=1}^K p(Y_i = k \mid X_i) \log p(\hat{Y}_i = k \mid X_i) = -\frac{1}{N} \sum_{i=1}^N \log p(\hat{Y}_i = y_i \mid X_i)

    where p(Yi=k∣Xi)=1p(Y_i = k \mid X_i) = 1 if k=yik = y_i and 00 otherwise.

    A key property and limitation of cross-entropy is that it is detached from the discrete decision rule and the confusion matrix: it depends solely on the predicted probability assigned to the single true class yiy_i. Consequently, cross-entropy is indifferent to how the remaining probability mass is distributed across incorrect classes. Two classifiers assigning the identical probability p(Y^i=yi∣Xi)p(\hat{Y}_i = y_i \mid X_i) receive the exact same cross-entropy loss for instance ii, even if one classifier would correctly predict class yiy_i under an argmax decision rule (assigning lower probabilities to all other classes) while the other would misclassify the instance by assigning a higher probability to an incorrect class.

  8. Knowl 8 — Weighted Balanced Accuracy

    equation

    Weighted Balanced Accuracy incorporates class frequencies to adjust the per-class recall contributions according to the sample size of each class. For KK classes, let TPk=CkkTP_k = C_{kk} be the number of true positives for class kk, Totalrow,k=∑j=1KCkj\text{Total}_{\text{row}, k} = \sum_{j=1}^K C_{kj} be the total number of ground-truth instances of class kk, wkw_k be the frequency (weight) of class kk in the dataset, and W=∑k=1KwkW = \sum_{k=1}^K w_k be the sum of class weights. The metric is defined as:

    Balanced Accuracy Weighted=∑k=1KTPkTotalrow,k⋅wkK⋅W\text{Balanced Accuracy Weighted} = \frac{\sum_{k=1}^K \frac{TP_k}{\text{Total}_{\text{row}, k}} \cdot w_k}{K \cdot W}

    By weighting the recall of each class by its prevalence, this formulation prevents low-frequency classes from disproportionately distorting the evaluation while maintaining visibility into per-class performance during model development.

Coverage note — None was omitted; all core evaluation metrics for multi-class classification, their mathematical formulations, properties, and trade-offs presented in the paper were included.

References

  1. 1.Sabri Boughorbel, Fethi Jarray, and Mohammed El-Anbari. “Optimal classifier for imbalanced data using Matthews Correlation Coefficient metric”. In: PLOS ONE 12.6 (June 2017), pp. 1–17. DOI: 10.1371/journal.pone.0177678. URL: https://doi.org/10.1371/journal.pone.0177678.
  2. 2.K. H. Brodersen et al. “The Balanced Accuracy and Its Posterior Distribution”. In: 2010 20th International Conference on Pattern Recognition. 2010, pp. 3121–3124.
  3. 3.J. B. Brown. “Classifiers and their Metrics Quantified”. In: Molecular Informatics 37.1-2 (2018), p. 1700127. DOI: 10.1002/minf.201700127. eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/minf.201700127. URL: https://onlinelibrary.wiley.com/doi/abs/10.1002/minf.201700127.
  4. 4.Davide Chicco and Giuseppe Jurman. “The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation”. eng. In: BMC genomics 21.1 (Jan. 2020), pp. 6–6. ISSN: 1471-2164. DOI: 10.1186/s12864-019-6413-7. URL: https://doi.org/10.1186/s12864-019-6413-7.
  5. 5.S.C. CHOI. “DISCRIMINATION AND CLASSIFICATION: OVERVIEW”. In: Statistical Methods of Discrimination and Classification. Ed. by SUNG C. CHOI. Pergamon, 1986, pp. 173–177. ISBN: 978-0-08-034000-5. DOI: https://doi.org/10.1016/B978-0-08-034000-5.50005-8. URL: http://www.sciencedirect.com/science/article/pii/B9780080340005500058.
  6. 6.J. Gorodkin. “Comparing two K-category assignments by a K-category correlation coefficient”. In: Computational Biology and Chemistry 28.5 (2004), pp. 367–374. ISSN: 1476-9271. DOI: https://doi.org/10.1016/j.compbiolchem . 2004 . 09 . 006. URL: http : / / www . sciencedirect . com / science / article / pii /S1476927104000799.
  7. 7.B.W. Matthews. “Comparison of the predicted and observed secondary structure of T4 phage lysozyme”. In: Biochimica et Biophysica Acta (BBA) - Protein Structure 405.2 (1975), pp. 442–451. ISSN: 0005-2795. DOI: https : / / doi . org / 10 . 1016 / 0005 - 2795(75 ) 90109 - 9. URL: http : / / www . sciencedirect . com /science/article/pii/0005279575901099.
  8. 8.Nechushtan Moran. Ok I got it. . . Mar. 2020. URL: https://medium.com/@morannechushtan/ok-i-gotit-9445e36d6c95.
  9. 9.Juri Opitz and Sebastian Burst. Macro F1 and Macro F1. 2019. arXiv: 1911.03347 [cs.LG].
  10. 10.Priya Ranganathan, C. S. Pramesh, and Rakesh Aggarwal. “Common pitfalls in statistical analysis: Measures of agreement”. eng. In: Perspectives in clinical research 8.4 (2017). PCR-8-187[PII], pp. 187–191. ISSN: 2229-3485. DOI: 10.4103/picr.PICR_123_17. URL: https://doi.org/10.4103/picr.PICR_123_17.
  11. 11.Yutaka Sasaki et al. The truth of the f-measure. 2007. 2007.
  12. 12.Boaz Shmueli. Multi-Class Metrics Made Simple, Part II: the F1-score. July 2019. URL: https : / /towardsdatascience . com / multi - class - metrics - made - simple - part - ii - the - f1 - score -ebe8b2c2ca1.
  13. 13.Antonio J. Tallón-Ballesteros and José Riquelme. “Data Mining Methods Applied to a Digital Forensics Task for Supervised Machine Learning”. In: Studies in Computational Intelligence 555 (Jan. 2014), pp. 413–428. DOI: 10.1007/978-3-319-05885-6-17.

Citation

MLA
Grandini, M., et al. “Metrics for Multi-Class Classification: An Overview”. arXiv, 2020, http://arxiv.org/abs/2008.05756v1.
APA
Grandini, M., Bagli, E., & Visani, G. (2020). Metrics for Multi-Class Classification: an Overview. arXiv. http://arxiv.org/abs/2008.05756v1
Chicago
Grandini, M., E. Bagli, and G. Visani. 2020. “Metrics for Multi-Class Classification: An Overview”. arXiv. http://arxiv.org/abs/2008.05756v1.
Harvard
Grandini, M., Bagli, E. and Visani, G. (2020) “Metrics for Multi-Class Classification: an Overview”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2008.05756v1.
Vancouver
1. Grandini M, Bagli E, Visani G (2020) Metrics for Multi-Class Classification: an Overview. arXiv

BibTeX

@article{grandini2020metrics,
  title = {Metrics for Multi-Class Classification: an Overview},
  author = {Grandini, Margherita and Bagli, Enrico and Visani, Giorgio},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2008.05756v1},
  eprint = {2008.05756}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors