keyword
class distribution
Class distribution refers to the frequency, proportion, or relative representation of instances belonging to each target category or label within a dataset. In machine learning classification tasks, this distribution describes how data samples are partitioned across different classes, ranging from binary to multi-class scenarios. When each category contains approximately the same number of examples, the dataset exhibits a balanced class distribution, whereas large disparities between majority and minority categories create an imbalanced or skewed distribution. The nature of the class distribution significantly influences algorithm design, the necessity of data resampling techniques, and the choice of appropriate performance metrics, as standard accuracy measures can become misleading when classes are disproportionately represented.
3 items

Metrics for Multi-Class Classification: an Overview
Margherita Grandini, Enrico Bagli, Giorgio Visani
Why you should read this
Analyzes key multi-class classification metrics by detailing their individual strengths, limitations, and practical applications across model selection and hyperparameter tuning.
Classification tasks in machine learning involving more than two classes are known by the name of "multi-class classification". Performance indicators are very useful when the aim is to evaluate and compare different classification models or machine learning techniques. Many metrics come in handy to test the ability of a multi-class classifier. Those metrics turn out to be useful at different stage of the development process, e.g. comparing the performance of two different models or analysing the behaviour of the same model by tuning different parameters. In this white paper we review a list of the most promising multi-class metrics, we highlight their advantages and disadvantages and show their possible usages during the development of a classification model.
Added
2026-09-25

Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning
Guillaume Lemaitre, Fernando Nogueira, Christos K. Aridas
Why you should read this
Presents imbalanced-learn, a Python library integrated with scikit-learn that provides standardized implementations of over-sampling, under-sampling, and ensemble algorithms to effectively train machine learning models on skewed class distributions.
Imbalanced-learn is an open-source python toolbox aiming at providing a wide range of methods to cope with the problem of imbalanced dataset frequently encountered in machine learning and pattern recognition. The implemented state-of-the-art methods can be categorized into 4 groups: (i) under-sampling, (ii) over-sampling, (iii) combination of over- and under-sampling, and (iv) ensemble learning methods. The proposed toolbox only depends on numpy, scipy, and scikit-learn and is distributed under MIT license. Furthermore, it is fully compatible with scikit-learn and is part of the scikit-learn-contrib supported project. Documentation, unit tests as well as integration tests are provided to ease usage and contribution. The toolbox is publicly available in GitHub: this https URL.
Added
2026-09-14

The relationship between Precision-Recall and ROC curves
Jesse Davis, Mark Goadrich
Why you should read this
Demonstrates mathematically why Precision-Recall curves fundamentally outperform ROC curves for imbalanced datasets, establishing core principles of metric selection.
Receiver Operator Characteristic (ROC) curves are commonly used to present results for binary decision problems in machine learning. However, when dealing with highly skewed datasets, Precision-Recall (PR) curves give a more informative picture of an algorithm's performance. We show that a deep connection exists between ROC space and PR space, such that a curve dominates in ROC space if and only if it dominates in PR space. A corollary is the notion of an achievable PR curve, which has properties much like the convex hull in ROC space; we show an efficient algorithm for computing this curve. Finally, we also note differences in the two types of curves are significant for algorithm design. For example, in PR space it is incorrect to linearly interpolate between points. Furthermore, algorithms that optimize the area under the ROC curve are not guaranteed to optimize the area under the PR curve.
Added
2026-03-22
