keyword
classification noise
Classification noise is a type of data corruption in machine learning where the assigned class labels of training examples are incorrectly recorded or flipped, even though their underlying feature values remain accurate. This imperfection frequently arises from human annotation mistakes, subjective judgments, sensor faults, or imperfect labeling pipelines. Unlike attribute noise, which alters the input measurements, classification noise directly distorts the target supervision signal, leading models to learn incorrect decision boundaries and suffer reduced predictive accuracy. Depending on the underlying process, label corruption can occur uniformly at random or conditionally based on the true class category. Addressing classification noise typically requires robust learning algorithms, symmetric or adjusted loss functions, and data-filtering techniques designed to prevent models from overfitting to mislabeled data points.
4 items

Learning with Noisy Labels
Nagarajan Natarajan, I. Dhillon, Pradeep Ravikumar, Ambuj Tewari
Why you should read this
Establishes theoretical guarantees and practical surrogate loss modifications that enable standard classifiers like biased support vector machines and weighted logistic regression to learn effectively from class-conditional noisy labels.
In this paper, we theoretically study the problem of binary classification in the presence of random classification noise — the learner, instead of seeing the true labels, sees labels that have independently been flipped with some small probability. Moreover, random label noise is class-conditional — the flip probability depends on the class. We provide two approaches to suitably modify any given surrogate loss function. First, we provide a simple unbiased estimator of any loss, and obtain performance bounds for empirical risk minimization in the presence of iid data with noisy labels. If the loss function satisfies a simple symmetry condition, we show that the method leads to an efficient algorithm for empirical minimization. Second, by leveraging a reduction of risk minimization under noisy labels to classification with weighted 0-1 loss, we suggest the use of a simple weighted surrogate loss, for which we are able to obtain strong empirical risk bounds. This approach has a very remarkable consequence — methods used in practice such as biased SVM and weighted logistic regression are provably noise-tolerant. On a synthetic non-separable dataset, our methods achieve over 88% accuracy even when 40% of the labels are corrupted, and are competitive with respect to recently proposed methods for dealing with label noise in several benchmark datasets.
Added
2026-09-25

A streaming ensemble algorithm (SEA) for large-scale classification
W. Street, YongSeog Kim
Why you should read this
Proposes a fast, constant-memory streaming ensemble algorithm that processes continuous data chunks and uses a targeted replacement heuristic to match batch classifier accuracy while rapidly adapting to concept drift.
Ensemble methods have recently garnered a great deal of attention in the machine learning community. Techniques such as Boosting and Bagging have proven to be highly effective but require repeated resampling of the training data, making them inappropriate in a data mining context. The methods presented in this paper take advantage of plentiful data, building separate classifiers on sequential chunks of training points. These classifiers are combined into a fixed-size ensemble using a heuristic replacement strategy. The result is a fast algorithm for large-scale or streaming data that classifies as well as a single decision tree built on all the data, requires approximately constant memory, and adjusts quickly to concept drift.
Added
2026-09-25

Learning in the Presence of Concept Drift and Hidden Contexts
G. Widmer, M. Kubát
Why you should read this
Presents the FLORA framework of incremental learning algorithms that dynamically adjust sample windows and reuse past concept descriptions to handle recurring hidden contexts and concept drift in continuous data streams.
On-line learning in domains where the target concept depends on some hidden context poses serious problems. A changing context can induce changes in the target concepts, producing what is known as concept drift. We describe a family of learning algorithms that flexibly react to concept drift and can take advantage of situations where contexts reappear. The general approach underlying all these algorithms consists of (1) keeping only a window of currently trusted examples and hypotheses; (2) storing concept descriptions and re-using them when a previous context re-appears; and (3) controlling both of these functions by a heuristic that constantly monitors the system's behavior. The paper reports on experiments that test the systems' performance under various conditions such as different levels of noise and different extent and rate of concept drift.
Added
2026-09-24

An Experimental Comparison of Three Methods for Constructing Ensembles of Decision Trees: Bagging, Boosting, and Randomization
Thomas G. Dietterich
Why you should read this
Demonstrates across 33 benchmark datasets that while boosting generates the most accurate decision tree ensembles on clean data, bagging remains superior under classification noise because boosting assigns excessive weight to mislabeled training instances.
Bagging and boosting are methods that generate a diverse ensemble of classifiers by manipulating the training data given to a "base" learning algorithm. Breiman has pointed out that they rely for their effectiveness on the instability of the base learning algorithm. An alternative approach to generating an ensemble is to randomize the internal decisions made by the base algorithm. This general approach has been studied previously by Ali and Pazzani and by Dietterich and Kong. This paper compares the effectiveness of randomization, bagging, and boosting for improving the performance of the decision-tree algorithm C4.5. The experiments show that in situations with little or no classification noise, randomization is competitive with (and perhaps slightly superior to) bagging but not as accurate as boosting. In situations with substantial classification noise, bagging is much better than boosting, and sometimes better than randomization.
Added
2026-09-12
