keyword
multiclass learning
Multiclass learning is a supervised machine learning paradigm in which a model is trained to assign each input instance to exactly one category from a set of three or more mutually exclusive classes. Unlike binary classification, which distinguishes between only two outcomes, or multilabel learning, where multiple categories can apply simultaneously, multiclass learning requires selecting a single correct label from a discrete target set. Practitioners typically address multiclass tasks either through decomposition strategies or through direct optimization methods. Decomposition approaches divide the multiclass problem into a collection of simpler binary classification tasks using schemes such as one-versus-all, one-versus-one, or error-correcting output codes, while direct methods generalize decision boundaries and objective functions within algorithms like support vector machines, neural networks, and decision trees to evaluate all candidate classes simultaneously.
4 items

A Brief Introduction to Boosting
R. Schapire
Why you should read this
Explains the theoretical foundations and mechanics of AdaBoost, providing clear proofs of exponential training error reduction alongside a margin-based explanation for why boosting resists overfitting even after achieving zero training error.
Boosting is a general method for improving the accuracy of any given learning algorithm. This short paper introduces the boosting algorithm AdaBoost, and explains the underlying theory of boosting, including an explanation of why boosting often does not suffer from overfitting. Some examples of recent applications of boosting are also described.
Added
2026-09-24

Reducing Multiclass to Binary: A Unifying Approach for Margin Classifiers
Erin L. Allwein, Rob Schapire, Y. Singer
Why you should read this
Unifies standard multiclass-to-binary reductions under a single framework by introducing margin- and loss-based decoding techniques backed by rigorous error bounds for algorithms like AdaBoost and support vector machines.
We present a unifying framework for studying the solution of multiclass categorization problems by reducing them to multiple binary problems that are then solved using a margin-based binary learning algorithm. The proposed framework unifies some of the most popular approaches in which each class is compared against all others, or in which all pairs of classes are compared to each other, or in which output codes with error-correcting properties are used. We propose a general method for combining the classifiers generated on the binary problems, and we prove a general empirical multiclass loss bound given the empirical loss of the individual binary learning algorithms. The scheme and the corresponding bounds apply to many popular classification learning algorithms including support-vector machines, AdaBoost, regression, logistic regression and decision-tree algorithms. We also give a multiclass generalization error analysis for general output codes with AdaBoost as the binary learner. Experimental results with SVM and AdaBoost show that our scheme provides a viable alternative to the most commonly used multiclass algorithms.
Added
2026-09-24

On the Algorithmic Implementation of Multiclass Kernel-based Vector Machines
Koby Crammer, Yoram Singer
Why you should read this
Develops a direct multiclass support vector machine framework based on a generalized margin that reduces large-scale quadratic optimization into small, single-example subproblems solved via a provably convergent fixed-point algorithm.
In this paper we describe the algorithmic implementation of multiclass kernel-based vector machines. Our starting point is a generalized notion of the margin to multiclass problems. Using this notion we cast multiclass categorization problems as a constrained optimization problem with a quadratic objective function. Unlike most of previous approaches which typically decompose a multiclass problem into multiple independent binary classification tasks, our notion of margin yields a direct method for training multiclass predictors. By using the dual of the optimization problem we are able to incorporate kernels with a compact set of constraints and decompose the dual problem into multiple optimization problems of reduced size. We describe an efficient fixed-point algorithm for solving the reduced optimization problems and prove its convergence. We then discuss technical details that yield significant running time improvements for large datasets. Finally, we describe various experiments with our approach comparing it to previously studied kernel-based methods. Our experiments indicate that for multiclass problems we attain state-of-the-art accuracy.
Added
2026-09-15

Boosting the margin: A new explanation for the effectiveness of voting methods
Robert E. Schapire, Yoav Freund, Peter Barlett, Wee Sun Lee
Why you should read this
Explains why boosting continues to improve generalization error even after reaching zero training error by proving theoretical bounds based on the classification margin distribution rather than ensemble complexity.
One of the surprising recurring phenomena observed in experiments with boosting is that the test error of the generated classifier usually does not increase as its size becomes very large, and often is observed to decrease even after the training error reaches zero. In this paper, we show that this phenomenon is related to the distribution of margins of the training examples with respect to the generated voting classification rule, where the margin of an example is simply the difference between the number of correct votes and the maximum number of votes received by any incorrect label. We show that techniques used in the analysis of Vapnik’s support vector classifiers and of neural networks with small weights can be applied to voting methods to relate the margin distribution to the test error. We also show theoretically and experimentally that boosting is especially effective at increasing the margins of the training examples. Finally, we compare our explanation to those based on the bias-variance decomposition.
Added
2026-09-12
