keyword
multiclass classification
Multiclass classification is a supervised machine learning task that involves assigning an input instance to exactly one of three or more distinct, mutually exclusive categories. Unlike binary classification, which distinguishes between only two outcomes, and multilabel classification, where an input can be associated with multiple labels simultaneously, multiclass problems restrict the final prediction to a single target class per instance. Models can address multiclass tasks directly using architectures optimized with generalized loss functions and categorical probability distributions, such as neural networks with softmax activation, or indirectly by decomposing the overall problem into multiple binary subproblems through strategies such as one-versus-all or one-versus-one.
9 items

Cross-Entropy Loss Functions: Theoretical Analysis and Applications
Anqi Mao, Mehryar Mohri, Yutao Zhong
Why you should read this
Establishes the first tight non-asymptotic -consistency bounds for cross-entropy and general comp-sum loss functions, using these theoretical guarantees to develop new adversarial training objectives that improve defense against attacks without sacrificing standard accuracy.
Cross-entropy is a widely used loss function in applications. It coincides with the logistic loss applied to the outputs of a neural network, when the softmax is used. But, what guarantees can we rely on when using cross-entropy as a surrogate loss? We present a theoretical analysis of a broad family of loss functions, comp-sum losses, that includes cross-entropy (or logistic loss), generalized cross-entropy, the mean absolute error and other cross-entropy-like loss functions. We give the first -consistency bounds for these loss functions. These are non-asymptotic guarantees that upper bound the zero-one loss estimation error in terms of the estimation error of a surrogate loss, for the specific hypothesis set used. We further show that our bounds are tight. These bounds depend on quantities called minimizability gaps. To make them more explicit, we give a specific analysis of these gaps for comp-sum losses. We also introduce a new family of loss functions, smooth adversarial comp-sum losses, that are derived from their comp-sum counterparts by adding in a related smooth term. We show that these loss functions are beneficial in the adversarial setting by proving that they admit -consistency bounds. This leads to new adversarial robustness algorithms that consist of minimizing a regularized smooth adversarial comp-sum loss. While our main purpose is a theoretical analysis, we also present an extensive empirical analysis comparing comp-sum losses. We further report the results of a series of experiments demonstrating that our adversarial robustness algorithms outperform the current state-of-the-art, while also achieving a superior non-adversarial accuracy.
Added
2026-10-05

Calibrated Learning to Defer with One-vs-All Classifiers
Rajeev Verma, Eric T. Nalisnick
Why you should read this
Develops a consistent one-vs-all surrogate loss for multiclass learning-to-defer systems that overcomes the uncalibrated, degenerate probability estimates of standard softmax formulations while matching or exceeding classification accuracy across diverse domains.
The learning to defer (L2D) framework has the potential to make AI systems safer. For a given input, the system can defer the decision to a human if the human is more likely than the model to take the correct action. We study the calibration of L2D systems, investigating if the probabilities they output are sound. We find that Mozannar & Sontag’s (2020) multiclass framework is not calibrated with respect to expert correctness. Moreover, it is not even guaranteed to produce valid probabilities due to its parameterization being degenerate for this purpose. We propose an L2D system based on one-vs-all classifiers that is able to produce calibrated probabilities of expert correctness. Furthermore, our loss function is also a consistent surrogate for multiclass L2D, like Mozannar & Sontag’s (2020). Our experiments verify that not only is our system calibrated, but this benefit comes at no cost to accuracy. Our model’s accuracy is always comparable (and often superior) to Mozannar & Sontag’s (2020) model’s in tasks ranging from hate speech detection to galaxy classification to diagnosis of skin lesions.
Added
2026-10-02

API design for machine learning software: experiences from the scikit-learn project
Lars Buitinck, Gilles Louppe, Mathieu Blondel, Fabian Pedregosa, Andreas Mueller, Olivier Grisel, Vlad Niculae, Peter Prettenhofer, Alexandre Gramfort, Jaques Grobler, Robert Layton, Jake Vanderplas, Arnaud Joly, Brian Holt, Gaël Varoquaux
Why you should read this
Presents the core API design principles behind scikit-learn, providing practical lessons on building consistent, composable, and user-friendly machine learning software.
Scikit-learn is an increasingly popular machine learning li- brary. Written in Python, it is designed to be simple and efficient, accessible to non-experts, and reusable in various contexts. In this paper, we present and discuss our design choices for the application programming interface (API) of the project. In particular, we describe the simple and elegant interface shared by all learning and processing units in the library and then discuss its advantages in terms of composition and reusability. The paper also comments on implementation details specific to the Python ecosystem and analyzes obstacles faced by users and developers of the library.
Added
2026-09-25

Support vector machine learning for interdependent and structured output spaces
Ioannis Tsochantaridis, Thomas Hofmann, Thorsten Joachims, Yasemin Altun
Why you should read this
Presents a maximum-margin framework that extends Support Vector Machines to complex, structured prediction tasks and solves the resulting exponential-sized optimization problem via an efficient cutting-plane algorithm.
Learning general functional dependencies is one of the main goals in machine learning. Recent progress in kernel-based methods has focused on designing flexible and powerful input representations. This paper addresses the complementary issue of problems involving complex outputs such as multiple dependent output variables and structured output spaces. We propose to generalize multiclass Support Vector Machine learning in a formulation that involves features extracted jointly from inputs and outputs. The resulting optimization problem is solved efficiently by a cutting plane algorithm that exploits the sparseness and structural decomposition of the problem. We demonstrate the versatility and effectiveness of our method on problems ranging from supervised grammar learning and named-entity recognition, to taxonomic text classification and sequence alignment.
Added
2026-09-25

Large Margin DAGs for Multiclass Classification
John C. Platt, N. Cristianini, J. Shawe-Taylor
Why you should read this
Introduces the Decision Directed Acyclic Graph architecture and the DAGSVM algorithm, providing dimension-independent generalization error bounds and substantially faster multiclass SVM evaluation by requiring only linear-time path traversals without sacrificing accuracy.
We present a new learning architecture: the Decision Directed Acyclic Graph (DDAG), which is used to combine many two-class classifiers into a multiclass classifier. For an N-class problem, the DDAG contains N(N − 1)/2 classifiers, one for each pair of classes. We present a VC analysis of the case when the node classifiers are hyperplanes; the resulting bound on the test error depends on N and on the margin achieved at the nodes, but not on the dimension of the space. This motivates an algorithm, DAGSVM, which operates in a kernel-induced feature space and uses two-class maximal margin hyperplanes at each decision-node of the DDAG. The DAGSVM is substantially faster to train and evaluate than either the standard algorithm or Max Wins, while maintaining comparable accuracy to both of these algorithms.
Added
2026-09-18

Multiple Kernel Learning Algorithms
Mehmet Gönen, Ethem Alpaydin
Why you should read this
Presents a comprehensive taxonomy and empirical comparison of multiple kernel learning algorithms across six key dimensions to guide the selection of combination methods based on computational complexity, solution sparsity, and kernel types.
In recent years, several methods have been proposed to combine multiple kernels instead of using a single one. These different kernels may correspond to using different notions of similarity or may be using information coming from multiple sources (different representations or different feature subsets). In trying to organize and highlight the similarities and differences between them, we give a taxonomy of and review several multiple kernel learning algorithms. We perform experiments on real data sets for better illustration and comparison of existing algorithms. We see that though there may not be large differences in terms of accuracy, there is difference between them in complexity as given by the number of stored support vectors, the sparsity of the solution as given by the number of used kernels, and training time complexity. We see that overall, using multiple kernels instead of a single one is useful and believe that combining kernels in a nonlinear or data-dependent way seems more promising than linear combination in fusing information provided by simple linear kernels, whereas linear methods are more reasonable when combining complex Gaussian kernels.
Added
2026-09-17

In Defense of One-Vs-All Classification
Ryan Rifkin, Aldebaro Klautau
Why you should read this
Demonstrates through rigorous empirical evaluations and theoretical analysis that simple one-vs-all multiclass strategies match the accuracy of far more complex methods when using well-tuned regularized binary classifiers.
We consider the problem of multiclass classification. Our main thesis is that a simple “one-vs-all” scheme is as accurate as any other approach, assuming that the underlying binary classifiers are well-tuned regularized classifiers such as support vector machines. This thesis is interesting in that it disagrees with a large body of recent published work on multiclass classification. We support our position by means of a critical review of the existing literature, a substantial collection of carefully controlled experimental work, and theoretical arguments.
Added
2026-09-16

A systematic study of the class imbalance problem in convolutional neural networks
Mateusz Buda, Atsuto Maki, Maciej A. Mazurowski
Why you should read this
Demonstrates that complete oversampling consistently outperforms undersampling and thresholding for class-imbalanced convolutional neural networks across standard vision benchmarks without causing the overfitting typical in classical machine learning.
In this study, we systematically investigate the impact of class imbalance on classification performance of convolutional neural networks (CNNs) and compare frequently used methods to address the issue. Class imbalance is a common problem that has been comprehensively studied in classical machine learning, yet very limited systematic research is available in the context of deep learning. In our study, we use three benchmark datasets of increasing complexity, MNIST, CIFAR-10 and ImageNet, to investigate the effects of imbalance on classification and perform an extensive comparison of several methods to address the issue: oversampling, undersampling, two-phase training, and thresholding that compensates for prior class probabilities. Our main evaluation metric is area under the receiver operating characteristic curve (ROC AUC) adjusted to multi-class tasks since overall accuracy metric is associated with notable difficulties in the context of imbalanced data. Based on results from our experiments we conclude that (i) the effect of class imbalance on classification performance is detrimental; (ii) the method of addressing class imbalance that emerged as dominant in almost all analyzed scenarios was oversampling; (iii) oversampling should be applied to the level that completely eliminates the imbalance, whereas the optimal undersampling ratio depends on the extent of imbalance; (iv) as opposed to some classical machine learning models, oversampling does not cause overfitting of CNNs; (v) thresholding should be applied to compensate for prior class probabilities when overall number of properly classified cases is of interest.
Added
2026-09-12

One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context
Skanda Athreya, Yutong Wang
We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifiers in the binary setting to the multiclass case. By leveraging the simplex encoding, we show that one-layer transformers with an argmax classification head behave identically to a one-nearest-neighbor classifier in the multiclass setting. This closes a gap left by prior work, whose multiclass result relied on a non-standard rounding-based approach rather than the typical argmax head used in practice.
Added
2026-09-02
