Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multiclass classification

Multiclass classification is a supervised machine learning task that involves assigning an input instance to exactly one of three or more distinct, mutually exclusive categories. Unlike binary classification, which distinguishes between only two outcomes, and multilabel classification, where an input can be associated with multiple labels simultaneously, multiclass problems restrict the final prediction to a single target class per instance. Models can address multiclass tasks directly using architectures optimized with generalized loss functions and categorical probability distributions, such as neural networks with softmax activation, or indirectly by decomposing the overall problem into multiple binary subproblems through strategies such as one-versus-all or one-versus-one.

9 items

Cross-Entropy Loss Functions: Theoretical Analysis and Applications

Cross-Entropy Loss Functions: Theoretical Analysis and Applications

Anqi Mao, Mehryar Mohri, Yutao Zhong

OrganizationsGoogleNew York University

Why you should read this

Establishes the first tight non-asymptotic HH-consistency bounds for cross-entropy and general comp-sum loss functions, using these theoretical guarantees to develop new adversarial training objectives that improve defense against attacks without sacrificing standard accuracy.

Cross-entropy is a widely used loss function in applications. It coincides with the logistic loss applied to the outputs of a neural network, when the softmax is used. But, what guarantees can we rely on when using cross-entropy as a surrogate loss? We present a theoretical analysis of a broad family of loss functions, comp-sum losses, that includes cross-entropy (or logistic loss), generalized cross-entropy, the mean absolute error and other cross-entropy-like loss functions. We give the first HH-consistency bounds for these loss functions. These are non-asymptotic guarantees that upper bound the zero-one loss estimation error in terms of the estimation error of a surrogate loss, for the specific hypothesis set HH used. We further show that our bounds are tight. These bounds depend on quantities called minimizability gaps. To make them more explicit, we give a specific analysis of these gaps for comp-sum losses. We also introduce a new family of loss functions, smooth adversarial comp-sum losses, that are derived from their comp-sum counterparts by adding in a related smooth term. We show that these loss functions are beneficial in the adversarial setting by proving that they admit HH-consistency bounds. This leads to new adversarial robustness algorithms that consist of minimizing a regularized smooth adversarial comp-sum loss. While our main purpose is a theoretical analysis, we also present an extensive empirical analysis comparing comp-sum losses. We further report the results of a series of experiments demonstrating that our adversarial robustness algorithms outperform the current state-of-the-art, while also achieving a superior non-adversarial accuracy.

Added

2026-10-05

Calibrated Learning to Defer with One-vs-All Classifiers

Calibrated Learning to Defer with One-vs-All Classifiers

Rajeev Verma, Eric T. Nalisnick

Why you should read this

Develops a consistent one-vs-all surrogate loss for multiclass learning-to-defer systems that overcomes the uncalibrated, degenerate probability estimates of standard softmax formulations while matching or exceeding classification accuracy across diverse domains.

The learning to defer (L2D) framework has the potential to make AI systems safer. For a given input, the system can defer the decision to a human if the human is more likely than the model to take the correct action. We study the calibration of L2D systems, investigating if the probabilities they output are sound. We find that Mozannar & Sontag’s (2020) multiclass framework is not calibrated with respect to expert correctness. Moreover, it is not even guaranteed to produce valid probabilities due to its parameterization being degenerate for this purpose. We propose an L2D system based on one-vs-all classifiers that is able to produce calibrated probabilities of expert correctness. Furthermore, our loss function is also a consistent surrogate for multiclass L2D, like Mozannar & Sontag’s (2020). Our experiments verify that not only is our system calibrated, but this benefit comes at no cost to accuracy. Our model’s accuracy is always comparable (and often superior) to Mozannar & Sontag’s (2020) model’s in tasks ranging from hate speech detection to galaxy classification to diagnosis of skin lesions.

Added

2026-10-02

Multiple Kernel Learning Algorithms

Multiple Kernel Learning Algorithms

Mehmet Gönen, Ethem Alpaydin

OrganizationsBoğaziçi University

Why you should read this

Presents a comprehensive taxonomy and empirical comparison of multiple kernel learning algorithms across six key dimensions to guide the selection of combination methods based on computational complexity, solution sparsity, and kernel types.

In recent years, several methods have been proposed to combine multiple kernels instead of using a single one. These different kernels may correspond to using different notions of similarity or may be using information coming from multiple sources (different representations or different feature subsets). In trying to organize and highlight the similarities and differences between them, we give a taxonomy of and review several multiple kernel learning algorithms. We perform experiments on real data sets for better illustration and comparison of existing algorithms. We see that though there may not be large differences in terms of accuracy, there is difference between them in complexity as given by the number of stored support vectors, the sparsity of the solution as given by the number of used kernels, and training time complexity. We see that overall, using multiple kernels instead of a single one is useful and believe that combining kernels in a nonlinear or data-dependent way seems more promising than linear combination in fusing information provided by simple linear kernels, whereas linear methods are more reasonable when combining complex Gaussian kernels.

Added

2026-09-17

A systematic study of the class imbalance problem in convolutional neural networks

A systematic study of the class imbalance problem in convolutional neural networks

Mateusz Buda, Atsuto Maki, Maciej A. Mazurowski

OrganizationsDuke UniversityKTH Royal Institute of Technology

Why you should read this

Demonstrates that complete oversampling consistently outperforms undersampling and thresholding for class-imbalanced convolutional neural networks across standard vision benchmarks without causing the overfitting typical in classical machine learning.

In this study, we systematically investigate the impact of class imbalance on classification performance of convolutional neural networks (CNNs) and compare frequently used methods to address the issue. Class imbalance is a common problem that has been comprehensively studied in classical machine learning, yet very limited systematic research is available in the context of deep learning. In our study, we use three benchmark datasets of increasing complexity, MNIST, CIFAR-10 and ImageNet, to investigate the effects of imbalance on classification and perform an extensive comparison of several methods to address the issue: oversampling, undersampling, two-phase training, and thresholding that compensates for prior class probabilities. Our main evaluation metric is area under the receiver operating characteristic curve (ROC AUC) adjusted to multi-class tasks since overall accuracy metric is associated with notable difficulties in the context of imbalanced data. Based on results from our experiments we conclude that (i) the effect of class imbalance on classification performance is detrimental; (ii) the method of addressing class imbalance that emerged as dominant in almost all analyzed scenarios was oversampling; (iii) oversampling should be applied to the level that completely eliminates the imbalance, whereas the optimal undersampling ratio depends on the extent of imbalance; (iv) as opposed to some classical machine learning models, oversampling does not cause overfitting of CNNs; (v) thresholding should be applied to compensate for prior class probabilities when overall number of properly classified cases is of interest.

Added

2026-09-12