Built independently by an author, for readers. Read the story and support ChapterPal

keyword

risk minimization

Risk minimization is a fundamental principle in machine learning and statistical decision theory that aims to find a model or decision rule that minimizes expected loss, or statistical risk, across all potential data. In this framework, a loss function quantifies the error or penalty incurred when a model prediction deviates from the true outcome. Because the true underlying probability distribution of the data is typically unknown, learning algorithms frequently rely on empirical risk minimization, which approximates the expected risk by minimizing the average loss over an observed training dataset. To address computational challenges and prevent overfitting, practical approaches often optimize continuous surrogate loss functions and integrate regularization techniques, thereby ensuring that the resulting model generalizes effectively to unseen instances.

7 items

Cross-Entropy Loss Functions: Theoretical Analysis and Applications

Cross-Entropy Loss Functions: Theoretical Analysis and Applications

Anqi Mao, Mehryar Mohri, Yutao Zhong

OrganizationsGoogleNew York University

Why you should read this

Establishes the first tight non-asymptotic HH-consistency bounds for cross-entropy and general comp-sum loss functions, using these theoretical guarantees to develop new adversarial training objectives that improve defense against attacks without sacrificing standard accuracy.

Cross-entropy is a widely used loss function in applications. It coincides with the logistic loss applied to the outputs of a neural network, when the softmax is used. But, what guarantees can we rely on when using cross-entropy as a surrogate loss? We present a theoretical analysis of a broad family of loss functions, comp-sum losses, that includes cross-entropy (or logistic loss), generalized cross-entropy, the mean absolute error and other cross-entropy-like loss functions. We give the first HH-consistency bounds for these loss functions. These are non-asymptotic guarantees that upper bound the zero-one loss estimation error in terms of the estimation error of a surrogate loss, for the specific hypothesis set HH used. We further show that our bounds are tight. These bounds depend on quantities called minimizability gaps. To make them more explicit, we give a specific analysis of these gaps for comp-sum losses. We also introduce a new family of loss functions, smooth adversarial comp-sum losses, that are derived from their comp-sum counterparts by adding in a related smooth term. We show that these loss functions are beneficial in the adversarial setting by proving that they admit HH-consistency bounds. This leads to new adversarial robustness algorithms that consist of minimizing a regularized smooth adversarial comp-sum loss. While our main purpose is a theoretical analysis, we also present an extensive empirical analysis comparing comp-sum losses. We further report the results of a series of experiments demonstrating that our adversarial robustness algorithms outperform the current state-of-the-art, while also achieving a superior non-adversarial accuracy.

Added

2026-10-05

Optimal Strategies for Reject Option Classifiers

Optimal Strategies for Reject Option Classifiers

Vojtech Franc, Daniel Prusa, Václav Vorácek

OrganizationsCzech Technical University in Prague

Why you should read this

Unifies cost-based, bounded-improvement, and bounded-abstention selective classification models by proving they share the same optimal strategy, while developing two Fisher consistent algorithms to learn optimal rejection functions for arbitrary black-box classifiers across diverse prediction tasks.

In classification with a reject option, the classifier is allowed in uncertain cases to abstain from prediction. The classical cost-based model of a reject option classifier requires the rejection cost to be defined explicitly. The alternative bounded-improvement model and the bounded-abstention model avoid the notion of the reject cost. The bounded-improvement model seeks a classifier with a guaranteed selective risk and maximal cover. The bounded-abstention model seeks a classifier with guaranteed cover and minimal selective risk. We prove that despite their different formulations the three rejection models lead to the same prediction strategy: the Bayes classifier endowed with a randomized Bayes selection function. We define the notion of a proper uncertainty score as a scalar summary of the prediction uncertainty sufficient to construct the randomized Bayes selection function. We propose two algorithms to learn the proper uncertainty score from examples for an arbitrary black-box classifier. We prove that both algorithms provide Fisher consistent estimates of the proper uncertainty score and demonstrate their efficiency in different prediction problems, including classification, ordinal regression, and structured output classification.

Added

2026-09-26

Learning with Noisy Labels

Learning with Noisy Labels

Nagarajan Natarajan, I. Dhillon, Pradeep Ravikumar, Ambuj Tewari

OrganizationsUniversity of MichiganUniversity of Texas at Austin

Why you should read this

Establishes theoretical guarantees and practical surrogate loss modifications that enable standard classifiers like biased support vector machines and weighted logistic regression to learn effectively from class-conditional noisy labels.

In this paper, we theoretically study the problem of binary classification in the presence of random classification noise — the learner, instead of seeing the true labels, sees labels that have independently been flipped with some small probability. Moreover, random label noise is class-conditional — the flip probability depends on the class. We provide two approaches to suitably modify any given surrogate loss function. First, we provide a simple unbiased estimator of any loss, and obtain performance bounds for empirical risk minimization in the presence of iid data with noisy labels. If the loss function satisfies a simple symmetry condition, we show that the method leads to an efficient algorithm for empirical minimization. Second, by leveraging a reduction of risk minimization under noisy labels to classification with weighted 0-1 loss, we suggest the use of a simple weighted surrogate loss, for which we are able to obtain strong empirical risk bounds. This approach has a very remarkable consequence — methods used in practice such as biased SVM and weighted logistic regression are provably noise-tolerant. On a synthetic non-separable dataset, our methods achieve over 88% accuracy even when 40% of the labels are corrupted, and are competitive with respect to recently proposed methods for dealing with label noise in several benchmark datasets.

Added

2026-09-25