Cross-Entropy Loss Functions: Theoretical Analysis and Applications
Anqi MaoMehryar MohriYutao Zhong
Establishes the first tight non-asymptotic -consistency bounds for cross-entropy and general comp-sum loss functions, using these theoretical guarantees to develop new adversarial training objectives that improve defense against attacks without sacrificing standard accuracy.
Modern machine learning systems rely heavily on classification models trained with surrogate objectives, primarily the cross-entropy (multinomial logistic) loss, because directly minimizing classification errors is computationally intractable. While cross-entropy is known to be statistically consistent in an asymptotic, unconstrained setting, this property offers no concrete performance guarantees when using real-world, restricted model families such as deep neural networks on finite samples. Furthermore, standard neural networks remain critically vulnerable to small, imperceptible adversarial input perturbations. The article addresses these core challenges by establishing rigorous, non-asymptotic theoretical guarantees for cross-entropy and related loss functions, and by deriving theoretically grounded algorithms for adversarial defense.
The main objective of the article is to provide the first tight, model-specific error bounds—known as hypothesis-set consistency bounds—for a broad class of composed loss functions (including standard cross-entropy, generalized cross-entropy, and mean absolute error), and to demonstrate how these principles can be extended to design superior, adversarially robust learning algorithms.
To achieve this, the article analyzes a unified mathematical family called composite-sum (comp-sum) losses parameterized by a scalar value that captures various standard classification objectives. The authors derive exact non-asymptotic bounds that connect surrogate training errors directly to actual classification errors across realistic, symmetric, and complete hypothesis sets, without imposing restrictive distribution assumptions. They analyze key structural quantities called minimizability gaps that capture approximation quality. Additionally, the authors formulate a new class of smooth adversarial comp-sum losses with matching theoretical guarantees and evaluate both standard and robust models empirically across image benchmark datasets (CIFAR-10, CIFAR-100, and SVHN) using various WideResNet and ResNet architectures against established attacks.
The findings reveal several crucial theoretical and practical insights. First, the article proves tight hypothesis-set consistency bounds for cross-entropy and related functions, showing that cross-entropy achieves a square-root error-scaling relationship with classification error. Second, while objectives like the mean absolute error possess a theoretically linear error rate, their bounds degrade with the total number of classes and face severe optimization difficulties in practice, explaining why cross-entropy strikes the best operational balance. Third, the newly introduced defense algorithm, ADV-COMP-SUM, consistently outperforms the current state-of-the-art adversarial defense benchmark (TRADES) across all tested architectures and datasets. Specifically, ADV-COMP-SUM achieves up to a 1.45% increase in robust accuracy under standard margin attacks and up to 1.12% higher robust accuracy under AutoAttack, while simultaneously improving clean (non-adversarial) classification accuracy by up to 2.54%.
These results provide a solid theoretical justification for the long-standing empirical dominance of cross-entropy in standard machine learning workflows. Crucially, in adversarial defense, where improving robustness has traditionally required sacrificing clean test accuracy, the findings demonstrate that smooth adversarial comp-sum losses eliminate this trade-off. This enhances model reliability and reduces deployment risk in safety-critical applications without compromising regular predictive performance.
Organizations developing machine learning models should continue prioritizing cross-entropy-style objectives for multi-class classification and adopt smooth adversarial comp-sum loss formulations for adversarial training pipelines. When implementing these robust algorithms, practitioners can tune key regularization and margin hyperparameters through standard cross-validation to maximize accuracy. Future work should focus on extending these theoretical bounds to incomplete hypothesis sets, exploring noisy label environments, and addressing general neural network generalization challenges under adversarial attacks.
Confidence in these findings is high due to the mathematical proofs of tightness and rigorous empirical comparisons matching benchmark protocols without data augmentation artifacts. However, users should note that the primary non-adversarial theoretical guarantees assume complete hypothesis sets whose generated output scores span the real space, and the empirical validations were conducted primarily on standard vision benchmark datasets.
- Paper: Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels, Zhilu Zhang et al. (2018). Its Lq framework connects cross-entropy to mean absolute error, giving useful groundwork for the source’s comparison of these losses and their theoretical guarantees.
- Paper: Theoretically Principled Trade-off between Robustness and Accuracy, Hongyang Zhang et al. (2019). TRADES supplies the adversarial-training objective and benchmark that the source’s smooth comp-sum defense develops beyond.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, Aleksander Madry et al. (2017). Its robust-optimization formulation establishes the adversarial-training setup that the source uses when deriving robust loss guarantees.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). Its account of gradient-based adversarial examples and adversarial training provides the basic threat and defense context for the source’s robust-loss analysis.
No sufficiently relevant recommendations were found.
