Built independently by an author, for readers. Read the story and support ChapterPal

keyword

loss approximation

Loss approximation is the process of estimating or replacing a complex, computationally expensive, or mathematically intractable loss function with a simpler surrogate model that is easier to optimize or analyze. In machine learning and mathematical optimization, complex loss landscapes, such as non-convex objectives, are frequently approximated locally using simpler mathematical formulations like quadratic functions or Taylor expansions. This technique allows optimization algorithms to efficiently compute parameter updates, derive theoretical convergence guarantees, evaluate error bounds, and select representative data subsets while preserving the essential geometric properties of the underlying learning objective.

1 item

Towards Sustainable Learning: Coresets for Data-efficient Deep Learning

Towards Sustainable Learning: Coresets for Data-efficient Deep Learning

Yu Yang, Hao Kang, Baharan Mirzasoleiman

Why you should read this

Develops CREST, a scalable coreset selection framework with theoretical convergence guarantees for non-convex optimization that models loss as piecewise quadratic sub-regions and filters learned examples to accelerate deep neural network training by up to 2.5x with minimal accuracy loss.

To improve the efficiency and sustainability of learning deep models, we propose CREST, the first scalable framework with rigorous theoretical guarantees to identify the most valuable examples for training non-convex models, particularly deep networks. To guarantee convergence to a stationary point of a non-convex function, CREST models the non-convex loss as a series of quadratic functions and extracts a coreset for each quadratic sub-region. In addition, to ensure faster convergence of stochastic gradient methods such as (mini-batch) SGD, CREST iteratively extracts multiple mini-batch coresets from larger random subsets of training data, to ensure nearly-unbiased gradients with small variances. Finally, to further improve scalability and efficiency, CREST identifies and excludes the examples that are learned from the coreset selection pipeline. Our extensive experiments on several deep networks trained on vision and NLP datasets, including CIFAR-10, CIFAR-100, TinyImageNet, and SNLI, confirm that CREST speeds up training deep networks on very large datasets, by 1.7x to 2.5x with minimum loss in the performance. By analyzing the learning difficulty of the subsets selected by CREST, we show that deep models benefit the most by learning from subsets of increasing difficulty levels 1.

Added

2026-10-03