Adaptive Second Order Coresets for Data-efficient Machine Learning
Omead PooladzandiDavid DaviniBaharan Mirzasoleiman
Proposes ADACORE, a data-selection method that dynamically approximates loss curvature through exponentially averaged Hessian estimates to construct weighted training subsets with provable convergence guarantees and over 2.9x training speedups across convex and deep learning models.
Training modern machine learning models on massive datasets incurs substantial computational, financial, and environmental costs. While training on smaller subsets of data (coresets) can reduce these overheads, existing data-selection methods lack theoretical convergence guarantees, rely on expensive proxy models, or perform poorly on complex architectures because they fail to capture the loss landscape's geometry.
The article demonstrates ADACORE (Adaptive Second-order Coresets), a data-selection method that leverages loss landscape curvature to extract small, high-quality subsets for efficient model training with proven convergence guarantees.
To scale efficiently without the prohibitive cost of calculating full Hessian second-derivative matrices, the approach approximates the loss curvature using Hessian-free methods and Hutchinson’s diagonal estimation, combined with exponential moving averages to smooth gradient noise. The subset selection is framed as a submodular facility location problem solved via an efficient greedy algorithm. The researchers proved theoretical convergence rates for both convex and non-convex overparameterized models, and evaluated performance across image classification benchmarks (MNIST, CIFAR-10, CIFAR-100, BDD100k) using various neural network architectures and optimization algorithms.
Key findings show that ADACORE achieves substantial computational acceleration, delivering over 2.9-fold speedups compared to training on full datasets and up to 4.5-fold speedups over random data selection. When training neural networks on 1% subsets, it outperformed leading coreset baselines by 6% to 16.8% in test accuracy. Additionally, ADACORE identified more diverse and informative data points, automatically prioritizing uncertain and forgettable examples while discarding redundant samples, which prevented catastrophic forgetting when subset updates were spaced across multiple training epochs.
These results demonstrate that incorporating second-order curvature information significantly improves data-efficient training pipelines. For organizations managing large-scale machine learning workloads, this approach reduces hardware resource consumption, shortens iteration cycles, and lowers energy usage and carbon emissions without sacrificing generalization performance.
Organizations training resource-intensive vision models should consider piloting ADACORE within their existing training workflows to reduce infrastructure compute costs. When deploying the method, teams should favor moderate mini-batch sizes for Hessian estimation and test periodic subset re-selection to balance coreset calculation overhead against optimization speed.
Confidence in these findings is strong across the evaluated standard benchmark vision datasets and residual neural network architectures. However, decision-makers should note that the empirical evaluations focused primarily on image classification tasks, meaning performance across other modalities (such as large language models or tabular data) will require further empirical validation.
- Paper: Active Learning for Convolutional Neural Networks: A Core-Set Approach, Ozan Sener et al. (2018). It establishes the foundational geometric core-set and submodular selection framework for deep neural networks that ADACORE adapts and improves using second-order curvature.
- Paper: Optimizing Neural Networks with Kronecker-factored Approximate Curvature, James Martens et al. (2015). It provides essential background on practical, scalable approximations of second-order curvature and natural gradient statistics in deep learning optimization.
- Paper: Learning to Reweight Examples for Robust Deep Learning, Mengye Ren et al. (2018). It introduces dynamic gradient-based sample evaluation and reweighting to optimize data efficiency and robustness during neural network training.
- Paper: Improving Generalization with Active Learning, David Cohn et al. (1994). It establishes classic active learning principles of selectively querying uncertain and informative samples to maximize training efficiency.
- Paper: The Tradeoffs of Large Scale Learning, Léon Bottou et al. (2007). It provides the foundational theoretical trade-offs between computation time, sample size, and approximate optimization in large-scale machine learning.
- Paper: Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation, Jiawei Du et al. (2023). It extends the concept of loss landscape geometry and trajectory matching to dataset distillation, refining synthetic data generation via flatter loss curvature.
- Paper: Adam-mini: Use Fewer Learning Rates To Gain More, Yushun Zhang et al. (2025). It leverages structured second-order Hessian approximations to drastically reduce optimizer memory overhead across large neural architectures.
