Towards Sustainable Learning: Coresets for Data-efficient Deep Learning
Yu YangHao KangBaharan Mirzasoleiman
Develops CREST, a scalable coreset selection framework with theoretical convergence guarantees for non-convex optimization that models loss as piecewise quadratic sub-regions and filters learned examples to accelerate deep neural network training by up to 2.5x with minimal accuracy loss.
Training modern deep neural networks requires massive datasets and compute resources, incurring high financial costs, significant energy consumption, and substantial environmental footprints. While subset selection (coreset) techniques can accelerate training by identifying the most valuable training examples, existing methods have been limited to simpler, convex models. When applied to deep learning, conventional coreset algorithms fail because non-convex training dynamics shift rapidly over time, mini-batch stochastic gradient descent introduces high bias and variance, and re-selecting subsets from full datasets is computationally prohibitive.
The article introduces CREST, a scalable coreset selection framework designed to provide rigorous theoretical convergence guarantees and data-efficient training for deep neural networks. The objective of the article is to demonstrate how modeling non-convex loss functions and extracting mini-batch coresets can significantly reduce training time while maintaining high model accuracy.
The authors evaluated the framework through mathematical convergence analysis and empirical testing on image classification and natural language processing tasks. The experiments included training ResNet-20 on CIFAR-10, ResNet-18 on CIFAR-100, ResNet-50 on TinyImageNet, and fine-tuning RoBERTa on the 570,000-example Stanford Natural Language Inference dataset. CREST approximates the loss surface as piece-wise quadratic regions using gradient and curvature information, iteratively extracts mini-batch coresets from smaller random subsets, and permanently filters out examples once they are consistently learned.
The analysis yielded four major findings. First, CREST accelerated model training by 1.7x to 2.5x compared to training on full datasets, achieving the lowest relative error among all evaluated coreset baselines. Second, CREST scaled to large-scale natural language processing benchmarks where previous full-dataset coreset methods failed due to computational overhead. Third, the framework drastically reduced coreset selection overhead, cutting the required subset update frequency by 74% to 98% relative to greedy mini-batch selection without compromising accuracy. Fourth, behavioral analysis of selected data revealed that neural networks benefit most from curriculum-like learning: models prioritize easier examples early in training and transition toward more difficult examples later, while completely dropping learned examples without degrading final test accuracy.
These findings indicate that organizations training deep learning models can achieve substantial reductions in compute costs, operational timelines, and carbon emissions. By replacing heuristic data pruning with theoretically grounded mini-batch selection, teams can optimize training throughput without sacrificing final generalization performance.
Organizations seeking to optimize deep learning pipelines should consider piloting CREST on large datasets where training costs are substantial. Implementation should prioritize hyperparameter tuning for threshold tolerances and update intervals, paired with efficient data-loading pipelines to maximize real-world wall-clock speedups.
The primary limitation of the method is that its relative performance advantage over random sampling diminishes when the allowable training budget expands beyond constrained compute regimes. Furthermore, because empirical validation was conducted on standard benchmarks up to 570,000 examples, practitioners should validate scalability and data-loading efficiency on multi-billion-parameter models and enterprise-scale datasets.
- Paper: Adaptive Second Order Coresets for Data-efficient Machine Learning, Omead Pooladzandi et al. (2022). ADACORE establishes curvature-aware coresets and convergence guarantees that CREST adapts to deep-network training with mini-batch selection.
- Paper: Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt, Sören Mindermann et al. (2022). RHO-LOSS introduces online selection of valuable, not-yet-learned examples, a key precursor to CREST’s mini-batch filtering and learned-example removal.
- Paper: Active Learning for Convolutional Neural Networks: A Core-Set Approach, Ozan Sener et al. (2018). This core-set formulation grounds the subset-selection problem that CREST extends from geometric active learning to efficient deep-network training.
No sufficiently relevant recommendations were found.
