Built independently by an author, for readers. Read the story and support ChapterPal

keyword

coreset selection

Coreset selection is the process of identifying and extracting a small, representative, and often weighted subset of a larger dataset such that a model trained on that subset achieves optimization and generalization performance comparable to training on the complete dataset. By selecting the most informative data points based on criteria such as geometric representations, gradient approximations, loss curvature, or statistical importance, this technique significantly reduces computational overhead, storage demands, and training time. It serves as a foundational approach for data-efficient machine learning, enabling scalable training and evaluation for complex models while preserving theoretical convergence guarantees and predictive accuracy.

4 items

Towards Sustainable Learning: Coresets for Data-efficient Deep Learning

Towards Sustainable Learning: Coresets for Data-efficient Deep Learning

Yu Yang, Hao Kang, Baharan Mirzasoleiman

Why you should read this

Develops CREST, a scalable coreset selection framework with theoretical convergence guarantees for non-convex optimization that models loss as piecewise quadratic sub-regions and filters learned examples to accelerate deep neural network training by up to 2.5x with minimal accuracy loss.

To improve the efficiency and sustainability of learning deep models, we propose CREST, the first scalable framework with rigorous theoretical guarantees to identify the most valuable examples for training non-convex models, particularly deep networks. To guarantee convergence to a stationary point of a non-convex function, CREST models the non-convex loss as a series of quadratic functions and extracts a coreset for each quadratic sub-region. In addition, to ensure faster convergence of stochastic gradient methods such as (mini-batch) SGD, CREST iteratively extracts multiple mini-batch coresets from larger random subsets of training data, to ensure nearly-unbiased gradients with small variances. Finally, to further improve scalability and efficiency, CREST identifies and excludes the examples that are learned from the coreset selection pipeline. Our extensive experiments on several deep networks trained on vision and NLP datasets, including CIFAR-10, CIFAR-100, TinyImageNet, and SNLI, confirm that CREST speeds up training deep networks on very large datasets, by 1.7x to 2.5x with minimum loss in the performance. By analyzing the learning difficulty of the subsets selected by CREST, we show that deep models benefit the most by learning from subsets of increasing difficulty levels 1.

Added

2026-10-03

UniPROT: Uniform Prototype Selection via Partial Optimal Transport with Submodular Guarantees

UniPROT: Uniform Prototype Selection via Partial Optimal Transport with Submodular Guarantees

Prateek Chanda, Prayas Agrawal, Karthik S. Gurumoorthy, Ganesh Ramakrishnan, Bamdev Mishra, Pratik Jawanpuria

OrganizationsIndian Institute of Technology BombayMicrosoftWalmart Labs

Why you should read this

Introduces UniPROT, a subset selection framework that reformulates optimal transport into a submodular objective with a (1 - 1/e) greedy approximation guarantee, effectively preserving minority-class representation in imbalanced classification and language model training.

Selecting prototypical examples from a source distribution to represent a target data distribution is a fundamental problem in machine learning. Existing subset selection methods often rely on implicit importance scores, which can be skewed towards majority classes and lead to low-quality prototypes for minority classes. We present \methodprop\methodprop, a novel subset selection framework that minimizes the optimal transport (OT) distance between a uniformly weighted prototypical distribution and the target distribution. While intuitive, this formulation leads to a cardinality-constrained maximization of a \emph{super-additive} objective, which is generally intractable to approximate efficiently. To address this, we propose a principled reformulation of the OT marginal constraints, yielding a partial optimal transport-based submodular objective. We prove that this reformulation enables a greedy algorithm with a (1−1/e)(1-1/e) approximation guarantee relative to the original super-additive maximization problem. Empirically, we showcase that enforcing uniform prototype weights in UniPROT consistently improves minority-class representation in imbalanced classification benchmarks without compromising majority-class accuracy. In both finetuning and pretraining regimes for large language models under domain imbalance, UniPROT enforces uniform source contributions, yielding robust performance gains. Our results establish UniPROT as a scalable, theoretically grounded solution for uniform-weighted prototype selection. Our code is publicly available at GitHub\footnote{Code: this https URL}

Added

2026-09-29

From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Ming Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, Jing Xiao

OrganizationsPing An Technology (Shenzhen) Co., Ltd.University of Maryland

Why you should read this

Proposes the Instruction-Following Difficulty metric to let large language models autonomously select high-impact training samples, outperforming full-dataset fine-tuning on Alpaca and WizardLM while using only ten percent of the data.

In the realm of Large Language Models (LLMs), the balance between instruction data quality and quantity is a focal point. Recognizing this, we introduce a self-guided methodology for LLMs to autonomously discern and select cherry samples from open-source datasets, effectively minimizing manual curation and potential cost for instruction tuning an LLM. Our key innovation, the Instruction-Following Difficulty (IFD) metric, emerges as a pivotal metric to identify discrepancies between a model's expected responses and its intrinsic generation capability. Through the application of IFD, cherry samples can be pinpointed, leading to a marked uptick in model training efficiency. Empirical validations on datasets like Alpaca and WizardLM underpin our findings; with a mere 10% of original data input, our strategy showcases improved results. This synthesis of self-guided cherry-picking and the IFD metric signifies a transformative leap in the instruction tuning of LLMs, promising both efficiency and resource-conscious advancements. Codes, data, and models are available.

Added

2026-09-28

Adaptive Second Order Coresets for Data-efficient Machine Learning

Adaptive Second Order Coresets for Data-efficient Machine Learning

Omead Pooladzandi, David Davini, Baharan Mirzasoleiman

OrganizationsUniversity of California, Los Angeles

Why you should read this

Proposes ADACORE, a data-selection method that dynamically approximates loss curvature through exponentially averaged Hessian estimates to construct weighted training subsets with provable convergence guarantees and over 2.9x training speedups across convex and deep learning models.

Training machine learning models on massive datasets incurs substantial computational costs. To alleviate such costs, there has been a sustained effort to develop data-efficient training methods that can carefully select subsets of the training examples that generalize on par with the full training data. However, existing methods are limited in providing theoretical guarantees for the quality of the models trained on the extracted subsets, and may perform poorly in practice. We propose ADACORE, a method that leverages the geometry of the data to extract subsets of the training examples for efficient machine learning. The key idea behind our method is to dynamically approximate the curvature of the loss function via an exponentially-averaged estimate of the Hessian to select weighted subsets (coresets) that provide a close approximation of the full gradient preconditioned with the Hessian. We prove rigorous guarantees for the convergence of various first and second-order methods applied to the subsets chosen by ADACORE. Our extensive experiments show that ADACORE extracts coresets with higher quality compared to baselines and speeds up training of convex and non-convex machine learning models, such as logistic regression and neural networks, by over 2.9x over the full data and 4.5x over random subsets1.

Added

2026-09-26