Towards Sustainable Learning: Coresets for Data-efficient Deep Learning

Yu YangHao KangBaharan Mirzasoleiman

article2023ICML68 citations

Develops CREST, a scalable coreset selection framework with theoretical convergence guarantees for non-convex optimization that models loss as piecewise quadratic sub-regions and filters learned examples to accelerate deep neural network training by up to 2.5x with minimal accuracy loss.

Listen

Training modern deep neural networks requires massive datasets and compute resources, incurring high financial costs, significant energy consumption, and substantial environmental footprints. While subset selection (coreset) techniques can accelerate training by identifying the most valuable training examples, existing methods have been limited to simpler, convex models. When applied to deep learning, conventional coreset algorithms fail because non-convex training dynamics shift rapidly over time, mini-batch stochastic gradient descent introduces high bias and variance, and re-selecting subsets from full datasets is computationally prohibitive.

The article introduces CREST, a scalable coreset selection framework designed to provide rigorous theoretical convergence guarantees and data-efficient training for deep neural networks. The objective of the article is to demonstrate how modeling non-convex loss functions and extracting mini-batch coresets can significantly reduce training time while maintaining high model accuracy.

The authors evaluated the framework through mathematical convergence analysis and empirical testing on image classification and natural language processing tasks. The experiments included training ResNet-20 on CIFAR-10, ResNet-18 on CIFAR-100, ResNet-50 on TinyImageNet, and fine-tuning RoBERTa on the 570,000-example Stanford Natural Language Inference dataset. CREST approximates the loss surface as piece-wise quadratic regions using gradient and curvature information, iteratively extracts mini-batch coresets from smaller random subsets, and permanently filters out examples once they are consistently learned.

The analysis yielded four major findings. First, CREST accelerated model training by 1.7x to 2.5x compared to training on full datasets, achieving the lowest relative error among all evaluated coreset baselines. Second, CREST scaled to large-scale natural language processing benchmarks where previous full-dataset coreset methods failed due to computational overhead. Third, the framework drastically reduced coreset selection overhead, cutting the required subset update frequency by 74% to 98% relative to greedy mini-batch selection without compromising accuracy. Fourth, behavioral analysis of selected data revealed that neural networks benefit most from curriculum-like learning: models prioritize easier examples early in training and transition toward more difficult examples later, while completely dropping learned examples without degrading final test accuracy.

These findings indicate that organizations training deep learning models can achieve substantial reductions in compute costs, operational timelines, and carbon emissions. By replacing heuristic data pruning with theoretically grounded mini-batch selection, teams can optimize training throughput without sacrificing final generalization performance.

Organizations seeking to optimize deep learning pipelines should consider piloting CREST on large datasets where training costs are substantial. Implementation should prioritize hyperparameter tuning for threshold tolerances and update intervals, paired with efficient data-loading pipelines to maximize real-world wall-clock speedups.

The primary limitation of the method is that its relative performance advantage over random sampling diminishes when the allowable training budget expands beyond constrained compute regimes. Furthermore, because empirical validation was conducted on standard benchmarks up to 570,000 examples, practitioners should validate scalability and data-loading efficiency on multi-billion-parameter models and enterprise-scale datasets.

arXiv: 2306.01244

No sufficiently relevant recommendations were found.

Cover for Towards Sustainable Learning: Coresets for Data-efficient Deep Learning

Abstract

To improve the efficiency and sustainability of learning deep models, we propose CREST, the first scalable framework with rigorous theoretical guarantees to identify the most valuable examples for training non-convex models, particularly deep networks. To guarantee convergence to a stationary point of a non-convex function, CREST models the non-convex loss as a series of quadratic functions and extracts a coreset for each quadratic sub-region. In addition, to ensure faster convergence of stochastic gradient methods such as (mini-batch) SGD, CREST iteratively extracts multiple mini-batch coresets from larger random subsets of training data, to ensure nearly-unbiased gradients with small variances. Finally, to further improve scalability and efficiency, CREST identifies and excludes the examples that are learned from the coreset selection pipeline. Our extensive experiments on several deep networks trained on vision and NLP datasets, including CIFAR-10, CIFAR-100, TinyImageNet, and SNLI, confirm that CREST speeds up training deep networks on very large datasets, by 1.7x to 2.5x with minimum loss in the performance. By analyzing the learning difficulty of the subsets selected by CREST, we show that deep models benefit the most by learning from subsets of increasing difficulty levels 1.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Problem Formulation and Background
  • 4. Coresets for Training Non-convex Models
  • 4.1. Modeling the Non-convex Loss Function
  • 4.2. Coresets for (Mini-batch) Stochastic GD
  • 4.3. Further Improving Efficiency of Coreset Selection
  • 5. Experiments
  • 5.1. Evaluating Accuracy and Speedup
  • 5.2. Ablation Study
  • 6. Conclusion
  • References
  • A. Appendix
  • A.1. Proofs
  • A.2. Experimental details

Knowls

  1. Knowl 1 — Piecewise-quadratic checks keep coresets valid during non-convex training

    model/method

    CREST handles changing example gradients in a non-convex model by selecting a coreset at parameters wtℓw_{t_\ell}, then using it only while a local quadratic approximation remains accurate. For displacement δ\delta from wtℓw_{t_\ell}, the coreset approximation is Fℓ(δ)=L(wtℓ)+gtℓ,SℓTδ+12δTHtℓ,SℓδF^\ell(\delta)=\mathcal{L}(w_{t_\ell})+g_{t_\ell,S_\ell}^{\mathsf T}\delta+\tfrac12\delta^{\mathsf T}H_{t_\ell,S_\ell}\delta, where L\mathcal{L} is the training loss, gtℓ,Sℓg_{t_\ell,S_\ell} is the coreset gradient, and Htℓ,SℓH_{t_\ell,S_\ell} is its curvature estimate. CREST compares this prediction with an estimate of the actual loss on a small random training-data sample VrV_r:

    ρtℓ=∣Fℓ(δ)−Lr(wtℓ+δ)∣Lr(wtℓ+δ),\rho_{t_\ell}=\frac{|F^\ell(\delta)-\mathcal{L}^{r}(w_{t_\ell}+\delta)|}{\mathcal{L}^{r}(w_{t_\ell}+\delta)},

    where Lr\mathcal{L}^{r} is the loss estimated on VrV_r. If ρtℓ≤τ\rho_{t_\ell}\leq\tau, the coreset remains in use; if the ratio exceeds threshold τ\tau, CREST selects a replacement and forms a new local approximation. This adapts the coreset's period of use to the loss landscape: it can be refreshed often when the approximation is only locally valid and less often when the loss is well approximated over a larger region.

  2. Knowl 2 — Random-subset mini-batch coresets target low-bias, low-variance stochastic gradients

    model/method

    At a coreset-selection step, CREST draws PP subsets VpV_p uniformly from the current training examples, each of size rr. For every subset, it selects a mini-batch coreset SpS_p of size at most mm by maximizing C−∑i∈Vpmin⁡j∈Sp∥giL−gjL∥C-\sum_{i\in V_p}\min_{j\in S_p}\|g^L_i-g^L_j\|, where giLg^L_i is the gradient of example ii with respect to the network's last-layer input, and CC is a constant. The resulting union S=⋃p=1PSpS=\bigcup_{p=1}^{P}S_p is used to construct the local quadratic approximation; during training, mini-batches are drawn from the selected mini-batch coresets.

    Each VpV_p is a random estimate of the full training data, while SpS_p is selected to represent its gradient. Thus, the mini-batch coreset gradients are nearly unbiased estimates of the full gradient; the paper reports that their variance is approximately r/mr/m times smaller than that of ordinary random mini-batches of size mm, when the subset coresets represent their random subsets closely. Selecting coresets from smaller random subsets also reduces selection cost: the paper gives total greedy-selection complexity O(rk)O(rk) for a total selected size kk, instead of O(nk)O(nk) when selecting from all nn training examples.

  3. Knowl 3 — Convergence guarantee for CREST's coreset stochastic gradients

    theoretical result

    Suppose the training objective L\mathcal{L} has LsmL_{\rm sm}-Lipschitz gradients, full-data stochastic gradients have variance at most σ2\sigma^2, and the random subsets used by CREST have size rr. Let w0w_0 be the initial parameter vector, L∗\mathcal{L}^* a lower bound on the objective, and ν>0\nu>0 the target stationarity tolerance. The theorem bounds the number of iterations until a run visits a point with gradient norm at most ν\nu, with probability at least 1−λ1-\lambda.

    In the nearly unbiased case, assume that at coreset-selection parameters wtℓw_{t_\ell} the expected norm of the coreset-gradient error ξtℓ\xi_{t_\ell} is at most ϵ∥∇L(wtℓ)∥\epsilon\|\nabla\mathcal{L}(w_{t_\ell})\|, where 0≤ϵ≤min⁡{1,∥∇L(wtℓ+δℓ)∥/(3∥∇L(wtℓ)∥)}0\leq\epsilon\leq\min\{1,\|\nabla\mathcal{L}(w_{t_\ell}+\delta_\ell)\|/(3\|\nabla\mathcal{L}(w_{t_\ell})\|)\}. If the quadratic-loss threshold is small enough to keep the gradient error controlled throughout each accepted neighborhood, CREST has iteration complexity

    O~ ⁣(Lsm(L(w0)−L∗)ν2(1+σ2rν2)).\widetilde{O}\!\left(\frac{L_{\rm sm}(\mathcal{L}(w_0)-\mathcal{L}^*)}{\nu^2}\left(1+\frac{\sigma^2}{r\nu^2}\right)\right).

    In particular, when r≤σ2/ν2r\leq\sigma^2/\nu^2, the paper reports a linear speedup in rr; compared with mini-batch SGD using batch size mm, this corresponds to a factor of r/mr/m under the stated conditions.

    For an absolute coreset-gradient bias bounded by ϵ\epsilon, the corresponding bound is O~ ⁣(Lsm(L(w0)−L∗)ν2−ϵ(1+σ2+rϵ2r(ν2−ϵ)))\widetilde{O}\!\left(\frac{L_{\rm sm}(\mathcal{L}(w_0)-\mathcal{L}^*)}{\nu^2-\epsilon}\left(1+\frac{\sigma^2+r\epsilon^2}{r(\nu^2-\epsilon)}\right)\right) when ϵ<ν2\epsilon<\nu^2. The theorem does not guarantee convergence when ϵ≥ν2\epsilon\geq\nu^2. Here O~\widetilde{O} suppresses logarithmic factors.

  4. Knowl 4 — Estimating coreset curvature without forming a full Hessian

    model/method

    CREST uses a diagonal curvature estimate for the coreset's local quadratic loss rather than explicitly forming a full Hessian matrix. For a Hessian HH and a random vector zz with independent Rademacher entries (each entry is +1+1 or −1-1 with equal probability), the diagonal is estimated using the Hutchinson identity diag⁡(H)=E[z⊙(Hz)]\operatorname{diag}(H)=\mathbb{E}[z\odot(Hz)], where ⊙\odot denotes elementwise multiplication. The product HzHz can be computed through Hessian-vector products obtained by differentiating the gradient, avoiding explicit construction of HH.

    Because neural-network gradients and curvature can be noisy, CREST smooths gradient estimates with an exponential moving average and smooths curvature using a root-mean-square exponential average of Hessian-diagonal estimates. For very large networks, the paper proposes applying these estimates to the input of the penultimate layer rather than to all model parameters.

  5. Knowl 5 — CREST procedure and experimental hyperparameters

    algorithm

    CREST alternates between selecting representative mini-batches, training on them while the local quadratic approximation remains valid, and removing examples that have consistently small loss. The selection objective uses last-layer-input gradients; learned examples are excluded only after their loss remains below threshold α\alpha throughout a monitoring interval.

    Input: Training examples V; initial parameters w0; mini-batch size m; random-subset size r; learning rate η; training-iteration limit N; thresholds τ and α; monitoring interval T2; multipliers b and h.
    Initialize t = 0, T1 = 1, update = true, and P = b × T1.
    While t < N:
        If update is true:
            For p = 1,...,P:
                Draw Vp uniformly at random from the remaining training examples, with |Vp| = r.
                Greedily select Sp ⊆ Vp, |Sp| ≤ m, to maximize
                    C - sum over i in Vp of min over j in Sp of ||gL_i - gL_j||.
            Let S be the union of the selected Sp and form its local quadratic loss approximation.
            Set update = false.
        Train for T1 iterations, drawing a mini-batch coreset from the selected Sp at each iteration;
        update parameters by the stochastic-gradient optimizer and increment t.
        At each non-overlapping T2-iteration monitoring interval, remove examples whose observed
        loss is below α throughout that interval.
        Evaluate ρ by comparing the quadratic loss prediction with the loss on a random reference sample.
        If ρ > τ:
            Set update = true.
            Set T1 = h × ||H0|| / ||Ht||, where H0 and Ht are the initial and current
            Hessian-diagonal estimates, and set P = b × T1.
    Output: The trained parameters.

    In the reported experiments, b=5b=5, T2=20T_2=20, and α=0.1\alpha=0.1. The random-subset and reference-sample sizes were r=0.01nr=0.01n for the vision datasets and r=0.005nr=0.005n for SNLI, where nn is the number of training examples. The settings (τ,h)(\tau,h) were (0.05,1)(0.05,1) for CIFAR-10, (0.01,10)(0.01,10) for CIFAR-100, (0.005,1)(0.005,1) for TinyImageNet, and (0.05,4)(0.05,4) for SNLI.

  6. Knowl 6 — Benchmark relative error under a 10% training budget

    empirical result

    CREST was evaluated against random mini-batches and coreset baselines using a 10% training budget. The reported metric is relative test-accuracy error, ∣accselected−accfull∣/accfull|\mathrm{acc}_{\rm selected}-\mathrm{acc}_{\rm full}|/\mathrm{acc}_{\rm full}, in percent; lower is better. For the vision tasks, training used 200 epochs of SGD with momentum 0.90.9, batch size 128, and learning-rate drops by a factor of 0.1 at 60% and 85% of training. RoBERTa fine-tuning on SNLI used AdamW, learning rate 10−510^{-5}, batch size 32, and eight epochs. Experiments used one NVIDIA RTX A6000 GPU. Each entry below is mean ±\pm reported variation; dashes indicate unavailable baseline results.

    Dataset Model SGD†^\dagger Random CRAIG GradMatch GLISTER CREST
    CIFAR-10 ResNet-20 21.3±\pm8.0 7.2±\pm1.4 13.0±\pm5.1 6.0±\pm0.1 7.0±\pm0.1 5.5±\pm0.2
    CIFAR-100 ResNet-18 36.5±\pm2.9 11.7±\pm0.4 17.2±\pm4.5 12.7±\pm0.9 27.6±\pm4.0 9.4±\pm0.3
    TinyImageNet ResNet-50 32.8±\pm2.1 16.0±\pm0.5 28.5±\pm0.6 27.7±\pm0.2 32.8±\pm2.1 15.4±\pm0.6
    SNLI RoBERTa (fine-tune) 1.2±\pm0.3 1.2±\pm0.3 – – – 0.8±\pm0.2

    CREST had the lowest reported relative error on each task. On SNLI, with 570,000 training examples, it was the only coreset baseline reported as applicable; the other full-data coreset-selection baselines did not scale to that dataset.

  7. Knowl 7 — Training speed and measured selection overhead

    empirical result

    Across the reported vision and language benchmarks, CREST accelerated training by 1.7× to 2.5× relative to training on the full data while retaining comparatively low relative accuracy error. The largest reported speedup was 2.5×. On CIFAR-100 with ResNet-18 and batch size 128, the mean measured times per operation were 0.006 seconds to select a CREST mini-batch coreset, 0.089 seconds to select a CRAIG coreset, 0.115 seconds to calculate the quadratic loss approximation, and 0.796 seconds to check its threshold on a random sample. The comparison shows that mini-batch selection from smaller random subsets was faster than selecting a 10%-of-data coreset from the full dataset, although approximation checks themselves incurred measurable overhead.

  8. Knowl 8 — Ablations show the value of quadratic scheduling, smoothing, and example removal

    empirical result

    On ResNet-20/CIFAR-10, the full CREST configuration achieved relative error 4.33% with 185 coreset updates. Removing example exclusion increased these to 4.61% and 346 updates; removing gradient and curvature smoothing gave 7.44% and 369 updates; replacing CREST's scheduling with the first-order variant (“CREST-First”) gave 7.45% and 343 updates. Thus, the full configuration reduced both error and update count in this comparison.

    Against greedy mini-batch selection that selects a new coreset for every mini-batch, CREST used 2% as many updates on CIFAR-100 while retaining 98% of its performance, 3% on CIFAR-10 while retaining 99%, 19% on TinyImageNet while retaining 95%, and 26% on SNLI while retaining 99%. The experiments also found that quadratic regions grew more usable later in training; more frequent updates early in training mattered more for final accuracy than extra updates late in training.

  9. Knowl 9 — Selected examples become harder over training, while learned examples can be dropped

    empirical result

    The paper measured example difficulty by forgettability: the number of times an example was misclassified after it had previously been classified correctly. The average forgettability of CREST-selected examples increased as training proceeded, whereas randomly selected subsets maintained a lower, comparatively constant difficulty. Excluding already-learned examples made CREST focus still more on difficult examples. Selection frequencies were long-tailed, indicating that examples were not chosen equally often.

    On ResNet-20/CIFAR-10, the accuracy on examples removed from the selection pool was about 92% earlier in training and later rose above 99%, even though those examples were not selected again. The authors report this as evidence that training on the remaining examples could recover performance on dropped examples.

  10. Knowl 10 — Benefits are strongest under limited training budgets

    limitation

    CREST's relative advantage is most pronounced when the training budget is limited; the authors report a smaller accuracy gap over the Random baseline at a 20% budget than at a 10% budget. At 20% budget, relative errors (%) for CREST, Random, and standard SGD trained for the corresponding number of iterations were respectively 2.32, 2.87, and 16.47 on CIFAR-10/ResNet-20; 3.37, 3.66, and 32.68 on CIFAR-100/ResNet-18; and 3.05, 3.51, and 47.43 on TinyImageNet/ResNet-50. The authors also note that more efficient data loading could further improve coreset-selection speed.

Coverage note — No substantial contributed material was omitted; background and proof-only derivations were excluded.

References

  1. 1.Alain, G., Lamb, A., Sankar, C., Courville, A., and Bengio, Y. Variance reduction in sgd by distributed importance sampling. arXiv preprint arXiv:1511.06481, 2015.
  2. 2.Asi, H. and Duchi, J. C. The importance of better models in stochastic optimization. arXiv preprint arXiv:1903.08619, 2019.
  3. 3.Bertsekas, D. P. Convexification procedures and decomposition methods for nonconvex optimization problems. Journal of Optimization Theory and Applications, 29(2): 169–197, 1979.
  4. 4.Birodkar, V., Mobahi, H., and Bengio, S. Semantic redundancies in image-classification datasets: The 10% you don’t need. arXiv preprint arXiv:1901.11409, 2019.
  5. 5.Bollapragada, R., Byrd, R. H., and Nocedal, J. Exact and inexact subsampled newton methods for optimization. IMA Journal of Numerical Analysis, 39(2):545–578, 2019.
  6. 6.Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 632–642, Lisbon, Portugal, September 2015. Association for Computational Linguistics. doi: 10.18653/v1/D15-1075. URL https://aclanthology.org/D15-1075.
  7. 7.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  8. 8.Carmon, Y., Duchi, J. C., Hinder, O., and Sidford, A. Accelerated methods for nonconvex optimization. SIAM Journal on Optimization, 28(2):1751–1772, 2018.
  9. 9.Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M. Selection via proxy: Efficient data selection for deep learning. In International Conference on Learning Representations (ICLR), 2020.
  10. 10.Defazio, A. and Bottou, L. On the ineffectiveness of variance reduced optimization for deep learning. Advances in Neural Information Processing Systems, 32:1755–1765, 2019.
  11. 11.Dembo, R. S., Eisenstat, S. C., and Steihaug, T. Inexact newton methods. SIAM Journal on Numerical analysis, 19(2):400–408, 1982.
  12. 12.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  13. 13.Ghadimi, S. and Lan, G. Stochastic first-and zeroth-order methods for nonconvex stochastic programming. SIAM Journal on Optimization, 23(4):2341–2368, 2013.
  14. 14.Hutchinson, M. F. A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines. Communications in Statistics-Simulation and Computation, 18(3):1059–1076, 1989.
  15. 15.Jin, C., Netrapalli, P., Ge, R., Kakade, S. M., and Jordan, M. I. On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points. Journal of the ACM (JACM), 68(2):1–29, 2021.
  16. 16.Katharopoulos, A. and Fleuret, F. Not all samples are created equal: Deep learning with importance sampling. In International conference on machine learning, pp. 2525–2534. PMLR, 2018.
  17. 17.Killamsetty, K., Durga, S., Ramakrishnan, G., De, A., and Iyer, R. Grad-match: Gradient matching based data subset selection for efficient deep model training. In International Conference on Machine Learning, pp. 5464–5474. PMLR, 2021a.
  18. 18.Killamsetty, K., Sivasubramanian, D., Ramakrishnan, G., and Iyer, R. Glister: Generalization based data subset selection for efficient and robust learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 8110–8118, 2021b.
  19. 19.Krizhevsky, A., Nair, V., and Hinton, G. Cifar-10 (canadian institute for advanced research). 2009. URL http://www.cs.toronto.edu/~kriz/cifar.html.
  20. 20.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  21. 21.Loshchilov, I. and Hutter, F. Online batch selection for faster training of neural networks. arXiv preprint arXiv:1511.06343, 2015.
  22. 22.Marquardt, D. W. An algorithm for least-squares estimation of nonlinear parameters. Journal of the society for Industrial and Applied Mathematics, 11(2):431–441, 1963.
  23. 23.Martens, J. and Grosse, R. Optimizing neural networks with kronecker-factored approximate curvature. In International conference on machine learning, pp. 2408–2417. PMLR, 2015.
  24. 24.Mindermann, S., Brauner, J. M., Razzak, M. T., Sharma, M., Kirsch, A., Xu, W., Höltgen, B., Gomez, A. N., Morisot, A., Farquhar, S., et al. Prioritized training on points that are learnable, worth learning, and not yet learnt. In International Conference on Machine Learning, pp. 15630–15649. PMLR, 2022.
  25. 25.Mirzasoleiman, B., Karbasi, A., Sarkar, R., and Krause, A. Distributed submodular maximization: Identifying representative elements in massive data. In Advances in Neural Information Processing Systems, pp. 2049–2057, 2013.
  26. 26.Mirzasoleiman, B., Bilmes, J., and Leskovec, J. Coresets for data-efficient training of machine learning models. In International Conference on Machine Learning, pp. 6950–6960. PMLR, 2020.
  27. 27.Paul, M., Ganguli, S., and Dziugaite, G. K. Deep learning on a data diet: Finding important examples early in training. Advances in Neural Information Processing Systems, 34: 20596–20607, 2021.
  28. 28.Pooladzandi, O., Davini, D., and Mirzasoleiman, B. Adaptive second order coresets for data-efficient machine learning. In International Conference on Machine Learning, pp. 17848–17869. PMLR, 2022.
  29. 29.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015. doi: 10.1007/s11263-015-0816-y.
  30. 30.Schaul, T., Quan, J., Antonoglou, I., and Silver, D. Prioritized experience replay. arXiv preprint arXiv:1511.05952, 2015.
  31. 31.Schwartz, R., Dodge, J., Smith, N. A., and Etzioni, O. Green ai. arXiv preprint arXiv:1907.10597, 2019.
  32. 32.Strubell, E., Ganesh, A., and McCallum, A. Energy and policy considerations for deep learning in nlp. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 3645–3650, 2019.
  33. 33.Toneva, M., Sordoni, A., des Combes, R. T., Trischler, A., Bengio, Y., and Gordon, G. J. An empirical study of example forgetting during deep neural network learning. In International Conference on Learning Representations, 2018.
  34. 34.Wang, W. and Srebro, N. Stochastic nonconvex optimization with large minibatches. In Algorithmic Learning Theory, pp. 857–882. PMLR, 2019.
  35. 35.Xiao, T., Xia, T., Yang, Y., Huang, C., and Wang, X. Learning from massive noisy labeled data for image classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2691–2699, 2015.
  36. 36.Xu, P., Roosta, F., and Mahoney, M. W. Newton-type methods for non-convex optimization under inexact hessian information. Mathematical Programming, 184(1):35–70, 2020.
  37. 37.Yao, Z., Xu, P., Roosta-Khorasani, F., and Mahoney, M. W. Inexact non-convex newton-type methods. arXiv preprint arXiv:1802.06925, 2018.
  38. 38.Yao, Z., Gholami, A., Shen, S., Keutzer, K., and Mahoney, M. W. Adahessian: An adaptive second order optimizer for machine learning. arXiv preprint arXiv:2006.00719, 2020.
  39. 39.Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L. Scaling vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12104–12113, 2022.

Citation

MLA
Yang, Y., et al. “Towards Sustainable Learning: Coresets for Data-efficient Deep Learning”. International Conference on Machine Learning, vol. 202, 2023, pp. 39314–30, https://proceedings.mlr.press/v202/yang23g.html.
APA
Yang, Y., Kang, H., & Mirzasoleiman, B. (2023). Towards Sustainable Learning: Coresets for Data-efficient Deep Learning. International Conference on Machine Learning, 202, 39314–39330. https://proceedings.mlr.press/v202/yang23g.html
Chicago
Yang, Y., H. Kang, and B. Mirzasoleiman. 2023. “Towards Sustainable Learning: Coresets for Data-efficient Deep Learning”. International Conference on Machine Learning 202: 39314–30. https://proceedings.mlr.press/v202/yang23g.html.
Harvard
Yang, Y., Kang, H. and Mirzasoleiman, B. (2023) “Towards Sustainable Learning: Coresets for Data-efficient Deep Learning”, International Conference on Machine Learning. PMLR, pp. 39314–39330. Available at: https://proceedings.mlr.press/v202/yang23g.html.
Vancouver
1. Yang Y, Kang H, Mirzasoleiman B (2023) Towards Sustainable Learning: Coresets for Data-efficient Deep Learning. In: International Conference on Machine Learning. PMLR, pp 39314–39330

BibTeX

@InProceedings{pmlr-v202-yang23g,
  title = 	 {Towards Sustainable Learning: Coresets for Data-efficient Deep Learning},
  author =       {Yang, Yu and Kang, Hao and Mirzasoleiman, Baharan},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {39314--39330},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/yang23g/yang23g.pdf},
  url = 	 {https://proceedings.mlr.press/v202/yang23g.html},
  abstract = 	 {To improve the efficiency and sustainability of learning deep models, we propose CREST, the first scalable framework with rigorous theoretical guarantees to identify the most valuable examples for training non-convex models, particularly deep networks. To guarantee convergence to a stationary point of a non-convex function, CREST models the non-convex loss as a series of quadratic functions and extracts a coreset for each quadratic sub-region. In addition, to ensure faster convergence of stochastic gradient methods such as (mini-batch) SGD, CREST iteratively extracts multiple mini-batch coresets from larger random subsets of training data, to ensure nearly-unbiased gradients with small variances. Finally, to further improve scalability and efficiency, CREST identifies and excludes the examples that are learned from the coreset selection pipeline. Our extensive experiments on several deep networks trained on vision and NLP datasets, including CIFAR-10, CIFAR-100, TinyImageNet, and SNLI, confirm that CREST speeds up training deep networks on very large datasets, by 1.7x to 2.5x with minimum loss in the performance. By analyzing the learning difficulty of the subsets selected by CREST, we show that deep models benefit the most by learning from subsets of increasing difficulty levels.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/