Learning to Optimize: A Primer and A Benchmark

Tianlong ChenXiaohan ChenWuyang ChenHoward HeatonJialin LiuZhangyang WangWotao Yin

article2022JMLR386 citations

Presents a comprehensive survey and standardized benchmarking framework for learning-to-optimize methods in continuous optimization, providing practical categorizations, open challenges, and the open-source Open-L2O package for reproducible evaluation.

Listen

Modern industrial and engineering systems frequently solve repetitive continuous optimization problems across domains such as medical imaging, signal processing, and machine learning model training. Traditional optimization algorithms are manually designed around worst-case mathematical theories, which often results in slow iterative runtimes and suboptimal performance on specific data distributions. Learning to Optimize (L2O) has emerged as a data-driven paradigm that trains machine learning models to discover or refine optimization update rules automatically. The article evaluates the state of this emerging field, introduces a comprehensive taxonomy of techniques, and benchmarks leading L2O approaches to determine their practical performance and reliability across representative optimization tasks.

The article systematically analyzes two main classes of learned optimizers: model-free methods, which learn black-box update rules using neural networks such as recurrent architectures, and model-based methods, which embed machine learning components into established analytical algorithms via algorithm unrolling or plug-and-play operators. To resolve inconsistent evaluation standards across prior studies, the authors established the Open-L2O benchmark platform. They conducted standardized empirical evaluations spanning three representative testbeds: convex sparse optimization, non-convex landscape minimization using the generalized Rastrigin function, and multi-layer neural network training across unseen architectures and dataset shifts.

The benchmark results reveal three primary findings. First, model-based L2O methods substantially outperform both classical solvers and model-free techniques when problem structure is available; for example, unrolled sparse solvers reached target precision in approximately 16 iterations compared to hundreds or thousands of steps needed by standard iterative algorithms. Second, model-free L2O methods exhibited inconsistent gains; while population-based swarm methods navigated complex non-convex landscapes more effectively than standard gradient techniques, standard model-free optimizers failed to match basic analytical solvers on structured convex tasks. Third, model-free optimizers demonstrated poor generalization and instability when deployed on longer iteration horizons, higher problem dimensions, or unseen neural network architectures, frequently diverging due to truncation bias and optimizer training artifacts.

These findings imply that learned optimization is not yet a universal replacement for classic general-purpose solvers, but it offers significant speedups and computational cost reductions for specialized, repetitive tasks. Model-based approaches represent a practical sweet spot by combining theoretical algorithmic structure with data-driven tuning, making them well-suited for high-throughput operational pipelines like image reconstruction. Conversely, purely model-free optimizers carry high operational risks due to memory bottlenecks during training and erratic performance on tasks that deviate from the training distribution.

Organizations evaluating L2O should focus near-term adoption on model-based methods for well-structured inverse problems while treating general model-free optimizers as exploratory. Future efforts should prioritize hybrid architectures, memory-efficient training procedures to handle large-scale models, and automated safeguarding mechanisms that fall back to classic algorithms when out-of-distribution inputs are detected. Because theoretical guarantees remain largely restricted to specialized model-based settings and empirical testing was limited to moderate problem scales, stakeholders should exercise caution and conduct thorough distributional validation before integrating learned optimizers into mission-critical systems.

Cover for Learning to Optimize: A Primer and A Benchmark

Abstract

Learning to optimize (L2O) is an emerging approach that leverages machine learning to develop optimization methods, aiming at reducing the laborious iterations of hand engineering. It automates the design of an optimization method based on its performance on a set of training problems. This data-driven procedure generates methods that can efficiently solve problems similar to those in training. In sharp contrast, the typical and traditional designs of optimization methods are theory-driven, so they obtain performance guarantees over the classes of problems specified by the theory. The difference makes L2O suitable for repeatedly solving a particular optimization problem over a specific distribution of data, while it typically fails on out-of-distribution problems. The practicality of L2O depends on the type of target optimization, the chosen architecture of the method to learn, and the training procedure. This new paradigm has motivated a community of researchers to explore L2O and report their findings.

This article is poised to be the first comprehensive survey and benchmark of L2O for continuous optimization. We set up taxonomies, categorize existing works and research directions, present insights, and identify open challenges. We benchmarked many existing L2O approaches on a few representative optimization problems. For reproducible research and fair benchmarking purposes, we released our software implementation and data in the package Open-L2O at https://github.com/VITA-Group/Open-L2O.

Table of Contents

  • 1. Introduction
  • 1.1 Background and Motivation
  • 1.2 Preliminaries
  • 1.3 Broader Contexts
  • 1.4 Paper Scope and Organization
  • 2. Model-Free L2O Approaches
  • 2.1 LSTM Optimizer for Continuous Minimization: Basic Idea and Variants
  • 2.2 Other Common Implementations for Model-Free L2O
  • 2.3 More Optimization Tasks for Model-Free L2O
  • 3. Model-Based L2O Approaches
  • 3.1 Plug and Play
  • 3.2 Algorithm Unrolling
  • 3.2.1 Different Target Problems
  • 3.2.2 Different Algorithms Unrolled
  • 3.2.3 Objective-Based v.s. Inverse Problems
  • 3.2.4 Learned Parameter Roles
  • 3.3 Applications
  • 3.4 Theoretical Efforts
  • 3.4.1 Capacity
  • 3.4.2 Interpretability
  • 3.4.3 Generalization
  • 4. The Open-L2O Benchmark
  • 4.1 Test 1: Convex Sparse Optimization
  • 4.1.1 Learning to Perform Sparse Optimization and Sparse-Signal Recovery
  • 4.1.2 Learning to Minimize the Lasso Model
  • 4.2 Test 2: Minimization of Non-Convex function Rastrigin
  • 4.3 Test 3: Neural Network Training
  • 4.4 Take-Home Messages
  • 5. Concluding Remarks
  • Acknowledgments
  • Appendix A. List of Abbreviations
  • References

Knowls

  1. Knowl 1 — No paper content was available for extraction

    limitation

    The source document referenced for knowl extraction provided no readable text, so the paper's methods, models, theory, experiments, and results could not be identified or restated. No claim about the paper's content is made here.

Coverage note — The attached document (7ca71290-2e59-4cb2-aa7d-ec42c549204d.pdf) contains no extractable text content in this session, so no knowls could be extracted from the paper's contribution; nothing was deliberately omitted, as no contributed material was available to process.

Citation

MLA
Chen, T., et al. “Learning to Optimize: A Primer and A Benchmark”. Journal of Machine Learning Research, vol. 23, no. 189, 2022, pp. 1–9, https://www.jmlr.org/papers/v23/21-0308.html.
APA
Chen, T., Chen, X., Chen, W., Heaton, H., Liu, J., Wang, Z., & Yin, W. (2022). Learning to Optimize: A Primer and A Benchmark. Journal of Machine Learning Research, 23(189), 1–59. https://www.jmlr.org/papers/v23/21-0308.html
Chicago
Chen, T., X. Chen, W. Chen, et al. 2022. “Learning to Optimize: A Primer and A Benchmark”. Journal of Machine Learning Research 23 (189): 1–59. https://www.jmlr.org/papers/v23/21-0308.html.
Harvard
Chen, T. et al. (2022) “Learning to Optimize: A Primer and A Benchmark”, Journal of Machine Learning Research, 23(189), pp. 1–59. Available at: https://www.jmlr.org/papers/v23/21-0308.html.
Vancouver
1. Chen T, Chen X, Chen W, Heaton H, Liu J, Wang Z, Yin W (2022) Learning to Optimize: A Primer and A Benchmark. Journal of Machine Learning Research 23:1–59

BibTeX

@article{JMLR:v23:21-0308,
  author  = {Tianlong Chen and Xiaohan Chen and Wuyang Chen and Howard Heaton and Jialin Liu and Zhangyang Wang and Wotao Yin},
  title   = {Learning to Optimize: A Primer and A Benchmark},
  journal = {Journal of Machine Learning Research},
  year    = {2022},
  volume  = {23},
  number  = {189},
  pages   = {1--59},
  url     = {http://jmlr.org/papers/v23/21-0308.html}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/