Decision-Focused Learning: Foundations, State of the Art, Benchmark and Future Opportunities
Jayanta MandiJames KotarySenne BerdenMaxime MulambaVictor BucareyTias GunsFerdinando Fioretto
Presents a comprehensive taxonomy of decision-focused learning techniques alongside an open-source benchmark evaluating eleven methods across seven optimization problems to guide end-to-end integration of machine learning and constrained optimization under uncertainty.
Many critical real-world operations—such as supply chain planning, energy scheduling, financial portfolio management, and transportation routing—require solving complex constrained optimization problems where key parameters are uncertain. Traditionally, organizations use a two-stage "predict-then-optimize" approach: a machine learning model first predicts the uncertain parameters (e.g., travel times or future prices) by minimizing standard prediction errors like mean squared error, and an optimization algorithm subsequently prescribes the optimal decision. However, because standard prediction metrics treat all forecasting errors equally regardless of how they affect the final operational outcome, small prediction inaccuracies can lead to severely suboptimal and costly business decisions.
The main objective of the article is to provide a comprehensive evaluation and survey of Decision-Focused Learning (DFL), an emerging paradigm that integrates machine learning and constrained optimization into an end-to-end training system. By embedding the optimization problem directly into the learning loop, DFL trains predictive models to minimize downstream operational losses—such as decision regret—rather than intermediate statistical prediction errors.
To assess the state of the art, the article analyzes both gradient-based and gradient-free DFL methodologies across diverse algorithmic designs. The authors establish a systematic classification of gradient-based techniques into four primary categories: analytical differentiation of optimization mappings, analytical smoothing, smoothing via random perturbations, and differentiation using surrogate loss functions. To evaluate practical performance, the study establishes an open benchmark spanning seven distinct optimization tasks—including shortest path routing, portfolio optimization, energy-cost-aware scheduling, knapsack allocation, and diverse bipartite matching—using both synthetic benchmarks and real-world empirical datasets.
The investigation yields several critical findings. First, DFL consistently produces higher-quality decisions and lower regret than traditional two-stage methods whenever machine learning models are misspecified or subject to real-world data noise and parameter correlations. Second, direct differentiation through optimization problems faces a fundamental technical barrier: linear and combinatorial optimization models produce piecewise-constant solution mappings with zero gradients almost everywhere, necessitating specialized smoothing or surrogate approximations. Third, while DFL substantially improves operational performance, solving and differentiating optimization models during every training iteration introduces significant computational overhead. Finally, benchmark evaluations demonstrate that surrogate loss techniques (such as SPO+) and perturbation-based methods (such as DBB and I-MLE) offer flexible, solver-agnostic implementations that perform well across combinatorial problems without requiring specialized solver modifications.
These findings indicate that adopting DFL can significantly improve operational efficiency, minimize financial suboptimality, and mitigate risk in automated decision systems. However, implementing DFL requires navigating clear engineering trade-offs between model complexity, solver differentiation techniques, and training runtimes. Standard two-stage prediction pipelines remain computationally faster and asymptotically effective only when the predictive model perfectly captures the data distribution, a condition rarely met in practice.
For practitioners and technology leaders, the article supports adopting DFL in high-stakes operational settings where decision regret directly drives costs. For complex combinatorial tasks, teams should leverage solver-agnostic surrogate losses or perturbation methods, or utilize solution-caching techniques that solve the underlying optimization problem only on a subset of iterations to reduce computational bottlenecks by up to 95%. When moving toward deployment, organizations must conduct careful hyperparameter tuning and evaluate model robustness against label noise and data poisoning. Further research and validation remain necessary for problems with uncertain constraints that risk generating infeasible decisions, as well as for scaling end-to-end learning to large, non-linear integer optimization problems.
- Paper: Machine Learning for Combinatorial Optimization: a Methodological Tour d'Horizon, Yoshua Bengio et al. (2018). This survey establishes how machine learning can be integrated with combinatorial optimization, the methodological foundation on which decision-focused learning builds.
- Paper: Learning to Branch with Tree MDPs, Lara Scavuzzo et al. (2022). Its learned branching policies show a concrete route for using learning inside a combinatorial solver, helping frame the decision-focused methods reviewed in the source.
No sufficiently relevant recommendations were found.
