Diffusion Models for Black-Box Optimization

Siddarth KrishnamoorthySatvik Mehul MashkariaAditya Grover

article2023ICML110 citations

Proposes Denoising Diffusion Optimization Models, an inverse approach that pairs conditional diffusion with objective reweighting and classifier-free guidance to generate candidate optima that surpass the best observations in offline datasets across diverse continuous and discrete benchmarks.

Listen

Many critical applications in science and engineering, such as drug design, materials discovery, and robotics, require optimizing complex systems where real-world testing is expensive, slow, or hazardous. To bypass costly live interactions, practitioners rely on offline black-box optimization, which attempts to find optimal designs using only fixed, historical datasets. However, standard methods struggle with limited data coverage; conventional forward surrogate models often produce inaccurate predictions outside the logged data distribution, while inverse generative models (such as generative adversarial networks) frequently suffer from unstable training and repetitive, low-diversity outputs.

The main objective of the article is to demonstrate that conditional diffusion models can serve as an effective, stable inverse framework for offline black-box optimization. To achieve this, the authors introduce Denoising Diffusion Optimization Models (DDOM), an approach that maps target performance values directly back to high-dimensional design inputs.

The authors evaluated DDOM across synthetic benchmarks and six complex tasks from the standard Design-Bench suite, encompassing continuous domains like robot morphology and superconductor design as well as discrete domains like DNA sequence binding and chemical property optimization. The approach integrates a reweighted training loss to bias learning toward higher-performing historical data points without discarding lower-tier data. During evaluation, candidate solutions are generated by conditioning the model on top-tier performance targets and using classifier-free guidance—a technique that prioritizes target compliance over broad output diversity.

The experimental findings show that DDOM delivers state-of-the-art performance across diverse problem domains. First, DDOM achieved the best average rank of 2.8 across all evaluated baselines, placing first or second on four of the six Design-Bench tasks. Second, in the superconductor optimization task, DDOM outperformed the nearest baseline by 11% and generated valid candidates that exceeded the best values found in the training dataset. Third, ablation studies demonstrated that loss reweighting consistently improved solution quality across all tasks compared to unweighted training. Finally, classifier-free guidance proved vital; omitting guidance resulted in significantly worse optimization results, whereas appropriate guidance allowed the model to reliably steer generated designs toward high-value regions.

These findings indicate that diffusion-based inverse modeling provides a more dependable, resilient framework for data-driven engineering and design. By avoiding the training instability of alternative generative approaches and providing stable, low-variance predictions, DDOM reduces the operational risk and computational cost of generating viable candidates from historical data. This capability enables organizations to extract greater value from existing experimental logs without investing in costly additional wet-lab or physical trials.

Organizations evaluating data-driven optimization workflows should consider piloting conditional diffusion models for high-dimensional design problems. Implementation should incorporate loss reweighting and classifier-free guidance to balance sample fidelity against target objectives. For deployment, teams should note that sampling candidates from diffusion models is computationally slower than from single-step generative alternatives; while acceptable for offline design cycles, accelerated sampling strategies or hybrid workflows that use DDOM to warm-start online optimization should be explored before deploying in latency-critical environments. Given the consistent experimental performance across varied domains, confidence in DDOM’s offline optimization capability is high, with the primary caution revolving around sampling runtimes and reliance on dataset quality.

arXiv: 2306.07180
  • Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). Classifier-free guidance provides the foundational mechanism leveraged directly by DDOM to bias generated designs toward high-performance target objectives without requiring an external classifier.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This seminal work establishes the foundational denoising diffusion probabilistic model framework upon which conditional diffusion optimization algorithms are formulated.
  • Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Diffuser pioneered formulating offline decision-making and planning as conditional generative trajectory diffusion, establishing the conceptual precedent for diffusion-based black-box optimization.
  • Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). This paper develops discrete denoising diffusion probabilistic models, providing the essential algorithmic basis for DDOM's handling of discrete design optimization benchmarks such as DNA sequence binding.
  • Paper: Diffusion Models: A Comprehensive Survey of Methods and Applications, Ling Yang et al. (2022). This comprehensive survey outlines the unified principles, forward and reverse trajectories, and optimization trade-offs of diffusion models necessary for understanding inverse design frameworks.
Cover for Diffusion Models for Black-Box Optimization

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Problem Statement
  • 2.2. Diffusion Models
  • 2.3. Classifier-free Guidance
  • 3. Denoising Diffusion Optimization Models
  • 4. Experiments
  • 4.1. Toy Branin Task
  • 4.2. Design-Bench
  • 4.3. Ablations
  • 5. Related Work
  • 6. Summary
  • Acknowledgements
  • References
  • A. Notation and Experimental Details
  • A.1. Notation
  • A.2. Additional Experimental details
  • A.3. Unnormalized results
  • A.4. HopperController
  • B. Additional Ablations and Analysis
  • B.1. Effect of evaluation budget
  • B.2. Ablation on reweighting parameters
  • B.3. Visualizing the Design-Bench datasets
  • B.4. Plotting the contours of the score function
  • B.5. Effect of adding randomness to data
  • B.6. Effect of size of the offline dataset
  • C. Societal Impact

Knowls

  1. Knowl 1 — Denoising Diffusion Optimization Models Framework

    model/method

    Denoising Diffusion Optimization Models (DDOM) formulate offline black-box optimization (BBO) as learning an inverse conditional generative mapping from objective values yy to candidate design points x∈X⊆Rd\mathbf{x} \in \mathcal{X} \subseteq \mathbb{R}^d. In offline BBO, an algorithm receives a static dataset D={(xi,yi)}i=1n\mathcal{D} = \{(\mathbf{x}_i, y_i)\}_{i=1}^n of evaluations of an unknown expensive objective function f:X→Rf: \mathcal{X} \to \mathbb{R}, where yi=f(xi)y_i = f(\mathbf{x}_i), with no active oracle evaluations allowed during training. The objective is to return a set of candidate solutions X\mathbf{X} within a query budget ∣X∣≤Q|\mathbf{X}| \le Q that maximizes ff.

    Because multiple inputs can map to identical function values, the inverse mapping is one-to-many and parameterized as a conditional probability distribution p(x∣y)p(\mathbf{x}|y). DDOM adopts a continuous-time Variance Preserving (VP) Stochastic Differential Equation (SDE) for forward diffusion: dx=−12βtx dt+βt dwd\mathbf{x} = -\frac{1}{2}\beta_t \mathbf{x} \, dt + \sqrt{\beta_t} \, d\mathbf{w} where t∈[0,1]t \in [0, 1], βt=βmin⁡+(βmax⁡−βmin⁡)t\beta_t = \beta_{\min} + (\beta_{\max} - \beta_{\min})t, and w\mathbf{w} is a standard Wiener process. A time-dependent score neural network ϵθ(xt,t,y)\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t, y) parameterizes the reverse-time denoising process, learning to invert Gaussian noise into high-dimensional points x0\mathbf{x}_0 concentrated around regions where f(x0)≈yf(\mathbf{x}_0) \approx y.

  2. Knowl 2 — Dataset Loss Reweighting for High-Value Optimization Bias

    equation

    To bias conditional diffusion models toward regions of high objective value without discarding lower-value samples that provide score gradient signals, DDOM uses a bin-based loss reweighting scheme. The offline dataset D\mathcal{D} is partitioned into NBN_B bins {B1,…,BNB}\{B_1, \dots, B_{N_B}\} of equal interval width over the observed range of yy. Each bin BiB_i is assigned a scalar weight wiw_i: wi=∣Bi∣∣Bi∣+Kexp⁡(−∣y^−ybi∣τ)w_i = \frac{|B_i|}{|B_i| + K} \exp\left(-\frac{|\hat{y} - y_{b_i}|}{\tau}\right) where y^=max⁡(x,y)∈Dy\hat{y} = \max_{(\mathbf{x}, y) \in \mathcal{D}} y is the maximum objective value in the dataset, ∣Bi∣|B_i| is the number of samples in bin BiB_i, ybiy_{b_i} is the midpoint of the function value interval for bin BiB_i, K>0K > 0 is a smoothing parameter balancing bin frequency versus function quality (typically K=0.01nK = 0.01 n for dataset size nn), and τ>0\tau > 0 is a temperature hyperparameter controlling the penalization of lower-value bins.

    The conditional score network ϵθ\boldsymbol{\epsilon}_\theta is trained by minimizing the reweighted denoising score matching objective: Et[λ(t) E(x0,y)∼D[w(y) Ext∣x0[∥ϵθ(xt,t,y)−∇xtlog⁡pt(xt∣x0)∥22]]]\mathbb{E}_t \left[ \lambda(t) \, \mathbb{E}_{(\mathbf{x}_0, y) \sim \mathcal{D}} \left[ w(y) \, \mathbb{E}_{\mathbf{x}_t | \mathbf{x}_0} \left[ \|\boldsymbol{\epsilon}_\theta(\mathbf{x}_t, t, y) - \nabla_{\mathbf{x}_t} \log p_t(\mathbf{x}_t | \mathbf{x}_0)\|_2^2 \right] \right] \right] where w(y)=wiw(y) = w_i when y∈Biy \in B_i, and λ(t)\lambda(t) is the standard time-dependent loss weighting of the diffusion process.

  3. Knowl 3 — Classifier-Free Guidance for Offline BBO Reverse Sampling

    equation

    Standard conditional diffusion models often trade conditioning adherence for sample diversity. In offline BBO, where the goal is to optimize strictly toward high function values, DDOM applies classifier-free guidance to steer the reverse diffusion process toward the target condition.

    During training, the conditioning variable yy is randomly dropped (set to zero) with a fixed dropout probability (such as pdrop=0.15p_{\text{drop}} = 0.15) so that a single network ϵθ\boldsymbol{\epsilon}_\theta jointly estimates conditional scores ϵcond(x,t,y)\boldsymbol{\epsilon}_{\text{cond}}(\mathbf{x}, t, y) and unconditional scores ϵuncond(x,t)\boldsymbol{\epsilon}_{\text{uncond}}(\mathbf{x}, t). At inference time, conditioning is set to the dataset maximum ytest=max⁡(x,y)∈Dyy_{\text{test}} = \max_{(\mathbf{x}, y) \in \mathcal{D}} y, and the effective score is computed as: ϵθ(x,t,ytest)=(1+γ)ϵcond(x,t,ytest)−γϵuncond(x,t)\boldsymbol{\epsilon}_\theta(\mathbf{x}, t, y_{\text{test}}) = (1 + \gamma) \boldsymbol{\epsilon}_{\text{cond}}(\mathbf{x}, t, y_{\text{test}}) - \gamma \boldsymbol{\epsilon}_{\text{uncond}}(\mathbf{x}, t) where γ≥0\gamma \ge 0 is the guidance scale parameter (e.g., γ=2.0\gamma = 2.0). Higher values of γ\gamma prioritize fidelity to the high-yy conditioning over sample coverage, enabling generated designs to extrapolate and attain objective values exceeding the dataset maximum.

  4. Knowl 4 — Denoising Diffusion Optimization Models Algorithm

    algorithm

    DDOM operates in two phases: Phase 1 partitions the offline dataset into equal-width value bins, assigns reweighting coefficients, and trains the conditional score network. Phase 2 conditions on the maximum dataset value and generates QQ candidate designs using a second-order Heun reverse SDE solver with classifier-free guidance.

    Input: Offline dataset D={(xi,yi)}i=1n\mathcal{D} = \{(\mathbf{x}_i, y_i)\}_{i=1}^n, query budget QQ, smoothing parameter KK, temperature τ\tau, bin count NBN_B, guidance weight γ\gamma, total diffusion steps TT
    Output: Proposed candidate set X\mathbf{X} with ∣X∣≤Q|\mathbf{X}| \le Q
    // Phase 1: Training
    Partition the range $[
    \min_{(\mathbf{x}, y) \in \mathcal{D}} y, \max_{(\mathbf{x}, y) \in \mathcal{D}} y]into into N_Bequal−widthbins equal-width bins \{B_1, \dots, B_{N_B}\}$
    for each bin BiB_i do
        wi←∣Bi∣∣Bi∣+Kexp⁡(−∣max⁡(x,y)∈Dy−ybi∣τ)w_i \leftarrow \frac{|B_i|}{|B_i| + K} \exp\left(-\frac{|\max_{(\mathbf{x}, y) \in \mathcal{D}} y - y_{b_i}|}{\tau}\right)
    end for
    Initialize score network parameters θ\theta
    Train ϵθ\boldsymbol{\epsilon}_\theta by minimizing the reweighted loss using bin weights w(y)=wiw(y) = w_i and conditioning dropout
    // Phase 2: Evaluation
    ytest←max⁡(x,y)∈Dyy_{\text{test}} \leftarrow \max_{(\mathbf{x}, y) \in \mathcal{D}} y
    X←∅\mathbf{X} \leftarrow \emptyset
    for i=1i = 1 to QQ do
        Sample xT∼N(0,I)\mathbf{x}_T \sim \mathcal{N}(0, \mathbf{I})
        for t=T−1t = T-1 down to 00 do
            xt←HEUN-SAMPLER(xt+1,θ,ytest,γ)\mathbf{x}_t \leftarrow \text{HEUN-SAMPLER}(\mathbf{x}_{t+1}, \theta, y_{\text{test}}, \gamma)
        end for
        X←X∪{x0}\mathbf{X} \leftarrow \mathbf{X} \cup \{\mathbf{x}_0\}
    end for
    return X\mathbf{X}
  5. Knowl 5 — Continuous Relaxation for Discrete Black-Box Design Domains

    model/method

    To apply continuous-time diffusion modeling to discrete optimization domains (such as DNA transcription factor binding or categorical molecular descriptors) without training an auxiliary variational autoencoder, DDOM converts discrete sequences into continuous log-probability space.

    For a dd-dimensional discrete input where each position takes one of cc categorical values, each token is mapped to a one-hot vector in Rc×d\mathbb{R}^{c \times d}. Continuous logit approximations are formed by linearly interpolating between a uniform probability distribution and the one-hot representation using a fixed mixing factor α=0.6\alpha = 0.6: x~=αxone-hot+(1−α)1c1c×d\tilde{\mathbf{x}} = \alpha \mathbf{x}_{\text{one-hot}} + (1 - \alpha) \frac{1}{c} \mathbf{1}_{c \times d} This maps categorical states to continuous vectors with non-zero coordinates, enabling standard continuous VP-SDE score matching during training and continuous reverse Heun sampling during candidate generation.

  6. Knowl 6 — Design-Bench Benchmark Evaluation Results

    data/table

    DDOM was evaluated against seven baseline methods on six high-dimensional offline BBO tasks from Design-Bench using an evaluation query budget Q=256Q = 256, with scores normalized against unseen full-range dataset statistics (ynorm=y−ymin⁡ymax⁡−ymin⁡y_{\text{norm}} = \frac{y - y_{\min}}{y_{\max} - y_{\min}}) and averaged across 5 random seeds.

    Baseline TFBIND8 TFBIND10 SUPERCON. ANT D'KITTY CHEMBL Mean Score Mean Rank
    D\mathcal{D} (best) 0.439 0.467 0.399 0.565 0.884 0.605 - -
    CbAS 0.958 ±\pm 0.018 0.657 ±\pm 0.017 0.450 ±\pm 0.083 0.876 ±\pm 0.015 0.896 ±\pm 0.016 0.640 ±\pm 0.005 0.746 ±\pm 0.003 5.5
    GP-qEI 0.824 ±\pm 0.086 0.635 ±\pm 0.011 0.501 ±\pm 0.021 0.887 ±\pm 0.000 0.896 ±\pm 0.000 0.633 ±\pm 0.000 0.729 ±\pm 0.019 6.2
    CMA-ES 0.933 ±\pm 0.035 0.679 ±\pm 0.034 0.491 ±\pm 0.004 1.436 ±\pm 0.928 0.725 ±\pm 0.002 0.636 ±\pm 0.004 0.816 ±\pm 0.168 4.5
    Grad. Ascent 0.981 ±\pm 0.015 0.659 ±\pm 0.039 0.504 ±\pm 0.005 0.340 ±\pm 0.034 0.906 ±\pm 0.017 0.647 ±\pm 0.020 0.672 ±\pm 0.021 3.5
    REINFORCE 0.959 ±\pm 0.013 0.640 ±\pm 0.028 0.481 ±\pm 0.017 0.261 ±\pm 0.042 0.474 ±\pm 0.202 0.636 ±\pm 0.023 0.575 ±\pm 0.054 6.3
    MINs 0.938 ±\pm 0.047 0.659 ±\pm 0.044 0.484 ±\pm 0.017 0.942 ±\pm 0.018 0.944 ±\pm 0.009 0.653 ±\pm 0.002 0.770 ±\pm 0.023 3.5
    COMs 0.964 ±\pm 0.020 0.654 ±\pm 0.020 0.423 ±\pm 0.033 0.949 ±\pm 0.021 0.948 ±\pm 0.006 0.648 ±\pm 0.005 0.764 ±\pm 0.018 3.7
    DDOM 0.971 ±\pm 0.005 0.688 ±\pm 0.092 0.560 ±\pm 0.044 0.957 ±\pm 0.012 0.926 ±\pm 0.009 0.633 ±\pm 0.007 0.787 ±\pm 0.034 2.8

    DDOM achieves the best average rank (2.8) across all benchmarks. On Superconductor, DDOM outperforms the closest baseline by 11% (0.560 vs. 0.504). While CMA-ES achieves a higher nominal mean score (0.816), its variance is five times larger (±0.168\pm 0.168), driven by instability on Ant.

  7. Knowl 7 — Ablation on DDOM Loss Reweighting

    data/table

    To measure the isolated contribution of the bin-based loss reweighting mechanism, DDOM was trained and evaluated with and without reweighting across all six Design-Bench tasks (averaged over 5 seeds with query budget Q=256Q = 256):

    Configuration D'KITTY ANT TFBIND8 TFBIND10 SUPERCON. CHEMBL
    No reweighting 0.926 ±\pm 0.008 0.888 ±\pm 0.018 0.957 ±\pm 0.000 0.644 ±\pm 0.011 0.553 ±\pm 0.051 0.633 ±\pm 0.000
    Reweighting 0.930 ±\pm 0.003 0.960 ±\pm 0.015 0.971 ±\pm 0.005 0.688 ±\pm 0.092 0.560 ±\pm 0.044 0.633 ±\pm 0.000

    Loss reweighting improves optimization performance across continuous and discrete domains, with large gains on Ant (0.888→0.9600.888 \to 0.960) and TFBind10 (0.644→0.6880.644 \to 0.688).

  8. Knowl 8 — Extrapolation and Generalization Beyond Truncated Dataset Maxima

    empirical result

    On the 2D Branin benchmark function (fbr(x1,x2)f_{\text{br}}(x_1, x_2) with range x1∈[−5,10],x2∈[0,15]x_1 \in [-5, 10], x_2 \in [0, 15] and maximum −0.397887-0.397887), DDOM was trained on a truncated dataset where the top 10% highest-value points were removed from a uniformly sampled dataset DUnif\mathcal{D}_{\text{Unif}}. When conditioned at test time on ytest=max⁡(x,y)∈Dyy_{\text{test}} = \max_{(\mathbf{x}, y) \in \mathcal{D}} y, the conditional reverse diffusion process with classifier-free guidance generated candidate designs that systematically exceeded the maximum objective value present in the truncated training data. Both the peak candidate score and the mean score of 256 generated samples reached their highest performance when conditioned on the dataset maximum.

  9. Knowl 9 — Impact of Guidance Scale on Reverse SDE Trajectories

    empirical result

    Evaluating the objective values of generated samples along the 1000-step reverse diffusion trajectory for the Superconductor task under varying classifier-free guidance weights γ∈{0.0,4.0,10.0,20.0}\gamma \in \{0.0, 4.0, 10.0, 20.0\} shows that unguided conditional diffusion (γ=0\gamma = 0) achieves poor optimization performance (mean objective values remaining below 0.20). Increasing the guidance weight to γ≥4.0\gamma \ge 4.0 substantially accelerates the rate of objective value improvement across timesteps, enabling final generated designs to exceed mean objective values of 0.35.

  10. Knowl 10 — Sampling Latency in Diffusion-Based Inverse Optimization

    limitation

    Generating candidate designs in DDOM requires executing multi-step numerical integration (e.g., T=1000T = 1000 steps with a second-order Heun SDE solver) for each generated batch of candidates. This makes inference slower than single-step feed-forward surrogate optimizers or one-step GAN-based inverse mappings. While manageable for offline BBO with small candidate evaluation budgets (such as Q=256Q = 256), this sampling latency limits real-time optimization applications.

Coverage note — Omitted the HopperController task results and the NAS domain because Hopper was excluded by the authors due to known oracle inconsistencies in Design-Bench, and NAS was excluded due to compute limitations.

References

  1. 1.Attia, P., Grover, A., Jin, N., Severson, K., Cheong, B., Liao, J., Chen, M. H., Perkins, N., Yang, Z., Herring, P., Aykol, M., Harris, S., Braatz, R., Ermon, S., and Chueh, W. Closed-loop optimization of extreme fast charging for batteries using machine learning. Nature, 2020.
  2. 2.Bansal, A., Borgnia, E., Chu, H.-M., Li, J. S., Kazemi, H., Huang, F., Goldblum, M., Geiping, J., and Goldstein, T. Cold diffusion: Inverting arbitrary image transforms without noise. arXiv preprint arXiv:2208.09392, 2022.
  3. 3.Bansal, H. and Grover, A. Leaving reality to imagination: Robust classification via generated datasets. arXiv preprint arXiv:2302.02503, 2023.
  4. 4.Brookes, D., Park, H., and Listgarten, J. Conditioning by adaptive sampling for robust design. In International Conference on Machine Learning, 2019.
  5. 5.Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. Decision transformer: Reinforcement learning via sequence modeling. arXiv preprint arXiv:2106.01345, 2021.
  6. 6.Dhariwal, P. and Nichol, A. Q. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, 2021.
  7. 7.Dutordoir, V., Saul, A., Ghahramani, Z., and Simpson, F. Neural diffusion processes. arXiv preprint arXiv:2206.03992, 2022.
  8. 8.Fannjiang, C. and Listgarten, J. Autofocused oracles for model-based design. In Advances in Neural Information Processing Systems, volume 33, 2020.
  9. 9.Garivier, A., Kaufmann, E., and Lattimore, T. On explore-then-commit strategies. In Proceedings of the 30th International Conference on Neural Information Processing Systems, 2016.
  10. 10.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial nets. In Advances in Neural Information Processing Systems, 2014.
  11. 11.Grover, A., Markov, T., Attia, P., Jin, N., Perkins, N., Cheong, B., Chen, M., Yang, Z., Harris, S., Chueh, W., and Ermon, S. Best arm identification in multi-armed bandits with delayed feedback. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2018.
  12. 12.Guo, W., Agrawal, K. K., Grover, A., Muthukumar, V., and Pananjady, A. Learning from an exploring demonstrator: Optimal reward estimation for bandits. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2021.
  13. 13.Hansen, N. The cma evolution strategy: a comparing review. Towards a new evolutionary computation: Advances in the estimation of distribution algorithms, pp. 75–102, 2006.
  14. 14.Hansen, N. The cma evolution strategy: A tutorial. arXiv preprint arXiv:1604.00772, 2016.
  15. 15.Ho, J. and Salimans, T. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021.
  16. 16.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, 2020.
  17. 17.Ho, J., Salimans, T., Gritsenko, A. A., Chan, W., Norouzi, M., and Fleet, D. J. Video diffusion models. In Advances in Neural Information Processing Systems, 2022.
  18. 18.Hoogeboom, E. and Salimans, T. Blurring diffusion models. arXiv preprint arXiv:2209.05557, 2022.
  19. 19.Huang, C.-W., Lim, J. H., and Courville, A. C. A variational perspective on diffusion-based generative models and score matching. Advances in Neural Information Processing Systems, 34:22863–22876, 2021.
  20. 20.Janner, M., Du, Y., Tenenbaum, J. B., and Levine, S. Planning with diffusion for flexible behavior synthesis. arXiv preprint arXiv:2205.09991, 2022.
  21. 21.Jeong, M., Kim, H., Cheon, S. J., Choi, B. J., and Kim, N. S. Diff-tts: A denoising diffusion model for text-to-speech. arXiv preprint arXiv:2104.01409, 2021.
  22. 22.Joachims, T., Swaminathan, A., and de Rijke, M. Deep learning with logged bandit feedback. In International Conference on Learning Representations, 2018.
  23. 23.Jolicoeur-Martineau, A., Li, K., Piché-Taillefer, R., Kachman, T., and Mitliagkas, I. Gotta go fast when generating data with score-based models. arXiv preprint arXiv:2105.14080, 2021.
  24. 24.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the design space of diffusion-based generative models. arXiv preprint arXiv:2206.00364, 2022.
  25. 25.Kim, H., Kim, S., and Yoon, S. Guided-tts: A diffusion model for text-to-speech via classifier guidance. In International Conference on Machine Learning. PMLR, 2022.
  26. 26.Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B. Diffwave: A versatile diffusion model for audio synthesis. In International Conference on Learning Representations, 2021.
  27. 27.Krishnamoorthy, S., Mashkaria, S. M., and Grover, A. Generative pretraining for black-box optimization. In International Conference on Machine Learning (ICML), 2023.
  28. 28.Kumar, A. and Levine, S. Model inversion networks for model-based optimization. In Advances in Neural Information Processing Systems, 2020.
  29. 29.Liu, F., Liu, H., Grover, A., and Abbeel, P. Masked autoencoding for scalable and generalizable decision making. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  30. 30.Mirza, M. and Osindero, S. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
  31. 31.Nguyen, T. and Grover, A. Transformer neural processes: Uncertainty-aware meta learning via sequence modeling. In International Conference on Machine Learning (ICML), 2022.
  32. 32.Nguyen, T., Zheng, Q., and Grover, A. Reliable conditioning of behavioral cloning for offline reinforcement learning. arXiv preprint arXiv:2210.05158, 2022.
  33. 33.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. In International Conference on Machine Learning, 2021.
  34. 34.Riquelme, C., Tucker, G., and Snoek, J. Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling. arXiv preprint arXiv:1802.09127, 2018.
  35. 35.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 35: 36479–36494, 2022.
  36. 36.Shahriari, B., Swersky, K., Wang, Z., Adams, R. P., and de Freitas, N. Taking the human out of the loop: A review of bayesian optimization. Proceedings of the IEEE, 104 (1):148–175, 2016. doi: 10.1109/JPROC.2015.2494218.
  37. 37.Snoek, J., Larochelle, H., and Adams, R. P. Practical bayesian optimization of machine learning algorithms. Neural information processing systems, 2012.
  38. 38.Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, 2015.
  39. 39.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.
  40. 40.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021.
  41. 41.Srinivas, N., Krause, A., Kakade, S., and Seeger, M. W. Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design. In Proceedings of the 27th International Conference on Machine Learning, pp. 1015–1022. Omnipress, 2010.
  42. 42.Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y. Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems, 1999.
  43. 43.Swaminathan, A. and Joachims, T. Batch learning from logged bandit feedback through counterfactual risk minimization. Journal of Machine Learning Research, 16 (52):1731–1755, 2015.
  44. 44.Trabucco, B., Kumar, A., Geng, X., and Levine, S. Conservative objective models for effective offline model-based optimization. In International Conference on Machine Learning, 2021.
  45. 45.Trabucco, B., Geng, X., Kumar, A., and Levine, S. Design-bench: Benchmarks for data-driven offline model-based optimization. In International Conference on Machine Learning, pp. 21658–21676. PMLR, 2022.
  46. 46.Vincent, P. A connection between score matching and denoising autoencoders. Neural computation, 23(7):1661–1674, 2011.
  47. 47.Zheng, Q., Zhang, A., and Grover, A. Online decision transformer. In International Conference on Machine Learning (ICML), pp. 27042–27059, 2022.
  48. 48.Zheng, Q., Henaff, M., Amos, B., and Grover, A. Semi-supervised offline reinforcement learning with action-free trajectories. In International Conference on Machine Learning (ICML), 2023.
  49. 49.Zhu, B., Dang, M., and Grover, A. Scaling pareto-efficient decision making via offline multi-objective rl. In International Conference on Learning Representations (ICLR), 2023.

Citation

MLA
Krishnamoorthy, S., et al. “Diffusion Models for Black-Box Optimization”. International Conference on Machine Learning, vol. 202, 2023, pp. 17842–57, https://proceedings.mlr.press/v202/krishnamoorthy23a.html.
APA
Krishnamoorthy, S., Mashkaria, S. M., & Grover, A. (2023). Diffusion Models for Black-Box Optimization. International Conference on Machine Learning, 202, 17842–17857. https://proceedings.mlr.press/v202/krishnamoorthy23a.html
Chicago
Krishnamoorthy, S., S. M. Mashkaria, and A. Grover. 2023. “Diffusion Models for Black-Box Optimization”. International Conference on Machine Learning 202: 17842–57. https://proceedings.mlr.press/v202/krishnamoorthy23a.html.
Harvard
Krishnamoorthy, S., Mashkaria, S.M. and Grover, A. (2023) “Diffusion Models for Black-Box Optimization”, International Conference on Machine Learning. PMLR, pp. 17842–17857. Available at: https://proceedings.mlr.press/v202/krishnamoorthy23a.html.
Vancouver
1. Krishnamoorthy S, Mashkaria SM, Grover A (2023) Diffusion Models for Black-Box Optimization. In: International Conference on Machine Learning. PMLR, pp 17842–17857

BibTeX

@InProceedings{pmlr-v202-krishnamoorthy23a,
  title = 	 {Diffusion Models for Black-Box Optimization},
  author =       {Krishnamoorthy, Siddarth and Mashkaria, Satvik Mehul and Grover, Aditya},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {17842--17857},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/krishnamoorthy23a/krishnamoorthy23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/krishnamoorthy23a.html},
  abstract = 	 {The goal of offline black-box optimization (BBO) is to optimize an expensive black-box function using a fixed dataset of function evaluations. Prior works consider forward approaches that learn surrogates to the black-box function and inverse approaches that directly map function values to corresponding points in the input domain of the black-box function. These approaches are limited by the quality of the offline dataset and the difficulty in learning one-to-many mappings in high dimensions, respectively. We propose Denoising Diffusion Optimization Models (DDOM), a new inverse approach for offline black-box optimization based on diffusion models. Given an offline dataset, DDOM learns a conditional generative model over the domain of the black-box function conditioned on the function values. We investigate several design choices in DDOM, such as reweighting the dataset to focus on high function values and the use of classifier-free guidance at test-time to enable generalization to function values that can even exceed the dataset maxima. Empirically, we conduct experiments on the Design-Bench benchmark (Trabucco et al., 2022) and show that DDOM achieves results competitive with state-of-the-art baselines.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/