Deep Equilibrium Approaches to Diffusion Models

Ashwini PokleZhengyang GengJ. Zico Kolter

article2022NeurIPS62 citations

Formulates the diffusion sampling chain as a joint fixed-point system using deep equilibrium models, enabling parallel multi-GPU image generation and memory-efficient backpropagation for faster model inversion.

Listen

Diffusion-based generative models produce exceptionally high-quality images, often surpassing alternative methods in visual fidelity. However, their practical deployment in real-world workflows like image editing, restoration, and real-time generation is heavily restricted by slow, sequential generation steps. Standard diffusion models iteratively apply hundreds or thousands of serial denoising steps to convert random noise into an image. This serial structure fails to fully utilize modern multi-processor hardware and makes model inversion—the optimization process of discovering the specific starting noise that reproduces an existing image—computationally prohibitive and memory-intensive.

The article demonstrates that the sampling process of diffusion models can be reformulated as a joint, multi-variate fixed-point system using deep equilibrium frameworks. This approach evaluates whether modeling the entire generation chain simultaneously enables parallelized image generation and constant-memory optimization for model inversion, spanning deterministic and stochastic diffusion variants.

To achieve this, the authors restructured the generation chain so that all intermediate states are estimated together within a unified equilibrium state, solved using numerical fixed-point algorithms such as Anderson acceleration. This design replaces serial dependencies with parallel computations across graphics processing units. For model inversion, the formulation uses single-step damped implicit differentiation, enabling gradient updates without storing the complete generation trajectory in memory. The method was evaluated using standard benchmark image datasets across varying resolutions, including CIFAR-10, CelebA, and LSUN Bedroom and Church scenes.

The findings show substantial improvements in efficiency and inversion accuracy. For single-image generation on lower-resolution datasets, the equilibrium approach achieved up to a two-fold wall-clock speedup over sequential baselines—generating CIFAR-10 images in approximately 2.91 seconds compared to 20.16 seconds—while maintaining comparable or slightly improved image quality scores. In model inversion tasks, the equilibrium framework consistently converged in fewer iterations and attained dramatically lower reconstruction error across all datasets. For example, on 100-step CIFAR-10 inversion, the average error dropped from 15.74 in the baseline to 0.76, while reducing total inversion time from roughly 49 minutes to 13 minutes and capturing finer visual textures such as foliage and facial details.

These results provide a practical path to reducing computational bottlenecks and memory overheads in production-grade generative modeling. By decoupling generation and optimization from strictly serial chains, organizations can perform image manipulation and latent inversion tasks at a lower computational cost without requiring specialized differential equation solvers or excessive hardware memory. The method is orthogonal to other acceleration techniques, meaning it can readily integrate with existing model distillation or step-reduction strategies.

Engineering and research teams utilizing diffusion pipelines should consider adopting equilibrium solvers for workflows that require single-instance generation or frequent latent space inversions. Before broad deployment, teams should conduct pilot benchmarks on their target hardware and image resolutions. The authors note that the speed advantage diminishes on high-resolution images with very short diffusion chains and during large-batch processing, where memory scaling requirements across all time states can offset parallelization gains. However, within single-instance generation and model inversion workflows, the findings demonstrate a high level of empirical reliability and consistent performance gains.

arXiv: 2210.12867
Cover for Deep Equilibrium Approaches to Diffusion Models

Abstract

Diffusion-based generative models are extremely effective in generating high-quality images, with generated samples often surpassing the quality of those produced by other models under several metrics. One distinguishing feature of these models, however, is that they typically require long sampling chains to produce high-fidelity images. This presents a challenge not only from the lenses of sampling time, but also from the inherent difficulty in backpropagating through these chains in order to accomplish tasks such as model inversion, i.e., approximately finding latent states that generate known images. In this paper, we look at diffusion models through a different perspective, that of a (deep) equilibrium (DEQ) fixed point model. Specifically, we extend the recent denoising diffusion implicit model (DDIM) [68], and model the entire sampling chain as a joint, multi-variate fixed point system. This setup provides an elegant unification of diffusion and equilibrium models, and shows benefits in 1) single image sampling, as it replaces the fully-serial typical sampling process with a parallel one; and 2) model inversion, where we can leverage fast gradients in the DEQ setting to much more quickly find the noise that generates a given image. The approach is also orthogonal and thus complementary to other methods used to reduce the sampling time, or improve model inversion. We demonstrate our method’s strong performance across several datasets, including CIFAR10, CelebA, and LSUN Bedroom and Churches.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 A Deep Equilibrium Approach to DDIMs
  • 3.1 A DEQ formulation of DDIMs (DEQ-DDIM)
  • 4 Efficient Inversion of DDIM
  • 4.1 Problem Setup
  • 4.2 Inverting DDIM: The Naive Approach
  • 4.3 Efficient Inversion of DDIM with DEQs
  • 5 Experiments
  • 5.1 Convergence of DEQ-DDIM
  • 5.2 Sample quality of images generated with DEQ
  • 5.3 Model Inversion of DDIM with DEQs
  • 6 Related Work
  • 7 Conclusion
  • 8 Acknowledgements
  • References
  • Checklist

Knowls

  1. Knowl 1 — Joint multi-timestep equilibrium formulation of DDIM

    model/method

    DEQ-DDIM represents the complete deterministic DDIM sampling chain as one fixed-point problem over the latent images x0:T−1x_{0:T-1}, rather than applying one denoising update at a time. Let TT be the number of diffusion steps, let eta_t\in(0,1) be the variance schedule, define α0=1\alpha_0=1 and αt=∏s=1t(1−βs)\alpha_t=\prod_{s=1}^{t}(1-\beta_s), and let ϵθ(xt,t)\epsilon_\theta(x_t,t) be the trained noise-prediction network. For deterministic DDIM sampling, define

    ct=1−αt−1−αt−1(1−αt)αt.c_t=\sqrt{1-\alpha_{t-1}}-\sqrt{\frac{\alpha_{t-1}(1-\alpha_t)}{\alpha_t}}.

    The entire chain can then be written, for k∈{0,…,T}k\in\{0,\ldots,T\}, as

    xT−k=αT−kαTxT+∑t=T−kT−1αT−kαt ct+1 ϵθ(xt+1,t+1),x_{T-k}=\sqrt{\frac{\alpha_{T-k}}{\alpha_T}}x_T+\sum_{t=T-k}^{T-1}\sqrt{\frac{\alpha_{T-k}}{\alpha_t}}\,c_{t+1}\,\epsilon_\theta(x_{t+1},t+1),

    where xtx_t is an image-shaped latent tensor and the sum is empty when k=0k=0. Thus, the value of each xtx_t depends on all later states xt+1:Tx_{t+1:T}, not only on xt+1x_{t+1}. Let h~(x0:T−1;xT)\widetilde h(x_{0:T-1};x_T) denote the vector-valued function that evaluates these equations for every timestep simultaneously, with xT∼N(0,I)x_T\sim\mathcal N(0,I) treated as an input injection. DEQ-DDIM solves

    g(x0:T−1;xT)=h~(x0:T−1;xT)−x0:T−1=0g(x_{0:T-1};x_T)=\widetilde h(x_{0:T-1};x_T)-x_{0:T-1}=0

    using a black-box root solver, producing the equilibrium states x0:T∗x^*_{0:T}.

  2. Knowl 2 — Parallel augmented reverse process

    model/method

    The DEQ-DDIM fixed point is an augmented reverse process whose update for xtx_t uses the denoising predictions at every subsequent state xt+1:Tx_{t+1:T}. At its solution, this upper-triangular system exactly represents the deterministic DDIM chain, while allowing all timestep updates to be evaluated in a batched, parallel computation. Anderson acceleration is used to solve the joint equilibrium by repeatedly updating the whole latent-state tensor; the timestep dimension can be distributed across multiple GPUs. This reduces the serial dependence of ordinary DDIM sampling and is particularly useful for batch-size-one generation, although each equilibrium iteration is more computationally expensive and all TT latent images must be stored simultaneously.

  3. Knowl 3 — Stochastic DEQ extension and CIFAR-10 results

    data/table

    The DEQ construction extends to stochastic DDIM and DDPM-like samplers by drawing one Gaussian noise tensor for each diffusion step and treating the collection as an additional fixed input. If η>0\eta>0 controls the stochasticity and ξ1:T\xi_{1:T} are independent standard-Gaussian noise tensors, the stochastic equilibrium is

    x0:T∗=RootSolver⁡ ⁣(g(x0:T−1;xT,ξ1:T)),xT∼N(0,I),ξt∼N(0,I).x^*_{0:T}=\operatorname{RootSolver}\!\left(g(x_{0:T-1};x_T,\xi_{1:T})\right),\qquad x_T\sim\mathcal N(0,I),\quad \xi_t\sim\mathcal N(0,I).

    On CIFAR-10, the stochastic DEQ sampler used Anderson acceleration and was compared with sequential DDIM for T∈{20,50}T\in\{20,50\} over three values of η\eta. FID was computed on generated images, with lower values better; time is the average generation time in seconds.

    η\eta TT FID Time (s)
    DDIM DEQ-sDDIM DDIM DEQ-sDDIM
    0.2 20 7.19 6.99 0.33 0.51
    0.5 20 8.35 8.22 0.35 0.51
    1 20 18.37 17.72 0.34 0.93
    0.2 50 4.69 4.44 0.88 0.88
    0.5 50 5.26 4.99 0.83 1.00
    1 50 8.02 7.85 0.83 1.58

    DEQ-sDDIM consistently achieved comparable or slightly better FID than DDIM, but became slower as η\eta increased because more equilibrium iterations were needed. For comparison, the paper reports DDPM FIDs of 133.37133.37 for T=20T=20 and 32.7232.72 for T=50T=50 under the larger-variance setting reported by the DDIM authors.

  4. Knowl 4 — DEQ-based DDIM inversion

    algorithm

    Given a target image x0x_0, DEQ-DDIM inversion optimizes an initial latent xTx_T so that the equilibrium-generated image x0∗x^*_0 matches the target under the squared Frobenius loss

    L(x0,x0∗)=∥x0−x0∗∥F2.L(x_0,x^*_0)=\left\|x_0-x^*_0\right\|_F^2.

    The exact implicit gradient through the equilibrium, for any latent variable ϕ\phi being differentiated and especially for the optimized noise state ϕ=xT\phi=x_T, is

    ∂L∂ϕ=−∂L∂x0:T∗(Jg−1∣x0:T∗)∂h~(x0:T−1∗;xT)∂ϕ,\frac{\partial L}{\partial \phi}=-\frac{\partial L}{\partial x^*_{0:T}}\left(J_g^{-1}\big|_{x^*_{0:T}}\right)\frac{\partial\widetilde h(x^*_{0:T-1};x_T)}{\partial\phi},

    where Jg−1∣x0:T∗J_g^{-1}\big|_{x^*_{0:T}} is the inverse Jacobian of the fixed-point residual with respect to the equilibrium latent states. Because forming this inverse is impractical for high-dimensional image states, DEQ-DDIM uses the approximation

    ∂L∂ϕ≈−∂L∂x0:T∗M∂h~(x0:T−1∗;xT)∂ϕ,\frac{\partial L}{\partial \phi}\approx-\frac{\partial L}{\partial x^*_{0:T}}M\frac{\partial\widetilde h(x^*_{0:T-1};x_T)}{\partial\phi},

    where MM approximates the inverse Jacobian. The paper sets M=IM=I (the one-step gradient) and uses a damping factor τ=0.1\tau=0.1 in experiments. After solving the equilibrium with gradients disabled, the differentiable backward surrogate is

    xˉ0:T∗=τ h~(x0:T−1∗;xT∗)+(1−τ)x0:T∗,\bar x^*_{0:T}=\tau\,\widetilde h(x^*_{0:T-1};x^*_T)+(1-\tau)x^*_{0:T},

    through which standard automatic differentiation computes the approximate gradient. The resulting inversion procedure is:

    Input: target image x0x_0, trained noise predictor ϵθ\epsilon_\theta, timestep count TT, epoch count NN, root solver, damping factor τ=0.1\tau=0.1
    Initialize x^0:T\hat x_{0:T} with Gaussian noise
    for epoch from 1 to NN
        Disable gradient computation
        x0:T∗x^*_{0:T} = RootSolver(g(x^0:T−1;x^T)g(\hat x_{0:T-1};\hat x_T))
        Enable gradient computation
        Construct xˉ0:T∗=τh~(x0:T−1∗;xT∗)+(1−τ)x0:T∗\bar x^*_{0:T}=\tau\widetilde h(x^*_{0:T-1};x^*_T)+(1-\tau)x^*_{0:T}
        Compute L=∥x0−xˉ0∗∥F2L=\|x_0-\bar x^*_0\|_F^2
        Compute the one-step gradient of LL with respect to x^T\hat x_T
        Update x^T\hat x_T with a gradient-descent or optimizer step
    end for
    Output: recovered latent xT∗x^*_T and reconstructed image x0∗x^*_0

    Unlike naive inversion, which sequentially runs all TT denoising steps and stores the full computational graph for backpropagation, the DEQ backward pass has constant memory complexity with respect to the effective sampling depth and requires only one approximate backward step.

  5. Knowl 5 — Experimental protocol

    experimental setup

    Experiments covered CIFAR-10 at 32×3232\times32 resolution, CelebA at 64×6464\times64, LSUN Bedroom at 256×256256\times256, and LSUN Outdoor Church at 256×256256\times256. The experiments used pretrained diffusion models from Ho et al. for CIFAR-10, LSUN Bedroom, and LSUN Outdoor Church, and from Song et al. for CelebA. Anderson acceleration was the default fixed-point solver. All experiments ran on NVIDIA RTX A6000 GPUs, and the inversion backward pass used the damped one-step gradient with τ=0.1\tau=0.1. Deterministic DDIM generation used at most 15 Anderson iterations per image; stochastic DEQ-sDDIM generation used at most 50. FID was evaluated on 50,000 generated images, generation time was averaged over 500 images, and inversion results were averaged over 100 target images.

  6. Knowl 6 — Observed fixed-point convergence

    empirical result

    DEQ-DDIM reached useful equilibrium solutions for diffusion chains much longer than the number of Anderson iterations used at inference. In experiments on CIFAR-10 and CelebA with several chain lengths, high-quality images were obtained in as few as 15 Anderson steps even when the underlying diffusion model had been trained with hundreds or thousands of timesteps. Shorter chains reached simultaneous equilibrium more easily than longer chains. For larger TT, the full latent-state residual could approach a limit cycle, but the final few denoising states—including the generated image x0x_0—still converged sufficiently well for image generation. The residual could be reduced further with stronger root solvers such as Broyden's method.

  7. Knowl 7 — Deterministic single-image generation performance

    data/table

    The deterministic DEQ-DDIM sampler was compared with DDPM and sequential DDIM for single-image generation. The comparison reports FID, where lower is better, and average GPU-inclusive time per image. DEQ-DDIM preserved or improved FID relative to DDIM on all four datasets and produced its largest timing gains on the smaller-resolution datasets.

    Dataset TT DDPM DDIM DEQ-DDIM
    FID Time FID Time FID Time
    CIFAR10 1000 3.17 24.45s 4.07 20.16s 3.79 2.91s
    CelebA 500 5.32 14.95s 3.66 10.31s 2.92 5.12s
    LSUN Bedroom 25 184.05 1.72s 8.76 1.19s 8.73 3.82s
    LSUN Church 25 122.18 1.77s 13.44 1.68s 13.55 3.99s

    For CIFAR-10 and CelebA, DEQ-DDIM was substantially faster than sequential DDIM while attaining FIDs of 3.793.79 and 2.922.92, respectively. For the 256×256256\times256 LSUN datasets, DEQ-DDIM was slower than sequential DDIM because the expensive equilibrium updates outweighed the benefit of parallelizing a short chain.

  8. Knowl 8 — Model inversion accuracy and speed

    data/table

    DEQ-DDIM inversion was compared with a sequential baseline that repeatedly generates an image through all DDIM steps and backpropagates through the resulting chain. The metric is the minimum squared Frobenius reconstruction loss, and the time is the average minutes required to generate a reconstruction; both are reported with mean ±\pm variation over 100 target images. Lower values are better for both metrics.

    Dataset TT Baseline DEQ-DDIM
    Min loss Avg Time (mins) Min loss Avg Time (mins)
    CIFAR10 100 15.74 ±\pm 8.7 49.07 ±\pm 1.76 0.76 ±\pm 0.35 12.99 ±\pm 0.97
    CIFAR10 10 2.59 ±\pm 3.67 14.36 ±\pm 0.26 0.68 ±\pm 0.32 2.54 ±\pm 0.41
    CelebA 20 14.13 ±\pm 5.04 30.09 ±\pm 0.57 1.03 ±\pm 0.37 28.09 ±\pm 1.76
    Bedroom 10 1114.49 ±\pm 795.86 26.41 ±\pm 0.17 36.37 ±\pm 22.86 33.7 ±\pm 1.05
    Church 10 1674.68 ±\pm 1432.54 29.7 ±\pm 0.75 47.94 ±\pm 24.78 33.54 ±\pm 3.02

    DEQ-DDIM achieved substantially lower reconstruction losses on every dataset and converged in fewer optimization epochs. Reconstructions from the DEQ latent states retained finer image details, including foliage textures and surface crevices, than reconstructions obtained by the sequential baseline.

  9. Knowl 9 — Operational limitations of DEQ diffusion sampling

    limitation

    DEQ-DDIM does not uniformly improve wall-clock sampling. Its joint state requires storing all TT latent images, so memory usage grows with the number of diffusion states; for large inference batches, sequential sampling can be faster because DEQ processing effectively handles a batch of size BTBT rather than BB. On short diffusion chains and high-resolution images, the number of Anderson iterations needed for equilibrium becomes comparable to TT, while each DEQ iteration is more expensive, eliminating the speed advantage. Increasing stochasticity also requires more solver iterations and increases generation time. Finally, the paper leaves further acceleration through jointly training a DEQ noise-prediction model and latent variables as future work.

Coverage note — No substantial contributed component was omitted; background diffusion-model training objectives and related work were excluded because they are not part of the paper's own contribution.

References

  1. 1.Rameen Abdal, Yipeng Qin, and Peter Wonka. Image2stylegan: How to embed images into the stylegan latent space? In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4432–4441, 2019.
  2. 2.Rameen Abdal, Peihao Zhu, Niloy J Mitra, and Peter Wonka. Styleflow: Attribute-conditioned exploration of stylegan-generated images using conditional continuous normalizing flows. ACM Transactions on Graphics (TOG), 40(3):1–21, 2021.
  3. 3.Brandon Amos. Tutorial on amortized optimization for learning to optimize over continuous domains. arXiv preprint arXiv:2202.00665, 2022.
  4. 4.Brandon Amos and J. Zico Kolter. OptNet: Differentiable optimization as a layer in neural networks. In International Conference on Machine Learning (ICML), 2017.
  5. 5.Donald G Anderson. Iterative procedures for nonlinear integral equations. Journal of the ACM (JACM), 1965.
  6. 6.Shaojie Bai, J Zico Kolter, and Vladlen Koltun. Deep equilibrium models. Neural Information Processing Systems (NeurIPS), 2019.
  7. 7.Shaojie Bai, Vladlen Koltun, and J Zico Kolter. Multiscale deep equilibrium models. Neural Information Processing Systems (NeurIPS), 2020.
  8. 8.Shaojie Bai, Vladlen Koltun, and J Zico Kolter. Stabilizing equilibrium models by jacobian regularization. arXiv preprint arXiv:2106.14342, 2021.
  9. 9.Shaojie Bai, Zhengyang Geng, Yash Savani, and J Zico Kolter. Deep equilibrium optical flow estimation. arXiv preprint arXiv:2204.08442, 2022.
  10. 10.Shaojie Bai, Vladlen Koltun, and J Zico Kolter. Neural deep equilibrium solvers. In International Conference on Learning Representations, 2022.
  11. 11.David Bau, Hendrik Strobelt, William Peebles, Jonas Wulff, Bolei Zhou, Jun-Yan Zhu, and Antonio Torralba. Semantic photo manipulation with a generative image prior. arXiv preprint arXiv:2005.07727, 2020.
  12. 12.Ashish Bora, Ajil Jalal, Eric Price, and Alexandros G Dimakis. Compressed sensing using generative models. In International Conference on Machine Learning, pages 537–546. PMLR, 2017.
  13. 13.Charles G Broyden. A class of methods for solving nonlinear simultaneous equations. Mathematics of computation, 1965.
  14. 14.Kelvin CK Chan, Xintao Wang, Xiangyu Xu, Jinwei Gu, and Chen Change Loy. Glean: Generative latent bank for large-factor image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14245–14254, 2021.
  15. 15.Qi Chen, Yifei Wang, Yisen Wang, Jiansheng Yang, and Zhouchen Lin. Optimization-induced graph implicit nonlinear diffusion. In International Conference on Machine Learning, pages 3648–3661. PMLR, 2022.
  16. 16.Tian Qi Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud. Neural ordinary differential equations. In Neural Information Processing Systems (NeurIPS), 2018.
  17. 17.Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon. Ilvr: Conditioning method for denoising diffusion probabilistic models. arXiv preprint arXiv:2108.02938, 2021.
  18. 18.Hyungjin Chung, Byeongsu Sim, and Jong Chul Ye. Come-closer-diffuse-faster: Accelerating conditional diffusion models for inverse problems through stochastic contraction. arXiv preprint arXiv:2112.05146, 2021.
  19. 19.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34, 2021.
  20. 20.Josip Djolonga and Andreas Krause. Differentiable learning of submodular models. Advances in Neural Information Processing Systems, 30, 2017.
  21. 21.Priya L. Donti, David Rolnick, and J Zico Kolter. DC3: A learning method for optimization with hard constraints. In International Conference on Learning Representations (ICLR), 2021.
  22. 22.Emilien Dupont, Arnaud Doucet, and Yee Whye Teh. Augmented neural ODEs. In Neural Information Processing Systems (NeurIPS), 2019.
  23. 23.Laurent El Ghaoui, Fangda Gu, Bertrand Travacca, and Armin Askari. Implicit deep learning. arXiv:1908.06315, 2019.
  24. 24.Thorsten Falk, Dominic Mai, Robert Bensch, Özgün Çiçek, Ahmed Abdulkadir, Yassine Marrakchi, Anton Böhm, Jan Deubner, Zoe Jäckel, Katharina Seiwald, et al. U-net: deep learning for cell counting, detection, and morphometry. Nature methods, 16(1):67–70, 2019.
  25. 25.Zhili Feng and J Zico Kolter. On the neural tangent kernel of equilibrium models, 2021.
  26. 26.Samy Wu Fung, Howard Heaton, Qiuwei Li, Daniel McKenzie, Stanley Osher, and Wotao Yin. Fixed point networks: Implicit depth models with jacobian-free backprop. arXiv e-prints, pages arXiv–2103, 2021.
  27. 27.Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li, Ke Wei, and Zhouchen Lin. Is attention better than matrix decomposition? In International Conference on Learning Representations (ICLR), 2021.
  28. 28.Zhengyang Geng, Xin-Yu Zhang, Shaojie Bai, Yisen Wang, and Zhouchen Lin. On training implicit models. Neural Information Processing Systems (NeurIPS), 2021.
  29. 29.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in neural information processing systems, 27, 2014.
  30. 30.Albert Gu, Karan Goel, and Christopher Re. Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations (ICLR), 2022.
  31. 31.Fangda Gu, Heng Chang, Wenwu Zhu, Somayeh Sojoudi, and Laurent El Ghaoui. Implicit Graph Neural Networks. In Neural Information Processing Systems (NeurIPS), pages 11984–11995, 2020.
  32. 32.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  33. 33.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Neural Information Processing Systems (NeurIPS), 2020.
  34. 34.Minyoung Huh, Richard Zhang, Jun-Yan Zhu, Sylvain Paris, and Aaron Hertzmann. Transforming and projecting images into class-conditional generative networks. In European Conference on Computer Vision, pages 17–34. Springer, 2020.
  35. 35.Thibaut Issenhuth, Ugo Tanielian, Jérémie Mary, and David Picard. Edibert, a generative model for image editing. arXiv preprint arXiv:2111.15264, 2021.
  36. 36.Ajil Jalal, Marius Arvinte, Giannis Daras, Eric Price, Alexandros G Dimakis, and Jon Tamir. Robust compressed sensing mri with deep generative priors. Advances in Neural Information Processing Systems, 34:14938–14954, 2021.
  37. 37.Zahra Kadkhodaie and Eero P Simoncelli. Solving linear inverse problems using the prior implicit in a denoiser. arXiv preprint arXiv:2007.13640, 2020.
  38. 38.Kenji Kawaguchi. On the Theory of Implicit Deep Learning: Global Convergence with Implicit Layers. In International Conference on Learning Representations (ICLR), 2020.
  39. 39.Bahjat Kawar, Gregory Vaksman, and Michael Elad. Snips: Solving noisy inverse problems stochastically. Advances in Neural Information Processing Systems, 34:21757–21769, 2021.
  40. 40.Gwanghyun Kim and Jong Chul Ye. Diffusionclip: Text-guided image manipulation using diffusion models. arXiv preprint arXiv:2110.02711, 2021.
  41. 41.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  42. 42.Diederik P Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. arXiv preprint arXiv:2107.00630, 2021.
  43. 43.J. Zico Kolter, David Duvenaud, and Matthew Johnson. Deep implicit layers tutorial - neural ODEs, deep equilibirum models, and beyond. Neural Information Processing Systems Tutorial, 2020.
  44. 44.Zhifeng Kong and Wei Ping. On fast sampling of diffusion probabilistic models. arXiv preprint arXiv:2106.00132, 2021.
  45. 45.Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. Diffwave: A versatile diffusion model for audio synthesis. In International Conference on Learning Representations (ICLR), 2021.
  46. 46.Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, Citeseer, 2009.
  47. 47.Haoying Li, Yifan Yang, Meng Chang, Shiqi Chen, Huajun Feng, Zhihai Xu, Qi Li, and Yueting Chen. Srdiff: Single image super-resolution with diffusion probabilistic models. Neurocomputing, 2022.
  48. 48.Mingjie Li, Yisen Wang, and Zhouchen Lin. Cerdeq: Certifiable deep equilibrium model. In International Conference on Machine Learning, 2022.
  49. 49.Huan Ling, Karsten Kreis, Daiqing Li, Seung Wook Kim, Antonio Torralba, and Sanja Fidler. Editgan: High-precision semantic image editing. Advances in Neural Information Processing Systems, 34:16331–16345, 2021.
  50. 50.Zenan Ling, Xingyu Xie, Qiuhao Wang, Zongpeng Zhang, and Zhouchen Lin. Global convergence of over-parameterized deep equilibrium models. arXiv preprint arXiv:2205.13814, 2022.
  51. 51.Juncheng Liu, Kenji Kawaguchi, Bryan Hooi, Yiwei Wang, and Xiaokui Xiao. EIGNN: Efficient infinite-depth graph neural networks. In Neural Information Processing Systems (NeurIPS), 2021.
  52. 52.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV), December 2015.
  53. 53.Cheng Lu, Jianfei Chen, Chongxuan Li, Qiuhao Wang, and Jun Zhu. Implicit normalizing flows. In International Conference on Learning Representations (ICLR), 2021.
  54. 54.Eric Luhman and Troy Luhman. Knowledge distillation in iterative generative models for improved sampling speed. ArXiv, abs/2101.02388, 2021.
  55. 55.Chenlin Meng, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. Sdedit: Image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073, 2021.
  56. 56.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  57. 57.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pages 8162–8171. PMLR, 2021.
  58. 58.Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, and Anima Anandkumar. Diffusion models for adversarial purification. arXiv preprint arXiv:2205.07460, 2022.
  59. 59.Junyoung Park, Jinhyun Choo, and Jinkyoo Park. Convergent graph solvers. arXiv preprint arXiv:2106.01680, 2021.
  60. 60.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017.
  61. 61.Guim Perarnau, Joost Van De Weijer, Bogdan Raducanu, and Jose M Álvarez. Invertible conditional gans for image editing. arXiv preprint arXiv:1611.06355, 2016.
  62. 62.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  63. 63.Ardavan Saeedi, Matthew Hoffman, Stephen DiVerdi, Asma Ghandeharioun, Matthew Johnson, and Ryan Adams. Multimodal prediction and personalization of photo edits with deep generative models. In International Conference on Artificial Intelligence and Statistics, pages 1309–1317. PMLR, 2018.
  64. 64.Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David L Fleet, and Mohammad Norouzi. Image super-resolution via iterative refinement. arXiv preprint arXiv:2104.07636, 2021.
  65. 65.Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations (ICLR), 2022.
  66. 66.Hiroshi Sasaki, Chris G Willcocks, and Toby P Breckon. Unit-ddpm: Unpaired image translation with denoising diffusion probabilistic models. arXiv preprint arXiv:2104.05358, 2021.
  67. 67.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning (ICML), 2015.
  68. 68.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.
  69. 69.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. Advances in Neural Information Processing Systems, 32, 2019.
  70. 70.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations (ICLR), 2021.
  71. 71.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations (ICLR), 2021.
  72. 72.Po-Wei Wang, Priya Donti, Bryan Wilder, and Zico Kolter. Satnet: Bridging deep learning and logical reasoning using a differentiable satisfiability solver. In International Conference on Machine Learning (ICML), 2019.
  73. 73.Tiancai Wang, Xiangyu Zhang, and Jian Sun. Implicit Feature Pyramid Network for Object Detection. arXiv preprint arXiv:2012.13563, 2020.
  74. 74.Colin Wei and J Zico Kolter. Certified robustness for deep equilibrium models via interval bound propagation. In International Conference on Learning Representations, 2022.
  75. 75.Ezra Winston and J. Zico Kolter. Monotone operator equilibrium networks. In Neural Information Processing Systems (NeurIPS), 2020.
  76. 76.Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao. LSUN: construction of a large-scale image dataset using deep learning with humans in the loop. CoRR, abs/1506.03365, 2015.
  77. 77.Jiapeng Zhu, Yujun Shen, Deli Zhao, and Bolei Zhou. In-domain gan inversion for real image editing. In European conference on computer vision, pages 592–608. Springer, 2020.
  78. 78.Jun-Yan Zhu, Philipp Krähenbühl, Eli Shechtman, and Alexei A Efros. Generative visual manipulation on the natural image manifold. In European conference on computer vision, pages 597–613. Springer, 2016.

Citation

MLA
Pokle, A., et al. “Deep Equilibrium Approaches to Diffusion Models”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 37975–90, https://proceedings.neurips.cc/paper_files/paper/2022/file/f7f47a73d631c0410cbc2748a8015241-Paper-Conference.pdf.
APA
Pokle, A., Geng, Z., & Kolter, J. Z. (2022). Deep Equilibrium Approaches to Diffusion Models. Advances in Neural Information Processing Systems, 35, 37975–37990. https://proceedings.neurips.cc/paper_files/paper/2022/file/f7f47a73d631c0410cbc2748a8015241-Paper-Conference.pdf
Chicago
Pokle, A., Z. Geng, and J. Z. Kolter. 2022. “Deep Equilibrium Approaches to Diffusion Models”. Advances in Neural Information Processing Systems 35: 37975–90. https://proceedings.neurips.cc/paper_files/paper/2022/file/f7f47a73d631c0410cbc2748a8015241-Paper-Conference.pdf.
Harvard
Pokle, A., Geng, Z. and Kolter, J.Z. (2022) “Deep Equilibrium Approaches to Diffusion Models”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 37975–37990. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/f7f47a73d631c0410cbc2748a8015241-Paper-Conference.pdf.
Vancouver
1. Pokle A, Geng Z, Kolter JZ (2022) Deep Equilibrium Approaches to Diffusion Models. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 37975–37990

BibTeX

@inproceedings{pokle2022deep,
  title = {Deep Equilibrium Approaches to Diffusion Models},
  author = {Pokle, Ashwini and Geng, Zhengyang and Kolter, J. Zico},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {37975-37990},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/f7f47a73d631c0410cbc2748a8015241-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors