Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise

Arpit BansalEitan BorgniaHong-Min ChuJie LiHamid KazemiFurong HuangMicah GoldblumJonas GeipingTom Goldstein

article2023NeurIPS468 citations

Demonstrates that generative diffusion models do not require Gaussian noise or stochasticity, showing that iterative restoration algorithms can synthesize realistic images by inverting arbitrary deterministic transformations such as blurring, downsampling, and masking.

Listen

Modern generative artificial intelligence relies heavily on diffusion models to synthesize high-quality images. The prevailing theoretical foundation assumes that these models require random Gaussian noise during both training and generation to navigate complex data distributions. This article investigates whether random noise is truly essential, demonstrating that generative behavior can be achieved using completely deterministic and arbitrary image transformations, a framework termed cold diffusion.

The article evaluates both conditional restoration tasks and unconditional image generation from scratch across standard benchmark datasets, including MNIST, CIFAR-10, CelebA, and AFHQ. To invert non-random transformations such as Gaussian blur, pixel masking, resolution downsampling, color desaturation, and snow corruption, the researchers developed Transformation Agnostic Cold Sampling (TACoS). This sampling algorithm corrects errors introduced by imperfect neural network approximations by subtracting and re-applying degradation steps iteratively rather than relying on naive update rules.

The findings show that generalized diffusion models successfully reconstruct and generate images without noise. In conditional tasks, TACoS consistently outperformed direct single-step restoration and naive sampling; for instance, on CelebA deblurring, TACoS achieved an image quality score of 26.14 compared to 36.37 for direct reconstruction and 299.61 for naive sampling (where lower scores indicate closer alignment with real data). In unconditional generation from scratch, starting from low-dimensional representations such as average color values, cold diffusion with blur produced competitive results (a score of 49.45 on CelebA), especially when minor variations were introduced to break pixel symmetry. The analysis also proved that naive sampling accumulates compounding mathematical errors under deterministic transformations, explaining why earlier methods failed without noise.

These results demonstrate that iterative diffusion is a general mathematical framework for reversing degradation processes rather than a technique strictly tied to noise or thermodynamics. For practitioners and decision-makers, this expands generative modeling to domain-specific physical processes, potentially reducing training constraints and enabling specialized restoration pipelines in imaging, security, and industrial design. While the empirical results firmly establish feasibility, generation diversity currently trails state-of-the-art noise-based models unless symmetry-breaking steps are applied, indicating that further work on larger datasets and refined scheduling is necessary before deploying cold diffusion in production generative systems.

arXiv: 2208.09392
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Cold Diffusion directly challenges and generalizes the foundational thermodynamic and Gaussian noise-based generative framework established in this seminal paper.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Understanding deterministic reverse sampling and non-Markovian generation formulations provides the mathematical basis for analyzing error accumulation in deterministic degradation inversion.
  • Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). This text provides the continuous-time stochastic and ordinary differential equation perspective that Cold Diffusion contrasts against when defining non-stochastic, operator-based degradation trajectories.
  • Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). Reading this paper introduces standard inverse problem setups in diffusion modeling, highlighting the limitations of relying on noise-driven schedules for general linear degradations.
  • Paper: Diffusion Posterior Sampling for General Noisy Inverse Problems, Hyungjin Chung et al. (2022). This work establishes how classical diffusion models perform posterior sampling for image inverse problems, motivating Cold Diffusion's noise-free degradation and restoration alternative.
  • Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). This paper establishes structured forward corruption processes beyond isotropic Gaussian perturbations, serving as an important conceptual step toward arbitrary deterministic transformations.
  • Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). This work clarifies the components and sampling dynamics of noise-based diffusion models, contextualizing Cold Diffusion's departure from standard noise schedules.
  • Paper: Palette: Image-to-Image Diffusion Models, Chitwan Saharia et al. (2021). Understanding image-to-image conditional diffusion models clarifies the benchmark restoration tasks—such as inpainting, deblurring, and colorization—that Cold Diffusion reformulates without noise.
Cover for Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise

Abstract

Standard diffusion models involve an image transform – adding Gaussian noise – and an image restoration operator that inverts this degradation. We observe that the generative behavior of diffusion models is not strongly dependent on the choice of image degradation, and in fact an entire family of generative models can be constructed by varying this choice. Even when using completely deterministic degradations (e.g., blur, masking, and more), the training and test-time update rules that underlie diffusion models can be easily generalized to create generative models. The success of these fully deterministic models calls into question the community’s understanding of diffusion models, which relies on noise in either gradient Langevin dynamics or variational inference, and paves the way for generalized diffusion models that invert arbitrary processes.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Generalized Diffusion
  • 3.1 Model components and training
  • 3.2 Sampling from the model
  • 3.3 Properties of Algorithm
  • 4 Generalized Diffusions with Various Transformations
  • 4.1 Deblurring
  • 4.2 Inpainting
  • 4.3 Super-Resolution
  • 4.4 Snowification
  • 5 Cold Generation
  • 5.1 Generation using deterministic noise degradation
  • 5.2 Image generation using blur
  • 5.3 Generation using other transformations
  • 6 Conclusion
  • References
  • A Appendix
  • A.1 Deblurring
  • A.2 Inpainting
  • A.3 Super-Resolution
  • A.4 Colorization
  • A.5 Image Snow
  • A.6 Generation using noise : Further Details
  • A.7 Generation using blur transformation: Further Details

Knowls

  1. Knowl 1 — Generalized Cold Diffusion Framework

    model/method

    Standard diffusion models rely on injecting additive Gaussian noise in the forward process and training a denoising neural network for the reverse process. Generalized (or "cold") diffusion extends diffusion modeling to arbitrary deterministic or stochastic image degradation operators.

    Let x0∈RNx_0 \in \mathbb{R}^N denote a clean image sampled from a data distribution X\mathcal{X}. A forward degradation operator D(x0,t)D(x_0, t) applies a transformation with severity parameter t∈[0,T]t \in [0, T], such that D(x0,0)=x0D(x_0, 0) = x_0 and the amount of retained information decreases monotonically as tt increases from 00 to TT.

    A restoration neural network Rθ(xt,t)R_\theta(x_t, t), parameterized by weights θ\theta, is trained to invert the degradation at step tt by estimating the original clean image x^0≈x0\hat{x}_0 \approx x_0. The model is optimized using an ℓ1\ell_1 loss: min⁡θEx∼X,t∼[1,T]∥Rθ(D(x,t),t)−x∥1\min_\theta \mathbb{E}_{x \sim \mathcal{X}, t \sim [1, T]} \| R_\theta(D(x, t), t) - x \|_1 This framework removes the requirement that the forward process must be stochastic Brownian motion or Gaussian noise, enabling diffusion architectures to invert arbitrary deterministic transformations such as blurring, downsampling, spatial masking, desaturation, and cross-dataset morphing.

  2. Knowl 2 — Transformation Agnostic Cold Sampling Algorithm

    algorithm

    Transformation Agnostic Cold Sampling (TACoS) is an iterative reverse sampling algorithm designed to reconstruct clean data from a severely degraded state xt=D(x0,t)x_t = D(x_0, t) using a trained restoration network R(xs,s)R(x_s, s) and the degradation operator D(x,s)D(x, s). Unlike naive reverse sampling (xs−1=D(R(xs,s),s−1)x_{s-1} = D(R(x_s, s), s-1)), TACoS applies an error-correction update at each step s→s−1s \to s-1.

    Input: Degraded sample xt∈RNx_t \in \mathbb{R}^N, starting severity step tt, degradation operator DD, restoration model RR
    Output: Reconstructed clean sample x0∈RNx_0 \in \mathbb{R}^N
    for s=t,t−1,…,1s = t, t-1, \dots, 1 do
        x^0←R(xs,s)\hat{x}_0 \leftarrow R(x_s, s)
        xs−1←xs−D(x^0,s)+D(x^0,s−1)x_{s-1} \leftarrow x_s - D(\hat{x}_0, s) + D(\hat{x}_0, s - 1)
    end for
    return x0x_0

    At step ss, the model estimates the clean image x^0=R(xs,s)\hat{x}_0 = R(x_s, s), evaluates the difference between degrading x^0\hat{x}_0 to severity s−1s-1 and severity ss, and updates the current state xsx_s with this differential. In the case of Gaussian blur, this difference corresponds to re-injecting the bandpass frequency components removed between step s−1s-1 and step ss.

  3. Knowl 3 — Error Invariance of TACoS under Linear Degradation Approximations

    theoretical result

    Let x0∈RNx_0 \in \mathbb{R}^N be an initial image, and let D(x,s)D(x, s) be a degradation operator with severity s≥0s \ge 0 satisfying D(x,0)=xD(x, 0) = x. For small values of ss, smooth degradation operators can be approximated by a first-order Taylor expansion around s=0s = 0: D(x,s)≈x+s⋅eD(x, s) \approx x + s \cdot e where e∈RNe \in \mathbb{R}^N is a constant vector independent of xx.

    Under this first-order degradation model, for any restoration operator R(xs,s)R(x_s, s) (whether perfect or imperfect), the TACoS update rule: xs−1=xs−D(R(xs,s),s)+D(R(xs,s),s−1)x_{s-1} = x_s - D(R(x_s, s), s) + D(R(x_s, s), s-1) produces exact intermediate iterates xs−1=D(x0,s−1)x_{s-1} = D(x_0, s-1) for all s≤ts \le t by mathematical induction.

    In contrast, the naive sampling rule xs−1=D(R(xs,s),s−1)x_{s-1} = D(R(x_s, s), s-1) lacks this fixed-point property whenever RR incurs error (since x0≠D(R(x,0),0)=R(x,0)x_0 \neq D(R(x, 0), 0) = R(x, 0)), causing reconstruction errors to compound over iterations.

  4. Knowl 4 — Step-wise Error Bound for TACoS versus Naive Sampling in Frequency Removal

    theoretical result

    Let a signal X∈RNX \in \mathbb{R}^N have Fourier mode decomposition X=∑i=0TxiX = \sum_{i=0}^T x_i, where each xix_i represents Fourier mode T−iT-i. Consider a degradation operator D(X,t)=∑i=tTxiD(X, t) = \sum_{i=t}^T x_i that removes one frequency mode for each unit increment of t∈{1,…,T}t \in \{1, \dots, T\}, such that at t=Tt = T, D(X,T)=xTD(X, T) = x_T is the constant DC component.

    Let Xt=∑i=tTxiX_t = \sum_{i=t}^T x_i be the true degraded signal at step tt, and let X^=R(Xt,t)=∑n=0Tx^n\hat{X} = R(X_t, t) = \sum_{n=0}^T \hat{x}_n be the reconstructed signal estimated by RR. The step error is defined as Et2=∥Xt−1−X^t−1∥2E_t^2 = \| X_{t-1} - \hat{X}_{t-1} \|^2.

    1. For Naive Sampling, where X^t−1=D(X^,t−1)=∑i=t−1Tx^i\hat{X}_{t-1} = D(\hat{X}, t-1) = \sum_{i=t-1}^T \hat{x}_i: Et,Naive2=∥∑i=t−1Txi−∑i=t−1Tx^i∥2=∑i=t−1T∥xi−x^i∥2E_{t, \text{Naive}}^2 = \left\| \sum_{i=t-1}^T x_i - \sum_{i=t-1}^T \hat{x}_i \right\|^2 = \sum_{i=t-1}^T \| x_i - \hat{x}_i \|^2
    2. For TACoS, where X^t−1=Xt−D(X^,t)+D(X^,t−1)\hat{X}_{t-1} = X_t - D(\hat{X}, t) + D(\hat{X}, t-1): Et,TACoS2=∥(D(X^,t)−D(X^,t−1))−(Xt−Xt−1)∥2=∥x^t−1−xt−1∥2E_{t, \text{TACoS}}^2 = \| (D(\hat{X}, t) - D(\hat{X}, t-1)) - (X_t - X_{t-1}) \|^2 = \| \hat{x}_{t-1} - x_{t-1} \|^2

    Consequently, whenever the restoration network RR is imperfect for more than one frequency mode i≥t−1i \ge t-1, the error of Naive Sampling is strictly greater than TACoS (Et,Naive2>Et,TACoS2E_{t, \text{Naive}}^2 > E_{t, \text{TACoS}}^2), because Naive Sampling accumulates estimation errors across all remaining lower frequencies, while TACoS isolates error strictly to the single frequency mode x^t−1\hat{x}_{t-1} reintroduced at that step.

  5. Knowl 5 — Equivalence of TACoS and DDIM Deterministic Sampling under Additive Gaussian Noise

    theoretical result

    When cold diffusion is applied to standard additive Gaussian noise degradation: D(x0,t)=αtx0+1−αtzD(x_0, t) = \sqrt{\alpha_t} x_0 + \sqrt{1 - \alpha_t} z where z∼N(0,I)z \sim \mathcal{N}(0, I) and αt∈(0,1)\alpha_t \in (0, 1) is the noise schedule, the Transformation Agnostic Cold Sampling (TACoS) update rule is mathematically equivalent to the deterministic sampling equation of Denoising Diffusion Implicit Models (DDIM).

    Given the noisy state xtx_t and the network's estimate of the clean image x^0=R(xt,t)\hat{x}_0 = R(x_t, t), the estimated noise vector z^\hat{z} is: z^(xt,t)=xt−αtx^01−αt\hat{z}(x_t, t) = \frac{x_t - \sqrt{\alpha_t} \hat{x}_0}{\sqrt{1 - \alpha_t}} Evaluating the TACoS update xt−1=xt−D(x^0,t)+D(x^0,t−1)x_{t-1} = x_t - D(\hat{x}_0, t) + D(\hat{x}_0, t-1) with D(x^0,t)=αtx^0+1−αtz^D(\hat{x}_0, t) = \sqrt{\alpha_t} \hat{x}_0 + \sqrt{1-\alpha_t} \hat{z} yields: xt−1=xt−(αtx^0+1−αtz^)+(αt−1x^0+1−αt−1z^)=αt−1x^0+1−αt−1z^x_{t-1} = x_t - \left(\sqrt{\alpha_t} \hat{x}_0 + \sqrt{1-\alpha_t} \hat{z}\right) + \left(\sqrt{\alpha_{t-1}} \hat{x}_0 + \sqrt{1-\alpha_{t-1}} \hat{z}\right) = \sqrt{\alpha_{t-1}} \hat{x}_0 + \sqrt{1 - \alpha_{t-1}} \hat{z} This is identical to DDIM deterministic sampling with stochastic parameter σt=0\sigma_t = 0.

  6. Knowl 6 — Unconditional Generation from Deterministic Degraded States

    model/method

    Unconditional image generation with cold diffusion generates images from scratch by sampling from a low-dimensional parametric distribution representing completely degraded images xTx_T, then running Transformation Agnostic Cold Sampling (TACoS) back to t=0t=0.

    1. Blur Generation: As T→∞T \to \infty, an image x0x_0 blurred with a large Gaussian filter converges to a uniform image where every pixel equals the channel-wise RGB mean vector μ∈R3\mu \in \mathbb{R}^3. A 1-component Gaussian Mixture Model (GMM) is fitted to the 3D RGB mean vectors of the training dataset. At test time, a 3D vector sampled from the GMM is expanded across spatial dimensions H×W×3H \times W \times 3, perturbed by low-magnitude additive Gaussian noise (standard deviation σ=0.002\sigma = 0.002) to break spatial pixel symmetry, and restored using TACoS over T=300T=300 deblurring steps.
    2. Inpainting Generation: The degradation blends the image toward a uniform solid color c∈R3c \in \mathbb{R}^3 sampled uniformly: xt=Gt⊙x0+(1−Gt)⊙cx_t = G_t \odot x_0 + (1 - G_t) \odot c, where GtG_t is a cumulative Gaussian mask. Sampling begins from a flat, randomly colored image cc.
    3. Super-Resolution Generation: Images are downsampled to 2×2×32 \times 2 \times 3 (12 values). A Gaussian distribution is fitted to these 12-dimensional vectors across training images. A sampled 12-dimensional vector is expanded and iteratively upsampled via TACoS.
    4. Animorphosis Generation: Data interpolation between two image distributions (e.g., CelebA human faces X\mathcal{X} and AFHQ animal faces Z\mathcal{Z}) uses xt=αtx+1−αtzx_t = \sqrt{\alpha_t} x + \sqrt{1-\alpha_t} z. Sampling from the human face manifold is performed by drawing a real animal image z∼Zz \sim \mathcal{Z} at t=Tt=T and applying TACoS to iteratively transform it into a human face.
  7. Knowl 7 — Non-Noise Forward Degradation Operators in Cold Diffusion

    model/method

    Cold diffusion implements forward degradations D(x0,t)D(x_0, t) using structured deterministic or non-Gaussian operators:

    1. Deblurring: Successive 2D Gaussian convolution filters MsM_s are applied: xt=Mt∘⋯∘M1∘x0=Mˉt∘x0=Gˉt∗x0x_t = M_t \circ \dots \circ M_1 \circ x_0 = \bar{M}_t \circ x_0 = \bar{G}_t * x_0, where the Gaussian kernel standard deviation increases with step tt (e.g., 11×1111 \times 11 kernel with standard deviation σt=0.01t+0.35\sigma_t = 0.01 t + 0.35 for CIFAR-10, or 15×1515 \times 15 kernel growing exponentially at rate 0.010.01 for CelebA).
    2. Inpainting: Spatial Gaussian masks zβi∈[0,1]H×Wz_{\beta_i} \in [0, 1]^{H \times W} are iteratively applied: D(x0,t)=x0⊙∏i=1tzβiD(x_0, t) = x_0 \odot \prod_{i=1}^t z_{\beta_i}, where βi\beta_i parameterizes Gaussian curve variance to progressively enlarge the grayed-out region.
    3. Super-Resolution / Downsampling: Images are downsampled by factors of 2 per step and resized to original dimensions using nearest-neighbor interpolation, reducing information until reaching a minimum base resolution (4×44 \times 4 for MNIST/CIFAR-10, 2×22 \times 2 for CelebA).
    4. Colorization / Desaturation: Three-channel 1×11 \times 1 convolution filters z(α)z(\alpha) blend RGB channels toward their arithmetic mean 13(R+G+B)\frac{1}{3}(R+G+B) via a schedule αt=t/T\alpha_t = t/T, producing a grayscale image at t=Tt=T.
    5. Snowification: Adaptations of the ImageNet-C snow corruption apply motion-blurred spline noise patterns S+S′S + S' parameterized by severity c0(t)c_0(t) and windiness c1(t)c_1(t) linearly scheduled over t∈[1,T]t \in [1, T].
  8. Knowl 8 — Reconstruction Performance across Vision Inverse Tasks

    data/table

    Generalized cold diffusion models using TACoS were evaluated on deblurring, inpainting, super-resolution, and colorization across MNIST, CIFAR-10, and CelebA. Across all tasks, iterative sampling with TACoS consistently achieves lower (better) Fréchet Inception Distance (FID) than direct one-shot reconstruction R(D(x0,T),T)R(D(x_0, T), T), demonstrating that iterative inversion aligns generated outputs more closely with the true data distribution despite trade-offs in pixel-level metrics (SSIM, RMSE).

    Task Dataset Sampled (TACoS) Direct Reconstruction
    FID ↓\downarrow SSIM ↑\uparrow RMSE ↓\downarrow FID ↓\downarrow SSIM ↑\uparrow RMSE ↓\downarrow
    Deblurring MNIST 4.69 0.718 0.154 5.10 0.757 0.142
    CIFAR-10 80.08 0.773 0.075 83.69 0.775 0.071
    CelebA 26.14 0.568 0.093 36.37 0.607 0.083
    Inpainting MNIST 1.61 0.941 0.068 2.24 0.948 0.060
    CIFAR-10 8.92 0.859 0.068 9.97 0.869 0.063
    CelebA 5.73 0.917 0.043 7.74 0.922 0.039
    Super-Resolution MNIST 4.33 0.820 0.115 4.05 0.823 0.114
    CIFAR-10 152.76 0.411 0.155 169.94 0.420 0.152
    CelebA 96.92 0.381 0.201 112.84 0.400 0.196
    Colorization CIFAR-10 45.74 0.942 0.069 - - -
    CelebA 17.50 0.973 0.042 - - -
  9. Knowl 9 — Unconditional Generation Quality and the Effect of Symmetry Breaking

    data/table

    Unconditional image generation on 128×128128 \times 128 CelebA and AFHQ datasets was evaluated comparing standard noise diffusion (Hot Diffusion) against blur-based generalized diffusion (Cold Diffusion). In blur-based generation, the initial state xTx_T is sampled from a Gaussian Mixture Model (GMM) fitted over channel-wise RGB means.

    Sampling under perfect spatial symmetry (assigning all pixels in a channel the identical GMM mean value) produces images with high visual fidelity but limited diversity, leading to elevated FID scores. Adding a small Gaussian perturbation (standard deviation σ=0.002\sigma = 0.002) to break spatial pixel symmetry substantially improves FID scores.

    Dataset Hot Diffusion (FID ↓\downarrow) Cold Diffusion (FID ↓\downarrow)
    Fixed Noise Estimated Noise (DDIM) Perfect Symmetry Broken Symmetry (σ=0.002\sigma=0.002)
    CelebA (128×128128 \times 128) 59.91 23.11 97.00 49.45
    AFHQ (128×128128 \times 128) 25.62 20.59 93.05 54.68

    For other cold transformations on CelebA (128×128128 \times 128), unconditional generation achieves FID scores of 90.14 for inpainting, 92.91 for super-resolution, and 48.51 for animorphosis (inverting AFHQ animal images to CelebA human faces).

  10. Knowl 10 — Superiority of TACoS over Naive Sampling in Cold Diffusion

    empirical result

    When inverting smooth, deterministic degradations (such as Gaussian blurring), Naive Sampling (xs−1=D(R(xs,s),s−1)x_{s-1} = D(R(x_s, s), s-1)) causes severe compounding reconstruction errors, resulting in high-frequency artifacts and failure to converge to the target distribution.

    For Gaussian blur inversion across MNIST, CIFAR-10, and CelebA, Naive Sampling yields significantly worse FID scores than both TACoS and direct single-step reconstruction:

    • MNIST: Naive Sampling FID = 8.24 vs. Direct Reconstruction FID = 5.10 vs. TACoS FID = 4.69.
    • CIFAR-10: Naive Sampling FID = 97.89 vs. Direct Reconstruction FID = 83.69 vs. TACoS FID = 80.08.
    • CelebA: Naive Sampling FID = 299.61 vs. Direct Reconstruction FID = 36.37 vs. TACoS FID = 26.14.

    In unconditional generation on 128×128128 \times 128 datasets with broken symmetry, direct one-step reconstruction yields FID scores of 257.69 on CelebA and 214.24 on AFHQ, whereas TACoS achieves 49.45 and 54.68 respectively, while Naive Sampling fails entirely.

Coverage note — None was omitted; all key theoretical formulations, sampling algorithms, mathematical proofs, experimental task designs, and quantitative results for both conditional restoration and unconditional generation are represented in the knowls.

References

  1. 1.Shane Barratt and Rishi Sharma. A note on the inception score. arXiv preprint arXiv:1801.01973, 2018.
  2. 2.Mikołaj Bińkowski, Danica J Sutherland, Michael Arbel, and Arthur Gretton. Demystifying mmd gans. arXiv preprint arXiv:1801.01401, 2018.
  3. 3.Tong Che, Yanran Li, Athul Paul Jacob, Yoshua Bengio, and Wenjie Li. Mode regularized generative adversarial networks. arXiv preprint arXiv:1612.02136, 2016.
  4. 4.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat gans on image synthesis. volume 34, 2021.
  5. 5.Dan Hendrycks and Thomas G. Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations, ICLR 2019. OpenReview.net, 2019.
  6. 6.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  7. 7.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 32, 2020.
  8. 8.Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation. J. Mach. Learn. Res., 23, 2022.
  9. 9.Alexia Jolicoeur-Martineau, Rémi Piché-Taillefer, Ioannis Mitliagkas, and Remi Tachet des Combes. Adversarial score matching and improved sampling for image generation. International Conference on Learning Representations, 2021.
  10. 10.Zahra Kadkhodaie and Eero Simoncelli. Stochastic solutions for linear inverse problems using the prior implicit in a denoiser. Advances in Neural Information Processing Systems, 34, 2021.
  11. 11.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. arXiv preprint arXiv:2206.00364, 2022.
  12. 12.Bahjat Kawar, Gregory Vaksman, and Michael Elad. Stochastic image denoising by sampling from the posterior distribution. International Conference on Computer Vision Workshops, 2021.
  13. 13.Bahjat Kawar, Michael Elad, Stefano Ermon, and Jiaming Song. Denoising diffusion restoration models. arXiv preprint arXiv:2201.11793, 2022.
  14. 14.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  15. 15.Diederik P. Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. On density estimation with diffusion models. Advances in Neural Information Processing Systems, 34, 2021.
  16. 16.Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, 2009.
  17. 17.Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proc. IEEE, 86(11):2278–2324, 1998.
  18. 18.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV), December 2015.
  19. 19.Christopher A. Metzler, Ali Mousavi, and Richard G. Baraniuk. Learned D-AMP: principled neural network based compressive image recovery. Advances in Neural Information Processing Systems, 30, 2017.
  20. 20.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8162–8171, 2021.
  21. 21.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  22. 22.Severi Rissanen, Markus Heinonen, and Arno Solin. Generative modelling with inverse heat dissipation. arXiv preprint arXiv:2206.13397, 2022.
  23. 23.Yaniv Romano, Michael Elad, and Peyman Milanfar. The little engine that could: Regularization by denoising (RED). arXiv preprint arXiv:1611.02862, 2016.
  24. 24.Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J. Fleet, and Mohammad Norouzi. Image super-resolution via iterative refinement. arXiv preprint arXiv:2104.07636, 2021.
  25. 25.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. Advances in neural information processing systems, 29, 2016.
  26. 26.Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, volume 37 of JMLR Workshop and Conference Proceedings, 2015.
  27. 27.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. International Conference on Learning Representations, 2021a.
  28. 28.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. Advances in Neural Information Processing Systems, 32, 2019.
  29. 29.Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. International Conference on Learning Representations, 2021b.
  30. 30.Jay Whang, Mauricio Delbracio, Hossein Talebi, Chitwan Saharia, Alexandros G. Dimakis, and Peyman Milanfar. Deblurring via stochastic refinement. arXiv preprint arXiv:2112.02475, 2021.

Citation

MLA
Bansal, A., et al. “Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 41259–82, https://proceedings.neurips.cc/paper_files/paper/2023/file/80fe51a7d8d0c73ff7439c2a2554ed53-Paper-Conference.pdf.
APA
Bansal, A., Borgnia, E., Chu, H.-M., Li, J., Kazemi, H., Huang, F., Goldblum, M., Geiping, J., & Goldstein, T. (2023). Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise. Advances in Neural Information Processing Systems, 36, 41259–41282. https://proceedings.neurips.cc/paper_files/paper/2023/file/80fe51a7d8d0c73ff7439c2a2554ed53-Paper-Conference.pdf
Chicago
Bansal, A., E. Borgnia, H.-M. Chu, et al. 2023. “Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise”. Advances in Neural Information Processing Systems 36: 41259–82. https://proceedings.neurips.cc/paper_files/paper/2023/file/80fe51a7d8d0c73ff7439c2a2554ed53-Paper-Conference.pdf.
Harvard
Bansal, A. et al. (2023) “Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 41259–41282. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/80fe51a7d8d0c73ff7439c2a2554ed53-Paper-Conference.pdf.
Vancouver
1. Bansal A, Borgnia E, Chu H-M, Li J, Kazemi H, Huang F, Goldblum M, Geiping J, Goldstein T (2023) Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 41259–41282

BibTeX

@inproceedings{bansal2023cold,
  title = {Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise},
  author = {Bansal, Arpit and Borgnia, Eitan and Chu, Hong-Min and Li, Jie and Kazemi, Hamid and Huang, Furong and Goldblum, Micah and Geiping, Jonas and Goldstein, Tom},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {41259-41282},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/80fe51a7d8d0c73ff7439c2a2554ed53-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors