Improved Techniques for Training Score-Based Generative Models

Yang SongStefano Ermon

article2020NeurIPS1,510 citations

Develops practical training methods and theoretical principles that scale score-based generative models up to 256x256 image synthesis, matching the visual quality of state-of-the-art GANs without adversarial training.

Listen

Generating realistic, high-resolution synthetic data is critical for applications in automated design, data synthesis, and anomaly detection. While score-based generative models offer a stable alternative to generative adversarial networks (GANs) by avoiding fragile adversarial training objectives, earlier implementations suffered from training instability and were restricted to low-resolution images (typically 32x32 pixels or lower). The article aims to evaluate the underlying mathematical failure modes of score-based models in high-dimensional spaces and demonstrates a principled framework to scale them to diverse, high-resolution datasets.

To address these limitations, the authors developed a set of five foundational techniques based on mathematical analyses of simplified data distributions and Langevin dynamics sampling. Specifically, they established analytical formulas to set the initial noise scale equal to the maximum pairwise distance between data points, arranged intermediate noise levels in a geometric progression with shared coverage, scaled network outputs inversely by noise intensity to eliminate memory-intensive conditioning, and optimized sampling step sizes under fixed compute budgets. They also implemented an exponential moving average (EMA) of model parameters to stabilize training. These methods were evaluated across standard benchmarks including CIFAR-10, CelebA, LSUN scenes, and FFHQ across image resolutions from 32x32 to 256x256 using standard automated metrics (such as Frechet Inception Distance, or FID) and human perceptual evaluations.

The findings show that this upgraded model (termed NCSNv2) successfully scales score-based generative models to 256x256 resolution, generating sharp, structurally consistent, and diverse images that rival top GANs. Quantitatively, the combined techniques with Tweedie denoising reduced the CIFAR-10 FID score from 25.32 down to 10.87 (where lower is better) and improved CelebA 64x64 FID from 25.30 to 10.23. Furthermore, human perceptual evaluation scores on CelebA 64x64 nearly doubled, rising from 19.8% to 37.3%, matching established GAN benchmarks. Applying exponential moving averages proved essential, eliminating prominent color-shift artifacts and stabilizing checkpoint quality throughout training.

These results demonstrate that generative modeling can achieve state-of-the-art visual quality without the high training instability or complex hyperparameter tuning typical of adversarial frameworks. By replacing ad-hoc tuning with closed-form mathematical choices for noise schedules and step sizes, the framework significantly reduces engineering trial-and-error. For technical decision-makers, this lowers the risk and operational cost of training robust generative pipelines across high-dimensional modalities, while opening opportunities for seamless data synthesis and representation learning.

Organizations should adopt these five design principles—especially EMA tracking and analytical noise scheduling—when deploying score-based architectures. When compute budgets are constrained, teams should set sampling steps according to available hardware and calculate the step size analytically using the article’s formula rather than running expensive grid searches. As a next step, practitioners should explore fine-tuning hyperparameters for specific use cases and evaluate extending these noise principles to non-image modalities, such as audio, sensor streams, or behavioral data.

While the analytical derivations rely on simplified Gaussian assumptions and standard Euclidean distances, the empirical outcomes robustly transfer to complex real-world image datasets. Users should exercise caution regarding computational inference latency, as iterative sampling remains slower than single-pass generators, and should account for inherited dataset biases when deploying synthetic models in sensitive operational settings.

arXiv: 2006.09011
  • Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). This work generalizes the discrete multi-scale noise levels analyzed in the source into a unified continuous stochastic differential equation (SDE) framework, incorporating the architectural and training improvements introduced here.
  • Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). This paper systematically dissects and refines the design space of diffusion and score-based models, directly building on the noise schedules, network preconditioning, and sampling stability techniques advanced by the source.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This seminal work demonstrates that diffusion probabilistic models trained with score-matching objectives can scale effectively to high-resolution synthesis, echoing and expanding upon the high-dimensional sampling findings of the source.
  • Paper: Improved Denoising Diffusion Probabilistic Models, Alex Nichol et al. (2021). This work builds on practical high-resolution training and sampling insights from score-based diffusion methods to optimize noise schedules and reduce required sampling steps.
  • Paper: Cascaded Diffusion Models for High Fidelity Image Generation, Jonathan Ho et al. (2021). This paper extends high-resolution score/diffusion generation to higher scales via cascaded pipelines, utilizing conditioning augmentations to mitigate errors in multi-stage generation.
  • Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). This work advances continuous generative flows and transformers for large-scale, high-resolution synthesis, extending beyond standard score-based and diffusion formulations.
Cover for Improved Techniques for Training Score-Based Generative Models

Abstract

Score-based generative models can produce high quality image samples comparable to GANs, without requiring adversarial optimization. However, existing training procedures are limited to images of low resolution (typically below 32x32), and can be unstable under some settings. We provide a new theoretical analysis of learning and sampling from score models in high dimensional spaces, explaining existing failure modes and motivating new solutions that generalize across datasets. To enhance stability, we also propose to maintain an exponential moving average of model weights. With these improvements, we can effortlessly scale score-based generative models to images with unprecedented resolutions ranging from 64x64 to 256x256. Our score-based models can generate high-fidelity samples that rival best-in-class GANs on various image datasets, including CelebA, FFHQ, and multiple LSUN categories.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Langevin dynamics
  • 2.2 Score-based generative modeling
  • 3 Choosing noise scales
  • 3.1 Initial noise scale
  • 3.2 Other noise scales
  • 3.3 Incorporating the noise information
  • 4 Configuring annealed Langevin dynamics
  • 5 Improving stability with moving average
  • 6 Combining all techniques together
  • 7 Conclusion
  • References
  • A Proofs
  • B Experimental details
  • B.1 Network architectures and hyperparameters
  • B.2 Additional settings
  • C Additional experimental results
  • C.1 Additional results without the denoising step
  • C.2 Training and sampling speed
  • C.3 Color shifts
  • C.4 Additional results on ablation studies
  • C.5 Generalization
  • C.5.1 Loss curves
  • C.5.2 Nearest neighbors
  • C.5.3 Additional interpolation results
  • C.6 Additional uncurated samples

Knowls

  1. Knowl 1 — Initial Noise Scale Selection Based on Maximum Pairwise Distance

    theoretical result

    For an empirical dataset distribution p^data(x)=1N∑i=1Nδ(x=x(i))\hat{p}_\text{data}(\mathbf{x}) = \frac{1}{N}\sum_{i=1}^N \delta(\mathbf{x} = \mathbf{x}^{(i)}) of NN data points in RD\mathbb{R}^D perturbed by Gaussian noise with variance σ12\sigma_1^2, the perturbed density is given by p^σ1(x)=1N∑i=1Np(i)(x)\hat{p}_{\sigma_1}(\mathbf{x}) = \frac{1}{N}\sum_{i=1}^N p^{(i)}(\mathbf{x}), where p(i)(x)=N(x∣x(i),σ12I)p^{(i)}(\mathbf{x}) = \mathcal{N}(\mathbf{x} \mid \mathbf{x}^{(i)}, \sigma_1^2 \mathbf{I}). Defining the component weight function

    r(i)(x)=p(i)(x)∑k=1Np(k)(x)r^{(i)}(\mathbf{x}) = \frac{p^{(i)}(\mathbf{x})}{\sum_{k=1}^N p^{(k)}(\mathbf{x})}

    the score function decomposes as ∇xlog⁡p^σ1(x)=∑i=1Nr(i)(x)∇xlog⁡p(i)(x)\nabla_\mathbf{x} \log \hat{p}_{\sigma_1}(\mathbf{x}) = \sum_{i=1}^N r^{(i)}(\mathbf{x}) \nabla_\mathbf{x} \log p^{(i)}(\mathbf{x}). The expected transition weight from component ii to component jj is upper-bounded by:

    Ep(i)(x)[r(j)(x)]≤12exp⁡(−∥x(i)−x(j)∥228σ12)\mathbb{E}_{p^{(i)}(\mathbf{x})}[r^{(j)}(\mathbf{x})] \le \frac{1}{2} \exp\left( -\frac{\|\mathbf{x}^{(i)} - \mathbf{x}^{(j)}\|_2^2}{8\sigma_1^2} \right)

    If the initial noise standard deviation σ1\sigma_1 is small relative to the Euclidean distance ∥x(i)−x(j)∥2\|\mathbf{x}^{(i)} - \mathbf{x}^{(j)}\|_2, the expectation decays exponentially, causing Langevin dynamics initialized near x(i)\mathbf{x}^{(i)} to ignore other modes x(j)\mathbf{x}^{(j)}. To ensure mode exploration and diverse sample generation, σ1\sigma_1 must be chosen to be as large as the maximum Euclidean distance between all pairs of training data points (or estimated on a random subset of data points when NN is very large).

  2. Knowl 2 — Geometric Noise Scale Schedule via High-Dimensional Radial Overlap

    theoretical result

    Let x∈RD∼N(0,σ2I)\mathbf{x} \in \mathbb{R}^D \sim \mathcal{N}(0, \sigma^2 \mathbf{I}) and r=∥x∥2r = \|\mathbf{x}\|_2. The radial density of x\mathbf{x} in hyperspherical coordinates is

    p(r)=12D/2−1Γ(D/2)rD−1σDexp⁡(−r22σ2)p(r) = \frac{1}{2^{D/2-1}\Gamma(D/2)} \frac{r^{D-1}}{\sigma^D} \exp\left( -\frac{r^2}{2\sigma^2} \right)

    As dimensionality D→∞D \to \infty, the radial variable converges in distribution according to r−Dσ→dN(0,σ2/2)r - \sqrt{D}\sigma \xrightarrow{d} \mathcal{N}(0, \sigma^2/2). For high-dimensional data, the radial density pσi(r)p_{\sigma_i}(r) for noise scale σi\sigma_i is concentrated within the three-sigma high-density interval Ii−1=[mi−1−3si−1,mi−1+3si−1]I_{i-1} = [m_{i-1} - 3s_{i-1}, m_{i-1} + 3s_{i-1}], where mi−1=Dσi−1m_{i-1} = \sqrt{D}\sigma_{i-1} and si−1=σi−1/2s_{i-1} = \sigma_{i-1}/\sqrt{2}.

    To ensure samples from noise level σi\sigma_i adequately cover the high-density regions of the preceding scale σi−1\sigma_{i-1}, the probability mass of pσi(r)p_{\sigma_i}(r) within Ii−1I_{i-1} is set to a constant overlap probability C∈(0,1)C \in (0, 1):

    pσi(r∈Ii−1)=Φ(2D(γi−1)+3γi)−Φ(2D(γi−1)−3γi)=Cp_{\sigma_i}(r \in I_{i-1}) = \Phi\left(\sqrt{2D}(\gamma_i - 1) + 3\gamma_i\right) - \Phi\left(\sqrt{2D}(\gamma_i - 1) - 3\gamma_i\right) = C

    where γi=σi−1/σi\gamma_i = \sigma_{i-1}/\sigma_i and Φ(⋅)\Phi(\cdot) is the cumulative distribution function of the standard normal distribution N(0,1)\mathcal{N}(0, 1). Enforcing a constant overlap C≈0.5C \approx 0.5 forces the ratio γi\gamma_i to be constant across all scale pairs (γ2=γ3=⋯=γL=γ\gamma_2 = \gamma_3 = \dots = \gamma_L = \gamma), proving that the sequence of noise scales {σi}i=1L\{\sigma_i\}_{i=1}^L should follow a geometric progression.

  3. Knowl 3 — Variance Evolution and Step Size Selection in Annealed Langevin Dynamics

    theoretical result

    Consider applying Langevin dynamics targeting pσi(x)=N(x∣0,σi2I)p_{\sigma_i}(\mathbf{x}) = \mathcal{N}(\mathbf{x} \mid \mathbf{0}, \sigma_i^2 \mathbf{I}) initialized with samples from the prior scale x0∼pσi−1(x)=N(0,σi−12I)\mathbf{x}_0 \sim p_{\sigma_{i-1}}(\mathbf{x}) = \mathcal{N}(\mathbf{0}, \sigma_{i-1}^2 \mathbf{I}). Using the step size αi=ϵ⋅σi2σL2\alpha_i = \epsilon \cdot \frac{\sigma_i^2}{\sigma_L^2}, the Langevin iteration updates as xt+1←xt+αi∇xlog⁡pσi(xt)+2αizt\mathbf{x}_{t+1} \leftarrow \mathbf{x}_t + \alpha_i \nabla_\mathbf{x}\log p_{\sigma_i}(\mathbf{x}_t) + \sqrt{2\alpha_i}\mathbf{z}_t with zt∼N(0,I)\mathbf{z}_t \sim \mathcal{N}(\mathbf{0}, \mathbf{I}). After TT iterations, the state distribution is xT∼N(0,sT2I)\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, s_T^2 \mathbf{I}), where the variance ratio is given in closed form by:

    sT2σi2=(1−ϵσL2)2T(γ2−2ϵσL2−σL2(1−ϵσL2)2)+2ϵσL2−σL2(1−ϵσL2)2\frac{s_T^2}{\sigma_i^2} = \left( 1 - \frac{\epsilon}{\sigma_L^2} \right)^{2T} \left( \gamma^2 - \frac{2\epsilon}{\sigma_L^2 - \sigma_L^2\left(1 - \frac{\epsilon}{\sigma_L^2}\right)^2} \right) + \frac{2\epsilon}{\sigma_L^2 - \sigma_L^2\left(1 - \frac{\epsilon}{\sigma_L^2}\right)^2}

    where γ=σi−1/σi\gamma = \sigma_{i-1}/\sigma_i is the common ratio of the geometric noise progression. The ratio sT2/σi2s_T^2/\sigma_i^2 is independent of the scale index ii and data dimension DD.

    To optimize mixing across noise scales, one fixes TT according to a target computation budget (e.g., T∈[3,5]T \in [3, 5] when the total number of noise levels LL is large) and selects the step size hyperparameter ϵ\epsilon via grid search such that sT2/σi2s_T^2/\sigma_i^2 is maximally close to 11.

  4. Knowl 4 — Noise-Conditional Score Parameterization via Output Rescaling

    model/method

    For an isotropic Gaussian perturbation pσ(x)=N(x∣0,σ2I)p_\sigma(\mathbf{x}) = \mathcal{N}(\mathbf{x} \mid \mathbf{0}, \sigma^2 \mathbf{I}) in RD\mathbb{R}^D, the expected gradient norm satisfies E[∥∇xlog⁡pσ(x)∥2]≈Dσ\mathbb{E}[\|\nabla_\mathbf{x} \log p_\sigma(\mathbf{x})\|_2] \approx \frac{\sqrt{D}}{\sigma}, indicating that the norm of the score function scales inversely with σ\sigma.

    Rather than modulating intermediate network activations via scale-and-bias parameters in conditional normalization layers (whose memory footprint scales linearly with the number of noise scales LL), the noise conditional score network sθ(x,σ)\mathbf{s}_\theta(\mathbf{x}, \sigma) is parameterized by directly dividing the output of an unconditional score network sθ(x)\mathbf{s}_\theta(\mathbf{x}) by σ\sigma:

    sθ(x,σ)=sθ(x)σ\mathbf{s}_\theta(\mathbf{x}, \sigma) = \frac{\mathbf{s}_\theta(\mathbf{x})}{\sigma}

    The network is trained using the multi-scale denoising score matching objective:

    L(θ)=12L∑i=1LEpdata(x)Epσi(x~∣x)[∥σisθ(x~,σi)+x~−xσi∥22]\mathcal{L}(\theta) = \frac{1}{2L} \sum_{i=1}^L \mathbb{E}_{p_\text{data}(\mathbf{x})}\mathbb{E}_{p_{\sigma_i}(\tilde{\mathbf{x}}\mid\mathbf{x})}\left[ \left\| \sigma_i \mathbf{s}_\theta(\tilde{\mathbf{x}}, \sigma_i) + \frac{\tilde{\mathbf{x}} - \mathbf{x}}{\sigma_i} \right\|_2^2 \right]

    This parameterization enables score estimation across thousands of discrete or continuous noise scales without increasing model memory or parameters.

  5. Knowl 5 — Exponential Moving Average of Parameters for Score-Based Models

    model/method

    During training of noise-conditional score networks, stochastic gradient updates produce high variance in the score estimation, resulting in severe sample quality volatility and global color shift artifacts across checkpoints.

    To stabilize the learned score field, an exponential moving average (EMA) of network parameters is maintained. Letting θi\theta_i denote the model parameters at optimization step ii, shadow parameters θ′\theta' are tracked via:

    θ′←mθ′+(1−m)θi\theta' \leftarrow m\theta' + (1 - m)\theta_i

    where m∈[0,1)m \in [0, 1) is the momentum decay factor (typically m=0.999m = 0.999). At inference time, score evaluation during annealed Langevin dynamics uses the averaged parameters sθ′(x,σ)\mathbf{s}_{\theta'}(\mathbf{x}, \sigma) instead of sθi(x,σ)\mathbf{s}_{\theta_i}(\mathbf{x}, \sigma). This eliminates visual color shifts and consistently lowers sample Fréchet Inception Distance (FID).

  6. Knowl 6 — Annealed Langevin Dynamics with Tweedie Denoising

    algorithm

    The sampling procedure generates images by progressively sampling from noise-perturbed distributions {pσi}i=1L\{p_{\sigma_i}\}_{i=1}^L with descending noise standard deviations σ1>σ2>⋯>σL\sigma_1 > \sigma_2 > \dots > \sigma_L. At each scale σi\sigma_i, TT steps of Langevin dynamics are performed using step size αi=ϵ⋅σi2/σL2\alpha_i = \epsilon \cdot \sigma_i^2 / \sigma_L^2. After the final Langevin step at scale σL\sigma_L, Tweedie's formula is applied as an analytical denoising step to subtract the residual Gaussian noise N(0,σL2I)\mathcal{N}(0, \sigma_L^2 \mathbf{I}).

    Input: Noise scales {σi\sigma_i}i=1L_{i=1}^L, step size scaling factor ϵ\epsilon, iterations per scale TT, score network sθ(x,σ)\mathbf{s}_\theta(\mathbf{x}, \sigma), initial sample x0\mathbf{x}_0
    Output: Denoised generated sample xfinal\mathbf{x}_\text{final}
    for i←1i \leftarrow 1 to LL do
        αi←ϵ⋅σi2/σL2\alpha_i \leftarrow \epsilon \cdot \sigma_i^2 / \sigma_L^2
        for t←1t \leftarrow 1 to TT do
            Draw zt∼N(0,I)\mathbf{z}_t \sim \mathcal{N}(\mathbf{0}, \mathbf{I})
            xt←xt−1+αisθ(xt−1,σi)+2αizt\mathbf{x}_t \leftarrow \mathbf{x}_{t-1} + \alpha_i \mathbf{s}_\theta(\mathbf{x}_{t-1}, \sigma_i) + \sqrt{2\alpha_i} \mathbf{z}_t
        x0←xT\mathbf{x}_0 \leftarrow \mathbf{x}_T
    xfinal←xT+σL2sθ(xT,σL)\mathbf{x}_\text{final} \leftarrow \mathbf{x}_T + \sigma_L^2 \mathbf{s}_\theta(\mathbf{x}_T, \sigma_L)
    return xfinal\mathbf{x}_\text{final}
  7. Knowl 7 — Latent Spherical Interpolation via Langevin Noise Trajectories

    model/method

    Continuous semantic interpolations between two generated samples x(1)\mathbf{x}^{(1)} and x(2)\mathbf{x}^{(2)} are constructed by interpolating the injected Gaussian noise sequences used during annealed Langevin dynamics.

    Let {zij}1≤i≤L,1≤j≤T\{\mathbf{z}_{ij}\}_{1 \le i \le L, 1 \le j \le T} denote the full tensor of standard Gaussian noise vectors injected at noise level ii and step jj during sampling. For two trajectories sharing the same initialization x0\mathbf{x}_0 but driven by noise realizations {zij(1)}\{\mathbf{z}_{ij}^{(1)}\} and {zij(2)}\{\mathbf{z}_{ij}^{(2)}\}, NN intermediate interpolated samples are produced. The kk-th intermediate sample (1≤k≤N1 \le k \le N) is generated by running annealed Langevin dynamics initialized at x0\mathbf{x}_0 with the spherically interpolated noise trajectory:

    zij(k)=cos⁡(kπ2(N+1))zij(1)+sin⁡(kπ2(N+1))zij(2)\mathbf{z}_{ij}^{(k)} = \cos\left(\frac{k\pi}{2(N+1)}\right) \mathbf{z}_{ij}^{(1)} + \sin\left(\frac{k\pi}{2(N+1)}\right) \mathbf{z}_{ij}^{(2)}

    This procedure produces smooth, high-fidelity semantic transitions across complex image domains without requiring an explicit latent encoder.

  8. Knowl 8 — Quantitative Generation Performance on CIFAR-10 and CelebA

    data/table

    NCSNv2 incorporates theoretical noise scale selection ({σi}\{\sigma_i\}), network output rescaling (1/σ1/\sigma), step size calibration ({αi}\{\alpha_i\}), parameter EMA, and Tweedie denoising. It achieves lower FID scores on unconditional CIFAR-10 (32×3232 \times 32) and CelebA (64×6464 \times 64) compared to the baseline NCSN and competitive likelihood-based and GAN models.

    Model Inception ↑\uparrow FID ↓\downarrow
    CIFAR-10 Unconditional
    PixelCNN 4.60 65.93
    IGEBM 6.02 40.58
    WGAN-GP 7.86 ±\pm .07 36.4
    SNGAN 8.22 ±\pm .05 21.7
    NCSN 8.87 ±\pm .12 25.32
    NCSN (w/ denoising) 7.32 ±\pm .12 29.8
    NCSNv2 (w/o denoising) 8.73 ±\pm .13 31.75
    NCSNv2 (w/ denoising) 8.40 ±\pm .07 10.87
    CelebA 64×6464 \times 64
    NCSN (w/o denoising) - 26.89
    NCSN (w/ denoising) - 25.30
    NCSNv2 (w/o denoising) - 28.86
    NCSNv2 (w/ denoising) - 10.23

    The combined techniques along with Tweedie denoising improve the CIFAR-10 FID of score-based models from 25.3225.32 (NCSN) down to 10.8710.87 (NCSNv2), and the CelebA 64×6464 \times 64 FID from 25.3025.30 down to 10.2310.23.

  9. Knowl 9 — Human Perceptual Evaluation (HYPE-infinity) on CelebA

    data/table

    Human perceptual fidelity measured by the HYPE∞\text{HYPE}_\infty benchmark (where higher percentages indicate higher realism, up to 50%50\%) shows that NCSNv2 achieves sample quality competitive with progressive GAN architectures on CelebA 64×6464 \times 64, whereas the original NCSN scores poorly due to color shifts.

    Model HYPE∞\text{HYPE}_\infty (%) ↑\uparrow Fakes Error (%) Reals Error (%) Std.
    StyleGAN* 50.7 62.2 39.3 1.3
    ProgressiveGAN 40.3 46.2 34.4 0.9
    BEGAN 10.0 6.2 13.8 1.6
    WGAN-GP 3.8 1.7 5.9 0.6
    NCSN 19.8 22.3 17.3 0.4
    NCSNv2 37.3 49.8 24.8 0.5

    *Evaluated with truncation tricks. While vanilla NCSN scored 19.8%19.8\%, NCSNv2 achieves 37.3%37.3\%, showing that the proposed techniques bring score-based generative sample fidelity to a level comparable to ProgressiveGAN (40.3%40.3\%).

  10. Knowl 10 — Hyperparameter and Scale Configurations for High-Resolution Datasets

    data/table

    NCSNv2 parameters σ1\sigma_1, LL, TT, and ϵ\epsilon are configured analytically across varying image resolutions using Techniques 1–4. The initial noise scale σ1\sigma_1 is set to the maximum pairwise dataset Euclidean distance, the number of noise levels LL is set to satisfy radial overlap C≈0.5C \approx 0.5, and step size parameter ϵ\epsilon is matched to TT using the closed-form variance relation.

    Model Dataset σ1\sigma_1 LL TT ϵ\epsilon Batch Size Iterations
    NCSN CIFAR-10 32232^2 1 10 100 2e-5 128 300k
    NCSN CelebA 64264^2 1 10 100 2e-5 128 210k
    NCSN LSUN church 96296^2 1 10 100 2e-5 128 200k
    NCSN LSUN bedroom 1282128^2 1 10 100 2e-5 64 150k
    NCSNv2 CIFAR-10 32232^2 50 232 5 6.2e-6 128 300k
    NCSNv2 CelebA 64264^2 90 500 5 3.3e-6 128 210k
    NCSNv2 LSUN church 96296^2 140 788 4 4.9e-6 128 200k
    NCSNv2 LSUN bedroom/tower 1282128^2 190 1086 3 1.8e-6 128 150k
    NCSNv2 FFHQ 2562256^2 348 2311 3 0.9e-7 32 80k

    For higher image resolutions, σ1\sigma_1 scales from 5050 to 348348, which requires LL to increase from 232232 to 23112311 geometric noise levels with T∈[3,5]T \in [3, 5] Langevin steps per level, scaling score-based generative modeling up to 256×256256 \times 256 images.

Coverage note — None was omitted; all primary theoretical propositions, derivation-free formulas (initial noise bound, radial Gaussian limit, variance ratio formula), algorithmic enhancements (Techniques 1-5), and key quantitative results across resolutions are fully represented.

References

  1. 1.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. In Advances in Neural Information Processing Systems, pages 11895–11907, 2019.
  2. 2.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  3. 3.Aapo Hyvärinen. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research, 6(Apr):695–709, 2005.
  4. 4.Pascal Vincent. A connection between score matching and denoising autoencoders. Neural computation, 23(7):1661–1674, 2011.
  5. 5.Martin Raphan and Eero P Simoncelli. Least squares estimation without priors or supervision. Neural computation, 23(2):374–420, 2011.
  6. 6.Gareth O Roberts, Richard L Tweedie, et al. Exponential convergence of langevin distributions and their discrete approximations. Bernoulli, 2(4):341–363, 1996.
  7. 7.Max Welling and Yee W Teh. Bayesian learning via stochastic gradient langevin dynamics. In Proceedings of the 28th international conference on machine learning (ICML-11), pages 681–688, 2011.
  8. 8.Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. Sliced score matching: A scalable approach to density and score estimation. In Proceedings of the Thirty-Fifth Conference on Uncertainty in Artificial Intelligence, UAI 2019, Tel Aviv, Israel, July 22-25, 2019, page 204, 2019.
  9. 9.Alexia Jolicoeur-Martineau, Rémi Piché-Taillefer, Ioannis Mitliagkas, and Rémi Tachet des Combes. Adversarial score matching and improved sampling for image generation. arXiv preprint arXiv:2009.05475, 2020.
  10. 10.Saeed Saremi and Aapo Hyvarinen. Neural empirical bayes. Journal of Machine Learning Research, 20:1–23, 2019.
  11. 11.Zengyi Li, Yubei Chen, and Friedrich T Sommer. Learning energy-based models in high-dimensional spaces with multi-scale denoising score matching. arXiv, pages arXiv–1910, 2019.
  12. 12.Zahra Kadkhodaie and Eero P Simoncelli. Solving linear inverse problems using the prior implicit in a denoiser. arXiv preprint arXiv:2007.13640, 2020.
  13. 13.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in neural information processing systems, pages 6626–6637, 2017.
  14. 14.Bradley Efron. Tweedie’s formula and selection bias. Journal of the American Statistical Association, 106(496):1602–1614, 2011.
  15. 15.Erik W Grafarend. Linear and nonlinear models: fixed effects, random effects, and mixed models. de Gruyter, 2006.
  16. 16.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV), December 2015.
  17. 17.Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al. Conditional image generation with pixelcnn decoders. In Advances in neural information processing systems, pages 4790–4798, 2016.
  18. 18.Yilun Du and Igor Mordatch. Implicit generation and generalization in energy-based models. arXiv preprint arXiv:1903.08689, 2019.
  19. 19.Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville. Improved training of wasserstein gans. In Advances in neural information processing systems, pages 5767–5777, 2017.
  20. 20.Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. In International Conference on Learning Representations, 2018.
  21. 21.Shane Barratt and Rishi Sharma. A note on the inception score. arXiv preprint arXiv:1801.01973, 2018.
  22. 22.Mehdi SM Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly. Assessing generative models via precision and recall. In Advances in Neural Information Processing Systems, pages 5228–5237, 2018.
  23. 23.Ali Razavi, Aaron van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with vq-vae-2. In Advances in Neural Information Processing Systems, pages 14837–14847, 2019.
  24. 24.Guosheng Lin, Anton Milan, Chunhua Shen, and Ian Reid. Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1925–1934, 2017.
  25. 25.Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter. Fast and accurate deep network learning by exponential linear units (elus). arXiv preprint arXiv:1511.07289, 2015.
  26. 26.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  27. 27.Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015.
  28. 28.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4401–4410, 2019.
  29. 29.Sharon Zhou, Mitchell Gordon, Ranjay Krishna, Austin Narcomey, Li F Fei-Fei, and Michael Bernstein. Hype: A benchmark for human eye perceptual evaluation of generative models. In Advances in Neural Information Processing Systems, pages 3444–3456, 2019.
  30. 30.Jiaming Song and Stefano Ermon. Bridging the gap between f-gans and wasserstein gans. arXiv preprint arXiv:1910.09779, 2019.
  31. 31.Sashank J Reddi, Satyen Kale, and Sanjiv Kumar. On the convergence of adam and beyond. arXiv preprint arXiv:1904.09237, 2019.
  32. 32.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196, 2017.
  33. 33.David Berthelot, Thomas Schumm, and Luke Metz. Began: Boundary equilibrium generative adversarial networks. arXiv preprint arXiv:1703.10717, 2017.

Citation

MLA
Song, Y., and S. Ermon. “Improved Techniques for Training Score-Based Generative Models”. Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 12438–48, https://proceedings.neurips.cc/paper_files/paper/2020/file/92c3b916311a5517d9290576e3ea37ad-Paper.pdf.
APA
Song, Y., & Ermon, S. (2020). Improved Techniques for Training Score-Based Generative Models. Advances in Neural Information Processing Systems, 33, 12438–12448. https://proceedings.neurips.cc/paper_files/paper/2020/file/92c3b916311a5517d9290576e3ea37ad-Paper.pdf
Chicago
Song, Y., and S. Ermon. 2020. “Improved Techniques for Training Score-Based Generative Models”. Advances in Neural Information Processing Systems 33: 12438–48. https://proceedings.neurips.cc/paper_files/paper/2020/file/92c3b916311a5517d9290576e3ea37ad-Paper.pdf.
Harvard
Song, Y. and Ermon, S. (2020) “Improved Techniques for Training Score-Based Generative Models”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 12438–12448. Available at: https://proceedings.neurips.cc/paper_files/paper/2020/file/92c3b916311a5517d9290576e3ea37ad-Paper.pdf.
Vancouver
1. Song Y, Ermon S (2020) Improved Techniques for Training Score-Based Generative Models. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 12438–12448

BibTeX

@inproceedings{song2020improved,
  title = {Improved Techniques for Training Score-Based Generative Models},
  author = {Song, Yang and Ermon, Stefano},
  year = {2020},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {33},
  pages = {12438-12448},
  url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/92c3b916311a5517d9290576e3ea37ad-Paper.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors