Demystifying MMD GANs

Mikołaj BińkowskiDanica J. SutherlandMichael ArbelArthur Gretton

article2018ICLR2,177 citations

Introduces the Kernel Inception Distance metric for generative model evaluation while demonstrating that Maximum Mean Discrepancy GANs resolve critic gradient biases to train faster and with smaller architectures than Wasserstein GANs.

Listen

Generative adversarial networks are powerful machine learning tools used to generate realistic synthetic data, such as images, by training a generator to trick a discriminator or critic. However, these systems are notoriously difficult to train, frequently suffering from numerical instability and training failure. Recent efforts have sought to stabilize training by using integral probability metrics, such as the Wasserstein distance and Maximum Mean Discrepancy (MMD), as critic loss functions. Despite these advances, theoretical confusion has persisted regarding statistical bias in gradient estimators, and existing models often demand large discriminator networks that impose high computational burdens.

This article set out to rigorously characterize gradient bias across integral probability metric-based generative networks and demonstrate that MMD-based models can achieve equal or superior generation performance with smaller architectures and faster training. Additionally, it aimed to develop an unbiased, robust evaluation metric for generative models to replace biased legacy measures.

To evaluate these questions, the authors conducted mathematical analyses of gradient estimators across deep feedforward neural architectures and performed extensive empirical benchmarks. They tested MMD networks with various kernel functions alongside leading alternatives, specifically gradient-penalized Wasserstein networks and Cramér networks. The experimental validation was conducted across four standard image benchmark datasets (MNIST, CIFAR-10, LSUN Bedrooms, and CelebA) using standard convolutional and residual deep learning architectures.

The findings provide key theoretical clarifications and practical performance improvements. First, the authors prove that gradient estimators for both MMD and Wasserstein networks are unbiased for any fixed critic representation, but learning the critic from finite samples unavoidably introduces bias relative to the optimal population critic in both frameworks. Second, MMD networks using a rational quadratic kernel consistently outperform or match competing models while using significantly smaller discriminator networks; for instance, on CIFAR-10, an MMD model with a small critic achieved performance comparable to a Wasserstein network requiring four times as many convolutional filters. Third, on larger benchmarks like CelebA and LSUN, MMD networks attained superior image fidelity scores compared to Wasserstein and Cramér alternatives. Fourth, the authors introduced the Kernel Inception Distance (KID), demonstrating that, unlike the widely used Fréchet Inception Distance (FID), KID provides an unbiased estimator that converges reliably without misleading sample-size dependencies.

These results demonstrate that MMD generative networks function as efficient hybrid models, using initial neural network layers to project complex data into simpler feature representations where kernel methods can operate effectively. This architectural advantage allows teams to cut the computational cost of training roughly in half without sacrificing output quality. Furthermore, demonstrating that the standard evaluation metric (FID) is inherently biased means organizations evaluating generative models may make incorrect model selections unless they account for sample sizes or switch to unbiased metrics.

For practical application, practitioners should adopt MMD networks with rational quadratic kernels rather than Gaussian RBF kernels, which suffer from severe gradient decay. Teams should also adopt the Kernel Inception Distance both for benchmarking models and for automating learning rate decay schedules during training. Where computational budgets are constrained, engineering teams can safely deploy smaller discriminator networks within MMD frameworks to reduce training runtimes.

The conclusions are supported with high confidence by rigorous theoretical proofs and consistent empirical results across multiple visual domains. However, users should note that performance depends on the choice of kernel and the application of gradient regularization penalties. Further empirical analysis on broader, non-visual data modalities is recommended before standardizing these architectures across all production generative workloads.

arXiv: 1801.01401mbinkowski/MMD-GAN
  • Paper: Wasserstein Generative Adversarial Networks, Martin Arjovsky et al. (2017). This paper introduces the Wasserstein GAN framework that the source source-item builds upon and analyzes alongside Maximum Mean Discrepancy (MMD) GANs.
  • Paper: A Kernel Two-Sample Test, Arthur Gretton et al. (2012). This foundational text establishes the Maximum Mean Discrepancy statistical framework that serves as the core critic mechanism investigated in the source paper.
  • Paper: Spectral Normalization for Generative Adversarial Networks, Takeru Miyato et al. (2018). This paper details spectral normalization techniques for stabilizing GAN discriminators, providing essential context for the training strategies discussed in the source.
  • Paper: cGANs with Projection Discriminator, Takeru Miyato et al. (2018). This work extends adversarial discriminator architectures using projection methods that build directly upon the principles of metric-based critics.
  • Paper: Alias-Free Generative Adversarial Networks, Tero Karras et al. (2021). This paper advances generative adversarial network architectures by eliminating aliasing artifacts in generators, continuing the line of high-fidelity synthesis research.
Cover for Demystifying MMD GANs

Abstract

We investigate the training and performance of generative adversarial networks using the Maximum Mean Discrepancy (MMD) as critic, termed MMD GANs. As our main theoretical contribution, we clarify the situation with bias in GAN loss functions raised by recent work: we show that gradient estimators used in the optimization process for both MMD GANs and Wasserstein GANs are unbiased, but learning a discriminator based on samples leads to biased gradients for the generator parameters. We also discuss the issue of kernel choice for the MMD critic, and characterize the kernel corresponding to the energy distance used for the Cramer GAN critic. Being an integral probability metric, the MMD benefits from training strategies recently developed for Wasserstein GANs. In experiments, the MMD GAN is able to employ a smaller critic network than the Wasserstein GAN, resulting in a simpler and faster-training algorithm with matching performance. We also propose an improved measure of GAN convergence, the Kernel Inception Distance, and show how to use it to dynamically adapt learning rates during GAN training.

Table of Contents

  • 1 Introduction
  • 2 Losses and witness functions
  • 2.1 Maximum Mean Discrepancy and witness functions
  • 2.2 Witness function and gradient penalties
  • 2.3 The energy distance and associated MMD
  • 2.4 Other related models
  • 3 Gradient bias
  • 4 Evaluation metrics
  • 4.1 Learning rate adaptation
  • 5 Experiments
  • References
  • A Score functions, divergences, and the Cramér GAN
  • B Bias of generalized IPM estimators
  • B.1 Generalized IPMs and data-splitting estimators
  • B.2 Estimator bias
  • B.3 Gradient estimator bias
  • B.4 WGANs
  • B.5 Maximal MMD estimator
  • C Proof of unbiased gradients
  • C.1 Network definition
  • C.2 Assumptions
  • C.3 Main results
  • C.4 Bounds on network growth
  • C.5 Critical parameters have zero measure
  • D FID estimator bias
  • D.1 Analytic example with one-dimensional normals
  • D.2 Empirical example with high-dimensional censored normals
  • D.3 Non-existence of an unbiased estimator
  • E Comparison of evaluation metrics’ resilience to noise
  • F Samples and detailed results for MNIST and CIFAR-10

Knowls

  1. Knowl 1 — Unbiasedness of Deep Network Gradient Estimators for MMD and WGAN Losses

    theoretical result

    Let Gψ:Z→XG_\psi : \mathcal{Z} \to \mathcal{X} be a generator neural network with parameters ψ∈Rmψ\psi \in \mathbb{R}^{m_\psi}, and let hθ:X→Rdh_\theta : \mathcal{X} \to \mathbb{R}^d be a critic representation network with parameters θ∈Rmθ\theta \in \mathbb{R}^{m_\theta}. Assume GψG_\psi and hθh_\theta are standard feedforward networks composed of linear/convolutional layers and Lipschitz, piecewise-analytic activation functions (including ReLU, ELU, max pooling, and batch normalization).

    Let P\mathbb{P} be a target data distribution on X\mathcal{X} and Z\mathcal{Z} a noise distribution on Z\mathcal{Z} such that the second moments EX∼P[∥X∥2]<∞\mathbb{E}_{X \sim \mathbb{P}}[\|X\|^2] < \infty and EZ∼Z[∥Z∥2]<∞\mathbb{E}_{Z \sim \mathcal{Z}}[\|Z\|^2] < \infty exist. Let k:Rd×Rd→Rk : \mathbb{R}^d \times \mathbb{R}^d \to \mathbb{R} be a kernel function satisfying polynomial growth conditions:

    ∣k(u,v)∣≤C0(∥u∥2+∥v∥2)α/2+C0|k(u, v)| \le C_0 \left( \|u\|^2 + \|v\|^2 \right)^{\alpha/2} + C_0

    ∥∇u,vk(u,v)∥≤C1(∥u∥2+∥v∥2)(α−1)/2+C1\|\nabla_{u, v} k(u, v)\| \le C_1 \left( \|u\|^2 + \|v\|^2 \right)^{(\alpha - 1)/2} + C_1

    for constants C0,C1>0C_0, C_1 > 0 and α∈[1,2]\alpha \in [1, 2] (which holds for dot product, Gaussian RBF, rational quadratic, and Euclidean distance-induced kernels).

    Then, for Lebesgue-almost all (ψ,θ)∈Rmψ+mθ(\psi, \theta) \in \mathbb{R}^{m_\psi + m_\theta}, expectations and parameter gradients can be interchanged:

    EX∼Pm,Z∼Zn[∂ψ,θMMDu2(hθ(X),hθ(Gψ(Z)))]=∂ψ,θMMD2(hθ(P),hθ(Gψ(Z)))\mathbb{E}_{X \sim \mathbb{P}^m, Z \sim \mathcal{Z}^n} \left[ \partial_{\psi, \theta} \mathrm{MMD}_u^2\big(h_\theta(X), h_\theta(G_\psi(Z))\big) \right] = \partial_{\psi, \theta} \mathrm{MMD}^2\big(h_\theta(\mathbb{P}), h_\theta(G_\psi(\mathcal{Z}))\big)

    where MMDu2\mathrm{MMD}_u^2 is the standard unbiased sample estimator of squared Maximum Mean Discrepancy. For fixed critic parameters θ\theta, the generator gradient estimator evaluated on a minibatch is an unbiased estimator of the population gradient for both MMD GANs and Wasserstein GANs.

  2. Knowl 2 — Downward Estimator Bias and Non-Existence of Unbiased Estimators for IPMs in GANs

    theoretical result

    For a generalized Integral Probability Metric (IPM) defined as:

    D(P,Q)=sup⁡f∈FJ(f,P,Q)\mathcal{D}(\mathbb{P}, \mathbb{Q}) = \sup_{f \in \mathcal{F}} J(f, \mathbb{P}, \mathbb{Q})

    consider a data-splitting estimator D^(X,Y)\hat{\mathcal{D}}(X, Y) that splits samples X∼Pm,Y∼QnX \sim \mathbb{P}^m, Y \sim \mathbb{Q}^n into training sets (Xtr,Ytr)(X^{\mathrm{tr}}, Y^{\mathrm{tr}}) and test sets (Xte,Yte)(X^{\mathrm{te}}, Y^{\mathrm{te}}), selects a critic function f^Xtr,Ytr∈F\hat{f}_{X^{\mathrm{tr}}, Y^{\mathrm{tr}}} \in \mathcal{F}, and computes an estimator J^(f^Xtr,Ytr,Xte,Yte)\hat{J}(\hat{f}_{X^{\mathrm{tr}}, Y^{\mathrm{tr}}}, X^{\mathrm{te}}, Y^{\mathrm{te}}) where J^\hat{J} is unbiased for any fixed f∈Ff \in \mathcal{F}.

    Then either the critic selection procedure is almost surely optimal:

    Pr⁡ ⁣(J(f^Xtr,Ytr,P,Q)=D(P,Q))=1\Pr\!\left( J(\hat{f}_{X^{\mathrm{tr}}, Y^{\mathrm{tr}}}, \mathbb{P}, \mathbb{Q}) = \mathcal{D}(\mathbb{P}, \mathbb{Q}) \right) = 1

    or the estimator is strictly biased downwards:

    E[D^(X,Y)]<D(P,Q)\mathbb{E}\big[\hat{\mathcal{D}}(X, Y)\big] < \mathcal{D}(\mathbb{P}, \mathbb{Q})

    Furthermore, for any standard IPM DF(P,Q)=sup⁡f∈F(EPf(X)−EQf(Y))\mathcal{D}_\mathcal{F}(\mathbb{P}, \mathbb{Q}) = \sup_{f \in \mathcal{F}} (\mathbb{E}_\mathbb{P} f(X) - \mathbb{E}_\mathbb{Q} f(Y)) over a family of distributions P\mathcal{P} that contains mixture segments {(1−α)P0+αP1:0≤α≤1}\{(1 - \alpha)\mathbb{P}_0 + \alpha \mathbb{P}_1 : 0 \le \alpha \le 1\} for distinct P0≠P1\mathbb{P}_0 \neq \mathbb{P}_1, there exists no unbiased estimator of DF\mathcal{D}_\mathcal{F} on P\mathcal{P} for any finite sample size.

  3. Knowl 3 — Generator Gradient Bias Induced by Sample-Based Critic Learning

    theoretical result

    Let D(ψ)=D(P,Qψ)\mathcal{D}(\psi) = \mathcal{D}(\mathbb{P}, \mathbb{Q}_\psi) be a discrepancy measure between a target distribution P\mathbb{P} and a generated distribution Qψ=Gψ(Z)\mathbb{Q}_\psi = G_\psi(\mathcal{Z}) parametrized by generator weights ψ∈Ψ⊆Rmψ\psi \in \Psi \subseteq \mathbb{R}^{m_\psi}. Let D^(ψ)\hat{\mathcal{D}}(\psi) be an almost surely differentiable estimator of D(ψ)\mathcal{D}(\psi).

    If the parameter gradient estimator is unbiased, E[∇ψD^(ψ)]=∇ψD(ψ)\mathbb{E}[\nabla_\psi \hat{\mathcal{D}}(\psi)] = \nabla_\psi \mathcal{D}(\psi), then on each connected component of Ψ\Psi, the expectation of the estimator satisfies:

    E[D^(ψ)]=D(ψ)+c\mathbb{E}\big[\hat{\mathcal{D}}(\psi)\big] = \mathcal{D}(\psi) + c

    for some constant c∈Rc \in \mathbb{R}.

    Because learning critic parameters θ\theta from finite samples introduces a downward bias in the IPM estimate that varies non-constantly with the generator parameter ψ\psi, the generator gradient estimator ∇ψD^(ψ)\nabla_\psi \hat{\mathcal{D}}(\psi) computed using a learned critic is necessarily biased relative to the true gradient of the optimal population objective ∇ψsup⁡θDθ(P,Qψ)\nabla_\psi \sup_\theta \mathcal{D}_\theta(\mathbb{P}, \mathbb{Q}_\psi). This applies equally to Wasserstein GANs and MMD GANs (including energy distance critics).

  4. Knowl 4 — MMD GAN Formulation with Gradient-Constrained Critic Witness

    model/method

    The MMD GAN optimizes a generator GψG_\psi to minimize the Maximum Mean Discrepancy (MMD) computed on representations output by a critic convolutional feature network hθ:X→Rdh_\theta : \mathcal{X} \to \mathbb{R}^d. For samples X={xi}i=1m∼PmX = \{x_i\}_{i=1}^m \sim \mathbb{P}^m and generated samples Y={yj}j=1n∼QnY = \{y_j\}_{j=1}^n \sim \mathbb{Q}^n, the squared MMD is computed using the unbiased U-statistic estimator:

    MMDu2(hθ(X),hθ(Y))=1m(m−1)∑i≠jmk(hθ(xi),hθ(xj))+1n(n−1)∑i≠jnk(hθ(yi),hθ(yj))−2mn∑i=1m∑j=1nk(hθ(xi),hθ(yj))\mathrm{MMD}_u^2(h_\theta(X), h_\theta(Y)) = \frac{1}{m(m - 1)} \sum_{i \neq j}^m k\big(h_\theta(x_i), h_\theta(x_j)\big) + \frac{1}{n(n - 1)} \sum_{i \neq j}^n k\big(h_\theta(y_i), h_\theta(y_j)\big) - \frac{2}{mn} \sum_{i=1}^m \sum_{j=1}^n k\big(h_\theta(x_i), h_\theta(y_j)\big)

    To regularize the critic representation and prevent trivial feature collapse during adversarial training, a gradient penalty is applied directly to the empirical MMD witness function:

    f^(x)=1m∑i=1mk(hθ(xi),hθ(x))−1n∑j=1nk(hθ(yj),hθ(x))\hat{f}(x) = \frac{1}{m} \sum_{i=1}^m k\big(h_\theta(x_i), h_\theta(x)\big) - \frac{1}{n} \sum_{j=1}^n k\big(h_\theta(y_j), h_\theta(x)\big)

    by penalizing (∥∇x^f^(x^)∥2−1)2(\|\nabla_{\hat{x}} \hat{f}(\hat{x})\|_2 - 1)^2 at random convex interpolations x^=αxi+(1−α)yj\hat{x} = \alpha x_i + (1 - \alpha) y_j with α∼Uniform(0,1)\alpha \sim \mathrm{Uniform}(0, 1).

  5. Knowl 5 — Characteristic Kernel Choices for MMD GANs: Rational Quadratic and Mixed RQ-Dot

    model/method

    In MMD GANs, the choice of positive definite kernel kk on top of critic features hθ(x)h_\theta(x) determines training dynamics:

    1. Gaussian RBF Kernel:

      kσrbf(x,y)=exp⁡ ⁣(−∥x−y∥22σ2)k_\sigma^{\mathrm{rbf}}(x, y) = \exp\!\left(-\frac{\|x - y\|^2}{2\sigma^2}\right)

      Both the kernel values and its gradients decay exponentially with distance, causing vanishing gradients when generated samples are far from target data and leading to persistent blurriness.

    2. Rational Quadratic (RQ) Kernel:

      kαrq(x,y)=(1+∥x−y∥22α)−αk_\alpha^{\mathrm{rq}}(x, y) = \left(1 + \frac{\|x - y\|^2}{2\alpha}\right)^{-\alpha}

      for α>0\alpha > 0, equivalent to a mixture of Gaussian RBF kernels with a Gamma(α,1)\mathrm{Gamma}(\alpha, 1) prior on the inverse squared lengthscale. The RQ kernel is characteristic and has polynomial tail decay, ensuring sustained gradients across larger distances.

    3. Mixed RQ-Dot Kernel (krq∗k^{\mathrm{rq*}}): A multiscale mixture combined with a linear dot-product term:

      krq∗(x,y)=∑α∈{0.2,0.5,1,2,5}kαrq(x,y)+⟨x,y⟩k^{\mathrm{rq*}}(x, y) = \sum_{\alpha \in \{0.2, 0.5, 1, 2, 5\}} k_\alpha^{\mathrm{rq}}(x, y) + \langle x, y \rangle

      The dot-product component penalizes differences in feature means while the rational quadratic mixture matches higher-order moments across multiple lengthscales.

  6. Knowl 6 — Equivalence of Energy Distance to MMD and Degeneracy of Cramér GAN Surrogate Critic

    theoretical result

    The energy distance between distributions P\mathbb{P} and Q\mathbb{Q} with metric ρβ(x,y)=∥x−y∥β\rho_\beta(x, y) = \|x - y\|^\beta for 0<β≤20 < \beta \le 2:

    De(P,Q)=−12EX,X′∼P[ρβ(X,X′)]−12EY,Y′∼Q[ρβ(Y,Y′)]+EX∼P,Y∼Q[ρβ(X,Y)]\mathcal{D}_e(\mathbb{P}, \mathbb{Q}) = -\frac{1}{2} \mathbb{E}_{X, X' \sim \mathbb{P}}[\rho_\beta(X, X')] - \frac{1}{2} \mathbb{E}_{Y, Y' \sim \mathbb{Q}}[\rho_\beta(Y, Y')] + \mathbb{E}_{X \sim \mathbb{P}, Y \sim \mathbb{Q}}[\rho_\beta(X, Y)]

    is an instance of Maximum Mean Discrepancy with the distance-induced kernel:

    kρ,z0dist(x,y)=12[ρ(x,z0)+ρ(y,z0)−ρ(x,y)]k_{\rho, z_0}^{\mathrm{dist}}(x, y) = \frac{1}{2} \big[\rho(x, z_0) + \rho(y, z_0) - \rho(x, y)\big]

    The Cramér GAN surrogate critic modifies this objective by replacing Y′Y' with the origin 00, yielding witness function fc(x)=EX∼P[ρ(x,X)]−ρ(x,0)f_c(x) = \mathbb{E}_{X \sim \mathbb{P}}[\rho(x, X)] - \rho(x, 0) and expected critic loss:

    Dc(P,Q)=E[ρ(X,X′)]+E[ρ(Y,0)]−E[ρ(X,0)]−E[ρ(X′,Y)]\mathcal{D}_c(\mathbb{P}, \mathbb{Q}) = \mathbb{E}[\rho(X, X')] + \mathbb{E}[\rho(Y, 0)] - \mathbb{E}[\rho(X, 0)] - \mathbb{E}[\rho(X', Y)]

    Because of this replacement, Dc\mathcal{D}_c is not a proper metric: distinct distributions can have zero loss. For example, if P=δ0\mathbb{P} = \delta_0 (point mass at origin in R\mathbb{R}) and Q=δt\mathbb{Q} = \delta_t (point mass at t≠0t \neq 0), then P≠Q\mathbb{P} \neq \mathbb{Q} but Dc(P,Q)=0\mathcal{D}_c(\mathbb{P}, \mathbb{Q}) = 0 because E[ρ(X,X′)]=E[ρ(X,0)]=0\mathbb{E}[\rho(X, X')] = \mathbb{E}[\rho(X, 0)] = 0 and E[ρ(X′,Y)]=E[ρ(Y,0)]=t\mathbb{E}[\rho(X', Y)] = \mathbb{E}[\rho(Y, 0)] = t.

  7. Knowl 7 — Kernel Inception Distance Metric for Generative Models

    definition

    The Kernel Inception Distance (KID) is a metric for evaluating generative models, defined as the squared Maximum Mean Discrepancy between feature representations extracted from the pool3 layer (d=2048d = 2048) of a pre-trained Inception-v3 network:

    KID(P,Q)=MMD2(ϕ(P),ϕ(Q))\mathrm{KID}(\mathbb{P}, \mathbb{Q}) = \mathrm{MMD}^2\big(\phi(\mathbb{P}), \phi(\mathbb{Q})\big)

    using a cubic polynomial kernel:

    k(x,y)=(1dxTy+1)3k(x, y) = \left( \frac{1}{d} x^T y + 1 \right)^3

    where d=2048d = 2048 is the representation dimension. KID evaluates differences in the mean, variance, and skewness of feature distributions without assuming a parametric form (unlike the Fréchet Inception Distance, which assumes activations follow Gaussian distributions). KID has an unbiased U-statistic estimator MMDu2\mathrm{MMD}_u^2, asymptotically normal error distributions, and computational complexity O(n2d)O(n^2 d) for sample size nn.

  8. Knowl 8 — Estimator Bias and Ranking Inversion in Fréchet Inception Distance

    theoretical result

    The Fréchet Inception Distance (FID) between distributions P\mathbb{P} and Q\mathbb{Q} with means μP,μQ\mu_\mathbb{P}, \mu_\mathbb{Q} and covariance matrices ΣP,ΣQ\Sigma_\mathbb{P}, \Sigma_\mathbb{Q} is:

    FID(P,Q)=∥μP−μQ∥2+Tr(ΣP)+Tr(ΣQ)−2Tr ⁣((ΣPΣQ)1/2)\mathrm{FID}(\mathbb{P}, \mathbb{Q}) = \|\mu_\mathbb{P} - \mu_\mathbb{Q}\|^2 + \mathrm{Tr}(\Sigma_\mathbb{P}) + \mathrm{Tr}(\Sigma_\mathbb{Q}) - 2\mathrm{Tr}\!\left( (\Sigma_\mathbb{P} \Sigma_\mathbb{Q})^{1/2} \right)

    The standard plug-in estimator evaluated on samples X∼Pm,Y∼QnX \sim \mathbb{P}^m, Y \sim \mathbb{Q}^n is biased. For one-dimensional Gaussians P=N(μP,σP2)\mathbb{P} = \mathcal{N}(\mu_\mathbb{P}, \sigma_\mathbb{P}^2) and Q=N(μQ,σQ2)\mathbb{Q} = \mathcal{N}(\mu_\mathbb{Q}, \sigma_\mathbb{Q}^2), the expected plug-in estimate is:

    E[FID(P^X,Q^Y)]=(μP−μQ)2+m+1mσP2+n+1nσQ2−2dmdnσPσQ\mathbb{E}\big[\mathrm{FID}(\hat{\mathbb{P}}_X, \hat{\mathbb{Q}}_Y)\big] = (\mu_\mathbb{P} - \mu_\mathbb{Q})^2 + \frac{m + 1}{m}\sigma_\mathbb{P}^2 + \frac{n + 1}{n}\sigma_\mathbb{Q}^2 - 2 d_m d_n \sigma_\mathbb{P} \sigma_\mathbb{Q}

    where dm=2Γ(m/2)m−1Γ((m−1)/2)<1d_m = \frac{\sqrt{2}\Gamma(m/2)}{\sqrt{m-1}\Gamma((m-1)/2)} < 1.

    Because dm<1d_m < 1, the bias can cause ranking inversions where FID(P1,Q)>FID(P2,Q)\mathrm{FID}(\mathbb{P}_1, \mathbb{Q}) > \mathrm{FID}(\mathbb{P}_2, \mathbb{Q}) but E[FID(P^1,Q)]<E[FID(P^2,Q)]\mathbb{E}[\mathrm{FID}(\hat{\mathbb{P}}_{1}, \mathbb{Q})] < \mathbb{E}[\mathrm{FID}(\hat{\mathbb{P}}_{2}, \mathbb{Q})] for finite sample size mm. Furthermore, no unbiased estimator for FID can exist over any class of distributions that includes two-component Gaussian mixtures.

  9. Knowl 9 — Dynamic Learning Rate Adaptation via Relative KID Similarity Test

    algorithm

    Dynamic learning rate reduction during GAN training uses a statistical hypothesis test of relative similarity to determine whether generator sample quality has plateaued against a validation set.

    Input: Target validation samples YvalY_{\mathrm{val}}, initial learning rate η0\eta_0, evaluation interval TT, reference step lag ΔT\Delta T, failure threshold KK, significance level α\alpha
    Output: Trained generator parameters ψ\psi
    failures ←0\leftarrow 0
    η←η0\eta \leftarrow \eta_0
    for step t=1,2,…t = 1, 2, \dots do
        Update critic and generator parameters using learning rate η\eta
        if t mod T=0t \bmod T = 0 and t>ΔTt > \Delta T then
            Sample Xt∼Gψt(Z)X_t \sim G_{\psi_t}(\mathcal{Z}) from current generator
            Sample Xprev∼Gψt−ΔT(Z)X_{\mathrm{prev}} \sim G_{\psi_{t - \Delta T}}(\mathcal{Z}) from previous generator
            Compute pp-value pp for null hypothesis H0:KID(Xt,Yval)≥KID(Xprev,Yval)H_0: \mathrm{KID}(X_t, Y_{\mathrm{val}}) \ge \mathrm{KID}(X_{\mathrm{prev}}, Y_{\mathrm{val}})
            if p>αp > \alpha then
                failures ←\leftarrow failures +1+ 1
            else
                failures ←0\leftarrow 0
            end if
            if failures ≥K\ge K then
                η←η/2\eta \leftarrow \eta / 2
                failures ←0\leftarrow 0
            end if
        end if
    end for

    In standard benchmark configurations, hyperparameters are set to T=2000T = 2000 generator steps (T=500T = 500 for MNIST), ΔT=20000\Delta T = 20000 steps (ΔT=5000\Delta T = 5000 for MNIST), failure threshold K=3K = 3, and initial learning rate η0=10−4\eta_0 = 10^{-4}.

  10. Knowl 10 — Generative Performance on LSUN Bedrooms and CelebA Across Critic Capacities

    data/table

    Generative evaluation scores on 64×6464 \times 64 LSUN Bedrooms and 160×160160 \times 160 CelebA datasets comparing MMD GANs (rational quadratic krqk^{\mathrm{rq}} or mixed krq∗k^{\mathrm{rq*}} kernel and distance kernel kdistk^{\mathrm{dist}}), Cramér GANs, and WGAN-GP across critic architectures (16 vs 64 convolutional filters in the first layer; 16 vs 256 neurons in the top feature layer).

    Dataset Loss Filters Top Layer FID KID
    LSUN rqrq 16 16 86.47 (0.29) 0.091 (0.002)
    LSUN rqrq 64 16 31.95 (0.28) 0.028 (0.002)
    LSUN distdist 16 256 104.85 (0.32) 0.109 (0.002)
    LSUN distdist 64 256 35.28 (0.21) 0.032 (0.001)
    LSUN Cramér GAN 16 256 122.03 (0.41) 0.132 (0.002)
    LSUN Cramér GAN 64 256 54.18 (0.39) 0.050 (0.002)
    LSUN WGAN-GP 16 1 292.77 (0.35) 0.370 (0.003)
    LSUN WGAN-GP 64 1 41.39 (0.25) 0.039 (0.002)
    LSUN Test Set – – 2.49 (0.02) 0.000 (0.000)
    CelebA rq∗rq^* 64 16 20.55 (0.25) 0.013 (0.001)
    CelebA Cramér GAN 64 256 31.30 (0.17) 0.025 (0.001)
    CelebA WGAN-GP 64 1 29.24 (0.22) 0.022 (0.001)
    CelebA Test Set – – 2.25 (0.04) 0.000 (0.000)

    These results demonstrate that MMD GAN with rational quadratic kernels outperforms Cramér GAN and WGAN-GP at matched critic capacities on complex image datasets. On LSUN Bedrooms, when the critic filter count is reduced from 64 to 16, MMD GAN maintains reasonable performance (FID 86.47, KID 0.091), whereas WGAN-GP fails completely (FID 292.77, KID 0.370). On CelebA (160×160160 \times 160), MMD GAN (rq∗rq^*) achieves superior FID (20.55 vs 29.24) and KID (0.013 vs 0.022) compared to WGAN-GP.

  11. Knowl 11 — Quantitative Generative Performance on CIFAR-10 and MNIST Benchmarks

    data/table

    Evaluation scores on 32×3232 \times 32 CIFAR-10 and 28×2828 \times 28 MNIST benchmark datasets across critic filter sizes (16 vs 64 filters in the first convolutional layer) and kernel variants.

    Dataset Loss Filters Top Layer Inception FID KID
    CIFAR-10 rqrq 16 16 5.86 (0.06) 48.10 (0.16) 0.032 (0.001)
    CIFAR-10 rqrq 64 16 6.51 (0.03) 39.90 (0.29) 0.027 (0.001)
    CIFAR-10 distdist 16 256 4.53 (0.03) 80.48 (0.19) 0.061 (0.001)
    CIFAR-10 distdist 64 256 6.39 (0.04) 40.25 (0.19) 0.028 (0.001)
    CIFAR-10 Cramér GAN 16 256 4.67 (0.02) 74.93 (0.32) 0.060 (0.001)
    CIFAR-10 Cramér GAN 64 256 6.39 (0.01) 40.27 (0.15) 0.028 (0.001)
    CIFAR-10 WGAN-GP 16 1 3.15 (0.01) 147.09 (0.31) 0.116 (0.002)
    CIFAR-10 WGAN-GP 64 1 6.53 (0.02) 37.52 (0.19) 0.026 (0.001)
    CIFAR-10 Test Set – – 11.21 (0.13) 6.11 (0.05) 0.000 (0.000)
    MNIST rqrq 16 16 9.11 (0.01) 4.206 (0.05) 0.005 (0.004)
    MNIST rbfrbf 16 16 8.98 (0.02) 8.264 (0.02) 0.011 (0.006)
    MNIST dotdot 16 16 8.86 (0.02) 6.245 (0.06) 0.006 (0.004)
    MNIST distdist 16 256 9.13 (0.004) 6.179 (0.05) 0.005 (0.004)
    MNIST Cramér GAN 16 256 9.25 (0.02) 3.385 (0.10) 0.006 (0.005)
    MNIST WGAN-GP 16 1 9.12 (0.02) 6.915 (0.10) 0.009 (0.004)
    MNIST Test Set – – 9.78 (0.02) 4.305 (0.16) 0.003 (0.003)

    On CIFAR-10, an MMD GAN with a small 16-filter critic achieves an Inception score of 5.86 and FID of 48.10, approaching the large 64-filter WGAN-GP (6.53 Inception, 37.52 FID) while requiring approximately half the training runtime per step. With 16 filters, WGAN-GP degrades severely (Inception 3.15, FID 147.09, KID 0.116). On MNIST, Gaussian RBF (rbfrbf) exhibits the worst FID (8.264) among MMD variants due to gradient decay in low-density regions.

Coverage note — Omitted technical measure-theoretic supporting lemmas (Lemmas 1 through 7) used exclusively to prove Theorem 5, as well as the empirical noise perturbation curves from Appendix E, to focus on the self-contained primary theoretical and empirical contributions.

References

  1. 1.M. Arjovsky and L. Bottou. Towards principled methods for training generative adversarial networks. In ICLR, 2017. arXiv:1701.04862.
  2. 2.M. Arjovsky, S. Chintala, and L. Bottou. Wasserstein generative adversarial networks. In ICML, 2017. arXiv:1701.07875.
  3. 3.S. Arora and Y. Zhang. Do GANs actually learn the distribution? An empirical study, 2017. arXiv:1706.08224.
  4. 4.S. Arora, R. Ge, Y. Liang, T. Ma, and Y. Zhang. Generalization and equilibrium in generative adversarial nets (GANs). In ICML, 2017. arXiv:1703.00573.
  5. 5.M. G. Bellemare, I. Danihelka, W. Dabney, S. Mohamed, B. Lakshminarayanan, S. Hoyer, and R. Munos. The Cramer distance as a solution to biased Wasserstein gradients, 2017. arXiv:1705.10743.
  6. 6.Y. Bengio, G. Mesnil, Y. Dauphin, and S. Rifai. Better mixing via deep representations. In ICML, 2013. arXiv:1207.4404.
  7. 7.D. Berthelot, T. Schumm, and L. Metz. BEGAN: Boundary equilibrium generative adversarial networks, 2017. arXiv:1703.10717.
  8. 8.P. J. Bickel and E. L. Lehmann. Unbiased estimation in convex families. The Annals of Mathematical Statistics, 40(5):1523–1535, 1969.
  9. 9.D. Bouchacourt, P. K. Mudigonda, and S. Nowozin. DISCO nets: DISsimilarity COefficients networks. In NIPS, pp. 352–360. 2016.
  10. 10.W. Bounliphone, E. Belilovsky, M. B. Blaschko, I. Antonoglou, and A. Gretton. A test of relative similarity for model selection in generative models. In ICLR, 2016. arXiv:1511.04581.
  11. 11.D.-A. Clevert, T. Unterthiner, and S. Hochreiter. Fast and accurate deep network learning by exponential linear units (ELUs). In ICLR, 2016. arXiv:1511.07289.
  12. 12.I. Danihelka, B. Lakshminarayanan, B. Uria, D. Wierstra, and P. Dayan. Comparison of maximum likelihood and GAN-based training of Real NVPs, 2017. arXiv:1705.05263.
  13. 13.G. K. Dziugaite, D. M. Roy, and Z. Ghahramani. Training generative neural networks via maximum mean discrepancy optimization. In UAI, 2015. arXiv:1505.03906.
  14. 14.W. Fedus, M. Rosca, B. Lakshminarayanan, A. M. Dai, S. Mohamed, and I. Goodfellow. Many paths to equilibrium: GANs do not need to decrease a divergence at every step. In ICLR, 2018. arXiv:1710.08446.
  15. 15.T. Gneiting and A. E. Raftery. Strictly proper scoring rules, prediction, and estimation. JASA, 102(477):359–378, 2007.
  16. 16.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In NIPS, 2014. arXiv:1406.2661.
  17. 17.A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. J. Smola. A kernel two-sample test. JMLR, 13, 2012.
  18. 18.I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. Courville. Improved training of Wasserstein GANs. In NIPS, 2017. arXiv:1704.00028.
  19. 19.M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, G. Klambauer, and S. Hochreiter. GANs trained by a two time-scale update rule converge to a Nash equilibrium. In NIPS, 2017. arXiv:1706.08500.
  20. 20.R. Huang, S. Zhang, T. Li, and R. He. Beyond face rotation: Global and local perception GAN for photorealistic and identity preserving frontal view synthesis. In ICCV, 2017a. arXiv:1704.04086.
  21. 21.X. Huang, Y. Li, O. Poursaeed, J. Hopcroft, and S. Belongie. Stacked generative adversarial networks. In CVPR, 2017b. arXiv:1612.04357.
  22. 22.Y. Jin, K. Zhang, M. Li, Y. Tian, H. Zhu, and Z. Fang. Towards the automatic anime characters creation with generative adversarial networks, 2017. arXiv:1708.05509.
  23. 23.D. Kingma and J. Ba. Adam: A method for stochastic optimization. In ICLR, 2015. arXiv:1412.6980.
  24. 24.A. Klenke. Probability Theory: A Comprehensive Course. World Publishing Corporation, 2008.
  25. 25.A. Krizhevsky. Learning multiple layers of features from tiny images, 2009.
  26. 26.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 1998.
  27. 27.C. Li, D. Alvarez-Melis, K. Xu, S. Jegelka, and S. Sra. Distributional adversarial networks, 2017a. arXiv:1706.09549.
  28. 28.C.-L. Li, W.-C. Chang, Y. Cheng, Y. Yang, and B. Póczos. MMD GAN: Towards deeper understanding of moment matching network. In NIPS, 2017b. arXiv:1705.08584.
  29. 29.Y. Li, K. Swersky, and R. Zemel. Generative moment matching networks. In ICML, 2015. arXiv:1502.02761.
  30. 30.L. Liu. On the two-sample statistic approach to generative adversarial networks. Master’s thesis, University of Princeton Senior Thesis, April 2017. URL http://arks.princeton.edu/ark:/88435/dsp0179408079v.
  31. 31.S. Liu, O. Bousquet, and K. Chaudhuri. Approximation and convergence properties of generative adversarial learning. In NIPS, 2017. arXiv:1705.08991.
  32. 32.Z. Liu, P. Luo, X. Wang, and X. Tang. Deep learning face attributes in the wild. In ICCV, 2015.
  33. 33.D. Lopez-Paz and M. Oquab. Revisiting classifier two-sample tests. In ICLR, 2017. arXiv:1610.06545.
  34. 34.R. Lyons. Distance covariance in metric spaces. The Annals of Probability, 41(5):3051–3696, 2013.
  35. 35.B. Mityagin. The zero set of a real analytic function, 2015. arXiv:1512.07276.
  36. 36.Y. Mroueh and T. Sercu. Fisher GAN. In NIPS, 2017. arXiv:1705.09675.
  37. 37.Y. Mroueh, T. Sercu, and V. Goel. McGan: Mean and covariance feature matching GAN. In ICML, 2017. arXiv:1702.08398.
  38. 38.A. Müller. Integral probability metrics and their generating classes of functions. Advances in Applied Probability, 29(2):429–443, 1997.
  39. 39.S. Nowozin, B. Cseke, and R. Tomioka. f-GAN: Training generative neural samplers using variational divergence minimization. In NIPS, 2016. arXiv:1606.00709.
  40. 40.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. JMLR, 12:2825–2830, 2011.
  41. 41.G. Piranian. The Set of Nondifferentiability of a Continuous Function. The American Mathematical Monthly, 73(4):57–61, 1966.
  42. 42.A. Radford, L. Metz, and S. Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks. In ICLR, 2016. arXiv:1511.06434.
  43. 43.C. E. Rasmussen and C. K. I. Williams. Gaussian Processes for Machine Learning. MIT Press, Cambridge, MA, 2006.
  44. 44.S. Rosenbaum. Moments of a truncated bivariate normal distribution. JRSS B, 23:405–408, 1961.
  45. 45.T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen. Improved techniques for training GANs. In NIPS, 2016. arXiv:1606.03498.
  46. 46.D. Sejdinovic, B. K. Sriperumbudur, A. Gretton, and K. Fukumizu. Equivalence of distance-based and RKHS-based statistics in hypothesis testing. The Annals of Stastistics, 41(5):2263–2291, 2013. arXiv:1207.6076.
  47. 47.B. K. Sriperumbudur, K. Fukumizu, A. Gretton, G. R. G. Lanckriet, and B. Schölkopf. Kernel choice and classifiability for RKHS embeddings of probability distributions. In NIPS, 2009a.
  48. 48.B. K. Sriperumbudur, K. Fukumizu, A. Gretton, B. Schölkopf, and G. R. G. Lanckriet. On integral probability metrics, phi-divergences and binary classification, 2009b. arXiv:0901.2698.
  49. 49.B. K. Sriperumbudur, A. Gretton, K. Fukumizu, G. R. G. Lanckriet, and B. Schölkopf. Hilbert space embeddings and metrics on probability measures. JMLR, 11:1517–1561, 2010. arXiv:0907.5309.
  50. 50.B. K. Sriperumbudur, K. Fukumizu, and G. R. G. Lanckriet. Universality, characteristic kernels and RKHS embedding of measures. JMLR, 12:2389–2410, 2011. arXiv:1003.0887.
  51. 51.B. K. Sriperumbudur, K. Fukumizu, A. Gretton, B. Schölkopf, and G. R. G. Lanckriet. On the empirical estimation of integral probability metrics. Electronic Journal of Statistics, 6:1550–1599, 2012.
  52. 52.I. Steinwart and A. Christmann. Support Vector Machines. Information Science and Statistics. Springer, 2008.
  53. 53.D. J. Sutherland. What are the mean and variance of a 0-censored multivariate normal? Cross Validated answer, 2018. URL https://stats.stackexchange.com/q/326347.
  54. 54.D. J. Sutherland, H.-Y. Tung, H. Strathmann, S. De, A. Ramdas, A. Smola, and A. Gretton. Generative models and model criticism via optimized maximum mean discrepancy. In International Conference on Learning Representations, 2017. arXiv:1611.04488.
  55. 55.C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus. Intriguing properties of neural networks. In ICLR, 2014. arXiv:1312.6199.
  56. 56.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the Inception architecture for computer vision. In CVPR, 2016. arXiv:1512.00567.
  57. 57.G. Székely and M. Rizzo. Testing for equal distributions in high dimension. InterStat, 5, 2004.
  58. 58.L. Theis, A. van den Oord, and M. Bethge. A note on the evaluation of generative models. In ICLR, 2016. arXiv:1511.01844.
  59. 59.F. Yu, Y. Zhang, S. Song, A. Seff, and J. Xiao. LSUN: Construction of a large-scale image dataset using deep learning with humans in the loop, 2015. arXiv:1506.03365.
  60. 60.Z. Zahorski. Sur l’ensemble des points de non-dérivabilité d’une fonction continue. Bulletin de la Société mathématique de France, 2:147–178, 1946.
  61. 61.W. Zaremba, A. Gretton, and M. B. Blaschko. B-tests: Low variance kernel two-sample tests. In NIPS, 2013. arXiv:1307.1954.
  62. 62.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV, 2017. arXiv:1703.10593.

Citation

MLA
Bińkowski, M., et al. “Demystifying MMD GANs”. arXiv, 2018, http://arxiv.org/abs/1801.01401v5.
APA
Bińkowski, M., Sutherland, D. J., Arbel, M., & Gretton, A. (2018). Demystifying MMD GANs. arXiv. http://arxiv.org/abs/1801.01401v5
Chicago
Bińkowski, M., D. J. Sutherland, M. Arbel, and A. Gretton. 2018. “Demystifying MMD GANs”. arXiv. http://arxiv.org/abs/1801.01401v5.
Harvard
Bińkowski, M. et al. (2018) “Demystifying MMD GANs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1801.01401v5.
Vancouver
1. Bińkowski M, Sutherland DJ, Arbel M, Gretton A (2018) Demystifying MMD GANs. arXiv

BibTeX

@article{binkowski2018demystifying,
  title = {Demystifying MMD GANs},
  author = {Bińkowski, Mikołaj and Sutherland, Danica J. and Arbel, Michael and Gretton, Arthur},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1801.01401v5},
  eprint = {1801.01401}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF