Generalization in diffusion models arises from geometry-adaptive harmonic representations

Zahra KadkhodaieFlorentin GuthEero P. SimoncelliStéphane Mallat

article2024ICLR162 citationsOutstanding Paper Award

Demonstrates that diffusion models generalize rather than memorize because neural network denoisers naturally learn geometry-adaptive harmonic bases that perform near-optimal shrinkage operations.

Listen

Generative diffusion models have achieved state-of-the-art results in generating complex, high-dimensional data such as photographic images. However, approximating a continuous high-dimensional probability density typically requires an exponentially large volume of data, leading to concerns that these models succeed merely by memorizing and replicating their training sets rather than learning true underlying data distributions.

The article investigates whether deep neural network denoisers used in diffusion models truly generalize or simply memorize training data, and evaluates the specific internal mathematical mechanisms—known as inductive biases—that enable them to learn effectively from feasible sample sizes.

To assess generalization, the authors trained bias-free neural network architectures (UNet and BF-CNN) on disjoint subsets of benchmark image datasets across varying sample sizes and noise levels. They analyzed convergence by comparing generated outputs from models trained on non-overlapping data, and performed an eigendecomposition of the denoisers' input-output mappings (Jacobians) across real photographs, synthetic geometric image classes with known theoretical limits, and low-dimensional manifold datasets.

The article establishes several key findings. First, diffusion models exhibit a clear phase transition: at small sample sizes (e.g., 1 to 100 images), networks strictly memorize training examples, but at larger sizes (around 100,000 images), models trained on completely separate datasets converge to essentially the same score function and generate nearly identical, novel images. Second, network denoisers operate by performing shrinkage operations in geometry-adaptive harmonic bases, which naturally form oscillating harmonic structures aligned along image contours and across smooth regions. Third, when evaluated on synthetic geometric image classes where these harmonic bases are theoretically optimal, the neural networks achieve near-optimal theoretical performance rates. Finally, when trained on low-dimensional curved manifolds or scrambled image data where harmonic representations are suboptimal, the networks still impose these harmonic structures, demonstrating that this adaptive harmonic representation is an inherent inductive bias of the architecture.

These findings provide strong evidence that high-performing diffusion models genuinely generalize rather than memorize data, provided training sets reach sufficient scale. This reduces legal, compliance, and performance risks associated with data memorization and replication. Furthermore, the results demonstrate that the success of diffusion models stems from architectural biases that closely match the physical geometry and structure of natural images, rather than pure brute-force statistical estimation.

Organizations developing or deploying diffusion models should monitor model variance and sample-size thresholds to ensure operations remain in the true generalization regime. When adapting these models to non-image or unstructured data domains, teams should anticipate performance trade-offs if the domain structure conflicts with the model's inherent geometric and harmonic biases. Further research is recommended to mathematically formalize how convolutional layers and non-linearities interact to produce these bases across different architectures and modalities.

The empirical findings rely primarily on downsampled image datasets and specific convolutional network designs without additive bias parameters. While confidence in the demonstrated transition from memorization to generalization is high for standard image data, readers should exercise caution when extending these conclusions to non-visual data types or vastly different network architectures.

Cover for Generalization in diffusion models arises from geometry-adaptive harmonic representations

Abstract

Deep neural networks (DNNs) trained for image denoising are able to generate high-quality samples with score-based reverse diffusion algorithms. These impressive capabilities seem to imply an escape from the curse of dimensionality, but recent reports of memorization of the training set raise the question of whether these networks are learning the "true" continuous density of the data. Here, we show that two DNNs trained on non-overlapping subsets of a dataset learn nearly the same score function, and thus the same density, when the number of training images is large enough. In this regime of strong generalization, diffusion-generated images are distinct from the training set, and are of high visual quality, suggesting that the inductive biases of the DNNs are well-aligned with the data density. We analyze the learned denoising functions and show that the inductive biases give rise to a shrinkage operation in a basis adapted to the underlying image. Examination of these bases reveals oscillating harmonic structures along contours and in homogeneous regions. We demonstrate that trained denoisers are inductively biased towards these geometry-adaptive harmonic bases since they arise not only when the network is trained on photographic images, but also when it is trained on image classes supported on low-dimensional manifolds for which the harmonic basis is suboptimal. Finally, we show that when trained on regular image classes for which the optimal basis is known to be geometry-adaptive and harmonic, the denoising performance of the networks is near-optimal.

Table of Contents

  • 1 Introduction
  • 2 Diffusion model variance and denoising generalization
  • 2.1 Diffusion models and denoising
  • 2.2 Transition from memorization to generalization
  • 3 Inductive biases
  • 3.1 Denoising as shrinkage in an adaptive basis
  • 3.2 Geometry-adaptive harmonic bases in DNNs
  • 4 Discussion
  • References
  • A Experimental details
  • A.1 Training and architecture details
  • A.2 Sampling algorithm
  • B Additional numerical results on generalization
  • B.1 Similarity between data subsets
  • B.2 Generalization of UNet model
  • B.2.1 Trained on CelebA dataset
  • B.2.2 Trained on LSUN bedroom dataset
  • B.3 Generalization of BF-CNN model
  • B.3.1 Trained on CelebA dataset
  • B.3.2 Trained on LSUN bedroom dataset
  • B.4 Convergence as a function of training set size NN and image resolution
  • C Additional numerical results on inductive biases
  • C.1 More 𝐂α{\bf C}^{\alpha} examples
  • C.2 Additional low-dimensional manifold examples
  • C.3 Shuffled faces
  • D Mathematical derivations
  • D.1 Miyasawa relationships
  • D.2 Control on Kullback-Leibler divergence
  • D.3 SURE objective
  • D.4 Optimal thresholding in a basis
  • E Geometric 𝐂α{\bf C}^{\alpha} images

Citation

MLA
Kadkhodaie, Z., et al. “Generalization in Diffusion Models Arises from Geometry-adaptive Harmonic Representations”. Int'l Conf on Learning Representations (ICLR), Vol.12, Vienna, May 2024. Outstanding Paper Award, 2023, http://arxiv.org/abs/2310.02557v3.
APA
Kadkhodaie, Z., Guth, F., Simoncelli, E. P., & Mallat, S. (2023). Generalization in diffusion models arises from geometry-adaptive harmonic representations. Int'l Conf on Learning Representations (ICLR), Vol.12, Vienna, May 2024. Outstanding Paper Award. http://arxiv.org/abs/2310.02557v3
Chicago
Kadkhodaie, Z., F. Guth, E. P. Simoncelli, and S. Mallat. 2023. “Generalization in Diffusion Models Arises from Geometry-adaptive Harmonic Representations”. Int'l Conf on Learning Representations (ICLR), Vol.12, Vienna, May 2024. Outstanding Paper Award. http://arxiv.org/abs/2310.02557v3.
Harvard
Kadkhodaie, Z. et al. (2023) “Generalization in diffusion models arises from geometry-adaptive harmonic representations”, Int'l Conf on Learning Representations (ICLR), vol.12, Vienna, May 2024. Outstanding Paper award [Preprint]. Available at: http://arxiv.org/abs/2310.02557v3.
Vancouver
1. Kadkhodaie Z, Guth F, Simoncelli EP, Mallat S (2023) Generalization in diffusion models arises from geometry-adaptive harmonic representations. Int'l Conf on Learning Representations (ICLR), vol.12, Vienna, May 2024. Outstanding Paper award

BibTeX

@article{kadkhodaie2023generalization,
  title = {Generalization in diffusion models arises from geometry-adaptive harmonic representations},
  author = {Kadkhodaie, Zahra and Guth, Florentin and Simoncelli, Eero P. and Mallat, Stéphane},
  year = {2023},
  journal = {Int'l Conf on Learning Representations (ICLR), vol.12, Vienna, May 2024. Outstanding Paper award},
  url = {http://arxiv.org/abs/2310.02557v3},
  eprint = {2310.02557}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors