Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Wasserstein distance

The Wasserstein distance, also known as the earth mover distance or Kantorovich-Rubinstein metric, is a measure of the distance between two probability distributions defined on a metric space. Derived from optimal transport theory, it quantifies the minimum total cost needed to transform one probability distribution into another, where cost is calculated as the amount of probability mass transported multiplied by the distance it travels across the underlying space. Unlike information-theoretic divergences that evaluate probability distributions through pointwise density comparisons, the Wasserstein distance incorporates the geometric structure of the domain, remaining continuous and informative even when comparing distributions with non-overlapping supports. It satisfies all mathematical criteria of a true metric, including symmetry and the triangle inequality, making it widely used across probability theory, statistics, and machine learning for tasks such as generative modeling and domain adaptation.

11 items

Improved Analysis of Score-based Generative Modeling: User-Friendly Bounds under Minimal Smoothness Assumptions

Improved Analysis of Score-based Generative Modeling: User-Friendly Bounds under Minimal Smoothness Assumptions

Hongrui Chen, Holden Lee, Jianfeng Lu

OrganizationsDuke UniversityJohns Hopkins UniversityPeking University

Why you should read this

Establishes tight polynomial convergence guarantees for score-based generative models assuming only an accurate L2L^2 score estimator and finite second moments, eliminating restrictive structural assumptions like log-concavity while offering concrete guidance on discretization schemes.

We give an improved theoretical analysis of score-based generative modeling. Under a score estimate with small L^2 error (averaged across timesteps), we provide efficient convergence guarantees for any data distribution with second-order moment, by either employing early stopping or assuming a smoothness condition on the score function of the data distribution. Our result does not rely on any log-concavity or functional inequality assumption and has a logarithmic dependence on the smoothness. In particular, we show that under only a finite second moment condition, approximating the following in reverse KL divergence in ε-accuracy can be done in Õ(d log(1/δ)/ε) steps: 1) the variance-δ Gaussian perturbation of any data distribution; 2) data distributions with 1/δ-smooth score functions. Our analysis also provides a quantitative comparison between different discrete approximations and may guide the choice of discretization points in practice.

Added

2026-09-28

Multisample Flow Matching: Straightening Flows with Minibatch Couplings

Multisample Flow Matching: Straightening Flows with Minibatch Couplings

Aram-Alexandre Pooladian, Heli Ben-Hamu, Carles Domingo-Enrich, Brandon Amos, Yaron Lipman, Ricky T. Q. Chen

OrganizationsMetaNew York UniversityWeizmann Institute of Science

Why you should read this

Proposes Multisample Flow Matching, a simulation-free training framework that couples minibatch data and noise distributions to straighten probability paths, reducing gradient variance during training and enabling faster generative sampling with fewer model evaluations.

Simulation-free methods for training continuous-time generative models construct probability paths that go between noise distributions and individual data samples. Recent works, such as Flow Matching, derived paths that are optimal for each data sample. However, these algorithms rely on independent data and noise samples, and do not exploit underlying structure in the data distribution for constructing probability paths. We propose Multisample Flow Matching, a more general framework that uses non-trivial couplings between data and noise samples while satisfying the correct marginal constraints. At very small overhead costs, this generalization allows us to (i) reduce gradient variance during training, (ii) obtain straighter flows for the learned vector field, which allows us to generate high-quality samples using fewer function evaluations, and (iii) obtain transport maps with lower cost in high dimensions, which has applications beyond generative modeling. Importantly, we do so in a completely simulation-free manner with a simple minimization objective. We show that our proposed methods improve sample consistency on downsampled ImageNet data sets, and lead to better low-cost sample generation.

Added

2026-09-28

Evaluation Metrics for Graph Generative Models: Problems, Pitfalls, and Practical Solutions

Evaluation Metrics for Graph Generative Models: Problems, Pitfalls, and Practical Solutions

Leslie O'Bray, Max Horn, Bastian Rieck, Karsten M. Borgwardt

OrganizationsETH ZurichHelmholtz MunichSwiss Institute of BioinformaticsTechnical University of Munich

Why you should read this

Demonstrates critical flaws in evaluating graph generative models with Maximum Mean Discrepancy and establishes practical guidelines to ensure reliable and standardized model benchmarking.

Graph generative models are a highly active branch of machine learning. Given the steady development of new models of ever-increasing complexity, it is necessary to provide a principled way to evaluate and compare them. In this paper, we enumerate the desirable criteria for such a comparison metric and provide an overview of the status quo of graph generative model comparison in use today, which predominantly relies on the maximum mean discrepancy (MMD). We perform a systematic evaluation of MMD in the context of graph generative model comparison, highlighting some of the challenges and pitfalls researchers inadvertently may encounter. After conducting a thorough analysis of the behaviour of MMD on synthetically-generated perturbed graphs as well as on recently-proposed graph generative models, we are able to provide a suitable procedure to mitigate these challenges and pitfalls. We aggregate our findings into a list of practical recommendations for researchers to use when evaluating graph generative models.

Added

2026-09-26

Optimal Transport for Domain Adaptation

Optimal Transport for Domain Adaptation

Nicolas Courty, Rémi Flamary, Devis Tuia, Alain Rakotomamonjy

Why you should read this

Proposes a regularized optimal transport framework that aligns probability distributions between distinct domains while preserving class structure, providing a principled geometric solution for visual domain adaptation tasks.

Domain adaptation from one data space (or domain) to another is one of the most challenging tasks of modern data analytics. If the adaptation is done correctly, models built on a specific data space become more robust when confronted to data depicting the same semantic concepts (the classes), but observed by another observation system with its own specificities. Among the many strategies proposed to adapt a domain to another, finding a common representation has shown excellent properties: by finding a common representation for both domains, a single classifier can be effective in both and use labelled samples from the source domain to predict the unlabelled samples of the target domain. In this paper, we propose a regularized unsupervised optimal transportation model to perform the alignment of the representations in the source and target domains. We learn a transportation plan matching both PDFs, which constrains labelled samples in the source domain to remain close during transport. This way, we exploit at the same time the few labeled information in the source and the unlabelled distributions observed in both domains. Experiments in toy and challenging real visual adaptation examples show the interest of the method, that consistently outperforms state of the art approaches.

Added

2026-09-25

Low Dose CT Image Denoising Using a Generative Adversarial Network with Wasserstein Distance and Perceptual Loss

Low Dose CT Image Denoising Using a Generative Adversarial Network with Wasserstein Distance and Perceptual Loss

Qingsong Yang, Pingkun Yan, Yanbo Zhang, Hengyong Yu, Yongyi Shi, Xuanqin Mou, Mannudeep K. Kalra, Yi Zhang, Ling Sun, Ge Wang

OrganizationsHarvard UniversityHuaxi MR Research CenterMassachusetts General HospitalRensselaer Polytechnic InstituteSichuan UniversityUniversity of Massachusetts LowellWest China HospitalXi'an Jiaotong University

Why you should read this

Proposes a Wasserstein generative adversarial framework with perceptual loss to suppress noise and artifacts in low-dose CT scans while preserving fine diagnostic structures.

In this paper, we introduce a new CT image denoising method based on the generative adversarial network (GAN) with Wasserstein distance and perceptual similarity. The Wasserstein distance is a key concept of the optimal transform theory, and promises to improve the performance of the GAN. The perceptual loss compares the perceptual features of a denoised output against those of the ground truth in an established feature space, while the GAN helps migrate the data noise distribution from strong to weak. Therefore, our proposed method transfers our knowledge of visual perception to the image denoising task, is capable of not only reducing the image noise level but also keeping the critical information at the same time. Promising results have been obtained in our experiments with clinical CT images.

Added

2026-09-25

Demystifying MMD GANs

Demystifying MMD GANs

Mikołaj Bińkowski, Danica J. Sutherland, Michael Arbel, Arthur Gretton

OrganizationsImperial College LondonUniversity College London

Why you should read this

Introduces the Kernel Inception Distance metric for generative model evaluation while demonstrating that Maximum Mean Discrepancy GANs resolve critic gradient biases to train faster and with smaller architectures than Wasserstein GANs.

We investigate the training and performance of generative adversarial networks using the Maximum Mean Discrepancy (MMD) as critic, termed MMD GANs. As our main theoretical contribution, we clarify the situation with bias in GAN loss functions raised by recent work: we show that gradient estimators used in the optimization process for both MMD GANs and Wasserstein GANs are unbiased, but learning a discriminator based on samples leads to biased gradients for the generator parameters. We also discuss the issue of kernel choice for the MMD critic, and characterize the kernel corresponding to the energy distance used for the Cramer GAN critic. Being an integral probability metric, the MMD benefits from training strategies recently developed for Wasserstein GANs. In experiments, the MMD GAN is able to employ a smaller critic network than the Wasserstein GAN, resulting in a simpler and faster-training algorithm with matching performance. We also propose an improved measure of GAN convergence, the Kernel Inception Distance, and show how to use it to dynamically adapt learning rates during GAN training.

Added

2026-09-16

Wasserstein Auto-Encoders

Wasserstein Auto-Encoders

Ilya Tolstikhin, Olivier Bousquet, Sylvain Gelly, Bernhard Schölkopf

OrganizationsGoogleMax Planck Institute for Intelligent Systems

Why you should read this

Reformulates the generative model objective using Optimal Transport (Wasserstein distance) instead of KL-divergence, offering a stable alternative training method.

We propose the Wasserstein Auto-Encoder (WAE)—a new algorithm for building a generative model of the data distribution. WAE minimizes a penalized form of the Wasserstein distance between the model distribution and the target distribution, which leads to a different regularizer than the one used by the Variational Auto-Encoder (VAE) [1]. This regularizer encourages the encoded training distribution to match the prior. We compare our algorithm with several other techniques and show that it is a generalization of adversarial auto-encoders (AAE) [2]. Our experiments show that WAE shares many of the properties of VAEs (stable training, encoder-decoder architecture, nice latent manifold structure) while generating samples of better quality, as measured by the FID score.

Added

2026-03-09

Wasserstein GAN

Wasserstein GAN

Martin Arjovsky, Soumith Chintala, Léon Bottou

OrganizationsMetaNew York University

Why you should read this

Demonstrates that Generative Adversarial Networks (GANs) can be made stable and robust, eliminating mode collapse and providing interpretable learning curves, by leveraging the Wasserstein distance.

We introduce a new algorithm named WGAN, an alternative to traditional GAN training. In this new model, we show that we can improve the stability of learning, get rid of problems like mode collapse, and provide meaningful learning curves useful for debugging and hyperparameter searches. Furthermore, we show that the corresponding optimization problem is sound, and provide extensive theoretical work highlighting the deep connections to other distances between distributions.

Added

2026-03-08

Creative Commons License
Distributional Reinforcement Learning with Quantile Regression

Distributional Reinforcement Learning with Quantile Regression

Will Dabney, Mark Rowland, Marc G. Bellemare, Rémi Munos

OrganizationsGoogleUniversity of Cambridge

Why you should read this

This paper introduces a novel distributional reinforcement learning algorithm, QR-DQN, which bridges theoretical gaps in reinforcement learning by leveraging quantile regression and the Wasserstein metric, achieving state-of-the-art performance on Atari 2600 games.

In reinforcement learning an agent interacts with the environment by taking actions and observing the next state and reward. When sampled probabilistically, these state transitions, rewards, and actions can all induce randomness in the observed long-term return. Traditionally, reinforcement learning algorithms average over this randomness to estimate the value function. In this paper, we build on recent work advocating a distributional approach to reinforcement learning in which the distribution over returns is modeled explicitly instead of only estimating the mean. That is, we examine methods of learning the value distribution instead of the value function. We give results that close a number of gaps between the theoretical and algorithmic results given by Bellemare, Dabney, and Munos (2017). First, we extend existing results to the approximate distribution setting. Second, we present a novel distributional reinforcement learning algorithm consistent with our theoretical formulation. Finally, we evaluate this new algorithm on the Atari 2600 games, observing that it significantly outperforms many of the recent improvements on DQN, including the related distributional algorithm C51.

Added

2026-02-21