Built independently by an author, for readers. Read the story and support ChapterPal

keyword

stochastic interpolants

Stochastic interpolants are a mathematical framework in generative modeling that constructs continuous-time processes to transport probability distributions from an arbitrary base density to a target data density in finite time. By defining an explicit, time-dependent interpolation between samples drawn from the two boundary distributions, often supplemented by a tunable latent noise process, this framework unifies and generalizes flow-based and diffusion-based generative models. The evolution of the interpolating distribution corresponds to transport and Fokker-Planck equations, allowing generation to be carried out either deterministically via probability flow ordinary differential equations or stochastically via stochastic differential equations with adjustable diffusion levels. The underlying velocity and score fields can be learned directly from data samples through simulation-free square-loss regression objectives, enabling efficient training, flexible source-to-target couplings, and exact finite-time distribution bridging.

15 items

Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes

Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes

Yifan Chen, Mark Goldstein, Mengjian Hua, Michael S. Albergo, Nicholas Matthew Boffi, Eric Vanden-Eijnden

OrganizationsNew York University

Why you should read this

Develops a generative framework using stochastic interpolants and Föllmer processes to construct stochastic differential equations that map point-mass state measurements directly to predictive distributions for high-dimensional dynamical systems and video forecasting.

We propose a framework for probabilistic forecasting of dynamical systems based on generative modeling. Given observations of the system state over time, we formulate the forecasting problem as sampling from the conditional distribution of the future system state given its current state. To this end, we leverage the framework of stochastic interpolants, which facilitates the construction of a generative model between an arbitrary base distribution and the target. We design a fictitious, non-physical stochastic dynamics that takes as initial condition the current system state and produces as output a sample from the target conditional distribution in finite time and without bias. This process therefore maps a point mass centered at the current state onto a probabilistic ensemble of forecasts. We prove that the drift coefficient entering the stochastic differential equation (SDE) achieving this task is non-singular, and that it can be learned efficiently by square loss regression over the time-series data. We show that the drift and the diffusion coefficients of this SDE can be adjusted after training, and that a specific choice that minimizes the impact of the estimation error gives a Föllmer process. We highlight the utility of our approach on several complex, high-dimensional forecasting problems, including stochastically forced Navier-Stokes and video prediction on the KTH and CLEVRER datasets. The code is available at https://github.com/interpolants/forecasting.

Added

2026-10-04

Neural Diffusion Models

Neural Diffusion Models

Grigory Bartosh, Dmitry P. Vetrov, Christian A. Naesseth

OrganizationsJacobs University BremenUniversity of Amsterdam

Why you should read this

Introduces Neural Diffusion Models, a generative framework that replaces rigid linear noising processes with learnable, time-dependent non-linear data transformations to achieve tighter likelihood bounds and state-of-the-art density estimation on image benchmarks.

Diffusion models have shown remarkable performance on many generative tasks. Despite recent success, most diffusion models are restricted in that they only allow linear transformation of the data distribution. In contrast, broader family of transformations can help train generative distributions more efficiently, simplifying the reverse process and closing the gap between the true negative log-likelihood and the variational approximation. In this paper, we present Neural Diffusion Models (NDMs), a generalization of conventional diffusion models that enables defining and learning time-dependent non-linear transformations of data. We show how to optimise NDMs using a variational bound in a simulation-free setting. Moreover, we derive a time-continuous formulation of NDMs, which allows fast and reliable inference using off-the-shelf numerical ODE and SDE solvers. Finally, we demonstrate the utility of NDMs through experiments on many image generation benchmarks, including MNIST, CIFAR-10, downsampled versions of ImageNet and CelebA-HQ. NDMs outperform conventional diffusion models in terms of likelihood, achieving state-of-the-art results on ImageNet and CelebA-HQ, and produces high-quality samples.

Added

2026-10-02

Stochastic Interpolants with Data-Dependent Couplings

Stochastic Interpolants with Data-Dependent Couplings

Michael S. Albergo, Mark Goldstein, Nicholas Matthew Boffi, Rajesh Ranganath, Eric Vanden-Eijnden

Why you should read this

Proposes a framework for building continuous-time generative models by coupling base and target distributions conditioned on data, enabling efficient simulation-free training for conditional image super-resolution and in-painting tasks.

Generative models inspired by dynamical transport of measure – such as flows and diffusions – construct a continuous-time map between two probability densities. Conventionally, one of these is the target density, only accessible through samples, while the other is taken as a simple base density that is data-agnostic. In this work, using the framework of stochastic interpolants, we formalize how to couple the base and the target densities, whereby samples from the base are computed conditionally given samples from the target in a way that is different from (but does not preclude) incorporating information about class labels or continuous embeddings. This enables us to construct dynamical transport maps that serve as conditional generative models. We show that these transport maps can be learned by solving a simple square loss regression problem analogous to the standard independent setting. We demonstrate the usefulness of constructing dependent couplings in practice through experiments in super-resolution and in-painting. The code is available at https://github.com/interpolants/couplings.

Added

2026-10-02

Any-Order Flexible Length Masked Diffusion

Any-Order Flexible Length Masked Diffusion

Jaeyeon Kim, C. Lee, Carles Domingo-Enrich, Yilun Du, S. Kakade, Timothy Ngotiaoco, Sitan Chen, M. S. Albergo

OrganizationsHarvard UniversityKemper Institute for the Study of Natural and Artificial IntelligenceMicrosoftThe NSF Institute for Artificial Intelligence and Fundamental Interactions

Why you should read this

Introduces Flexible Masked Diffusion Models (FlexMDMs) to eliminate the fixed-length constraint of discrete diffusion models by enabling dynamic token insertion alongside any-order generation, substantially improving mathematical reasoning and code infilling when fine-tuned on large language models.

Masked diffusion models (MDMs) have recently emerged as a promising alternative to autoregressive models over discrete domains. MDMs generate sequences in an any-order, parallel fashion, enabling fast inference and strong performance on non-causal tasks. However, a crucial limitation is that they do not support token insertions and are thus limited to fixed-length generations. To this end, we introduce Flexible Masked Diffusion Models (FlexMDMs), a discrete diffusion paradigm that simultaneously can model sequences of flexible length while provably retaining MDMs' flexibility of any-order inference. Grounded in an extension of the stochastic interpolant framework, FlexMDMs generate sequences by inserting mask tokens and unmasking them. Empirically, we show that FlexMDMs match MDMs in perplexity while modeling length statistics with much higher fidelity. On a synthetic maze planning task, they achieve ≈60%\approx 60 \% higher success rate than MDM baselines. Finally, we show pretrained MDMs can easily be retrofitted into FlexMDMs: on 16 H100s, it takes only three days to fine-tune LLaDA-8B into a FlexMDM, achieving superior performance on math (GSM8K, 58%→67%58\% \to 67\%) and code infilling performance (52%→65%52\% \to 65\%).

Added

2026-09-30

Thinking with Looped Flows

Thinking with Looped Flows

Ayhan Suleymanzade, Chanhyuk Lee, Floor Eijkelboom, Nicholas M. Boffi, İsmail İlkan Ceylan, Jinwoo Kim

OrganizationsAITHYRACarnegie Mellon UniversityÉcole Polytechnique Fédérale de LausanneKorea Advanced Institute of Science and TechnologyTU WienUniversity of AmsterdamUniversity of Oxford

Why you should read this

Introduces looped flows, a framework that trains recurrent architectures through progressive denoising objectives to overcome gradient truncation, enabling dynamic test-time compute scaling via probability flow integration and achieving leading accuracy on ARC-AGI reasoning benchmarks.

Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local denoising objectives. By imposing temporal association across denoising objectives through progressively decreasing noise levels and shared noise, the model is incentivized to learn recurrent states that transfer useful computation over time, even when gradients cover only a few updates. We then formulate inference as integrating the velocity of a probability flow parameterized by the learned denoiser, coupled with recurrent states. This allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples. Across six reasoning benchmarks including two multi-solution benchmarks, looped flows outperform prior state-of-the-art looped models overall, achieving 58.8% test accuracy on ARC-AGI-1 and 12.2% on ARC-AGI-2.

Added

2026-09-13

Creative Commons License
There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation

There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation

Gabe Guo, Elon Litman, Thanawat Sornwanee, Jose Blanchet, Stefano Ermon

OrganizationsStanford University

Why you should read this

Develops bidirectional diffusion bridges that directly interpolate between text and image representations, establishing a unified continuous-space framework for both text-to-image generation and image-to-text inversion.

Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) follow generative paths that do not directly represent the source modality, limiting the flexibility of some sampling algorithms; and (2) are unidirectional, preventing inversion (e.g., image-to-text). We propose BIT: Bidirectional Image-Text Diffusion Bridges. In contrast to previous approaches, BIT starts directly from text and interpolates into images, providing (1) a source-aware generative path that enables diverse and flexible sampling algorithms; and (2) an endpoint-conditioned process that can be traversed from image to text, providing a unified, bidirectional generative framework. BIT is derived through stochastic calculus, yielding SDE forms amenable to simulation and tractable loss functions that scale to high dimensions. Our experiments show that BIT is competitive with denoising-diffusion and deterministic-flow baselines, and outperforms them on several vision--language and natural-science evaluations.

Added

2026-09-02

Creative Commons License
Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden

OrganizationsNew York University

Why you should read this

Introduces a unifying framework of "stochastic interpolants" that bridges flow-based and diffusion-based generative models, allowing exact probability density function transformations in finite time with tunable noise levels.

A class of generative models that unifies flow-based and diffusion-based methods is introduced. These models extend the framework proposed in Albergo and Vanden-Eijnden (2023), enabling the use of a broad class of continuous-time stochastic processes called stochastic interpolants to bridge any two probability density functions exactly in finite time. These interpolants are built by combining data from the two prescribed densities with an additional latent variable that shapes the bridge in a flexible way. The time-dependent density function of the interpolant is shown to satisfy a transport equation as well as a family of forward and backward Fokker-Planck equations with tunable diffusion coefficient. Upon consideration of the time evolution of an individual sample, this viewpoint leads to both deterministic and stochastic generative models based on probability flow equations or stochastic differential equations with an adjustable level of noise. The drift coefficients entering these models are time-dependent velocity fields characterized as the unique minimizers of simple quadratic objective functions, one of which is a new objective for the score. We show that minimization of these quadratic objectives leads to control of the likelihood for generative models built upon stochastic dynamics, while likelihood control for deterministic dynamics is more stringent. We also construct estimators for the likelihood and the cross entropy of interpolant-based generative models, and we discuss connections with other methods such as score-based diffusion models, stochastic localization, probabilistic denoising, and rectifying flows. In addition, we demonstrate that stochastic interpolants recover the Schrödinger bridge between the two target densities when explicitly optimizing over the interpolant. Finally, algorithmic aspects are discussed and the approach is illustrated on numerical examples.

Added

2026-05-14

Creative Commons License
Mean Flows for One-step Generative Modeling

Mean Flows for One-step Generative Modeling

Zhengyang Geng, Mingyang Deng, Xingjian Bai, J. Zico Kolter, Kaiming He

OrganizationsCarnegie Mellon UniversityMassachusetts Institute of Technology

Why you should read this

This paper demonstrates a novel framework, MeanFlow, that achieves state-of-the-art one-step generative modeling performance without pre-training or distillation, significantly narrowing the gap with multi-step predecessors by introducing the principled concept of average velocity.

We propose a principled and effective framework for one-step generative modeling. We introduce the notion of average velocity to characterize flow fields, in contrast to instantaneous velocity modeled by Flow Matching methods. A well-defined identity between average and instantaneous velocities is derived and used to guide neural network training. Our method, termed the MeanFlow model, is self-contained and requires no pre-training, distillation, or curriculum learning. MeanFlow demonstrates strong empirical performance: it achieves an FID of 3.43 with a single function evaluation (1-NFE) on ImageNet 256x256 trained from scratch, significantly outperforming previous state-of-the-art one-step diffusion/flow models. Our study substantially narrows the gap between one-step diffusion/flow models and their multi-step predecessors, and we hope it will motivate future research to revisit the foundations of these powerful models.

Added

2026-01-20

Creative Commons License
Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training

Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training

Tony Bonnaire, Raphaël Urfin, Giulio Biroli, Marc Mézard

Why you should read this

Reveals how an implicit dynamical regularization in training prevents diffusion models from memorizing, maintaining generalization even in heavily overparameterized settings.

Diffusion models have achieved remarkable success across a wide range of generative tasks. A key challenge is understanding the mechanisms that prevent their memorization of training data and allow generalization. In this work, we investigate the role of the training dynamics in the transition from generalization to memorization. Through extensive experiments and theoretical analysis, we identify two distinct timescales: an early time τgen\tau_\mathrm{gen} at which models begin to generate high-quality samples, and a later time τmem\tau_\mathrm{mem} beyond which memorization emerges. Crucially, we find that τmem\tau_\mathrm{mem} increases linearly with the training set size nn, while τgen\tau_\mathrm{gen} remains constant. This creates a growing window of training times with nn where models generalize effectively, despite showing strong memorization if training continues beyond it. It is only when nn becomes larger than a model-dependent threshold that overfitting disappears at infinite training times. These findings reveal a form of implicit dynamical regularization in the training dynamics, which allow to avoid memorization even in highly overparameterized settings. Our results are supported by numerical experiments with standard U-Net architectures on realistic and synthetic datasets, and by a theoretical analysis using a tractable random features model studied in the high-dimensional limit.

Added

2025-12-27

Creative Commons License