keyword
stochastic interpolants
Stochastic interpolants are a mathematical framework in generative modeling that constructs continuous-time processes to transport probability distributions from an arbitrary base density to a target data density in finite time. By defining an explicit, time-dependent interpolation between samples drawn from the two boundary distributions, often supplemented by a tunable latent noise process, this framework unifies and generalizes flow-based and diffusion-based generative models. The evolution of the interpolating distribution corresponds to transport and Fokker-Planck equations, allowing generation to be carried out either deterministically via probability flow ordinary differential equations or stochastically via stochastic differential equations with adjustable diffusion levels. The underlying velocity and score fields can be learned directly from data samples through simulation-free square-loss regression objectives, enabling efficient training, flexible source-to-target couplings, and exact finite-time distribution bridging.
15 items

Probabilistic Forecasting with Stochastic Interpolants and Föllmer Processes
Yifan Chen, Mark Goldstein, Mengjian Hua, Michael S. Albergo, Nicholas Matthew Boffi, Eric Vanden-Eijnden
Why you should read this
Develops a generative framework using stochastic interpolants and Föllmer processes to construct stochastic differential equations that map point-mass state measurements directly to predictive distributions for high-dimensional dynamical systems and video forecasting.
We propose a framework for probabilistic forecasting of dynamical systems based on generative modeling. Given observations of the system state over time, we formulate the forecasting problem as sampling from the conditional distribution of the future system state given its current state. To this end, we leverage the framework of stochastic interpolants, which facilitates the construction of a generative model between an arbitrary base distribution and the target. We design a fictitious, non-physical stochastic dynamics that takes as initial condition the current system state and produces as output a sample from the target conditional distribution in finite time and without bias. This process therefore maps a point mass centered at the current state onto a probabilistic ensemble of forecasts. We prove that the drift coefficient entering the stochastic differential equation (SDE) achieving this task is non-singular, and that it can be learned efficiently by square loss regression over the time-series data. We show that the drift and the diffusion coefficients of this SDE can be adjusted after training, and that a specific choice that minimizes the impact of the estimation error gives a Föllmer process. We highlight the utility of our approach on several complex, high-dimensional forecasting problems, including stochastically forced Navier-Stokes and video prediction on the KTH and CLEVRER datasets. The code is available at https://github.com/interpolants/forecasting.
Added
2026-10-04

Neural Diffusion Models
Grigory Bartosh, Dmitry P. Vetrov, Christian A. Naesseth
Why you should read this
Introduces Neural Diffusion Models, a generative framework that replaces rigid linear noising processes with learnable, time-dependent non-linear data transformations to achieve tighter likelihood bounds and state-of-the-art density estimation on image benchmarks.
Diffusion models have shown remarkable performance on many generative tasks. Despite recent success, most diffusion models are restricted in that they only allow linear transformation of the data distribution. In contrast, broader family of transformations can help train generative distributions more efficiently, simplifying the reverse process and closing the gap between the true negative log-likelihood and the variational approximation. In this paper, we present Neural Diffusion Models (NDMs), a generalization of conventional diffusion models that enables defining and learning time-dependent non-linear transformations of data. We show how to optimise NDMs using a variational bound in a simulation-free setting. Moreover, we derive a time-continuous formulation of NDMs, which allows fast and reliable inference using off-the-shelf numerical ODE and SDE solvers. Finally, we demonstrate the utility of NDMs through experiments on many image generation benchmarks, including MNIST, CIFAR-10, downsampled versions of ImageNet and CelebA-HQ. NDMs outperform conventional diffusion models in terms of likelihood, achieving state-of-the-art results on ImageNet and CelebA-HQ, and produces high-quality samples.
Added
2026-10-02

Stochastic Interpolants with Data-Dependent Couplings
Michael S. Albergo, Mark Goldstein, Nicholas Matthew Boffi, Rajesh Ranganath, Eric Vanden-Eijnden
Why you should read this
Proposes a framework for building continuous-time generative models by coupling base and target distributions conditioned on data, enabling efficient simulation-free training for conditional image super-resolution and in-painting tasks.
Generative models inspired by dynamical transport of measure – such as flows and diffusions – construct a continuous-time map between two probability densities. Conventionally, one of these is the target density, only accessible through samples, while the other is taken as a simple base density that is data-agnostic. In this work, using the framework of stochastic interpolants, we formalize how to couple the base and the target densities, whereby samples from the base are computed conditionally given samples from the target in a way that is different from (but does not preclude) incorporating information about class labels or continuous embeddings. This enables us to construct dynamical transport maps that serve as conditional generative models. We show that these transport maps can be learned by solving a simple square loss regression problem analogous to the standard independent setting. We demonstrate the usefulness of constructing dependent couplings in practice through experiments in super-resolution and in-painting. The code is available at https://github.com/interpolants/couplings.
Added
2026-10-02

Action Matching: Learning Stochastic Dynamics from Samples
Kirill Neklyudov, Rob Brekelmans, Daniel Severo, Alireza Makhzani
Why you should read this
Proposes Action Matching, a tractable framework for learning continuous population dynamics directly from uncorrelated snapshot samples across time without requiring optimal transport solvers or backpropagation through differential equations.
Machine learning offers several techniques where systems learn and improve automatically from experience without being clearly programmed. Some set ups which require very fast adaptation needs less training samples. In meta-learning, we focus on problems dealing with tasks where a learning algorithm has to quickly adapt itself with limited number of labelled samples to execute new tasks extracted from similar distribution. We introduce our approach to leverage large scale deep neural networks along with context-aware parameter generation mechanism using affine transformations of embeddings achieved through convolutional feature maps -> scaled Learned Init...
Added
2026-10-02

Any-Order Flexible Length Masked Diffusion
Jaeyeon Kim, C. Lee, Carles Domingo-Enrich, Yilun Du, S. Kakade, Timothy Ngotiaoco, Sitan Chen, M. S. Albergo
Why you should read this
Introduces Flexible Masked Diffusion Models (FlexMDMs) to eliminate the fixed-length constraint of discrete diffusion models by enabling dynamic token insertion alongside any-order generation, substantially improving mathematical reasoning and code infilling when fine-tuned on large language models.
Masked diffusion models (MDMs) have recently emerged as a promising alternative to autoregressive models over discrete domains. MDMs generate sequences in an any-order, parallel fashion, enabling fast inference and strong performance on non-causal tasks. However, a crucial limitation is that they do not support token insertions and are thus limited to fixed-length generations. To this end, we introduce Flexible Masked Diffusion Models (FlexMDMs), a discrete diffusion paradigm that simultaneously can model sequences of flexible length while provably retaining MDMs' flexibility of any-order inference. Grounded in an extension of the stochastic interpolant framework, FlexMDMs generate sequences by inserting mask tokens and unmasking them. Empirically, we show that FlexMDMs match MDMs in perplexity while modeling length statistics with much higher fidelity. On a synthetic maze planning task, they achieve higher success rate than MDM baselines. Finally, we show pretrained MDMs can easily be retrofitted into FlexMDMs: on 16 H100s, it takes only three days to fine-tune LLaDA-8B into a FlexMDM, achieving superior performance on math (GSM8K, ) and code infilling performance ().
Added
2026-09-30

WarmPrior: Straightening Flow-Matching Policies with Temporal Priors
Sinjae Kang, Chanyoung Kim, Kaixin Wang, Li Zhao, Kimin Lee
Why you should read this
Introduces WarmPrior, a temporal source distribution derived from recent action history that straightens probability paths in flow-matching policies to boost robotic manipulation performance across imitation and reinforcement learning.
Generative policies based on diffusion and flow matching have become a dominant paradigm for visuomotor robotic control. We show that replacing the standard Gaussian source distribution with WarmPrior, a simple temporally grounded prior constructed from readily available recent action history, consistently improves success rates on robotic manipulation tasks. We trace this gain to markedly straighter probability paths, echoing the effect of optimal-transport couplings in Rectified Flow. Beyond standard behavior cloning, WarmPrior also reshapes the exploration distribution in prior-space reinforcement learning, improving both sample efficiency and final performance. Collectively, these results identify the source distribution as an important and underexplored design axis in generative robot control.
Added
2026-09-29

Inductive Moment Matching
Linqi Zhou, Stefano Ermon, Jiaming Song
Why you should read this
Introduces Inductive Moment Matching, a single-stage framework that trains few-step generative models from scratch without pre-training or distillation, achieving state-of-the-art fast sampling performance on ImageNet and CIFAR-10 with guaranteed distribution-level convergence.
Diffusion models and Flow Matching generate high-quality samples but are slow at inference, and distilling them into few-step models often leads to instability and extensive tuning. To resolve these trade-offs, we propose Inductive Moment Matching (IMM), a new class of generative models for one- or few-step sampling with a single-stage training procedure. Unlike distillation, IMM does not require pre-training initialization and optimization of two networks; and unlike Consistency Models, IMM guarantees distribution-level convergence and remains stable under various hyperparameters and standard model architectures. IMM surpasses diffusion models on ImageNet-256×256 with 1.99 FID using only 8 inference steps and achieves state-of-the-art 2-step FID of 1.98 on CIFAR-10 for a model trained from scratch.
Added
2026-09-26

Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control
Carles Domingo-Enrich, Michal Drozdzal, Brian Karrer, Ricky T. Q. Chen
Why you should read this
Introduces Adjoint Matching, a regression-based algorithm that solves stochastic optimal control with a memoryless noise schedule to effectively fine-tune flow and diffusion generative models for reward alignment while preserving sample diversity.
Dynamical generative models that produce samples through an iterative process, such as Flow Matching and denoising diffusion models, have seen widespread use, but there have not been many theoretically-sound methods for improving these models with reward fine-tuning. In this work, we cast reward fine-tuning as stochastic optimal control (SOC). Critically, we prove that a very specific memoryless noise schedule must be enforced during fine-tuning, in order to account for the dependency between the noise variable and the generated samples. We also propose a new algorithm named Adjoint Matching which outperforms existing SOC algorithms, by casting SOC problems as a regression problem. We find that our approach significantly improves over existing methods for reward fine-tuning, achieving better consistency, realism, and generalization to unseen human preference reward models, while retaining sample diversity.
Added
2026-09-26

Thinking with Looped Flows
Ayhan Suleymanzade, Chanhyuk Lee, Floor Eijkelboom, Nicholas M. Boffi, İsmail İlkan Ceylan, Jinwoo Kim
Why you should read this
Introduces looped flows, a framework that trains recurrent architectures through progressive denoising objectives to overcome gradient truncation, enabling dynamic test-time compute scaling via probability flow integration and achieving leading accuracy on ARC-AGI reasoning benchmarks.
Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local denoising objectives. By imposing temporal association across denoising objectives through progressively decreasing noise levels and shared noise, the model is incentivized to learn recurrent states that transfer useful computation over time, even when gradients cover only a few updates. We then formulate inference as integrating the velocity of a probability flow parameterized by the learned denoiser, coupled with recurrent states. This allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples. Across six reasoning benchmarks including two multi-solution benchmarks, looped flows outperform prior state-of-the-art looped models overall, achieving 58.8% test accuracy on ARC-AGI-1 and 12.2% on ARC-AGI-2.
Added
2026-09-13


There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation
Gabe Guo, Elon Litman, Thanawat Sornwanee, Jose Blanchet, Stefano Ermon
Why you should read this
Develops bidirectional diffusion bridges that directly interpolate between text and image representations, establishing a unified continuous-space framework for both text-to-image generation and image-to-text inversion.
Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) follow generative paths that do not directly represent the source modality, limiting the flexibility of some sampling algorithms; and (2) are unidirectional, preventing inversion (e.g., image-to-text). We propose BIT: Bidirectional Image-Text Diffusion Bridges. In contrast to previous approaches, BIT starts directly from text and interpolates into images, providing (1) a source-aware generative path that enables diverse and flexible sampling algorithms; and (2) an endpoint-conditioned process that can be traversed from image to text, providing a unified, bidirectional generative framework. BIT is derived through stochastic calculus, yielding SDE forms amenable to simulation and tractable loss functions that scale to high dimensions. Our experiments show that BIT is competitive with denoising-diffusion and deterministic-flow baselines, and outperforms them on several vision--language and natural-science evaluations.
Added
2026-09-02


A Mathematical Introduction to Diffusion Models
Jianfeng Lu
Why you should read this
Develops a self-contained, proof-oriented mathematical foundation for diffusion models by connecting classical Langevin sampling dynamics to modern continuous and discrete score-based samplers, sampling error bounds, and inference-time control.
These notes give a proof-oriented introduction to diffusion models from the viewpoint of sampling, tracing a single arc from classical sampling dynamics to modern diffusion samplers, their error analysis, and inference-time control. Throughout, the material is layered into core definitions and identities proved in full, representative estimates proved under simplifying assumptions, and research-level theorems stated with a proof roadmap. The intended audience is beginning graduate students with a background in probability but no prior exposure to stochastic differential equations, stochastic numerics, or diffusion models.
Added
2026-09-01


Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
Michael S. Albergo, Nicholas M. Boffi, Eric Vanden-Eijnden
Why you should read this
Introduces a unifying framework of "stochastic interpolants" that bridges flow-based and diffusion-based generative models, allowing exact probability density function transformations in finite time with tunable noise levels.
A class of generative models that unifies flow-based and diffusion-based methods is introduced. These models extend the framework proposed in Albergo and Vanden-Eijnden (2023), enabling the use of a broad class of continuous-time stochastic processes called stochastic interpolants to bridge any two probability density functions exactly in finite time. These interpolants are built by combining data from the two prescribed densities with an additional latent variable that shapes the bridge in a flexible way. The time-dependent density function of the interpolant is shown to satisfy a transport equation as well as a family of forward and backward Fokker-Planck equations with tunable diffusion coefficient. Upon consideration of the time evolution of an individual sample, this viewpoint leads to both deterministic and stochastic generative models based on probability flow equations or stochastic differential equations with an adjustable level of noise. The drift coefficients entering these models are time-dependent velocity fields characterized as the unique minimizers of simple quadratic objective functions, one of which is a new objective for the score. We show that minimization of these quadratic objectives leads to control of the likelihood for generative models built upon stochastic dynamics, while likelihood control for deterministic dynamics is more stringent. We also construct estimators for the likelihood and the cross entropy of interpolant-based generative models, and we discuss connections with other methods such as score-based diffusion models, stochastic localization, probabilistic denoising, and rectifying flows. In addition, we demonstrate that stochastic interpolants recover the Schrödinger bridge between the two target densities when explicitly optimizing over the interpolant. Finally, algorithmic aspects are discussed and the approach is illustrated on numerical examples.
Added
2026-05-14


Generative Modeling via Drifting
Mingyang Deng, He Li, Tianhong Li, Yilun Du, Kaiming He
Why you should read this
Proposes Drifting Models, a novel generative modeling paradigm that achieves state-of-the-art image generation with one-step inference by evolving the pushforward distribution during training, significantly improving efficiency for high-quality results.
Generative modeling can be formulated as learning a mapping f such that its pushforward distribution matches the data distribution. The pushforward behavior can be carried out iteratively at inference time, for example in diffusion and flow-based models. In this paper, we propose a new paradigm called Drifting Models, which evolve the pushforward distribution during training and naturally admit one-step inference. We introduce a drifting field that governs the sample movement and achieves equilibrium when the distributions match. This leads to a training objective that allows the neural network optimizer to evolve the distribution. In experiments, our one-step generator achieves state-of-the-art results on ImageNet at 256 x 256 resolution, with an FID of 1.54 in latent space and 1.61 in pixel space. We hope that our work opens up new opportunities for high-quality one-step generation.
Added
2026-04-05


Mean Flows for One-step Generative Modeling
Zhengyang Geng, Mingyang Deng, Xingjian Bai, J. Zico Kolter, Kaiming He
Why you should read this
This paper demonstrates a novel framework, MeanFlow, that achieves state-of-the-art one-step generative modeling performance without pre-training or distillation, significantly narrowing the gap with multi-step predecessors by introducing the principled concept of average velocity.
We propose a principled and effective framework for one-step generative modeling. We introduce the notion of average velocity to characterize flow fields, in contrast to instantaneous velocity modeled by Flow Matching methods. A well-defined identity between average and instantaneous velocities is derived and used to guide neural network training. Our method, termed the MeanFlow model, is self-contained and requires no pre-training, distillation, or curriculum learning. MeanFlow demonstrates strong empirical performance: it achieves an FID of 3.43 with a single function evaluation (1-NFE) on ImageNet 256x256 trained from scratch, significantly outperforming previous state-of-the-art one-step diffusion/flow models. Our study substantially narrows the gap between one-step diffusion/flow models and their multi-step predecessors, and we hope it will motivate future research to revisit the foundations of these powerful models.
Added
2026-01-20


Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training
Tony Bonnaire, Raphaël Urfin, Giulio Biroli, Marc Mézard
Why you should read this
Reveals how an implicit dynamical regularization in training prevents diffusion models from memorizing, maintaining generalization even in heavily overparameterized settings.
Diffusion models have achieved remarkable success across a wide range of generative tasks. A key challenge is understanding the mechanisms that prevent their memorization of training data and allow generalization. In this work, we investigate the role of the training dynamics in the transition from generalization to memorization. Through extensive experiments and theoretical analysis, we identify two distinct timescales: an early time at which models begin to generate high-quality samples, and a later time beyond which memorization emerges. Crucially, we find that increases linearly with the training set size , while remains constant. This creates a growing window of training times with where models generalize effectively, despite showing strong memorization if training continues beyond it. It is only when becomes larger than a model-dependent threshold that overfitting disappears at infinite training times. These findings reveal a form of implicit dynamical regularization in the training dynamics, which allow to avoid memorization even in highly overparameterized settings. Our results are supported by numerical experiments with standard U-Net architectures on realistic and synthetic datasets, and by a theoretical analysis using a tractable random features model studied in the high-dimensional limit.
Added
2025-12-27

