Built independently by an author, for readers. Read the story and support ChapterPal

keyword

flow matching

Flow matching is a machine learning framework used in generative modeling to train continuous normalizing flows that transform a simple prior distribution, such as Gaussian noise, into a complex target data distribution. It works by training a neural network using standard regression to learn a time-dependent vector field, which defines the velocity and direction needed to guide samples along predefined continuous probability trajectories connecting noise and real data. Unlike conventional continuous flow training methods that require computationally expensive differential equation simulations during training, flow matching enables simulation-free optimization by conditioning the trajectories on individual data points, often forming straight-line or optimal-transport paths. Once trained, new samples are generated by drawing random noise and numerically integrating the learned vector field forward in time with an ordinary differential equation solver, offering faster sampling and training stability compared to traditional diffusion models.

22 items

Image Restoration Through Generalized Ornstein-Uhlenbeck Bridge

Image Restoration Through Generalized Ornstein-Uhlenbeck Bridge

Conghan Yue, Zhengwei Peng, Junlong Ma, Shiyan Du, Pengxu Wei, Dongyu Zhang

OrganizationsPeng Cheng LaboratorySun Yat-sen University

Why you should read this

Proposes a Generalized Ornstein-Uhlenbeck Bridge framework that applies Doob's h-transform to establish direct point-to-point diffusion mappings between low- and high-quality images, unifying existing bridge methods and achieving state-of-the-art performance across inpainting, deraining, and super-resolution.

Diffusion models exhibit powerful generative capabilities enabling noise mapping to data via reverse stochastic differential equations. However, in image restoration, the focus is on the mapping relationship from low-quality to high-quality images. Regarding this issue, we introduce the Generalized Ornstein-Uhlenbeck Bridge (GOUB) model. By leveraging the natural mean-reverting property of the generalized OU process and further eliminating the variance of its steady-state distribution through the Doob's h-transform, we achieve diffusion mappings from point to point enabling the recovery of high-quality images from low-quality ones. Moreover, we unravel the fundamental mathematical essence shared by various bridge models, all of which are special instances of GOUB and empirically demonstrate the optimality of our proposed models. Additionally, we present the corresponding Mean-ODE model adept at capturing both pixel-level details and structural perceptions. Experimental outcomes showcase the state-of-the-art performance achieved by both models across diverse tasks, including inpainting, deraining, and super-resolution. Code is available at https://github.com/Hammour-steak/GOUB.

Added

2026-10-05

One Diffusion to Generate Them All

One Diffusion to Generate Them All

Duong H. Le, Tuan Pham, Sangho Lee, Christopher Clark, Aniruddha Kembhavi, Stephan Mandt, Ranjay Krishna, Jiasen Lu

OrganizationsAllen Institute for AIUniversity of California, IrvineUniversity of Washington

Why you should read this

Presents OneDiffusion, a unified diffusion model that integrates bidirectional image synthesis and visual understanding into a single architecture by formulating diverse vision tasks as multi-frame sequences with variable noise scales.

We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables conditional generation from inputs such as text, depth, pose, layout, and semantic maps, while also handling tasks like image deblurring, upscaling, and reverse processes such as depth estimation and segmentation. Additionally, OneDiffusion allows for multi-view generation, camera pose estimation, and instant personalization using sequential image inputs. Our model takes a straightforward yet effective approach by treating all tasks as frame sequences with varying noise scales during training, allowing any frame to act as a conditioning image at inference time. Our unified training framework removes the need for specialized architectures, supports scalable multi-task training, and adapts smoothly to any resolution, enhancing both generalization and scalability. Experimental results demonstrate competitive performance across tasks in both generation and prediction such as text-to-image, multiview generation, ID preservation, depth estimation and camera pose estimation despite relatively small training dataset. Our code and checkpoint are freely available at this https URL

Added

2026-10-03

EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers

EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers

Daiheng Gao, Shilin Lu, Wenbo Zhou, Jiaming Chu, Jie Zhang, Mengxi Jia, Bang Zhang, Zhaoxin Fan, Weiming Zhang

Why you should read this

Proposes a bi-level optimization framework combining LoRA tuning, attention map regularization, and self-contrastive learning to successfully remove unwanted concepts from modern flow-matching transformer models like Flux without degrading unrelated visual outputs.

Removing unwanted concepts from large-scale text-to-image (T2I) diffusion models while maintaining their overall generative quality remains an open challenge. This difficulty is especially pronounced in emerging paradigms, such as Stable Diffusion (SD) v3 and Flux, which incorporate flow matching and transformer-based architectures. These advancements limit the transferability of existing concept-erasure techniques that were originally designed for the previous T2I paradigm (e.g., SD v1.4). In this work, we introduce EraseAnything, the first method specifically developed to address concept erasure within the latest flow-based T2I framework. We formulate concept erasure as a bi-level optimization problem, employing LoRA-based parameter tuning and an attention map regularizer to selectively suppress undesirable activations. Furthermore, we propose a self-contrastive learning strategy to ensure that removing unwanted concepts does not inadvertently harm performance on unrelated ones. Experimental results demonstrate that EraseAnything successfully fills the research gap left by earlier methods in this new T2I paradigm, achieving SOTA performance across a wide range of concept erasure tasks.

Added

2026-10-02

Stochastic Interpolants with Data-Dependent Couplings

Stochastic Interpolants with Data-Dependent Couplings

Michael S. Albergo, Mark Goldstein, Nicholas Matthew Boffi, Rajesh Ranganath, Eric Vanden-Eijnden

Why you should read this

Proposes a framework for building continuous-time generative models by coupling base and target distributions conditioned on data, enabling efficient simulation-free training for conditional image super-resolution and in-painting tasks.

Generative models inspired by dynamical transport of measure – such as flows and diffusions – construct a continuous-time map between two probability densities. Conventionally, one of these is the target density, only accessible through samples, while the other is taken as a simple base density that is data-agnostic. In this work, using the framework of stochastic interpolants, we formalize how to couple the base and the target densities, whereby samples from the base are computed conditionally given samples from the target in a way that is different from (but does not preclude) incorporating information about class labels or continuous embeddings. This enables us to construct dynamical transport maps that serve as conditional generative models. We show that these transport maps can be learned by solving a simple square loss regression problem analogous to the standard independent setting. We demonstrate the usefulness of constructing dependent couplings in practice through experiments in super-resolution and in-painting. The code is available at https://github.com/interpolants/couplings.

Added

2026-10-02

Learning GFlowNets From Partial Episodes For Improved Convergence And Stability

Learning GFlowNets From Partial Episodes For Improved Convergence And Stability

Kanika Madan, Jarrid Rector-Brooks, Maksym Korablyov, Emmanuel Bengio, Moksh Jain, Andrei Cristian Nica, Tom Bosc, Yoshua Bengio, Nikolay Malkin

OrganizationsCIFARMila – Québec Artificial Intelligence InstituteNational University of Science and Technology Politehnica BucharestRecursionUniversité de Montréal

Why you should read this

Introduces Subtrajectory Balance, a GFlowNet training objective inspired by TD(λ\lambda) that learns from partial action sequences to balance gradient bias and variance, accelerating convergence and enabling effective training in long-horizon, sparse-reward environments.

Generative flow networks (GFlowNets) are a family of algorithms for training a sequential sampler of discrete objects under an unnormalized target density and have been successfully used for various probabilistic modeling tasks. Existing training objectives for GFlowNets are either local to states or transitions, or propagate a reward signal over an entire sampling trajectory. We argue that these alternatives represent opposite ends of a gradient bias-variance tradeoff and propose a way to exploit this tradeoff to mitigate its harmful effects. Inspired by the TD(λ) algorithm in reinforcement learning, we introduce subtrajectory balance or SubTB(λ), a GFlowNet training objective that can learn from partial action subsequences of varying lengths. We show that SubTB(λ) accelerates sampler convergence in previously studied and new environments and enables training GFlowNets in environments with longer action sequences and sparser reward landscapes than what was possible before. We also perform a comparative analysis of stochastic gradient dynamics, shedding light on the bias-variance tradeoff in GFlowNet training and the advantages of subtrajectory balance.

Added

2026-10-01

Any-Order Flexible Length Masked Diffusion

Any-Order Flexible Length Masked Diffusion

Jaeyeon Kim, C. Lee, Carles Domingo-Enrich, Yilun Du, S. Kakade, Timothy Ngotiaoco, Sitan Chen, M. S. Albergo

OrganizationsHarvard UniversityKemper Institute for the Study of Natural and Artificial IntelligenceMicrosoftThe NSF Institute for Artificial Intelligence and Fundamental Interactions

Why you should read this

Introduces Flexible Masked Diffusion Models (FlexMDMs) to eliminate the fixed-length constraint of discrete diffusion models by enabling dynamic token insertion alongside any-order generation, substantially improving mathematical reasoning and code infilling when fine-tuned on large language models.

Masked diffusion models (MDMs) have recently emerged as a promising alternative to autoregressive models over discrete domains. MDMs generate sequences in an any-order, parallel fashion, enabling fast inference and strong performance on non-causal tasks. However, a crucial limitation is that they do not support token insertions and are thus limited to fixed-length generations. To this end, we introduce Flexible Masked Diffusion Models (FlexMDMs), a discrete diffusion paradigm that simultaneously can model sequences of flexible length while provably retaining MDMs' flexibility of any-order inference. Grounded in an extension of the stochastic interpolant framework, FlexMDMs generate sequences by inserting mask tokens and unmasking them. Empirically, we show that FlexMDMs match MDMs in perplexity while modeling length statistics with much higher fidelity. On a synthetic maze planning task, they achieve ≈60%\approx 60 \% higher success rate than MDM baselines. Finally, we show pretrained MDMs can easily be retrofitted into FlexMDMs: on 16 H100s, it takes only three days to fine-tune LLaDA-8B into a FlexMDM, achieving superior performance on math (GSM8K, 58%→67%58\% \to 67\%) and code infilling performance (52%→65%52\% \to 65\%).

Added

2026-09-30

Diffeomorphic Optimization

Diffeomorphic Optimization

Ludwig Winkler, Andrew Leaver-Fay, Joseph Kleinhenz, Pan Kessel

OrganizationsGenentechMicrosoft

Why you should read this

Introduces diffeomorphic optimization to perform Riemannian gradient descent through the base space of generative models, extending the framework to Lie groups to achieve superior accuracy and speed in computational protein design.

Generative models learn data distributions that reside on a low-dimensional manifold within a higher-dimensional ambient space. Optimizing differentiable objectives on this manifold is challenging: the ambient loss landscape is high-dimensional, rugged, and non-convex. Direct gradient descent, blind to the manifold's geometry, quickly drifts off it. Diffeomorphic optimization starts from the observation that diffusion and flow models provide a map from the data manifold to a much simpler base space in which we perform gradient descent. Using differential geometry, we show this is equivalent to Riemannian gradient descent on the data manifold up to O(λ2)\mathcal{O}(\lambda^2) corrections, keeping trajectories on-manifold by construction and yielding a smoother optimization surface. For protein design, we extend diffeomorphic optimization to the matrix Lie groups SO(3)\mathrm{SO}(3) and SE(3)\mathrm{SE}(3), deriving an autograd-compatible SO(3)\mathrm{SO}(3) gradient and a generalized adjoint-state method for backpropagation through Lie-group ODE solvers. Diffeomorphic optimization improves over tuned guidance on secondary-structure targeting with FrameFlow (91.3%91.3\% vs. 63.3%63.3\% of residues in the Ramachandran target), outperforms OC-Flow on peptide binding affinity at 2×2\times the speed, and reduces Rosetta energies by thousands of units across the PDB test set for structures with hundreds of residues.

Added

2026-09-29

Normalizing Flows are Capable Generative Models

Normalizing Flows are Capable Generative Models

Shuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot, Jiatao Gu, Huangjie Zheng, Tianrong Chen, Miguel ngel Bautista, Navdeep Jaitly, Joshua M. Susskind

OrganizationsApple

Why you should read this

Demonstrates that normalizing flows can rival diffusion models in image generation quality while achieving state-of-the-art exact likelihood estimation through a scalable Transformer-based architecture paired with noise augmentation and guidance.

Normalizing Flows (NFs) are likelihood-based models for continuous inputs. They have demonstrated promising results on both density estimation and generative modeling tasks, but have received relatively little attention in recent years. In this work, we demonstrate that NFs are more powerful than previously believed. We present TARFLOW: a simple and scalable architecture that enables highly performant NF models. TARFLOW can be thought of as a Transformer-based variant of Masked Autoregressive Flows (MAFs): it consists of a stack of autoregressive Transformer blocks on image patches, alternating the autoregression direction between layers. TARFLOW is straightforward to train end-to-end, and capable of directly modeling and generating pixels. We also propose three key techniques to improve sample quality: Gaussian noise augmentation during training, a post training denoising procedure, and an effective guidance method for both class-conditional and unconditional settings. Putting these together, TARFLOW sets new state-of-the-art results on likelihood estimation for images, beating the previous best methods by a large margin, and generates samples with quality and diversity comparable to diffusion models, for the first time with a stand-alone NF model. We make our code available at https://github.com/apple/ml-tarflow.

Added

2026-09-26

π0: A Vision-Language-Action Flow Model for General Robot Control

π0: A Vision-Language-Action Flow Model for General Robot Control

Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, Ury Zhilinsky

Why you should read this

Presents π₀, a vision-language-action model that combines flow matching with pre-trained vision-language representations to enable generalist, highly dexterous control across single-arm, dual-arm, and mobile manipulation platforms.

Robot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artificial intelligence. However, bringing robot learning to the level of generality required for effective real-world systems faces major obstacles in terms of data, generalization, and robustness. In this paper, we discuss how generalist robot policies (i.e., robot foundation models) can address these challenges, and how we can design effective generalist robot policies for complex and highly dexterous tasks. We propose a novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge. We then discuss how this model can be trained on a large and diverse dataset from multiple dexterous robot platforms, including single-arm robots, dual-arm robots, and mobile manipulators. We evaluate our model in terms of its ability to perform tasks in zero shot after pre-training, follow language instructions from people and from a high-level VLM policy, and its ability to acquire new skills via fine-tuning. Our results cover a wide variety of tasks, such as laundry folding, table cleaning, and assembling boxes.

Added

2026-09-14

Thinking with Looped Flows

Thinking with Looped Flows

Ayhan Suleymanzade, Chanhyuk Lee, Floor Eijkelboom, Nicholas M. Boffi, İsmail İlkan Ceylan, Jinwoo Kim

OrganizationsAITHYRACarnegie Mellon UniversityÉcole Polytechnique Fédérale de LausanneKorea Advanced Institute of Science and TechnologyTU WienUniversity of AmsterdamUniversity of Oxford

Why you should read this

Introduces looped flows, a framework that trains recurrent architectures through progressive denoising objectives to overcome gradient truncation, enabling dynamic test-time compute scaling via probability flow integration and achieving leading accuracy on ARC-AGI reasoning benchmarks.

Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local denoising objectives. By imposing temporal association across denoising objectives through progressively decreasing noise levels and shared noise, the model is incentivized to learn recurrent states that transfer useful computation over time, even when gradients cover only a few updates. We then formulate inference as integrating the velocity of a probability flow parameterized by the learned denoiser, coupled with recurrent states. This allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples. Across six reasoning benchmarks including two multi-solution benchmarks, looped flows outperform prior state-of-the-art looped models overall, achieving 58.8% test accuracy on ARC-AGI-1 and 12.2% on ARC-AGI-2.

Added

2026-09-13

Creative Commons License
Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning

Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning

Mariia Drozdova, Aidan Sirbu, Pietro Miotti, Robert Obryk, Mayalen Etcheverry, Eyvind Niklasson, Blake Richards

OrganizationsCIFARGoogleMcGill UniversityMila – Québec Artificial Intelligence InstituteMontreal Neurological InstituteUniversity of Geneva

Why you should read this

Demonstrates that diffusion acts as a training curriculum rather than an inference procedure, allowing recurrent networks with persistent hidden states to achieve near-perfect accuracy on hard reasoning tasks through arbitrary-depth iteration and constant noise injection without requiring search, verifiers, or denoising schedules.

Diffusion models and recursive reasoners are both iterative, but they carry information across iterations differently. We add a persistent hidden state to a diffusion denoiser and remove its timestep conditioning, leaving a single shared update that can be run to arbitrary depth. The result is an anytime solver: accuracy keeps improving with inference depth far beyond the rollout lengths and backpropagation window used in training, reaching 99.90% exact solve on Sudoku-Extreme. We also obtain 98.93% solve rate on Maze-Unique. Surprisingly, progressive denoising is unnecessary at inference: holding corruption at its maximum by replacing every non-clue variable with fresh Gaussian noise at each step retains near-perfect solving and converges to stable solutions. This simple noise-injection mechanism enables a single trajectory to efficiently explore the solution space and settle on the correct answer without parallel rollouts, candidate selection, or external verifiers required by prior reasoning models. Nonetheless, ordered annealed corruption remains critical during training, which suggests that diffusion's primary contribution to our anytime solver is not a sampling procedure at inference, but a denoising training curriculum.

Added

2026-09-10

Creative Commons License
Flow Reasoning Models: Turning Flows Into Efficient Recurrent Reasoners

Flow Reasoning Models: Turning Flows Into Efficient Recurrent Reasoners

Alec Helbling, Andrey Bryutkin, Mauro Martino, Duen Horng Chau, Nima Dehmamy, Hendrik Strobelt

OrganizationsGeorgia Institute of TechnologyIBMMassachusetts Institute of TechnologyMIT-IBM Watson AI Lab

Why you should read this

Introduces Flow Reasoning Models, a continuous-flow framework that iteratively refines discrete decisions in parallel to solve complex structured reasoning benchmarks with state-of-the-art accuracy using 44 times fewer inference FLOPs than competing diffusion baselines.

Structured reasoning requires making and revising interdependent decisions to reach a globally consistent solution. Existing architectures struggle with this: autoregressive models commit sequentially and cannot revise earlier decisions, while masked diffusion models often require careful decoding schemes to coordinate interdependent predictions. We introduce Flow Reasoning Models (FRMs), a novel framework for structured reasoning that adapts continuous flows over discrete structured outputs with a simple recurrent refinement mechanism. By self-conditioning a flow model on its own past outputs, we turn one-shot denoising into iterative solution refinement. This lets FRMs make and revise decisions in parallel, efficiently coordinating interdependent choices across solutions. Yet conventional self-conditioning becomes unreliable at greater recurrent depth due to exposure bias between one-step training predictions and recursively generated inference states. We address this mismatch with Fixed-Point Forcing (FPF), which trains FRMs on states produced by their own inference dynamics while preserving the standard flow-matching objective. FRMs achieve solve rates of 99.5%99.5\%, 100.0%100.0\%, and 99.9%99.9\% on Sudoku-Extreme, Zebra, and Maze-Unique, respectively. On Sudoku-Extreme, FRMs achieve higher peak accuracy than the evaluated masked-diffusion and specialized reasoning baselines while remaining highly compute-efficient, matching the next-best method's 98.7%98.7\% peak solve rate with 44×44\times fewer inference FLOPs.

Added

2026-09-05

Creative Commons License
ELF: Embedded Language Flows

ELF: Embedded Language Flows

Keya Hu, Linlu Qiu, Yiyang Lu, Hanhong Zhao, Tianhong Li, Yoon Kim, Jacob Andreas, Kaiming He

OrganizationsMassachusetts Institute of Technology

Why you should read this

Proposes Embedded Language Flows (ELF), a novel class of continuous diffusion language models that achieve superior generation quality with fewer sampling steps than existing discrete and continuous DLMs by operating predominantly in continuous embedding space and leveraging continuous-time Flow Matching.

Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their success has attracted growing interest in applying them to language modeling. Unlike their image-domain counterparts, today's leading diffusion language models (DLMs) primarily operate over discrete tokens. In this paper, we show that continuous DLMs can be made effective with minimal adaptation to the discrete domain. We propose Embedded Language Flows (ELF), a class of diffusion models in continuous embedding space based on continuous-time Flow Matching. Unlike existing DLMs, ELF predominantly stays within the continuous embedding space until the final time step, where it maps to discrete tokens using a shared-weight network. This formulation makes it straightforward to adapt established techniques from image-domain diffusion models, e.g., classifier-free guidance (CFG). Experiments show that ELF substantially outperforms leading discrete and continuous DLMs, achieving better generation quality with fewer sampling steps. These results suggest that ELF offers a promising path toward effective continuous DLMs.

Added

2026-05-18

Creative Commons License
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, Saining Xie

OrganizationsKorea Advanced Institute of Science and TechnologyKorea UniversityNew York UniversityScaled Foundations

Why you should read this

Introduces REPA, a simple regularization framework that achieves over 17.5x faster training convergence and state-of-the-art image generation quality by aligning diffusion transformer representations with those from pretrained self-supervised visual encoders.

Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that one main bottleneck in training large-scale diffusion models for generation lies in effectively learning these representations. Moreover, training can be made easier by incorporating high-quality external visual representations, rather than relying solely on the diffusion models to learn them independently. We study this by introducing a straightforward regularization called REPresentation Alignment (REPA), which aligns the projections of noisy input hidden states in denoising networks with clean image representations obtained from external, pretrained visual encoders. The results are striking: our simple strategy yields significant improvements in both training efficiency and generation quality when applied to popular diffusion and flow-based transformers, such as DiTs and SiTs. For instance, our method can speed up SiT training by over 17.5×\times, matching the performance (without classifier-free guidance) of a SiT-XL model trained for 7M steps in less than 400K steps. In terms of final generation quality, our approach achieves state-of-the-art results of FID=1.42 using classifier-free guidance with the guidance interval.

Added

2026-05-18

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Robin Rombach

OrganizationsStability AI

Why you should read this

Proposes a multimodal transformer architecture and tailored noise sampling strategy that enables rectified flow models to scale efficiently, outperforming existing state-of-the-art systems in high-resolution image synthesis and typographic accuracy.

Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos. Rectified flow is a recent generative model formulation that connects data and noise in a straight line. Despite its better theoretical properties and conceptual simplicity, it is not yet decisively established as standard practice. In this work, we improve existing noise sampling techniques for training rectified flow models by biasing them towards perceptually relevant scales. Through a large-scale study, we demonstrate the superior performance of this approach compared to established diffusion formulations for high-resolution text-to-image synthesis. Additionally, we present a novel transformer-based architecture for text-to-image generation that uses separate weights for the two modalities and enables a bidirectional flow of information between image and text tokens, improving text comprehension, typography, and human preference ratings. We demonstrate that this architecture follows predictable scaling trends and correlates lower validation loss to improved text-to-image synthesis as measured by various metrics and human evaluations. Our largest models outperform state-of-the-art models, and we will make our experimental data, code, and model weights publicly available.

Added

2026-05-13

Flow Matching for Generative Modeling

Flow Matching for Generative Modeling

Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, Matt Le

OrganizationsMetaWeizmann Institute of Science

Why you should read this

Generalizes diffusion into a Continuous Normalizing Flow (CNF) framework trained with a simulation-free objective.

We introduce a new paradigm for generative modeling built on Continuous Normalizing Flows (CNFs), allowing us to train CNFs at unprecedented scale. Specifically, we present the notion of Flow Matching (FM), a simulation-free approach for training CNFs based on regressing vector fields of fixed conditional probability paths. Flow Matching is compatible with a general family of Gaussian probability paths for transforming between noise and data samples -- which subsumes existing diffusion paths as specific instances. Interestingly, we find that employing FM with diffusion paths results in a more robust and stable alternative for training diffusion models. Furthermore, Flow Matching opens the door to training CNFs with other, non-diffusion probability paths. An instance of particular interest is using Optimal Transport (OT) displacement interpolation to define the conditional probability paths. These paths are more efficient than diffusion paths, provide faster training and sampling, and result in better generalization. Training CNFs using Flow Matching on ImageNet leads to consistently better performance than alternative diffusion-based methods in terms of both likelihood and sample quality, and allows fast and reliable sample generation using off-the-shelf numerical ODE solvers.

Added

2026-02-25