SE(3)-Stochastic Flow Matching for Protein Backbone Generation
Avishek Joey BoseTara Akhound-SadeghGuillaume HuguetKilian FatrasJarrid Rector-BrooksCheng-Hao LiuAndrei Cristian NicaMaksym KorablyovMichael M. BronsteinAlexander Tong
Introduces FoldFlow, a generative framework combining SE(3) flow matching and Riemannian optimal transport to produce designable, diverse protein backbones with faster training and greater stability than diffusion models.
Designing novel protein structures computationally is a critical frontier for biotechnology and drug discovery, offering solutions to major global health challenges by enabling targeted therapeutics and molecular binders. However, current machine learning approaches for 3D protein backbone generation face steep engineering bottlenecks: diffusion-based models require expensive numerical simulations, demand massive computing infrastructure, and remain fundamentally restricted to starting from standard, uninformed random noise distributions. These limitations hinder practical workflows, particularly when modeling dynamic structural changes or training models under standard computational budgets.
The article develops and evaluates FoldFlow, a new family of continuous generative models tailored to the geometric symmetries of 3D rigid motions in protein backbones. Its main objective is to establish a faster, simulation-free modeling framework that generates highly designable, diverse, and novel protein backbones across sequence lengths of up to 300 amino acids, while supporting flexible generation from informative starting distributions.
To achieve this, the authors introduce three flow-matching models: FoldFlow-Base, which learns deterministic trajectories without numerical simulation; FoldFlow-OT, which incorporates optimal transport to straighten trajectories and stabilize training; and FoldFlow-SFM, which integrates stochastic dynamics via a simulation-free approximation of a Brownian bridge. The models are evaluated on a benchmark dataset of 22,248 protein structures from the Protein Data Bank and tested against established diffusion baselines. In addition, the framework is applied to an equilibrium conformation task using molecular dynamics trajectories of the bovine pancreatic trypsin inhibitor protein to assess multi-state conformational modeling.
The findings show that FoldFlow trains more than twice as fast per step as the leading non-pretrained baseline, FrameDiff, completing full training in approximately 2.5 days on four graphics processing units compared to ten or more days for baselines. In structural quality, FoldFlow-OT achieves an 82.0% in-silico designability rate, outperforming all non-pretrained models. In generating novel and designable structures, FoldFlow-SFM achieves a 54.4% success rate, exceeding all comparable baselines. In dynamic equilibrium modeling, using an informed starting distribution with FoldFlow significantly improves structural distribution matching over random priors and baselines, successfully capturing multiple flexible conformational states that standard static predictors fail to resolve.
These results demonstrate that flow matching provides a substantially more computationally efficient and functionally versatile alternative to diffusion models for structural biology. By reducing training times and hardware requirements, the approach lowers computational costs and accelerates prototyping timelines for drug candidate design. Furthermore, the capacity to start from informative prior distributions opens practical avenues for modeling flexible molecular interactions, protein-ligand docking, and conformational transitions without full-scale, expensive molecular dynamics simulations.
Organizations advancing machine-learning-driven protein engineering should consider adopting simulation-free flow-matching frameworks to streamline generative pipelines and improve structural novelty. When maximum designability is required, deterministic optimal transport models should be selected, whereas stochastic variants should be deployed when discovering diverse and structurally unique drug candidates is the primary goal. Teams should also conduct pilot implementations on conditional generation tasks, such as incorporating target binding sites and functional constraints directly into training.
While the results are strong across standard computational metrics, confidence is subject to clear experimental boundaries: evaluation relies on in-silico self-consistency filters using external predictive models rather than wet-lab synthesis. Additionally, performance metrics drop slightly on longer chains of 250 to 300 amino acids. Stakeholders should therefore validate high-priority generative candidates with laboratory assays before advancing them into formal development pipelines.
- Paper: Flow Matching for Generative Modeling, Yaron Lipman et al. (2023). Introduces the simulation-free Flow Matching framework with optimal transport probability paths that FoldFlow generalizes to the SE(3) group.
- Paper: SE(3) diffusion model with application to protein backbone generation, Jason Yim et al. (2023). Establishes the foundation of framing protein backbone generation as rigid body transformations over SE(3) using SE(3) diffusion models (FrameDiff), which FoldFlow directly builds upon and improves with flow matching.
- Paper: Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow, Xingchao Liu et al. (2023). Provides fundamental principles of continuous-time flow generation between arbitrary distributions via straight trajectories and optimal transport.
- Paper: Equivariant Diffusion for Molecule Generation in 3D, Emiel Hoogeboom et al. (2022). Pioneers the use of SE(3) and E(3) equivariant generative modeling for 3D molecular structures, setting the geometric principles leveraged in rigid-body protein modeling.
- Paper: Learning inverse folding from millions of predicted structures, Chloe Hsu et al. (2022). Provides the machine learning methods for evaluating generated backbones via fixed-backbone inverse folding and sequence recovery metrics.
- Paper: Stochastic Interpolants: A Unifying Framework for Flows and Diffusions, Michael S. Albergo et al. (2025). Extends and unifies deterministic flow matching and continuous-time stochastic interpolants within a comprehensive theoretical framework.
- Paper: Mean Flows for One-step Generative Modeling, Zhengyang Geng et al. (2025). Advances flow matching dynamics toward single-step generative modeling by learning mean velocity vector fields.
- Paper: Variational Flow Maps: Make Some Noise for One-Step Conditional Generation, Abbas Mammadov et al. (2026). Extends continuous flow matching models to efficient, few-step conditional generation and inverse problems.
