SE(3)-Stochastic Flow Matching for Protein Backbone Generation

Avishek Joey BoseTara Akhound-SadeghGuillaume HuguetKilian FatrasJarrid Rector-BrooksCheng-Hao LiuAndrei Cristian NicaMaksym KorablyovMichael M. BronsteinAlexander Tong

article2024ICLR200 citations

Introduces FoldFlow, a generative framework combining SE(3) flow matching and Riemannian optimal transport to produce designable, diverse protein backbones with faster training and greater stability than diffusion models.

Listen

Designing novel protein structures computationally is a critical frontier for biotechnology and drug discovery, offering solutions to major global health challenges by enabling targeted therapeutics and molecular binders. However, current machine learning approaches for 3D protein backbone generation face steep engineering bottlenecks: diffusion-based models require expensive numerical simulations, demand massive computing infrastructure, and remain fundamentally restricted to starting from standard, uninformed random noise distributions. These limitations hinder practical workflows, particularly when modeling dynamic structural changes or training models under standard computational budgets.

The article develops and evaluates FoldFlow, a new family of continuous generative models tailored to the geometric symmetries of 3D rigid motions in protein backbones. Its main objective is to establish a faster, simulation-free modeling framework that generates highly designable, diverse, and novel protein backbones across sequence lengths of up to 300 amino acids, while supporting flexible generation from informative starting distributions.

To achieve this, the authors introduce three flow-matching models: FoldFlow-Base, which learns deterministic trajectories without numerical simulation; FoldFlow-OT, which incorporates optimal transport to straighten trajectories and stabilize training; and FoldFlow-SFM, which integrates stochastic dynamics via a simulation-free approximation of a Brownian bridge. The models are evaluated on a benchmark dataset of 22,248 protein structures from the Protein Data Bank and tested against established diffusion baselines. In addition, the framework is applied to an equilibrium conformation task using molecular dynamics trajectories of the bovine pancreatic trypsin inhibitor protein to assess multi-state conformational modeling.

The findings show that FoldFlow trains more than twice as fast per step as the leading non-pretrained baseline, FrameDiff, completing full training in approximately 2.5 days on four graphics processing units compared to ten or more days for baselines. In structural quality, FoldFlow-OT achieves an 82.0% in-silico designability rate, outperforming all non-pretrained models. In generating novel and designable structures, FoldFlow-SFM achieves a 54.4% success rate, exceeding all comparable baselines. In dynamic equilibrium modeling, using an informed starting distribution with FoldFlow significantly improves structural distribution matching over random priors and baselines, successfully capturing multiple flexible conformational states that standard static predictors fail to resolve.

These results demonstrate that flow matching provides a substantially more computationally efficient and functionally versatile alternative to diffusion models for structural biology. By reducing training times and hardware requirements, the approach lowers computational costs and accelerates prototyping timelines for drug candidate design. Furthermore, the capacity to start from informative prior distributions opens practical avenues for modeling flexible molecular interactions, protein-ligand docking, and conformational transitions without full-scale, expensive molecular dynamics simulations.

Organizations advancing machine-learning-driven protein engineering should consider adopting simulation-free flow-matching frameworks to streamline generative pipelines and improve structural novelty. When maximum designability is required, deterministic optimal transport models should be selected, whereas stochastic variants should be deployed when discovering diverse and structurally unique drug candidates is the primary goal. Teams should also conduct pilot implementations on conditional generation tasks, such as incorporating target binding sites and functional constraints directly into training.

While the results are strong across standard computational metrics, confidence is subject to clear experimental boundaries: evaluation relies on in-silico self-consistency filters using external predictive models rather than wet-lab synthesis. Additionally, performance metrics drop slightly on longer chains of 250 to 300 amino acids. Stakeholders should therefore validate high-priority generative candidates with laboratory assays before advancing them into formal development pipelines.

Cover for SE(3)-Stochastic Flow Matching for Protein Backbone Generation

Abstract

The computational design of novel protein structures has the potential to impact numerous scientific disciplines greatly. Toward this goal, we introduce FoldFlow, a series of novel generative models of increasing modeling power based on the flow-matching paradigm over 3D3\mathrm{D} rigid motions -- i.e. the group SE(3)\text{SE}(3) -- enabling accurate modeling of protein backbones. We first introduce FoldFlow-Base, a simulation-free approach to learning deterministic continuous-time dynamics and matching invariant target distributions on SE(3)\text{SE}(3). We next accelerate training by incorporating Riemannian optimal transport to create FoldFlow-OT, leading to the construction of both more simple and stable flows. Finally, we design FoldFlow-SFM, coupling both Riemannian OT and simulation-free training to learn stochastic continuous-time dynamics over SE(3)\text{SE}(3). Our family of FoldFlow, generative models offers several key advantages over previous approaches to the generative modeling of proteins: they are more stable and faster to train than diffusion-based approaches, and our models enjoy the ability to map any invariant source distribution to any invariant target distribution over SE(3)\text{SE}(3). Empirically, we validate FoldFlow, on protein backbone generation of up to 300300 amino acids leading to high-quality designable, diverse, and novel samples.

Citation

MLA
Bose, A. J., et al. “SE(3)-Stochastic Flow Matching for Protein Backbone Generation”. arXiv, 2023, http://arxiv.org/abs/2310.02391v4.
APA
Bose, A. J., Akhound-Sadegh, T., Huguet, G., Fatras, K., Rector-Brooks, J., Liu, C.-H., Nica, A. C., Korablyov, M., Bronstein, M., & Tong, A. (2023). SE(3)-Stochastic Flow Matching for Protein Backbone Generation. arXiv. http://arxiv.org/abs/2310.02391v4
Chicago
Bose, A. J., T. Akhound-Sadegh, G. Huguet, et al. 2023. “SE(3)-Stochastic Flow Matching for Protein Backbone Generation”. arXiv. http://arxiv.org/abs/2310.02391v4.
Harvard
Bose, A.J. et al. (2023) “SE(3)-Stochastic Flow Matching for Protein Backbone Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.02391v4.
Vancouver
1. Bose AJ, Akhound-Sadegh T, Huguet G, Fatras K, Rector-Brooks J, Liu C-H, Nica AC, Korablyov M, Bronstein M, Tong A (2023) SE(3)-Stochastic Flow Matching for Protein Backbone Generation. arXiv

BibTeX

@article{bose2023stochastic,
  title = {SE(3)-Stochastic Flow Matching for Protein Backbone Generation},
  author = {Bose, Avishek Joey and Akhound-Sadegh, Tara and Huguet, Guillaume and Fatras, Kilian and Rector-Brooks, Jarrid and Liu, Cheng-Hao and Nica, Andrei Cristian and Korablyov, Maksym and Bronstein, Michael and Tong, Alexander},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.02391v4},
  eprint = {2310.02391}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors