SE(3) diffusion model with application to protein backbone generation

Jason YimBrian L. TrippeValentin De BortoliEmile MathieuArnaud DoucetRegina BarzilayTommi S. Jaakkola

article2023ICML307 citations

Introduces FrameDiff, an SE(3) frame diffusion model that generates novel, designable protein backbones up to 500 amino acids long without relying on pretrained structure prediction networks.

Listen

Designing novel protein backbones computationally is a vital step toward creating targeted therapeutics and biomaterials, but traditional engineering requires laborious laboratory experiments and deep domain expertise. While recent generative artificial intelligence models simulate 3D rigid-body transformations (rotations and translations) to design proteins, they rely on heuristic training techniques and expensive pretrained networks rather than mathematically rigorous foundations.

To resolve this gap, the article establishes a theoretical framework for diffusion modeling over rigid 3D motions and introduces FrameDiff, a generative method designed to sample monomeric protein backbones. FrameDiff operates by separating translational and rotational Brownian motions, enforcing geometric rotational invariance by pinning the diffusion process at the center of mass, and predicting coordinates using an equivariant neural network combined with auxiliary structural loss functions to eliminate physical errors like atomic clashes.

The evaluation evaluated FrameDiff on 20,312 experimentally determined protein structures from the Protein Data Bank. Key findings show that FrameDiff produces highly designable proteins—structures for which stable amino acid sequences can be successfully identified—reaching up to an 84% in-silico designability rate without any pretraining on protein structure prediction. Reducing sampling noise significantly improved designability across protein lengths from 100 to 500 amino acids while maintaining high structural diversity (>0.5 unique cluster proportions). FrameDiff also generated completely novel backbones structurally distant from known natural proteins and executed sample generation over ten times faster than leading alternative tools.

These results demonstrate that principled generative diffusion models can design biologically viable and novel protein backbones with roughly one-fourth the parameters of prior models and substantially lower computational overhead. By proving that expensive pretraining is unnecessary for high-quality protein backbone generation, this framework substantially reduces the cost, training timelines, and barrier to entry for developing computational biotherapeutics.

Before adopting these generated structures for clinical or real-world biomanufacturing pipelines, stakeholders should conduct wet-lab experimental characterization to validate that computationally designable backbones fold correctly and function in physical biological settings. Technical teams should pursue extensions into conditional design, such as motif scaffolding and multimeric complexes. Given that the current evaluation relies purely on in-silico structure-prediction algorithms and exhibits reduced accuracy on proteins exceeding 400 residues, decision-makers should treat purely computational success metrics as strong pilot evidence rather than guaranteed physical validation.

  • Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). Demonstrates how to scale diffusion models by replacing traditional backbones with transformer architectures, providing insights for scaling SE(3) diffusion models beyond standard graph backbones.
  • Paper: Stochastic Interpolants: A Unifying Framework for Flows and Diffusions, Michael S. Albergo et al. (2025). Presents a generalized continuous-time framework uniting flows and diffusions on finite horizons, offering potential mathematical generalizations for SE(3) manifold generation.
Cover for SE(3) diffusion model with application to protein backbone generation

Abstract

The design of novel protein structures remains a challenge in protein engineering for applications across biomedicine and chemistry. In this line of work, a diffusion model over rigid bodies in 3D (referred to as frames) has shown success in generating novel, functional protein backbones that have not been observed in nature. However, there exists no principled methodological framework for diffusion on SE(3), the space of orientation preserving rigid motions in R3, that operates on frames and confers the group invariance. We address these shortcomings by developing theoretical foundations of SE(3) invariant diffusion models on multiple frames followed by a novel framework, FrameDiff, for learning the SE(3) equivariant score over multiple frames. We apply FrameDiff on monomer backbone generation and find it can generate designable monomers up to 500 amino acids without relying on a pretrained protein structure prediction network that has been integral to previous methods. We find our samples are capable of generalizing beyond any known protein structure.

Citation

MLA
Yim, J., et al. “SE(3) Diffusion Model with Application to Protein Backbone Generation”. International Conference of Machine Learning (ICML) 2023, 2023, http://arxiv.org/abs/2302.02277v3.
APA
Yim, J., Trippe, B. L., Bortoli, V. D., Mathieu, E., Doucet, A., Barzilay, R., & Jaakkola, T. (2023). SE(3) diffusion model with application to protein backbone generation. International Conference of Machine Learning (ICML) 2023. http://arxiv.org/abs/2302.02277v3
Chicago
Yim, J., B. L. Trippe, V. D. Bortoli, et al. 2023. “SE(3) Diffusion Model with Application to Protein Backbone Generation”. International Conference of Machine Learning (ICML) 2023. http://arxiv.org/abs/2302.02277v3.
Harvard
Yim, J. et al. (2023) “SE(3) diffusion model with application to protein backbone generation”, International Conference of Machine Learning (ICML) 2023 [Preprint]. Available at: http://arxiv.org/abs/2302.02277v3.
Vancouver
1. Yim J, Trippe BL, Bortoli VD, Mathieu E, Doucet A, Barzilay R, Jaakkola T (2023) SE(3) diffusion model with application to protein backbone generation. International Conference of Machine Learning (ICML) 2023

BibTeX

@article{yim2023diffusion,
  title = {SE(3) diffusion model with application to protein backbone generation},
  author = {Yim, Jason and Trippe, Brian L. and Bortoli, Valentin De and Mathieu, Emile and Doucet, Arnaud and Barzilay, Regina and Jaakkola, Tommi},
  year = {2023},
  journal = {International Conference of Machine Learning (ICML) 2023},
  url = {http://arxiv.org/abs/2302.02277v3},
  eprint = {2302.02277}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/