Learning inverse folding from millions of predicted structures
Chloe HsuRobert VerkuilJason LiuZeming LinBrian HieTom SercuAdam LererAlexander Rives
Demonstrates that training a geometric transformer on 12 million AlphaFold2-predicted structures improves fixed-backbone sequence recovery by nearly ten percentage points and generalizes to multi-state proteins, complexes, and variant effect prediction.
Designing novel amino acid sequences that fold into predefined three-dimensional structures—a challenge known as inverse folding or fixed-backbone protein design—is a critical capability in bioengineering and therapeutic development. While machine learning offers an attractive alternative to traditional physics-based energy modeling, existing deep learning models have been bottlenecked by the limited number of experimentally determined structures. Currently, experimental structural databases cover less than 0.1% of known protein sequences, restricting model scale and generalization.
The article demonstrates that augmenting training data with millions of computationally predicted structures enables larger deep learning models to achieve state-of-the-art accuracy across fixed-backbone design and mutation effect prediction tasks. Specifically, the researchers evaluated whether predicting structures for 12 million sequences using AlphaFold2 could effectively overcome the scarcity of experimental training data.
To conduct this evaluation, the team predicted structures for 12 million UniRef50 sequences, generating a synthetic dataset roughly 750 times larger than the standard experimental training corpus from the CATH database. They trained autoregressive encoder-decoder architectures, including graph neural networks and a hybrid architecture that pairs geometric vector perceptron encoding with standard transformer layers. Crucially, the authors introduced coordinate span masking and added minor Gaussian coordinate noise to prevent synthetic artifacts from biasing the models, while rigorously splitting datasets by topology and structure similarity to avoid data leakage.
The findings show that training with predicted structures provides a substantial performance leap when paired with scaled architectures. First, the hybrid model trained on predicted structures achieved a 51.6% native sequence recovery rate on structurally held-out test backbones (and 72% recovery on buried core residues), representing an improvement of nearly 10 percentage points over baseline models trained strictly on experimental data. Second, smaller legacy models degraded when exposed to predicted structures, proving that increased parameter scale is essential to leverage synthetic structural data effectively. Third, models trained with coordinate span masking successfully generalized to complex design scenarios, including multi-chain protein complexes, multi-state dynamic conformations, and partially masked backbones. Finally, in zero-shot mutational effect prediction, the trained models achieved strong correlation with experimental measurements, including a 0.69 rank correlation in predicting human receptor binding affinity for the SARS-CoV-2 spike protein without requiring task-specific fine-tuning.
These results establish that the bottleneck in computational protein design is largely data availability rather than model architecture alone. By utilizing predicted structures analogously to back-translation in natural language processing, engineering teams can dramatically improve design fidelity, shorten development cycles for novel biologics, and better predict functional mutation impacts without incurring the severe time and financial costs of experimental structural characterization.
Organizations developing computational protein design workflows should adopt large-scale hybrid transformer architectures pre-trained on high-confidence predicted structures. Implementers should incorporate coordinate noise and span masking to enable robust modeling of flexible regions, binding interfaces, and structural multi-states. Further prospective wet-lab experimental validation is recommended to confirm in vitro and in vivo activity for fully de novo designed sequences.
Confidence in these findings is supported by rigorous topology-level and structural similarity holdouts across diverse benchmarks. However, users should note that the approach relies on the baseline accuracy of structure prediction models, and performance naturally declines on unconstrained, highly exposed surface residues compared to tightly packed protein cores.
- Paper: Highly accurate protein structure prediction with AlphaFold, John Jumper et al. (2021). AlphaFold2 provides the foundational protein structure prediction system used by the source paper to generate the 12 million synthetic training structures for inverse folding.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). This paper introduces the sequence-to-sequence Transformer architecture that forms the core predictive framework adapted with geometric layers in the source model.
- Paper: E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials, Simon Batzner et al. (2021). This work establishes 3D geometric invariant and equivariant representations for molecular structures, informing the design of invariant geometric input processing layers for protein backbones.
No sufficiently relevant recommendations were found.
