Learning inverse folding from millions of predicted structures

Chloe HsuRobert VerkuilJason LiuZeming LinBrian HieTom SercuAdam LererAlexander Rives

article2022ICML548 citationsOutstanding Paper Runner Up

Demonstrates that training a geometric transformer on 12 million AlphaFold2-predicted structures improves fixed-backbone sequence recovery by nearly ten percentage points and generalizes to multi-state proteins, complexes, and variant effect prediction.

Listen

Designing novel amino acid sequences that fold into predefined three-dimensional structures—a challenge known as inverse folding or fixed-backbone protein design—is a critical capability in bioengineering and therapeutic development. While machine learning offers an attractive alternative to traditional physics-based energy modeling, existing deep learning models have been bottlenecked by the limited number of experimentally determined structures. Currently, experimental structural databases cover less than 0.1% of known protein sequences, restricting model scale and generalization.

The article demonstrates that augmenting training data with millions of computationally predicted structures enables larger deep learning models to achieve state-of-the-art accuracy across fixed-backbone design and mutation effect prediction tasks. Specifically, the researchers evaluated whether predicting structures for 12 million sequences using AlphaFold2 could effectively overcome the scarcity of experimental training data.

To conduct this evaluation, the team predicted structures for 12 million UniRef50 sequences, generating a synthetic dataset roughly 750 times larger than the standard experimental training corpus from the CATH database. They trained autoregressive encoder-decoder architectures, including graph neural networks and a hybrid architecture that pairs geometric vector perceptron encoding with standard transformer layers. Crucially, the authors introduced coordinate span masking and added minor Gaussian coordinate noise to prevent synthetic artifacts from biasing the models, while rigorously splitting datasets by topology and structure similarity to avoid data leakage.

The findings show that training with predicted structures provides a substantial performance leap when paired with scaled architectures. First, the hybrid model trained on predicted structures achieved a 51.6% native sequence recovery rate on structurally held-out test backbones (and 72% recovery on buried core residues), representing an improvement of nearly 10 percentage points over baseline models trained strictly on experimental data. Second, smaller legacy models degraded when exposed to predicted structures, proving that increased parameter scale is essential to leverage synthetic structural data effectively. Third, models trained with coordinate span masking successfully generalized to complex design scenarios, including multi-chain protein complexes, multi-state dynamic conformations, and partially masked backbones. Finally, in zero-shot mutational effect prediction, the trained models achieved strong correlation with experimental measurements, including a 0.69 rank correlation in predicting human receptor binding affinity for the SARS-CoV-2 spike protein without requiring task-specific fine-tuning.

These results establish that the bottleneck in computational protein design is largely data availability rather than model architecture alone. By utilizing predicted structures analogously to back-translation in natural language processing, engineering teams can dramatically improve design fidelity, shorten development cycles for novel biologics, and better predict functional mutation impacts without incurring the severe time and financial costs of experimental structural characterization.

Organizations developing computational protein design workflows should adopt large-scale hybrid transformer architectures pre-trained on high-confidence predicted structures. Implementers should incorporate coordinate noise and span masking to enable robust modeling of flexible regions, binding interfaces, and structural multi-states. Further prospective wet-lab experimental validation is recommended to confirm in vitro and in vivo activity for fully de novo designed sequences.

Confidence in these findings is supported by rigorous topology-level and structural similarity holdouts across diverse benchmarks. However, users should note that the approach relies on the baseline accuracy of structure prediction models, and performance naturally declines on unconstrained, highly exposed surface residues compared to tightly packed protein cores.

No sufficiently relevant recommendations were found.

Cover for Learning inverse folding from millions of predicted structures

Abstract

We consider the problem of predicting a protein sequence from its backbone atom coordinates. Machine learning approaches to this problem to date have been limited by the number of available experimentally determined protein structures. We augment training data by nearly three orders of magnitude by predicting structures for 12M protein sequences using AlphaFold2. Trained with this additional data, a sequence-to-sequence transformer with invariant geometric input processing layers achieves 51% native sequence recovery on structurally held-out backbones with 72% recovery for buried residues, an overall improvement of almost 10 percentage points over existing methods. The model generalizes to a variety of more complex tasks including design of protein complexes, partially masked structures, binding interfaces, and multiple states.

Table of Contents

  • 1. Introduction
  • 2. Learning inverse folding from predicted structures
  • 2.1. Data
  • 2.2. Model architectures
  • 2.3. Training
  • 3. Results
  • 3.1. Fixed backbone protein design
  • 3.2. Zero-shot predictions
  • 4. Related work
  • 5. Conclusions
  • References
  • A. Additional details on datasets, training procedures, and model architectures
  • A.1. Details on dataset of predicted structures
  • A.2. Details on span masking
  • A.3. Details on model architectures
  • B. TM-score-based test set
  • C. Additional results and details

Knowls

  1. Knowl 1 — GVP-Transformer Architecture for Invariant Inverse Protein Folding

    model/method

    Inverse protein folding models the conditional distribution p(Y∣X)p(Y \mid X) of an amino acid sequence Y=(y1,…,yn)Y = (y_1, \dots, y_n) given backbone 3D coordinates X=(x1,…,x3n)X = (x_1, \dots, x_{3n}) of the protein's NN, CαC_\alpha, and CC atoms. To ensure invariance under any 3D Euclidean transformation T=(R,t)T = (R, t) such that p(Y∣X)=p(Y∣TX)p(Y \mid X) = p(Y \mid TX), the 142M-parameter GVP-Transformer combines geometric vector perceptron graph layers with a standard Transformer encoder-decoder:

    1. Equivariant Geometric Feature Extraction: An initial structural encoder of 4 Geometric Vector Perceptron Graph Neural Network (GVP-GNN) layers processes translation-invariant geometric inputs, maintaining rotation equivariance for vector features and rotation invariance for scalar features across 30 nearest spatial neighbors.
    2. Change of Basis to Local Reference Frames: For each residue ii, a local orthonormal reference frame (ui,vi,wi)(u_i, v_i, w_i) is constructed from its NN, CαC_\alpha, and CC coordinates. The rotation-equivariant vector feature outputs from the GVP-GNN encoder are projected into this local frame via a change of basis, yielding rotation-invariant local vector features.
    3. Feature Concatenation: The rotated vector features are flattened and concatenated with the GVP-GNN scalar features to construct a translation- and rotation-invariant feature matrix in Rn×512\mathbb{R}^{n \times 512}.
    4. Generic Transformer Encoder-Decoder: The invariant features are processed by an 8-layer Transformer encoder and decoded autoregressively by an 8-layer Transformer decoder with learned positional embeddings, pre-layer normalization, and 8 attention heads:

    p(Y∣X)=∏i=1np(yi∣yi−1,…,y1;X)p(Y \mid X) = \prod_{i=1}^n p(y_i \mid y_{i-1}, \dots, y_1; X)

  2. Knowl 2 — AlphaFold2 Synthetic Data Generation and Curation Pipeline for Inverse Folding

    model/method

    To overcome the data bottleneck of experimental protein structures (~16,000 non-redundant CATH chains), 12 million synthetic backbone structures were generated from the UniRef50 sequence database:

    1. Sequence Prioritization: Distograms for multiple sequence alignments (MSAs) across all UniRef50 sequences were predicted using MSA Transformer. Sequences of length ≤500\le 500 amino acids were ranked by predicted distogram lDDT scores as a quality proxy, selecting the top 12 million.
    2. Structure Folding: Single-model AlphaFold2 (CASP14 Model 1 weights) was run with 3 recycling iterations on UniRef100 MSAs generated via HHblits, followed by Amber force-field energy relaxation.
    3. Test Set Leakage Prevention: Validation and test domains from the CATH v4.3 80/10/10 topology split were mapped to Gene3D profile HMMs. UniRef50 sequences and UniRef100 sequences used for MSA construction that matched any validation or test Gene3D HMM via hmmsearch were strictly filtered out.
    4. Noise Injection & Confidence Masking: Any predicted residue coordinate with an AlphaFold2 predicted local distance difference test (pLDDT) confidence score below 90 (~25% of coordinates) was masked out. The per-residue pLDDT score was provided as an input feature encoded via Gaussian radial basis functions. To prevent models from learning unphysical sub-Angstrom coordinate artifacts specific to AlphaFold2, zero-mean Gaussian coordinate noise with standard deviation σ=0.1 A˚\sigma = 0.1\,\text{\AA} was added to predicted backbone coordinates during training.
    5. Data Mixing: In each training epoch, the ~16,000 experimental structures were mixed with a 10% random sample of the 12 million predicted structures, establishing an experimental:predicted training ratio of 1:80.
  3. Knowl 3 — Backbone Coordinate Span Masking Scheme for Partial Structure Infilling

    model/method

    To train inverse folding models to design sequences for partially specified backbones, functional motifs, or loops without complete coordinates, a coordinate span masking protocol is used during training:

    • Continuous residue segments of length up to 30 amino acids are randomly chosen until 15% of the total input backbone coordinates are masked.
    • Span lengths are sampled from a geometric distribution Geo(p)\text{Geo}(p) with parameter p=0.05p = 0.05 (yielding an expected average span length of 1/p=201/p = 20), with starting positions drawn uniformly at random.
    • In the GVP-GNN layers of the GVP-Transformer, masked residues are omitted as nodes in the geometric message-passing graph. When passing representations to the generic Transformer encoder, geometric scalar and vector features for masked positions are imputed as zero vectors, preserving global translation and rotation invariance while supplying a binary mask indicator token.
  4. Knowl 4 — Fixed-Backbone Sequence Recovery Benchmark on CATH Topology Splits

    data/table

    Models are evaluated on the CATH v4.3 topology-split test set (1,797 structurally held-out chains) across three length subsets: Short (≤100 \le 100 residues), Single-chain, and All chains. Evaluation metrics are per-residue perplexity (lower is better) and sequence recovery accuracy (percentage of positions matching native sequence under low-temperature sampling T=10−6T = 10^{-6}):

    Perplexity Recovery %
    Model Training Data Short Single-chain All Short Single-chain All
    Natural frequencies - 18.12 18.03 17.97 9.6% 9.0% 9.5%
    Structured GNN CATH 7.91 6.48 6.49 31.5% 37.1% 37.1%
    GVP-GNN (1M) CATH 7.14 5.36 5.43 34.0% 42.7% 42.2%
    GVP-GNN (1M) CATH + AF2 8.55 6.17 6.06 29.5% 38.2% 38.6%
    GVP-GNN-large (21M) CATH 7.68 6.12 6.17 32.6% 39.4% 39.2%
    GVP-GNN-large (21M) CATH + AF2 6.11 4.09 4.08 38.3% 50.8% 50.8%
    GVP-Transformer (142M) CATH 8.18 6.33 6.44 31.3% 38.5% 38.3%
    GVP-Transformer (142M) CATH + AF2 6.05 4.00 4.01 38.1% 51.5% 51.6%

    The 142M-parameter GVP-Transformer trained on CATH combined with 12M AlphaFold2 predicted structures improves native sequence recovery by 9.4 percentage points (51.6% vs. 42.2%) and reduces perplexity from 5.43 to 4.01 compared to the previous state-of-the-art GVP-GNN trained only on CATH.

  5. Knowl 5 — Residue Recovery Disparity Across Core and Surface Environments

    empirical result

    Inverse folding sequence recovery and perplexity depend heavily on local solvent accessibility and packing density:

    • Core Residues: Defined as residues having ≥24\ge 24 neighboring CαC_\alpha atoms within a 10 A˚10\,\text{\AA} sphere, core positions achieve 72%72\% native sequence recovery under GVP-Transformer (+AF2).
    • Surface Residues: Defined as residues having <16< 16 neighboring CαC_\alpha atoms within 10 A˚10\,\text{\AA}, surface positions achieve 39%39\% native sequence recovery.
    • Solvent Accessibility Correlation: Perplexity increases monotonically with solvent-accessible surface area (SASA).
    • Hydrophobic Placement: Despite the degeneracy of surface mutations, sampled sequences preserve wildtype hydrophobic spatial distributions, correctly burying hydrophobic residues (defined as I, V, L, F, C, M, A) in low-SASA environments while placing polar residues on the solvent-exposed surface.
  6. Knowl 6 — Multi-Conformation Sequence Design via Geometric Average Likelihood

    model/method

    To design amino acid sequences compatible with multiple distinct conformational states AA and BB (such as open vs. closed enzyme or receptor states), the dual-state conditional distribution p(Y∣A,B)p(Y \mid A, B) is approximated by the geometric mean of the single-state conditional likelihoods:

    p(Y∣A,B)=1Zp(Y∣A)⋅p(Y∣B)p(Y \mid A, B) = \frac{1}{Z} \sqrt{p(Y \mid A) \cdot p(Y \mid B)}

    Equivalently, the model scores sequence tokens autoregressively using the arithmetic mean of log probabilities 12(log⁡p(yi∣y<i;A)+log⁡p(yi∣y<i;B))\frac{1}{2}\left(\log p(y_i \mid y_{<i}; A) + \log p(y_i \mid y_{<i}; B)\right).

    When evaluated on 87 structurally held-out proteins from the PDBFlex dataset that adopt conformations separated by ≥5 A˚\ge 5\,\text{\AA} overall RMSD, dual-state conditioning systematically lowers sequence perplexity at locally flexible residues (local RMSD >1 A˚> 1\,\text{\AA}) relative to conditioning on either individual conformation. On the SARS-CoV-2 spike receptor-binding domain (RBD), dual-state conditioning across open (PDB: 6XRA) and closed (PDB: 6VXX) states yields 4.06 perplexity and 53.6% recovery, outperforming open-state alone (4.50 perplexity, 49.2% recovery) and closed-state alone (4.96 perplexity, 48.1% recovery).

  7. Knowl 7 — Generalization to Protein Complexes and Zero-Shot Binding Affinity Prediction

    empirical result

    Despite training strictly on single-chain structures, inverse folding models generalize to multi-chain protein complexes represented by concatenating individual chains separated by 10 mask coordinate tokens:

    • Complex Perplexity: On 796 test complexes, conditioning on the entire complex structure lowers sequence perplexity on constituent chains from 6.32 (chain alone) to 3.81 (full complex) for GVP-Transformer (+AF2), and from 7.80 to 5.37 for GVP-GNN.
    • Zero-Shot SARS-CoV-2 RBD-ACE2 Binding Affinity: Zero-shot mutational effect predictions on the receptor-binding motif (RBM, 69 residues in direct contact with human ACE2; PDB: 6M0J) were evaluated against deep mutational scanning affinity data from Starr et al. (2020) across four structural contexts (Spearman rank correlation ρ\rho):
    Model No coords (Seq only) No RBM coords No ACE2 coords All coords
    ESM-1v 0.03 - - -
    ESM-1b 0.02 - - -
    ESM-MSA-1b (few-shot) 0.51 - - -
    GVP-GNN - -0.10 0.50 0.60
    GVP-GNN-large + AF2 - -0.05 0.52 0.69
    GVP-Transformer + AF2 - -0.06 0.53 0.64

    Conditioning on both target and partner backbones produces a Spearman correlation of 0.69, which drops when partner coordinates are excluded (0.52) or when RBM coordinates are masked (-0.05).

  8. Knowl 8 — Zero-Shot Mutation and Insertion Effect Scoring via Log-Likelihood Ratios

    model/method

    Autoregressive inverse folding models predict the functional effect of sequence variations on an experimental wildtype backbone XX by evaluating the log-likelihood difference between mutant sequence YmutY_{\text{mut}} and wildtype sequence YwtY_{\text{wt}}:

    Δlog⁡p=log⁡p(Ymut∣X)−log⁡p(Ywt∣X)=∑i=1nlog⁡p(yi,mut∣y<i,mut;X)p(yi,wt∣y<i,wt;X)\Delta \log p = \log p(Y_{\text{mut}} \mid X) - \log p(Y_{\text{wt}} \mid X) = \sum_{i=1}^n \log \frac{p(y_{i,\text{mut}} \mid y_{<i,\text{mut}}; X)}{p(y_{i,\text{wt}} \mid y_{<i,\text{wt}}; X)}

    • De Novo Mini-Protein Stability: On single-point mutational scans across 10 de novo mini-protein folds (Rocklin et al.), GVP-Transformer (+AF2) achieves an average Pearson correlation of r=0.48r = 0.48 with experimental stability (improving over CATH-only GVP-GNN, r=0.42r = 0.42, on 8 of 10 folds).
    • Complex Interface Stability: On the SKEMPI single-point mutation benchmark (Atom3D binary classification of ΔΔG>0\Delta\Delta G > 0), zero-shot GVP-GNN on complex backbones achieves an AUROC of 0.71, matching supervised transfer learning models.
    • AAV Capsid Sequence Insertions: To score insertion mutations within the 28-amino acid variable region of adeno-associated virus (AAV) capsids (Bryant et al.), masked coordinate tokens are inserted into the wildtype backbone (PDB: 1LP3) at the insertion positions. On variant subsets with ≥8\ge 8 mutations, GVP-Transformer (+AF2) achieves a Spearman correlation of ρ=0.55\rho = 0.55 with DNA packaging fitness, compared to ρ=0.20\rho = 0.20 for sequence-only masked language model ESM-1v.
  9. Knowl 9 — Scaling Dynamics: Model Capacity Thresholds for Predicted Structure Utilization

    empirical result

    Effective utilization of synthetic AlphaFold2 training structures requires sufficient model capacity:

    • Small Model Degradation: Augmenting the 1M-parameter GVP-GNN architecture with 12M AlphaFold2 predicted structures degrades test sequence recovery on native backbones from 42.2% to 38.6% (perplexity worsens from 5.43 to 6.06).
    • Large Model Scaling: Increasing model capacity to 21M parameters (GVP-GNN-large) and 142M parameters (GVP-Transformer) enables positive transfer from synthetic structures, increasing sequence recovery from 39.2% to 50.8% and from 38.3% to 51.6%, respectively.
    • Data Volume Returns: Model performance improves steeply as synthetic data scales up to 1 million predicted structures (~75 times the CATH experimental dataset size), with diminishing marginal improvements observed between 1M and 12M structures.
    • Mixing Ratio: Increasing the training mixing ratio of predicted to experimental structures up to 80:1 prevents larger models from overfitting the small experimental training set.
  10. Knowl 10 — Strict Structural Hold-Out Benchmark Filtered by Pairwise TM-Score

    data/table

    Because 54% of chains in the standard CATH v4.3 topology split test set share a TM-score ≥0.5\ge 0.5 with training structures, a stringent structural hold-out test set of 223 protein chains was constructed using Foldseek pairwise TMalign, enforcing TM-score <0.5< 0.5 against all training chains:

    Perplexity Recovery %
    Model Training Data Short Single-chain All Short Single-chain All
    Structured GNN CATH 10.08 7.04 7.06 27.8% 35.1% 35.4%
    GVP-GNN (1M) CATH 8.13 5.76 5.86 31.5% 41.1% 40.4%
    GVP-GNN (1M) CATH + AF2 9.87 6.61 6.50 26.3% 36.3% 36.8%
    GVP-GNN-large (21M) CATH 8.87 6.62 6.68 31.0% 37.2% 37.4%
    GVP-GNN-large (21M) CATH + AF2 7.08 4.46 4.39 34.1% 48.2% 48.7%
    GVP-Transformer (142M) CATH 8.80 6.78 6.97 28.5% 36.7% 36.3%
    GVP-Transformer (142M) CATH + AF2 6.99 4.36 4.34 33.0% 48.9% 49.5%

    Performance trends on the TM-filtered test set mirror the standard topology split: GVP-Transformer (+AF2) achieves 49.5% sequence recovery (perplexity 4.34), outperforming CATH-only models by over 9 percentage points and demonstrating true generalization to structurally distinct folds.

Coverage note — Omitted hardware profiling wall-clock sampling speeds (Table C.6) and qualitative visual comparisons of BLOSUM62 confusion matrix logs (Figure C.4) as they are standard auxiliary implementation verifications rather than core methodology or primary empirical findings.

References

  1. 1.Alford, R. F., Leaver-Fay, A., Jeliazkov, J. R., O’Meara, M. J., DiMaio, F. P., Park, H., Shapovalov, M. V., Renfrew, P. D., Mulligan, V. K., Kappel, K., et al. The rosetta all-atom energy function for macromolecular modeling and design. Journal of chemical theory and computation, 13(6):3031–3048, 2017.
  2. 2.Alley, E. C., Khimulya, G., Biswas, S., AlQuraishi, M., and Church, G. M. Unified rational protein engineering with sequence-based deep representation learning. Nature methods, 16(12):1315–1322, 2019.
  3. 3.Anand, N. and Huang, P. Generative modeling for protein structures. Advances in neural information processing systems, 31, 2018.
  4. 4.Anand-Achim, N., Eguchi, R. R., Mathews, I. I., Perez, C. P., Derry, A., Altman, R. B., and Huang, P.-S. Protein sequence design with a learned potential. Biorxiv, pp. 2020–01, 2021.
  5. 5.Angermueller, C., Dohan, D., Belanger, D., Deshpande, R., Murphy, K., and Colwell, L. Model-based reinforcement learning for biological sequence design. In International conference on learning representations, 2019.
  6. 6.Anishchenko, I., Pellock, S. J., Chidyausiku, T. M., Ramelot, T. A., Ovchinnikov, S., Hao, J., Bafna, K., Norn, C., Kang, A., Bera, A. K., et al. De novo protein design by deep network hallucination. Nature, 600(7889):547–552, 2021.
  7. 7.Baek, M., DiMaio, F., Anishchenko, I., Dauparas, J., Ovchinnikov, S., Lee, G. R., Wang, J., Cong, Q., Kinch, L. N., Schaeffer, R. D., Millán, C., Park, H., Adams, C., Glassman, C. R., DeGiovanni, A., Pereira, J. H., Rodrigues, A. V., van Dijk, A. A., Ebrecht, A. C., Opperman, D. J., Sagmeister, T., Buhlheller, C., Pavkov-Keller, T., Rathinaswamy, M. K., Dalwadi, U., Yip, C. K., Burke, J. E., Garcia, K. C., Grishin, N. V., Adams, P. D., Read, R. J., and Baker, D. Accurate prediction of protein structures and interactions using a three-track neural network. Science, 373(6557):871–876, 2021. doi: 10.1126/science.abj8754.
  8. 8.Bepler, T. and Berger, B. Learning protein sequence embeddings using information from structure. arXiv preprint arXiv:1902.08661, 2019.
  9. 9.Berman, H. M., Westbrook, J., Feng, Z., Gilliland, G., Bhat, T. N., Weissig, H., Shindyalov, I. N., and Bourne, P. E. The protein data bank. Nucleic acids research, 28(1): 235–242, 2000.
  10. 10.Boomsma, W. and Frellsen, J. Spherical convolutions and their application in molecular modelling. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper/2017/file/1113d7a76ffceca1bb350bfe145467c6-Paper.pdf.
  11. 11.Brookes, D., Park, H., and Listgarten, J. Conditioning by adaptive sampling for robust design. In International conference on machine learning, pp. 773–782. PMLR, 2019.
  12. 12.Bryant, D. H., Bashir, A., Sinai, S., Jain, N. K., Ogden, P. J., Riley, P. F., Church, G. M., Colwell, L. J., and Kelsic, E. D. Deep diversification of an aav capsid protein by machine learning. Nature Biotechnology, 39(6):691–696, 2021.
  13. 13.Chen, S., Sun, Z., Lin, L., Liu, Z., Liu, X., Chong, Y., Lu, Y., Zhao, H., and Yang, Y. To improve protein sequence profile prediction through image captioning on pairwise residue distance map. Journal of chemical information and modeling, 60(1):391–399, 2019.
  14. 14.Dahiyat, B. I. and Mayo, S. L. Probing the role of packing specificity in protein design. Proceedings of the National Academy of Sciences, 94(19):10172–10177, 1997.
  15. 15.Dallago, C., Mou, J., Johnston, K. E., Wittmann, B. J., Bhattacharya, N., Goldman, S., Madani, A., and Yang, K. K. Flip: Benchmark tasks in fitness landscape inference for proteins. bioRxiv, 2021.
  16. 16.DeGrado, W. F., Raleigh, D. P., and Handel, T. De novo protein design: what are we learning? Current Opinion in Structural Biology, 1(6):984–993, 1991.
  17. 17.Edunov, S., Ott, M., Auli, M., and Grangier, D. Understanding back-translation at scale. arXiv preprint arXiv:1808.09381, 2018.
  18. 18.Eguchi, R. R., Anand, N., Choe, C. A., and Huang, P.-S. Ig-vae: generative modeling of immunoglobulin proteins by direct 3d coordinate generation. bioRxiv, 2020.
  19. 19.Elnaggar, A., Heinzinger, M., Dallago, C., Rehawi, G., Yu, W., Jones, L., Gibbs, T., Feher, T., Angerer, C., Steinegger, M., Bhowmik, D., and Rost, B. Prottrans: Towards cracking the language of lifes code through selfsupervised deep learning and high performance computing. IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–1, 2021. doi: 10.1109/TPAMI.2021.3095381.
  20. 20.Gligorijevic, V., Berenberg, D., Ra, S., Watkins, A., Kelow, S., Cho, K., and Bonneau, R. Function-guided protein design by deep manifold sampling. bioRxiv, 2021.
  21. 21.Harbury, P. B., Plecs, J. J., Tidor, B., Alber, T., and Kim, P. S. High-resolution protein design with backbone freedom. Science, 282(5393):1462–1467, 1998.
  22. 22.Heinzinger, M., Elnaggar, A., Wang, Y., Dallago, C., Nechaev, D., Matthes, F., and Rost, B. Modeling aspects of the language of life through transfer-learning protein sequences. BMC bioinformatics, 20(1):1–17, 2019.
  23. 23.Hornak, V., Abel, R., Okur, A., Strockbine, B., Roitberg, A., and Simmerling, C. Comparison of multiple amber force fields and development of improved protein backbone parameters. Proteins: Structure, Function, and Bioinformatics, 65(3):712–725, 2006.
  24. 24.Hrabe, T., Li, Z., Sedova, M., Rotkiewicz, P., Jaroszewski, L., and Godzik, A. Pdbflex: exploring flexibility in protein structures. Nucleic acids research, 44(D1):D423–D428, 2016.
  25. 25.Huang, P.-S., Boyken, S. E., and Baker, D. The coming of age of de novo protein design. Nature, 537(7620):320–327, 2016.
  26. 26.Ingraham, J., Garg, V. K., Barzilay, R., and Jaakkola, T. S. Generative models for graph-based protein design. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 15794–15805, 2019.
  27. 27.Jankauskaitė, J., Jiménez-García, B., Dapkūnas, J., Fernández-Recio, J., and Moal, I. H. Skempi 2.0: an updated benchmark of changes in protein–protein binding energy, kinetics and thermodynamics upon mutation. Bioinformatics, 35(3):462–469, 2019.
  28. 28.Jin, W., Wohlwend, J., Barzilay, R., and Jaakkola, T. Iterative refinement graph neural network for antibody sequence-structure co-design. arXiv preprint arXiv:2110.04624, 2021.
  29. 29.Jing, B., Eismann, S., Soni, P. N., and Dror, R. O. Equivariant graph neural networks for 3d macromolecular structure. Proceedings of the International Conference on Machine Learning, 2021a.
  30. 30.Jing, B., Eismann, S., Suriana, P., Townshend, R. J. L., and Dror, R. O. Learning from protein structure with geometric vector perceptrons. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021b.
  31. 31.Joshi, M., Chen, D., Liu, Y., Weld, D. S., Zettlemoyer, L., and Levy, O. Spanbert: Improving pre-training by representing and predicting spans. Transactions of the Association for Computational Linguistics, 8:64–77, 2020.
  32. 32.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  33. 33.Kim, S., van Kempen, M., Söding, J., and Steinegger, M. foldseek. https://github.com/steineggerlab/foldseek, 2021.
  34. 34.Kunzmann, P. and Hamacher, K. Biotite: a unifying open source computational biology framework in python. BMC bioinformatics, 19(1):1–8, 2018.
  35. 35.Lan, J., Ge, J., Yu, J., Shan, S., Zhou, H., Fan, S., Zhang, Q., Shi, X., Wang, Q., Zhang, L., et al. Structure of the sars-cov-2 spike receptor-binding domain bound to the ace2 receptor. Nature, 581(7807):215–220, 2020.
  36. 36.Langan, R. A., Boyken, S. E., Ng, A. H., Samson, J. A., Dods, G., Westbrook, A. M., Nguyen, T. H., Lajoie, M. J., Chen, Z., Berger, S., et al. De novo design of bioactive protein switches. Nature, 572(7768):205–210, 2019.
  37. 37.Lees, J., Yeats, C., Perkins, J., Sillitoe, I., Rentzsch, R., Dessailly, B. H., and Orengo, C. Gene3d: a domain-based resource for comparative genomics, functional annotation and protein network analysis. Nucleic acids research, 40(D1):D465–D471, 2012.
  38. 38.Li, B., Yang, Y. T., Capra, J. A., and Gerstein, M. B. Predicting changes in protein thermodynamic stability upon point mutation with deep 3d convolutional neural networks. PLoS computational biology, 16(11):e1008291, 2020.
  39. 39.Li, Z., Yang, Y., Faraggi, E., Zhan, J., and Zhou, Y. Direct prediction of profiles of sequences compatible with a protein structure by neural networks with fragment-based local and energy-based nonlocal profiles. Proteins: Structure, Function, and Bioinformatics, 82(10):2565–2573, 2014.
  40. 40.Madani, A., McCann, B., Naik, N., Keskar, N. S., Anand, N., Eguchi, R. R., Huang, P.-S., and Socher, R. Progen: Language modeling for protein generation. arXiv preprint arXiv:2004.03497, 2020.
  41. 41.Madani, A., Krause, B., Greene, E. R., Subramanian, S., Mohr, B. P., Holton, J. M., Olmos, J. L., Xiong, C., Sun, Z. Z., Socher, R., et al. Deep neural language modeling enables functional protein generation across families. bioRxiv, 2021.
  42. 42.Meier, J., Rao, R., Verkuil, R., Liu, J., Sercu, T., and Rives, A. Language models enable zero-shot prediction of the effects of mutations on protein function. Advances in Neural Information Processing Systems, 34, 2021.
  43. 43.Mirdita, M., von den Driesch, L., Galiez, C., Martin, M. J., Söding, J., and Steinegger, M. Uniclust databases of clustered and deeply annotated protein sequences and alignments. Nucleic acids research, 45(D1):D170–D176, 2017.
  44. 44.Norn, C., Wicky, B. I., Juergens, D., Liu, S., Kim, D., Tischer, D., Koepnick, B., Anishchenko, I., Baker, D., and Ovchinnikov, S. Protein sequence design by conformational landscape optimization. Proceedings of the National Academy of Sciences, 118(11), 2021.
  45. 45.O’Connell, J., Li, Z., Hanson, J., Heffernan, R., Lyons, J., Paliwal, K., Dehzangi, A., Yang, Y., and Zhou, Y. Spin2: Predicting sequence profiles from protein structures using deep neural networks. Proteins: Structure, Function, and Bioinformatics, 86(6):629–633, 2018.
  46. 46.Orengo, C. A., Michie, A. D., Jones, S., Jones, D. T., Swindells, M. B., and Thornton, J. M. Cath–a hierarchic classification of protein domain structures. Structure, 5(8):1093–1109, 1997.
  47. 47.Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M. fairseq: A fast, extensible toolkit for sequence modeling. arXiv preprint arXiv:1904.01038, 2019.
  48. 48.Potter, S. C., Luciani, A., Eddy, S. R., Park, Y., Lopez, R., and Finn, R. D. Hmmer web server: 2018 update. Nucleic acids research, 46(W1):W200–W204, 2018.
  49. 49.Qi, Y. and Zhang, J. Z. Densecpd: improving the accuracy of neural-network-based computational protein sequence design with densenet. Journal of Chemical Information and Modeling, 60(3):1245–1252, 2020.
  50. 50.Quijano-Rubio, A., Yeh, H.-W., Park, J., Lee, H., Langan, R. A., Boyken, S. E., Lajoie, M. J., Cao, L., Chow, C. M., Miranda, M. C., et al. De novo design of modular and tunable protein biosensors. Nature, 591(7850):482–487, 2021.
  51. 51.Rao, R., Bhattacharya, N., Thomas, N., Duan, Y., Chen, P., Canny, J., Abbeel, P., and Song, Y. Evaluating protein transfer learning with tape. Advances in neural information processing systems, 32, 2019.
  52. 52.Rao, R., Liu, J., Verkuil, R., Meier, J., Canny, J. F., Abbeel, P., Sercu, T., and Rives, A. Msa transformer. bioRxiv, 2021.
  53. 53.Rives, A., Meier, J., Sercu, T., Goyal, S., Lin, Z., Liu, J., Guo, D., Ott, M., Zitnick, C. L., Ma, J., et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences, 118(15), 2021.
  54. 54.Rocklin, G. J., Chidyausiku, T. M., Goreshnik, I., Ford, A., Houliston, S., Lemak, A., Carter, L., Ravichandran, R., Mulligan, V. K., Chevalier, A., et al. Global analysis of protein folding using massively parallel design, synthesis, and testing. Science, 357(6347):168–175, 2017.
  55. 55.Senior, A. W., Evans, R., Jumper, J., Kirkpatrick, J., Sifre, L., Green, T., Qin, C., Žídek, A., Nelson, A. W., Bridgland, A., et al. Improved protein structure prediction using potentials from deep learning. Nature, 577(7792):706–710, 2020.
  56. 56.Sennrich, R., Haddow, B., and Birch, A. Improving neural machine translation models with monolingual data. arXiv preprint arXiv:1511.06709, 2015.
  57. 57.Shin, J.-E., Riesselman, A. J., Kollasch, A. W., McMahon, C., Simon, E., Sander, C., Manglik, A., Kruse, A. C., and Marks, D. S. Protein design and variant prediction using autoregressive generative models. Nature communications, 12(1):1–11, 2021.
  58. 58.Shroff, R., Cole, A. W., Diaz, D. J., Morrow, B. R., Donnell, I., Annapareddy, A., Gollihar, J., Ellington, A. D., and Thyer, R. Discovery of novel gain-of-function mutations guided by structure-based deep learning. ACS synthetic biology, 9(11):2927–2935, 2020.
  59. 59.Sinai, S., Wang, R., Whatley, A., Slocum, S., Locane, E., and Kelsic, E. D. Adalead: A simple and robust adaptive greedy search algorithm for sequence design. arXiv preprint arXiv:2010.02141, 2020.
  60. 60.Starr, T. N., Greaney, A. J., Hilton, S. K., Ellis, D., Crawford, K. H., Dingens, A. S., Navarro, M. J., Bowen, J. E., Tortorici, M. A., Walls, A. C., et al. Deep mutational scanning of sars-cov-2 receptor binding domain reveals constraints on folding and ace2 binding. Cell, 182(5):1295–1310, 2020.
  61. 61.Steinegger, M., Meier, M., Mirdita, M., Vöhringer, H., Haunsberger, S. J., and Söding, J. Hh-suite3 for fast remote homology detection and deep protein annotation. BMC bioinformatics, 20(1):1–15, 2019.
  62. 62.Street, A. G. and Mayo, S. L. Computational protein design. Structure, 7(5):R105–R109, 1999.
  63. 63.Strokach, A., Becerra, D., Corbi-Verge, C., Perez-Riba, A., and Kim, P. M. Fast and flexible protein design using deep graph neural networks. Cell Systems, 11(4):402–411, 2020.
  64. 64.Suzek, B. E., Wang, Y., Huang, H., McGarvey, P. B., Wu, C. H., and Consortium, U. Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics, 31(6):926–932, 2015.
  65. 65.Townshend, R. J. L., Vögele, M., Suriana, P., Derry, A., Powers, A., Laloudakis, Y., Balachandar, S., Anderson, B. M., Eismann, S., Kondor, R., Altman, R. B., and Dror, R. O. ATOM3D: tasks on molecules in three dimensions. CoRR, abs/2012.04035, 2020.
  66. 66.Trinquier, J., Uguzzoni, G., Pagnani, A., Zamponi, F., and Weigt, M. Efficient generative modeling of protein sequences using simple autoregressive models. arXiv preprint arXiv:2103.03292, 2021.
  67. 67.Turc, I., Chang, M.-W., Lee, K., and Toutanova, K. Well-read students learn better: On the importance of pre-training compact models. arXiv preprint arXiv:1908.08962, 2019.
  68. 68.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.
  69. 69.Walls, A. C., Park, Y.-J., Tortorici, M. A., Wall, A., McGuire, A. T., and Veesler, D. Structure, function, and antigenicity of the sars-cov-2 spike glycoprotein. Cell, 181(2):281–292, 2020.
  70. 70.Wang, J., Cao, H., Zhang, J. Z., and Qi, Y. Computational protein design with deep learning neural networks. Scientific reports, 8(1):1–9, 2018.
  71. 71.Wang, J., Lisanza, S., Juergens, D., Tischer, D., Anishchenko, I., Baek, M., Watson, J. L., Chun, J. H., Milles, L. F., Dauparas, J., et al. Deep learning methods for designing proteins scaffolding functional sites. bioRxiv, 2021.
  72. 72.Wu, Z., Johnston, K. E., Arnold, F. H., and Yang, K. K. Protein sequence design with deep generative models. Current Opinion in Chemical Biology, 65:18–27, 2021.
  73. 73.Yang, J., Anishchenko, I., Park, H., Peng, Z., Ovchinnikov, S., and Baker, D. Improved protein structure prediction using predicted interresidue orientations. Proceedings of the National Academy of Sciences, 117(3):1496–1503, 2020.
  74. 74.Yang, K. K., Wu, Z., and Arnold, F. H. Machine-learning-guided directed evolution for protein engineering. Nature methods, 16(8):687–694, 2019.
  75. 75.Zhang, Y., Chen, Y., Wang, C., Lo, C.-C., Liu, X., Wu, W., and Zhang, J. Prodconn: Protein design using a convolutional neural network. Proteins: Structure, Function, and Bioinformatics, 88(7):819–829, 2020.
  76. 76.Zhou, J., Panaitiu, A. E., and Grigoryan, G. A general-purpose protein design framework based on mining sequence–structure relationships in known protein structures. Proceedings of the National Academy of Sciences, 117(2):1059–1068, 2020.

Citation

MLA
Hsu, C., et al. “Learning Inverse Folding from Millions of Predicted Structures”. International Conference on Machine Learning, vol. 162, 2022, pp. 8946–70, https://proceedings.mlr.press/v162/hsu22a.html.
APA
Hsu, C., Verkuil, R., Liu, J., Lin, Z., Hie, B., Sercu, T., Lerer, A., & Rives, A. (2022). Learning inverse folding from millions of predicted structures. International Conference on Machine Learning, 162, 8946–8970. https://proceedings.mlr.press/v162/hsu22a.html
Chicago
Hsu, C., R. Verkuil, J. Liu, et al. 2022. “Learning Inverse Folding from Millions of Predicted Structures”. International Conference on Machine Learning 162: 8946–70. https://proceedings.mlr.press/v162/hsu22a.html.
Harvard
Hsu, C. et al. (2022) “Learning inverse folding from millions of predicted structures”, International Conference on Machine Learning. PMLR, pp. 8946–8970. Available at: https://proceedings.mlr.press/v162/hsu22a.html.
Vancouver
1. Hsu C, Verkuil R, Liu J, Lin Z, Hie B, Sercu T, Lerer A, Rives A (2022) Learning inverse folding from millions of predicted structures. In: International Conference on Machine Learning. PMLR, pp 8946–8970

BibTeX

@InProceedings{pmlr-v162-hsu22a,
  title = 	 {Learning inverse folding from millions of predicted structures},
  author =       {Hsu, Chloe and Verkuil, Robert and Liu, Jason and Lin, Zeming and Hie, Brian and Sercu, Tom and Lerer, Adam and Rives, Alexander},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {8946--8970},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/hsu22a/hsu22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/hsu22a.html},
  abstract = 	 {We consider the problem of predicting a protein sequence from its backbone atom coordinates. Machine learning approaches to this problem to date have been limited by the number of available experimentally determined protein structures. We augment training data by nearly three orders of magnitude by predicting structures for 12M protein sequences using AlphaFold2. Trained with this additional data, a sequence-to-sequence transformer with invariant geometric input processing layers achieves 51% native sequence recovery on structurally held-out backbones with 72% recovery for buried residues, an overall improvement of almost 10 percentage points over existing methods. The model generalizes to a variety of more complex tasks including design of protein complexes, partially masked structures, binding interfaces, and multiple states.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/