Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval

Pascal NotinMafalda DiasJonathan FrazerJavier Marchena-HurtadoAidan N. GomezDebora S. MarksYarin Gal

article2022ICML268 citations

Introduces an autoregressive transformer architecture that combines multi-scale attention with inference-time homology retrieval to score insertions, deletions, and complex substitutions across diverse protein families without requiring multiple sequence alignments during training.

Listen

Predicting the functional impact of genetic mutations on protein fitness is essential for drug discovery, clinical diagnosis of disease variants, and biotherapeutic design. Traditional state-of-the-art computational methods rely heavily on training deep generative models on protein-specific multiple sequence alignments. However, this dependence creates significant bottlenecks: alignments cannot readily handle complex variations such as insertions and deletions, fail entirely on hard-to-align regions like disordered proteins, and perform poorly when protein families lack deep evolutionary records. Meanwhile, existing language models trained on massive unaligned databases avoid alignment constraints but generally underperform specialized alignment models when predicting mutation effects.

The main objective of the article is to demonstrate and evaluate Tranception, a novel autoregressive transformer architecture that combines large-scale pretraining on non-aligned protein sequences with lightweight evolutionary retrieval at inference time. The authors also establish ProteinGym, a large-scale, diverse benchmark designed to rigorously evaluate protein fitness prediction models across diverse taxa, mutation depths, and variation types.

To develop Tranception, the authors trained a 700-million-parameter autoregressive model on roughly 250 million unaligned sequences from the UniRef100 database, incorporating specialized attention mechanisms designed to capture contiguous subsequences at varying lengths. At inference time, the model scores protein fitness by merging its sequence log-likelihood predictions with positional frequency statistics retrieved from homologous evolutionary sequences. To evaluate predictive capability, the authors established ProteinGym, compiling 87 amino-acid substitution assays covering approximately 1.5 million variants and 7 insertion-deletion assays covering around 300,000 variants across human, viral, prokaryotic, and other eukaryotic targets.

The evaluation produced several critical findings. First, Tranception with inference-time retrieval achieved top overall predictive accuracy on the ProteinGym substitution benchmark, attaining an average Spearman rank correlation of 0.451 and surpassing leading specialized alignment models like EVE and existing protein language models like ESM-1v. Second, Tranception exhibited its most substantial performance advantages on proteins with shallow alignments, where conventional alignment-based methods degraded sharply. Third, the model demonstrated superior extrapolation capabilities on complex multiple-mutant sequences, achieving a correlation of 0.499 on variants with five or more simultaneous mutations compared to 0.420 for EVE. Fourth, Tranception effectively scored sequence insertions and deletions out-of-the-box, reaching a correlation of 0.463 and outperforming the only competing baseline capable of evaluating indels. Finally, model ensembling revealed that pairing Tranception with alignment-based models boosted correlation to 0.473, demonstrating high complementarity between autoregressive sequence modeling and evolutionary alignments.

These findings indicate that protein fitness prediction can be reliably scaled to the entire proteome without sacrificing accuracy or requiring time-consuming, protein-specific retraining. By decoupling model training from alignment availability, Tranception substantially reduces computational overhead, improves coverage across difficult-to-model disordered regions, and mitigates the risk of prediction failure on understudied or emerging pathogens. Furthermore, its demonstrated strength on multiple-substitution variants offers direct performance benefits for machine-learning-guided protein engineering and sequence design pipelines.

For practical implementation, organizations should leverage hybrid scoring frameworks by deploying autoregressive models augmented with lightweight inference retrieval. When maximum predictive accuracy is required for well-characterized proteins, practitioners should consider ensembling Tranception with specialized models like EVE. For generative sequence design or evaluating disordered regions and insertion-deletions, Tranception serves as an effective standalone tool. Future efforts should focus on expanding the architecture's capacity through scaling model size, integrating more diverse metagenomic sequence datasets, and generating additional multi-mutant experimental assays to refine design workflows.

Confidence in these findings is bolstered by the extensive scale and diversity of the ProteinGym benchmark, which consistently verified Tranception across multiple metrics. A primary limitation remains the benchmark's historical bias toward single-substitution mutations relative to multi-mutant and insertion-deletion profiles, as well as the model context window of 1,024 amino acids, which requires sequence slicing for the small fraction of proteins exceeding that length.

Cover for Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval

Abstract

The ability to accurately model the fitness landscape of protein sequences is critical to a wide range of applications, from quantifying the effects of human variants on disease likelihood, to predicting immune-escape mutations in viruses and designing novel biotherapeutic proteins. Deep generative models of protein sequences trained on multiple sequence alignments have been the most successful approaches so far to address these tasks. The performance of these methods is however contingent on the availability of sufficiently deep and diverse alignments for reliable training. Their potential scope is thus limited by the fact many protein families are hard, if not impossible, to align. Large language models trained on massive quantities of non-aligned protein sequences from diverse families address these problems and show potential to eventually bridge the performance gap. We introduce Tranception, a novel transformer architecture leveraging autoregressive predictions and retrieval of homologous sequences at inference to achieve state-of-the-art fitness prediction performance. Given its markedly higher performance on multiple mutants, robustness to shallow alignments and ability to score indels, our approach offers significant gain of scope over existing approaches. To enable more rigorous model testing across a broader range of protein families, we develop ProteinGym – an extensive set of multiplexed assays of variant effects, substantially increasing both the number and diversity of assays compared to existing benchmarks.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Mutation effect prediction with aligned sequences
  • 2.2. Modeling proteins without alignments
  • 2.3. Deep Mutational Scanning benchmarks
  • 2.4. Retrieval
  • 3. Tranception
  • 3.1. Tranception attention
  • 3.2. Grouped ALiBi position encoding
  • 3.3. Data processing and augmentations
  • 3.4. Scoring sequences for fitness prediction
  • 4. Inference-time retrieval
  • 4.1. Multiple sequence alignments
  • 4.2. Two modes of inference
  • 5. ProteinGym
  • 6. Results
  • 6.1. Baselines
  • 6.2. ProteinGym substitution benchmark
  • 6.3. ProteinGym indel benchmark
  • 7. Discussion
  • 8. Conclusion
  • Acknowledgements
  • References
  • Appendix
  • A. Glossary
  • B. Tranception model architecture and training details
  • B.1. Ablation studies
  • B.2. Data processing
  • B.3. Model training
  • B.4. Scoring protein sequences
  • B.5. Retrieval
  • C. Multiple Sequence Alignments
  • D. Baselines
  • D.1. ESM-1v and MSA Transformer
  • D.2. EVE and DeepSequence
  • E. Detailed performance analysis
  • E.1. Performance reporting methodology
  • E.2. Detailed results
  • E.3. MSA Filtering analyses
  • E.4. Full protein Vs domain-specific models
  • E.5. Model ensembling
  • F. ProteinGym curation

Knowls

  1. Knowl 1 — Tranception Attention Mechanism

    model/method

    Tranception attention is a multi-head self-attention mechanism designed for autoregressive protein language models that explicitly extracts contiguous subsequence motifs (kk-mers) at multiple spatial resolutions.

    At each transformer layer, the attention heads are partitioned into four distinct groups:

    1. 1-mer group: Standard self-attention projections with no spatial convolution applied after the Query (QQ), Key (KK), and Value (VV) linear projections.
    2. 3-mer group: Separate 1D spatial depthwise convolutions with a kernel size of 3 applied to the QQ, KK, and VV projections prior to attention score computation.
    3. 5-mer group: Separate 1D spatial depthwise convolutions with a kernel size of 5 applied to QQ, KK, and VV.
    4. 7-mer group: Separate 1D spatial depthwise convolutions with a kernel size of 7 applied to QQ, KK, and VV.

    To preserve strict causality in the autoregressive architecture, causal left-padding is applied before each convolution. The model uses squared ReLU non-linear activations in the feed-forward layers. This multi-scale grouping incentivizes different attention heads to specialize in capturing evolutionary sequence patterns at differing kk-mer widths.

  2. Knowl 2 — Inference-Time Retrieval-Augmented Fitness Scoring

    equation

    The fitness effect FxF_x of a mutated protein sequence xmutx^{\text{mut}} relative to a wild-type reference sequence xwtx^{\text{wt}} is scored by the log-likelihood ratio:

    Fx=log⁡P(xmut)−log⁡P(xwt)F_x = \log P(x^{\text{mut}}) - \log P(x^{\text{wt}})

    To compute the probability of a sequence x=(x1,x2,…,xl)x = (x_1, x_2, \dots, x_l) of length ll, Tranception combines an autoregressive generative model distribution PAP_A with an empirical position-specific distribution PRP_R derived from an alignment of homologous sequences (Multiple Sequence Alignment, MSA) retrieved at inference time:

    log⁡P(x)=∑i=1l[(1−α)log⁡PA(xi∣x<i)+αlog⁡PR(xi)]\log P(x) = \sum_{i=1}^{l} \left[ (1 - \alpha) \log P_A(x_i \mid x_{<i}) + \alpha \log P_R(x_i) \right]

    where:

    • α∈[0,1]\alpha \in [0, 1] is the retrieval weighting parameter (optimized to α=0.6\alpha = 0.6 on validation assays).
    • PA(xi∣x<i)P_A(x_i \mid x_{<i}) is the autoregressive conditional token probability output by the transformer.
    • PR(xi)P_R(x_i) is the empirical amino acid distribution at aligned coordinate position ii computed across sequences in the retrieved MSA, weighted by sequence similarity to downweight over-represented clades, with Laplace pseudocount smoothing (10−510^{-5}) and gap tokens excluded.

    For inserted positions or positions not covered in the retrieved MSA, the retrieval component αlog⁡PR(xi)\alpha \log P_R(x_i) is omitted, and the score at position ii defaults purely to the autoregressive term log⁡PA(xi∣x<i)\log P_A(x_i \mid x_{<i}).

  3. Knowl 3 — Grouped ALiBi Position Encoding

    model/method

    Grouped Attention with Linear Biases (Grouped ALiBi) is a positional bias scheme that removes learned or sinusoidal input position embeddings from transformer token inputs and directly modifies the query-key attention computation.

    For token indices ii and jj, an attention head computes the dot product between query qiq_i and key kjk_j and applies a static, non-learned linear penalty proportional to their distance ∣i−j∣|i - j|:

    Attention Logitij=qikjT−m⋅∣i−j∣\text{Attention Logit}_{ij} = q_i k_j^T - m \cdot |i - j|

    where mm is a head-specific slope scalar. In Grouped ALiBi, the geometric progression of slopes mm is assigned independently within each of the four convolution-kernel attention groups. This allows each kernel group to learn distance-dependent contextual associations tailored to its specific kk-mer resolution while providing inductive bias for length extrapolation beyond the pretraining context window.

  4. Knowl 4 — Bidirectional Sequence Mirror Scoring

    model/method

    While autoregressive protein transformers generate tokens in a single causal direction (typically NN-to-CC terminus), functional protein constraints operate bidirectionally. Tranception implements bidirectional scoring via sequence mirroring:

    1. Pretraining Augmentation: During pretraining, sequences within each batch are reversed at random with probability p=0.5p = 0.5, teaching the network parameters to evaluate sequence likelihood in both N→CN \to C (canonical) and C→NC \to N (reversed) directions.
    2. Inference Aggregation: The fitness prediction FxF_x for a variant is evaluated by taking the arithmetic mean of the log-likelihood ratios obtained by scoring the sequence in the canonical left-to-right order and the reversed right-to-left order:

    Fx=12(Fxcanonical+Fxreversed)F_x = \frac{1}{2} \left( F_x^{\text{canonical}} + F_x^{\text{reversed}} \right)

    1. Context Windowing: For proteins exceeding the 1024-token context length, a contiguous sub-window is extracted that centers around the barycenter of the mutated positions to maximize available contextual sequence on both flanks.
  5. Knowl 5 — ProteinGym Benchmark Suite for Variant Effect Prediction

    definition

    ProteinGym is an evaluation benchmark suite for zero-shot protein fitness prediction, consisting of 94 Deep Mutational Scanning (DMS) assays across 77 publications divided into two tracks:

    1. Substitution Benchmark: 87 DMS assays encompassing approximately 1.5 million missense variants (0.36 million single substitutions and 1.26 million multiple substitutions) spanning human proteins (33 assays), non-human eukaryotes (14 assays), prokaryotes (24 assays), and viruses (22 assays).
    2. Indel Benchmark: 7 DMS assays covering approximately 270,000 to 300,000 insertion and deletion (indel) variants across diverse functional properties.

    Raw DMS measurements are standardized by setting higher experimental values to denote higher fitness, averaging duplicates, and dropping missing or silent mutations. When multiple distinct assays evaluate the same underlying protein (such as different phenotypes for p53 or TEM-1 β\beta-lactamase), metrics are averaged at the UniProt ID level with uniform weighting per protein. Evaluation metrics include Spearman's rank correlation ρ\rho, Area Under the Receiver Operating Characteristic (AUC), and Matthews Correlation Coefficient (MCC).

  6. Knowl 6 — Zero-Shot Substitution Fitness Prediction Across Multiple Sequence Alignment Depths

    data/table

    Tranception (700M parameter model with retrieval weight α=0.6\alpha = 0.6) was evaluated on the 87 substitution assays of ProteinGym across varying Multiple Sequence Alignment (MSA) depth categories. Alignment depth is quantified by Neff/LN_{\text{eff}}/L, where NeffN_{\text{eff}} is the effective sequence count in the MSA and LL is the sequence length: Low (Neff/L<1N_{\text{eff}}/L < 1), Medium (1≤Neff/L≤1001 \le N_{\text{eff}}/L \le 100), and High (Neff/L>100N_{\text{eff}}/L > 100).

    Model Spearman's rank correlation (ρ\rho) ↑\uparrow AUC ↑\uparrow
    Low Depth Medium Depth High Depth All Assays All Assays
    Site independent 0.428 0.403 0.350 0.397 0.725
    Wavenet 0.319 0.398 0.469 0.398 0.725
    DeepSequence 0.375 0.397 0.506 0.415 0.733
    EVmutation 0.401 0.421 0.468 0.427 0.738
    EVE 0.408 0.440 0.507 0.448 0.751
    ESM-1v 0.321 0.348 0.484 0.371 0.713
    MSA Transformer 0.373 0.418 0.482 0.422 0.737
    Tranception (w/o retrieval) 0.394 0.398 0.439 0.406 0.728
    Tranception (w/ retrieval) 0.453 0.438 0.488 0.451 0.754

    Tranception with retrieval achieves the highest overall Spearman correlation (0.451) and AUC (0.754). Its performance advantage over alignment-only models (such as EVE) and unaligned models (such as ESM-1v) is most pronounced in the low-depth regime (ho=0.453 ho = 0.453).

  7. Knowl 7 — Extrapolation to Multiple Mutants Across Mutation Depths

    data/table

    Model performance was evaluated across different mutation depths (the number of amino acid substitutions in a variant relative to the wild-type sequence) on ProteinGym substitution assays.

    Model Spearman's rank correlation (ρ\rho) by mutation depth ↑\uparrow
    1 2 3 4 5+ All
    Site independent 0.396 0.325 0.286 0.319 0.421 0.397
    Wavenet 0.394 0.344 0.329 0.281 0.396 0.398
    DeepSequence 0.415 0.394 0.372 0.304 0.418 0.415
    EVmutation 0.427 0.392 0.379 0.319 0.433 0.427
    EVE 0.448 0.392 0.375 0.334 0.420 0.448
    ESM-1v 0.372 0.291 0.190 0.160 0.245 0.371
    MSA Transformer 0.423 0.359 0.390 0.327 0.431 0.422
    Tranception (w/o retrieval) 0.397 0.412 0.425 0.335 0.479 0.406
    Tranception (w/ retrieval) 0.448 0.435 0.443 0.368 0.499 0.451

    Masked language models scoring multiple mutants via additive masked marginals (e.g., ESM-1v) degrade significantly as mutation depth increases (0.3720.372 at depth 1 down to 0.1600.160 at depth 4). In contrast, Tranception autoregressively models epistatic co-dependencies, sustaining high correlations across higher mutation orders (0.4990.499 at depth 5+).

  8. Knowl 8 — Fitness Prediction on Protein Insertions and Deletions

    data/table

    Predicting the fitness effects of insertion and deletion (indel) mutations requires variable-length sequence scoring. Standard alignment-based models (EVE, DeepSequence, EVmutation) cannot score indels due to fixed coordinate columns, and masked language models (ESM-1v, MSA Transformer) require the target position to exist in the wild-type reference. Tranception dynamically adjusts the alignment columns at inference and scores indels natively.

    Assay Wavenet Tranception (w/o retrieval) Tranception (w/ retrieval)
    A0A1J4YT16 9PROT 0.117 0.178 0.191
    B1LPA6 ECOSM 0.385 0.321 0.415
    BLAT ECOLX 0.546 0.296 0.357
    PTEN HUMAN 0.699 0.563 0.598
    CAPSD AAV2S 0.457 0.549 0.586
    HIS7 YEAST 0.680 0.707 0.692
    P53 HUMAN 0.001 0.395 0.401
    Average Spearman (ρ\rho) 0.412 0.430 0.463
    Average AUC 0.724 0.740 0.759

    Tranception with retrieval outperforms Wavenet across the 7 indel assays in ProteinGym, reaching an average Spearman ρ\rho of 0.463 and AUC of 0.759.

  9. Knowl 9 — Complementary Cross-Architecture Ensembling of Tranception and EVE

    data/table

    Ensembling distinct model families combines orthogonal biological signals: autoregressive representation learning across unaligned sequences (Tranception) and explicit co-evolutionary modeling from protein-specific sequence alignments (EVE).

    Model Configuration Average Spearman Correlation (ρ\rho)
    Tranception w/o retrieval (single model) 0.406
    ESM-1v (ensemble of 5 models) 0.401
    MSA Transformer (ensemble of 5 models) 0.434
    EVE (ensemble of 5 models) 0.452
    Tranception w/ retrieval (single model) 0.451
    Tranception w/o retrieval + ESM-1v 0.427
    Tranception w/o retrieval + MSA Transformer 0.449
    Tranception w/o retrieval + EVE 0.473
    Tranception w/ retrieval + EVE 0.475

    A cross-architecture ensemble pairing single-seed Tranception (without retrieval) with single-seed EVE yields ρ=0.473\rho = 0.473, outperforming a 5-seed ensemble of EVE models (0.4520.452) and a single Tranception model with retrieval (0.4510.451).

  10. Knowl 10 — Robustness of Tranception to Multiple Sequence Alignment Shallowing

    empirical result

    When Multiple Sequence Alignments (MSAs) are downsampled by removing sequences based on their minimum sequence identity to the query wild-type, alignment-based methods and MSA-dependent language models exhibit sharp drops in mutation effect prediction accuracy. EVE and MSA Transformer performance degrades markedly as alignment diversity decreases (Spearman correlation dropping below ρ=0.20\rho = 0.20 as the minimum similarity threshold increases toward 80%).

    In contrast, Tranception exhibits high stability under alignment shallowing. Because Tranception is pretrained on unaligned sequence databases and relies on retrieval only for aggregate per-position pseudocounts, its Spearman correlation degrades gracefully as sequences are removed, asymptotically approaching its zero-retrieval baseline performance (Spearman ρ=0.406\rho = 0.406).

  11. Knowl 11 — Impact of Pretraining Sequence Granularity on Autoregressive Protein Models

    empirical result

    Autoregressive transformer architectures for protein fitness prediction benefit from maximizing sequence granularity in the pretraining dataset rather than aggressive identity clustering.

    When pretraining the Tranception Small architecture (85M parameters, 150k steps) on different similarity-clustered splits of UniRef:

    • UniRef100 (unclustered dataset containing ~249M sequences): achieves Spearman ρ=0.335\rho = 0.335 on the full ProteinGym substitution set.
    • UniRef90 (clustered at 90% identity): achieves Spearman ρ=0.275\rho = 0.275.
    • UniRef50 (clustered at 50% identity): achieves Spearman ρ=0.247\rho = 0.247.

    Although UniRef50 exhibits lower redundancy and lower cross-entropy during pretraining, the greater diversity and exact sequence variants preserved in UniRef100 transfer to superior downstream zero-shot fitness prediction.

Coverage note — Domain-specific hyperparameter sweeps on the 10-assay validation subset and hardware specifications were omitted as secondary implementation details.

References

  1. 1.Aakre, C., Herrou, J., Phung, T., Perchuk, B., Crosson, S., and Laub, M. Evolving New Protein-Protein Interaction Specificity through Promiscuous Intermediates. Cell, 163(3):594–606, October 2015. ISSN 00928674. doi: 10.1016/j.cell.2015.09.055. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867415012726.
  2. 2.Adkar, B., Tripathi, A., Sahoo, A., Bajaj, K., Goswami, D., Chakrabarti, P., Swarnkar, M., Gokhale, R., and Varadarajan, R. Protein Model Discrimination Using Mutational Sensitivity Derived from Deep Sequencing. Structure, 20(2):371–381, February 2012. ISSN 09692126. doi: 10.1016/j.str.2011.11.021. URL https://linkinghub.elsevier.com/retrieve/pii/S0969212612000068.
  3. 3.Ahler, E., Register, A. C., Chakraborty, S., Fang, L., Dieter, E. M., Sitko, K. A., Vidadala, R. S. R., Trevillian, B. M., Golkowski, M., Gelman, H., Stephany, J. J., Rubin, A. F., Merritt, E. A., Fowler, D. M., and Maly, D. J. A Combined Approach Reveals a Regulatory Mechanism Coupling Src’s Kinase Activity, Localization, and Phosphotransferase-Independent Functions. Molecular Cell, 74(2):393–408.e20, April 2019. ISSN 10972765. doi: 10.1016/j.molcel.2019.02.003. URL https://linkinghub.elsevier.com/retrieve/pii/S1097276519300930.
  4. 4.Alley, E. C., Khimulya, G., Biswas, S., AlQuraishi, M., and Church, G. M. Unified rational protein engineering with sequence-only deep representation learning. bioRxiv, pp. 589333, 2019.
  5. 5.Amorosi, C. J., Chiasson, M. A., McDonald, M. G., Wong, L. H., Sitko, K. A., Boyle, G., Kowalski, J. P., Rettie, A. E., Fowler, D. M., and Dunham, M. J. Massively parallel characterization of CYP2C9 variant enzyme activity and abundance. The American Journal of Human Genetics, 108(9):1735–1751, September 2021. ISSN 00029297. doi: 10.1016/j.ajhg.2021.07.001. URL https://linkinghub.elsevier.com/retrieve/pii/S000292972100269X.
  6. 6.Araya, C. L., Fowler, D. M., Chen, W., Muniez, I., Kelly, J. W., and Fields, S. A fundamental protein property, thermodynamic stability, revealed solely from large-scale measurements of protein function. Proceedings of the National Academy of Sciences, 109(42):16858–16863, October 2012. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1209751109. URL http://www.pnas.org/cgi/doi/10.1073/pnas.1209751109.
  7. 7.Baek, M., Dimaio, F., Anishchenko, I. V., Dauparas, J., Ovchinnikov, S., Lee, G. R., Wang, J., Cong, Q., Kinch, L. N., Schaeffer, R. D., Millan, C., Park, H., Adams, C., Glassman, C. R., DeGiovanni, A. M., Pereira, J. H., Rodrigues, A. V., van Dijk, A. A., Ebrecht, A. C., Opperman, D. J., Sagmeister, T., Buhlheller, C., Pavkov-Keller, T., Rathinaswamy, M. K., Dalwadi, U., Yip, C. K., Burke, J. E., Garcia, K. C., Grishin, N. V., Adams, P. D., Read, R. J., and Baker, D. Accurate prediction of protein structures and interactions using a 3-track neural network. Science (New York, N.Y.), 373:871 – 876, 2021.
  8. 8.Bandaru, P., Shah, N. H., Bhattacharyya, M., Barton, J. P., Kondo, Y., Cofsky, J. C., Gee, C. L., Chakraborty, A. K., Kortemme, T., Ranganathan, R., and Kuriyan, J. Deconstruction of the Ras switching cycle through saturation mutagenesis. eLife, 6:e27810, July 2017. ISSN 2050-084X. doi: 10.7554/eLife.27810. URL https://elifesciences.org/articles/27810.
  9. 9.Bepler, T. and Berger, B. Learning protein sequence embeddings using information from structure. arXiv preprint arXiv:1902.08661, 2019.
  10. 10.Biswas, S., Khimulya, G., Alley, E. C., Esvelt, K. M., and Church, G. M. Low-n protein engineering with data-efficient deep learning. Nature Methods, 18(4):389–396, 2021.
  11. 11.Bolognesi, B., Faure, A. J., Seuma, M., Schmiedel, J. M., Tartaglia, G. G., and Lehner, B. The mutational landscape of a prion-like domain. Nature communications, 10(1): 1–12, 2019.
  12. 12.Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., van den Driessche, G., Lespiau, J.-B., Damoc, B., Clark, A., de Las Casas, D., Guy, A., Menick, J., Ring, R., Hennigan, T. W., Huang, S., Maggiore, L., Jones, C., Cassirer, A., Brock, A., Paganini, M., Irving, G., Vinyals, O., Osindero, S., Simonyan, K., Rae, J. W., Elsen, E., and Sifre, L. Improving language models by retrieving from trillions of tokens. ArXiv, abs/2112.04426, 2021.
  13. 13.Boucher, J. I., Bolon, D. N., and Tawfik, D. S. Quantifying and understanding the fitness effects of protein mutations: Laboratory versus nature. Protein Science, 25(7):1219–1226, 2016.
  14. 14.Brenan, L., Andreev, A., Cohen, O., Pantel, S., Kamburov, A., Cacchiarelli, D., Persky, N., Zhu, C., Bagul, M., Goetz, E., Burgin, A., Garraway, L., Getz, G., Mikkelsen, T., Piccioni, F., Root, D., and Johannessen, C. Phenotypic Characterization of a Comprehensive Set of MAPK1 /ERK2 Missense Mutants. Cell Reports, 17(4):1171–1183, October 2016. ISSN 22111247. doi: 10.1016/j.celrep.2016.09.061. URL https://linkinghub.elsevier.com/retrieve/pii/S2211124716313171.
  15. 15.Bridgford, J. L., Lee, S. M., Lee, C. M. M., Guglielmelli, P., Rumi, E., Pietra, D., Wilcox, S., Chhabra, Y., Rubin, A. F., Cazzola, M., Vannucchi, A. M., Brooks, A. J., Call, M. E., and Call, M. J. Novel drivers and modifiers of MPL-dependent oncogenic transformation identified by deep mutational scanning. Blood, 135(4):287–292, January 2020. ISSN 0006-4971, 1528-0020. doi: 10.1182/blood.2019002561. URL https://ashpublications.org/blood/article/135/4/287/381157/Novel-drivers-and-modifiers-of-MPLdependent.
  16. 16.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners, 2020.
  17. 17.Chan, Y. H., Venev, S. V., Zeldovich, K. B., and Matthews, C. R. Correlation of fitness landscapes from three orthologous TIM barrels originates from sequence and structure constraints. Nature Communications, 8(1): 14614, April 2017. ISSN 2041-1723. doi: 10.1038/ncomms14614. URL http://www.nature.com/articles/ncomms14614.
  18. 18.Chen, J. Z., Fowler, D. M., and Tokuriki, N. Comprehensive exploration of the translocation, stability and substrate recognition requirements in VIM-2 lactamase. eLife, 9:e56707, June 2020. ISSN 2050-084X. doi: 10.7554/eLife.56707. URL https://elifesciences.org/articles/56707.
  19. 19.Chiasson, M. A., Rollins, N. J., Stephany, J. J., Sitko, K. A., Matreyek, K. A., Verby, M., Sun, S., Roth, F. P., DeSloover, D., Marks, D. S., Rettie, A. E., and Fowler, D. M. Multiplexed measurement of variant abundance and activity reveals VKOR topology, active site and human variant impact. eLife, 9: e58026, September 2020. ISSN 2050-084X. doi: 10.7554/eLife.58026. URL https://elifesciences.org/articles/58026.
  20. 20.Dallago, C., Mou, J., Johnston, K. E., Wittmann, B. J., Bhattacharya, N., Goldman, S., Madani, A., and Yang, K. K. Flip: Benchmark tasks in fitness landscape inference for proteins. 2021.
  21. 21.Dandage, R., Pandey, R., Jayaraj, G., Rai, M., Berger, D., and Chakraborty, K. Differential strengths of molecular determinants guide environment specific mutational fates. PLOS Genetics, 14(5):e1007419, May 2018. ISSN 1553-7404. doi: 10.1371/journal.pgen.1007419. URL https://dx.plos.org/10.1371/journal.pgen.1007419.
  22. 22.Davidi, D., Shamshoum, M., Guo, Z., Bar-On, Y. M., Prywes, N., Oz, A., Jablonska, J., Flamholz, A., Wernick, D. G., Antonovsky, N., et al. Highly active rubiscos discovered by systematic interrogation of natural sequence diversity. The EMBO journal, 39(18):e104081, 2020.
  23. 23.Deng, Z., Huang, W., Bakkalbasi, E., Brown, N. G., Adamski, C. J., Rice, K., Muzny, D., Gibbs, R. A., and Palzkill, T. Deep Sequencing of Systematic Combinatorial Libraries Reveals β-Lactamase Sequence Constraints at High Resolution. Journal of Molecular Biology, 424(3-4):150–167, December 2012. ISSN 00222836. doi: 10.1016/j.jmb.2012.09.014. URL https://linkinghub.elsevier.com/retrieve/pii/S0022283612007711.
  24. 24.Doud, M. and Bloom, J. Accurate Measurement of the Effects of All Amino-Acid Mutations on Influenza Hemagglutinin. Viruses, 8(6):155, June 2016. ISSN 1999-4915. doi: 10.3390/v8060155. URL http://www.mdpi.com/1999-4915/8/6/155.
  25. 25.Doud, M. B., Ashenberg, O., and Bloom, J. D. Site-Specific Amino Acid Preferences Are Mostly Conserved in Two Closely Related Protein Homologs. Molecular Biology and Evolution, 32(11):2944–2960, November 2015. ISSN 0737-4038, 1537-1719. doi: 10.1093/molbev/msv167. URL https://academic.oup.com/mbe/article-lookup/doi/10.1093/molbev/msv167.
  26. 26.Duenas-Decamp, M., Jiang, L., Bolon, D., and Clapham, P. R. Saturation mutagenesis of the hiv-1 envelope cd4 binding loop reveals residues controlling distinct trimer conformations. PLoS pathogens, 12(11):e1005988, 2016.
  27. 27.Eddy, S. R. Accelerated profile hmm searches. PLoS computational biology, 7(10):e1002195, 2011.
  28. 28.Edgar, R. C. Muscle: multiple sequence alignment with high accuracy and high throughput. Nucleic acids research, 32 5:1792–7, 2004.
  29. 29.Elnaggar, A., Heinzinger, M., Dallago, C., Rihawi, G., Wang, Y., Jones, L., Gibbs, T., Feher, T., Angerer, C., Steinegger, M., et al. Prottrans: towards cracking the language of life’s code through self-supervised deep learning and high performance computing. arXiv preprint arXiv:2007.06225, 2020.
  30. 30.Faure, A. J., Domingo, J., Schmiedel, J. M., Hidalgo-Carcedo, C., Diss, G., and Lehner, B. Mapping the energetic and allosteric landscapes of protein binding domains. Nature, 604(7904):175–183, 2022.
  31. 31.Fernandes, J. D., Faust, T. B., Strauli, N. B., Smith, C., Crosby, D. C., Nakamura, R. L., Hernandez, R. D., and Frankel, A. D. Functional Segregation of Overlapping Genes in HIV. Cell, 167(7):1762–1773.e12, December 2016. ISSN 00928674. doi: 10.1016/j.cell.2016.11.031. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867416316038.
  32. 32.Findlay, G. M., Daza, R. M., Martin, B., Zhang, M. D., Leith, A. P., Gasperini, M., Janizek, J. D., Huang, X., Starita, L. M., and Shendure, J. Accurate classification of BRCA1 variants with saturation genome editing. Nature, 562(7726):217–222, October 2018. ISSN 0028-0836, 1476-4687. doi: 10.1038/s41586-018-0461-z. URL http://www.nature.com/articles/s41586-018-0461-z.
  33. 33.Firnberg, E., Labonte, J. W., Gray, J. J., and Ostermeier, M. A Comprehensive, High-Resolution Map of a Gene’s Fitness Landscape. Molecular Biology and Evolution, 31 (6):1581–1592, June 2014. ISSN 1537-1719, 0737-4038. doi: 10.1093/molbev/msu081. URL https://academic.oup.com/mbe/article-lookup/doi/10.1093/molbev/msu081.
  34. 34.Flynn, J. M., Rossouw, A., Cote-Hammarlof, P., Fragata, I., Mavor, D., Hollins, C., Bank, C., and Bolon, D. N. Comprehensive fitness maps of Hsp90 show widespread environmental dependence. eLife, 9:e53810, March 2020. ISSN 2050-084X. doi: 10.7554/eLife.53810. URL https://elifesciences.org/articles/53810.
  35. 35.Fowler, D. M. and Fields, S. Deep mutational scanning: a new style of protein science. Nature methods, 11(8): 801–807, 2014.
  36. 36.Frazer, J., Notin, P., Dias, M., Gomez, A., Min, J. K., Brock, K. P., Gal, Y., and Marks, D. S. Disease variant prediction with deep generative models of evolutionary data. Nature, 2021.
  37. 37.Giacomelli, A. O., Yang, X., Lintner, R. E., McFarland, J. M., Duby, M., Kim, J., Howard, T. P., Takeda, D. Y., Ly, S. H., Kim, E., Gannon, H. S., Hurhula, B., Sharpe, T., Goodale, A., Fritchman, B., Steelman, S., Vazquez, F., Tsherniak, A., Aguirre, A. J., Doench, J. G., Piccioni, F., Roberts, C. W. M., Meyerson, M., Getz, G., Johannessen, C. M., Root, D. E., and Hahn, W. C. Mutational processes shape the landscape of TP53 mutations in human cancer. Nature Genetics, 50(10):1381–1387, October 2018. ISSN 1061-4036, 1546-1718. doi: 10.1038/s41588-018-0204-y. URL http://www.nature.com/articles/s41588-018-0204-y.
  38. 38.Glazer, A. M., Kroncke, B. M., Matreyek, K. A., Yang, T., Wada, Y., Shields, T., Salem, J.-E., Fowler, D. M., and Roden, D. M. Deep Mutational Scan of an SCN5A Voltage Sensor. Circulation: Genomic and Precision Medicine, 13(1), February 2020. ISSN 2574-8300. doi: 10.1161/CIRCGEN.119.002786. URL https://www.ahajournals.org/doi/10.1161/CIRCGEN.119.002786.
  39. 39.Gonzalez, C. E., Roberts, P., and Ostermeier, M. Fitness effects of single amino acid insertions and deletions in tem-1 β-lactamase. Journal of molecular biology, 431 (12):2320–2330, 2019.
  40. 40.Grave, E., Cisse, M., and Joulin, A. Unbounded cache model for online language modeling with open vocabulary. In NIPS, 2017.
  41. 41.Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M.-W. Realm: Retrieval-augmented language model pre-training. ArXiv, abs/2002.08909, 2020.
  42. 42.Haddox, H. K., Dingens, A. S., and Bloom, J. D. Experimental Estimation of the Effects of All Amino-Acid Mutations to HIV’s Envelope Protein on Viral Replication in Cell Culture. PLOS Pathogens, 12(12):e1006114, December 2016. ISSN 1553-7374. doi: 10.1371/journal.ppat.1006114. URL https://dx.plos.org/10.1371/journal.ppat.1006114.
  43. 43.Haddox, H. K., Dingens, A. S., Hilton, S. K., Overbaugh, J., and Bloom, J. D. Mapping mutational effects along the evolutionary landscape of HIV envelope. eLife, 7:e34420, March 2018. ISSN 2050-084X. doi: 10.7554/eLife.34420. URL https://elifesciences.org/articles/34420.
  44. 44.Heinzinger, M., Elnaggar, A., Wang, Y., Dallago, C., Nechaev, D., Matthes, F., and Rost, B. Modeling the language of life–deep learning protein sequences. bioRxiv, pp. 614313, 2019.
  45. 45.Hesslow, D., ed. Zanichelli, N., Notin, P., Poli, I., and Marks, D. S. Rita: a study on scaling up generative protein sequence models. ArXiv, abs/2205.05789, 2022.
  46. 46.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural Computation, 9:1735–1780, 1997.
  47. 47.Hopf, T. A., Scharfe, C. P., Rodrigues, J. P., Green, A. G., Kohlbacher, O., Sander, C., Bonvin, A. M., and Marks, D. S. Sequence co-evolution gives 3d contacts and structures of protein complexes. Elife, 3:e03430, 2014.
  48. 48.Hopf, T. A., Ingraham, J. B., Poelwijk, F. J., Scharfe, C. P., Springer, M., Sander, C., and Marks, D. S. Mutation effects predicted from sequence co-variation. Nature biotechnology, 35(2):128–135, 2017.
  49. 49.Jacquier, H., Birgy, A., Le Nagard, H., Mechulam, Y., Schmitt, E., Glodt, J., Bercot, B., Petit, E., Poulain, J., Barnaud, G., Gros, P.-A., and Tenaillon, O. Capturing the mutational landscape of the beta-lactamase TEM-1. Proceedings of the National Academy of Sciences, 110(32):13067–13072, August 2013. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1215206110. URL http://www.pnas.org/cgi/doi/10.1073/pnas.1215206110.
  50. 50.Jia, X., Burugula, B. B., Chen, V., Lemons, R. M., Jayakody, S., Maksutova, M., and Kitzman, J. O. Massively parallel functional testing of MSH2 missense variants conferring Lynch syndrome risk. The American Journal of Human Genetics, 108(1):163–175, January 2021. ISSN 00029297. doi: 10.1016/j.ajhg.2020.12.003. URL https://linkinghub.elsevier.com/retrieve/pii/S0002929720304390.
  51. 51.Jiang, L., Liu, P., Bank, C., Renzette, N., Prachanronarong, K., Yilmaz, L. S., Caffrey, D. R., Zeldovich, K. B., Schiffer, C. A., Kowalik, T. F., et al. A balance between inhibitor binding and substrate processing confers influenza drug resistance. Journal of molecular biology, 428(3): 538–553, 2016.
  52. 52.Jones, E. M., Lubock, N. B., Venkatakrishnan, A., Wang, J., Tseng, A. M., Paggi, J. M., Latorraca, N. R., Cancilla, D., Satyadi, M., Davis, J. E., Babu, M. M., Dror, R. O., and Kosuri, S. Structural and functional characterization of G protein–coupled receptors with deep mutational scanning. eLife, 9:e54895, October 2020. ISSN 2050-084X. doi: 10.7554/eLife.54895. URL https://elifesciences.org/articles/54895.
  53. 53.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Zˇ´ıdek, A., Potapenko, A., et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  54. 54.Jurafsky, D. and Martin, J. H. Speech and language processing, 2nd edition. 2008.
  55. 55.Kaplan, J., McCandlish, S., Henighan, T. J., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. ArXiv, abs/2001.08361, 2020.
  56. 56.Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L. Y., Edunov, S., Chen, D., and tau Yih, W. Dense passage retrieval for open-domain question answering. ArXiv, abs/2004.04906, 2020.
  57. 57.Kelsic, E. D., Chung, H., Cohen, N., Park, J., Wang, H. H., and Kishony, R. RNA Structural Determinants of Optimal Codons Revealed by MAGE-Seq. Cell Systems, 3(6):563–571.e6, December 2016. ISSN 24054712. doi: 10.1016/j.cels.2016.11.004. URL https://linkinghub.elsevier.com/retrieve/pii/S2405471216303684.
  58. 58.Kennouche, P., Charles-Orszag, A., Nishiguchi, D., Goussard, S., Imhaus, A., Dupre, M., Chamot-Rooke, J., and Dumenil, G. Deep mutational scanning of the Neisseria meningitidis major pilin reveals the importance of pilus tip-mediated adhesion. The EMBO Journal, 38(22), November 2019. ISSN 0261-4189, 1460-2075. doi: 10.15252/embj.2019102145. URL https://onlinelibrary.wiley.com/doi/10.15252/embj.2019102145.
  59. 59.Kessel, A. and Ben-Tal, N. Introduction to Proteins: Structure, Function, and Motion, SECOND EDITION (Chapman & Hall/CRC Mathematical and Computational Biology). 03 2018. ISBN 9781498747172. doi: 10.1201/9781315113876.
  60. 60.Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M. Generalization through memorization: Nearest neighbor language models. ArXiv, abs/1911.00172, 2020.
  61. 61.Kitzman, J. O., Starita, L. M., Lo, R. S., Fields, S., and Shendure, J. Massively parallel single-amino-acid mutagenesis. Nature Methods, 12(3):203–206, March 2015. ISSN 1548-7091, 1548-7105. doi: 10.1038/nmeth.3223. URL http://www.nature.com/articles/nmeth.3223.
  62. 62.Klesmith, J. R., Bacik, J.-P., Michalczyk, R., and Whitehead, T. A. Comprehensive Sequence-Flux Mapping of a Levoglucosan Utilization Pathway in E. coli. ACS Synthetic Biology, 4(11):1235–1243, November 2015. ISSN 2161-5063, 2161-5063. doi: 10.1021/acssynbio.5b00131. URL https://pubs.acs.org/doi/10.1021/acssynbio.5b00131.
  63. 63.Kotler, E., Shani, O., Goldfeld, G., Lotan-Pompan, M., Tarcic, O., Gershoni, A., Hopf, T. A., Marks, D. S., Oren, M., and Segal, E. A Systematic p53 Mutation Library Links Differential Functional Impact to Cancer Mutation Pattern and Evolutionary Conservation. Molecular Cell, 71(1):178–190.e8, July 2018. ISSN 10972765. doi: 10.1016/j.molcel.2018.06.012. URL https://linkinghub.elsevier.com/retrieve/pii/S1097276518304544.
  64. 64.Kozek, K. A., Glazer, A. M., Ng, C.-A., Blackwell, D., Egly, C. L., Vanags, L. R., Blair, M., Mitchell, D., Matreyek, K. A., Fowler, D. M., Knollmann, B. C., Vandenberg, J. I., Roden, D. M., and Kroncke, B. M. High-throughput discovery of trafficking-deficient variants in the cardiac potassium channel KV11.1. Heart Rhythm, 17(12):2180–2189, December 2020. ISSN 15475271. doi: 10.1016/j.hrthm.2020.05.041. URL https://linkinghub.elsevier.com/retrieve/pii/S1547527120305427.
  65. 65.Lee, J. M., Huddleston, J., Doud, M. B., Hooper, K. A., Wu, N. C., Bedford, T., and Bloom, J. D. Deep mutational scanning of hemagglutinin helps predict evolutionary fates of human H3N2 influenza variants. Proceedings of the National Academy of Sciences, 115(35):E8276–E8285, August 2018. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1806133115. URL http://www.pnas.org/lookup/doi/10.1073/pnas.1806133115.
  66. 66.Lee, K., Chang, M.-W., and Toutanova, K. Latent retrieval for weakly supervised open domain question answering. ArXiv, abs/1906.00300, 2019.
  67. 67.Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Kuttler, H., Lewis, M., tau Yih, W., Rocktaschel, T., Riedel, S., and Kiela, D. Retrieval-augmented generation for knowledge-intensive nlp tasks. ArXiv, abs/2005.11401, 2020.
  68. 68.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In ICLR, 2019.
  69. 69.Madani, A., McCann, B., Naik, N., Keskar, N. S., Anand, N., Eguchi, R. R., Huang, P.-S., and Socher, R. Progen: Language modeling for protein generation, 2020.
  70. 70.Manekar, S. C. and Sathe, S. R. A benchmark study of k-mer counting methods for high-throughput sequencing. GigaScience, 7(12):giy125, 2018.
  71. 71.Matreyek, K. A., Starita, L. M., Stephany, J. J., Martin, B., Chiasson, M. A., Gray, V. E., Kircher, M., Khechaduri, A., Dines, J. N., Hause, R. J., Bhatia, S., Evans, W. E., Relling, M. V., Yang, W., Shendure, J., and Fowler, D. M. Multiplex assessment of protein variant abundance by massively parallel sequencing. Nature Genetics, 50(6):874–882, June 2018. ISSN 1061-4036, 1546-1718. doi: 10.1038/s41588-018-0122-z. URL http://www.nature.com/articles/s41588-018-0122-z.
  72. 72.Matreyek, K. A., Stephany, J. J., Ahler, E., and Fowler, D. M. Integrating thousands of pten variant activity and abundance measurements reveals variant subgroups and new dominant negatives in cancers. Genome medicine, 13(1):1–17, 2021.
  73. 73.Mattenberger, F., Latorre, V., Tirosh, O., Stern, A., and Geller, R. Globally defining the effects of mutations in a picornavirus capsid. eLife, 10:e64256, January 2021. ISSN 2050-084X. doi: 10.7554/eLife.64256. URL https://elifesciences.org/articles/64256.
  74. 74.Mavor, D., Barlow, K., Thompson, S., Barad, B. A., Bonny, A. R., Cario, C. L., Gaskins, G., Liu, Z., Deming, L., Axen, S. D., Caceres, E., Chen, W., Cuesta, A., Gate, R. E., Green, E. M., Hulce, K. R., Ji, W., Kenner, L. R., Mensa, B., Morinishi, L. S., Moss, S. M., Mravic, M., Muir, R. K., Niekamp, S., Nnadi, C. I., Palovcak, E., Poss, E. M., Ross, T. D., Salcedo, E. C., See, S. K., Subramaniam, M., Wong, A. W., Li, J., Thorn, K. S., Conchuir, S. O., Roscoe, B. P., Chow, E. D., DeRisi, J. L., Kortemme, T., Bolon, D. N., and Fraser, J. S. Determination of ubiquitin fitness landscapes under different chemical stresses in a classroom setting. eLife, 5:e15802, April 2016. ISSN 2050-084X. doi: 10.7554/eLife.15802. URL https://elifesciences.org/articles/15802.
  75. 75.McLaughlin Jr, R. N., Poelwijk, F. J., Raman, A., Gosal, W. S., and Ranganathan, R. The spatial architecture of protein function and adaptation. Nature, 491(7422):138–142, November 2012. ISSN 0028-0836, 1476-4687. doi: 10.1038/nature11500. URL http://www.nature.com/articles/nature11500.
  76. 76.Meier, J., Rao, R., Verkuil, R., Liu, J., Sercu, T., and Rives, A. Language models enable zero-shot prediction of the effects of mutations on protein function. bioRxiv, 2021. doi: 10.1101/2021.07.09.450648. URL https://www.biorxiv.org/content/early/2021/07/10/2021.07.09.450648.
  77. 77.Melamed, D., Young, D. L., Gamble, C. E., Miller, C. R., and Fields, S. Deep mutational scanning of an RRM domain of the Saccharomyces cerevisiae poly(A)-binding protein. RNA, 19(11):1537–1551, November 2013. ISSN 1355-8382, 1469-9001. doi: 10.1261/rna.040709.113. URL http://rnajournal.cshlp.org/lookup/doi/10.1261/rna.040709.113.
  78. 78.Melnikov, A., Rogov, P., Wang, L., Gnirke, A., and Mikkelsen, T. S. Comprehensive mutational scanning of a kinase in vivo reveals substrate-dependent fitness landscapes. Nucleic Acids Research, 42(14):e112–e112, August 2014. ISSN 0305-1048, 1362-4962. doi: 10.1093/nar/gku511. URL https://academic.oup.com/nar/article-lookup/doi/10.1093/nar/gku511.
  79. 79.Mighell, T. L., Evans-Dutson, S., and O’Roak, B. J. A Saturation Mutagenesis Approach to Understanding PTEN Lipid Phosphatase Activity and Genotype-Phenotype Relationships. The American Journal of Human Genetics, 102(5):943–955, May 2018. ISSN 00029297. doi: 10.1016/j.ajhg.2018.03.018. URL https://linkinghub.elsevier.com/retrieve/pii/S0002929718301071.
  80. 80.Mishra, P., Flynn, J., Starr, T., and Bolon, D. Systematic Mutant Analyses Elucidate General and Client-Specific Aspects of Hsp90 Function. Cell Reports, 15(3):588–598, April 2016. ISSN 22111247. doi: 10.1016/j.celrep.2016.03.046. URL https://linkinghub.elsevier.com/retrieve/pii/S2211124716303175.
  81. 81.Mitchell, A. L., Almeida, A., Beracochea, M., Boland, M. A., Burgin, J., Cochrane, G., Crusoe, M. R., Kale, V., Potter, S. C., Richardson, L. J., Sakharova, E. A., Scheremetjew, M., Korobeynikov, A. I., Shlemov, A., Kunyavskaya, O., Lapidus, A. L., and Finn, R. D. Mgnify: the microbiome analysis resource in 2020. Nucleic Acids Research, 48:D570 – D578, 2020.
  82. 82.Nambiar, A., Heflin, M., Liu, S., Maslov, S., Hopkins, M., and Ritz, A. Transforming the language of life: transformer neural networks for protein prediction tasks. In Proceedings of the 11th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics, pp. 1–8, 2020.
  83. 83.Newberry, R. W., Arhar, T., Costello, J., Hartoularos, G. C., Maxwell, A. M., Naing, Z. Z. C., Pittman, M., Reddy, N. R., Schwarz, D. M. C., Wassarman, D. R., Wu, T. S., Barrero, D., Caggiano, C., Catching, A., Cavazos, T. B., Estes, L. S., Faust, B., Fink, E. A., Goldman, M. A., Gomez, Y. K., Gordon, M. G., Gunsalus, L. M., Hoppe, N., Jaime-Garza, M., Johnson, M. C., Jones, M. G., Kung, A. F., Lopez, K. E., Lumpe, J., Martyn, C., McCarthy, E. E., Miller-Vedam, L. E., Navarro, E. J., Palar, A., Pellegrino, J., Saylor, W., Stephens, C. A., Strickland, J., Torosyan, H., Wankowicz, S. A., Wong, D. R., Wong, G., Redding, S., Chow, E. D., DeGrado, W. F., and Kampmann, M. Robust Sequence Determinants of α-Synuclein Toxicity in Yeast Implicate Membrane Binding. ACS Chemical Biology, 15(8):2137–2153, August 2020. ISSN 1554-8929, 1554-8937. doi: 10.1021/acschembio.0c00339. URL https://pubs.acs.org/doi/10.1021/acschembio.0c00339.
  84. 84.Ng, P. C. and Henikoff, S. Predicting deleterious amino acid substitutions. Genome research, 11(5):863–874, 2001.
  85. 85.Nutschel, C., Fulton, A., Zimmermann, O., Schwaneberg, U., Jaeger, K.-E., and Gohlke, H. Systematically Scrutinizing the Impact of Substitution Sites on Thermostability and Detergent Tolerance for Bacillus subtilis Lipase A. Journal of Chemical Information and Modeling, 60(3): 1568–1584, March 2020. ISSN 1549-9596, 1549-960X. doi: 10.1021/acs.jcim.9b00954. URL https://pubs.acs.org/doi/10.1021/acs.jcim.9b00954.
  86. 86.Olson, C., Wu, N., and Sun, R. A Comprehensive Biophysical Description of Pairwise Epistasis throughout an Entire Protein Domain. Current Biology, 24(22):2643–2651, November 2014. ISSN 09609822. doi: 10.1016/j.cub.2014.09.072. URL https://linkinghub.elsevier.com/retrieve/pii/S0960982214012688.
  87. 87.Pokusaeva, V. O., Usmanova, D. R., Putintseva, E. V., Espinar, L., Sarkisyan, K. S., Mishin, A. S., Bogatyreva, N. S., Ivankov, D. N., Akopyan, A. V., Avvakumov, S. Y., Povolotskaya, I. S., Filion, G. J., Carey, L. B., and Kondrashov, F. A. An experimental assay of the interactions of amino acids from orthologous sequences shaping a complex fitness landscape. PLOS Genetics, 15(4):e1008079, April 2019. ISSN 1553-7404. doi: 10.1371/journal.pgen.1008079. URL https://dx.plos.org/10.1371/journal.pgen.1008079.
  88. 88.Press, O., Smith, N. A., and Lewis, M. Train short, test long: Attention with linear biases enables input length extrapolation, 2021.
  89. 89.Qi, H., Olson, C. A., Wu, N. C., Ke, R., Loverdo, C., Chu, V., Truong, S., Remenyi, R., Chen, Z., Du, Y., Su, S.-Y., Al-Mawsawi, L. Q., Wu, T.-T., Chen, S.-H., Lin, C.-Y., Zhong, W., Lloyd-Smith, J. O., and Sun, R. A Quantitative High-Resolution Genetic Profile Rapidly Identifies Sequence Determinants of Hepatitis C Viral Fitness and Drug Sensitivity. PLoS Pathogens, 10(4):e1004064, April 2014. ISSN 1553-7374. doi: 10.1371/journal.ppat.1004064. URL https://dx.plos.org/10.1371/journal.ppat.1004064.
  90. 90.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2018. URL https://d4mucfpksywv.cloudfront.net/better-language-models/language-models.pdf.
  91. 91.Radivojac, P., Obradovic, Z., Smith, D. K., Zhu, G., Vucetic, S., Brown, C. J., Lawson, J. D., and Dunker, A. K. Protein flexibility and intrinsic disorder. Protein Science, 13(1): 71–80, 2004.
  92. 92.Ramensky, V., Bork, P., and Sunyaev, S. Human non-synonymous snps: server and survey. Nucleic acids research, 30(17):3894–3900, 2002.
  93. 93.Rao, R., Bhattacharya, N., Thomas, N., Duan, Y., Chen, X., Canny, J. F., Abbeel, P., and Song, Y. S. Evaluating protein transfer learning with TAPE. CoRR, abs/1906.08230, 2019. URL http://arxiv.org/abs/1906.08230.
  94. 94.Rao, R., Meier, J., Sercu, T., Ovchinnikov, S., and Rives, A. Transformer protein language models are unsupervised structure learners. In International Conference on Learning Representations, 2020.
  95. 95.Rao, R., Liu, J., Verkuil, R., Meier, J., Canny, J. F., Abbeel, P., Sercu, T., and Rives, A. Msa transformer. bioRxiv, 2021. doi: 10.1101/2021.02.12.430858. URL https://www.biorxiv.org/content/early/2021/02/13/2021.02.12.430858.
  96. 96.Remmert, M., Biegert, A., Hauser, A., and Soding, J. Hh-blits: lightning-fast iterative protein sequence searching by hmm-hmm alignment. Nature Methods, 9:173–175, 2012.
  97. 97.Reva, B., Antipin, Y., and Sander, C. Predicting the functional impact of protein mutations: application to cancer genomics. Nucleic acids research, 39(17):e118–e118, 2011.
  98. 98.Riesselman, A. J., Ingraham, J. B., and Marks, D. S. Deep generative models of genetic variation capture the effects of mutations. Nature methods, 15(10):816–822, 2018.
  99. 99.Rives, A., Meier, J., Sercu, T., Goyal, S., Lin, Z., Liu, J., Guo, D., Ott, M., Zitnick, C. L., Ma, J., et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences, 118(15), 2021.
  100. 100.Robertson, S. E. and Zaragoza, H. The probabilistic relevance framework: Bm25 and beyond. Found. Trends Inf. Retr., 3:333–389, 2009.
  101. 101.Rockah-Shmuel, L., Toth-Petroczy, A., and Tawfik, D. S. Systematic Mapping of Protein Mutational Space by Prolonged Drift Reveals the Deleterious Effects of Seemingly Neutral Mutations. PLOS Computational Biology, 11(8):e1004421, August 2015. ISSN 1553-7358. doi: 10.1371/journal.pcbi.1004421. URL https://dx.plos.org/10.1371/journal.pcbi.1004421.
  102. 102.Romero, P. A., Tran, T. M., and Abate, A. R. Dissecting enzyme function with microfluidic-based deep mutational scanning. Proceedings of the National Academy of Sciences, 112(23):7159–7164, June 2015. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1422285112. URL http://www.pnas.org/lookup/doi/10.1073/pnas.1422285112.
  103. 103.Roscoe, B. P. and Bolon, D. N. Systematic Exploration of Ubiquitin Sequence, E1 Activation Efficiency, and Experimental Fitness in Yeast. Journal of Molecular Biology, 426(15):2854–2870, July 2014. ISSN 00222836. doi: 10.1016/j.jmb.2014.05.019. URL https://linkinghub.elsevier.com/retrieve/pii/S0022283614002587.
  104. 104.Roscoe, B. P., Thayer, K. M., Zeldovich, K. B., Fushman, D., and Bolon, D. N. Analyses of the Effects of All Ubiquitin Point Mutants on Yeast Growth Rate. Journal of Molecular Biology, 425(8):1363–1377, April 2013. ISSN 00222836. doi: 10.1016/j.jmb.2013.01.032. URL https://linkinghub.elsevier.com/retrieve/pii/S0022283613000636.
  105. 105.Russ, W. P., Figliuzzi, M., Stocker, C., Barrat-Charlaix, P., Socolich, M., Kast, P., Hilvert, D., Monasson, R., Cocco, S., Weigt, M., et al. An evolution-based model for designing chorismate mutase enzymes. Science, 369 (6502):440–445, 2020.
  106. 106.Sarkisyan, K. S., Bolotin, D. A., Meer, M. V., Usmanova, D. R., Mishin, A. S., Sharonov, G. V., Ivankov, D. N., Bozhanova, N. G., Baranov, M. S., Soylemez, O., et al. Local fitness landscape of the green fluorescent protein. Nature, 533(7603):397–401, 2016.
  107. 107.Seuma, M., Faure, A. J., Badia, M., Lehner, B., and Bolognesi, B. The genetic landscape for amyloid beta fibril nucleation accurately discriminates familial Alzheimer’s disease mutations. eLife, 10:e63364, February 2021. ISSN 2050-084X. doi: 10.7554/eLife.63364. URL https://elifesciences.org/articles/63364.
  108. 108.Shin, J.-E., Riesselman, A. J., Kollasch, A. W., McMahon, C., Simon, E., Sander, C., Manglik, A., Kruse, A. C., and Marks, D. S. Protein design and variant prediction using autoregressive generative models. Nature communications, 12(1):1–11, 2021.
  109. 109.Sievers, F., Wilm, A., Dineen, D., Gibson, T. J., Karplus, K., Li, W., Lopez, R., McWilliam, H., Remmert, M., Soding, J., Thompson, J. D., and Higgins, D. G. Fast, scalable generation of high-quality protein multiple sequence alignments using clustal omega. Molecular Systems Biology, 7:539 – 539, 2011.
  110. 110.Sinai, S., Jain, N., Church, G. M., and Kelsic, E. D. Generative aav capsid diversification by latent interpolation. bioRxiv, 2021.
  111. 111.So, D. R., Manke, W., Liu, H., Dai, Z., Shazeer, N., and Le, Q. V. Primer: Searching for efficient transformers for language modeling, 2021.
  112. 112.Soh, Y. S., Moncla, L. H., Eguia, R., Bedford, T., and Bloom, J. D. Comprehensive mapping of adaptation of the avian influenza polymerase protein PB2 to humans. eLife, 8:e45079, April 2019. ISSN 2050-084X. doi: 10.7554/eLife.45079. URL https://elifesciences.org/articles/45079.
  113. 113.Sourisseau, M., Lawrence, D. J. P., Schwarz, M. C., Storrs, C. H., Veit, E. C., Bloom, J. D., and Evans, M. J. Deep Mutational Scanning Comprehensively Maps How Zika Envelope Protein Mutations Affect Viral Growth and Antibody Escape. Journal of Virology, 93(23), December 2019. ISSN 0022-538X, 1098-5514. doi: 10.1128/JVI.01291-19. URL https://journals.asm.org/doi/10.1128/JVI.01291-19.
  114. 114.Staller, M. V., Holehouse, A. S., Swain-Lenz, D., Das, R. K., Pappu, R. V., and Cohen, B. A. A high-throughput mutational scan of an intrinsically disordered acidic transcriptional activation domain. Cell systems, 6(4):444–455, 2018.
  115. 115.Starita, L. M., Pruneda, J. N., Lo, R. S., Fowler, D. M., Kim, H. J., Hiatt, J. B., Shendure, J., Brzovic, P. S., Fields, S., and Klevit, R. E. Activity-enhancing mutations in an E3 ubiquitin ligase identified by high-throughput mutagenesis. Proceedings of the National Academy of Sciences, 110(14):E1263–E1272, April 2013. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1303309110. URL http://www.pnas.org/cgi/doi/10.1073/pnas.1303309110.
  116. 116.Starr, T. N., Greaney, A. J., Hilton, S. K., Ellis, D., Crawford, K. H., Dingens, A. S., Navarro, M. J., Bowen, J. E., Tortorici, M. A., Walls, A. C., King, N. P., Veesler, D., and Bloom, J. D. Deep Mutational Scanning of SARS-CoV-2 Receptor Binding Domain Reveals Constraints on Folding and ACE2 Binding. Cell, 182(5):1295–1310.e20, September 2020. ISSN 00928674. doi: 10.1016/j.cell.2020.08.012. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867420310035.
  117. 117.Steinegger, M. and Soding, J. Clustering huge protein sequence sets in linear time. Nature Communications, 9, 2018.
  118. 118.Steinegger, M., Meier, M., Mirdita, M., Vohringer, H., Haunsberger, S. J., and Soding, J. Hh-suite3 for fast remote homology detection and deep protein annotation. BMC Bioinformatics, 20, 2019.
  119. 119.Stiffler, M., Hekstra, D., and Ranganathan, R. Evolvability as a Function of Purifying Selection in TEM-1 β-Lactamase. Cell, 160(5):882–892, February 2015. ISSN 00928674. doi: 10.1016/j.cell.2015.01.035. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867415000781.
  120. 120.Suiter, C. C., Moriyama, T., Matreyek, K. A., Yang, W., Scaletti, E. R., Nishii, R., Yang, W., Hoshitsuki, K., Singh, M., Trehan, A., Parish, C., Smith, C., Li, L., Bhojwani, D., Yuen, L. Y. P., Li, C.-k., Li, C.-h., Yang, Y.-l., Walker, G. J., Goodhand, J. R., Kennedy, N. A., Klussmann, F. A., Bhatia, S., Relling, M. V., Kato, M., Hori, H., Bhatia, P., Ahmad, T., Yeoh, A. E. J., Stenmark, P., Fowler, D. M., and Yang, J. J. Massively parallel variant characterization identifies NUDT15 alleles associated with thiopurine toxicity. Proceedings of the National Academy of Sciences, 117(10):5394–5401, March 2020. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1915680117. URL http://www.pnas.org/lookup/doi/10.1073/pnas.1915680117.
  121. 121.Suzek, B. E., Wang, Y., Huang, H., McGarvey, P. B., Wu, C. H., and the UniProt Consortium. UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics, 31(6): 926–932, 11 2014. ISSN 1367-4803. doi: 10.1093/bioinformatics/btu739. URL https://doi.org/10.1093/bioinformatics/btu739.
  122. 122.Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. Going deeper with convolutions, 2014.
  123. 123.Thompson, J. D., Higgins, D. G., and Gibson, T. J. Clustal w: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice. Nucleic acids research, 22 22:4673–80, 1994.
  124. 124.Thompson, J. D., Gibson, T. J., Plewniak, F., Jeanmougin, F., and Higgins, D. G. The clustal x windows interface: flexible strategies for multiple sequence alignment aided by quality analysis tools. Nucleic acids research, 25 24: 4876–82, 1997.
  125. 125.Thompson, S., Zhang, Y., Ingle, C., Reynolds, K. A., and Kortemme, T. Altered expression of a quality control protease in E. coli reshapes the in vivo mutational landscape of a model enzyme. eLife, 9:e53476, July 2020. ISSN 2050-084X. doi: 10.7554/eLife.53476. URL https://elifesciences.org/articles/53476.
  126. 126.Toth-Petroczy, A., Palmedo, P., Ingraham, J., Hopf, T. A., Berger, B., Sander, C., and Marks, D. S. Structured states of disordered proteins from genomic sequences. Cell, 167 (1):158–170, 2016.
  127. 127.Tripathi, A., Gupta, K., Khare, S., Jain, P. C., Patel, S., Kumar, P., Pulianmackal, A. J., Aghera, N., and Varadarajan, R. Molecular Determinants of Mutant Phenotypes, Inferred from Saturation Mutagenesis Data. Molecular Biology and Evolution, 33(11):2960–2975, November 2016. ISSN 0737-4038, 1537-1719. doi: 10.1093/molbev/msw182. URL https://academic.oup.com/mbe/article-lookup/doi/10.1093/molbev/msw182.
  128. 128.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need, 2017.
  129. 129.Wang, S., Yu, M., Guo, X., Wang, Z., Klinger, T., Zhang, W., Chang, S., Tesauro, G., Zhou, B., and Jiang, J. R3: Reinforced ranker-reader for open-domain question answering. In AAAI, 2018.
  130. 130.Weile, J., Sun, S., Cote, A. G., Knapp, J., Verby, M., Mellor, J. C., Wu, Y., Pons, C., Wong, C., Lieshout, N., Yang, F., Tasan, M., Tan, G., Yang, S., Fowler, D. M., Nussbaum, R., Bloom, J. D., Vidal, M., Hill, D. E., Aloy, P., and Roth, F. P. A framework for exhaustively mapping functional missense variants. Molecular Systems Biology, 13(12):957, December 2017. ISSN 1744-4292, 1744-4292. doi: 10.15252/msb.20177908. URL https://onlinelibrary.wiley.com/doi/10.15252/msb.20177908.
  131. 131.Weinstein, E. N. and Marks, D. S. A structured observation distribution for generative biological sequence prediction and forecasting. bioRxiv, pp. 2020–07, 2021.
  132. 132.Wrenbeck, E. E., Azouz, L. R., and Whitehead, T. A. Single-mutation fitness landscapes for an enzyme on multiple substrates reveal specificity is globally encoded. Nature Communications, 8(1):15695, August 2017. ISSN 2041-1723. doi: 10.1038/ncomms15695. URL http://www.nature.com/articles/ncomms15695.
  133. 133.Wu, N. C., Young, A. P., Al-Mawsawi, L. Q., Olson, C. A., Feng, J., Qi, H., Chen, S.-H., Lu, I.-H., Lin, C.-Y., Chin, R. G., Luan, H. H., Nguyen, N., Nelson, S. F., Li, X., Wu, T.-T., and Sun, R. High-throughput profiling of influenza A virus hemagglutinin gene at single-nucleotide resolution. Scientific Reports, 4(1):4942, May 2014. ISSN 2045-2322. doi: 10.1038/srep04942. URL http://www.nature.com/articles/srep04942.
  134. 134.Wu, N. C., Olson, C. A., Du, Y., Le, S., Tran, K., Remenyi, R., Gong, D., Al-Mawsawi, L. Q., Qi, H., Wu, T.-T., and Sun, R. Functional Constraint Profiling of a Viral Protein Reveals Discordance of Evolutionary Conservation and Functionality. PLOS Genetics, 11(7):e1005310, July 2015. ISSN 1553-7404. doi: 10.1371/journal.pgen.1005310. URL https://dx.plos.org/10.1371/journal.pgen.1005310.
  135. 135.Young, H. J., Chan, M., Selvam, B., Szymanski, S. K., Shukla, D., and Procko, E. Deep Mutagenesis of a Transporter for Uptake of a Non-Native Substrate Identifies Conformationally Dynamic Regions. preprint, Biochemistry, April 2021. URL http://biorxiv.org/lookup/doi/10.1101/2021.04.19.440442.

Citation

MLA
Notin, P., et al. “Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval”. International Conference on Machine Learning, vol. 162, 2022, pp. 16990–7017, https://proceedings.mlr.press/v162/notin22a.html.
APA
Notin, P., Dias, M., Frazer, J., Marchena-Hurtado, J., Gomez, A. N., Marks, D., & Gal, Y. (2022). Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval. International Conference on Machine Learning, 162, 16990–17017. https://proceedings.mlr.press/v162/notin22a.html
Chicago
Notin, P., M. Dias, J. Frazer, et al. 2022. “Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval”. International Conference on Machine Learning 162: 16990–17017. https://proceedings.mlr.press/v162/notin22a.html.
Harvard
Notin, P. et al. (2022) “Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval”, International Conference on Machine Learning. PMLR, pp. 16990–17017. Available at: https://proceedings.mlr.press/v162/notin22a.html.
Vancouver
1. Notin P, Dias M, Frazer J, Marchena-Hurtado J, Gomez AN, Marks D, Gal Y (2022) Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval. In: International Conference on Machine Learning. PMLR, pp 16990–17017

BibTeX

@InProceedings{pmlr-v162-notin22a,
  title = 	 {Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval},
  author =       {Notin, Pascal and Dias, Mafalda and Frazer, Jonathan and Marchena-Hurtado, Javier and Gomez, Aidan N and Marks, Debora and Gal, Yarin},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {16990--17017},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/notin22a/notin22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/notin22a.html},
  abstract = 	 {The ability to accurately model the fitness landscape of protein sequences is critical to a wide range of applications, from quantifying the effects of human variants on disease likelihood, to predicting immune-escape mutations in viruses and designing novel biotherapeutic proteins. Deep generative models of protein sequences trained on multiple sequence alignments have been the most successful approaches so far to address these tasks. The performance of these methods is however contingent on the availability of sufficiently deep and diverse alignments for reliable training. Their potential scope is thus limited by the fact many protein families are hard, if not impossible, to align. Large language models trained on massive quantities of non-aligned protein sequences from diverse families address these problems and show potential to eventually bridge the performance gap. We introduce Tranception, a novel transformer architecture leveraging autoregressive predictions and retrieval of homologous sequences at inference to achieve state-of-the-art fitness prediction performance. Given its markedly higher performance on multiple mutants, robustness to shallow alignments and ability to score indels, our approach offers significant gain of scope over existing approaches. To enable more rigorous model testing across a broader range of protein families, we develop ProteinGym – an extensive set of multiplexed assays of variant effects, substantially increasing both the number and diversity of assays compared to existing benchmarks.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/