ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design

Pascal NotinAaron KollaschDaniel RitterLood van NiekerkSteffanie PaulHan SpinnerNathan J. RollinsAda ShawRose OrenbuchRuben Weitzman

article2023NeurIPS300 citations

Establishes a standardized, large-scale benchmarking suite of over 250 deep mutational scanning assays and clinical datasets to systematically evaluate more than 70 machine learning models across zero-shot and supervised protein fitness prediction and design tasks.

Listen

Understanding and engineering protein functions holds immense potential for addressing critical challenges in healthcare, agriculture, and climate solutions. While machine learning has emerged as a transformative technology for predicting how mutations alter protein behavior, evaluating these computational models has historically been difficult. Existing validation efforts have relied on small, fragmented, and artificial datasets that fail to reflect the true diversity of protein families and experimental environments, obscuring which models are genuinely effective for clinical and industrial use.

The article introduces ProteinGym, an expansive benchmarking platform designed to rigorously assess and compare machine learning models across protein fitness prediction and protein design tasks. It systematically evaluates both zero-shot methods, which predict mutational outcomes without prior experimental labels, and supervised methods across standardized real-world benchmarks.

To conduct this evaluation, the authors compiled over 250 deep mutational scanning assays covering more than 2.7 million mutated sequences across 200 protein families, spanning diverse taxa, functional categories, and mutation types including substitutions, insertions, and deletions. This experimental dataset is paired with expert-curated clinical records covering approximately 65,000 human mutations from the ClinVar and gnomAD databases. Across this comprehensive suite, the authors standardized and tested more than 70 high-performing computational models spanning alignment-based methods, large protein language models, structure-based inverse folding architectures, and hybrid systems under multiple cross-validation strategies.

The analysis produced several vital findings. First, hybrid architectures such as TranceptEVE demonstrated the strongest overall zero-shot performance across deep mutational scanning and clinical datasets, though specialized alignment-based models like GEMME excelled in specific niches such as viral proteins and shallow sequence alignments. Second, in supervised settings where labeled training data is available, non-parametric transformer models like ProteinNPT outperformed all alternatives by jointly analyzing sequence context and assay labels. Third, unsupervised models frequently matched or outperformed supervised models on human clinical benchmarks, highlighting how supervised predictors often suffer from data leakage and overfitting to known disease genes. Finally, model rankings varied meaningfully depending on the objective: some models excelled at rank-ordering all mutations across a full distribution, while others were significantly better at prioritizing top-performing candidates for protein design.

These findings demonstrate that no single computational model fits every application, meaning organizations must align model selection directly with operational goals. Teams seeking to optimize high-performing proteins for manufacturing or drug design should prioritize models optimized for top-tier retrieval metrics, whereas clinical diagnostics require models capable of accurate broad-spectrum pathogenicity scoring. Relying on the wrong predictive framework risks pursuing unviable therapeutic targets, inflating laboratory validation costs, and lengthening project timelines.

Decision-makers should immediately utilize standardized, multi-metric benchmarking suites like ProteinGym to select and validate computational biology models prior to committing wet-lab resources. Organizations should also adopt hybrid and autoregressive architectures as foundational starting points for mutational effect scoring and explore specialized transformer frameworks for supervised property optimization. Before major deployment in production, teams should conduct focused pilot evaluations on the specific protein class or target phenotype of interest.

Confidence in these findings is high due to the unprecedented breadth of evaluated assays, though certain limitations remain. Deep mutational scanning assays inherently contain measurement noise, dynamic range limits, and selection biases toward well-studied disease targets and specific protein families. Furthermore, existing clinical datasets present risks of circularity and label bias. Ongoing validation on unstudied protein families and expanded testing into non-coding regulatory sequences will be essential to ensure continuous predictive reliability.

Cover for ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design

Abstract

Predicting the effects of mutations in proteins is critical to many applications, from understanding genetic disease to designing novel proteins that can address our most pressing challenges in climate, agriculture and healthcare. Despite a surge in machine learning-based protein models to tackle these questions, an assessment of their respective benefits is challenging due to the use of distinct, often contrived, experimental datasets, and the variable performance of models across different protein families. Addressing these challenges requires scale. To that end we introduce ProteinGym, a large-scale and holistic set of benchmarks specifically designed for protein fitness prediction and design. It encompasses both a broad collection of over 250 standardized deep mutational scanning assays, spanning millions of mutated sequences, as well as curated clinical datasets providing high-quality expert annotations about mutation effects. We devise a robust evaluation framework that combines metrics for both fitness prediction and design, factors in known limitations of the underlying experimental methods, and covers both zero-shot and supervised settings. We report the performance of a diverse set of over 70 high-performing models from various subfields (eg., alignment-based, inverse folding) into a unified benchmark suite. We open source the corresponding codebase, datasets, MSAs, structures, model predictions and develop a user-friendly website that facilitates data access and analysis.

Table of Contents

  • 1 Introduction
  • 2 Related Work and Background
  • 3 ProteinGym benchmarks
  • 3.1 Mutation types
  • 3.2 Dataset types
  • 3.3 Model training regime
  • 4 Evaluation framework
  • 4.1 Zero-shot benchmarks
  • 4.2 Supervised benchmarks
  • 5 Results
  • 5.1 Substitution benchmarks
  • 5.2 Indel benchmarks
  • 6 Resources
  • 7 Conclusion
  • Acknowledgments and Disclosure of Funding
  • References
  • A Appendix
  • A.1 Social Impact
  • A.2 Limitations
  • A.3 Datasets
  • A.3.1 DMS assays
  • A.3.2 Clinical datasets
  • A.3.3 Access
  • A.3.4 License
  • A.4 Baselines
  • A.4.1 Zero-shot baselines
  • A.4.2 Supervised baselines
  • A.4.3 Clinical baselines
  • A.5 Detailed performance results
  • A.5.1 DMS substitution benchmarks
  • A.5.2 DMS indel benchmarks
  • A.5.3 Clinical substitution benchmarks
  • A.5.4 Clinical indel benchmarks

Knowls

  1. Knowl 1 — Structure and Composition of the ProteinGym Benchmark Suite

    experimental setup

    ProteinGym is a large-scale, standardized benchmark suite designed to evaluate machine learning models on protein fitness prediction and protein design tasks. The benchmark encompasses two complementary ground-truth data sources across two distinct mutation types:

    1. Deep Mutational Scanning (DMS) Assays:
      • Substitutions: 217 distinct experimental assays spanning approximately 2.4 million mutated protein sequences across more than 200 protein families. Assays cover diverse biological taxa (human, other eukaryotes, prokaryotes, viruses), varied multiple sequence alignment (MSA) depths, and include both single and multi-mutant variants.
      • Indels (Insertions and Deletions): 66 assays covering approximately 289,000 to 300,000 mutated sequences.
    2. Clinical Datasets (ClinVar and gnomAD):
      • Substitutions: Expert-curated clinical pathogenicity annotations from ClinVar covering 2,525 human genes and 63,000 missense variants (filtered for at least one star of clinical review evidence and mapped to the GRCh38 human genome assembly).
      • Indels: 1,555 human genes comprising approximately 3,000 short in-frame indels (length ≤3\le 3 amino acids). Because 84% of ClinVar indels are labeled pathogenic, common indels from gnomAD v2.1.1 (allele frequency >0.5%> 0.5\%) are included as benign pseudocontrols.

    In total, the benchmark comprises 3,422 unique proteins and over 2.7 million mutated sequence observations, evaluated across both zero-shot and supervised prediction regimes.

  2. Knowl 2 — Evaluation Metrics and Functional Group Weighting in ProteinGym

    model/method

    ProteinGym employs evaluation metrics tailored to both overall fitness landscape prediction and protein engineering (design optimization):

    • Spearman's Rank Correlation Coefficient (ρ\rho): Evaluates monotonic concordance between predicted fitness scores and experimental readouts, accounting for potential non-linear relationships between experimental assay signals and organismal fitness.
    • Normalized Discounted Cumulative Gain at Top 10% (NDCG@10%\text{NDCG@10\%}): Assesses the ranking quality specifically on the highest-fitness variants by up-weighting correct predictions among top experimental performers, reflecting protein engineering objectives.
    • Top 10% Recall: Measures the fraction of the top 10% highest-fitness experimental variants correctly identified within the top 10% of model-ranked variants.
    • Area Under the Receiver Operating Characteristic Curve (AUC) & Matthews Correlation Coefficient (MCC): Used to evaluate predictions on bimodal DMS assays (binarized at assay-specific thresholds) and binary clinical labels (pathogenic vs. benign).

    Functional Group Weighting ('Corrected Average'): To avoid biasing aggregate benchmark scores toward functional assays that are overrepresented in the literature (such as thermostability screens), substitution assays are grouped into five non-overlapping functional categories: Activity, Binding, Expression, Organismal Fitness, and Stability. Each metric is calculated as the unweighted mean of the average performance scores across these five functional classes.

  3. Knowl 3 — Zero-Shot Mutation Fitness Scoring Formulation Across Model Families

    model/method

    In the zero-shot regime, models score the fitness effect of a variant without training on experimental measurements for the target protein. For a wild-type sequence xwtx_{\text{wt}} and a mutated sequence xmutx_{\text{mut}}, the fitness prediction score s(xmut)s(x_{\text{mut}}) is calculated as the log-odds ratio: s(xmut)=log⁡p(xmut)p(xwt)=log⁡p(xmut)−log⁡p(xwt)s(x_{\text{mut}}) = \log \frac{p(x_{\text{mut}})}{p(x_{\text{wt}})} = \log p(x_{\text{mut}}) - \log p(x_{\text{wt}})

    Scoring across specific model paradigms is implemented as follows:

    1. Alignment-Based Generative Models (e.g., DeepSequence, EVE): The sequence probability p(x)p(x) under a variational autoencoder (VAE) parameterized by θ\theta is approximated by the evidence lower bound (ELBO): log⁡p(x∣θ)≥Eqϕ(z∣x)[log⁡pθ(x∣z)]−DKL(qϕ(z∣x)∥p(z))\log p(x \mid \theta) \ge \mathbb{E}_{q_\phi(z \mid x)}\left[\log p_\theta(x \mid z)\right] - D_{\mathrm{KL}}\left(q_\phi(z \mid x) \parallel p(z)\right)

    2. Autoregressive Protein Language Models (e.g., Tranception, ProGen2, RITA): Decompose the sequence probability into causal conditionals across length LL: log⁡p(x)=∑i=1Llog⁡p(xi∣x<i)\log p(x) = \sum_{i=1}^L \log p(x_i \mid x_{<i}) This formulation natively scores sequences of variable lengths, enabling zero-shot indel evaluation.

    3. Masked Protein Language Models (e.g., ESM-1v, ESM-2, MSA Transformer): For substitution mutations at positions M⊂{1,…,L}M \subset \{1, \dots, L\}, fitness is computed via masked-marginal log-odds: s(xmut)=∑i∈M[log⁡p(xi=xmut,i∣x\M)−log⁡p(xi=xwt,i∣x\M)]s(x_{\text{mut}}) = \sum_{i \in M} \left[ \log p(x_i = x_{\text{mut}, i} \mid x_{\backslash M}) - \log p(x_i = x_{\text{wt}, i} \mid x_{\backslash M}) \right]

    4. Inverse Folding Models (e.g., ProteinMPNN, ESM-IF1, MIF-ST): Learn conditional distributions p(x∣C)p(x \mid C) given 3D backbone coordinates CC (predicted using AlphaFold2 when experimental structures are unavailable): s(xmut)=log⁡p(xmut∣C)p(xwt∣C)s(x_{\text{mut}}) = \log \frac{p(x_{\text{mut}} \mid C)}{p(x_{\text{wt}} \mid C)}

  4. Knowl 4 — Cross-Validation Schemes for Supervised Protein Fitness Prediction

    model/method

    To evaluate supervised fitness prediction models and assess generalization to unobserved sequence positions, ProteinGym defines three distinct 5-fold cross-validation (CV) schemes for single-substitution DMS datasets:

    1. Random Split: Mutants are randomly partitioned into 5 folds. This benchmarks model interpolation in sequence space when mutations at the same residue positions have been observed during training.
    2. Contiguous Split: Mutated residue positions along the protein chain are divided into 5 contiguous segments of equal length. All mutations occurring within a given segment are assigned exclusively to one fold. This evaluates a model's capacity to extrapolate predictions across unobserved contiguous structural and functional regions.
    3. Modulo Split: Mutated positions are assigned to fold k∈{0,1,2,3,4}k \in \{0, 1, 2, 3, 4\} according to the residue position index ii via the modulo operation i(mod5)i \pmod 5. This ensures complete separation of positions between training and test sets while distributing held-out sites uniformly across the sequence length.

    Performance across all splits is reported using Spearman's rank correlation (ρ\rho) and Mean Squared Error (MSE) between model predictions and experimental measurements.

  5. Knowl 5 — Zero-Shot Substitution DMS Benchmark Results

    data/table

    ProteinGym compares zero-shot models on 217 substitution DMS assays. Performance is summarized by the corrected average across functional groups for Spearman rank correlation (ρ\rho), AUC, MCC, NDCG@10%, and top 10% Recall.

    Model Type Model Name Spearman AUC MCC NDCG Recall
    Alignment-based Site-Independent 0.359 0.696 0.286 0.747 0.201
    Alignment-based WaveNet 0.373 0.707 0.294 0.761 0.203
    Alignment-based EVmutation 0.395 0.716 0.305 0.777 0.222
    Alignment-based DeepSequence (ensemble) 0.419 0.729 0.328 0.776 0.226
    Alignment-based EVE (ensemble) 0.439 0.741 0.342 0.783 0.230
    Alignment-based GEMME 0.455 0.749 0.352 0.777 0.211
    Protein language UniRep 0.190 0.605 0.147 0.647 0.139
    Protein language CARP (640M) 0.368 0.701 0.285 0.748 0.208
    Protein language ESM-1b 0.394 0.719 0.311 0.747 0.203
    Protein language ESM-2 (15B) 0.401 0.720 0.314 0.759 0.208
    Protein language RITA XL 0.372 0.707 0.293 0.751 0.193
    Protein language ESM-1v (ensemble) 0.407 0.723 0.320 0.749 0.211
    Protein language ProGen2 XL 0.391 0.717 0.306 0.767 0.199
    Protein language VESPA 0.436 0.742 0.346 0.775 0.201
    Hybrid UniRep evotuned 0.347 0.693 0.274 0.739 0.181
    Hybrid MSA Transformer (ensemble) 0.434 0.738 0.340 0.779 0.224
    Hybrid Tranception L 0.434 0.739 0.341 0.779 0.220
    Hybrid TranceptEVE L 0.456 0.751 0.356 0.786 0.230
    Inverse Folding ESM-IF1 0.422 0.730 0.331 0.748 0.223
    Inverse Folding MIF-ST 0.401 0.718 0.311 0.766 0.226
    Inverse Folding ProteinMPNN 0.258 0.639 0.196 0.713 0.186

    Key takeaways:

    • The hybrid model TranceptEVE L achieves the top overall performance across all five metrics (Spearman 0.456, AUC 0.751, MCC 0.356, NDCG 0.786, Recall 0.230).
    • The alignment-based model GEMME performs second best overall (Spearman 0.455) and leads across viral proteins, non-human eukaryotic proteins, and low-to-medium MSA depths.
    • Alignment-based and co-evolutionary models tend to rank higher in NDCG relative to Spearman, indicating strong utility for selecting high-activity mutants for protein design, whereas protein language models exhibit stronger relative performance across the entire fitness distribution.
  6. Knowl 6 — Supervised Substitution DMS Benchmark Performance

    data/table

    ProteinGym benchmarks supervised models across 217 substitution DMS assays using Random (Rand.), Contiguous (Contig.), and Modulo (Mod.) 5-fold cross-validation splits. Baselines include one-hot encoded (OHE) ridge regression, pre-trained protein language model mean-pooled embeddings augmented with zero-shot fitness scores, and ProteinNPT.

    Model Type Model Name Spearman (↑\uparrow) MSE (↓\downarrow)
    Contig. Mod. Rand. Avg. Contig. Mod. Rand. Avg.
    OHE None 0.064 0.027 0.579 0.224 1.158 1.125 0.898 1.061
    OHE DeepSequence 0.400 0.400 0.521 0.440 0.967 0.940 0.767 0.891
    OHE ESM-1v 0.367 0.368 0.514 0.417 0.977 0.949 0.764 0.897
    OHE MSAT 0.410 0.412 0.536 0.453 0.963 0.934 0.749 0.882
    OHE Tranception 0.419 0.419 0.535 0.458 0.985 0.934 0.766 0.895
    OHE TranceptEVE 0.441 0.440 0.550 0.477 0.953 0.914 0.743 0.870
    Embeddings ESM-1v 0.481 0.506 0.639 0.542 0.937 0.861 0.563 0.787
    Embeddings MSAT 0.525 0.538 0.642 0.568 0.836 0.795 0.573 0.735
    Embeddings Tranception 0.490 0.526 0.696 0.571 0.972 0.833 0.503 0.769
    NPT ProteinNPT 0.547 0.564 0.730 0.613 0.820 0.771 0.459 0.683

    Key takeaways:

    • Unaugmented OHE ridge regression fails on position-extrapolation splits (Contiguous Spearman 0.064, Modulo Spearman 0.027), while zero-shot score augmentation restores extrapolation ability (0.417 to 0.477 average Spearman).
    • Pre-trained sequence embeddings outperform augmented OHE models across all cross-validation schemes (0.542 to 0.571 average Spearman).
    • ProteinNPT obtains the highest overall performance across all splits and metrics (average Spearman of 0.613, average MSE of 0.683), demonstrating the benefit of axial self-attention applied jointly across batch sequence rows and token/label columns.
  7. Knowl 7 — Zero-Shot Indel DMS Benchmark Results

    data/table

    ProteinGym evaluates indel-compatible models on 66 indel DMS assays. The benchmark splits assays into unbiased random mutational libraries ("Library") and model-designed sequences biased toward natural distributions ("Designed/Natural").

    Model Type Model Name Spearman by DMS Type (↑\uparrow) AUC (↑\uparrow)
    Library Designed/Natural All All
    Alignment models HMM 0.373 0.518 0.389 0.744
    Alignment models WaveNet 0.323 0.597 0.368 0.720
    Alignment models PROVEAN 0.306 0.585 0.347 0.725
    Protein language RITA L 0.443 0.519 0.457 0.773
    Protein language ProtGPT2 0.185 0.128 0.191 0.620
    Protein language ProGen2 M 0.472 0.205 0.465 0.776
    Hybrid models Tranception M 0.395 0.544 0.394 0.733
    Hybrid models Tranception L 0.387 0.563 0.395 0.741
    Hybrid models TranceptEVE M 0.426 0.587 0.424 0.754

    Key takeaways:

    • ProGen2 M attains the top rank on unbiased random Library assays (ρ=0.472\rho = 0.472, AUC = 0.776) but underperforms on Designed/Natural assays (ρ=0.205\rho = 0.205).
    • Alignment decoders (WaveNet and PROVEAN) perform best on Designed/Natural assays (ρ=0.597\rho = 0.597 and 0.5850.585, respectively) but perform poorly on random Library assays (ρ=0.323\rho = 0.323 and 0.3060.306).
    • Hybrid models such as TranceptEVE M provide balanced performance across both assay types (ρ=0.426\rho = 0.426 Library, 0.5870.587 Designed, 0.4240.424 All).
  8. Knowl 8 — Clinical Variant Effect Prediction and Supervised Model Generalization

    empirical result

    ProteinGym compares supervised clinical meta-predictors and unsupervised generative/language models on ClinVar pathogenic/benign classifications and human-variant DMS assays:

    1. Clinical Benchmark Performance:

      • Supervised models trained explicitly on ClinVar annotations achieve the highest AUROC scores on ClinVar test sets: ClinPred (0.981), MetaRNN (0.977), BayesDel (0.972), VEST4 (0.929), and REVEL (0.928).
      • Unsupervised models achieve competitive ClinVar AUROC performance without access to clinical labels during training: TranceptEVE (0.920), GEMME (0.919), EVE (0.917), and ESM-1b (0.892).
      • On ClinVar indels with gnomAD pseudocontrols, the alignment-based model PROVEAN achieves the highest AUROC (0.926) and AUPRC (0.947), followed by RITA XL (AUROC 0.923, AUPRC 0.954).
    2. Generalization to Experimental Assays:

      • When evaluated against experimental DMS assays assessing variant effects in human disease proteins, unsupervised models (TranceptEVE, GEMME, EVE) consistently outperform supervised clinical models.
      • Supervised clinical models are prone to data leakage from training-test gene overlap, ascertainment biases in ClinVar (e.g., European ancestry and cancer gene overrepresentation), and population frequency confounding in benign classifications.
  9. Knowl 9 — Functional Classification Scheme for Deep Mutational Scanning Assays

    definition

    ProteinGym categorizes substitution deep mutational scanning assays into five non-overlapping functional categories:

    1. Activity (43 assays): Directly or indirectly measures catalytic, enzymatic, or biochemical function.
    2. Binding (14 assays): Measures binding affinity, specificity, or target molecular interaction.
    3. Expression (17 assays): Measures intracellular protein abundance or cell-surface presentation.
    4. Organismal Fitness (77 assays): Measures cellular proliferation, growth rate, or viral replication in organismal systems.
    5. Stability (66 assays): Measures thermodynamic folding stability or thermostability.

    Disambiguation Rules:

    • Cell Growth Assays: Standard complementation assays where mutant functionality directly determines organismal growth are categorized as Organismal Fitness. When growth readouts are artificially engineered to couple to specific enzymatic rates, protein abundance, or binding affinity, assays are recategorized into Activity, Expression, or Binding respectively.
    • Cell Sorting Assays (e.g., FACS): Assays that fluorescently label the target protein to quantify abundance or membrane presentation are classified as Expression (or Stability when abundance depends directly on folding stability). Assays labeling downstream enzymatic products or binding partner recruitment are categorized as Activity or Binding respectively.
  10. Knowl 10 — Limitations of Deep Mutational Scanning and Clinical Benchmarks

    limitation

    ProteinGym highlights several structural limitations inherent to experimental DMS assays and clinical databases:

    1. DMS Assay Limitations:

      • Measurement Noise and Dynamic Range: Assays often impose artificial floor or ceiling effects (censoring dynamic range) and exhibit experimental replicate noise, placing an upper limit on achievable model correlation.
      • Protein Selection Bias: Assays over-represent well-studied disease-related genes (e.g., viral proteins, oncogenes) and structured proteins, while intrinsically disordered proteins are under-represented.
      • Phenotype Representativeness: Individual assays measure isolated molecular properties (e.g., binding or abundance) in specific environments, which may not capture multidimensional organismal fitness.
      • Processing Heterogeneity: Reported fitness values are derived using varied assay-specific statistical pipelines across different laboratories.
    2. Clinical Database Limitations:

      • Noise and Ascertainment Bias: ClinVar contains conflicting variant classifications and is heavily biased toward European ancestries and cancer genes.
      • Circularity and Leakage: Supervised clinical predictors exhibit training-test data leakage across homologous proteins, while unsupervised evolutionary models face subtle circularity because evolutionary conservation is often used as clinical evidence to assign ClinVar labels.
    3. Sequence Scope: ProteinGym benchmarks are restricted to protein-coding variants (substitutions and in-frame indels) and do not evaluate non-coding regulatory regions (promoters, enhancers, introns, UTRs).

Coverage note — Omitted materials include exhaustive appendix tables of per-assay results (Tables A15, A19, A20), individual ablation model sub-variants, and full multi-page model architecture implementation details in the appendix, in accordance with the requirement to distill standalone benchmark definitions, evaluation frameworks, core benchmark results, and methodological insights.

References

  1. 1.Christopher D. Aakre, Julien Herrou, Tuyen N. Phung, Barrett S. Perchuk, Sean Crosson, and Michael T. Laub. Evolving New Protein-Protein Interaction Specificity through Promiscuous Intermediates. Cell, 163(3): 594–606, October 2015. ISSN 00928674. doi: 10.1016/j.cell.2015.09.055. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867415012726.
  2. 2.Bharat V. Adkar, Arti Tripathi, Anusmita Sahoo, Kanika Bajaj, Devrishi Goswami, Purbani Chakrabarti, Mohit K. Swarnkar, Rajesh S. Gokhale, and Raghavan Varadarajan. Protein Model Discrimination Using Mutational Sensitivity Derived from Deep Sequencing. Structure, 20(2):371–381, February 2012. ISSN 09692126. doi: 10.1016/j.str.2011.11.021. URL https://linkinghub.elsevier.com/retrieve/pii/S0969212612000068.
  3. 3.Ivan A Adzhubei, Steffen Schmidt, Leonid Peshkin, Vasily E Ramensky, Anna Gerasimova, Peer Bork, Alexey S Kondrashov, and Shamil R Sunyaev. A method and server for predicting damaging missense mutations. Nature Methods, 7(4):248–249, April 2010. ISSN 1548-7091, 1548-7105. doi: 10.1038/nmeth0410-248. URL http://www.nature.com/articles/nmeth0410-248.
  4. 4.Ethan Ahler, Ames C. Register, Sujata Chakraborty, Linglan Fang, Emily M. Dieter, Katherine A. Sitko, Rama Subba Rao Vidadala, Bridget M. Trevillian, Martin Golkowski, Hannah Gelman, Jason J. Stephany, Alan F. Rubin, Ethan A. Merritt, Douglas M. Fowler, and Dustin J. Maly. A Combined Approach Reveals a Regulatory Mechanism Coupling Src’s Kinase Activity, Localization, and Phosphotransferase-Independent Functions. Molecular Cell, 74(2):393–408.e20, April 2019. ISSN 10972765. doi: 10.1016/j.molcel.2019.02.003. URL https://linkinghub.elsevier.com/retrieve/pii/S1097276519300930.
  5. 5.Najmeh Alirezaie, Kristin D Kernohan, Taila Hartley, Jacek Majewski, and Toby Dylan Hocking. Clinpred: prediction tool to identify disease-relevant nonsynonymous single-nucleotide variants. The American Journal of Human Genetics, 103(4):474–483, 2018.
  6. 6.Ethan C. Alley, Grigory Khimulya, Surojit Biswas, Mohammed AlQuraishi, and George M. Church. Unified rational protein engineering with sequence-based deep representation learning. Nature Methods, pages 1–8, 2019a.
  7. 7.Ethan C Alley, Grigory Khimulya, Surojit Biswas, Mohammed AlQuraishi, and George M Church. Unified rational protein engineering with sequence-based deep representation learning. Nature methods, 16(12): 1315–1322, 2019b.
  8. 8.Clara J. Amorosi, Melissa A. Chiasson, Matthew G. McDonald, Lai Hong Wong, Katherine A. Sitko, Gabriel Boyle, John P. Kowalski, Allan E. Rettie, Douglas M. Fowler, and Maitreya J. Dunham. Massively parallel characterization of CYP2C9 variant enzyme activity and abundance. The American Journal of Human Genetics, 108(9):1735–1751, September 2021. ISSN 00029297. doi: 10.1016/j.ajhg.2021.07.001. URL https://linkinghub.elsevier.com/retrieve/pii/S000292972100269X.
  9. 9.Bryan Andrews and Stanley Fields. Distinct patterns of mutational sensitivity for λ resistance and maltodextrin transport in escherichia coli LamB. Microbial Genomics, 6(4), April 2020.
  10. 10.Carlos L. Araya, Douglas M. Fowler, Wentao Chen, Ike Muniez, Jeffery W. Kelly, and Stanley Fields. A fundamental protein property, thermodynamic stability, revealed solely from large-scale measurements of protein function. Proceedings of the National Academy of Sciences, 109(42):16858–16863, October 2012. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1209751109. URL https://pnas.org/doi/full/10.1073/pnas.1209751109.
  11. 11.Pradeep Bandaru, Neel H Shah, Moitrayee Bhattacharyya, John P Barton, Yasushi Kondo, Joshua C Cofsky, Christine L Gee, Arup K Chakraborty, Tanja Kortemme, Rama Ranganathan, and John Kuriyan. Deconstruction of the Ras switching cycle through saturation mutagenesis. eLife, 6:e27810, July 2017. ISSN 2050-084X. doi: 10.7554/eLife.27810. URL https://elifesciences.org/articles/27810.
  12. 12.Jesse D Bloom. An experimentally determined evolutionary model dramatically improves phylogenetic fit. Molecular Biology and Evolution, 31(8):1956–1978, August 2014.
  13. 13.Benedetta Bolognesi, Andre J. Faure, Mireia Seuma, Jörn M. Schmiedel, Gian Gaetano Tartaglia, and Ben Lehner. The mutational landscape of a prion-like domain. Nature Communications, 10(1):4162, December 2019. ISSN 2041-1723. doi: 10.1038/s41467-019-12101-z. URL http://www.nature.com/articles/s41467-019-12101-z.
  14. 14.Jeffrey I Boucher, Daniel NA Bolon, and Dan S Tawfik. Quantifying and understanding the fitness effects of protein mutations: Laboratory versus nature. Protein Science, 25(7):1219–1226, 2016.
  15. 15.Nadav Brandes and Vasilis Ntranos. ESM variants - data & code for analysis and figures, June 2023. URL https://doi.org/10.5281/zenodo.8088402.
  16. 16.Nadav Brandes, Grant Goldman, Charlotte H Wang, Chun Jimmie Ye, and Vasilis Ntranos. Genome-wide prediction of disease variant effects with a deep protein language model. Nature Genetics, 55(9):1512–1522, 2023.
  17. 17.Lisa Brenan, Aleksandr Andreev, Ofir Cohen, Sasha Pantel, Atanas Kamburov, Davide Cacchiarelli, Nicole S. Persky, Cong Zhu, Mukta Bagul, Eva M. Goetz, Alex B. Burgin, Levi A. Garraway, Gad Getz, Tarjei S. Mikkelsen, Federica Piccioni, David E. Root, and Cory M. Johannessen. Phenotypic Characterization of a Comprehensive Set of MAPK1 /ERK2 Missense Mutants. Cell Reports, 17(4):1171–1183, October 2016. ISSN 22111247. doi: 10.1016/j.celrep.2016.09.061. URL https://linkinghub.elsevier.com/retrieve/pii/S2211124716313171.
  18. 18.Jessica L. Bridgford, Su Min Lee, Christine M. M. Lee, Paola Guglielmelli, Elisa Rumi, Daniela Pietra, Stephen Wilcox, Yash Chhabra, Alan F. Rubin, Mario Cazzola, Alessandro M. Vannucchi, Andrew J. Brooks, Matthew E. Call, and Melissa J. Call. Novel drivers and modifiers of MPL-dependent oncogenic transformation identified by deep mutational scanning. Blood, 135(4):287–292, January 2020. ISSN 0006-4971, 1528-0020. doi: 10.1182/blood.2019002561. URL https://ashpublications.org/blood/article/135/4/287/381157/Novel-drivers-and-modifiers-of-MPLdependent.
  19. 19.Ian J. Campbell, Joshua T. Atkinson, Matthew D. Carpenter, Dru Myerscough, Lin Su, Caroline M. Ajo-Franklin, and Jonathan J. Silberg. Determinants of Multiheme Cytochrome Extracellular Electron Transfer Uncovered by Systematic Peptide Insertion. Biochemistry, 61(13):1337–1350, July 2022. ISSN 0006-2960, 1520-4995. doi: 10.1021/acs.biochem.2c00148. URL https://pubs.acs.org/doi/10.1021/acs.biochem.2c00148.
  20. 20.Henriette Capel, Robin Weiler, Maurits Dijkstra, Reinier Vleugels, Peter Bloem, and K. Anton Feenstra. ProteinGLUE multi-task benchmark suite for self-supervised protein modeling. Scientific Reports, 12(1): 16047, September 2022. ISSN 2045-2322. doi: 10.1038/s41598-022-19608-4. URL https://www.nature.com/articles/s41598-022-19608-4. Number: 1 Publisher: Nature Publishing Group.
  21. 21.Hannah Carter, Christopher Douville, Peter D Stenson, David N Cooper, and Rachel Karchin. Identifying mendelian disease genes with the variant effect scoring tool. BMC genomics, 14(3):1–16, 2013.
  22. 22.Sujata Chakraborty, Ethan Ahler, Jessica J Simon, Linglan Fang, Zachary E Potter, Katherine A Sitko, Jason J Stephany, Miklos Guttman, Douglas M Fowler, and Dustin J Maly. Profiling of the drug resistance of thousands of src tyrosine kinase mutants uncovers a regulatory network that couples autoinhibition to catalytic domain dynamics. December 2021.
  23. 23.Kui K. Chan, Danielle Dorosky, Preeti Sharma, Shawn A. Abbasi, John M. Dye, David M. Kranz, Andrew S. Herbert, and Erik Procko. Engineering human ACE2 to optimize binding to the spike protein of SARS coronavirus 2. Science, 369(6508):1261–1265, September 2020. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.abc0870. URL https://www.science.org/doi/10.1126/science.abc0870.
  24. 24.Yvonne H. Chan, Sergey V. Venev, Konstantin B. Zeldovich, and C. Robert Matthews. Correlation of fitness landscapes from three orthologous TIM barrels originates from sequence and structure constraints. Nature Communications, 8(1):14614, April 2017. ISSN 2041-1723. doi: 10.1038/ncomms14614. URL http://www.nature.com/articles/ncomms14614.
  25. 25.John Z Chen, Douglas M Fowler, and Nobuhiko Tokuriki. Comprehensive exploration of the translocation, stability and substrate recognition requirements in VIM-2 lactamase. eLife, 9:e56707, June 2020. ISSN 2050-084X. doi: 10.7554/eLife.56707. URL https://elifesciences.org/articles/56707.
  26. 26.Tianlong Chen, Chengyue Gong, Daniel Jesus Diaz, Xuxi Chen, Jordan Tyler Wells, Qiang Liu, Zhangyang Wang, Andrew Ellington, Alex Dimakis, and Adam Klivans. HotProtein: A Novel Framework for Protein Thermostability Prediction and Editing. October 2022. URL https://openreview.net/forum?id=RtV_iEbWeGE.
  27. 27.Yongcan Chen, Ruyun Hu, Keyi Li, Yating Zhang, Lihao Fu, Jianzhi Zhang, and Tong Si. Deep mutational scanning of an Oxygen-Independent fluorescent protein CreiLOV for comprehensive profiling of mutational and epistatic effects. ACS Synthetic Biology, 12(5):1461–1473, May 2023.
  28. 28.Melissa A Chiasson, Nathan J Rollins, Jason J Stephany, Katherine A Sitko, Kenneth A Matreyek, Marta Verby, Song Sun, Frederick P Roth, Daniel DeSloover, Debora S Marks, Allan E Rettie, and Douglas M Fowler. Multiplexed measurement of variant abundance and activity reveals VKOR topology, active site and human variant impact. eLife, 9:e58026, September 2020. ISSN 2050-084X. doi: 10.7554/eLife.58026. URL https://elifesciences.org/articles/58026.
  29. 29.Yongwook Choi, Gregory E Sims, Sean Murphy, Jason R Miller, and Agnes P Chan. Predicting the functional effect of amino acid substitutions and indels. PLoS One, 7(10):e46688, October 2012.
  30. 30.Sung Chun and Justin C Fay. Identification of deleterious mutations within three human genomes. Genome research, 19(9):1553–1561, 2009.
  31. 31.Lene Clausen, Vasileios Voutsinos, Matteo Cagiada, Kristoffer E Johansson, Martin Grønbæk-Thygesen, Snehal Nariya, Rachel L Powell, Magnus K N Have, Vibe H Oestergaard, Amelie Stein, Douglas M Fowler, Kresten Lindorff-Larsen, and Rasmus Hartmann-Petersen. A mutational atlas for parkin proteostasis. June 2023.
  32. 32.Willow Coyote-Maestas, David Nedrud, Yungui He, and Daniel Schmidt. Determinants of trafficking, conduction, and disease within a K+ channel revealed through multiparametric deep mutational scanning. eLife, 11: e76903, May 2022. ISSN 2050-084X. doi: 10.7554/eLife.76903. URL https://elifesciences.org/articles/76903.
  33. 33.Christian Dallago, Jody Mou, Kadina E Johnston, Bruce J Wittmann, Nicholas Bhattacharya, Samuel Goldman, Ali Madani, and Kevin K Yang. FLIP: Benchmark tasks in fitness landscape inference for proteins. 2021.
  34. 34.Rohan Dandage, Rajesh Pandey, Gopal Jayaraj, Manish Rai, David Berger, and Kausik Chakraborty. Differential strengths of molecular determinants guide environment specific mutational fates. PLOS Genetics, 14(5): e1007419, May 2018. ISSN 1553-7404. doi: 10.1371/journal.pgen.1007419. URL https://dx.plos.org/10.1371/journal.pgen.1007419.
  35. 35.J Dauparas, I Anishchenko, N Bennett, H Bai, R J Ragotte, L F Milles, B I M Wicky, A Courbet, R J de Haas, N Bethel, P J Y Leung, T F Huddy, S Pellock, D Tischer, F Chan, B Koepnick, H Nguyen, A Kang, B Sankaran, A K Bera, N P King, and D Baker. Robust deep learning-based protein sequence design using ProteinMPNN. Science, 378(6615):49–56, October 2022.
  36. 36.Natalie L. Dawson, Tony E. Lewis, Sayoni Das, Jonathan G. Lees, David A. Lee, Paul Ashford, Christine A. Orengo, and Ian P. W. Sillitoe. Cath: an expanded resource to predict protein function through structure and sequence. Nucleic Acids Research, 45:D289 – D295, 2016. URL https://api.semanticscholar.org/CorpusID:9356024.
  37. 37.Zhifeng Deng, Wanzhi Huang, Erol Bakkalbasi, Nicholas G. Brown, Carolyn J. Adamski, Kacie Rice, Donna Muzny, Richard A. Gibbs, and Timothy Palzkill. Deep Sequencing of Systematic Combinatorial Libraries Reveals beta-Lactamase Sequence Constraints at High Resolution. Journal of Molecular Biology, 424(3-4): 150–167, December 2012. ISSN 00222836. doi: 10.1016/j.jmb.2012.09.014. URL https://linkinghub.elsevier.com/retrieve/pii/S0022283612007711.
  38. 38.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding, 2019.
  39. 39.David Ding, Ada Shaw, Sam Sinai, Nathan Rollins, Noam Prywes, David F Savage, Michael T Laub, and Debora S Marks. Protein design using structure-based residue preferences. June 2023.
  40. 40.Michael Doud and Jesse Bloom. Accurate Measurement of the Effects of All Amino-Acid Mutations on Influenza Hemagglutinin. Viruses, 8(6):155, June 2016. ISSN 1999-4915. doi: 10.3390/v8060155. URL http://www.mdpi.com/1999-4915/8/6/155.
  41. 41.Michael B. Doud, Orr Ashenberg, and Jesse D. Bloom. Site-Specific Amino Acid Preferences Are Mostly Conserved in Two Closely Related Protein Homologs. Molecular Biology and Evolution, 32(11):2944–2960, November 2015. ISSN 0737-4038, 1537-1719. doi: 10.1093/molbev/msv167. URL https://academic.oup.com/mbe/article-lookup/doi/10.1093/molbev/msv167.
  42. 42.Maria Duenas-Decamp, Li Jiang, Daniel Bolon, and Paul R. Clapham. Saturation Mutagenesis of the HIV-1 Envelope CD4 Binding Loop Reveals Residues Controlling Distinct Trimer Conformations. PLOS Pathogens, 12(11):e1005988, November 2016. ISSN 1553-7374. doi: 10.1371/journal.ppat.1005988. URL https://dx.plos.org/10.1371/journal.ppat.1005988.
  43. 43.Richard Durbin, Sean Eddy, Anders Krogh, and Graeme Mitchison. Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids. Cambridge University Press, 1998.
  44. 44.Sean R Eddy. Accelerated profile HMM searches. PLoS Comput. Biol., 7(10):e1002195, October 2011.
  45. 45.Assaf Elazar, Jonathan Weinstein, Ido Biran, Yearit Fridman, Eitan Bibi, and Sarel Jacob Fleishman. Mutational scanning reveals the determinants of protein insertion and association energetics in the plasma membrane. eLife, 5:e12125, January 2016. ISSN 2050-084X. doi: 10.7554/eLife.12125. URL https://elifesciences.org/articles/12125.
  46. 46.Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Wang Yu, Llion Jones, Tom Gibbs, Tamas B. Fehér, Christoph Angerer, Martin Steinegger, Debsindhu Bhowmik, and Burkhard Rost. Prottrans: Towards cracking the language of lifes code through self-supervised deep learning and high performance computing. IEEE transactions on pattern analysis and machine intelligence, PP, 2021.
  47. 47.Stefan Engelen, Ladislas A Trojan, Sophie Sacquin-Mora, Richard Lavery, and Alessandra Carbone. Joint evolutionary trees: a large-scale method to predict protein interfaces based on sequence sampling. PLoS computational biology, 5(1):e1000267, 2009.
  48. 48.Steven Erwood, Teija M. I. Bily, Jason Lequyer, Joyce Yan, Nitya Gulati, Reid A. Brewer, Liangchi Zhou, Laurence Pelletier, Evgueni A. Ivakine, and Ronald D. Cohn. Saturation variant interpretation using CRISPR prime editing. Nature Biotechnology, 40(6):885–895, June 2022. ISSN 1087-0156, 1546-1696. doi: 10.1038/s41587-021-01201-1. URL https://www.nature.com/articles/s41587-021-01201-1.
  49. 49.Gabriella O. Estevam, Edmond M. Linossi, Christian B. Macdonald, Carla A. Espinoza, Jennifer M. Michaud, Willow Coyote-Maestas, Eric A. Collisson, Natalia Jura, and James S. Fraser. Conserved regulatory motifs in the juxtamembrane domain and kinase N-lobe revealed through deep mutational scanning of the MET receptor tyrosine kinase domain. preprint, Molecular Biology, August 2023. URL http://biorxiv.org/lookup/doi/10.1101/2023.08.03.551866.
  50. 50.Andre J. Faure, Júlia Domingo, Jörn M. Schmiedel, Cristina Hidalgo-Carcedo, Guillaume Diss, and Ben Lehner. Mapping the energetic and allosteric landscapes of protein binding domains. Nature, 604(7904): 175–183, April 2022. ISSN 0028-0836, 1476-4687. doi: 10.1038/s41586-022-04586-4. URL https://www.nature.com/articles/s41586-022-04586-4.
  51. 51.Bing-Jian Feng. PERCH: a unified framework for disease gene prioritization. Human mutation, 38(3):243–251, 2017.
  52. 52.Jason D. Fernandes, Tyler B. Faust, Nicolas B. Strauli, Cynthia Smith, David C. Crosby, Robert L. Nakamura, Ryan D. Hernandez, and Alan D. Frankel. Functional Segregation of Overlapping Genes in HIV. Cell, 167(7):1762–1773.e12, December 2016. ISSN 00928674. doi: 10.1016/j.cell.2016.11.031. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867416316038.
  53. 53.Noelia Ferruz, Steffen Schmidt, and Birte Höcker. ProtGPT2 is a deep unsupervised language model for protein design. Nature Communications, 13, 2022.
  54. 54.Gregory M. Findlay, Riza M. Daza, Beth Martin, Melissa D. Zhang, Anh P. Leith, Molly Gasperini, Joseph D. Janizek, Xingfan Huang, Lea M. Starita, and Jay Shendure. Accurate classification of BRCA1 variants with saturation genome editing. Nature, 562(7726):217–222, October 2018. ISSN 0028-0836, 1476-4687. doi: 10.1038/s41586-018-0461-z. URL http://www.nature.com/articles/s41586-018-0461-z.
  55. 55.Elad Firnberg, Jason W. Labonte, Jeffrey J. Gray, and Marc Ostermeier. A Comprehensive, High-Resolution Map of a Gene’s Fitness Landscape. Molecular Biology and Evolution, 31(6):1581–1592, June 2014. ISSN 1537-1719, 0737-4038. doi: 10.1093/molbev/msu081. URL https://academic.oup.com/mbe/article-lookup/doi/10.1093/molbev/msu081.
  56. 56.Julia M Flynn, Ammeret Rossouw, Pamela Cote-Hammarlof, Inês Fragata, David Mavor, Carl Hollins, Claudia Bank, and Daniel Na Bolon. Comprehensive fitness maps of Hsp90 show widespread environmental dependence. eLife, 9:e53810, March 2020. ISSN 2050-084X. doi: 10.7554/eLife.53810. URL https://elifesciences.org/articles/53810.
  57. 57.Julia M. Flynn, Neha Samant, Gily Schneider-Nachum, David T. Barkan, Nese Kurt Yilmaz, Celia A. Schiffer, Stephanie A. Moquin, Dustin Dovala, and Daniel N.A. Bolon. Comprehensive fitness landscape of SARS-CoV-2 M pro reveals insights into viral resistance mechanisms. preprint, Molecular Biology, January 2022. URL http://biorxiv.org/lookup/doi/10.1101/2022.01.26.477860.
  58. 58.Jonathan Frazer, Pascal Notin, Mafalda Dias, Aidan Gomez, Joseph K Min, Kelly P. Brock, Yarin Gal, and Debora S. Marks. Disease variant prediction with deep generative models of evolutionary data. Nature, 2021.
  59. 59.Kiran S Gajula, Peter J Huwe, Charlie Y Mo, Daniel J Crawford, James T Stivers, Ravi Radhakrishnan, and Rahul M Kohli. High-throughput mutagenesis reveals functional determinants for DNA targeting by activation-induced deaminase. Nucleic Acids Research, 42(15):9964–9975, September 2014.
  60. 60.Zhangyang Gao, Cheng Tan, and Stan Z. Li. Pifold: Toward effective and efficient protein inverse folding. ArXiv, abs/2209.12643, 2022. URL https://api.semanticscholar.org/CorpusID:252596302.
  61. 61.Sarah Gersing, Matteo Cagiada, Marinella Gebbia, Anette P. Gjesing, Atina G. Coté, Gireesh Seesankar, Roujia Li, Daniel Tabet, Amelie Stein, Anna L. Gloyn, Torben Hansen, Frederick P. Roth, Kresten Lindorff-Larsen, and Rasmus Hartmann-Petersen. A comprehensive map of human glucokinase variant activity. preprint, Genetics, May 2022. URL http://biorxiv.org/lookup/doi/10.1101/2022.05.04.490571.
  62. 62.Sarah Gersing, Thea K Schulze, Matteo Cagiada, Amelie Stein, Frederick P Roth, Kresten Lindorff-Larsen, and Rasmus Hartmann-Petersen. Characterizing glucokinase variant mechanisms using a multiplexed abundance assay. bioRxiv, May 2023.
  63. 63.Dia A Ghose, Kaitlyn E Przydzial, Emily M Mahoney, Amy E Keating, and Michael T Laub. Marginal specificity in protein interactions constrains evolution of a paralogous family. Proceedings of the National Academy of Sciences of the United States of America, 120(18):e2221163120, May 2023.
  64. 64.Andrew O. Giacomelli, Xiaoping Yang, Robert E. Lintner, James M. McFarland, Marc Duby, Jaegil Kim, Thomas P. Howard, David Y. Takeda, Seav Huong Ly, Eejung Kim, Hugh S. Gannon, Brian Hurhula, Ted Sharpe, Amy Goodale, Briana Fritchman, Scott Steelman, Francisca Vazquez, Aviad Tsherniak, Andrew J. Aguirre, John G. Doench, Federica Piccioni, Charles W. M. Roberts, Matthew Meyerson, Gad Getz, Cory M. Johannessen, David E. Root, and William C. Hahn. Mutational processes shape the landscape of TP53 mutations in human cancer. Nature Genetics, 50(10):1381–1387, October 2018. ISSN 1061-4036, 1546-1718. doi: 10.1038/s41588-018-0204-y. URL https://www.nature.com/articles/s41588-018-0204-y.
  65. 65.Kevin S Gill, Kritika Mehta, Jeremiah D Heredia, Vishnu V Krishnamurthy, Kai Zhang, and Erik Procko. Multiple mechanisms of self-association of chemokine receptors CXCR4 and CCR5 demonstrated by deep mutagenesis. bioRxiv, March 2023.
  66. 66.Andrew M. Glazer, Brett M. Kroncke, Kenneth A. Matreyek, Tao Yang, Yuko Wada, Tiffany Shields, Joe-Elie Salem, Douglas M. Fowler, and Dan M. Roden. Deep Mutational Scan of an SCN5A Voltage Sensor. Circulation: Genomic and Precision Medicine, 13(1):e002786, February 2020. ISSN 2574-8300. doi: 10.1161/CIRCGEN.119.002786. URL https://www.ahajournals.org/doi/10.1161/CIRCGEN.119.002786.
  67. 67.Courtney E. Gonzalez, Paul Roberts, and Marc Ostermeier. Fitness Effects of Single Amino Acid Insertions and Deletions in TEM-1 beta-Lactamase. Journal of Molecular Biology, 431(12):2320–2330, May 2019. ISSN 00222836. doi: 10.1016/j.jmb.2019.04.030. URL https://linkinghub.elsevier.com/retrieve/pii/S0022283619302372.
  68. 68.Louisa Gonzalez Somermeyer, Aubin Fleiss, Alexander S Mishin, Nina G Bozhanova, Anna A Igolkina, Jens Meiler, Maria-Elisenda Alaball Pujol, Ekaterina V Putintseva, Karen S Sarkisyan, and Fyodor A Kondrashov. Heterogeneity of the GFP fitness landscape and data-driven protein design. eLife, 11:e75842, May 2022. ISSN 2050-084X. doi: 10.7554/eLife.75842. URL https://elifesciences.org/articles/75842.
  69. 69.Vanessa E Gray, Katherine Sitko, Floriane Z Ngako Kameni, Miriam Williamson, Jason J Stephany, Nicholas Hasle, and Douglas M Fowler. Elucidating the molecular determinants of Aβ aggregation with deep mutational scanning. G3, 9(11):3683–3689, November 2019.
  70. 70.Dominik G Grimm, Chloé-Agathe Azencott, Fabian Aicheler, Udo Gieraths, Daniel G MacArthur, Kaitlin E Samocha, David N Cooper, Peter D Stenson, Mark J Daly, Jordan W Smoller, Laramie E Duncan, and Karsten M Borgwardt. The evaluation of tools used to predict the impact of missense variants is hindered by two types of circularity. Hum. Mutat., 36(5):513–523, May 2015.
  71. 71.Hugh K Haddox, Adam S Dingens, Sarah K Hilton, Julie Overbaugh, and Jesse D Bloom. Mapping mutational effects along the evolutionary landscape of HIV envelope. eLife, 7:e34420, March 2018. ISSN 2050-084X. doi: 10.7554/eLife.34420. URL https://elifesciences.org/articles/34420.
  72. 72.Michael Heinzinger, Ahmed Elnaggar, Yu Wang, Christian Dallago, Dmitrii Nechaev, Florian Matthes, and Burkhard Rost. Modeling aspects of the language of life through transfer-learning protein sequences. BMC Bioinformatics, 20(1):723, December 2019. ISSN 1471-2105. doi: 10.1186/s12859-019-3220-8. URL https://doi.org/10.1186/s12859-019-3220-8.
  73. 73.Daniel Hesslow, N. ed. Zanichelli, Pascal Notin, Iacopo Poli, and Debora S. Marks. RITA: a study on scaling up generative protein sequence models. ArXiv, abs/2205.05789, 2022.
  74. 74.Ryan T Hietpas, Jeffrey D Jensen, and Daniel N A Bolon. Experimental illumination of a fitness landscape. Proceedings of the National Academy of Sciences of the United States of America, 108(19):7896–7901, May 2011.
  75. 75.Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. Axial attention in multidimensional transformers. ArXiv, abs/1912.12180, 2019a. URL https://api.semanticscholar.org/CorpusID:209323787.
  76. 76.Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. Axial attention in multidimensional transformers. arXiv preprint arXiv:1912.12180, 2019b.
  77. 77.Helen T. Hobbs, Neel H. Shah, Sophie R. Shoemaker, Jeanine F. Amacher, Susan Marqusee, and John Kuriyan. Saturation mutagenesis of a predicted ancestral Syk-family kinase. Protein Science, 31(10), October 2022. ISSN 0961-8368, 1469-896X. doi: 10.1002/pro.4411. URL https://onlinelibrary.wiley.com/doi/10.1002/pro.4411.
  78. 78.Nancy Hom, Lauren Gentles, Jesse D Bloom, and Kelly K Lee. Deep mutational scan of the highly conserved influenza a virus M1 matrix protein reveals substantial intrinsic mutational tolerance. Journal of Virology, 93 (13), July 2019.
  79. 79.Jirí Hon, Martin Marusiak, Tomáš Martínek, Antonin Kunka, Jaroslav Zendulka, David Bednáˇr, and Jiˇrí Damborský. SoluProt: prediction of soluble protein expression in escherichia coli. Bioinformatics, 37:23 – 28, 2020.
  80. 80.Thomas A Hopf, John B Ingraham, Frank J Poelwijk, Charlotta PI Schärfe, Michael Springer, Chris Sander, and Debora S Marks. Mutation effects predicted from sequence co-variation. Nature biotechnology, 35(2): 128–135, 2017.
  81. 81.Chloe Hsu, Hunter Nisonoff, Clara Fannjiang, and Jennifer Listgarten. Learning protein fitness models from evolutionary and assay-labeled data. Nature Biotechnology, 40(7):1114–1122, July 2022a. ISSN 1087-0156, 1546-1696. doi: 10.1038/s41587-021-01146-5. URL https://www.nature.com/articles/s41587-021-01146-5.
  82. 82.Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives. Learning inverse folding from millions of predicted structures. April 2022b.
  83. 83.Po-Ssu Huang, Scott E. Boyken, and David Baker. The coming of age of de novo protein design. Nature, 537: 320–327, 2016. URL https://api.semanticscholar.org/CorpusID:205251398.
  84. 84.Zachary M Huttinger, Laura M Haynes, Andrew Yee, Colin A Kretz, Matthew L Holding, David R Siemieniak, Daniel A Lawrence, and David Ginsburg. Deep mutational scanning of the plasminogen activator inhibitor-1 functional landscape. Scientific Reports, 11(1):18827, September 2021.
  85. 85.John Ingraham, Vikas Garg, Regina Barzilay, and Tommi Jaakkola. Generative models for graph-based protein design. Advances in neural information processing systems, 32, 2019.
  86. 86.Nilah M Ioannidis, Joseph H Rothstein, Vikas Pejaver, Sumit Middha, Shannon K McDonnell, Saurabh Baheti, Anthony Musolf, Qing Li, Emily Holzinger, Danielle Karyadi, et al. REVEL: an ensemble method for predicting the pathogenicity of rare missense variants. The American Journal of Human Genetics, 99(4): 877–885, 2016.
  87. 87.Hervé Jacquier, André Birgy, Hervé Le Nagard, Yves Mechulam, Emmanuelle Schmitt, Jérémy Glodt, Beatrice Bercot, Emmanuelle Petit, Julie Poulain, Guilène Barnaud, Pierre-Alexis Gros, and Olivier Tenaillon. Capturing the mutational landscape of the beta-lactamase TEM-1. Proceedings of the National Academy of Sciences, 110(32):13067–13072, August 2013. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1215206110. URL https://pnas.org/doi/full/10.1073/pnas.1215206110.
  88. 88.Milind Jagota, Chengzhong Ye, Ruchir Rastogi, Carlos Albors, Antoine Koehl, Nilah M. Ioannidis, and Yun S. Song. Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects. 2022. URL https://api.semanticscholar.org/CorpusID:253628877.
  89. 89.Xiaoyan Jia, Bala Bharathi Burugula, Victor Chen, Rosemary M. Lemons, Sajini Jayakody, Mariam Maksutova, and Jacob O. Kitzman. Massively parallel functional testing of MSH2 missense variants conferring Lynch syndrome risk. The American Journal of Human Genetics, 108(1):163–175, January 2021. ISSN 00029297. doi: 10.1016/j.ajhg.2020.12.003. URL https://linkinghub.elsevier.com/retrieve/pii/S0002929720304390.
  90. 90.Li Jiang, Ping Liu, Claudia Bank, Nicholas Renzette, Kristina Prachanronarong, Lutfu S. Yilmaz, Daniel R. Caffrey, Konstantin B. Zeldovich, Celia A. Schiffer, Timothy F. Kowalik, Jeffrey D. Jensen, Robert W. Finberg, Jennifer P. Wang, and Daniel N.A. Bolon. A Balance between Inhibitor Binding and Substrate Processing Confers Influenza Drug Resistance. Journal of Molecular Biology, 428(3):538–553, February 2016. ISSN 00222836. doi: 10.1016/j.jmb.2015.11.027. URL https://linkinghub.elsevier.com/retrieve/pii/S0022283615006907.
  91. 91.Rosanna Junchen Jiang. Exhaustive Mapping of Missense Variation in Coronary Heart Disease-related Genes. PhD thesis, University of Toronto, November 2019. URL https://hdl.handle.net/1807/98076.
  92. 92.Bowen Jing, Stephan Eismann, Patricia Suriana, Raphael J L Townshend, and Ron Dror. Learning from protein structure with geometric vector perceptrons. September 2020.
  93. 93.Eric M Jones, Nathan B Lubock, Aj Venkatakrishnan, Jeffrey Wang, Alex M Tseng, Joseph M Paggi, Naomi R Latorraca, Daniel Cancilla, Megan Satyadi, Jessica E Davis, M Madan Babu, Ron O Dror, and Sriram Kosuri. Structural and functional characterization of G protein–coupled receptors with deep mutational scanning. eLife, 9:e54895, October 2020. ISSN 2050-084X. doi: 10.7554/eLife.54895. URL https://elifesciences.org/articles/54895.
  94. 94.John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A A Kohl, Andrew J Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas Berghammer, Sebastian Bodenstein, David Silver, Oriol Vinyals, Andrew W Senior, Koray Kavukcuoglu, Pushmeet Kohli, and Demis Hassabis. Highly accurate protein structure prediction with AlphaFold. Nature, July 2021.
  95. 95.Konrad J Karczewski, Laurent C Francioli, Grace Tiao, Beryl B Cummings, Jessica Alföldi, Qingbo Wang, Ryan L Collins, Kristen M Laricchia, Andrea Ganna, Daniel P Birnbaum, Laura D Gauthier, Harrison Brand, Matthew Solomonson, Nicholas A Watts, Daniel Rhodes, Moriel Singer-Berk, Eleina M England, Eleanor G Seaby, Jack A Kosmicki, Raymond K Walters, Katherine Tashman, Yossi Farjoun, Eric Banks, Timothy Poterba, Arcturus Wang, Cotton Seed, Nicola Whiffin, Jessica X Chong, Kaitlin E Samocha, Emma Pierce-Hoffman, Zachary Zappala, Anne H O’Donnell-Luria, Eric Vallabh Minikel, Ben Weisburd, Monkol Lek, James S Ware, Christopher Vittal, Irina M Armean, Louis Bergelson, Kristian Cibulskis, Kristen M Connolly, Miguel Covarrubias, Stacey Donnelly, Steven Ferriera, Stacey Gabriel, Jeff Gentry, Namrata Gupta, Thibault Jeandet, Diane Kaplan, Christopher Llanwarne, Ruchi Munshi, Sam Novod, Nikelle Petrillo, David Roazen, Valentin Ruano-Rubio, Andrea Saltzman, Molly Schleicher, Jose Soto, Kathleen Tibbetts, Charlotte Tolonen, Gordon Wade, Michael E Talkowski, Genome Aggregation Database Consortium, Benjamin M Neale, Mark J Daly, and Daniel G MacArthur. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature, 581(7809):434–443, May 2020.
  96. 96.Eric D. Kelsic, Hattie Chung, Niv Cohen, Jimin Park, Harris H. Wang, and Roy Kishony. RNA Structural Determinants of Optimal Codons Revealed by MAGE-Seq. Cell Systems, 3(6):563–571.e6, December 2016. ISSN 24054712. doi: 10.1016/j.cels.2016.11.004. URL https://linkinghub.elsevier.com/retrieve/pii/S2405471216303684.
  97. 97.Paul Kennouche, Arthur Charles-Orszag, Daiki Nishiguchi, Sylvie Goussard, Anne-Flore Imhaus, Mathieu Dupré, Julia Chamot-Rooke, and Guillaume Duménil. Deep mutational scanning of the Neisseria meningitidis major pilin reveals the importance of pilus tip-mediated adhesion. The EMBO Journal, 38(22):e102145, November 2019. ISSN 0261-4189, 1460-2075. doi: 10.15252/embj.2019102145. URL https://www.embopress.org/doi/10.15252/embj.2019102145.
  98. 98.Jacob O Kitzman, Lea M Starita, Russell S Lo, Stanley Fields, and Jay Shendure. Massively parallel single-amino-acid mutagenesis. Nature Methods, 12(3):203–206, March 2015. ISSN 1548-7091, 1548-7105. doi: 10.1038/nmeth.3223. URL http://www.nature.com/articles/nmeth.3223.
  99. 99.Justin R. Klesmith, John-Paul Bacik, Ryszard Michalczyk, and Timothy A. Whitehead. Comprehensive Sequence-Flux Mapping of a Levoglucosan Utilization Pathway in E. coli. ACS Synthetic Biology, 4 (11):1235–1243, November 2015. ISSN 2161-5063, 2161-5063. doi: 10.1021/acssynbio.5b00131. URL https://pubs.acs.org/doi/10.1021/acssynbio.5b00131.
  100. 100.Justin R. Klesmith, Lihe Su, Lan Wu, Ian A. Schrack, Fay J. Dufort, Alyssa Birt, Christine Ambrose, Benjamin J. Hackel, Roy R. Lobb, and Paul D. Rennert. Retargeting CD19 Chimeric Antigen Receptor T Cells via Engineered CD19-Fusion Proteins. Molecular Pharmaceutics, 16(8):3544–3558, August 2019. ISSN 1543-8384, 1543-8392. doi: 10.1021/acs.molpharmaceut.9b00418. URL https://pubs.acs.org/doi/10.1021/acs.molpharmaceut.9b00418.
  101. 101.Jannik Kossen, Neil Band, Clare Lyle, Aidan N. Gomez, Tom Rainforth, and Yarin Gal. Self-Attention Between Datapoints: Going Beyond Individual Input-Output Pairs in Deep Learning, February 2022. URL http://arxiv.org/abs/2106.02584. arXiv:2106.02584 [cs, stat] version: 2.
  102. 102.Eran Kotler, Odem Shani, Guy Goldfeld, Maya Lotan-Pompan, Ohad Tarcic, Anat Gershoni, Thomas A. Hopf, Debora S. Marks, Moshe Oren, and Eran Segal. A Systematic p53 Mutation Library Links Differential Functional Impact to Cancer Mutation Pattern and Evolutionary Conservation. Molecular Cell, 71(1):178–190.e8, July 2018. ISSN 10972765. doi: 10.1016/j.molcel.2018.06.012. URL https://linkinghub.elsevier.com/retrieve/pii/S1097276518304544.
  103. 103.Krystian A. Kozek, Andrew M. Glazer, Chai-Ann Ng, Daniel Blackwell, Christian L. Egly, Loren R. Vanags, Marcia Blair, Devyn Mitchell, Kenneth A. Matreyek, Douglas M. Fowler, Bjorn C. Knollmann, Jamie I. Vandenberg, Dan M. Roden, and Brett M. Kroncke. High-throughput discovery of trafficking-deficient variants in the cardiac potassium channel KV11.1. Heart Rhythm, 17(12):2180–2189, December 2020. ISSN 15475271. doi: 10.1016/j.hrthm.2020.05.041. URL https://linkinghub.elsevier.com/retrieve/pii/S1547527120305427.
  104. 104.Andriy Kryshtafovych, Torsten Schwede, Maya Topf, Krzysztof Fidelis, and John Moult. Critical assessment of methods of protein structure prediction (CASP)—Round XIV. Proteins: Structure, 89:1607 – 1617, 2021.
  105. 105.Jason J. Kwon, Behnoush Hajian, Yuemin Bian, Lucy C. Young, Alvaro J. Amor, James R. Fuller, Cara V. Fraley, Abbey M. Sykes, Jonathan So, Joshua Pan, Laura Baker, Sun Joo Lee, Douglas B. Wheeler, David L. Mayhew, Nicole S. Persky, Xiaoping Yang, David E. Root, Anthony M. Barsotti, Andrew W. Stamford, Charles K. Perry, Alex Burgin, Frank McCormick, Christopher T. Lemke, William C. Hahn, and Andrew J. Aguirre. Structure–function analysis of the SHOC2–MRAS–PP1C holophosphatase complex. Nature, 609 (7926):408–415, September 2022. ISSN 0028-0836, 1476-4687. doi: 10.1038/s41586-022-04928-2. URL https://www.nature.com/articles/s41586-022-04928-2.
  106. 106.Elodie Laine, Yasaman Karami, and Alessandra Carbone. GEMME: A Simple and Fast Global Epistatic Model Predicting Mutational Effects. Molecular Biology and Evolution, 36(11):2604–2619, November 2019. ISSN 0737-4038. doi: 10.1093/molbev/msz179. URL https://doi.org/10.1093/molbev/msz179.
  107. 107.Melissa J. Landrum and Brandi L. Kattman. ClinVar at five years: Delivering on the promise. Human Mutation, 39(11):1623–1630, November 2018. ISSN 1098-1004. doi: 10.1002/humu.23641.
  108. 108.Juhye M. Lee, John Huddleston, Michael B. Doud, Kathryn A. Hooper, Nicholas C. Wu, Trevor Bedford, and Jesse D. Bloom. Deep mutational scanning of hemagglutinin helps predict evolutionary fates of human H3N2 influenza variants. Proceedings of the National Academy of Sciences, 115(35), August 2018. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1806133115. URL https://pnas.org/doi/full/10.1073/pnas.1806133115.
  109. 109.Ruipeng Lei, Andrea Hernandez Garcia, Timothy J C Tan, Qi Wen Teo, Yiquan Wang, Xiwen Zhang, Shitong Luo, Satish K Nair, Jian Peng, and Nicholas C Wu. Mutational fitness landscape of human influenza H3N2 neuraminidase. Cell Reports, 42(1):111951, January 2023.
  110. 110.Biao Li, Vidhya G Krishnan, Matthew E Mort, Fuxiao Xin, Kishore K Kamati, David N Cooper, Sean D Mooney, and Predrag Radivojac. Automated inference of molecular mechanisms of disease from amino acid substitutions. Bioinformatics, 25(21):2744–2750, 2009.
  111. 111.Chang Li, Degui Zhi, Kai Wang, and Xiaoming Liu. MetaRNN: differentiating rare pathogenic and rare benign missense SNVs and InDels using deep learning. Genome Medicine, 14(1):115, October 2022. ISSN 1756-994X. doi: 10.1186/s13073-022-01120-z. URL https://genomemedicine.biomedcentral.com/articles/10.1186/s13073-022-01120-z.
  112. 112.Yuan Li, Sarah Arcos, Kimberly R. Sabsay, Aartjan J.W. Te Velthuis, and Adam S. Lauring. Deep mutational scanning reveals the functional constraints and evolutionary potential of the influenza A virus PB1 protein. preprint, Microbiology, August 2023. URL http://biorxiv.org/lookup/doi/10.1101/2023.08.27.554986.
  113. 113.Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Salvatore Candido, and Alexander Rives. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123–1130, March 2023. doi: 10.1126/science.ade2574.
  114. 114.Xiaoming Liu, Xueqiu Jian, and Eric Boerwinkle. dbNSFP: a lightweight database of human nonsynonymous SNPs and their functional predictions. Human mutation, 32(8):894–899, 2011.
  115. 115.Xiaoming Liu, Chang Li, Chengcheng Mou, Yibo Dong, and Yicheng Tu. dbNSFP v4: a comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs. Genome medicine, 12(1):1–8, 2020.
  116. 116.Benjamin J Livesey and Joseph A Marsh. Updated benchmarking of variant effect predictors using deep mutational scanning. Molecular Systems Biology, page e11474, 2023.
  117. 117.Russell S Lo, Gareth A Cromie, Michelle Tang, Kevin Teng, Katherine Owens, Amy Sirr, J Nathan Kutz, Hiroki Morizono, Ljubica Caldovic, Nicholas Ah Mew, Andrea Gropman, and Aimée M Dudley. The functional impact of 1,570 individual amino acid substitutions in human OTC. American Journal of Human Genetics, 110(5):863–879, May 2023.
  118. 118.Christian B. Macdonald, David Nedrud, Patrick Rockefeller Grimes, Donovan Trinidad, James S. Fraser, and Willow Coyote-Maestas. DIMPLE: deep insertion, deletion, and missense mutation libraries for exploring protein variation in evolution, disease, and biology. Genome Biology, 24(1):36, February 2023. ISSN 1474-760X. doi: 10.1186/s13059-023-02880-6. URL https://genomebiology.biomedcentral.com/articles/10.1186/s13059-023-02880-6.
  119. 119.Mark R MacRae, Dhenesh Puvanendran, Max A B Haase, Nicolas Coudray, Ljuvica Kolich, Cherry Lam, Minkyung Baek, Gira Bhabha, and Damian C Ekiert. Protein-protein interactions in the mla lipid transport system probed by computational structure prediction and deep mutational scanning. Journal of Biological Chemistry, 299(6):104744, June 2023.
  120. 120.Ali Madani, Bryan McCann, Nikhil Naik, Nitish Shirish Keskar, Namrata Anand, Raphael R. Eguchi, Po-Ssu Huang, and Richard Socher. ProGen: Language modeling for protein generation, 2020.
  121. 121.Nawar Malhis, Matthew Jacobson, Steven JM Jones, and Jörg Gsponer. LIST-S2: taxonomy based sorting of deleterious missense mutations across species. Nucleic acids research, 48(W1):W154–W161, 2020.
  122. 122.Céline Marquet, Michael Heinzinger, Tobias Olenyi, Christian Dallago, Kyra Erckert, Michael Bernhofer, Dmitrii Nechaev, and Burkhard Rost. Embeddings from protein language models predict conservation and variant effects. Human Genetics, 141(10):1629–1647, October 2022. ISSN 1432-1203. doi: 10.1007/s00439-021-02411-y.
  123. 123.Kenneth A. Matreyek, Lea M. Starita, Jason J. Stephany, Beth Martin, Melissa A. Chiasson, Vanessa E. Gray, Martin Kircher, Arineh Khechaduri, Jennifer N. Dines, Ronald J. Hause, Smita Bhatia, William E. Evans, Mary V. Relling, Wenjian Yang, Jay Shendure, and Douglas M. Fowler. Multiplex assessment of protein variant abundance by massively parallel sequencing. Nature Genetics, 50(6):874–882, June 2018. ISSN 1061-4036, 1546-1718. doi: 10.1038/s41588-018-0122-z. URL https://www.nature.com/articles/s41588-018-0122-z.
  124. 124.Kenneth A. Matreyek, Jason J. Stephany, Ethan Ahler, and Douglas M. Fowler. Integrating thousands of PTEN variant activity and abundance measurements reveals variant subgroups and new dominant negatives in cancers. Genome Medicine, 13(1):165, December 2021. ISSN 1756-994X. doi: 10.1186/s13073-021-00984-x. URL https://genomemedicine.biomedcentral.com/articles/10.1186/s13073-021-00984-x.
  125. 125.Florian Mattenberger, Victor Latorre, Omer Tirosh, Adi Stern, and Ron Geller. Globally defining the effects of mutations in a picornavirus capsid. eLife, 10:e64256, January 2021. ISSN 2050-084X. doi: 10.7554/eLife.64256. URL https://elifesciences.org/articles/64256.
  126. 126.David Mavor, Kyle Barlow, Samuel Thompson, Benjamin A Barad, Alain R Bonny, Clinton L Cario, Garrett Gaskins, Zairan Liu, Laura Deming, Seth D Axen, Elena Caceres, Weilin Chen, Adolfo Cuesta, Rachel E Gate, Evan M Green, Kaitlin R Hulce, Weiyue Ji, Lillian R Kenner, Bruk Mensa, Leanna S Morinishi, Steven M Moss, Marco Mravic, Ryan K Muir, Stefan Niekamp, Chimno I Nnadi, Eugene Palovcak, Erin M Poss, Tyler D Ross, Eugenia C Salcedo, Stephanie K See, Meena Subramaniam, Allison W Wong, Jennifer Li, Kurt S Thorn, Shane Ó Conchúir, Benjamin P Roscoe, Eric D Chow, Joseph L DeRisi, Tanja Kortemme, Daniel N Bolon, and James S Fraser. Determination of ubiquitin fitness landscapes under different chemical stresses in a classroom setting. eLife, 5:e15802, April 2016. ISSN 2050-084X. doi: 10.7554/eLife.15802. URL https://elifesciences.org/articles/15802.
  127. 127.Richard N. McLaughlin, Jr., Frank J. Poelwijk, Arjun Raman, Walraj S. Gosal, and Rama Ranganathan. The spatial architecture of protein function and adaptation. Nature, 491(7422):138–142, November 2012. ISSN 0028-0836, 1476-4687. doi: 10.1038/nature11500. URL https://www.nature.com/articles/nature11500.
  128. 128.Gianmarco Meier, Sujani Thavarasah, Kai Ehrenbolger, Cedric A J Hutter, Lea M Hürlimann, Jonas Barandun, and Markus A Seeger. Deep mutational scan of a drug efflux pump reveals its structure-function landscape. Nature Chemical Biology, 19(4):440–450, April 2023.
  129. 129.Joshua Meier, Roshan Rao, Robert Verkuil, Jason Liu, Tom Sercu, and Alexander Rives. Language models enable zero-shot prediction of the effects of mutations on protein function. bioRxiv, 2021. doi: 10.1101/2021.07.09.450648. URL https://www.biorxiv.org/content/early/2021/07/10/2021.07.09.450648.
  130. 130.Iana Meitlis, Eric J. Allenspach, Bradly M. Bauman, Isabelle Q. Phan, Gina Dabbah, Erica G. Schmitt, Nathan D. Camp, Troy R. Torgerson, Deborah A. Nickerson, Michael J. Bamshad, David Hagin, Christopher R. Luthers, Jeffrey R. Stinson, Jessica Gray, Ingrid Lundgren, Joseph A. Church, Manish J. Butte, Mike B. Jordan, Seema S. Aceves, Daniella M. Schwartz, Joshua D. Milner, Susan Schuval, Suzanne Skoda-Smith, Megan A. Cooper, Lea M. Starita, David J. Rawlings, Andrew L. Snow, and Richard G. James. Multiplexed Functional Assessment of Genetic Variants in CARD11. The American Journal of Human Genetics, 107(6):1029–1043, December 2020. ISSN 00029297. doi: 10.1016/j.ajhg.2020.10.015. URL https://linkinghub.elsevier.com/retrieve/pii/S0002929720303736.
  131. 131.Daniel Melamed, David L. Young, Caitlin E. Gamble, Christina R. Miller, and Stanley Fields. Deep mutational scanning of an RRM domain of the Saccharomyces cerevisiae poly(A)-binding protein. RNA, 19(11): 1537–1551, November 2013. ISSN 1355-8382, 1469-9001. doi: 10.1261/rna.040709.113. URL http://rnajournal.cshlp.org/lookup/doi/10.1261/rna.040709.113.
  132. 132.Alexandre Melnikov, Peter Rogov, Li Wang, Andreas Gnirke, and Tarjei S. Mikkelsen. Comprehensive mutational scanning of a kinase in vivo reveals substrate-dependent fitness landscapes. Nucleic Acids Research, 42(14):e112–e112, August 2014. ISSN 0305-1048, 1362-4962. doi: 10.1093/nar/gku511. URL https://academic.oup.com/nar/article-lookup/doi/10.1093/nar/gku511.
  133. 133.Taylor L. Mighell, Sara Evans-Dutson, and Brian J. O’Roak. A Saturation Mutagenesis Approach to Understanding PTEN Lipid Phosphatase Activity and Genotype-Phenotype Relationships. The American Journal of Human Genetics, 102(5):943–955, May 2018. ISSN 00029297. doi: 10.1016/j.ajhg.2018.03.018. URL https://linkinghub.elsevier.com/retrieve/pii/S0002929718301071.
  134. 134.Peter G. Miller, Murugappan Sathappa, Jamie A. Moroco, Wei Jiang, Yue Qian, Sumaiya Iqbal, Qi Guo, Andrew O. Giacomelli, Subrata Shaw, Camille Vernier, Besnik Bajrami, Xiaoping Yang, Cerise Raffier, Adam S. Sperling, Christopher J. Gibson, Josephine Kahn, Cyrus Jin, Matthew Ranaghan, Alisha Caliman, Merissa Brousseau, Eric S. Fischer, Robert Lintner, Federica Piccioni, Arthur J. Campbell, David E. Root, Colin W. Garvie, and Benjamin L. Ebert. Allosteric inhibition of PPM1D serine/threonine phosphatase via an altered conformational state. Nature Communications, 13(1):3778, June 2022. ISSN 2041-1723. doi: 10.1038/s41467-022-30463-9. URL https://www.nature.com/articles/s41467-022-30463-9.
  135. 135.Parul Mishra, Julia M. Flynn, Tyler N. Starr, and Daniel N.A. Bolon. Systematic Mutant Analyses Elucidate General and Client-Specific Aspects of Hsp90 Function. Cell Reports, 15(3):588–598, April 2016. ISSN 22111247. doi: 10.1016/j.celrep.2016.03.046. URL https://linkinghub.elsevier.com/retrieve/pii/S2211124716303175.
  136. 136.Iain H. Moal and Juan Fernández-Recio. SKEMPI: a structural kinetic and energetic database of mutant protein interactions and its use in empirical models. Bioinformatics, 28 20:2600–7, 2012.
  137. 137.Ayesha Muhammad, Maria E Calandranis, Bian Li, Tao Yang, Daniel J Blackwell, M Lorena Harvey, Jeremy E Smith, Ashli E Chew, John A Capra, Kenneth A Matreyek, Douglas M Fowler, Dan M Roden, and Andrew M Glazer. High-throughput functional mapping of variants in an arrhythmia gene, KCNE1 , reveals novel biology. bioRxiv, April 2023.
  138. 138.Robert W. Newberry, Taylor Arhar, Jean Costello, George C. Hartoularos, Alison M. Maxwell, Zun Zar Chi Naing, Maureen Pittman, Nishith R. Reddy, Daniel M. C. Schwarz, Douglas R. Wassarman, Taia S. Wu, Daniel Barrero, Christa Caggiano, Adam Catching, Taylor B. Cavazos, Laurel S. Estes, Bryan Faust, Elissa A. Fink, Miriam A. Goldman, Yessica K. Gomez, M. Grace Gordon, Laura M. Gunsalus, Nick Hoppe, Maru Jaime-Garza, Matthew C. Johnson, Matthew G. Jones, Andrew F. Kung, Kyle E. Lopez, Jared Lumpe, Calla Martyn, Elizabeth E. McCarthy, Lakshmi E. Miller-Vedam, Erik J. Navarro, Aji Palar, Jenna Pellegrino, Wren Saylor, Christina A. Stephens, Jack Strickland, Hayarpi Torosyan, Stephanie A. Wankowicz, Daniel R. Wong, Garrett Wong, Sy Redding, Eric D. Chow, William F. DeGrado, and Martin Kampmann. Robust Sequence Determinants of alpha-Synuclein Toxicity in Yeast Implicate Membrane Binding. ACS Chemical Biology, 15(8):2137–2153, August 2020. ISSN 1554-8929, 1554-8937. doi: 10.1021/acschembio.0c00339. URL https://pubs.acs.org/doi/10.1021/acschembio.0c00339.
  139. 139.Pauline C Ng and Steven Henikoff. Accounting for human polymorphisms predicted to affect protein function. Genome Res., 12(3):436–446, March 2002.
  140. 140.Thuy N Nguyen, Christine Ingle, Samuel Thompson, and Kimberly A Reynolds. The genetic landscape of a metabolic interaction. May 2023a.
  141. 141.Vanessa Nguyen, Ethan Ahler, Katherine A Sitko, Jason J Stephany, Dustin J Maly, and Douglas M Fowler. Molecular determinants of hsp90 dependence of src kinase revealed by deep mutational scanning. Protein Science, 32(7):e4656, July 2023b.
  142. 142.Erik Nijkamp, Jeffrey A. Ruffolo, Eli N. Weinstein, Nikhil Naik, and Ali Madani. ProGen2: Exploring the boundaries of protein language models. ArXiv, abs/2206.13517, 2022.
  143. 143.Pascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado, Aidan N. Gomez, Debora S. Marks, and Yarin Gal. Tranception: protein fitness prediction with autoregressive transformers and inference-time retrieval. In ICML, 2022a.
  144. 144.Pascal Notin, Lood Van Niekerk, Aaron W. Kollasch, Daniel Ritter, Yarin Gal, and Debora Susan Marks. TranceptEVE: Combining Family-specific and Family-agnostic Models of Protein Sequences for Improved Fitness Prediction. December 2022b. URL https://openreview.net/forum?id=l7Oo9DcLmR1.
  145. 145.Pascal Notin, Ruben Weitzman, Debora S. Marks, and Yarin Gal. Proteinnpt: Improving protein property prediction and design with non-parametric transformers. Advances in Neural Information Processing Systems, 37, 2023.
  146. 146.Christina Nutschel, Alexander Fulton, Olav Zimmermann, Ulrich Schwaneberg, Karl-Erich Jaeger, and Holger Gohlke. Systematically Scrutinizing the Impact of Substitution Sites on Thermostability and Detergent Tolerance for Bacillus subtilis Lipase A. Journal of Chemical Information and Modeling, 60(3):1568–1584, March 2020. ISSN 1549-9596, 1549-960X. doi: 10.1021/acs.jcim.9b00954. URL https://pubs.acs.org/doi/10.1021/acs.jcim.9b00954.
  147. 147.C. Anders Olson, Nicholas C. Wu, and Ren Sun. A Comprehensive Biophysical Description of Pairwise Epistasis throughout an Entire Protein Domain. Current Biology, 24(22):2643–2651, November 2014. ISSN 09609822. doi: 10.1016/j.cub.2014.09.072. URL https://linkinghub.elsevier.com/retrieve/pii/S0960982214012688.
  148. 148.Martin K Ostermaier, Christian Peterhans, Rolf Jaussi, Xavier Deupi, and Jörg Standfuss. Functional map of arrestin-1 at single amino acid resolution. Proceedings of the National Academy of Sciences, 111(5): 1825–1830, February 2014.
  149. 149.Vikas Pejaver, Alicia B Byrne, Bing-Jian Feng, Kymberleigh A Pagel, Sean D Mooney, Rachel Karchin, Anne O’Donnell-Luria, Steven M Harrison, Sean V Tavtigian, Marc S Greenblatt, et al. Calibration of computational tools for missense variant pathogenicity classification and clingen recommendations for pp3/bp4 criteria. The American Journal of Human Genetics, 109(12):2163–2177, 2022.
  150. 150.Victoria O. Pokusaeva, Dinara R. Usmanova, Ekaterina V. Putintseva, Lorena Espinar, Karen S. Sarkisyan, Alexander S. Mishin, Natalya S. Bogatyreva, Dmitry N. Ivankov, Arseniy V. Akopyan, Sergey Ya. Avvakumov, Inna S. Povolotskaya, Guillaume J. Filion, Lucas B. Carey, and Fyodor A. Kondrashov. An experimental assay of the interactions of amino acids from orthologous sequences shaping a complex fitness landscape. PLOS Genetics, 15(4):e1008079, April 2019. ISSN 1553-7404. doi: 10.1371/journal.pgen.1008079. URL https://dx.plos.org/10.1371/journal.pgen.1008079.
  151. 151.Hangfei Qi, C. Anders Olson, Nicholas C. Wu, Ruian Ke, Claude Loverdo, Virginia Chu, Shawna Truong, Roland Remenyi, Zugen Chen, Yushen Du, Sheng-Yao Su, Laith Q. Al-Mawsawi, Ting-Ting Wu, Shu-Hua Chen, Chung-Yen Lin, Weidong Zhong, James O. Lloyd-Smith, and Ren Sun. A Quantitative High-Resolution Genetic Profile Rapidly Identifies Sequence Determinants of Hepatitis C Viral Fitness and Drug Sensitivity. PLoS Pathogens, 10(4):e1004064, April 2014. ISSN 1553-7374. doi: 10.1371/journal.ppat.1004064. URL https://dx.plos.org/10.1371/journal.ppat.1004064.
  152. 152.Daniel Quang, Yifei Chen, and Xiaohui Xie. DANN: a deep learning approach for annotating the pathogenicity of genetic variants. Bioinformatics, 31(5):761–763, 2015.
  153. 153.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019. URL https://api.semanticscholar.org/CorpusID:160025533.
  154. 154.Daniele Raimondi, Ibrahim Tanyalcin, Julien Ferté, Andrea Gazzo, Gabriele Orlando, Tom Lenaerts, Marianne Rooman, and Wim Vranken. DEOGEN2: prediction and interactive visualization of single amino acid variant deleteriousness in human proteins. Nucleic acids research, 45(W1):W201–W206, 2017.
  155. 155.Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Xi Chen, John Canny, Pieter Abbeel, and Yun S. Song. Evaluating Protein Transfer Learning with TAPE, June 2019. URL http://arxiv.org/abs/1906.08230. arXiv:1906.08230 [cs, q-bio, stat].
  156. 156.Roshan Rao, Jason Liu, Robert Verkuil, Joshua Meier, John F. Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives. MSA transformer. bioRxiv, 2021. doi: 10.1101/2021.02.12.430858. URL https://www.biorxiv.org/content/early/2021/02/13/2021.02.12.430858.
  157. 157.Philipp Rentzsch, Daniela Witten, Gregory M Cooper, Jay Shendure, and Martin Kircher. Cadd: predicting the deleteriousness of variants throughout the human genome. Nucleic acids research, 47(D1):D886–D894, 2019.
  158. 158.Boris Reva, Yevgeniy Antipin, and Chris Sander. Predicting the functional impact of protein mutations: application to cancer genomics. Nucleic Acids Research, 39(17):e118, September 2011. ISSN 1362-4962. doi: 10.1093/nar/gkr407.
  159. 159.Adam J Riesselman, John B Ingraham, and Debora S Marks. Deep generative models of genetic variation capture the effects of mutations. Nature Methods, 15(10):816–822, 2018.
  160. 160.Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences, 118(15), 2021.
  161. 161.Liat Rockah-Shmuel, Ágnes Tóth-Petróczy, and Dan S. Tawfik. Systematic Mapping of Protein Mutational Space by Prolonged Drift Reveals the Deleterious Effects of Seemingly Neutral Mutations. PLOS Computational Biology, 11(8):e1004421, August 2015. ISSN 1553-7358. doi: 10.1371/journal.pcbi.1004421. URL https://dx.plos.org/10.1371/journal.pcbi.1004421.
  162. 162.Philip A. Romero, Andreas Krause, and Frances H. Arnold. Navigating the protein fitness landscape with gaussian processes. Proceedings of the National Academy of Sciences, 110:E193 – E201, 2012. URL https://api.semanticscholar.org/CorpusID:1093192.
  163. 163.Philip A. Romero, Tuan M. Tran, and Adam R. Abate. Dissecting enzyme function with microfluidic-based deep mutational scanning. Proceedings of the National Academy of Sciences, 112(23):7159–7164, June 2015. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1422285112. URL https://pnas.org/doi/full/10.1073/pnas.1422285112.
  164. 164.Benjamin P. Roscoe and Daniel N.A. Bolon. Systematic Exploration of Ubiquitin Sequence, E1 Activation Efficiency, and Experimental Fitness in Yeast. Journal of Molecular Biology, 426(15):2854–2870, July 2014. ISSN 00222836. doi: 10.1016/j.jmb.2014.05.019. URL https://linkinghub.elsevier.com/retrieve/pii/S0022283614002587.
  165. 165.Benjamin P. Roscoe, Kelly M. Thayer, Konstantin B. Zeldovich, David Fushman, and Daniel N.A. Bolon. Analyses of the Effects of All Ubiquitin Point Mutants on Yeast Growth Rate. Journal of Molecular Biology, 425(8):1363–1377, April 2013. ISSN 00222836. doi: 10.1016/j.jmb.2013.01.032. URL https://linkinghub.elsevier.com/retrieve/pii/S0022283613000636.
  166. 166.Hridindu Roychowdhury and Philip A Romero. Microfluidic deep mutational scanning of the human executioner caspases reveals differences in structure and regulation. Cell Death Discovery, 8(1):7, January 2022.
  167. 167.Alan F. Rubin, Joseph K Min, Nathan J. Rollins, Estelle Y Da, Daniel Esposito, Matthew Harrington, Jeremy Stone, Aisha Haley Bianchi, Mafalda Dias, Jonathan Frazer, Yunfan Fu, Molly Gallaher, Iris Li, Olivia Moscatelli, Jesslyn YL Ong, Joshua E Rollins, Matthew J. Wakefield, Shenyi “Sunny” Ye, Amy Sze Pui Tam, Abbye E. McEwen, Lea M. Starita, Vanessa L. Bryant, Debora S. Marks, and Douglas M. Fowler. MaveDB v2: a curated community database with over three million variant effects from multiplexed functional assays. bioRxiv, 2021.
  168. 168.William P. Russ, Matteo Figliuzzi, Christian Stocker, Pierre Barrat-Charlaix, Michael Socolich, Peter Kast, Donald Hilvert, Remi Monasson, Simona Cocco, Martin Weigt, and Rama Ranganathan. An evolution-based model for designing chorismate mutase enzymes. Science, 369(6502):440–445, July 2020. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.aba3304. URL https://www.sciencemag.org/lookup/doi/10.1126/science.aba3304.
  169. 169.Kaitlin E Samocha, Jack A Kosmicki, Konrad J Karczewski, Anne H O’Donnell-Luria, Emma Pierce-Hoffman, Daniel G MacArthur, Benjamin M Neale, and Mark J Daly. Regional missense constraint improves variant deleteriousness prediction. BioRxiv, page 148353, 2017.
  170. 170.Karen S. Sarkisyan, Dmitry A. Bolotin, Margarita V. Meer, Dinara R. Usmanova, Alexander S. Mishin, George V. Sharonov, Dmitry N. Ivankov, Nina G. Bozhanova, Mikhail S. Baranov, Onuralp Soylemez, Natalya S. Bogatyreva, Peter K. Vlasov, Evgeny S. Egorov, Maria D. Logacheva, Alexey S. Kondrashov, Dmitry M. Chudakov, Ekaterina V. Putintseva, Ilgar Z. Mamedov, Dan S. Tawfik, Konstantin A. Lukyanov, and Fyodor A. Kondrashov. Local fitness landscape of the green fluorescent protein. Nature, 533(7603):397–401, May 2016. ISSN 0028-0836, 1476-4687. doi: 10.1038/nature17995. URL http://www.nature.com/articles/nature17995.
  171. 171.Jana Marie Schwarz, Christian Rödelsperger, Markus Schuelke, and Dominik Seelow. MutationTaster evaluates disease-causing potential of sequence alterations. Nature methods, 7(8):575–576, 2010.
  172. 172.Mireia Seuma, Andre J Faure, Marta Badia, Ben Lehner, and Benedetta Bolognesi. The genetic landscape for amyloid beta fibril nucleation accurately discriminates familial Alzheimer’s disease mutations. eLife, 10: e63364, February 2021. ISSN 2050-084X. doi: 10.7554/eLife.63364. URL https://elifesciences.org/articles/63364.
  173. 173.Mireia Seuma, Ben Lehner, and Benedetta Bolognesi. An atlas of amyloid aggregation: the impact of substitutions, insertions, deletions and truncations on amyloid beta fibril nucleation. Nature Communications, 13(1): 7084, November 2022.
  174. 174.Hashem A Shihab, Julian Gough, David N Cooper, Peter D Stenson, Gary LA Barker, Keith J Edwards, Ian NM Day, and Tom R Gaunt. Predicting the functional, molecular, and phenotypic consequences of amino acid substitutions using hidden markov models. Human mutation, 34(1):57–65, 2013.
  175. 175.Jung-Eun Shin, Adam J Riesselman, Aaron W Kollasch, Conor McMahon, Elana Simon, Chris Sander, Aashish Manglik, Andrew C Kruse, and Debora S Marks. Protein design and variant prediction using autoregressive generative models. Nature communications, 12(1):1–11, 2021.
  176. 176.D Shortle and J Sondek. The emerging role of insertions and deletions in protein engineering. Curr. Opin. Biotechnol., 6(4):387–393, August 1995.
  177. 177.Rachel A. Silverstein, Song Sun, Marta Verby, Jochen Weile, Yingzhou Wu, Marinella Gebbia, Iosifina Fotiadou, Julia Kitaygorodsky, and Frederick P. Roth. A systematic genotype-phenotype map for missense variants in the human intellectual disability-associated gene GDI1. preprint, Genetics, October 2021. URL http://biorxiv.org/lookup/doi/10.1101/2021.10.06.463360.
  178. 178.Sam Sinai, Nina Jain, George M Church, and Eric D Kelsic. Generative AAV capsid diversification by latent interpolation. preprint, Synthetic Biology, April 2021. URL http://biorxiv.org/lookup/doi/10.1101/2021.04.16.440236.
  179. 179.Yq Shirleen Soh, Louise H Moncla, Rachel Eguia, Trevor Bedford, and Jesse D Bloom. Comprehensive mapping of adaptation of the avian influenza polymerase protein PB2 to humans. eLife, 8:e45079, April 2019. ISSN 2050-084X. doi: 10.7554/eLife.45079. URL https://elifesciences.org/articles/45079.
  180. 180.Marion Sourisseau, Daniel J. P. Lawrence, Megan C. Schwarz, Carina H. Storrs, Ethan C. Veit, Jesse D. Bloom, and Matthew J. Evans. Deep Mutational Scanning Comprehensively Maps How Zika Envelope Protein Mutations Affect Viral Growth and Antibody Escape. Journal of Virology, 93(23):e01291–19, December 2019. ISSN 0022-538X, 1098-5514. doi: 10.1128/JVI.01291-19. URL https://journals.asm.org/doi/10.1128/JVI.01291-19.
  181. 181.Jeffrey M. Spencer and Xiaoliu Zhang. Deep mutational scanning of S. pyogenes Cas9 reveals important functional domains. Scientific Reports, 7(1):16836, December 2017. ISSN 2045-2322. doi: 10.1038/s41588-017-17081-y. URL https://www.nature.com/articles/s41598-017-17081-y.
  182. 182.Tobias Stadelmann, Daniel Heid, Michael Jendrusch, Jan Mathony, Stéphane Rosset, Bruno E. Correia, and Dominik Niopek. A deep mutational scanning platform to characterize the fitness landscape of anti-CRISPR proteins. preprint, Synthetic Biology, August 2021. URL http://biorxiv.org/lookup/doi/10.1101/2021.08.21.457204.
  183. 183.Max V. Staller, Alex S. Holehouse, Devjanee Swain-Lenz, Rahul K. Das, Rohit V. Pappu, and Barak A. Cohen. A High-Throughput Mutational Scan of an Intrinsically Disordered Acidic Transcriptional Activation Domain. Cell Systems, 6(4):444–455.e6, April 2018. ISSN 24054712. doi: 10.1016/j.cels.2018.01.015. URL https://linkinghub.elsevier.com/retrieve/pii/S2405471218300528.
  184. 184.Lea M. Starita, Jonathan N. Pruneda, Russell S. Lo, Douglas M. Fowler, Helen J. Kim, Joseph B. Hiatt, Jay Shendure, Peter S. Brzovic, Stanley Fields, and Rachel E. Klevit. Activity-enhancing mutations in an E3 ubiquitin ligase identified by high-throughput mutagenesis. Proceedings of the National Academy of Sciences, 110(14), April 2013. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1303309110. URL https://pnas.org/doi/full/10.1073/pnas.1303309110.
  185. 185.Tyler N. Starr, Allison J. Greaney, Sarah K. Hilton, Daniel Ellis, Katharine H.D. Crawford, Adam S. Dingens, Mary Jane Navarro, John E. Bowen, M. Alejandra Tortorici, Alexandra C. Walls, Neil P. King, David Veesler, and Jesse D. Bloom. Deep Mutational Scanning of SARS-CoV-2 Receptor Binding Domain Reveals Constraints on Folding and ACE2 Binding. Cell, 182(5):1295–1310.e20, September 2020. ISSN 00928674. doi: 10.1016/j.cell.2020.08.012. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867420310035.
  186. 186.Martin Steinegger and Johannes Söding. Clustering huge protein sequence sets in linear time. Nature Communications, 9(1):2542, Jun 2018. ISSN 2041-1723. doi: 10.1038/s41467-018-04964-5. URL https://doi.org/10.1038/s41467-018-04964-5.
  187. 187.Michael A. Stiffler, Doeke R. Hekstra, and Rama Ranganathan. Evolvability as a Function of Purifying Selection in TEM-1 beta-Lactamase. Cell, 160(5):882–892, February 2015. ISSN 00928674. doi: 10.1016/j.cell.2015.01.035. URL https://linkinghub.elsevier.com/retrieve/pii/S0092867415000781.
  188. 188.Jan Stourac, Juraj Dúbrava, Miloš Musil, Jana Horácková, Ji ˇ ˇrí Damborský, S. Mazurenko, and David Bednáˇr. FireProtDB: database of manually curated protein stability data. Nucleic Acids Research, 49:D319 – D324, 2020.
  189. 189.Chase C. Suiter, Takaya Moriyama, Kenneth A. Matreyek, Wentao Yang, Emma Rose Scaletti, Rina Nishii, Wenjian Yang, Keito Hoshitsuki, Minu Singh, Amita Trehan, Chris Parish, Colton Smith, Lie Li, Deepa Bhojwani, Liz Y. P. Yuen, Chi-kong Li, Chak-ho Li, Yung-li Yang, Gareth J. Walker, James R. Goodhand, Nicholas A. Kennedy, Federico Antillon Klussmann, Smita Bhatia, Mary V. Relling, Motohiro Kato, Hiroki Hori, Prateek Bhatia, Tariq Ahmad, Allen E. J. Yeoh, Pål Stenmark, Douglas M. Fowler, and Jun J. Yang. Massively parallel variant characterization identifies NUDT15 alleles associated with thiopurine toxicity. Proceedings of the National Academy of Sciences, 117(10):5394–5401, March 2020. ISSN 0027-8424, 1091-6490. doi: 10.1073/pnas.1915680117. URL https://pnas.org/doi/full/10.1073/pnas.1915680117.
  190. 190.Song Sun, Jochen Weile, Marta Verby, Yingzhou Wu, Yang Wang, Atina G. Cote, Iosifina Fotiadou, Julia Kitaygorodsky, Marc Vidal, Jasper Rine, Pavel Ješina, Viktor Kožich, and Frederick P. Roth. A proactive genotypeto-patient-phenotype map for cystathionine beta-synthase. Genome Medicine, 12(1):13, December 2020. ISSN 1756-994X. doi: 10.1186/s13073-020-0711-1. URL https://genomemedicine.biomedcentral.com/articles/10.1186/s13073-020-0711-1.
  191. 191.Laksshman Sundaram, Hong Gao, Samskruthi Reddy Padigepati, Jeremy F McRae, Yanjun Li, Jack A Kosmicki, Nondas Fritzilas, Jörg Hakenberg, Anindita Dutta, John Shon, et al. Predicting the clinical impact of human mutation with deep neural networks. Nature genetics, 50(8):1161–1170, 2018.
  192. 192.Amporn Suphatrakul, Pratsaneeyaporn Posiri, Nittaya Srisuk, Rapirat Nantachokchawapan, Suppachoke Onnome, Juthathip Mongkolsapaya, and Bunpote Siridechadilok. Functional analysis of flavivirus replicase by deep mutational scanning of dengue NS5. March 2023.
  193. 193.Baris E. Suzek, Yuqi Wang, Hongzhan Huang, Peter B. McGarvey, and Cathy H. Wu. UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics, 31:926 – 932, 2014.
  194. 194.Baris E Suzek, Yuqi Wang, Hongzhan Huang, Peter B McGarvey, Cathy H Wu, and UniProt Consortium. UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics, 31(6):926–932, 2015.
  195. 195.Timothy J C Tan, Zongjun Mou, Ruipeng Lei, Wenhao O Ouyang, Meng Yuan, Ge Song, Raiees Andrabi, Ian A Wilson, Collin Kieffer, Xinghong Dai, Kenneth A Matreyek, and Nicholas C Wu. High-throughput identification of prefusion-stabilizing mutations in SARS-CoV-2 spike. Nature Communications, 14(1):2003, April 2023.
  196. 196.Samuel Thompson, Yang Zhang, Christine Ingle, Kimberly A Reynolds, and Tanja Kortemme. Altered expression of a quality control protease in E. coli reshapes the in vivo mutational landscape of a model enzyme. eLife, 9: e53476, July 2020. ISSN 2050-084X. doi: 10.7554/eLife.53476. URL https://elifesciences.org/articles/53476.
  197. 197.Bargavi Thyagarajan and Jesse D Bloom. The inherent mutational tolerance and antigenic evolvability of influenza hemagglutinin. Elife, 3, July 2014.
  198. 198.Agnes Tóth-Petróczy and Dan S Tawfik. Protein insertions and deletions enabled by neutral roaming in sequence space. Mol. Biol. Evol., 30(4):761–771, April 2013.
  199. 199.Arti Tripathi, Kritika Gupta, Shruti Khare, Pankaj C. Jain, Siddharth Patel, Prasanth Kumar, Ajai J. Pulianmackal, Nilesh Aghera, and Raghavan Varadarajan. Molecular Determinants of Mutant Phenotypes, Inferred from Saturation Mutagenesis Data. Molecular Biology and Evolution, 33(11):2960–2975, November 2016. ISSN 0737-4038, 1537-1719. doi: 10.1093/molbev/msw182. URL https://academic.oup.com/mbe/article-lookup/doi/10.1093/molbev/msw182.
  200. 200.Kotaro Tsuboyama, Justas Dauparas, Jonathan Chen, Elodie Laine, Yasser Mohseni Behbahani, Jonathan J. Weinstein, Niall M. Mangan, Sergey Ovchinnikov, and Gabriel J. Rocklin. Mega-scale experimental analysis of protein folding stability in biology and design. Nature, 620(7973):434–444, August 2023. ISSN 0028-0836, 1476-4687. doi: 10.1038/s41586-023-06328-6. URL https://www.nature.com/articles/s41586-023-06328-6.
  201. 201.UK Monogenic Diabetes Consortium, Myocardial Infarction Genetics Consortium, UK Congenital Lipodystrophy Consortium, Amit R Majithia, Ben Tsuda, Maura Agostini, Keerthana Gnanapradeepan, Robert Rice, Gina Peloso, Kashyap A Patel, Xiaolan Zhang, Marjoleine F Broekema, Nick Patterson, Marc Duby, Ted Sharpe, Eric Kalkhoven, Evan D Rosen, Inês Barroso, Sian Ellard, Sekar Kathiresan, Stephen O’Rahilly, Krishna Chatterjee, Jose C Florez, Tarjei Mikkelsen, David B Savage, and David Altshuler. Prospective functional classification of all possible missense variants in PPARG. Nature Genetics, 48(12):1570–1575, December 2016. ISSN 1061-4036, 1546-1718. doi: 10.1038/ng.3700. URL https://www.nature.com/articles/ng.3700.
  202. 202.Fabio Urbina, Filippa Lentzos, Cédric Invernizzi, and Sean Ekins. Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence, 4(3):189–191, 2022.
  203. 203.Oana Ursu, James T Neal, Emily Shea, Pratiksha I Thakore, Livnat Jerby-Arnon, Lan Nguyen, Danielle Dionne, Celeste Diaz, Julia Bauman, Mariam Mounir Mosaad, Christian Fagre, April Lo, Maria McSharry, Andrew O Giacomelli, Seav Huong Ly, Orit Rozenblatt-Rosen, William C Hahn, Andrew J Aguirre, Alice H Berger, Aviv Regev, and Jesse S Boehm. Massively parallel phenotyping of coding variants in cancer with perturb-seq. Nature Biotechnology, 40(6):896–905, June 2022.
  204. 204.Warren van Loggerenberg, Shahin Sowlati-Hashjin, Jochen Weile, Rayna Hamilton, Aditya Chawla, Marinella Gebbia, Nishka Kishore, Laure Frésard, Sami Mustajoki, Elena Pischik, Elena Di Pierro, Michela Barbaro, Ylva Floderus, Caroline Schmitt, Laurent Gouya, Alexandre Colavin, Robert Nussbaum, Edith C H Friesema, Raili Kauppinen, Jordi To-Figueras, Aasne K Aarsand, Robert J Desnick, Michael Garton, and Frederick P Roth. Systematically testing human HMBS missense variants to reveal mechanism and pathogenic variation. bioRxiv, February 2023.
  205. 205.Rosario Vanella, Christoph Küng, Alexandre A Schoepfer, Vanni Doffini, Jin Ren, and Michael A Nash. Understanding Activity-Stability tradeoffs in biocatalysts by enzyme proximity sequencing. March 2023.
  206. 206.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2017.
  207. 207.Veeramohan Veerapandian, Jan Ole Ackermann, Yogesh Srivastava, Vikas Malik, Mingxi Weng, Xiaoxiao Yang, and Ralf Jauch. Directed evolution of reprogramming factors by cell selection and sequencing. Stem Cell Reports, 11(2):593–606, August 2018.
  208. 208.Aliete Wan, Emily Place, Eric A. Pierce, and Jason Comander. Characterizing variants of unknown significance in rhodopsin: A functional genomics approach. Human Mutation, 40(8):1127–1144, August 2019. ISSN 1059-7794, 1098-1004. doi: 10.1002/humu.23762. URL https://onlinelibrary.wiley.com/doi/10.1002/humu.23762.
  209. 209.Ryan Weeks and Marc Ostermeier. Fitness and functional landscapes of the e. coli RNase III gene rnc. Molecular Biology and Evolution, 40(3), March 2023.
  210. 210.Jochen Weile, Song Sun, Atina G Cote, Jennifer Knapp, Marta Verby, Joseph C Mellor, Yingzhou Wu, Carles Pons, Cassandra Wong, Natascha Lieshout, Fan Yang, Murat Tasan, Guihong Tan, Shan Yang, Douglas M Fowler, Robert Nussbaum, Jesse D Bloom, Marc Vidal, David E Hill, Patrick Aloy, and Frederick P Roth. A framework for exhaustively mapping functional missense variants. Molecular Systems Biology, 13(12):957, December 2017. ISSN 1744-4292, 1744-4292. doi: 10.15252/msb.20177908. URL https://onlinelibrary.wiley.com/doi/10.15252/msb.20177908.
  211. 211.Jochen Weile, Nishka Kishore, Song Sun, Ranim Maaieh, Marta Verby, Roujia Li, Iosifina Fotiadou, Julia Kitaygorodsky, Yingzhou Wu, Alexander Holenstein, Céline Bürer, Linnea Blomgren, Shan Yang, Robert Nussbaum, Rima Rozen, David Watkins, Marinella Gebbia, Viktor Kozich, Michael Garton, D Sean Froese, and Frederick P Roth. Shifting landscapes of human MTHFR missense-variant effects. American Journal of Human Genetics, 108(7):1283–1300, July 2021.
  212. 212.Chenchun Weng, Andre J Faure, and Ben Lehner. The energetic and allosteric landscape for KRAS inhibition. December 2022.
  213. 213.Emily E. Wrenbeck, Laura R. Azouz, and Timothy A. Whitehead. Single-mutation fitness landscapes for an enzyme on multiple substrates reveal specificity is globally encoded. Nature Communications, 8(1):15695, August 2017. ISSN 2041-1723. doi: 10.1038/ncomms15695. URL http://www.nature.com/articles/ncomms15695.
  214. 214.Emily E Wrenbeck, Matthew A Bedewitz, Justin R Klesmith, Syeda Noshin, Cornelius S Barry, and Timothy A Whitehead. An automated Data-Driven pipeline for improving heterologous enzyme expression. ACS Synthetic Biology, 8(3):474–481, March 2019.
  215. 215.Nicholas C. Wu, Arthur P. Young, Laith Q. Al-Mawsawi, C. Anders Olson, Jun Feng, Hangfei Qi, Shu-Hwa Chen, I.-Hsuan Lu, Chung-Yen Lin, Robert G. Chin, Harding H. Luan, Nguyen Nguyen, Stanley F. Nelson, Xinmin Li, Ting-Ting Wu, and Ren Sun. High-throughput profiling of influenza A virus hemagglutinin gene at single-nucleotide resolution. Scientific Reports, 4(1):4942, December 2014. ISSN 2045-2322. doi: 10.1038/srep04942. URL https://www.nature.com/articles/srep04942.
  216. 216.Nicholas C. Wu, C. Anders Olson, Yushen Du, Shuai Le, Kevin Tran, Roland Remenyi, Danyang Gong, Laith Q. Al-Mawsawi, Hangfei Qi, Ting-Ting Wu, and Ren Sun. Functional Constraint Profiling of a Viral Protein Reveals Discordance of Evolutionary Conservation and Functionality. PLOS Genetics, 11(7):e1005310, July 2015. ISSN 1553-7404. doi: 10.1371/journal.pgen.1005310. URL https://dx.plos.org/10.1371/journal.pgen.1005310.
  217. 217.Nicholas C Wu, Lei Dai, C Anders Olson, James O Lloyd-Smith, and Ren Sun. Adaptation in protein fitness landscapes is facilitated by indirect paths. eLife, 5:e16965, July 2016. ISSN 2050-084X. doi: 10.7554/eLife.16965. URL https://elifesciences.org/articles/16965.
  218. 218.Yingzhou Wu, Hanqing Liu, Roujia Li, Song Sun, Jochen Weile, and Frederick P Roth. Improved pathogenicity prediction for rare human missense variants. The American Journal of Human Genetics, 108(10):1891–1906, 2021.
  219. 219.Zachary Wu, S. B. Jennifer Kan, Russell D. Lewis, Bruce J. Wittmann, and Frances H. Arnold. Machine learning-assisted directed protein evolution with combinatorial libraries. Proceedings of the National Academy of Sciences, 116:8852 – 8858, 2019. URL https://api.semanticscholar.org/CorpusID:67770057.
  220. 220.Michael J Xie, Gareth A Cromie, Katherine Owens, Martin S Timour, Michelle Tang, J Nathan Kutz, Ayman W El-Hattab, Richard N McLaughlin, and Aimée M Dudley. Predicting the functional effect of compound heterozygous genotypes from large scale variant effect maps. bioRxiv, January 2023.
  221. 221.Minghao Xu, Zuobai Zhang, Jiarui Lu, Zhaocheng Zhu, Yangtian Zhang, Chang Ma, Runcheng Liu, and Jian Tang. PEER: A Comprehensive and Multi-Task Benchmark for Protein Sequence Understanding, September 2022. URL http://arxiv.org/abs/2206.02096. arXiv:2206.02096 [cs].
  222. 222.Kevin Kaichuang Yang, Zachary Wu, and Frances H. Arnold. Machine-learning-guided directed evolution for protein engineering. Nature Methods, pages 1–8, 2018. URL https://api.semanticscholar.org/CorpusID:128342395.
  223. 223.Kevin Kaichuang Yang, Alex X. Lu, and Nicoló Fusi. Convolutions are competitive with transformers for protein sequence pretraining. bioRxiv, 2023a. URL https://api.semanticscholar.org/CorpusID:248990392.
  224. 224.Kevin Kaichuang Yang, Niccoló Zanichelli, and Hugh Yeh. Masked inverse folding with sequence transfer for protein representation learning. bioRxiv, 2023b. URL https://api.semanticscholar.org/CorpusID:249241961.
  225. 225.Sook Wah Yee, Christian Macdonald, Darko Mitrovic, Xujia Zhou, Megan L Koleske, Jia Yang, Dina Buitrago Silva, Patrick Rockefeller Grimes, Donovan Trinidad, Swati S More, Linda Kachuri, John S Witte, Lucie Delemotte, Kathleen M Giacomini, and Willow Coyote-Maestas. The full spectrum of OCT1 (SLC22A1) mutations bridges transporter biophysics to drug pharmacogenomics. bioRxiv, June 2023.
  226. 226.Heather J. Young, Matthew Chan, Balaji Selvam, Steven K. Szymanski, Diwakar Shukla, and Erik Procko. Deep Mutagenesis of a Transporter for Uptake of a Non-Native Substrate Identifies Conformationally Dynamic Regions. preprint, Biochemistry, April 2021. URL http://biorxiv.org/lookup/doi/10.1101/2021.04.19.440442.
  227. 227.Haicang Zhang, Michelle S Xu, Xiao Fan, Wendy K Chung, and Yufeng Shen. Predicting functional effect of missense variants using graph attention neural networks. Nature Machine Intelligence, 4(11):1017–1028, 2022.
  228. 228.Naihui Zhou, Yuxiang Jiang, Timothy Bergquist, Alexandra J. Lee, Balint Z. Kacsoh, Alex Crocker, Kimberley A. Lewis, George E. Georghiou, Huy N. Nguyen, Nafiz Imtiaz Bin Hamid, Larry Davis, Tunca Dogan, Volkan Atalay, Ahmet Sureyya Rifaioglu, Alperen Dalkiran, Rengul Cetin-Atalay, Chengxin Zhang, Rebecca L. Hurto, Peter L. Freddolino, Yang Zhang, Prajwal Bhat, Fran Supek, José María Fernández, Branislava Gemovic, Vladimir Perovic, Radoslav Davidovic, Neven Sumonja, Nevena Veljkovic, Ehsaneddin Asgari, ´ Mohammad R. K. Mofrad, Giuseppe Profiti, Castrense Savojardo, Pier Luigi Martelli, Rita Casadio, Florian Boecker, Indika Kahanda, Natalie Thurlby, Alice Mchardy, Alexandre Renaux, Rabie Saidi, Julian Gough, Alex Alves Freitas, Magdalena Antczak, Fábio Fabris, Mark N. Wass, Jie Hou, Jianlin Cheng, Zheng Wang, Alfonso E. Romero, Alberto Paccanaro, Haixuan Yang, Tatyana Goldberg, Chenguang Zhao, Liisa Holm, Petri Törönen, Alan Medlar, Elaine Zosa, Itamar Borukhov, Ilya B. Novikov, Angela D. Wilkins, Olivier Lichtarge, Po-Han Chi, Wei-Cheng Tseng, Michal Linial, Peter W. Rose, Christophe Dessimoz, Vedrana Vidulin, Sašo Džeroski, Ian P. W. Sillitoe, Sayoni Das, Jonathan G. Lees, David T. Jones, Cen Wan, Domenico Cozzetto, Rui Fa, Mateo Torres, Alex Warwick Vesztrocy, Jose Manuel Rodriguez, Michael L. Tress, Marco Frasca, Marco Notaro, Giuliano Grossi, Alessandro Petrini, Matteo Ré, Giorgio Valentini, Marco Mesiti, Daniel B. Roche, Jonas Reeb, David W. Ritchie, Sabeur Aridhi, Seyed Ziaeddin Alborzi, Marie-Dominique Devignes, Da Chen Emily Koo, Richard Bonneau, Vladimir Gligorijevic, Meet Barot, Hai Fang, Stefano Toppo, Enrico ´ Lavezzo, Marco Falda, Michele Berselli, Silvio C. E. Tosatto, Marco Carraro, Damiano Piovesan, Hafeez ur Rehman, Qizhong Mao, Shanshan Zhang, Slobodan Vucetic, Gage S Black, Dane Jo, Dallas J. Larsen, Ashton Omdahl, Luke Sagers, Erica Suh, Jonathan B. Dayton, Liam James McGuffin, Danielle Allison Brackenridge, Patricia C. Babbitt, Jeffrey M. Yunes, Paolo Fontana, Feng Zhang, Shanfeng Zhu, Ronghui You, Zihan Zhang, Suyang Dai, Shuwei Yao, Weidong Tian, Renzhi Cao, Caleb Chandler, Miguel Amezola, Devon Johnson, Jia-Ming Chang, Wen-Hung Liao, Yi-Wei Liu, Stefano Pascarelli, Yotam Frank, R. Hoehndorf, Maxat Kulmanov, Imane Boudellioua, Gianfranco Politano, Stefano Di Carlo, Alfredo Benso, Kai Hakala, Filip Ginter, Farrokh Mehryary, Suwisa Kaewphan, Jari Björne, Hans Moen, Martti Tolvanen, Tapio Salakoski, Daisuke Kihara, Aashish Jain, Tomislav Šmuc, Adrian M. Altenhoff, Asa Ben-Hur, Burkhard Rost, Steven E. Brenner, Christine A. Orengo, Constance J. Jeffery, Giovanni Bosco, Deborah A. Hogan, Maria Jesus Martin, Claire O’Donovan, Sean D. Mooney, Casey S. Greene, Predrag Radivojac, and Iddo Friedberg. The CAFA challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens. Genome Biology, 20, 2019.
  229. 229.Julia Zinkus-Boltz, Craig DeValk, and Bryan C. Dickinson. A Phage-Assisted Continuous Selection Approach for Deep Mutational Scanning of Protein–Protein Interactions. ACS Chemical Biology, 14(12):2757–2767, December 2019. ISSN 1554-8929, 1554-8937. doi: 10.1021/acschembio.9b00669. URL https://pubs.acs.org/doi/10.1021/acschembio.9b00669.

Citation

MLA
Notin, P., et al. “ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design”. Advances in Neural Information Processing Systems 36, 2023, pp. 64331–79, https://doi.org/10.52202/075280-2810.
APA
Notin, P., Kollasch, A., Ritter, D., Niekerk, L. V., Paul, S., Spinner, H., Rollins, N., Shaw, A., Orenbuch, R., Weitzman, R., Frazer, J., Dias, M., Franceschi, D., Gal, Y., & Marks, D. (2023). ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design. Advances in Neural Information Processing Systems 36, 64331–64379. https://doi.org/10.52202/075280-2810
Chicago
Notin, P., A. Kollasch, D. Ritter, et al. 2023. “ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design”. Advances in Neural Information Processing Systems 36, 64331–79. https://doi.org/10.52202/075280-2810.
Harvard
Notin, P. et al. (2023) “ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design”, Advances in Neural Information Processing Systems 36. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp. 64331–64379. Available at: https://doi.org/10.52202/075280-2810.
Vancouver
1. Notin P, Kollasch A, Ritter D, et al (2023) ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design. In: Advances in Neural Information Processing Systems 36. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp 64331–64379

BibTeX

@inproceedings{Notin_2023, series={NeurIPS 2023}, title={ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design}, url={http://dx.doi.org/10.52202/075280-2810}, DOI={10.52202/075280-2810}, booktitle={Advances in Neural Information Processing Systems 36}, publisher={Neural Information Processing Systems Foundation, Inc. (NeurIPS)}, author={Notin, Pascal and Kollasch, Aaron and Ritter, Daniel and Niekerk, Lood Van and Paul, Steffanie and Spinner, Han and Rollins, Nathan and Shaw, Ada and Orenbuch, Rose and Weitzman, Ruben and Frazer, Jonathan and Dias, Mafalda and Franceschi, Dinko and Gal, Yarin and Marks, Debora}, year={2023}, pages={64331–64379}, collection={NeurIPS 2023} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors