ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design
Pascal NotinAaron KollaschDaniel RitterLood van NiekerkSteffanie PaulHan SpinnerNathan J. RollinsAda ShawRose OrenbuchRuben Weitzman
Establishes a standardized, large-scale benchmarking suite of over 250 deep mutational scanning assays and clinical datasets to systematically evaluate more than 70 machine learning models across zero-shot and supervised protein fitness prediction and design tasks.
Understanding and engineering protein functions holds immense potential for addressing critical challenges in healthcare, agriculture, and climate solutions. While machine learning has emerged as a transformative technology for predicting how mutations alter protein behavior, evaluating these computational models has historically been difficult. Existing validation efforts have relied on small, fragmented, and artificial datasets that fail to reflect the true diversity of protein families and experimental environments, obscuring which models are genuinely effective for clinical and industrial use.
The article introduces ProteinGym, an expansive benchmarking platform designed to rigorously assess and compare machine learning models across protein fitness prediction and protein design tasks. It systematically evaluates both zero-shot methods, which predict mutational outcomes without prior experimental labels, and supervised methods across standardized real-world benchmarks.
To conduct this evaluation, the authors compiled over 250 deep mutational scanning assays covering more than 2.7 million mutated sequences across 200 protein families, spanning diverse taxa, functional categories, and mutation types including substitutions, insertions, and deletions. This experimental dataset is paired with expert-curated clinical records covering approximately 65,000 human mutations from the ClinVar and gnomAD databases. Across this comprehensive suite, the authors standardized and tested more than 70 high-performing computational models spanning alignment-based methods, large protein language models, structure-based inverse folding architectures, and hybrid systems under multiple cross-validation strategies.
The analysis produced several vital findings. First, hybrid architectures such as TranceptEVE demonstrated the strongest overall zero-shot performance across deep mutational scanning and clinical datasets, though specialized alignment-based models like GEMME excelled in specific niches such as viral proteins and shallow sequence alignments. Second, in supervised settings where labeled training data is available, non-parametric transformer models like ProteinNPT outperformed all alternatives by jointly analyzing sequence context and assay labels. Third, unsupervised models frequently matched or outperformed supervised models on human clinical benchmarks, highlighting how supervised predictors often suffer from data leakage and overfitting to known disease genes. Finally, model rankings varied meaningfully depending on the objective: some models excelled at rank-ordering all mutations across a full distribution, while others were significantly better at prioritizing top-performing candidates for protein design.
These findings demonstrate that no single computational model fits every application, meaning organizations must align model selection directly with operational goals. Teams seeking to optimize high-performing proteins for manufacturing or drug design should prioritize models optimized for top-tier retrieval metrics, whereas clinical diagnostics require models capable of accurate broad-spectrum pathogenicity scoring. Relying on the wrong predictive framework risks pursuing unviable therapeutic targets, inflating laboratory validation costs, and lengthening project timelines.
Decision-makers should immediately utilize standardized, multi-metric benchmarking suites like ProteinGym to select and validate computational biology models prior to committing wet-lab resources. Organizations should also adopt hybrid and autoregressive architectures as foundational starting points for mutational effect scoring and explore specialized transformer frameworks for supervised property optimization. Before major deployment in production, teams should conduct focused pilot evaluations on the specific protein class or target phenotype of interest.
Confidence in these findings is high due to the unprecedented breadth of evaluated assays, though certain limitations remain. Deep mutational scanning assays inherently contain measurement noise, dynamic range limits, and selection biases toward well-studied disease targets and specific protein families. Furthermore, existing clinical datasets present risks of circularity and label bias. Ongoing validation on unstudied protein families and expanded testing into non-coding regulatory sequences will be essential to ensure continuous predictive reliability.
- Paper: Highly accurate protein structure prediction with AlphaFold, John Jumper et al. (2021). Provides the foundational AlphaFold structure prediction architecture that supplies the computational 3D structures leveraged by inverse folding and structure-aware fitness models benchmarked in ProteinGym.
- Paper: Learning inverse folding from millions of predicted structures, Chloe Hsu et al. (2022). Establishes inverse folding and mutational effect scoring trained on predicted structures, representing a core class of zero-shot models evaluated within ProteinGym.
- Paper: ProtTrans: Towards Cracking the Language of Life’s Code Through Self-Supervised Deep Learning and High Performance Computing, Ahmed Elnaggar et al. (2020). Introduces large-scale self-supervised protein language modeling, which forms the fundamental basis for the sequence-based zero-shot and supervised baseline models assessed across the benchmark.
- Paper: ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention, Mingchen Li et al. (2024). Develops a structure-quantized protein language model (ProSST) and validates its zero-shot mutation effect predictions directly on the ProteinGym benchmark.
- Paper: ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts, Minghao Xu et al. (2023). Extends protein representation learning by integrating functional biomedical text descriptions with sequence pretraining for enhanced functional and property prediction.
