Built independently by an author, for readers. Read the story and support ChapterPal

keyword

ProteinGym Substitution benchmark

The ProteinGym Substitution benchmark is a collection of experimental measurements used to evaluate how well computational methods predict the effects of amino-acid substitutions on protein fitness. It comprises multiplexed assays that test variant effects across diverse proteins, enabling comparison of prediction methods on measured outcomes.

2 items

Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval

Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval

Pascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado, Aidan N. Gomez, Debora S. Marks, Yarin Gal

OrganizationsCohereHarvard UniversityUniversity of Oxford

Why you should read this

Introduces an autoregressive transformer architecture that combines multi-scale attention with inference-time homology retrieval to score insertions, deletions, and complex substitutions across diverse protein families without requiring multiple sequence alignments during training.

The ability to accurately model the fitness landscape of protein sequences is critical to a wide range of applications, from quantifying the effects of human variants on disease likelihood, to predicting immune-escape mutations in viruses and designing novel biotherapeutic proteins. Deep generative models of protein sequences trained on multiple sequence alignments have been the most successful approaches so far to address these tasks. The performance of these methods is however contingent on the availability of sufficiently deep and diverse alignments for reliable training. Their potential scope is thus limited by the fact many protein families are hard, if not impossible, to align. Large language models trained on massive quantities of non-aligned protein sequences from diverse families address these problems and show potential to eventually bridge the performance gap. We introduce Tranception, a novel transformer architecture leveraging autoregressive predictions and retrieval of homologous sequences at inference to achieve state-of-the-art fitness prediction performance. Given its markedly higher performance on multiple mutants, robustness to shallow alignments and ability to score indels, our approach offers significant gain of scope over existing approaches. To enable more rigorous model testing across a broader range of protein families, we develop ProteinGym – an extensive set of multiplexed assays of variant effects, substantially increasing both the number and diversity of assays compared to existing benchmarks.

Added

2026-09-28

ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts

ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts

Minghao Xu, Xinyu Yuan, Santiago Miret, Jian Tang

OrganizationsCIFARHEC MontréalIntelMila – Québec Artificial Intelligence InstituteUniversité de Montréal

Why you should read this

Proposes a multimodal pre-training framework that aligns protein sequences with biomedical text descriptions to improve protein language models and enable zero-shot protein classification and functional retrieval.

Current protein language models (PLMs) learn protein representations mainly based on their sequences, thereby well capturing co-evolutionary information, but they are unable to explicitly acquire protein functions, which is the end goal of protein representation learning. Fortunately, for many proteins, their textual property descriptions are available, where their various functions are also described. Motivated by this fact, we first build the ProtDescribe dataset to augment protein sequences with text descriptions of their functions and other important properties. Based on this dataset, we propose the ProtST framework to enhance Protein Sequence pre-training and understanding by biomedical Texts. During pre-training, we design three types of tasks, i.e., unimodal mask prediction, multimodal representation alignment and multimodal mask prediction, to enhance a PLM with protein property information with different granularities and, at the same time, preserve the PLM’s original representation power. On downstream tasks, ProtST enables both supervised learning and zero-shot prediction. We verify the superiority of ProtST-induced PLMs over previous ones on diverse representation learning benchmarks. Under the zero-shot setting, we show the effectiveness of ProtST on zero-shot protein classification, and ProtST also enables functional protein retrieval from a large-scale database without any function annotation. Source code and model weights are available at https://github.com/DeepGraphLearning/ProtST.

Added

2026-09-26