Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time Retrieval
Pascal NotinMafalda DiasJonathan FrazerJavier Marchena-HurtadoAidan N. GomezDebora S. MarksYarin Gal
Introduces an autoregressive transformer architecture that combines multi-scale attention with inference-time homology retrieval to score insertions, deletions, and complex substitutions across diverse protein families without requiring multiple sequence alignments during training.
Predicting the functional impact of genetic mutations on protein fitness is essential for drug discovery, clinical diagnosis of disease variants, and biotherapeutic design. Traditional state-of-the-art computational methods rely heavily on training deep generative models on protein-specific multiple sequence alignments. However, this dependence creates significant bottlenecks: alignments cannot readily handle complex variations such as insertions and deletions, fail entirely on hard-to-align regions like disordered proteins, and perform poorly when protein families lack deep evolutionary records. Meanwhile, existing language models trained on massive unaligned databases avoid alignment constraints but generally underperform specialized alignment models when predicting mutation effects.
The main objective of the article is to demonstrate and evaluate Tranception, a novel autoregressive transformer architecture that combines large-scale pretraining on non-aligned protein sequences with lightweight evolutionary retrieval at inference time. The authors also establish ProteinGym, a large-scale, diverse benchmark designed to rigorously evaluate protein fitness prediction models across diverse taxa, mutation depths, and variation types.
To develop Tranception, the authors trained a 700-million-parameter autoregressive model on roughly 250 million unaligned sequences from the UniRef100 database, incorporating specialized attention mechanisms designed to capture contiguous subsequences at varying lengths. At inference time, the model scores protein fitness by merging its sequence log-likelihood predictions with positional frequency statistics retrieved from homologous evolutionary sequences. To evaluate predictive capability, the authors established ProteinGym, compiling 87 amino-acid substitution assays covering approximately 1.5 million variants and 7 insertion-deletion assays covering around 300,000 variants across human, viral, prokaryotic, and other eukaryotic targets.
The evaluation produced several critical findings. First, Tranception with inference-time retrieval achieved top overall predictive accuracy on the ProteinGym substitution benchmark, attaining an average Spearman rank correlation of 0.451 and surpassing leading specialized alignment models like EVE and existing protein language models like ESM-1v. Second, Tranception exhibited its most substantial performance advantages on proteins with shallow alignments, where conventional alignment-based methods degraded sharply. Third, the model demonstrated superior extrapolation capabilities on complex multiple-mutant sequences, achieving a correlation of 0.499 on variants with five or more simultaneous mutations compared to 0.420 for EVE. Fourth, Tranception effectively scored sequence insertions and deletions out-of-the-box, reaching a correlation of 0.463 and outperforming the only competing baseline capable of evaluating indels. Finally, model ensembling revealed that pairing Tranception with alignment-based models boosted correlation to 0.473, demonstrating high complementarity between autoregressive sequence modeling and evolutionary alignments.
These findings indicate that protein fitness prediction can be reliably scaled to the entire proteome without sacrificing accuracy or requiring time-consuming, protein-specific retraining. By decoupling model training from alignment availability, Tranception substantially reduces computational overhead, improves coverage across difficult-to-model disordered regions, and mitigates the risk of prediction failure on understudied or emerging pathogens. Furthermore, its demonstrated strength on multiple-substitution variants offers direct performance benefits for machine-learning-guided protein engineering and sequence design pipelines.
For practical implementation, organizations should leverage hybrid scoring frameworks by deploying autoregressive models augmented with lightweight inference retrieval. When maximum predictive accuracy is required for well-characterized proteins, practitioners should consider ensembling Tranception with specialized models like EVE. For generative sequence design or evaluating disordered regions and insertion-deletions, Tranception serves as an effective standalone tool. Future efforts should focus on expanding the architecture's capacity through scaling model size, integrating more diverse metagenomic sequence datasets, and generating additional multi-mutant experimental assays to refine design workflows.
Confidence in these findings is bolstered by the extensive scale and diversity of the ProteinGym benchmark, which consistently verified Tranception across multiple metrics. A primary limitation remains the benchmark's historical bias toward single-substitution mutations relative to multi-mutant and insertion-deletion profiles, as well as the model context window of 1,024 amino acids, which requires sequence slicing for the small fraction of proteins exceeding that length.
- Paper: ProtTrans: Towards Cracking the Language of Life’s Code Through Self-Supervised Deep Learning and High Performance Computing, Ahmed Elnaggar et al. (2020). ProtTrans establishes the foundational methodology of pre-training large self-supervised transformer language models on raw protein sequences without pre-alignment, which Tranception builds directly upon.
- Paper: Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, Ofir Press et al. (2022). This paper introduces ALiBi relative positional encodings, an architectural mechanism foundational for long-sequence autoregressive transformer modeling that Tranception leverages.
- Paper: ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design, Pascal Notin et al. (2023). This paper expands and formalizes the ProteinGym benchmark introduced alongside Tranception into a large-scale, standalone standardized platform evaluating over 70 protein fitness models.
- Paper: ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention, Mingchen Li et al. (2024). ProSST extends zero-shot protein fitness prediction by incorporating quantized 3D structural tokens alongside sequence representations, evaluating its performance directly against the ProteinGym benchmark.
