ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention
Mingchen LiYang TanXinzhu MaBozitao ZhongHuiqun YuZiyi ZhouWanli OuyangBingxin ZhouPan TanLiang Hong
Proposes a protein language model that combines sequence data with 3D structural information by quantizing local residue geometries into discrete tokens and coupling them through disentangled attention, achieving state-of-the-art accuracy on zero-shot mutation effect predictions and downstream protein function benchmarks.
Predicting protein function is essential for advancing drug discovery, biotechnology, and basic life sciences. While recent computational advances have established protein language models as fundamental analytical tools, most existing models analyze only linear amino acid sequences and neglect three-dimensional structural information. Because a protein's biological function is largely determined by its spatial structure, sequence-only methods often fail to capture critical functional regions such as catalytic sites and binding pockets.
The article evaluates a new computational model, named ProSST (Protein Sequence-Structure Transformer), designed to systematically integrate three-dimensional structural data with primary sequence information. The objective is to demonstrate that explicitly modeling local residue environments and their interactions with sequence data significantly improves accuracy across both zero-shot mutation effect predictions and diverse supervised downstream tasks.
The researchers developed a two-part framework comprising a structure quantizer and a disentangled attention transformer. The quantizer converts complex three-dimensional local structures (incorporating up to the nearest 40 neighboring residues) into discrete structural tokens using a geometric vector perceptron encoder and clustering across millions of structural fragments from curated databases. The transformer architecture explicitly decouples and computes attention between residues, structures, and relative spatial positions. The model was pre-trained on a non-redundant database of 18.8 million predicted protein structures using a masked language modeling objective.
The analysis reveals several key findings. First, ProSST achieved state-of-the-art accuracy in zero-shot mutation effect prediction on the ProteinGym benchmark, securing a Spearman rank correlation of 0.504 compared to previous best-performing baselines ranging from 0.434 to 0.458. Second, the model demonstrated superior predictive capability in protein stability, binding, and expression subsets. Third, in supervised downstream tasks, the 110-million-parameter ProSST matched or outperformed baseline models that were up to six times larger (over 650 million parameters), leading the field in protein sub-cellular localization (94.32% accuracy) and metal ion binding prediction (76.37% accuracy). Finally, ablation studies confirmed that performance gains stem directly from the structure quantization codebook (optimally sized at 2,048 tokens) and the residue-to-structure attention mechanism rather than mere parameter scaling.
These results demonstrate that incorporating discrete structural micro-environments into language models significantly improves functional predictive power while maintaining compact model sizes. This efficiency reduces high-performance computing costs, accelerates screening timelines for protein engineering, and lowers experimental failure risks in laboratory pipelines by providing more accurate early-stage computational filtering.
Organizations evaluating computational biology pipelines should consider integrating structure-aware language modeling for mutation screening and functional annotation. For datasets where experimental or predicted structures are missing, the article outlines two viable pathways: generating structures via reliable prediction algorithms such as AlphaFold2 or utilizing masked-structure training variants designed for sequence-only inputs. Future research should prioritize optimizing the computational speed of structure quantization, scaling training to even larger structural databases, and refining performance on intrinsically disordered proteins where structure prediction confidence is typically lower.
- Paper: Learning inverse folding from millions of predicted structures, Chloe Hsu et al. (2022). This paper establishes the methodology of pretraining geometric vector perceptron-based architectures on millions of predicted structures, providing the direct foundation for ProSST's structural encoding pipeline.
- Paper: ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design, Pascal Notin et al. (2023). This work introduces the comprehensive ProteinGym fitness prediction benchmark that serves as the primary evaluation suite for validating ProSST's zero-shot mutation effect performance.
- Paper: Protein Representation Learning by Geometric Structure Pretraining, Zuobai Zhang et al. (2023). This study demonstrates pretraining representation models directly on 3D geometric structures, establishing key principles for integrating structural micro-environments into protein language models.
- Paper: ProtTrans: Towards Cracking the Language of Life’s Code Through Self-Supervised Deep Learning and High Performance Computing, Ahmed Elnaggar et al. (2020). This paper outlines foundational large-scale sequence-based protein language modeling and downstream functional transfer benchmarks that ProSST builds upon and compares against.
- Paper: Highly accurate protein structure prediction with AlphaFold, John Jumper et al. (2021). AlphaFold provides the accurate computational 3D structure predictions required to build the massive structural databases and codebooks utilized by ProSST.
- Paper: Self-Attention with Relative Position Representations, Peter Shaw et al. (2018). This paper introduces relative position representations into transformer self-attention, foundational to ProSST's disentangled attention mechanism across sequence and spatial positions.
- Paper: A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery, Yu Zhang et al. (2024). This survey contextualizes structure-aware protein transformers like ProSST within the broader landscape of multimodal foundation models and masked sequence pre-training across scientific disciplines.
