ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention

Mingchen LiYang TanXinzhu MaBozitao ZhongHuiqun YuZiyi ZhouWanli OuyangBingxin ZhouPan TanLiang Hong

article2024NeurIPS71 citations

Proposes a protein language model that combines sequence data with 3D structural information by quantizing local residue geometries into discrete tokens and coupling them through disentangled attention, achieving state-of-the-art accuracy on zero-shot mutation effect predictions and downstream protein function benchmarks.

Listen

Predicting protein function is essential for advancing drug discovery, biotechnology, and basic life sciences. While recent computational advances have established protein language models as fundamental analytical tools, most existing models analyze only linear amino acid sequences and neglect three-dimensional structural information. Because a protein's biological function is largely determined by its spatial structure, sequence-only methods often fail to capture critical functional regions such as catalytic sites and binding pockets.

The article evaluates a new computational model, named ProSST (Protein Sequence-Structure Transformer), designed to systematically integrate three-dimensional structural data with primary sequence information. The objective is to demonstrate that explicitly modeling local residue environments and their interactions with sequence data significantly improves accuracy across both zero-shot mutation effect predictions and diverse supervised downstream tasks.

The researchers developed a two-part framework comprising a structure quantizer and a disentangled attention transformer. The quantizer converts complex three-dimensional local structures (incorporating up to the nearest 40 neighboring residues) into discrete structural tokens using a geometric vector perceptron encoder and clustering across millions of structural fragments from curated databases. The transformer architecture explicitly decouples and computes attention between residues, structures, and relative spatial positions. The model was pre-trained on a non-redundant database of 18.8 million predicted protein structures using a masked language modeling objective.

The analysis reveals several key findings. First, ProSST achieved state-of-the-art accuracy in zero-shot mutation effect prediction on the ProteinGym benchmark, securing a Spearman rank correlation of 0.504 compared to previous best-performing baselines ranging from 0.434 to 0.458. Second, the model demonstrated superior predictive capability in protein stability, binding, and expression subsets. Third, in supervised downstream tasks, the 110-million-parameter ProSST matched or outperformed baseline models that were up to six times larger (over 650 million parameters), leading the field in protein sub-cellular localization (94.32% accuracy) and metal ion binding prediction (76.37% accuracy). Finally, ablation studies confirmed that performance gains stem directly from the structure quantization codebook (optimally sized at 2,048 tokens) and the residue-to-structure attention mechanism rather than mere parameter scaling.

These results demonstrate that incorporating discrete structural micro-environments into language models significantly improves functional predictive power while maintaining compact model sizes. This efficiency reduces high-performance computing costs, accelerates screening timelines for protein engineering, and lowers experimental failure risks in laboratory pipelines by providing more accurate early-stage computational filtering.

Organizations evaluating computational biology pipelines should consider integrating structure-aware language modeling for mutation screening and functional annotation. For datasets where experimental or predicted structures are missing, the article outlines two viable pathways: generating structures via reliable prediction algorithms such as AlphaFold2 or utilizing masked-structure training variants designed for sequence-only inputs. Future research should prioritize optimizing the computational speed of structure quantization, scaling training to even larger structural databases, and refining performance on intrinsically disordered proteins where structure prediction confidence is typically lower.

Cover for ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention

Abstract

Protein language models (PLMs) have shown remarkable capabilities in various protein function prediction tasks. However, while protein function is intricately tied to structure, most existing PLMs do not incorporate protein structure information. To address this issue, we introduce ProSST, a Transformer-based protein language model that seamlessly integrates both protein sequences and structures. ProSST incorporates a structure quantization module and a Transformer architecture with disentangled attention. The structure quantization module translates a 3D protein structure into a sequence of discrete tokens by first serializing the protein structure into residue-level local structures and then embeds them into dense vector space. These vectors are then quantized into discrete structure tokens by a pre-trained clustering model. These tokens serve as an effective protein structure representation. Furthermore, ProSST explicitly learns the relationship between protein residue token sequences and structure token sequences through the sequence-structure disentangled attention. We pre-train ProSST on millions of protein structures using a masked language model objective, enabling it to learn comprehensive contextual representations of proteins. To evaluate the proposed ProSST, we conduct extensive experiments on the zero-shot mutation effect prediction and several supervised downstream tasks, where ProSST achieves the state-of-the-art performance among all baselines. Our code and pre-trained models are publicly available 2.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Protein Representation Models
  • 2.2 Protein Structure Quantization
  • 3 Method
  • 3.1 Structure Quantization Module
  • 3.2 Sequence-Structure Disentangled Attention
  • 3.3 Pre-Training Objective
  • 4 Experiments
  • 4.1 Zero-Shot Mutant Effect Prediction
  • 4.2 Supervised Fine-Tuning Tasks
  • 4.3 Ablation Study
  • 5 Conclusion and Limitations
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Zero-Shot Scoring
  • A.2 Details of the Datasets and Metrics
  • A.3 Details of Implementations
  • A.4 Performance of models on the ProteinGYM benchmark separated by functional categories
  • A.5 Additional experiments on disentangled attention.
  • A.6 Solutions to Sequence-only Datasets
  • A.7 AlphaFold pLDDT versus Zero-shot mutant effect performance
  • NeurIPS Paper Checklist
  • 1. Claims
  • 2. Limitations
  • 3. Theory Assumptions and Proofs
  • 4. Experimental Result Reproducibility
  • 5. Open access to data and code
  • 6. Experimental Setting/Details
  • 7. Experiment Statistical Significance
  • 8. Experiments Compute Resources
  • 9. Code Of Ethics
  • 10. Broader Impacts
  • 11. Safeguards
  • 12. Licenses for existing assets
  • 13. New Assets
  • 14. Crowdsourcing and Research with Human Subjects
  • 15. Institutional Review Board (IRB) Approvals or Equivalent for Research with Human Subjects

Knowls

  1. Knowl 1 — Protein Structure Quantization Module

    model/method

    The structure quantization module in ProSST discretizes continuous 3D protein structures into sequences of discrete structure tokens s=(s1,s2,…,sL)\mathbf{s} = (s_1, s_2, \dots, s_L), where LL is the number of residues in the protein.

    Each residue ii is represented by its local structural environment within the complete protein structure. The local structure is formalized as a star-shaped graph gi=(Vi,Ei)g_i = (V_i, E_i) centered on residue ii and containing its k=40k=40 nearest spatial neighbors within a Euclidean distance of 10 A˚10\,\text{Å} based on CαC_\alpha atom coordinates. The graph nodes contain purely geometric coordinate representations without amino acid residue types, ensuring the structure encoder captures only backbone geometric cues.

    The local structure graph gig_i is embedded using a 6-layer Geometric Vector Perceptron (GVP) graph neural network with node embedding dimension 256 and edge embedding dimension 64. The GVP encoder is trained as an auto-encoder on the CATH43-S40 dataset (comprising 31,270 non-redundant crystal structures) using a denoising objective where CαC_\alpha coordinates are perturbed by 3D Gaussian noise and rotation matrices are perturbed by Brownian motion on the SO(3)SO(3) manifold. The decoder is discarded after training, and the mean-pooled node representation from the encoder yields a continuous local structure vector ri=1∣Vi∣∑v∈Viπθ(gi)v∈Rd\mathbf{r}_i = \frac{1}{|V_i|} \sum_{v \in V_i} \pi_\theta(g_i)_v \in \mathbb{R}^d.

    To discretize ri\mathbf{r}_i, KK-means clustering is fitted on 4,735,6774,735,677 local structure vectors extracted across CATH43-S40, producing a codebook of KK centroids {ek}k=1K\{\mathbf{e}_k\}_{k=1}^K (where K=2048K=2048 by default). Residue ii's local structure vector ri\mathbf{r}_i is mapped to discrete token si=arg⁡min⁡k∥ri−ek∥2s_i = \arg\min_{k} \|\mathbf{r}_i - \mathbf{e}_k\|_2.

  2. Knowl 2 — Sequence-Structure Disentangled Attention Mechanism

    equation

    In ProSST, the standard self-attention mechanism in the Transformer is replaced by a sequence-structure disentangled attention mechanism that explicitly computes cross-attention among residue hidden states, quantized structure token embeddings, and relative positions.

    Let Ri∈Rd\mathbf{R}_i \in \mathbb{R}^d be the residue hidden state at position ii, Si∈Rd\mathbf{S}_i \in \mathbb{R}^d be the embedding of the local structure token sis_i, and Pi∣j∈Rd\mathbf{P}_{i|j} \in \mathbb{R}^d be the embedding of the relative position between positions ii and jj. The relative distance index δ(i,j)\delta(i, j) clipped at maximum distance Lmax=1024L_{\text{max}} = 1024 is:

    δ(i,j)={0if i−j≤−Lmax,2Lmax−1if i−j≥Lmax,i−j+Lmaxotherwise.\delta(i, j) = \begin{cases} 0 & \text{if } i - j \le -L_{\text{max}}, \\ 2L_{\text{max}} - 1 & \text{if } i - j \ge L_{\text{max}}, \\ i - j + L_{\text{max}} & \text{otherwise.} \end{cases}

    Linear projections for residue (rr), structure (ss), and relative position (pp) query, key, and value vectors are computed via learnable projection matrices:

    Qr=RWrq,Kr=RWrk,Vr=RWrv\mathbf{Q}^r = \mathbf{R}\mathbf{W}^q_r, \quad \mathbf{K}^r = \mathbf{R}\mathbf{W}^k_r, \quad \mathbf{V}^r = \mathbf{R}\mathbf{W}^v_r

    Qs=SWsq,Ks=SWsk\mathbf{Q}^s = \mathbf{S}\mathbf{W}^q_s, \quad \mathbf{K}^s = \mathbf{S}\mathbf{W}^k_s

    Qp=PWpq,Kp=SWpk\mathbf{Q}^p = \mathbf{P}\mathbf{W}^q_p, \quad \mathbf{K}^p = \mathbf{S}\mathbf{W}^k_p

    The attention score A^i,j\hat{A}_{i,j} between residue ii and residue jj is decomposed into five components (omitting structure-to-structure, structure-to-position, position-to-structure, and position-to-position interactions):

    A^i,j=Qir(Kjr)⊤+Qir(Kjs)⊤+Qir(Kδ(i,j)p)⊤+Kjr(Qis)⊤+Kjr(Qδ(j,i)p)⊤\hat{A}_{i,j} = \mathbf{Q}^r_i (\mathbf{K}^r_j)^\top + \mathbf{Q}^r_i (\mathbf{K}^s_j)^\top + \mathbf{Q}^r_i (\mathbf{K}^p_{\delta(i,j)})^\top + \mathbf{K}^r_j (\mathbf{Q}^s_i)^\top + \mathbf{K}^r_j (\mathbf{Q}^p_{\delta(j,i)})^\top

    where Qir(Kjr)⊤\mathbf{Q}^r_i (\mathbf{K}^r_j)^\top represents residue-to-residue, Qir(Kjs)⊤\mathbf{Q}^r_i (\mathbf{K}^s_j)^\top is residue-to-structure, Qir(Kδ(i,j)p)⊤\mathbf{Q}^r_i (\mathbf{K}^p_{\delta(i,j)})^\top is residue-to-position, Kjr(Qis)⊤\mathbf{K}^r_j (\mathbf{Q}^s_i)^\top is structure-to-residue, and Kjr(Qδ(j,i)p)⊤\mathbf{K}^r_j (\mathbf{Q}^p_{\delta(j,i)})^\top is position-to-residue attention.

    To ensure training stability across the sum of five terms, a scaling factor of 15d\frac{1}{\sqrt{5d}} is applied. The resulting residue hidden state Ro\mathbf{R}_o is:

    Ro=softmax(A^5d)Vr\mathbf{R}_o = \text{softmax}\left(\frac{\hat{\mathbf{A}}}{\sqrt{5d}}\right) \mathbf{V}^r

  3. Knowl 3 — Structure-Conditioned Masked Language Modeling Pre-Training

    model/method

    ProSST is pre-trained using a structure-conditioned masked language modeling (MLM) objective on 18.818.8 million protein structures obtained from the 90%90\% non-redundant AlphaFoldDB cluster.

    Given an amino acid sequence x=(x1,…,xL)\mathbf{x} = (x_1, \dots, x_L) and its corresponding quantized structure token sequence s=(s1,…,sL)\mathbf{s} = (s_1, \dots, s_L), a subset of indices M⊂{1,…,L}M \subset \{1, \dots, L\} comprising 15%15\% of all sequence positions is randomly selected. At each selected index i∈Mi \in M, the residue is substituted with a special [MASK] token with 80%80\% probability, replaced with a random residue token with 10%10\% probability, or left unchanged with 10%10\% probability. The structure token sequence s\mathbf{s} is left uncorrupted, serving as clean contextual structural conditioning.

    The pre-training loss minimizes the negative log-likelihood of the corrupted residues conditioned on the corrupted sequence x/M\mathbf{x}_{/M} and structure sequence s\mathbf{s}:

    LMLM=Ex∼X,M∑i∈M−log⁡p(xi∣x/M,s)\mathcal{L}_{\text{MLM}} = \mathbb{E}_{\mathbf{x} \sim \mathcal{X}, M} \sum_{i \in M} -\log p(x_i \mid \mathbf{x}_{/M}, \mathbf{s})

    Model architecture specifications: 12 Transformer layers, 12 attention heads, hidden dimension d=768d = 768, and feed-forward dimension 3,172 with GELU activations (totaling approximately 110M parameters). Pre-training is performed using AdamW (β1=0.9,β2=0.999\beta_1 = 0.9, \beta_2 = 0.999, weight decay 0.0010.001, gradient clipping threshold 1.01.0, dropout 0.10.1) with a batch size of 8,192 tokens for 500,000 steps. The learning rate warms up from 0 to 2×10−42 \times 10^{-4} over the first 2,000 steps and decays to 0 following a cosine schedule.

  4. Knowl 4 — Zero-Shot Mutation Effect Scoring Formulation

    equation

    ProSST scores single-point and multi-point amino acid mutations in a zero-shot setting by computing the log-likelihood ratio between the mutant residues and the wild-type residues, conditioned on both the wild-type amino acid sequence and the wild-type quantized structure token sequence.

    Let x\mathbf{x} be the wild-type amino acid sequence, and s\mathbf{s} be the wild-type structure token sequence. A mutation F={(pi,fi,wi)}i=1∣F∣F = \{(p_i, f_i, w_i)\}_{i=1}^{|F|} is defined as a set of triplets where pi∈{1,…,L}p_i \in \{1, \dots, L\} is the mutated position index, wiw_i is the original wild-type amino acid at position pip_i, and fif_i is the mutated amino acid at position pip_i.

    The zero-shot fitness score is computed as:

    Score(F)=∑i=1∣F∣[log⁡P(xpi=fi∣x,s)−log⁡P(xpi=wi∣x,s)]\text{Score}(F) = \sum_{i=1}^{|F|} \left[\log P(x_{p_i} = f_i \mid \mathbf{x}, \mathbf{s}) - \log P(x_{p_i} = w_i \mid \mathbf{x}, \mathbf{s})\right]

    where P(xpi=⋅∣x,s)P(x_{p_i} = \cdot \mid \mathbf{x}, \mathbf{s}) denotes the conditional probability assigned by ProSST to a given amino acid at position pip_i given the wild-type sequence x\mathbf{x} and wild-type structure tokens s\mathbf{s}.

  5. Knowl 5 — Zero-Shot Mutation Effect Prediction Performance on ProteinGym

    data/table

    ProSST was evaluated on the ProteinGym zero-shot substitution mutation benchmark comprising 217 experimental assays using AlphaFold2-predicted wild-type structures. Performance was measured by Spearman's rank correlation (ρs\rho_s), Normalized Discounted Cumulative Gain (NDCG), and Top-recall.

    Model Model Type ρs↑\rho_s \uparrow NDCG ↑\uparrow Top-recall ↑\uparrow
    EVE Evolution-based 0.439 0.781 0.230
    EVmutation Evolution-based 0.395 0.777 0.222
    DeepSequence Evolution-based 0.407 0.774 0.225
    WaveNet Evolution-based 0.373 0.761 0.203
    GEMME Evolution-based 0.457 0.777 0.211
    MSA-Transformer Evolution-based 0.434 0.779 0.217
    Tranception Sequence-based 0.434 0.779 0.220
    RITA Sequence-based 0.372 0.751 0.193
    UniRep Sequence-based 0.190 0.647 0.139
    ESM-1v Sequence-based 0.374 0.732 0.211
    ESM-2 Sequence-based 0.414 0.747 0.217
    ProGen2 Sequence-based 0.391 0.767 0.199
    VESPA Sequence-based 0.394 0.759 0.201
    ESM-IF Inverse-folding 0.422 0.748 0.223
    MIF-ST Inverse-folding 0.401 0.765 0.226
    Tranception-EVE Ensemble Models 0.457 0.786 0.230
    ESM-1v* Ensemble Models 0.407 0.749 0.211
    DeepSequence* Ensemble Models 0.419 0.776 0.226
    SaProt Sequence-Structure 0.457 0.768 0.233
    ProSST Sequence-Structure 0.504 0.777 0.239

    ProSST achieves the highest overall Spearman correlation (ρs=0.504\rho_s = 0.504) and Top-recall (0.2390.239), outperforming all sequence-only models, evolution-based methods, inverse folding models, and the hybrid sequence-structure baseline SaProt (ρs=0.457\rho_s = 0.457). Bootstrap significance tests across 10,000 resamples show all standard errors of difference versus baselines are below 0.010.01.

  6. Knowl 6 — Supervised Downstream Task Performance

    data/table

    ProSST was evaluated across four supervised downstream tasks: DeepLoc (subcellular localization binary classification, accuracy), Metal Ion Binding (binary site classification, accuracy), Thermostability (FLIP human-cell split regression, Spearman correlation ρs\rho_s), and Gene Ontology (GO) annotation multi-label classification across Molecular Function (MF), Biological Process (BP), and Cellular Component (CC) evaluated by Max F1-score.

    Model # Params DeepLoc Acc% ↑\uparrow Metal Ion Binding Acc% ↑\uparrow Thermostability ρs↑\rho_s \uparrow GO-MF F1-Max ↑\uparrow GO-BP F1-Max ↑\uparrow GO-CC F1-Max ↑\uparrow
    ESM-2 650M 91.96 71.56 0.680 0.670 0.473 0.470
    ESM-1b 650M 92.83 73.57 0.708 0.656 0.451 0.466
    MIF-ST 643M 91.76 75.08 0.694 0.633 0.375 0.322
    GearNet 42M 89.18 71.26 0.571 0.644 0.481 0.476
    SaProt-35M 35M 91.97 74.29 0.692 0.642 0.431 0.418
    SaProt-650M 650M 93.55 75.75 0.724 0.682 0.486 0.479
    ESM-GearNet 690M 93.55 74.11 0.651 0.676 0.516 0.507
    ProSST 110M 94.32 (0.10) 76.37 (0.02) 0.726 (0.04) 0.682 (0.003) 0.492 (0.004) 0.501 (0.002)

    With 110M parameters, ProSST achieves the top performance on 4 of the 6 benchmark settings (DeepLoc, Metal Ion Binding, Thermostability, and GO-MF), matching or outperforming models with over 6×6\times more parameters, including SaProt-650M and ESM-GearNet-690M.

  7. Knowl 7 — Ablation of Structure Vocabulary Size and Quantization Method

    data/table

    Ablations on the structure codebook vocabulary size KK and comparison against alternative structure tokenizers (Foldseek and DSSP) were evaluated on DeepLoc accuracy, zero-shot ProteinGym metrics, and pre-training validation perplexity.

    Model Variant DeepLoc Acc% ↑\uparrow ProteinGYM ρs↑\rho_s \uparrow NDCG ↑\uparrow Top-Recall ↑\uparrow Perplexity ↓\downarrow
    ProSST (K=4096K=4096) 93.88 (0.15) 0.498 0.773 0.233 8.880
    ProSST (K=2048K=2048) 94.32 (0.10) 0.504 0.777 0.239 9.033
    ProSST (K=1024K=1024) 93.43 (0.15) 0.485 0.760 0.231 9.333
    ProSST (K=512K=512) 93.70 (0.16) 0.471 0.759 0.223 9.577
    ProSST (K=128K=128) 93.14 (0.04) 0.469 0.753 0.228 10.021
    ProSST (K=20K=20) 93.05 (0.13) 0.438 0.744 0.210 10.719
    ProSST (K=1K=1) 89.48 (0.24) 0.390 0.738 0.181 12.182
    ProSST (K=0K=0) 89.77 (0.26) 0.392 0.741 0.184 12.190
    ProSST (Foldseek) 93.08 (0.22) 0.468 0.759 0.228 10.049
    ProSST (DSSP) 93.16 (0.16) 0.439 0.760 0.204 10.009

    K=0K=0 denotes a sequence-only model without structure tokens; K=1K=1 feeds a constant dummy structure token into the disentangled attention module. Key observations:

    1. Increasing KK improves model performance up to K=2048K=2048, where downstream and zero-shot performance peaks.
    2. ProSST (K=1K=1) achieves nearly identical results to ProSST (K=0K=0), demonstrating that parameter additions in disentangled attention do not yield improvements without meaningful structural cues.
    3. The GVP-quantized vocabulary (K=2048K=2048, ρs=0.504\rho_s=0.504) significantly outperforms Foldseek (ρs=0.468\rho_s=0.468) and DSSP (ρs=0.439\rho_s=0.439) under identical model architectures.
  8. Knowl 8 — Ablation of Disentangled Attention Components

    data/table

    Ablation experiments evaluate the contribution of individual attention terms in the sequence-structure disentangled attention formulation against standard Transformer self-attention.

    Attention Variant DeepLoc Acc% ↑\uparrow ProteinGYM ρs↑\rho_s \uparrow NDCG ↑\uparrow Top-Recall ↑\uparrow Perplexity ↓\downarrow
    ProSST (Full) 94.32 (0.10) 0.504 0.777 0.239 9.033
    ProSST (−- P2R) 91.31 (0.14) 0.478 0.778 0.227 9.173
    ProSST (−- R2P) 92.17 (0.32) 0.466 0.772 0.216 9.410
    ProSST (−- R2S) 90.48 (0.41) 0.438 0.766 0.208 12.142
    ProSST (−- S2R) 91.27 (0.20) 0.475 0.779 0.226 9.355
    ProSST (−- PE) 86.05 (0.65) 0.095 0.634 0.126 13.885
    ProSST (self-attention) 90.37 (0.21) 0.401 0.728 0.189 12.346

    Removing residue-to-structure attention (−-R2S) causes the largest individual component degradation (perplexity jumps from 9.033 to 12.142; ρs\rho_s drops from 0.504 to 0.438). Position-to-residue (−-P2R) has the smallest impact on perplexity (9.173). Replacing the disentangled attention formulation with standard self-attention while retaining structure tokens causes a severe reduction in zero-shot performance (ρs=0.401\rho_s = 0.401 vs 0.5040.504), demonstrating that explicit multi-term interaction is necessary to exploit structural information effectively.

  9. Knowl 9 — ProteinGym Zero-Shot Performance by Functional Category

    data/table

    Zero-shot mutation effect prediction on ProteinGym categorized across five functional subsets: Activity, Binding, Expression, Organismal Fitness, and Stability, evaluated by Spearman's rank correlation (ρs\rho_s).

    Model Activity Binding Expression Organismal Fitness Stability
    EVE 0.464 0.386 0.408 0.447 0.491
    EVmutation 0.440 0.317 0.378 0.411 0.430
    DeepSequence 0.455 0.363 0.390 0.413 0.476
    WaveNet 0.379 0.325 0.350 0.365 0.449
    GEMME 0.482 0.383 0.438 0.452 0.519
    MSA-Transformer 0.469 0.337 0.446 0.421 0.495
    Tranception 0.465 0.349 0.450 0.436 0.471
    RITA 0.366 0.302 0.414 0.381 0.398
    UniRep 0.182 0.202 0.216 0.141 0.210
    ESM-1v 0.396 0.268 0.405 0.362 0.437
    ESM-2 0.425 0.337 0.415 0.369 0.523
    ProGen2 0.402 0.302 0.418 0.387 0.445
    VESPA 0.429 0.347 0.326 0.404 0.461
    ESM-IF 0.368 0.389 0.407 0.324 0.624
    MIF-ST 0.390 0.321 0.438 0.366 0.485
    Tranception-EVE 0.487 0.376 0.457 0.460 0.500
    ESM-1v (ensemble) 0.420 0.320 0.429 0.387 0.477
    DeepSequence (ensemble) 0.455 0.363 0.390 0.413 0.476
    SaProt 0.458 0.378 0.488 0.367 0.592
    ProSST 0.448 0.477 0.506 0.415 0.674

    ProSST achieves state-of-the-art Spearman correlations on Stability (ρs=0.674\rho_s = 0.674, surpassing SaProt's 0.5920.592 and ESM-IF's 0.6240.624), Binding (ρs=0.477\rho_s = 0.477, vs. ESM-IF's 0.3890.389), and Expression (ρs=0.506\rho_s = 0.506, vs. SaProt's 0.4880.488).

  10. Knowl 10 — Masked Structure Training for Sequence-Only Input Inference

    model/method

    To handle sequence-only inputs where 3D structures are unavailable, ProSST introduces Masked Structure Training (MST). During pre-training under the MST paradigm, each protein sample has a 50%50\% probability of having its structural token sequence entirely replaced by a constant masked token sequence [1,1,…,1][1, 1, \dots, 1], simulating missing structural inputs.

    At inference time on sequence-only inputs, the dummy sequence [1,1,…,1][1, 1, \dots, 1] is supplied as the structure token sequence.

    Evaluation on ProteinGym (ρs\rho_s), Binary Localization Prediction (BLP accuracy from the PEER benchmark), and validation perplexity compares input modalities:

    • ProSST (K=2048K=2048, AlphaFold2 structures): ProteinGym ρs=0.504\rho_s = 0.504, BLP =94.32%= 94.32\%, Perplexity =9.033= 9.033.
    • ProSST (K=2048K=2048, ESMFold structures): ProteinGym ρs=0.471\rho_s = 0.471, BLP =92.73%= 92.73\%, Perplexity =9.144= 9.144.
    • ProSST (MST, missing structure [1,…,1][1,\dots,1]): ProteinGym ρs=0.438\rho_s = 0.438, BLP =91.84%= 91.84\%, Perplexity =10.325= 10.325.
    • ProSST (MST, AlphaFold2 structures): ProteinGym ρs=0.456\rho_s = 0.456, BLP =92.31%= 92.31\%, Perplexity =9.447= 9.447.
    • ProSST (K=0K=0, sequence-only model): ProteinGym ρs=0.392\rho_s = 0.392, BLP =89.65%= 89.65\%, Perplexity =12.190= 12.190.

    ProSST (MST) applied to sequence-only inputs achieves higher performance (ρs=0.438\rho_s = 0.438, BLP =91.84%= 91.84\%) than training a sequence-only model from scratch (K=0K=0, ρs=0.392\rho_s = 0.392, BLP =89.65%= 89.65\%).

  11. Knowl 11 — Limitations and Sensitivity to Structural Prediction Confidence

    limitation

    ProSST exhibits two key limitations:

    1. Quantization Computation Overhead: Serializing protein structures into kk-NN residue graphs (k=40k=40) and running GVP message passing to assign discrete cluster tokens requires significant pre-processing time and memory, limiting throughput compared to raw sequence tokenization.

    2. Sensitivity to Predicted Structure Quality: When experimental structures are absent, ProSST relies on predicted coordinates (e.g., AlphaFold2). Zero-shot mutation effect prediction accuracy on ProteinGym positively correlates with AlphaFold predicted Local Distance Difference Test (pLDDT) confidence scores, with a Pearson correlation coefficient of ρp=0.30\rho_p = 0.30 (compared to ρp=0.31\rho_p = 0.31 for SaProt and ρp=0.42\rho_p = 0.42 for ESM-IF1). Consequently, performance deteriorates on intrinsically disordered protein regions or low-confidence structural predictions.

Coverage note — None was omitted; all contributed methodologies, model architectures, attention equations, zero-shot scoring formulations, experimental benchmark results, ablation studies, inference extensions (MST), and stated limitations are fully represented.

References

  1. 1.William P Jencks. Catalysis in chemistry and enzymology. Courier Corporation, 1987.
  2. 2.The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Research, 51(D1):D523–D531, 11 2022.
  3. 3.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  4. 4.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  5. 5.Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences, 118(15):e2016239118, 2021.
  6. 6.Joshua Meier, Roshan Rao, Robert Verkuil, Jason Liu, Tom Sercu, and Alex Rives. Language models enable zero-shot prediction of the effects of mutations on protein function. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 29287–29303. Curran Associates, Inc., 2021.
  7. 7.Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123–1130, 2023.
  8. 8.Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, et al. Prottrans: Toward understanding the language of life through self-supervised learning. IEEE transactions on pattern analysis and machine intelligence, 44(10):7112–7127, 2021.
  9. 9.Nadav Brandes, Dan Ofer, Yam Peleg, Nadav Rappoport, and Michal Linial. Proteinbert: a universal deep-learning model of protein sequence and function. Bioinformatics, 38(8):2102– 2110, 2022.
  10. 10.Hedi Hegyi and Mark Gerstein. The relationship between protein structure and function: a comprehensive survey with application to the yeast genome 11edited by g. von heijne. Journal of Molecular Biology, 288(1):147–164, 1999.
  11. 11.John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  12. 12.Minkyung Baek, Frank DiMaio, Ivan Anishchenko, Justas Dauparas, Sergey Ovchinnikov, Gyu Rie Lee, Jue Wang, Qian Cong, Lisa N Kinch, R Dustin Schaeffer, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science, 373(6557):871–876, 2021.
  13. 13.Mihaly Varadi, Stephen Anyango, Mandar Deshpande, Sreenath Nair, Cindy Natassia, Galabina Yordanova, David Yuan, Oana Stroe, Gemma Wood, Agata Laydon, et al. Alphafold protein structure database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic acids research, 50(D1):D439–D444, 2022.
  14. 14.Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, and Fajie Yuan. Saprot: Protein language modeling with structure-aware vocabulary. In The Twelfth International Conference on Learning Representations, 2024.
  15. 15.Michael Heinzinger, Konstantin Weissenow, Joaquin Gomez Sanchez, Adrian Henkel, Martin Steinegger, and Burkhard Rost. Prostt5: Bilingual language model for protein sequence and structure. bioRxiv, pages 2023–07, 2023.
  16. 16.Michel Van Kempen, Stephanie S Kim, Charlotte Tumescheit, Milot Mirdita, Jeongjae Lee, Cameron LM Gilchrist, Johannes Söding, and Martin Steinegger. Fast and accurate protein structure search with foldseek. Nature Biotechnology, pages 1–4, 2023.
  17. 17.Hongyuan Lu, Daniel J Diaz, Natalie J Czarnecki, Congzhi Zhu, Wantae Kim, Raghav Shroff, Daniel J Acosta, Bradley R Alexander, Hannah O Cole, Yan Zhang, et al. Machine learningaided engineering of hydrolases for pet depolymerization. Nature, 604(7907):662–667, 2022.
  18. 18.Bowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend, and Ron Dror. Learning from protein structure with geometric vector perceptrons. In International Conference on Learning Representations, 2021.
  19. 19.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017.
  20. 20.Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Peter Chen, John Canny, Pieter Abbeel, and Yun Song. Evaluating protein transfer learning with tape. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019.
  21. 21.Pascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena Hurtado, Aidan N Gomez, Debora Marks, and Yarin Gal. Tranception: protein fitness prediction with autoregressive transformers and inference-time retrieval. In International Conference on Machine Learning, pages 16990–17017. PMLR, 2022.
  22. 22.Erik Nijkamp, Jeffrey A Ruffolo, Eli N Weinstein, Nikhil Naik, and Ali Madani. Progen2: exploring the boundaries of protein language models. Cell systems, 14(11):968–978, 2023.
  23. 23.Noelia Ferruz, Steffen Schmidt, and Birte Höcker. Protgpt2 is a deep unsupervised language model for protein design. Nature communications, 13(1):4348, 2022.
  24. 24.Ahmed Elnaggar, Hazem Essam, Wafaa Salah-Eldin, Walid Moustafa, Mohamed Elkerdawy, Charlotte Rochereau, and Burkhard Rost. Ankh: Optimized protein language model unlocks general-purpose modelling. arXiv preprint arXiv:2301.06568, 2023.
  25. 25.Bo Chen, Xingyi Cheng, Pan Li, Yangli-ao Geng, Jing Gong, Shen Li, Zhilei Bei, Xu Tan, Boyan Wang, Xin Zeng, et al. xtrimopglm: unified 100b-scale pre-trained transformer for deciphering the language of protein. arXiv preprint arXiv:2401.06199, 2024.
  26. 26.Vladimir Gligorijević, P Douglas Renfrew, Tomasz Kosciolek, Julia Koehler Leman, Daniel Berenberg, Tommi Vatanen, Chris Chandler, Bryn C Taylor, Ian M Fisk, Hera Vlamakis, et al. Structure-based protein function prediction using graph convolutional networks. Nature communications, 12(1):3168, 2021.
  27. 27.Yang Tan, Jia Zheng, Liang Hong, and Bingxin Zhou. Protsolm: Protein solubility prediction with multi-modal features. arXiv:2406.19744, 2024.
  28. 28.Bingxin Zhou, Lirong Zheng, Banghao Wu, Kai Yi, Bozitao Zhong, Yang Tan, Qian Liu, Pietro Liò, and Liang Hong. A conditional protein diffusion model generates artificial programmable endonuclease sequences with enhanced activity. Cell Discovery, 10(1):95, 2024.
  29. 29.Bingxin Zhou, Lirong Zheng, Banghao Wu, Yang Tan, Outongyi Lv, Kai Yi, Guisheng Fan, and Liang Hong. Protein engineering with lightweight graph denoising neural networks. Journal of Chemical Information and Modeling, 64(9):3650–3661, 2024.
  30. 30.Yang Tan, Lirong Zheng, Bozitao Zhong, Liang Hong, and Bingxin Zhou. Protein representation learning with sequence information embedding: Does it always lead to a better performance? arXiv:2406.19755, 2024.
  31. 31.Ruidong Wu, Fan Ding, Rui Wang, Rui Shen, Xiwen Zhang, Shitong Luo, Chenpeng Su, Zuofan Wu, Qi Xie, Bonnie Berger, et al. High-resolution de novo structure prediction from primary sequence. BioRxiv, pages 2022–07, 2022.
  32. 32.Zuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan, Aurelie Lozano, Payel Das, and Jian Tang. Protein representation learning by geometric structure pretraining. In The Eleventh International Conference on Learning Representations, 2023.
  33. 33.Zichen Wang, Steven A Combs, Ryan Brand, Miguel Romero Calvo, Panpan Xu, George Price, Nataliya Golovach, Emmanuel O Salawu, Colby J Wise, Sri Priya Ponnapalli, et al. Lm-gvp: an extensible sequence and structure informed deep learning framework for protein property prediction. Scientific reports, 12(1):6832, 2022.
  34. 34.Zuobai Zhang, Minghao Xu, Aurelie Lozano, Vijil Chenthamarakshan, Payel Das, and Jian Tang. Enhancing protein language model with structure-based encoder and pre-training. In ICLR 2023 - Machine Learning for Drug Discovery workshop, 2023.
  35. 35.Yang Tan, Bingxin Zhou, Lirong Zheng, Guisheng Fan, and Liang Hong. Semantical and topological protein encoding toward enhanced bioactivity and thermostability. bioRxiv, pages 2023–12, 2023.
  36. 36.Vıctor Garcia Satorras, Emiel Hoogeboom, and Max Welling. E (n) equivariant graph neural networks. In International conference on machine learning, pages 9323–9332. PMLR, 2021.
  37. 37.Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives. Learning inverse folding from millions of predicted structures. In International conference on machine learning, pages 8946–8970. PMLR, 2022.
  38. 38.Kevin K Yang, Niccolò Zanichelli, and Hugh Yeh. Masked inverse folding with sequence transfer for protein representation learning. Protein Engineering, Design and Selection, 36:gzad015, 2023.
  39. 39.Wolfgang Kabsch and Christian Sander. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features. Biopolymers: Original Research on Biomolecules, 22(12):2577–2637, 1983.
  40. 40.Joseph L Watson, David Juergens, Nathaniel R Bennett, Brian L Trippe, Jason Yim, Helen E Eisenach, Woody Ahern, Andrew J Borst, Robert J Ragotte, Lukas F Milles, et al. De novo design of protein structure and function with rfdiffusion. Nature, 620(7976):1089–1100, 2023.
  41. 41.Ian Sillitoe, Nicola Bordin, Natalie Dawson, Vaishali P Waman, Paul Ashford, Harry M Scholes, Camilla SM Pang, Laurel Woodridge, Clemens Rauer, Neeladri Sen, et al. Cath: increased structural coverage of functional space. Nucleic acids research, 49(D1):D266–D273, 2021.
  42. 42.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations, 2021.
  43. 43.Pascal Notin, Aaron Kollasch, Daniel Ritter, Lood van Niekerk, Steffanie Paul, Han Spinner, Nathan Rollins, Ada Shaw, Rose Orenbuch, Ruben Weitzman, Jonathan Frazer, Mafalda Dias, Dinko Franceschi, Yarin Gal, and Debora Marks. Proteingym: Large-scale benchmarks for protein fitness prediction and design. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 64331–64379. Curran Associates, Inc., 2023.
  44. 44.Daniel Hesslow, Niccoló Zanichelli, Pascal Notin, Iacopo Poli, and Debora Marks. Rita: a study on scaling up generative protein sequence models. arXiv preprint arXiv:2205.05789, 2022.
  45. 45.Ethan C Alley, Grigory Khimulya, Surojit Biswas, Mohammed AlQuraishi, and George M Church. Unified rational protein engineering with sequence-based deep representation learning. Nature methods, 16(12):1315–1322, 2019.
  46. 46.Céline Marquet, Michael Heinzinger, Tobias Olenyi, Christian Dallago, Kyra Erckert, Michael Bernhofer, Dmitrii Nechaev, and Burkhard Rost. Embeddings from protein language models predict conservation and variant effects. Human genetics, 141(10):1629–1647, 2022.
  47. 47.Elodie Laine, Yasaman Karami, and Alessandra Carbone. Gemme: a simple and fast global epistatic model predicting mutational effects. Molecular biology and evolution, 36(11):2604– 2619, 2019.
  48. 48.Roshan M Rao, Jason Liu, Robert Verkuil, Joshua Meier, John Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives. Msa transformer. In International Conference on Machine Learning, pages 8844–8856. PMLR, 2021.
  49. 49.Jonathan Frazer, Pascal Notin, Mafalda Dias, Aidan Gomez, Joseph K Min, Kelly Brock, Yarin Gal, and Debora S Marks. Disease variant prediction with deep generative models of evolutionary data. Nature, 599(7883):91–95, 2021.
  50. 50.Jung-Eun Shin, Adam J Riesselman, Aaron W Kollasch, Conor McMahon, Elana Simon, Chris Sander, Aashish Manglik, Andrew C Kruse, and Debora S Marks. Protein design and variant prediction using autoregressive generative models. Nature communications, 12(1):2403, 2021.
  51. 51.Adam J Riesselman, John B Ingraham, and Debora S Marks. Deep generative models of genetic variation capture the effects of mutations. Nature methods, 15(10):816–822, 2018.
  52. 52.Pascal Notin, Lood Van Niekerk, Aaron W Kollasch, Daniel Ritter, Yarin Gal, and Debora Susan Marks. TranceptEVE: Combining family-specific and family-agnostic models of protein sequences for improved fitness prediction. In NeurIPS 2022 Workshop on Learning Meaningful Representations of Life, 2022.
  53. 53.Thomas A Hopf, John B Ingraham, Frank J Poelwijk, Charlotta PI Schärfe, Michael Springer, Chris Sander, and Debora S Marks. Mutation effects predicted from sequence co-variation. Nature biotechnology, 35(2):128–135, 2017.
  54. 54.Raghav Shroff, Austin W Cole, Daniel J Diaz, Barrett R Morrow, Isaac Donnell, Ankur Annapareddy, Jimmy Gollihar, Andrew D Ellington, and Ross Thyer. Discovery of novel gain-of-function mutations guided by structure-based deep learning. ACS synthetic biology, 9(11):2927–2935, 2020.
  55. 55.Christian Dallago, Jody Mou, Kadina E Johnston, Bruce Wittmann, Nick Bhattacharya, Samuel Goldman, Ali Madani, and Kevin K Yang. FLIP: Benchmark tasks in fitness landscape inference for proteins. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021.
  56. 56.José Juan Almagro Armenteros, Casper Kaae Sønderby, Søren Kaae Sønderby, Henrik Nielsen, and Ole Winther. Deeploc: prediction of protein subcellular localization using deep learning. Bioinformatics, 33(21):3387–3395, 2017.
  57. 57.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019.
  58. 58.Minghao Xu, Zuobai Zhang, Jiarui Lu, Zhaocheng Zhu, Yangtian Zhang, Ma Chang, Runcheng Liu, and Jian Tang. Peer: A comprehensive and multi-task benchmark for protein sequence understanding. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 35156–35173. Curran Associates, Inc., 2022.

Citation

MLA
Li, M., et al. “ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 35700–26, https://proceedings.neurips.cc/paper_files/paper/2024/file/3ed57b293db0aab7cc30c44f45262348-Paper-Conference.pdf.
APA
Li, M., Tan, Y., Ma, X., Zhong, B., Yu, H., Zhou, Z., Ouyang, W., Zhou, B., Tan, P., & Hong, L. (2024). ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention. Advances in Neural Information Processing Systems, 37, 35700–35726. https://proceedings.neurips.cc/paper_files/paper/2024/file/3ed57b293db0aab7cc30c44f45262348-Paper-Conference.pdf
Chicago
Li, M., Y. Tan, X. Ma, et al. 2024. “ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention”. Advances in Neural Information Processing Systems 37: 35700–35726. https://proceedings.neurips.cc/paper_files/paper/2024/file/3ed57b293db0aab7cc30c44f45262348-Paper-Conference.pdf.
Harvard
Li, M. et al. (2024) “ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 35700–35726. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/3ed57b293db0aab7cc30c44f45262348-Paper-Conference.pdf.
Vancouver
1. Li M, Tan Y, Ma X, Zhong B, Yu H, Zhou Z, Ouyang W, Zhou B, Tan P, Hong L (2024) ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 35700–35726

BibTeX

@inproceedings{li2024prosst,
  title = {ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention},
  author = {Li, Mingchen and Tan, Yang and Ma, Xinzhu and Zhong, Bozitao and Yu, Huiqun and Zhou, Ziyi and Ouyang, Wanli and Zhou, Bingxin and Tan, Pan and Hong, Liang},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {35700-35726},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/3ed57b293db0aab7cc30c44f45262348-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors