ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts

Minghao XuXinyu YuanSantiago MiretJian Tang

article2023ICML144 citations

Proposes a multimodal pre-training framework that aligns protein sequences with biomedical text descriptions to improve protein language models and enable zero-shot protein classification and functional retrieval.

Listen

Understanding protein functions and properties is central to advancing drug discovery and healthcare. While recent machine learning approaches use protein language models to learn directly from amino acid sequences, these conventional models rely primarily on evolutionary patterns and struggle to explicitly capture real-world functional behaviors and cellular locations. Meanwhile, vast amounts of rich biomedical text describing protein functions exist but remain largely unintegrated during model training.

The article demonstrates that enriching protein sequence pre-training with biomedical text descriptions significantly enhances functional understanding and model accuracy. Specifically, it introduces a multimodal training framework named ProtST, which jointly aligns protein sequences with their corresponding textual property descriptions to improve supervised predictions and enable zero-shot applications.

To evaluate this concept, the authors constructed a dataset of over 553,000 paired protein sequences and curated property texts covering names, functions, cellular locations, and protein families. The framework combines existing protein models with a biomedical text language model across three objectives: predicting masked sequence elements to retain baseline evolutionary patterns, aligning overall sequence and text representations via contrastive learning, and predicting masked residues and words through a cross-modal fusion layer. The models were tested across eleven benchmark tasks spanning localization, fitness landscape prediction, and function annotation, as well as zero-shot classification and text-based database retrieval.

The findings show substantial performance improvements across downstream applications. Enhancing baseline models with the multimodal framework improved performance on nearly all evaluated metrics, with the enhanced models setting new state-of-the-art benchmarks on protein fitness and function annotation. In zero-shot classification settings with no labeled training examples, the model matched or exceeded the accuracy of conventional models trained with several labeled examples per class. Furthermore, using zero-shot predictions to refine standard supervised classifiers consistently improved full-dataset accuracy, while text-to-protein retrieval successfully identified functional ligand binders from large databases using only natural language prompts.

These results indicate that integrating textual biomedical knowledge into sequence models reduces dependency on costly, labor-intensive experimental labeling while improving predictive accuracy. Organizations involved in protein engineering and therapeutic development can leverage this approach to accelerate lead discovery, rapidly query unannotated sequence repositories, and improve model performance even when experimental data is scarce.

Moving forward, adopting multimodal sequence-text models is recommended when developing protein representation pipelines or screening databases for specific functional properties. To build on these results, future efforts should expand the text dataset using broader biomedical literature, incorporate 3D protein structural coordinates into the multimodal framework, and explore text-guided generative protein design.

Confidence in these findings is supported by consistent performance gains across diverse standard benchmarks and multiple baseline architectures. However, decision-makers should note that the current training data relies on curated Swiss-Prot annotations, which limits coverage across the entire protein universe. Additionally, data quality proved critical, as pre-training on larger but lower-quality automated text annotations led to measurable performance declines.

Cover for ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts

Abstract

Current protein language models (PLMs) learn protein representations mainly based on their sequences, thereby well capturing co-evolutionary information, but they are unable to explicitly acquire protein functions, which is the end goal of protein representation learning. Fortunately, for many proteins, their textual property descriptions are available, where their various functions are also described. Motivated by this fact, we first build the ProtDescribe dataset to augment protein sequences with text descriptions of their functions and other important properties. Based on this dataset, we propose the ProtST framework to enhance Protein Sequence pre-training and understanding by biomedical Texts. During pre-training, we design three types of tasks, i.e., unimodal mask prediction, multimodal representation alignment and multimodal mask prediction, to enhance a PLM with protein property information with different granularities and, at the same time, preserve the PLM’s original representation power. On downstream tasks, ProtST enables both supervised learning and zero-shot prediction. We verify the superiority of ProtST-induced PLMs over previous ones on diverse representation learning benchmarks. Under the zero-shot setting, we show the effectiveness of ProtST on zero-shot protein classification, and ProtST also enables functional protein retrieval from a large-scale database without any function annotation. Source code and model weights are available at https://github.com/DeepGraphLearning/ProtST.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 2.1. Problem Definition
  • 2.2. Protein Language Models
  • 2.3. Biomedical Language Models
  • 3. Method
  • 3.1. Motivation and Overview
  • 3.2. Pre-training Tasks: Joint Modeling of Protein Sequences and Biomedical Texts
  • 3.3. Discussion
  • 4. Experiments
  • 4.1. Pre-training Setups
  • 4.2. Representation Learning
  • 4.2.1. EXPERIMENTAL SETUPS
  • 4.2.2. EXPERIMENTAL RESULTS
  • 4.3. Zero-shot Protein Classification
  • 4.3.1. EXPERIMENTAL SETUPS
  • 4.3.2. DATA EFFICIENCY OF ZERO-SHOT CLASSIFIER
  • 4.3.3. ENHANCING SUPERVISED LEARNING WITH ZERO-SHOT CLASSIFIER
  • 4.4. Zero-shot Text-to-Protein Retrieval
  • 4.5. Ablation Study
  • 5. Related Work
  • 6. Conclusions and Future Work
  • Acknowledgments
  • References
  • A. Model Architecture for Pre-training
  • B. More Experimental Setups
  • B.1. More Pre-training Setups
  • B.2. More Representation Learning Setups
  • B.3. More Zero-shot Protein Classification Setups
  • C. Experimental Results on ProteinGym
  • C.1. Comparisons of Protein Language Models (PLMs)
  • C.2. Comparisons with Alignment-based Methods
  • D. More Zero-shot Text-to-Protein Retrieval Results
  • E. More Ablation Study
  • E.1. Ablation Study of Pre-training Losses
  • E.2. Ablation Study of Biomedical Language Model
  • F. More Visualization

Knowls

  1. Knowl 1 — ProtST Pre-training Objective and Multi-Modal Architecture

    model/method

    ProtST is a multimodal pre-training framework that aligns protein sequences with biomedical text property descriptions. Given a protein P=(S,T)P = (S, T) consisting of an amino acid sequence S=[s1,s2,…,sn]S = [s_1, s_2, \dots, s_n] of length nn and a text description T=[t1,t2,…,tm]T = [t_1, t_2, \dots, t_m] of length mm, the model jointly optimizes three pre-training objectives:

    min⁡θLMPM+LGC+LMMP\min_\theta \mathcal{L}_{\text{MPM}} + \mathcal{L}_{\text{GC}} + \mathcal{L}_{\text{MMP}}

    where θ\theta denotes all trainable parameters, and the three losses serve distinct complementary purposes:

    1. Unimodal Masked Protein Modeling (LMPM\mathcal{L}_{\text{MPM}}): Predicts 15% masked amino acid residues using contextual representations from the protein language model (PLM), preserving the model's sequence-level evolutionary representation power.
    2. Multimodal Representation Alignment (LGC\mathcal{L}_{\text{GC}}): Employs a global contrastive InfoNCE loss to align whole-protein sequence embeddings from the PLM with text representations from a pre-trained biomedical language model (BLM).
    3. Multimodal Mask Prediction (LMMP\mathcal{L}_{\text{MMP}}): Uses a cross-attention fusion layer to model fine-grained residue-word interdependencies by predicting masked residues and words from cross-modal fused representations.

    During pre-training, the BLM (e.g., PubMedBERT) weights are kept frozen, while the PLM (e.g., ProtBert, ESM-1b, or ESM-2), the single-layer cross-attention fusion module, and task projection heads are trained.

  2. Knowl 2 — Global Contrastive Loss for Protein-Text Alignment

    equation

    To align protein sequence representations with their corresponding biomedical text descriptions across a batch of MM paired proteins {Pi=(Si,Ti)}i=1M\{P_i = (S_i, T_i)\}_{i=1}^M, the global contrastive (GC) loss LGC\mathcal{L}_{\text{GC}} is defined as a symmetric InfoNCE loss:

    LGC=−12M∑i=1M(log⁡exp⁡(ziS⋅ziT/τ)∑j=1Mexp⁡(ziS⋅zjT/τ)+log⁡exp⁡(ziS⋅ziT/τ)∑j=1Mexp⁡(zjS⋅ziT/τ))\mathcal{L}_{\text{GC}} = -\frac{1}{2M} \sum_{i=1}^M \left( \log \frac{\exp(z_i^S \cdot z_i^T / \tau)}{\sum_{j=1}^M \exp(z_i^S \cdot z_j^T / \tau)} + \log \frac{\exp(z_i^S \cdot z_i^T / \tau)}{\sum_{j=1}^M \exp(z_j^S \cdot z_i^T / \tau)} \right)

    where:

    • ziS∈Rdz_i^S \in \mathbb{R}^d is the normalized projection of the sequence representation for protein ii, produced by passing the PLM's pooled output through a two-layer MLP with ReLU activation.
    • ziT∈Rdz_i^T \in \mathbb{R}^d is the normalized projection of the text description representation for protein ii, produced by passing the BLM's pooled output through another two-layer MLP with ReLU activation.
    • τ∈R+\tau \in \mathbb{R}^+ is a learnable temperature parameter initialized to 0.070.07.
    • In multi-GPU distributed data-parallel training, negative pairs are gathered across all GPUs.
  3. Knowl 3 — Multimodal Cross-Attention Fusion and Mask Prediction

    model/method

    To capture token-level associations between amino acid residues and descriptive biomedical words (such as polar residue co-occurrence with solubility terms), ProtST uses a multimodal mask prediction (MMP) module. Given corrupted sequence SS and corrupted text TT with 15% random masking each, sequence residue representations ZS=[z1s,…,zns]∈Rn×dZ^S = [z_1^s, \dots, z_n^s] \in \mathbb{R}^{n \times d} and word representations ZT=[z1t,…,zmt]∈Rm×dZ^T = [z_1^t, \dots, z_m^t] \in \mathbb{R}^{m \times d} are extracted by the PLM and BLM, respectively.

    Query, key, and value matrices are projected using learnable parameter matrices WqS,WkS,WvS,WqT,WkT,WvT∈Rd×dW_q^S, W_k^S, W_v^S, W_q^T, W_k^T, W_v^T \in \mathbb{R}^{d \times d}: QS=ZSWqS,KS=ZSWkS,VS=ZSWvSQ^S = Z^S W_q^S, \quad K^S = Z^S W_k^S, \quad V^S = Z^S W_v^S QT=ZTWqT,KT=ZTWkT,VT=ZTWvTQ^T = Z^T W_q^T, \quad K^T = Z^T W_k^T, \quad V^T = Z^T W_v^T

    A single fusion layer with 8 attention heads combines self-attention and cross-attention: Z~S=12(MHA(QS,KS,VS)+MHA(QS,KT,VT))\tilde{Z}^S = \frac{1}{2} \left( \text{MHA}(Q^S, K^S, V^S) + \text{MHA}(Q^S, K^T, V^T) \right) Z~T=12(MHA(QT,KT,VT)+MHA(QT,KS,VS))\tilde{Z}^T = \frac{1}{2} \left( \text{MHA}(Q^T, K^T, V^T) + \text{MHA}(Q^T, K^S, V^S) \right)

    where MHA(⋅,⋅,⋅)\text{MHA}(\cdot, \cdot, \cdot) is multi-head attention. Two-layer MLPs with ReLU nonlinearity then predict the masked residue types from Z~S\tilde{Z}^S (with cross-entropy loss LMMPS\mathcal{L}_{\text{MMP}}^S) and masked word tokens from Z~T\tilde{Z}^T (with cross-entropy loss LMMPT\mathcal{L}_{\text{MMP}}^T), forming the joint loss LMMP=LMMPS+LMMPT\mathcal{L}_{\text{MMP}} = \mathcal{L}_{\text{MMP}}^S + \mathcal{L}_{\text{MMP}}^T.

  4. Knowl 4 — The ProtDescribe Dataset

    definition

    ProtDescribe is a paired protein sequence and textual property description dataset curated from Swiss-Prot, containing 553,052 protein sequence-description pairs. Each textual description combines four structured annotation fields with dedicated prefixes, concatenated in the following order (omitting absent fields):

    1. PROTEIN NAME: Recommended full protein name (100% dataset coverage, 553,052 samples).
    2. FUNCTION: Descriptions of molecular and cellular functions (83.3% dataset coverage, 460,936 samples).
    3. SUBCELLULAR LOCATION: Cellular location and topology of the mature protein (63.5% dataset coverage, 350,929 samples).
    4. SIMILARITY: Protein family and structural homology annotations (92.6% dataset coverage, 512,276 samples).
  5. Knowl 5 — Zero-Shot Protein Classification and Supervised-Zero-Shot Logit Ensembling

    model/method

    Due to the shared representation space between the PLM and BLM, ProtST performs zero-shot protein classification without task-specific training data. Given a test protein sequence SS and textual descriptions {Ti}i=1K\{T_i\}_{i=1}^K for KK classes, the zero-shot classification logit for class ii is:

    yizero=zS⋅ziTτy_i^{\text{zero}} = \frac{z^S \cdot z_i^T}{\tau}

    where zSz^S is the normalized PLM sequence embedding, ziTz_i^T is the normalized BLM text embedding of prompt TiT_i, and τ\tau is the learned temperature. Pre-training prompt formatting (e.g., SUBCELLULAR LOCATION: {label} for localization or FUNCTION: {Name} {AlterNames} for enzyme reactions) provides superior zero-shot accuracy compared to bare label names or free-form natural language templates.

    When labeled downstream data is available, zero-shot predictions can refine supervised model decision boundaries via logit ensembling:

    yk=yksup+αykzero,k∈{1,…,K}y_k = y_k^{\text{sup}} + \alpha y_k^{\text{zero}}, \quad k \in \{1, \dots, K\}

    where yksupy_k^{\text{sup}} is the logit from the supervised classifier, and the ensemble weight α\alpha is set to the ratio of the zero-shot classifier's validation performance to the supervised model's validation performance.

  6. Knowl 6 — Zero-Shot Text-to-Protein Retrieval

    model/method

    ProtST enables functional protein retrieval from an unannotated database of NN protein sequences without requiring function labels. The procedure operates as follows:

    1. Precompute normalized sequence embeddings {ziS}i=1N\{z_i^S\}_{i=1}^N for all database proteins using the ProtST-trained PLM.
    2. Given a natural language query describing a target function TT (formatted with the prefix FUNCTION: followed by the function definition), compute the query embedding zT=BLM(T)z^T = \text{BLM}(T).
    3. Compute similarity scores ϵi=ziS⋅zT\epsilon_i = z_i^S \cdot z^T for each protein i∈{1,…,N}i \in \{1, \dots, N\}.
    4. Rank database proteins in descending order of ϵi\epsilon_i.
  7. Knowl 7 — Downstream Performance Across Localization, Fitness, and Function Benchmarks

    data/table

    ProtST pre-training improves downstream performance across 11 benchmark tasks covering subcellular localization (DeepLoc Binary and Subcellular), fitness landscape prediction (β\beta-lactamase, AAV, Thermostability, Fluorescence, Stability), and protein function annotation (Enzyme Commission EC, Gene Ontology GO-BP, GO-MF, GO-CC). ProtST-ProtBert consistently outperforms OntoProtein, ProtST-ESM-1b achieves the highest overall fitness prediction scores, and ProtST-ESM-2 obtains top performance in localization and function annotation.

    Model Localization (Acc%) Fitness Prediction (Spearman's ρ\rho)
    Bin Sub β\beta-lac AAV Thermo Flu Sta Mean ρ\rho
    Fixed-Encoder Evaluation
    ProtBert 81.54 59.44 0.616 0.209 0.562 0.339 0.697 0.485
    OntoProtein 84.87 68.34 0.471 0.217 0.605 0.432 0.688 0.483
    ESM-1b 91.61 79.82 0.528 0.454 0.674 0.430 0.750 0.567
    ESM-2 91.32 80.84 0.559 0.374 0.677 0.456 0.746 0.562
    ProtST-ProtBert 92.29 78.49 0.569 0.219 0.621 0.376 0.719 0.501
    ProtST-ESM-1b 92.87 82.00 0.578 0.460 0.680 0.523 0.766 0.601
    ProtST-ESM-2 92.52 83.39 0.565 0.398 0.681 0.499 0.776 0.584
    Full-Model Tuning Evaluation
    ProtBert 91.32 76.53 0.731 0.794 0.660 0.679 0.771 0.727
    OntoProtein 92.47 77.59 0.757 0.791 0.662 0.630 0.731 0.714
    ESM-1b 92.40 78.13 0.839 0.821 0.669 0.679 0.694 0.740
    ESM-2 91.72 78.67 0.867 0.817 0.672 0.677 0.718 0.750
    ProtST-ProtBert 91.78 78.71 0.863 0.804 0.673 0.679 0.745 0.753
    ProtST-ESM-1b 92.35 78.73 0.895 0.850 0.681 0.682 0.751 0.772
    ProtST-ESM-2 92.52 80.22 0.879 0.825 0.682 0.682 0.738 0.761
    Model EC GO-BP GO-MF GO-CC
    (Full-Model Tuning) AUPR FmaxF_{\text{max}} AUPR FmaxF_{\text{max}} AUPR FmaxF_{\text{max}} AUPR FmaxF_{\text{max}}
    ProtBert 0.859 0.838 0.188 0.279 0.464 0.456 0.234 0.408
    OntoProtein 0.854 0.841 0.284 0.436 0.603 0.631 0.300 0.441
    ESM-1b 0.884 0.869 0.332 0.452 0.630 0.659 0.324 0.477
    ESM-2 0.888 0.874 0.340 0.472 0.643 0.662 0.350 0.472
    ProtST-ProtBert 0.876 0.856 0.286 0.440 0.615 0.648 0.314 0.449
    ProtST-ESM-1b 0.894 0.878 0.328 0.480 0.644 0.661 0.364 0.488
    ProtST-ESM-2 0.898 0.878 0.342 0.482 0.647 0.668 0.364 0.487
  8. Knowl 8 — Ablation Analysis of ProtST Pre-training Losses

    empirical result

    Ablation experiments on ProtST-ESM-1b evaluate the contribution of each pre-training objective (Unimodal Masked Protein Modeling LMPM\mathcal{L}_{\text{MPM}}, Global Contrastive alignment LGC\mathcal{L}_{\text{GC}}, and Multimodal Mask Prediction LMMP\mathcal{L}_{\text{MMP}}) across downstream tasks:

    • Full loss: Mean Localization Accuracy = 87.44% (Fix-enc) / 85.54% (Full-m); Mean Fitness Spearman ρ\rho = 0.601 (Fix-enc) / 0.772 (Full-m); Mean Function FmaxF_{\text{max}} = 0.627.
    • Without LMPM\mathcal{L}_{\text{MPM}}: Fixed-encoder fitness ρ\rho drops by 1.33% (to 0.593) and full-tuning localization drops by 0.49% (to 85.12%), confirming that unimodal masked modeling preserves sequence evolutionary priors.
    • Without LGC\mathcal{L}_{\text{GC}}: Causes the largest decay on fixed-encoder localization (−1.26%-1.26\%, falling to 86.34%) and fixed-encoder fitness (−3.66%-3.66\%, falling to 0.579), as well as reducing function FmaxF_{\text{max}} to 0.613 (−2.23%-2.23\%).
    • Without LMMP\mathcal{L}_{\text{MMP}}: Diminishes full-model fitness ρ\rho by 2.72% (to 0.751) and function FmaxF_{\text{max}} by 1.91% (to 0.615), showing that token-level cross-attention fine-grained supervision is crucial for residue-level tasks.

    Removing any single loss degrades performance across 16 to 20 out of 24 benchmark metrics, proving that all three losses are necessary.

  9. Knowl 9 — Curated Annotation Quality Dominates Pre-training Corpus Size

    empirical result

    Evaluating the pre-training data source on ProtST-ESM-1b demonstrates that annotation quality and field coverage outweigh raw sequence quantity for multimodal protein pre-training:

    • Swiss-Prot: High-quality human-curated annotations across 553,052 proteins with 83.3% function coverage, 63.5% location coverage, and 92.6% family coverage.
    • TrEMBL: Over 200M computationally annotated proteins with only 24.0% function coverage, 51.5% location coverage, and 78.0% family coverage.

    Downstream evaluation shows ProtST-ESM-1b pre-trained on Swiss-Prot outperforms the version pre-trained on TrEMBL on both fixed-encoder and full-model settings:

    • Mean Localization Accuracy: Swiss-Prot achieves 87.44% (Fix-enc) and 85.54% (Full-m), vs. TrEMBL's 86.68% (Fix-enc) and 85.13% (Full-m).
    • Mean Fitness Spearman ρ\rho: Swiss-Prot achieves 0.601 (Fix-enc) and 0.772 (Full-m), vs. TrEMBL's 0.597 (Fix-enc) and 0.762 (Full-m).
  10. Knowl 10 — ProteinGym Zero-Shot Fitness Prediction and Homology Hybrid Ensembling

    empirical result

    On the ProteinGym substitution benchmark evaluating deep mutational scanning fitness landscapes, ProtST-ESM-1b achieves a UniProt-level mean Spearman's ρ\rho of 0.412, outperforming other PLMs evaluated without retrieval: ESM-1b (0.358, a 15.1% relative improvement), ESM-1v (0.372), Tranception L w/o retrieval (0.401), and Progen2 XL (0.402).

    While multiple sequence alignment (MSA) homology-based methods like EVE (ρ=0.443\rho = 0.443) and GEMME (ρ=0.459\rho = 0.459) outperform standalone PLMs by exploiting homologous sequence alignments, combining normalized predictions of ProtST-ESM-1b and GEMME into a hybrid ensemble yields ρ=0.464\rho = 0.464, surpassing both individual alignment-based methods.

Coverage note — None was omitted; all key contributions including dataset construction, loss equations, architectural designs, downstream zero-shot/supervised methods, and empirical benchmark/ablation results are covered.

References

  1. 1.Almagro Armenteros, J. J., Sønderby, C. K., Sønderby, S. K., Nielsen, H., and Winther, O. Deeploc: prediction of protein subcellular localization using deep learning. Bioinformatics, 33(21):3387–3395, 2017.
  2. 2.Baek, M., DiMaio, F., Anishchenko, I., Dauparas, J., Ovchinnikov, S., Lee, G. R., Wang, J., Cong, Q., Kinch, L. N., Schaeffer, R. D., et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science, 373(6557):871–876, 2021.
  3. 3.Bairoch, A. and Apweiler, R. The swiss-prot protein sequence database and its supplement trembl in 2000. Nucleic acids research, 28(1):45–48, 2000.
  4. 4.Beltagy, I., Lo, K., and Cohan, A. Scibert: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676, 2019.
  5. 5.Bhardwaj, N. and Lu, H. Correlation between gene expression profiles and protein–protein interactions within and across genomes. Bioinformatics, 21(11):2730–2738, 2005.
  6. 6.Bileschi, M. L., Belanger, D., Bryant, D., Sanderson, T., Carter, B., Sculley, D., DePristo, M. A., and Colwell, L. J. Using deep learning to annotate the protein universe. BioRxiv, pp. 626507, 2019.
  7. 7.Biswas, S., Khimulya, G., Alley, E. C., Esvelt, K. M., and Church, G. M. Low-n protein engineering with data-efficient deep learning. Nature methods, 18(4):389–396, 2021.
  8. 8.Canese, K. and Weis, S. Pubmed: the bibliographic database. The NCBI handbook, 2(1), 2013.
  9. 9.Capaldi, R. A. and Vanderkooi, G. The low polarity of many membrane proteins. Proceedings of the National Academy of Sciences, 69(4):930–932, 1972.
  10. 10.Chang, A., Jeske, L., Ulbrich, S., Hofmann, J., Koblitz, J., Schomburg, I., Neumann-Schaal, M., Jahn, D., and Schomburg, D. Brenda, the elixir core data resource in 2021: new developments and updates. Nucleic Acids Research, 49(D1):D498–D508, 2021.
  11. 11.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pp. 1597–1607. PMLR, 2020.
  12. 12.Chung, Y.-A., Zhu, C., and Zeng, M. Splat: Speech-language joint pre-training for spoken language understanding. arXiv preprint arXiv:2010.02295, 2020.
  13. 13.Consortium, U. Uniprot: a worldwide hub of protein knowledge. Nucleic acids research, 47(D1):D506–D515, 2019.
  14. 14.Dallago, C., Mou, J., Johnston, K. E., Wittmann, B. J., Bhattacharya, N., Goldman, S., Madani, A., and Yang, K. K. Flip: Benchmark tasks in fitness landscape inference for proteins. bioRxiv, 2021.
  15. 15.Edwards, C., Zhai, C., and Ji, H. Text2mol: Cross-modal molecule retrieval with natural language queries. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 595–607, 2021.
  16. 16.Elnaggar, A., Heinzinger, M., Dallago, C., Rihawi, G., Wang, Y., Jones, L., Gibbs, T., Feher, T., Angerer, C., Steinegger, M., et al. Prottrans: towards cracking the language of life’s code through self-supervised deep learning and high performance computing. arXiv preprint arXiv:2007.06225, 2020.
  17. 17.Frazer, J., Notin, P., Dias, M., Gomez, A., Min, J. K., Brock, K., Gal, Y., and Marks, D. S. Disease variant prediction with deep generative models of evolutionary data. Nature, 599(7883):91–95, 2021.
  18. 18.Gainza, P., Sverrisson, F., Monti, F., Rodola, E., Boscaini, D., Bronstein, M., and Correia, B. Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning. Nature Methods, 17(2):184–192, 2020.
  19. 19.Gligorijević, V., Renfrew, P. D., Kosciolek, T., Leman, J. K., Berenberg, D., Vatanen, T., Chandler, C., Taylor, B. C., Fisk, I. M., Vlamakis, H., et al. Structure-based protein function prediction using graph convolutional networks. Nature communications, 12(1):1–14, 2021.
  20. 20.Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., and Poon, H. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), 3(1):1–23, 2021.
  21. 21.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9729–9738, 2020.
  22. 22.Hermosilla, P., Schäfer, M., Lang, M., Fackelmann, G., Vázquez, P. P., Kozlíková, B., Krone, M., Ritschel, T., and Ropinski, T. Intrinsic-extrinsic convolution and pooling for learning on 3d protein structures. arXiv preprint arXiv:2007.06252, 2020.
  23. 23.Jin, Q., Dhingra, B., Cohen, W. W., and Lu, X. Probing biomedical embeddings from language models. arXiv preprint arXiv:1904.02181, 2019.
  24. 24.Jing, B., Eismann, S., Suriana, P., Townshend, R. J., and Dror, R. Learning from protein structure with geometric vector perceptrons. arXiv preprint arXiv:2009.01411, 2020.
  25. 25.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al. Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589, 2021.
  26. 26.Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M. Generalization through memorization: Nearest neighbor language models. arXiv preprint arXiv:1911.00172, 2019.
  27. 27.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  28. 28.Kumar, S., Tsai, C.-J., and Nussinov, R. Factors enhancing protein thermostability. Protein engineering, 13(3):179–191, 2000.
  29. 29.Laine, E., Karami, Y., and Carbone, A. Gemme: a simple and fast global epistatic model predicting mutational effects. Molecular biology and evolution, 36(11):2604–2619, 2019.
  30. 30.Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., and Kang, J. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240, 2020.
  31. 31.Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., dos Santos Costa, A., Fazel-Zarandi, M., Sercu, T., Candido, S., et al. Language models of protein sequences at the scale of evolution enable accurate structure prediction. bioRxiv, 2022.
  32. 32.Liu, S., Nie, W., Wang, C., Lu, J., Qiao, Z., Liu, L., Tang, J., Xiao, C., and Anandkumar, A. Multi-modal molecule structure-text model for text-based retrieval and editing. arXiv preprint arXiv:2212.10789, 2022.
  33. 33.Lu, A. X., Zhang, H., Ghassemi, M., and Moses, A. M. Self-supervised contrastive learning of protein representations by mutual information maximization. BioRxiv, 2020.
  34. 34.Luo, H., Ji, L., Shi, B., Huang, H., Duan, N., Li, T., Li, J., Bharti, T., and Zhou, M. Univl: A unified video and language pre-training model for multimodal understanding and generation. arXiv preprint arXiv:2002.06353, 2020.
  35. 35.Madani, A., McCann, B., Naik, N., Keskar, N. S., Anand, N., Eguchi, R. R., Huang, P.-S., and Socher, R. Progen: Language modeling for protein generation. arXiv preprint arXiv:2004.03497, 2020.
  36. 36.Marquet, C., Heinzinger, M., Olenyi, T., Dallago, C., Erckert, K., Bernhofer, M., Nechaev, D., and Rost, B. Embeddings from protein language models predict conservation and variant effects. Human genetics, 141(10):1629–1647, 2022.
  37. 37.Meier, J., Rao, R., Verkuil, R., Liu, J., Sercu, T., and Rives, A. Language models enable zero-shot prediction of the effects of mutations on protein function. bioRxiv, 2021.
  38. 38.Murzin, A. G., Brenner, S. E., Hubbard, T., and Chothia, C. Scop: a structural classification of proteins database for the investigation of sequences and structures. Journal of molecular biology, 247(4):536–540, 1995.
  39. 39.Nijkamp, E., Ruffolo, J., Weinstein, E. N., Naik, N., and Madani, A. Progen2: exploring the boundaries of protein language models. arXiv preprint arXiv:2206.13517, 2022.
  40. 40.Notin, P., Dias, M., Frazer, J., Hurtado, J. M., Gomez, A. N., Marks, D., and Gal, Y. Tranception: protein fitness prediction with autoregressive transformers and inference-time retrieval. In International Conference on Machine Learning, pp. 16990–17017. PMLR, 2022.
  41. 41.Oord, A. v. d., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  42. 42.Organization, W. H. and University, U. N. Protein and amino acid requirements in human nutrition, volume 935. World Health Organization, 2007.
  43. 43.Qian, Y., Bianv, X., Shi, Y., Kanda, N., Shen, L., Xiao, Z., and Zeng, M. Speech-language pre-training for end-to-end spoken language understanding. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7458–7462. IEEE, 2021.
  44. 44.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pp. 8748–8763. PMLR, 2021.
  45. 45.Rao, R., Bhattacharya, N., Thomas, N., Duan, Y., Chen, P., Canny, J., Abbeel, P., and Song, Y. Evaluating protein transfer learning with tape. Advances in neural information processing systems, 32, 2019.
  46. 46.Rives, A., Meier, J., Sercu, T., Goyal, S., Lin, Z., Liu, J., Guo, D., Ott, M., Zitnick, C. L., Ma, J., et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences, 118(15), 2021.
  47. 47.Shanehsazzadeh, A., Belanger, D., and Dohan, D. Is transfer learning necessary for protein landscape prediction? arXiv preprint arXiv:2011.03443, 2020.
  48. 48.Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., and Kiela, D. Flava: A foundational language and vision alignment model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15638–15650, 2022.
  49. 49.Steinegger, M. and Söding, J. Clustering huge protein sequence sets in linear time. Nature communications, 9(1): 1–8, 2018.
  50. 50.Sverrisson, F., Feydy, J., Correia, B. E., and Bronstein, M. M. Fast end-to-end learning on protein surfaces. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15272–15281, 2021.
  51. 51.Teague, S. J. Implications of protein flexibility for drug discovery. Nature reviews Drug discovery, 2(7):527–541, 2003.
  52. 52.Trott, O. and Olson, A. J. Autodock vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. Journal of computational chemistry, 31(2):455–461, 2010.
  53. 53.Van der Maaten, L. and Hinton, G. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  54. 54.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  55. 55.Xu, H., Ghosh, G., Huang, P.-Y., Okhonko, D., Aghajanyan, A., Metze, F., Zettlemoyer, L., and Feichtenhofer, C. Videoclip: Contrastive pre-training for zero-shot video-text understanding. arXiv preprint arXiv:2109.14084, 2021.
  56. 56.Xu, M., Guo, Y., Xu, Y., Tang, J., Chen, X., and Tian, Y. Eurnet: Efficient multi-range relational modeling of spatial multi-relational data. arXiv preprint arXiv:2211.12941, 2022a.
  57. 57.Xu, M., Zhang, Z., Lu, J., Zhu, Z., Zhang, Y., Ma, C., Liu, R., and Tang, J. Peer: A comprehensive and multi-task benchmark for protein sequence understanding. arXiv preprint arXiv:2206.02096, 2022b.
  58. 58.Zhang, N., Bi, Z., Liang, X., Cheng, S., Hong, H., Deng, S., Lian, J., Zhang, Q., and Chen, H. Ontoprotein: Protein pretraining with gene ontology embedding. arXiv preprint arXiv:2201.11147, 2022a.
  59. 59.Zhang, Z., Xu, M., Jamasb, A., Chenthamarakshan, V., Lozano, A., Das, P., and Tang, J. Protein representation learning by geometric structure pretraining. arXiv preprint arXiv:2203.06125, 2022b.
  60. 60.Zhang, Z., Xu, M., Lozano, A., Chenthamarakshan, V., Das, P., and Tang, J. Physics-inspired protein encoder pretraining via siamese sequence-structure diffusion trajectory prediction. arXiv preprint arXiv:2301.12068, 2023.
  61. 61.Zhu, Z., Shi, C., Zhang, Z., Liu, S., Xu, M., Yuan, X., Zhang, Y., Chen, J., Cai, H., Lu, J., et al. Torchdrug: A powerful and flexible machine learning platform for drug discovery. arXiv preprint arXiv:2202.08320, 2022.

Citation

MLA
Xu, M., et al. “ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts”. International Conference on Machine Learning, vol. 202, 2023, pp. 38749–67, https://proceedings.mlr.press/v202/xu23t.html.
APA
Xu, M., Yuan, X., Miret, S., & Tang, J. (2023). ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts. International Conference on Machine Learning, 202, 38749–38767. https://proceedings.mlr.press/v202/xu23t.html
Chicago
Xu, M., X. Yuan, S. Miret, and J. Tang. 2023. “ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts”. International Conference on Machine Learning 202: 38749–67. https://proceedings.mlr.press/v202/xu23t.html.
Harvard
Xu, M. et al. (2023) “ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts”, International Conference on Machine Learning. PMLR, pp. 38749–38767. Available at: https://proceedings.mlr.press/v202/xu23t.html.
Vancouver
1. Xu M, Yuan X, Miret S, Tang J (2023) ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts. In: International Conference on Machine Learning. PMLR, pp 38749–38767

BibTeX

@InProceedings{pmlr-v202-xu23t,
  title = 	 {{P}rot{ST}: Multi-Modality Learning of Protein Sequences and Biomedical Texts},
  author =       {Xu, Minghao and Yuan, Xinyu and Miret, Santiago and Tang, Jian},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {38749--38767},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/xu23t/xu23t.pdf},
  url = 	 {https://proceedings.mlr.press/v202/xu23t.html},
  abstract = 	 {Current protein language models (PLMs) learn protein representations mainly based on their sequences, thereby well capturing co-evolutionary information, but they are unable to explicitly acquire protein functions, which is the end goal of protein representation learning. Fortunately, for many proteins, their textual property descriptions are available, where their various functions are also described. Motivated by this fact, we first build the ProtDescribe dataset to augment protein sequences with text descriptions of their functions and other important properties. Based on this dataset, we propose the ProtST framework to enhance Protein Sequence pre-training and understanding by biomedical Texts. During pre-training, we design three types of tasks, i.e., unimodal mask prediction, multimodal representation alignment and multimodal mask prediction, to enhance a PLM with protein property information with different granularities and, at the same time, preserve the PLM’s original representation power. On downstream tasks, ProtST enables both supervised learning and zero-shot prediction. We verify the superiority of ProtST-induced PLMs over previous ones on diverse representation learning benchmarks. Under the zero-shot setting, we show the effectiveness of ProtST on zero-shot protein classification, and ProtST also enables functional protein retrieval from a large-scale database without any function annotation.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/