ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts
Minghao XuXinyu YuanSantiago MiretJian Tang
Proposes a multimodal pre-training framework that aligns protein sequences with biomedical text descriptions to improve protein language models and enable zero-shot protein classification and functional retrieval.
Understanding protein functions and properties is central to advancing drug discovery and healthcare. While recent machine learning approaches use protein language models to learn directly from amino acid sequences, these conventional models rely primarily on evolutionary patterns and struggle to explicitly capture real-world functional behaviors and cellular locations. Meanwhile, vast amounts of rich biomedical text describing protein functions exist but remain largely unintegrated during model training.
The article demonstrates that enriching protein sequence pre-training with biomedical text descriptions significantly enhances functional understanding and model accuracy. Specifically, it introduces a multimodal training framework named ProtST, which jointly aligns protein sequences with their corresponding textual property descriptions to improve supervised predictions and enable zero-shot applications.
To evaluate this concept, the authors constructed a dataset of over 553,000 paired protein sequences and curated property texts covering names, functions, cellular locations, and protein families. The framework combines existing protein models with a biomedical text language model across three objectives: predicting masked sequence elements to retain baseline evolutionary patterns, aligning overall sequence and text representations via contrastive learning, and predicting masked residues and words through a cross-modal fusion layer. The models were tested across eleven benchmark tasks spanning localization, fitness landscape prediction, and function annotation, as well as zero-shot classification and text-based database retrieval.
The findings show substantial performance improvements across downstream applications. Enhancing baseline models with the multimodal framework improved performance on nearly all evaluated metrics, with the enhanced models setting new state-of-the-art benchmarks on protein fitness and function annotation. In zero-shot classification settings with no labeled training examples, the model matched or exceeded the accuracy of conventional models trained with several labeled examples per class. Furthermore, using zero-shot predictions to refine standard supervised classifiers consistently improved full-dataset accuracy, while text-to-protein retrieval successfully identified functional ligand binders from large databases using only natural language prompts.
These results indicate that integrating textual biomedical knowledge into sequence models reduces dependency on costly, labor-intensive experimental labeling while improving predictive accuracy. Organizations involved in protein engineering and therapeutic development can leverage this approach to accelerate lead discovery, rapidly query unannotated sequence repositories, and improve model performance even when experimental data is scarce.
Moving forward, adopting multimodal sequence-text models is recommended when developing protein representation pipelines or screening databases for specific functional properties. To build on these results, future efforts should expand the text dataset using broader biomedical literature, incorporate 3D protein structural coordinates into the multimodal framework, and explore text-guided generative protein design.
Confidence in these findings is supported by consistent performance gains across diverse standard benchmarks and multiple baseline architectures. However, decision-makers should note that the current training data relies on curated Swiss-Prot annotations, which limits coverage across the entire protein universe. Additionally, data quality proved critical, as pre-training on larger but lower-quality automated text annotations led to measurable performance declines.
- Paper: ProtTrans: Towards Cracking the Language of Life’s Code Through Self-Supervised Deep Learning and High Performance Computing, Ahmed Elnaggar et al. (2020). ProtTrans establishes the foundational large-scale self-supervised protein sequence language modeling paradigm that ProtST directly adapts and enhances with biomedical text descriptions.
- Paper: BioBERT: a pre-trained biomedical language representation model for biomedical text mining, Jinhyuk Lee et al. (2019). BioBERT introduces domain-specific pre-training on PubMed literature, providing the core biomedical text representations used by ProtST to align sequence embeddings with textual properties.
- Paper: BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining, Renqian Luo et al. (2022). BioGPT demonstrates generative biomedical text pre-training, serving as key background for understanding how language models encode specialized functional terminology for proteins.
- Paper: Deep Bidirectional Language-Knowledge Graph Pretraining, Michihiro Yasunaga et al. (2022). DRAGON models bidirectional cross-modal pre-training between text and biological knowledge, establishing pre-training methodologies that ProtST builds upon for protein-text alignment.
- Paper: Text Embeddings by Weakly-Supervised Contrastive Pre-training, Liang Wang et al. (2022). E5 outlines the contrastive pre-training principles for embedding alignment that form one of ProtST's three core training objectives.
- Paper: A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery, Yu Zhang et al. (2024). This comprehensive survey examines scientific large language models and synthesizes multi-modal alignment paradigms like ProtST across broader scientific domains.
- Paper: ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention, Mingchen Li et al. (2024). ProSST extends protein representation modeling by incorporating quantized three-dimensional structure alongside sequence data, fulfilling ProtST's recommended direction to integrate structural coordinates.
- Paper: ProteinGym: Large-Scale Benchmarks for Protein Fitness Prediction and Design, Pascal Notin et al. (2023). ProteinGym establishes an extensive benchmark suite for evaluating downstream zero-shot and supervised protein fitness predictions, providing a standardized testbed for multimodal protein representations.
