BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations
Qizhi PeiWei ZhangJinhua ZhuKehan WuKaiyuan GaoLijun WuYingce XiaRui Yan
Develops BioT5, a unified pre-training framework that integrates SELFIES molecular strings, protein sequences, and biomedical text to guarantee valid chemical generation and improve performance across diverse drug discovery and bio-entity prediction tasks.
Modern drug discovery increasingly depends on computational models that can understand and integrate data across small molecules, proteins, and scientific literature. However, existing language models in this domain encounter major obstacles: they frequently generate chemically invalid molecular structures due to fragile sequence representations like SMILES, struggle to separate biological and linguistic meanings when token vocabularies are shared, fail to fully exploit contextual information from biomedical text, and treat structured database records the same as unstructured literature text.
The article introduces and evaluates BioT5, a comprehensive multi-modal pre-training framework that integrates chemical knowledge, protein sequences, and natural language. BioT5 aims to overcome previous structural representation flaws, distinguish between structured and unstructured biological data, and improve performance across diverse biological prediction and generation tasks.
To achieve this, the approach employs a 252-million-parameter encoder-decoder Transformer architecture that incorporates distinct token vocabularies for text, protein sequences (FASTA format with dedicated prefix identifiers), and molecules using SELFIES—a representation format that guarantees 100% chemical validity for all possible character strings. The model is pre-trained across six multi-task objectives using large-scale datasets, including general text (C4), 27 million sampled proteins (UniRef50), small molecules (ZINC20), 33 million PubMed scientific articles containing entity-linked molecular and protein sequences ("wrapped" text), and structured entity-description pairs from PubChem and Swiss-Prot. BioT5 was subsequently fine-tuned and tested across 15 standard downstream benchmark tasks covering molecule property prediction, protein property prediction, drug-target interactions, protein-protein interactions, molecule captioning, and text-based molecule generation.
The key findings demonstrate that BioT5 achieves state-of-the-art performance on 10 downstream tasks and competitive results on the remaining 5. In text-based molecule generation, BioT5 achieved 100% validity and surpassed the leading baseline (MolT5-Large) with a 32.8% relative improvement in exact match score (reaching 0.413 compared to 0.311) despite using roughly one-third of the parameters. In molecule captioning, it outperformed all competing models across all evaluation metrics, reaching a Text2Mol cross-modal similarity score of 0.603, nearly matching the ground-truth reference score of 0.609. Across biological classification benchmarks, BioT5 consistently matched or outperformed specialized graph neural networks and much larger protein language models—such as ProtBert (420M parameters) and ESM-1b (652M parameters)—on protein solubility, protein localization, drug-target interaction, and protein-protein interaction benchmarks.
These results show that separating modal token vocabularies while unifying cross-modal learning allows a moderately sized language model to outperform much larger domain-specific models. Eliminating invalid molecular outputs reduces downstream screening risks and potential pipeline failures in computational drug discovery. The integration of literature context and structured databases allows the model to capture biochemical mechanisms that sequence data alone cannot provide.
Based on these findings, development teams should adopt SELFIES representations and separated vocabularies when building generative biomedical language models. Future work should focus on expanding the architecture to incorporate additional biological data modalities (such as genomics and transcriptomics) as well as 2D and 3D structural representations, while exploring efficient adaptation methods like instruction-tuning that avoid the need for full-parameter fine-tuning on each task.
Confidence in these findings is supported by consistent multi-run benchmarking across 15 standard evaluation datasets. However, current limitations include the requirement for full-parameter model fine-tuning for each task due to cross-task data leakage risks and limited generalization under zero-shot prompting, as well as the restriction to linear sequence representations rather than higher-order structural formats.
- Paper: Translation between Molecules and Natural Language, Carl Edwards et al. (2022). MolT5 establishes the joint molecular-structure and natural-language pretraining approach that BioT5 adapts to biological entities and richer contextual knowledge.
- Paper: Unifying Molecular and Textual Representations via Multi-task Language Modelling, Dimitrios Christofidellis et al. (2023). Text+Chem T5 provides a unified molecular-and-text language-modeling framework that helps clarify the cross-domain modeling approach BioT5 advances.
- Paper: BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining, Renqian Luo et al. (2022). BioGPT introduces generative pretraining on biomedical literature, a useful foundation for understanding BioT5’s use of natural-language associations from biological text.
No sufficiently relevant recommendations were found.
