BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations

Qizhi PeiWei ZhangJinhua ZhuKehan WuKaiyuan GaoLijun WuYingce XiaRui Yan

article2023EMNLP123 citations

Develops BioT5, a unified pre-training framework that integrates SELFIES molecular strings, protein sequences, and biomedical text to guarantee valid chemical generation and improve performance across diverse drug discovery and bio-entity prediction tasks.

Listen

Modern drug discovery increasingly depends on computational models that can understand and integrate data across small molecules, proteins, and scientific literature. However, existing language models in this domain encounter major obstacles: they frequently generate chemically invalid molecular structures due to fragile sequence representations like SMILES, struggle to separate biological and linguistic meanings when token vocabularies are shared, fail to fully exploit contextual information from biomedical text, and treat structured database records the same as unstructured literature text.

The article introduces and evaluates BioT5, a comprehensive multi-modal pre-training framework that integrates chemical knowledge, protein sequences, and natural language. BioT5 aims to overcome previous structural representation flaws, distinguish between structured and unstructured biological data, and improve performance across diverse biological prediction and generation tasks.

To achieve this, the approach employs a 252-million-parameter encoder-decoder Transformer architecture that incorporates distinct token vocabularies for text, protein sequences (FASTA format with dedicated prefix identifiers), and molecules using SELFIES—a representation format that guarantees 100% chemical validity for all possible character strings. The model is pre-trained across six multi-task objectives using large-scale datasets, including general text (C4), 27 million sampled proteins (UniRef50), small molecules (ZINC20), 33 million PubMed scientific articles containing entity-linked molecular and protein sequences ("wrapped" text), and structured entity-description pairs from PubChem and Swiss-Prot. BioT5 was subsequently fine-tuned and tested across 15 standard downstream benchmark tasks covering molecule property prediction, protein property prediction, drug-target interactions, protein-protein interactions, molecule captioning, and text-based molecule generation.

The key findings demonstrate that BioT5 achieves state-of-the-art performance on 10 downstream tasks and competitive results on the remaining 5. In text-based molecule generation, BioT5 achieved 100% validity and surpassed the leading baseline (MolT5-Large) with a 32.8% relative improvement in exact match score (reaching 0.413 compared to 0.311) despite using roughly one-third of the parameters. In molecule captioning, it outperformed all competing models across all evaluation metrics, reaching a Text2Mol cross-modal similarity score of 0.603, nearly matching the ground-truth reference score of 0.609. Across biological classification benchmarks, BioT5 consistently matched or outperformed specialized graph neural networks and much larger protein language models—such as ProtBert (420M parameters) and ESM-1b (652M parameters)—on protein solubility, protein localization, drug-target interaction, and protein-protein interaction benchmarks.

These results show that separating modal token vocabularies while unifying cross-modal learning allows a moderately sized language model to outperform much larger domain-specific models. Eliminating invalid molecular outputs reduces downstream screening risks and potential pipeline failures in computational drug discovery. The integration of literature context and structured databases allows the model to capture biochemical mechanisms that sequence data alone cannot provide.

Based on these findings, development teams should adopt SELFIES representations and separated vocabularies when building generative biomedical language models. Future work should focus on expanding the architecture to incorporate additional biological data modalities (such as genomics and transcriptomics) as well as 2D and 3D structural representations, while exploring efficient adaptation methods like instruction-tuning that avoid the need for full-parameter fine-tuning on each task.

Confidence in these findings is supported by consistent multi-run benchmarking across 15 standard evaluation datasets. However, current limitations include the requirement for full-parameter model fine-tuning for each task due to cross-task data leakage risks and limited generalization under zero-shot prompting, as well as the restriction to linear sequence representations rather than higher-order structural formats.

No sufficiently relevant recommendations were found.

Cover for BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations

Abstract

Recent advancements in biological research leverage the integration of molecules, proteins, and natural language to enhance drug discovery. However, current models exhibit several limitations, such as the generation of invalid molecular SMILES, underutilization of contextual information, and equal treatment of structured and unstructured knowledge. To address these issues, we propose BioT5, a comprehensive pre-training framework that enriches cross-modal integration in biology with chemical knowledge and natural language associations. BioT5 utilizes SELFIES for 100% robust molecular representations and extracts knowledge from the surrounding context of bio-entities in unstructured biological literature. Furthermore, BioT5 distinguishes between structured and unstructured knowledge, leading to more effective utilization of information. After fine-tuning, BioT5 shows superior performance across a wide range of tasks, demonstrating its strong capability of capturing underlying relations and properties of bio-entities. Our code is available at https://github.com/QizhiPei/BioT5.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Cross-modal Models in Biology
  • 2.2 Representations of Molecule and Protein
  • 3 BioT5
  • 3.1 Pre-training Corpus
  • 3.2 Separate Tokenization and Embedding
  • 3.3 Model and Training
  • 4 Experiments and Results
  • 4.1 Single-instance Prediction
  • 4.1.1 Molecule Property Prediction
  • 4.1.2 Protein Property Prediction
  • 4.2 Multi-instance Prediction
  • 4.2.1 Drug-target Interaction Prediction
  • 4.2.2 Protein-protein Interaction Prediction
  • 4.3 Cross-modal Generation
  • 4.3.1 Molecule Captioning
  • 4.3.2 Text-Based Molecule Generation
  • 5 Conclusions and Future Work
  • 6 Limitations
  • 7 Acknowledgements
  • References
  • A Reproducibility
  • B NER and Entity Linking Process
  • C Dictionary and SELFIES Conversion
  • D Molecule-Text Generation Metrics
  • D.1 Molecule Captioning Metrics
  • D.2 Text-based Molecule Generation Metrics
  • E Pre-training Details
  • E.1 Special Tokens
  • E.2 Hyper-parameters
  • F Fine-tuning Details
  • F.1 Single-instance Prediction
  • F.1.1 Molecule Property Prediction
  • F.1.2 Protein Property Prediction
  • F.2 Multi-instance Prediction
  • F.2.1 Drug-target Interaction Prediction
  • F.2.2 Protein-protein Interaction Prediction
  • F.3 Cross-modal Generation
  • F.3.1 Molecule Captioning
  • F.3.2 Text-based molecule generation
  • G Case Study

Knowls

  1. Knowl 1 — BioT5 jointly pre-trains across six biological and language tasks

    model/method

    BioT5 is a text-to-text Transformer with a shared encoder and decoder, pre-trained to connect natural language, molecule SELFIES, and protein FASTA sequences through six tasks. For SELFIES, FASTA, and general text separately, it uses the T5 span-corruption objective: replace masked spans with sentinel tokens and train the model to reconstruct the missing content. It applies the same objective to wrapped biological sentences, where text and embedded SELFIES or FASTA tokens can all be masked and recovered. For structured molecule–description pairs and protein–description pairs, it trains on translation in both directions: molecule or protein sequence to text description, and text description to sequence. The combination is intended to learn both within-modality representations and associations between bio-entities and their textual descriptions.

  2. Knowl 2 — Pre-training combines molecular, protein, literature, and database data

    model/method

    BioT5’s single-modality corpora comprise ZINC20 molecules converted from SMILES to SELFIES, 27 million proteins sampled from UniRef50 after filtering long sequences, and general text from C4. To create wrapped biological text, the authors process 33 million PubMed articles with BERN2 named-entity recognition and link molecule mentions to ChEBI or MeSH and protein mentions to NCBI Gene. They replace detected molecule names with the corresponding SELFIES; for a detected gene name, they append the linked protein FASTA sequence. When a sentence contains multiple protein entities, one is chosen at random for FASTA appendage, and wrapping is applied only to sentences with detected entities. Structured translation data consist of 339,000 PubChem molecule–description pairs, excluding molecules in the downstream ChEBI-20 dataset to reduce leakage, and 569,000 Swiss-Prot protein–description pairs. Database descriptions are formatted with entity-specific fields, including molecule name and description, and protein name, function, subcellular location, and protein families; absent fields are omitted.

  3. Knowl 3 — Separate biological tokenization preserves modality-specific meanings

    model/method

    BioT5 uses separate vocabularies for natural-language text, molecule SELFIES, and protein sequences rather than treating them as one shared token language. Each bracketed chemically meaningful SELFIES unit is tokenized as one token; for example, [C][=C][Br] becomes [C], [=C], and [Br]. Each protein amino acid is prefixed with the special token <p>, so the sequence <p>M<p>K<p>R is tokenized as three prefixed amino-acid tokens. Natural-language text uses the original T5 vocabulary. This distinction avoids splitting meaningful molecular units—for example, a bromine token—into pieces that can be confused with ordinary text or other chemical symbols, and prevents a character such as C from sharing one token meaning across text, molecules, and proteins.

  4. Knowl 4 — BioT5 uses a T5-base architecture and a unified fine-tuning format

    experimental setup

    The 252-million-parameter BioT5 model follows the T5-v1.1-base architecture and has a vocabulary of 35,073 tokens. Pre-training runs for 350,000 steps on eight 80-GB NVIDIA A100 GPUs, with a batch size of 96 per GPU and six data types in each batch. For each structured molecule–text or protein–text example, the translation direction is selected with probability 0.5. Optimization uses AdamW with RMS scaling, a cosine learning-rate schedule from 1e-2 to 1e-5, 10,000 warm-up steps, zero dropout, and a maximum pre-training input length of 512 tokens. Downstream tasks are fine-tuned as sequence generation using prompts, with binary classification outputs expressed as Yes or No. For binary classification metrics requiring probabilities, the Yes and No token probabilities are normalized over those two labels.

  5. Knowl 5 — Molecule-property prediction improves on most MoleculeNet comparisons

    empirical result

    On six MoleculeNet binary classification datasets, using challenging Bemis–Murcko scaffold splits and reporting mean ± standard deviation over three runs, BioT5 obtains AUROC scores of 77.7 ± 0.6 on BBBP, 77.9 ± 0.2 on Tox21, 95.4 ± 0.5 on ClinTox, 81.0 ± 0.1 on HIV, 89.4 ± 0.3 on BACE, and 73.2 ± 0.2 on SIDER; its reported average is 82.4. The datasets cover one BBBP task, 12 Tox21 tasks, two ClinTox tasks, one HIV task, one BACE task, and 27 SIDER tasks. BioT5 exceeds the strongest listed prior result on five datasets; on Tox21, its 77.9 ± 0.2 is below GEM’s 78.1 ± 0.1. Its average is higher than MolXPT’s reported average of 81.9.

  6. Knowl 6 — Protein property and interaction results show mixed comparison outcomes

    empirical result

    On PEER protein property prediction, BioT5 achieves 74.65 ± 0.49 accuracy for solubility and 91.69 ± 0.05 for localization, averaged over three runs. It is best on solubility among the listed methods, while ESM-1b scores 70.23 ± 0.75; on localization, ESM-1b scores 92.40 ± 0.35, higher than BioT5. Localization is the binary distinction between membrane-bound and soluble proteins. On protein–protein interaction prediction, BioT5 obtains 64.89 ± 0.43 accuracy on Yeast and 86.22 ± 0.53 on Human, also over three runs. The best listed comparator on each is ESM-1b with only its prediction head trained (66.07 ± 0.58 and 88.06 ± 0.24, respectively); BioT5 exceeds the ProtBert and ESM-1b results reported with their full parameters fine-tuned. Thus, the protein results do not establish that BioT5 is best in every setting, but show competitive results with a substantially smaller model than the 419.9-million-parameter ProtBert and 652.4-million-parameter ESM-1b.

  7. Knowl 7 — BioT5 improves drug–target interaction prediction across three datasets

    empirical result

    For binary drug–target interaction prediction, BioT5 is evaluated on BioSNAP, Human, and BindingDB; the respective train/validation/test sizes are 19,224/2,747/5,493, 4,197/600/1,200, and 50,149/5,604/5,505. Results are means ± standard deviations over five runs. In the metric order AUROC, AUPRC, and accuracy, BioT5 scores 0.937 ± 0.001, 0.937 ± 0.004, and 0.874 ± 0.001 on BioSNAP, compared with DrugBAN’s 0.903 ± 0.005, 0.902 ± 0.004, and 0.834 ± 0.008. On Human, BioT5’s AUROC and AUPRC are 0.989 ± 0.001 and 0.985 ± 0.002, versus DrugBAN’s 0.982 ± 0.002 and 0.980 ± 0.003. On BindingDB, BioT5 scores 0.963 ± 0.001, 0.952 ± 0.001, and 0.907 ± 0.003, versus DrugBAN’s 0.960 ± 0.001, 0.948 ± 0.002, and 0.904 ± 0.004. BioT5 therefore exceeds DrugBAN, the strongest listed baseline, on each reported metric for these datasets.

  8. Knowl 8 — BioT5 leads the reported molecule-captioning comparisons

    empirical result

    On the ChEBI-20 molecule-captioning task, the input is a molecule SELFIES sequence and the output is an English description. The dataset contains 33,010 molecule–text pairs, split into 26,407 training, 3,301 validation, and 3,300 test examples. BioT5 has 252 million parameters and scores BLEU-2 0.635, BLEU-4 0.556, ROUGE-1 0.692, ROUGE-2 0.559, ROUGE-L 0.633, METEOR 0.656, and Text2Mol 0.603. These scores exceed the listed baselines across all reported metrics; for comparison, the 350-million-parameter MolXPT scores 0.594, 0.505, 0.660, 0.511, 0.597, 0.626, and 0.594 in the same metric order. The Text2Mol score between the ground-truth molecule and its description is 0.609, close to BioT5’s 0.603.

  9. Knowl 9 — SELFIES generation achieves perfect measured validity on ChEBI-20

    empirical result

    On the reverse ChEBI-20 task, BioT5 receives an English molecule description and generates a SELFIES sequence. It scores BLEU 0.867, exact match 0.413, Levenshtein distance 15.097, MACCS fingerprint similarity 0.886, RDK fingerprint similarity 0.801, Morgan fingerprint similarity 0.734, FCD 0.43, Text2Mol 0.576, and validity 1.000. BioT5 has 252 million parameters, compared with 783 million for MolT5-large; their exact-match scores are 0.413 and 0.311, respectively, and their validity scores are 1.000 and 0.905. BioT5 is ahead on nearly all reported measures, but not every one: MolXPT’s Text2Mol score is 0.578, slightly above BioT5’s 0.576. BLEU, exact match, Levenshtein distance, and validity are computed over all generated molecules; fingerprint similarity, FCD, and Text2Mol are computed only for syntactically valid molecules. The SELFIES sequences are converted to SMILES for metric calculation.

  10. Knowl 10 — Evaluation is limited to sequence modalities and task-specific fine-tuning

    limitation

    BioT5 is fine-tuned on each downstream task rather than using one instruction-tuned model across tasks. The authors report that they did not observe generalization across downstream tasks with instruction tuning and that combining task data could cause leakage; as an example, they note overlap between BindingDB training data and the BioSNAP and Human test sets. The demonstrated modalities are text, molecule sequences, and protein sequences. DNA/RNA, cells, and other biological data types are not evaluated, nor are molecule or protein representations based on 2D or 3D structures.

Coverage note — The paper’s future-work proposals are omitted because they are plans rather than demonstrated contributions.

References

  1. 1.
    1. Uniprot: the universal protein knowledgebase in 2023. Nucleic Acids Research, 51(D1):D523–D531.
  2. 2.José Juan Almagro Armenteros, Casper Kaae Sønderby, Søren Kaae Sønderby, Henrik Nielsen, and Ole Winther. 2017. Deeploc: prediction of protein subcellular localization using deep learning. Bioinform., 33(21):3387–3395.
  3. 3.AstraZeneca. 2023. A big future for small molecules: targeting the undruggable.
  4. 4.Viraj Bagal, Rishal Aggarwal, P. K. Vinod, and U. Deva Priyakumar. 2022. Molgpt: Molecular generation using a transformer-decoder model. J. Chem. Inf. Model., 62(9):2064–2076.
  5. 5.Peizhen Bai, Filip Miljkovic, Yan Ge, Nigel Greene, Bino John, and Haiping Lu. 2021. Hierarchical clustering split for low-bias evaluation of drug-target interaction prediction. In 2021 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 641–644. IEEE.
  6. 6.Peizhen Bai, Filip Miljkovic, Bino John, and Haiping Lu. 2023. Interpretable bilinear attention network with domain adaptation improves drug–target prediction. Nature Machine Intelligence, 5(2):126–136.
  7. 7.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization@ACL 2005, Ann Arbor, Michigan, USA, June 29, 2005, pages 65–72. Association for Computational Linguistics.
  8. 8.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. Scibert: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 3613–3618. Association for Computational Linguistics.
  9. 9.Emmanuel Boutet, Damien Lieberherr, Michael Tognolli, Michel Schneider, and Amos Bairoch. 2007. Uniprotkb/swiss-prot: the manually annotated section of the uniprot knowledgebase. Plant bioinformatics: methods and protocols, pages 89–112.
  10. 10.J Rodney Brister, Danso Ako-Adjei, Yiming Bao, and Olga Blinkova. 2015. Ncbi viral genomes resource. Nucleic acids research, 43(D1):D571–D577.
  11. 11.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  12. 12.Darko Butina. 1999. Unsupervised data base clustering based on daylight’s fingerprint and tanimoto similarity: A fast and automated way to cluster small and large data sets. J. Chem. Inf. Comput. Sci., 39(4):747–750.
  13. 13.Kathi Canese and Sarah Weis. 2013. Pubmed: the bibliographic database. The NCBI handbook, 2(1).
  14. 14.Dong-Sheng Cao, Qing-Song Xu, and Yi-Zeng Liang. 2013. propy: a tool to generate various modes of chou’s pseaac. Bioinform., 29(7):960–962.
  15. 15.Lifan Chen, Xiaoqin Tan, Dingyan Wang, Feisheng Zhong, Xiaohong Liu, Tianbiao Yang, Xiaomin Luo, Kaixian Chen, Hualiang Jiang, and Mingyue Zheng. 2020. Transformercpi: improving compound–protein interaction prediction by sequence-based deep learning with self-attention mechanism and label reversal experiments. Bioinformatics, 36(16):4406–4414.
  16. 16.Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. 2020. Chemberta: Large-scale self-supervised pretraining for molecular property prediction. arXiv preprint arXiv:2010.09885.
  17. 17.Corinna Cortes and Vladimir Vapnik. 1995. Support-vector networks. Machine learning, 20(3):273–297.
  18. 18.Suresh Dara, Swetha Dhamercherla, Surender Singh Jadav, CH Madhu Babu, and Mohamed Jawed Ahsan. 2022. Machine learning in drug discovery: a review. Artificial Intelligence Review, 55(3):1947–1999.
  19. 19.Joseph L Durant, Burton A Leland, Douglas R Henry, and James G Nourse. 2002. Reoptimization of mdl keys for use in drug discovery. Journal of chemical information and computer sciences, 42(6):1273–1280.
  20. 20.Carl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. 2022. Translation between molecules and natural language. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 375–413. Association for Computational Linguistics.
  21. 21.Carl Edwards, ChengXiang Zhai, and Heng Ji. 2021. Text2mol: Cross-modal molecule retrieval with natural language queries. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 595–607. Association for Computational Linguistics.
  22. 22.Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, et al. 2021. Prottrans: Toward understanding the language of life through self-supervised learning. IEEE transactions on pattern analysis and machine intelligence, 44(10):7112–7127.
  23. 23.Xiaomin Fang, Lihang Liu, Jieqiong Lei, Donglong He, Shanzhuo Zhang, Jingbo Zhou, Fan Wang, Hua Wu, and Haifeng Wang. 2022. Geometry-enhanced molecular representation learning for property prediction. Nature Machine Intelligence, 4(2):127–134.
  24. 24.Zhi-Ping Feng and Chun-Ting Zhang. 2000. Prediction of membrane protein types based on the hydrophobic index of amino acids. Journal of protein chemistry, 19:269–275.
  25. 25.Noelia Ferruz, Steffen Schmidt, and Birte Höcker. 2022. Protgpt2 is a deep unsupervised language model for protein design. Nature communications, 13(1):4348.
  26. 26.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 3816–3830. Association for Computational Linguistics.
  27. 27.Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), 3(1):1–23.
  28. 28.Yanzhi Guo, Lezheng Yu, Zhining Wen, and Menglong Li. 2008. Using support vector machine combined with auto covariance to predict protein–protein interactions from protein sequences. Nucleic acids research, 36(9):3025–3030.
  29. 29.Janna Hastings, Gareth Owen, Adriano Dekker, Marcus Ennis, Namrata Kale, Venkatesh Muthukrishnan, Steve Turner, Neil Swainston, Pedro Mendes, and Christoph Steinbeck. 2016. Chebi in 2016: Improved services and an expanding collection of metabolites. Nucleic acids research, 44(D1):D1214–D1219.
  30. 30.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778.
  31. 31.Stephen Heller, Alan McNaught, Stephen Stein, Dmitrii Tchekhovskoi, and Igor Pletnev. 2013. Inchi-the worldwide chemical structure identifier standard. Journal of cheminformatics, 5(1):1–9.
  32. 32.Tin Kam Ho. 1995. Random decision forests. In Proceedings of 3rd International Conference on Document Analysis and Recognition, volume 1, pages 278–282.
  33. 33.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation, 9(8):1735–1780.
  34. 34.Kexin Huang, Cao Xiao, Lucas Glass, and Jimeng Sun. 2021. MolTrans: Molecular interaction transformer for drug–target interaction prediction. Bioinformatics, 37:830 – 836.
  35. 35.John J Irwin, Khanh G Tang, Jennifer Young, Chinzorig Dandarchuluun, Benjamin R Wong, Munkhzul Khurelbaatar, Yurii S Moroz, John Mayfield, and Roger A Sayle. 2020. Zinc20—a free ultralarge-scale chemical database for ligand discovery. Journal of chemical information and modeling, 60(12):6065–6073.
  36. 36.Sameer Khurana, Reda Rawi, Khalid Kunji, Gwo-Yu Chuang, Halima Bensmail, and Raghvendra Mall. 2018. Deepsol: a deep learning framework for sequence-based protein solubility prediction. Bioinform., 34(15):2605–2613.
  37. 37.Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al. 2023. Pubchem 2023 update. Nucleic Acids Research, 51(D1):D1373–D1380.
  38. 38.Sunghwan Kim, Paul A Thiessen, Tiejun Cheng, Jian Zhang, Asta Gindulyte, and Evan E Bolton. 2019. Pug-view: programmatic access to chemical annotations integrated in pubchem. Journal of cheminformatics, 11(1):1–11.
  39. 39.Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  40. 40.Mario Krenn, Qianxiang Ai, Senja Barthel, Nessa Carson, Angelo Frei, Nathan C Frey, Pascal Friederich, Théophile Gaudin, Alberto Alexander Gayle, Kevin Maik Jablonka, et al. 2022. Selfies and the future of molecular string representations. Patterns, 3(10):100588.
  41. 41.Mario Krenn, Florian Häse, AkshatKumar Nigam, Pascal Friederich, and Alan Aspuru-Guzik. 2020. Self-referencing embedded strings (selfies): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1(4):045024.
  42. 42.Taku Kudo and John Richardson. 2018. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018: System Demonstrations, Brussels, Belgium, October 31 - November 4, 2018, pages 66–71. Association for Computational Linguistics.
  43. 43.Greg Landrum. 2021. Rdkit: Open-source cheminformatics software. GitHub release.
  44. 44.Ingoo Lee, Jongsoo Keum, and Hojung Nam. 2019. DeepConv-DTI: Prediction of drug-target interactions via deep learning with convolution on protein sequences. PLoS Computational Biology, 15.
  45. 45.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240.
  46. 46.Jiatong Li, Yunqing Liu, Wenqi Fan, Xiao-Yong Wei, Hui Liu, Jiliang Tang, and Qing Li. 2023. Empowering molecule discovery for molecule-caption translation with large language models: A chatgpt perspective. arXiv preprint arXiv:2306.06615.
  47. 47.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  48. 48.Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Sal Candido, et al. 2022. Language models of protein sequences at the scale of evolution enable accurate structure prediction. BioRxiv.
  49. 49.David J Lipman and William R Pearson. 1985. Rapid and sensitive protein similarity searches. Science, 227(4693):1435–1441.
  50. 50.Carolyn E Lipscomb. 2000. Medical subject headings (mesh). Bulletin of the Medical Library Association, 88(3):265.
  51. 51.Hui Liu, Jianjiang Sun, Jihong Guan, Jie Zheng, and Shuigeng Zhou. 2015. Improving compound–protein interaction prediction by building up highly credible negative samples. Bioinformatics, 31(12):i221–i229.
  52. 52.Shengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby, Hongyu Guo, and Jian Tang. 2022. Pre-training molecular graph representation with 3d geometry. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  53. 53.Shengchao Liu, Yutao Zhu, Jiarui Lu, Zhao Xu, Weili Nie, Anthony Gitter, Chaowei Xiao, Jian Tang, Hongyu Guo, and Anima Anandkumar. 2023a. A text-guided protein design framework. arXiv preprint arXiv:2302.04611.
  54. 54.Tiqing Liu, Yuhmei Lin, Xin Wen, Robert N Jorissen, and Michael K Gilson. 2007. Bindingdb: a web-accessible database of experimentally determined protein–ligand binding affinities. Nucleic acids research, 35(suppl_1):D198–D201.
  55. 55.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  56. 56.Zequn Liu, Wei Zhang, Yingce Xia, Lijun Wu, Shufang Xie, Tao Qin, Ming Zhang, and Tie-Yan Liu. 2023b. Molxpt: Wrapping molecules with text for generative pre-training. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 1606–1616. Association for Computational Linguistics.
  57. 57.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  58. 58.Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 23(6).
  59. 59.Yizhen Luo, Kui Huang, Massimo Hong, Kai Yang, Jiahuan Zhang, Yushuai Wu, and Zaiqin Nie. 2023. Empowering ai drug discovery with explicit and implicit knowledge. arXiv preprint arXiv:2305.01523.
  60. 60.Larry R Medsker and LC Jain. 2001. Recurrent neural networks. Design and Applications, 5:64–67.
  61. 61.Frederic P Miller, Agnes F Vandome, and John McBrewster. 2009. Levenshtein distance: Information theory, computer science, string (computer science), string metric, damerau? levenshtein distance, spell checker, hamming distance.
  62. 62.Piotr Nawrot. 2023. nanoT5.
  63. 63.Thin Nguyen, Hang Le, T. Quinn, Tri Minh Nguyen, Thuc Duy Le, and Svetha Venkatesh. 2021. GraphDTA: Predicting drug-target binding affinity with graph neural networks. Bioinformatics, 37(8):1140–1147.
  64. 64.Noel O’Boyle and Andrew Dalke. 2018. Deepsmiles: an adaptation of smiles for use in machine-learning of chemical structures.
  65. 65.Keiron O’Shea and Ryan Nash. 2015. An introduction to convolutional neural networks. arXiv preprint arXiv:1511.08458.
  66. 66.Xiao-Yong Pan, Ya-Nan Zhang, and Hong-Bin Shen. 2010. Large-scale prediction of human protein- protein interactions from amino acid sequence based on latent topic features. Journal of proteome research, 9(10):4992–5001.
  67. 67.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA, pages 311–318. ACL.
  68. 68.William R Pearson and David J Lipman. 1988. Improved tools for biological sequence comparison. Proceedings of the National Academy of Sciences, 85(8):2444–2448.
  69. 69.Suraj Peri, J Daniel Navarro, Ramars Amanchy, Troels Z Kristiansen, Chandra Kiran Jonnalagadda, Vineeth Surendranath, Vidya Niranjan, Babylakshmi Muthusamy, TKB Gandhi, Mads Gronborg, et al. 2003. Development of human protein reference database as an initial platform for approaching systems biology in humans. Genome research, 13(10):2363–2371.
  70. 70.Kristina Preuer, Philipp Renz, Thomas Unterthiner, Sepp Hochreiter, and Günter Klambauer. 2018. Fréchet chemnet distance: A metric for generative models for molecules in drug discovery. J. Chem. Inf. Model., 58(9):1736–1741.
  71. 71.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018. Improving language understanding by generative pre-training.
  72. 72.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  73. 73.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  74. 74.Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. 2021. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences, 118(15):e2016239118.
  75. 75.Stephen E. Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: BM25 and beyond. Found. Trends Inf. Retr., 3(4):333–389.
  76. 76.David Rogers and Mathew Hahn. 2010a. Extended-connectivity fingerprints. Journal of chemical information and modeling, 50(5):742–754.
  77. 77.David Rogers and Mathew Hahn. 2010b. Extended-connectivity fingerprints. J. Chem. Inf. Model., 50(5):742–754.
  78. 78.Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying Wei, Wenbing Huang, and Junzhou Huang. 2020. Self-supervised graph transformer on large-scale molecular data. Advances in Neural Information Processing Systems, 33:12559–12571.
  79. 79.Vijayakumar Saravanan and Namasivayam Gautham. 2015. Harnessing computational biology for exact linear b-cell epitope prediction: a novel amino acid composition-based feature descriptor. Omics: a journal of integrative biology, 19(10):648–658.
  80. 80.Nadine Schneider, Roger A. Sayle, and Gregory A. Landrum. 2015. Get your atoms in order - an open-source implementation of a novel and robust molecular canonicalization algorithm. J. Chem. Inf. Model., 55(10):2111–2120.
  81. 81.Martin Steinegger and Johannes Söding. 2018. Clustering huge protein sequence sets in linear time. Nature communications, 9(1):2542.
  82. 82.Teague Sterling and John J. Irwin. 2015. ZINC 15 - ligand discovery for everyone. J. Chem. Inf. Model., 55(11):2324–2337.
  83. 83.Bing Su, Dazhao Du, Zhao Yang, Yujie Zhou, Jiangmeng Li, Anyi Rao, Hao Sun, Zhiwu Lu, and Ji-Rong Wen. 2022. A molecular multimodal foundation model associating molecule graphs with natural language. arXiv preprint arXiv:2209.05481.
  84. 84.Mujeen Sung, Minbyul Jeong, Yonghwa Choi, Donghyeon Kim, Jinhyuk Lee, and Jaewoo Kang. 2022. Bern2: an advanced neural biomedical named entity recognition and normalization tool. Bioinformatics, 38(20):4837–4839.
  85. 85.Baris E Suzek, Hongzhan Huang, Peter McGarvey, Raja Mazumder, and Cathy H Wu. 2007. Uniref: comprehensive and non-redundant uniprot reference clusters. Bioinformatics, 23(10):1282–1288.
  86. 86.Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085.
  87. 87.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  88. 88.Yuyang Wang, Jianren Wang, Zhonglin Cao, and Amir Barati Farimani. 2022. Molecular contrastive learning of representations via graph neural networks. Nature Machine Intelligence, 4(3):279–287.
  89. 89.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned language models are zero-shot learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  90. 90.David Weininger. 1988. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of chemical information and computer sciences, 28(1):31–36.
  91. 91.David Weininger, Arthur Weininger, and Joseph L Weininger. 1989. Smiles. 2. algorithm for generation of unique smiles notation. Journal of chemical information and computer sciences, 29(2):97–101.
  92. 92.David S Wishart, Yannick D Feunang, An C Guo, Elvis J Lo, Ana Marcu, Jason R Grant, Tanvir Sajed, Daniel Johnson, Carin Li, Zinat Sayeeda, et al. 2018. Drugbank 5.0: a major update to the drugbank database for 2018. Nucleic acids research, 46(D1):D1074–D1082.
  93. 93.Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. 2018. Moleculenet: a benchmark for molecular machine learning. Chemical science, 9(2):513–530.
  94. 94.Hanwen Xu, Addie Woicik, Hoifung Poon, Russ B Altman, and Sheng Wang. 2023a. Multilingual translation for zero-shot biomedical classification using biotranslator. Nature Communications, 14(1):738.
  95. 95.Minghao Xu, Xinyu Yuan, Santiago Miret, and Jian Tang. 2023b. Protst: Multi-modality learning of protein sequences and biomedical texts. arXiv preprint arXiv:2301.12040.
  96. 96.Minghao Xu, Zuobai Zhang, Jiarui Lu, Zhaocheng Zhu, Yangtian Zhang, Ma Chang, Runcheng Liu, and Jian Tang. 2022. Peer: a comprehensive and multi-task benchmark for protein sequence understanding. Advances in Neural Information Processing Systems, 35:35156–35173.
  97. 97.Zheni Zeng, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2022. A deep-learning system bridging molecule structure and biomedical text with comprehension comparable to human professionals. Nature communications, 13(1):862.
  98. 98.Zaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu, and Chee-Kong Lee. 2021. Motif-based graph self-supervised learning for molecular property prediction. Advances in Neural Information Processing Systems, 34:15870–15882.
  99. 99.Marinka Zitnik, Rok Sosic, and Jure Leskovec. 2018. Biosnap datasets: Stanford biomedical network dataset collection. Note: http://snap. stanford. edu/biodata Cited by, 5(1).

Citation

MLA
Pei, Q., et al. “BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1102–23, https://doi.org/10.18653/v1/2023.emnlp-main.70.
APA
Pei, Q., Zhang, W., Zhu, J., Wu, K., Gao, K., Wu, L., Xia, Y., & Yan, R. (2023). BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1102–1123. https://doi.org/10.18653/v1/2023.emnlp-main.70
Chicago
Pei, Q., W. Zhang, J. Zhu, et al. 2023. “BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1102–23. https://doi.org/10.18653/v1/2023.emnlp-main.70.
Harvard
Pei, Q. et al. (2023) “BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1102–1123. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.70.
Vancouver
1. Pei Q, Zhang W, Zhu J, Wu K, Gao K, Wu L, Xia Y, Yan R (2023) BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1102–1123

BibTeX

@inproceedings{pei-etal-2023-biot5,
    title = "{B}io{T}5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations",
    author = "Pei, Qizhi  and
      Zhang, Wei  and
      Zhu, Jinhua  and
      Wu, Kehan  and
      Gao, Kaiyuan  and
      Wu, Lijun  and
      Xia, Yingce  and
      Yan, Rui",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.70/",
    doi = "10.18653/v1/2023.emnlp-main.70",
    pages = "1102--1123"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/