MolXPT: Wrapping Molecules with Text for Generative Pre-training

Zequn LiuWei ZhangYingce XiaLijun WuShufang XieTao QinMing ZhangTie-Yan Liu

article2023ACL124 citations

Proposes a unified generative language model pre-trained on biomedical text where molecule names are replaced with SMILES sequences, achieving superior MoleculeNet property prediction and competitive text-molecule translation with fewer parameters.

Listen

Artificial intelligence holds significant promise for accelerating molecular discovery and drug development, yet existing computational models often struggle to fully bridge the gap between chemical structures and scientific literature. While textual publications provide rich contextual descriptions of molecular properties and behavior, standard molecular models typically process chemical sequences or graph structures in isolation. Earlier attempts to combine these modalities often failed to directly link a molecule's exact structural notation with its surrounding narrative description, leaving valuable contextual relationships untapped.

The article introduces and evaluates MolXPT, a unified generative language model designed to bridge molecular structures and scientific text. The primary objective is to demonstrate that pre-training a language model on scientific text intertwined with chemical sequences improves performance across downstream molecular prediction and cross-modal translation tasks, all while maintaining high parameter efficiency.

To achieve this, the authors constructed a pre-training corpus comprising 30 million biomedical paper titles and abstracts from PubMed, 30 million molecular sequences from PubChem using the standard simplified molecular-input line-entry system (known as SMILES), and 8 million "wrapped" sequences. These wrapped sequences were created by identifying chemical entity mentions within biomedical text and replacing the names directly with their corresponding SMILES representations. The authors then pre-trained a 24-layer, 350-million-parameter generative model on this combined dataset and evaluated its downstream performance using prompt-based finetuning across standard property prediction benchmarks and bidirectional text-molecule translation tasks.

The findings show that MolXPT achieved an average score of 81.9 across six MoleculeNet property classification tasks, outperforming specialized graph neural network baselines such as GEM (79.0) and large multimodal language models such as Galactica (69.5). On text-to-molecule and molecule-to-text translation tasks using the CheBI-20 benchmark, MolXPT performed comparably to or better than the leading baseline, MolT5-large, while using only 44% of that baseline's parameter count (350 million parameters compared to 800 million). In text-to-molecule generation, the model achieved a 98.3% valid molecule generation rate, significantly surpassing the 90.5% rate of the larger baseline. Furthermore, the model demonstrated zero-shot generation capabilities, successfully reproducing exact molecular structures from raw text prompts without any task-specific finetuning.

These results indicate that explicitly embedding chemical structure notations into surrounding textual narratives enables models to build richer, complementary representations of chemical components and functional properties. For organizations focused on drug discovery and molecular design, this approach provides a way to reduce model deployment costs and computational overhead without sacrificing predictive accuracy, while simultaneously opening possibilities for direct, text-guided molecular design.

Based on these outcomes, organizations should explore integrated molecule-text pre-training architectures for computer-aided drug design workflows and adopt prompt-based finetuning to streamline downstream property prediction. For future research, the authors recommend scaling up model parameters to assess further improvements in zero-shot learning, incorporating contrastive learning techniques to reinforce consistency across modalities, and expanding evaluations into complex tasks such as text-guided molecular optimization. Decision-makers should also institute rigorous data anonymization and privacy safeguards if training such models on clinical records.

The primary limitation of this study is the substantial computational resource requirement needed to pre-train large-scale generative models from scratch, although the release of pre-trained weights mitigates this for downstream users. While confidence in the benchmark performance is high, practitioners should note that zero-shot generation accuracy remains lower than fully finetuned pipelines, warranting appropriate validation when deploying zero-shot capabilities in production.

arXiv: 2305.10688
Cover for MolXPT: Wrapping Molecules with Text for Generative Pre-training

Abstract

Generative pre-trained Transformer (GPT) has demonstrates its great success in natural language processing and related techniques have been adapted into molecular modeling. Considering that text is the most important record for scientific discovery, in this paper, we propose MolXPT, a unified language model of text and molecules pre-trained on SMILES (a sequence representation of molecules) wrapped by text. Briefly, we detect the molecule names in each sequence and replace them to the corresponding SMILES. In this way, the SMILES could leverage the information from surrounding text, and vice versa. The above wrapped sequences, text sequences from PubMed and SMILES sequences from PubChem are all fed into a language model for pre-training. Experimental results demonstrate that MolXPT outperforms strong baselines of molecular property prediction on MoleculeNet, performs comparably to the best model in text-molecule translation while using less than half of its parameters, and enables zero-shot molecular generation without finetuning.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Our Method
  • 2.1 Pre-training corpus
  • 2.2 Model and training
  • 3 Experiments
  • 3.1 Results on MoleculeNet
  • 3.2 Results on text-molecule translation
  • 4 Conclusions and Future Work
  • Limitations
  • Broader Impacts
  • Acknowledgement
  • References
  • Appendix
  • A Datasets and Baselines of MoleculeNet
  • B Pre-training hyper-parameters
  • C Finetuning details of downstream tasks
  • C.1 Prompts for finetuning MoleculeNet
  • C.2 Details of finetuning MoleculeNet
  • C.3 Details of finetuning text-molecule generation
  • C.4 MoleculeNet finetuning strategy selection
  • D Zero-shot text-to-molecule generation

Knowls

  1. Knowl 1 — Replacing molecular mentions with SMILES creates text–molecule training sequences

    model/method

    MolXPT constructs “wrapped” sequences that place molecular representations directly in scientific text. For each text sequence, the pipeline uses BERN2 to detect molecular mentions and link them to entities in public knowledge bases such as ChEBI, retrieves the linked molecule’s SMILES, and replaces the mention with that SMILES. A wrapped sequence is retained only if it contains at least one SMILES. This gives the language model text context around a molecular representation, and molecular representations embedded in text context, during generative pre-training.

  2. Knowl 2 — MolXPT improves average MoleculeNet performance on scaffold-split classification tasks

    data/table

    MolXPT was evaluated on six MoleculeNet molecular-classification tasks using scaffold-based train, validation, and test splits. The metric is ROC-AUC, reported on a 0–100 scale; the table gives the reported results, including variability where available. MolXPT has the highest reported average among the listed models (81.9), and the highest task score on BBBP, ClinTox, BACE, and SIDER. It does not lead on Tox21 or HIV: GEM scores 78.1 and 80.6 on those tasks, respectively, compared with MolXPT’s 77.1 and 78.1. The molecule counts are the dataset sizes reported by the paper.

    Model BBBP Tox21 ClinTox HIV BACE SIDER Avg
    # Molecules 2039 7831 1478 41127 1513 1478 –
    G-Contextual 70.3 ± 1.6 75.2 ± 0.3 59.9 ± 8.2 75.9 ± 0.9 79.2 ± 0.3 58.4 ± 0.6 69.8
    G-Motif 66.4 ± 3.4 73.2 ± 0.8 77.8 ± 2.0 73.8 ± 1.4 73.4 ± 4.0 60.6 ± 1.1 70.9
    GROVERbase_{\text{base}} 70.0 ± 0.1 74.3 ± 0.1 81.2 ± 3.0 62.5 ± 0.9 82.6 ± 0.7 64.8 ± 0.6 72.6
    GROVERlarge_{\text{large}} 69.5 ± 0.1 73.5 ± 0.1 76.2 ± 3.7 68.2 ± 1.1 81.0 ± 1.4 65.4 ± 0.1 72.3
    GraphMVP 72.4 ± 1.6 75.9 ± 0.5 79.1 ± 2.8 77.0 ± 1.2 81.2 ± 0.9 63.9 ± 1.2 74.9
    MGSSL 70.5 ± 1.1 76.5 ± 0.3 80.7 ± 2.1 79.5 ± 1.1 79.7 ± 0.8 61.8 ± 0.8 74.8
    GEM 72.4 ± 0.4 78.1 ± 0.1 90.1 ± 1.3 80.6 ± 0.9 85.6 ± 1.1 67.2 ± 0.4 79.0
    KV-PLM 74.6 ± 0.9 72.7 ± 0.6 – 74.0 ± 1.2 – 61.5 ± 1.5 –
    Galactica 66.1 68.9 82.6 74.5 61.7 63.2 69.5
    MoMu 70.5 ± 2.0 75.6 ± 0.3 79.9 ± 4.1 76.2 ± 0.9 77.1 ± 1.4 60.5 ± 0.9 73.3
    MolXPT 80.0 ± 0.5 77.1 ± 0.2 95.3 ± 0.2 78.1 ± 0.4 88.4 ± 1.0 71.7 ± 0.2 81.9
  3. Knowl 3 — MolXPT is a 350M-parameter causal Transformer trained across text, SMILES, and wrapped sequences

    model/method

    MolXPT uses a GPT-style autoregressive Transformer with the GPT-2-medium configuration: 24 layers, hidden size 1024, and 16 attention heads. Its maximum input length is 2048 tokens, its vocabulary contains 44,536 tokens, and it has 350 million parameters. Pre-training combines scientific-text sequences, molecular SMILES sequences, and wrapped text–SMILES sequences under a causal negative log-likelihood objective. Let DD be the collection of training sequences; sequence ii contains tokens si,1,…,si,nis_{i,1},\ldots,s_{i,n_i}, and si,<js_{i,<j} denotes all preceding tokens in that sequence. The objective averages the summed token loss over sequences:

    L(θ)=−1∣D∣∑i=1∣D∣∑j=1nilog⁡Pθ(si,j∣si,<j).\mathcal{L}(\theta)=-\frac{1}{|D|}\sum_{i=1}^{|D|}\sum_{j=1}^{n_i}\log P_{\theta}(s_{i,j}\mid s_{i,<j}).

    Here, PθP_{\theta} is the next-token distribution of the model with parameters θ\theta.

  4. Knowl 4 — Pre-training uses millions of PubMed, PubChem, and entity-linked wrapped sequences

    experimental setup

    The pre-training corpus contains titles and abstracts from 30 million PubMed papers, 30 million randomly selected PubChem molecules, and 8 million wrapped sequences produced from text by replacing linked molecular mentions with SMILES. Text is segmented using byte-pair encoding with 40,000 merge operations. SMILES—including those embedded in wrapped sequences—are tokenized with the regular expression tokenizer of Schwaller et al. Each SMILES is marked with a start-of-molecule token, ⟨som⟩, and an end-of-molecule token, ⟨eom⟩.

  5. Knowl 5 — Prompt-based fine-tuning casts classification as next-token generation

    model/method

    MolXPT adapts downstream tasks by expressing inputs and outputs as prompted text and/or SMILES sequences and fine-tuning with the language-modeling objective. For MoleculeNet classification, the prompt states the property being predicted and includes the molecule between ⟨som⟩ and ⟨eom⟩; the class label is the final token. For example, a BBBP prompt asks whether the BBB penetration of the marked SMILES is true or false. The selected training strategy updates the model using the negative log-probability of the label token only, rather than the loss over every prompt token. At inference, the two class probabilities are normalized: for true and false probabilities ptruep_{\mathrm{true}} and pfalsep_{\mathrm{false}}, the positive-class probability is ptrue/(ptrue+pfalse)p_{\mathrm{true}}/(p_{\mathrm{true}}+p_{\mathrm{false}}), and the negative-class probability is pfalse/(ptrue+pfalse)p_{\mathrm{false}}/(p_{\mathrm{true}}+p_{\mathrm{false}}).

  6. Knowl 6 — MolXPT approaches MolT5-large on bidirectional molecule–text translation with fewer parameters

    data/table

    The evaluation uses the 33,010-pair ChEBI-20 dataset, split into 80% training, 10% validation, and 10% test data. For molecule-to-text generation, the prompt supplies a marked SMILES and asks for its description; for text-to-molecule generation, the description is followed by a prompt to generate a marked SMILES. Molecule-to-text quality is reported using BLEU, ROUGE, METEOR, and Text2Mol; text-to-molecule quality uses exact SMILES match, MACCS/RDK/Morgan fingerprint similarity, FCD, Text2Mol, and validity. Higher is better except for FCD, where lower is better. MolXPT (350M parameters) is comparable to MolT5-large (800M) on molecule-to-text and has the highest Text2Mol score in both directions. For text-to-molecule generation, MolXPT also leads on MACCS, RDK, FCD, and validity, while MolT5-large scores higher on exact match and Morgan similarity.

    Molecule-to-text model BLEU-2 BLEU-4 ROUGE-1 ROUGE-2 ROUGE-L METEOR Text2Mol
    MolT5-small (77M) 0.519 0.436 0.620 0.469 0.563 0.551 0.540
    MolT5-base (250M) 0.540 0.457 0.634 0.485 0.578 0.569 0.547
    MolT5-large (800M) 0.594 0.508 0.654 0.510 0.594 0.614 0.582
    MolXPT (350M) 0.594 0.505 0.660 0.511 0.597 0.626 0.594
  7. Knowl 7 — Pre-trained MolXPT generates molecules from text without task fine-tuning

    empirical result

    The authors tested zero-shot text-to-molecule generation by giving pre-trained MolXPT a text description and generating molecules without fine-tuning on the translation task. For KK generated candidates and reference molecule mm, top-KK fingerprint similarity is the maximum similarity between mm and any of the KK candidates. The table reports top-1 and top-5 results for zero-shot generation, and top-1 results after training on the full task data. Zero-shot performance is lower than full-data performance, but MolXPT produced 33 exact reference-molecule matches without fine-tuning.

    Setting MACCS RDK Morgan
    Zero-shot (Top-1) 0.540 0.383 0.228
    Zero-shot (Top-5) 0.580 0.423 0.423
    Full data (Top-1) 0.841 0.746 0.660
  8. Knowl 8 — Reported pre-training and downstream fine-tuning settings

    experimental setup

    MolXPT pre-training ran for 200,000 steps on eight A100 GPUs. Each GPU processed batches of 2,048 tokens; gradients were accumulated for 16 steps before an optimizer update. The optimizer was Adam, with peak learning rate 0.0005, 20,000 warm-up steps, inverse-square-root learning-rate decay, and dropout 0.1. For MoleculeNet, the authors searched learning rates 3×10−53\times10^{-5} and 5×10−55\times10^{-5}, dropout 0.1 and 0.3, and 30 or 50 epochs, selecting by validation performance. For molecule–text generation, they report fine-tuning for 100 steps on one P40 GPU, with 1,024 tokens and 16 accumulated steps per device, and also report 100 fine-tuning epochs; the learning rate was 0.0001. Dropout was selected from 0.1, 0.2, 0.3, 0.4, and 0.5, with 0.4 selected for molecule-to-text and 0.5 for text-to-molecule.

  9. Knowl 9 — Scaling MolXPT is limited by computational cost

    limitation

    The authors identify the computational resources required to train larger MolXPT models as a limitation, noting that the cost is relatively high. They state that releasing the pre-trained models would allow users to use them without repeating pre-training.

Coverage note — Future research directions and broader-impact discussion were omitted because they are prospective or contextual rather than demonstrated contributions; the stated computational limitation is included.

References

  1. 1.Viraj Bagal, Rishal Aggarwal, P. K. Vinod, and U. Deva Priyakumar. 2022. Molgpt: Molecular generation using a transformer-decoder model. Journal of Chemical Information and Modeling, 62(9):2064–2076.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the ACL workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72.
  3. 3.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3615–3620, Hong Kong, China. Association for Computational Linguistics.
  4. 4.Elliot Bolton, David Hall, Michihiro Yasunaga, Tony Lee, Chris Manning, and Percy Liang. 2022. PubMedGPT 2.7B.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  6. 6.Yulong Chen, Yang Liu, Li Dong, Shuohang Wang, Chenguang Zhu, Michael Zeng, and Yue Zhang. 2022. Adaprompt: Adaptive model training for prompt-based nlp. arXiv preprint arXiv:2202.04824.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  8. 8.Joseph L Durant, Burton A Leland, Douglas R Henry, and James G Nourse. 2002. Reoptimization of mdl keys for use in drug discovery. Journal of chemical information and computer sciences, 42(6):1273–1280.
  9. 9.Carl Edwards, Tuan Lai, Kevin Ros, Garrett Honke, and Heng Ji. 2022. Translation between molecules and natural language. arXiv preprint arXiv:2204.11817.
  10. 10.Carl Edwards, ChengXiang Zhai, and Heng Ji. 2021. Text2mol: Cross-modal molecule retrieval with natural language queries. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 595–607.
  11. 11.Xiaomin Fang, Lihang Liu, Jieqiong Lei, Donglong He, Shanzhuo Zhang, Jingbo Zhou, Fan Wang, Hua Wu, and Haifeng Wang. 2022. Geometry-enhanced molecular representation learning for property prediction. Nature Machine Intelligence, 4(2):127–134.
  12. 12.Minghao Feng, Bingqing Tang, Steven H Liang, and Xuefeng Jiang. 2016. Sulfur containing scaffolds in drugs: synthesis and application in medicinal chemistry. Current topics in medicinal chemistry, 16(11):1200–1216.
  13. 13.Noelia Ferruz, Steffen Schmidt, and Birte Höcker. 2022. Protgpt2 is a deep unsupervised language model for protein design. Nature Communications, 13(1):4348.
  14. 14.Nathan Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi, Rafael Gomez-Bombarelli, Connor Coley, and Vijay Gadepally. 2022. Neural scaling of deep chemical models. ChemRxiv.
  15. 15.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830.
  16. 16.Yuxian Gu, Xu Han, Zhiyuan Liu, and Minlie Huang. 2022. Ppt: Pre-trained prompt tuning for few-shot learning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8410–8423.
  17. 17.Janna Hastings, Gareth Owen, Adriano Dekker, Marcus Ennis, Namrata Kale, Venkatesh Muthukrishnan, Steve Turner, Neil Swainston, Pedro Mendes, and Christoph Steinbeck. 2016. Chebi in 2016: Improved services and an expanding collection of metabolites. Nucleic acids research, 44(D1):D1214–D1219.
  18. 18.David N Juurlink, Muhammad Mamdani, Alexander Kopp, Andreas Laupacis, and Donald A Redelmeier. 2003. Drug-drug interactions among elderly patients hospitalized for drug toxicity. Jama, 289(13):1652–1658.
  19. 19.Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, Leonid Zaslavsky, Jian Zhang, and Evan E Bolton. 2022. PubChem 2023 update. Nucleic Acids Research, 51(D1):D1373–D1380.
  20. 20.Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In ICLR (Poster).
  21. 21.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  22. 22.Shengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby, Hongyu Guo, and Jian Tang. 2022. Pre-training molecular graph representation with 3d geometry. In International Conference on Learning Representations.
  23. 23.Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. BioGPT: generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 23(6).
  24. 24.OpenAI. 2022. Chatgpt: Optimizing language models for dialogue. Technical blog.
  25. 25.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In ACL, pages 311–318.
  26. 26.Kristina Preuer, Philipp Renz, Thomas Unterthiner, Sepp Hochreiter, and Gunter Klambauer. 2018. Frechet chemnet distance: a metric for generative models for molecules in drug discovery. Journal of chemical information and modeling, 58(9):1736–1741.
  27. 27.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  28. 28.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  29. 29.David Rogers and Mathew Hahn. 2010. Extended-connectivity fingerprints. Journal of chemical information and modeling, 50(5):742–754.
  30. 30.Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying Wei, Wenbing Huang, and Junzhou Huang. 2020. Self-supervised graph transformer on large-scale molecular data. Advances in Neural Information Processing Systems, 33:12559–12571.
  31. 31.Nadine Schneider, Roger A Sayle, and Gregory A Landrum. 2015. Get your atoms in order: An open-source implementation of a novel and robust molecular canonicalization algorithm. Journal of chemical information and modeling, 55(10):2111–2120.
  32. 32.Philippe Schwaller, Theophile Gaudin, David Lanyi, Costas Bekas, and Teodoro Laino. 2018. “found in translation”: predicting outcomes of complex organic chemistry reactions using neural sequence-to-sequence models. Chemical science, 9(28):6091–6098.
  33. 33.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  34. 34.Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2022. Large language models encode clinical knowledge. arXiv preprint arXiv:2212.13138.
  35. 35.Bing Su, Dazhao Du, Zhao Yang, Yujie Zhou, Jiangmeng Li, Anyi Rao, Hao Sun, Zhiwu Lu, and Ji-Rong Wen. 2022. A molecular multimodal foundation model associating molecule graphs with natural language. arXiv preprint arXiv:2209.05481.
  36. 36.Mujeen Sung, Minbyul Jeong, Yonghwa Choi, Donghyeon Kim, Jinhyuk Lee, and Jaewoo Kang. 2022. Bern2: an advanced neural biomedical named entity recognition and normalization tool. arXiv preprint arXiv:2201.02080.
  37. 37.Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085.
  38. 38.Xiaochu Tong, Xiaohong Liu, Xiaoqin Tan, Xutong Li, Jiaxin Jiang, Zhaoping Xiong, Tingyang Xu, Hualiang Jiang, Nan Qiao, and Mingyue Zheng. 2021. Generative models for de novo drug design. Journal of Medicinal Chemistry, 64(19):14011–14027.
  39. 39.David Weininger. 1988. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of chemical information and computer sciences, 28(1):31–36.
  40. 40.Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. 2018. Moleculenet: a benchmark for molecular machine learning. Chemical science, 9(2):513–530.
  41. 41.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022. An explanation of in-context learning as implicit bayesian inference. In International Conference on Learning Representations.
  42. 42.Zheni Zeng, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2022. A deep-learning system bridging molecule structure and biomedical text with comprehension comparable to human professionals. Nature Communications, 13(1):862.
  43. 43.Zaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu, and Chee-Kong Lee. 2021. Motif-based graph self-supervised learning for molecular property prediction. Advances in Neural Information Processing Systems, 34:15870–15882.

Citation

MLA
Liu, Z., et al. “MolXPT: Wrapping Molecules with Text for Generative Pre-training”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2023, pp. 1606–16, https://doi.org/10.18653/v1/2023.acl-short.138.
APA
Liu, Z., Zhang, W., Xia, Y., Wu, L., Xie, S., Qin, T., Zhang, M., & Liu, T.-Y. (2023). MolXPT: Wrapping Molecules with Text for Generative Pre-training. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 1606–1616. https://doi.org/10.18653/v1/2023.acl-short.138
Chicago
Liu, Z., W. Zhang, Y. Xia, et al. 2023. “MolXPT: Wrapping Molecules with Text for Generative Pre-training”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 1606–16. https://doi.org/10.18653/v1/2023.acl-short.138.
Harvard
Liu, Z. et al. (2023) “MolXPT: Wrapping Molecules with Text for Generative Pre-training”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp. 1606–1616. Available at: https://doi.org/10.18653/v1/2023.acl-short.138.
Vancouver
1. Liu Z, Zhang W, Xia Y, Wu L, Xie S, Qin T, Zhang M, Liu T-Y (2023) MolXPT: Wrapping Molecules with Text for Generative Pre-training. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, pp 1606–1616

BibTeX

@inproceedings{liu-etal-2023-molxpt,
    title = "{M}ol{XPT}: Wrapping Molecules with Text for Generative Pre-training",
    author = "Liu, Zequn  and
      Zhang, Wei  and
      Xia, Yingce  and
      Wu, Lijun  and
      Xie, Shufang  and
      Qin, Tao  and
      Zhang, Ming  and
      Liu, Tie-Yan",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-short.138/",
    doi = "10.18653/v1/2023.acl-short.138",
    pages = "1606--1616"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/