MolXPT: Wrapping Molecules with Text for Generative Pre-training
Zequn LiuWei ZhangYingce XiaLijun WuShufang XieTao QinMing ZhangTie-Yan Liu
Proposes a unified generative language model pre-trained on biomedical text where molecule names are replaced with SMILES sequences, achieving superior MoleculeNet property prediction and competitive text-molecule translation with fewer parameters.
Artificial intelligence holds significant promise for accelerating molecular discovery and drug development, yet existing computational models often struggle to fully bridge the gap between chemical structures and scientific literature. While textual publications provide rich contextual descriptions of molecular properties and behavior, standard molecular models typically process chemical sequences or graph structures in isolation. Earlier attempts to combine these modalities often failed to directly link a molecule's exact structural notation with its surrounding narrative description, leaving valuable contextual relationships untapped.
The article introduces and evaluates MolXPT, a unified generative language model designed to bridge molecular structures and scientific text. The primary objective is to demonstrate that pre-training a language model on scientific text intertwined with chemical sequences improves performance across downstream molecular prediction and cross-modal translation tasks, all while maintaining high parameter efficiency.
To achieve this, the authors constructed a pre-training corpus comprising 30 million biomedical paper titles and abstracts from PubMed, 30 million molecular sequences from PubChem using the standard simplified molecular-input line-entry system (known as SMILES), and 8 million "wrapped" sequences. These wrapped sequences were created by identifying chemical entity mentions within biomedical text and replacing the names directly with their corresponding SMILES representations. The authors then pre-trained a 24-layer, 350-million-parameter generative model on this combined dataset and evaluated its downstream performance using prompt-based finetuning across standard property prediction benchmarks and bidirectional text-molecule translation tasks.
The findings show that MolXPT achieved an average score of 81.9 across six MoleculeNet property classification tasks, outperforming specialized graph neural network baselines such as GEM (79.0) and large multimodal language models such as Galactica (69.5). On text-to-molecule and molecule-to-text translation tasks using the CheBI-20 benchmark, MolXPT performed comparably to or better than the leading baseline, MolT5-large, while using only 44% of that baseline's parameter count (350 million parameters compared to 800 million). In text-to-molecule generation, the model achieved a 98.3% valid molecule generation rate, significantly surpassing the 90.5% rate of the larger baseline. Furthermore, the model demonstrated zero-shot generation capabilities, successfully reproducing exact molecular structures from raw text prompts without any task-specific finetuning.
These results indicate that explicitly embedding chemical structure notations into surrounding textual narratives enables models to build richer, complementary representations of chemical components and functional properties. For organizations focused on drug discovery and molecular design, this approach provides a way to reduce model deployment costs and computational overhead without sacrificing predictive accuracy, while simultaneously opening possibilities for direct, text-guided molecular design.
Based on these outcomes, organizations should explore integrated molecule-text pre-training architectures for computer-aided drug design workflows and adopt prompt-based finetuning to streamline downstream property prediction. For future research, the authors recommend scaling up model parameters to assess further improvements in zero-shot learning, incorporating contrastive learning techniques to reinforce consistency across modalities, and expanding evaluations into complex tasks such as text-guided molecular optimization. Decision-makers should also institute rigorous data anonymization and privacy safeguards if training such models on clinical records.
The primary limitation of this study is the substantial computational resource requirement needed to pre-train large-scale generative models from scratch, although the release of pre-trained weights mitigates this for downstream users. While confidence in the benchmark performance is high, practitioners should note that zero-shot generation accuracy remains lower than fully finetuned pipelines, warranting appropriate validation when deploying zero-shot capabilities in production.
- Paper: Translation between Molecules and Natural Language, Carl Edwards et al. (2022). MolT5 establishes the joint pretraining and bidirectional molecule–text translation setup that MolXPT adapts through text-wrapped molecular sequences.
- Paper: XLNet: Generalized Autoregressive Pretraining for Language Understanding, Zhilin Yang et al. (2019). XLNet clarifies the autoregressive GPT-style pretraining lineage underlying MolXPT’s language-modeling approach.
- Paper: MoleculeNet: a benchmark for molecular machine learning, Zhenqin Wu et al. (2017). MoleculeNet defines the benchmark suite used to assess molecular property prediction, making MolXPT’s reported results easier to interpret.
- Paper: BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations, Qizhi Pei et al. (2023). BioT5 carries text-wrapped biomedical pretraining forward by adding protein sequences, SELFIES, and richer cross-modal objectives.
- Paper: Unifying Molecular and Textual Representations via Multi-task Language Modelling, Dimitrios Christofidellis et al. (2023). Text+Chem T5 extends unified text-and-chemistry language modeling to balanced multitask training across reactions, molecular generation, and procedural text.
- Paper: MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter, Zhiyuan Liu et al. (2023). MolCA advances molecule–language generation beyond SMILES strings by conditioning language models directly on molecular graph structure.
