Translation between Molecules and Natural Language
Carl EdwardsTuan Manh LaiKevin RosGarrett HonkeKyunghyun ChoHeng Ji
Presents MolT5, a self-supervised framework pretrained on unlabeled text and chemical strings to bridge language and chemistry through bidirectional translation tasks like molecule captioning and text-conditioned molecule generation.
Designing new chemical compounds has historically relied on manual, trial-and-error laboratory development that costs billions of dollars and spans over a decade. While computational chemistry and deep learning have begun assisting in drug design, existing tools primarily target narrow numerical properties rather than functional, human-understandable requirements. Creating automated systems that can understand and generate chemical structures from text has been hindered by a severe shortage of paired molecule-and-text data, as manual chemical annotation requires deep domain expertise.
The article evaluates MolT5, a self-supervised learning framework designed to translate bidirectionally between molecular structures and natural language. Specifically, the article demonstrates two novel tasks: generating descriptive captions for molecular structures and generating new molecular structures directly from plain-text descriptions.
The researchers addressed the data bottleneck by pretraining an encoder-decoder model on massive unaligned datasets: standard English web text and 100 million molecular text representations from public chemistry databases. Using a denoising objective, the model learned the underlying rules of both modalities simultaneously without requiring initial pairwise alignments. The model was subsequently finetuned on a benchmark dataset of 33,010 expert-annotated molecule-description pairs. Evaluation was conducted using standard natural language generation metrics, chemical fingerprint similarity scores, and a specialized cross-modal retrieval model to measure how accurately generated outputs matched their corresponding inputs.
The primary finding is that joint pretraining on language and molecular structures enables effective translation in both directions, substantially outperforming traditional recurrent neural networks and standard architectures trained from scratch. In molecule captioning, the largest model variant achieved statistically significant improvements across language quality and cross-modal retrieval metrics, correctly identifying complex molecular classes and functional roles. In molecule generation from text, the framework achieved an exact chemical match rate of 31.1% on test samples—representing an approximate 11% relative increase over standard language models of equivalent scale—while maintaining over 90% chemical validity. Applying specialized diverse decoding strategies further raised the syntactic validity of generated molecules to as high as 99.6%.
These findings indicate that natural language can serve as an effective interface for chemical design, allowing scientists to specify functional requirements directly in plain text. Translating between natural language and molecular structures has the potential to shorten drug discovery timelines and reduce early-stage research costs. While larger general-purpose language models exhibit surprising chemical generation capabilities, adding explicit molecular pretraining provides critical improvements in generating valid, structurally accurate compounds.
Organizations should treat these findings as proof of concept for text-guided chemical design. Decision-makers should consider supporting pilot workflows where natural language models assist domain experts in exploring molecular candidates, provided that all model-generated molecules are rigorously validated through standard laboratory synthesis and clinical testing before real-world deployment.
A primary limitation of this work is the potential for bias inherited from public web corpora, which may influence generation outputs. Additionally, linear text representations of molecules can occasionally yield syntactically invalid structures, and the current benchmark evaluation relies on single-reference descriptions. The article's empirical comparisons provide moderate-to-high confidence in the model's computational capabilities, but decision-makers must exercise caution until prospective laboratory evaluations confirm biological efficacy.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). It introduces the multilingual T5 text-to-text transformer architecture and span-corruption pre-training framework that MolT5 directly adapts for cross-modal chemistry and text translation.
- Paper: Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules, Rafael Gómez-Bombarelli et al. (2016). It establishes the foundational approach of using string-based chemical representations (SMILES) in sequence models for generative molecular design.
- Paper: MoleculeNet: a benchmark for molecular machine learning, Zhenqin Wu et al. (2017). It introduces standard molecular machine learning datasets and benchmark evaluation protocols for computational chemistry tasks.
- Paper: LXMERT: Learning Cross-Modality Encoder Representations from Transformers, Hao Tan et al. (2019). It details foundational cross-modal transformer pre-training methodologies that inspire cross-modal translation and evaluation frameworks between text and non-text modalities.
- Paper: ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts, Minghao Xu et al. (2023). It extends the concept of bidirectional translation between biochemical sequences and natural language descriptions to the protein domain through multimodal alignment.
- Paper: A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery, Yu Zhang et al. (2024). It provides a comprehensive survey synthesizing multi-modal scientific language modeling paradigms, contextualizing MolT5 within the broader landscape of domain-specific scientific LLMs.
- Paper: LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery, Pingchuan Ma et al. (2024). It applies language-model-driven molecular generation within a bilevel optimization framework that incorporates physical simulations for targeted property design.
- Paper: Innovator-VL: A Multimodal Large Language Model for Scientific Discovery, Zichen Wen et al. (2026). It extends multimodal reasoning and scientific cross-modality learning to multimodal large language models operating over diverse scientific notations and visual representations.
