Unifying Molecular and Textual Representations via Multi-task Language Modelling
Dimitrios ChristofidellisGiorgio GiannoneJannis BornOle WintherTeodoro LainoMatteo Manica
Introduces Text+Chem T5, a unified multi-task language model that bridges natural language and chemical structures to execute both single-domain and cross-domain chemistry workflows without needing task-specific architectures or expensive single-modality pre-training.
Data-driven scientific discovery increasingly relies on generative language models to design molecules, predict chemical reactions, and automate laboratory procedures. However, existing artificial intelligence approaches typically require separate specialized models for each individual task or depend on costly pre-training across separate single-domain datasets. This separation creates a barrier between natural language text and chemical representations, complicating human-machine collaboration and preventing models from sharing knowledge across related tasks.
The article demonstrates and evaluates a unified, multi-domain, multi-task framework called Multitask Text and Chemistry T5 (Text+Chem T5). The primary objective is to enable a single model to process natural language and chemical notations simultaneously, performing text-based, chemistry-based, and cross-domain translation tasks without specialized task-specific heads or expensive single-domain pre-training.
The authors implemented their approach using a standard encoder-decoder language architecture initialized from natural language models (ranging from 60 million to 220 million parameters). They trained the model on balanced multi-task datasets comprising 11.5 million and 33.5 million samples gathered from chemical reaction records, procedural reaction literature, and chemical text-structure databases. The model was evaluated across five core benchmarks: forward reaction prediction, retrosynthesis, molecular captioning, text-conditional molecule generation, and procedural action extraction from text.
The evaluation yielded several key findings. First, Text+Chem T5 outperformed existing specialized models on cross-domain tasks. For text-to-molecule generation, the base model achieved an accuracy of 32.2%—a fourfold improvement over the specialized MolT5 baseline (8.1%)—alongside superior sequence similarity scores. Second, for molecular captioning, the system consistently scored highest across standard text evaluation metrics, achieving a BLEU-4 score of 0.542 compared to 0.457 for MolT5-base. Third, architectural ablations revealed that sharing and fine-tuning a unified encoder across all domains yielded much better results than using separate domain-specific encoders with late-stage attention merging. Fourth, the model retained competitive performance on purely chemical and textual benchmarks while scaling more effectively with model and dataset size than baseline approaches. Finally, end-to-end qualitative testing showed that Text+Chem T5 successfully navigated a multi-step discovery workflow (generating a molecule from description, planning its synthesis, and extracting laboratory actions), whereas general foundation models like ChatGPT and Galactica produced invalid or incorrect outputs.
These results demonstrate that a single, unified model can replace fragmented, task-specific pipelines in automated discovery workflows. By removing the need for costly separate pre-training runs and specialized engineering for each chemical task, organizations can substantially reduce computational expenses, shorten model development timelines, and streamline human-AI interaction in digital laboratories.
Organizations developing computational chemistry and automated synthesis platforms should consider adopting multi-task, shared-encoder architectures to consolidate their workflow infrastructure. Teams planning implementation should prioritize training on balanced, high-volume datasets, as expanded data directly improves chemical string accuracy. Further pilot testing is warranted to integrate these models with laboratory automation hardware for automated experimentation.
Decision-makers should note certain limitations: chemical representations rely on string-based notations (SMILES), which carry a minor risk of generating invalid molecular structures. In addition, the training data reflects inherent biases from historical literature, and all outputs remain intended for research purposes. Any computational predictions must undergo standard experimental and clinical validation prior to practical chemical or medical deployment.
- Paper: Translation between Molecules and Natural Language, Carl Edwards et al. (2022). MolT5 establishes bidirectional molecule–text translation through joint pretraining, the direct precursor to the source’s unified treatment of chemical and natural-language tasks.
- Paper: An Overview of Multi-Task Learning in Deep Neural Networks, Sebastian Ruder (2017). Ruder’s overview explains shared-parameter multi-task learning, the central mechanism the source applies across language and chemistry.
- Paper: A unified architecture for natural language processing: deep neural networks with multitask learning, Ronan Collobert et al. (2008). Collobert and Weston demonstrate an early unified, jointly trained NLP architecture, providing the multi-task learning precedent that informs the source’s cross-domain approach.
No sufficiently relevant recommendations were found.
