MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter
Zhiyuan LiuSihang LiYanchen LuoHao FeiYixin CaoKenji KawaguchiXiang WangTat-Seng Chua
Presents MolCA, a framework that couples 2D molecular graph encoders with language models using a Q-Former cross-modal projector and LoRA adapters to achieve state-of-the-art performance in molecule captioning, retrieval, and IUPAC name prediction.
Standard artificial intelligence language models have demonstrated broad scientific understanding, but they typically represent chemical molecules only as one-dimensional text strings. This string-based approach fails to capture two-dimensional topological graph structures, which are vital for human chemists to understand molecular connectivity and behavior. Conventional multi-modal approaches connect text models and graph encoders using contrastive learning, which works well for cross-modal search and retrieval but cannot handle open-ended text generation tasks such as describing a compound or predicting standard chemical names.
The article introduces and evaluates MolCA (Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter), an artificial intelligence framework designed to enable language models to understand two-dimensional molecular graph structures as direct inputs for text generation and retrieval tasks.
The researchers developed a three-component architecture connecting a molecular graph neural network to a base scientific language model (Galactica) using a Querying-Transformer (Q-Former) cross-modal projector. To maintain computational efficiency, they integrated a low-rank adapter (LoRA) that trains less than one percent of the language model's parameters during downstream fine-tuning. The framework was trained across three stages using PubChem324k, a newly curated dataset of 324,000 molecule-text pairs. The authors evaluated MolCA across molecule captioning, standardized chemical naming (IUPAC name prediction), cross-modal molecule-text retrieval, property prediction, and functional group counting benchmarks.
The evaluations demonstrated four central findings. First, MolCA established new state-of-the-art performance in molecule-to-text generation, improving molecule captioning by 7.6 BLEU-2 on PubChem324k and 2.1 BLEU-2 on CheBI-20 over competitive baselines. Second, in standardized IUPAC naming tasks, MolCA surpassed baselines by 10.0 BLEU-2, demonstrating a superior grasp of molecular topology. Third, the system improved molecule-text retrieval accuracy by more than 20% compared to prior models on the PubChem324k benchmark. Finally, ablation studies showed that combining two-dimensional graph representations with one-dimensional text representations significantly improved downstream property prediction and functional group counting compared to text-only approaches.
These findings demonstrate that bridging the gap between two-dimensional topological structures and natural language models substantially enhances automated biochemical reasoning. By enabling language models to interpret graph structures directly, organizations can lower computational and fine-tuning costs while automating complex chemical analysis and literature querying without training massive language models from scratch.
For future development, the authors recommend expanding molecular graph-language frameworks into three-dimensional molecular modeling and drug discovery pipelines. Organizations looking to adopt these methods should test generated descriptions thoroughly and explore weakly supervised data mining from biomedical literature to increase training volume, as larger datasets are required to achieve commercial-grade reliability.
Key limitations include the risk of standard language model hallucinations and factual inaccuracies in open-ended text generation. Additionally, the pretraining dataset size of roughly 324,000 pairs remains modest compared to multi-modal vision-language corpuses, meaning the current system is not yet fully sufficient for zero-defect production workflows without human expert oversight.
- Paper: Translation between Molecules and Natural Language, Carl Edwards et al. (2022). Its MolT5 framework establishes molecule-to-text captioning and cross-modal retrieval, providing the direct molecular-language precedent that MolCA adapts with graph inputs.
No sufficiently relevant recommendations were found.
