MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter

Zhiyuan LiuSihang LiYanchen LuoHao FeiYixin CaoKenji KawaguchiXiang WangTat-Seng Chua

article2023EMNLP184 citations

Presents MolCA, a framework that couples 2D molecular graph encoders with language models using a Q-Former cross-modal projector and LoRA adapters to achieve state-of-the-art performance in molecule captioning, retrieval, and IUPAC name prediction.

Listen

Standard artificial intelligence language models have demonstrated broad scientific understanding, but they typically represent chemical molecules only as one-dimensional text strings. This string-based approach fails to capture two-dimensional topological graph structures, which are vital for human chemists to understand molecular connectivity and behavior. Conventional multi-modal approaches connect text models and graph encoders using contrastive learning, which works well for cross-modal search and retrieval but cannot handle open-ended text generation tasks such as describing a compound or predicting standard chemical names.

The article introduces and evaluates MolCA (Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter), an artificial intelligence framework designed to enable language models to understand two-dimensional molecular graph structures as direct inputs for text generation and retrieval tasks.

The researchers developed a three-component architecture connecting a molecular graph neural network to a base scientific language model (Galactica) using a Querying-Transformer (Q-Former) cross-modal projector. To maintain computational efficiency, they integrated a low-rank adapter (LoRA) that trains less than one percent of the language model's parameters during downstream fine-tuning. The framework was trained across three stages using PubChem324k, a newly curated dataset of 324,000 molecule-text pairs. The authors evaluated MolCA across molecule captioning, standardized chemical naming (IUPAC name prediction), cross-modal molecule-text retrieval, property prediction, and functional group counting benchmarks.

The evaluations demonstrated four central findings. First, MolCA established new state-of-the-art performance in molecule-to-text generation, improving molecule captioning by 7.6 BLEU-2 on PubChem324k and 2.1 BLEU-2 on CheBI-20 over competitive baselines. Second, in standardized IUPAC naming tasks, MolCA surpassed baselines by 10.0 BLEU-2, demonstrating a superior grasp of molecular topology. Third, the system improved molecule-text retrieval accuracy by more than 20% compared to prior models on the PubChem324k benchmark. Finally, ablation studies showed that combining two-dimensional graph representations with one-dimensional text representations significantly improved downstream property prediction and functional group counting compared to text-only approaches.

These findings demonstrate that bridging the gap between two-dimensional topological structures and natural language models substantially enhances automated biochemical reasoning. By enabling language models to interpret graph structures directly, organizations can lower computational and fine-tuning costs while automating complex chemical analysis and literature querying without training massive language models from scratch.

For future development, the authors recommend expanding molecular graph-language frameworks into three-dimensional molecular modeling and drug discovery pipelines. Organizations looking to adopt these methods should test generated descriptions thoroughly and explore weakly supervised data mining from biomedical literature to increase training volume, as larger datasets are required to achieve commercial-grade reliability.

Key limitations include the risk of standard language model hallucinations and factual inaccuracies in open-ended text generation. Additionally, the pretraining dataset size of roughly 324,000 pairs remains modest compared to multi-modal vision-language corpuses, meaning the current system is not yet fully sufficient for zero-defect production workflows without human expert oversight.

  • Paper: Translation between Molecules and Natural Language, Carl Edwards et al. (2022). Its MolT5 framework establishes molecule-to-text captioning and cross-modal retrieval, providing the direct molecular-language precedent that MolCA adapts with graph inputs.

No sufficiently relevant recommendations were found.

Cover for MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter

Abstract

Language Models (LMs) have demonstrated impressive molecule understanding ability on various 1D text-related tasks. However, they inherently lack 2D graph perception — a critical ability of human professionals in comprehending molecules’ topological structures. To bridge this gap, we propose MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter. MolCA enables an LM (i.e., Galactica) to understand both text- and graph-based molecular contents via the cross-modal projector. Specifically, the cross-modal projector is implemented as a Q-Former to connect a graph encoder’s representation space and an LM’s text space. Further, MolCA employs a uni-modal adapter (i.e., LoRA) for the LM’s efficient adaptation to downstream tasks. Unlike previous studies that couple an LM with a graph encoder via cross-modal contrastive learning, MolCA retains the LM’s ability of open-ended text generation and augments it with 2D graph information. To showcase its effectiveness, we extensively benchmark MolCA on tasks of molecule captioning, IUPAC name prediction, and molecule-text retrieval, on which MolCA significantly outperforms the baselines. Our codes and checkpoints can be found at https://github.com/acharkq/MolCA.

Table of Contents

  • 1 Introduction
  • 2 Model Architecture
  • 3 Training Pipeline
  • 3.1 Pretrain Stage 1: Learning to Extract Text Relevant Molecule Representations
  • 3.2 Pretrain Stage 2: Aligning 2D Molecular Graphs to Texts via Language Modeling
  • 3.3 Fine-tune Stage: Uni-Modal Adapter for Efficient Downstream Adaptation
  • 4 Experiments
  • 4.1 Experimental Setting
  • 4.2 Molecule Captioning
  • 4.3 IUPAC Name Prediction
  • 4.4 Molecule-Text Retrieval
  • 4.5 Ablation Study on Representation Types
  • 5 Related Works
  • 6 Conclusion and Future Works
  • Limitations
  • Broader Impacts
  • Acknowledgement
  • References
  • A Complete Related Works
  • B Experimental Settings
  • C More Experimental Results

Knowls

  1. Knowl 1 — MolCA connects molecular graphs to a generative language model

    model/method

    MolCA represents a molecule both as a 2D graph and, where used, as a 1D SMILES string. A five-layer GINE graph encoder produces structure-aware representations for graph nodes; it is initialized by contrastive pretraining on 2 million ZINC15 molecules. A Q-Former cross-modal projector uses eight learnable query tokens and cross-attention to extract molecular information from the graph representations and map it into a form usable as soft prompts by Galactica, the decoder-only language model. The Q-Former is initialized from SciBERT, with its cross-attention modules randomly initialized. MolCA can condition generation on the graph-derived prompts together with the molecule’s SMILES, allowing the language model to use both structural and sequence information.

  2. Knowl 2 — MolCA trains in two alignment stages before downstream adaptation

    model/method

    MolCA uses a three-stage pipeline on paired molecule–text data. In stage 1, the graph encoder and Q-Former are jointly trained using molecule–text contrasting, molecule–text matching, and molecule captioning. In stage 2, graph-derived Q-Former representations and the molecule’s SMILES are supplied to a frozen language model; the graph encoder and Q-Former are trained to make the language model generate the paired description. In stage 3, the model is adapted to a downstream generation task using a task prompt and language-modeling loss, with LoRA updates to selected language-model weights rather than full language-model fine-tuning. The reported schedule is 50 epochs for stage 1, 10 for stage 2, and 100 for fine-tuning. Pretraining uses AdamW with weight decay 0.05, a peak learning rate of 1e-4, linear warmup for 1,000 steps, and cosine decay.

  3. Knowl 3 — Stage 1 combines contrastive, matching, and captioning objectives

    model/method

    During MolCA’s first pretraining stage, the graph encoder and Q-Former are optimized jointly with three complementary tasks. For molecule–text contrasting, the model separately encodes each text and each molecule; it compares a text representation with the most similar of the molecule’s query-token representations and uses temperature-scaled cosine similarity in both graph-to-text and text-to-graph directions. The temperature is 0.1. For molecule–text matching, the queries and text interact through self-attention, and a classifier over the mean-pooled query representations predicts whether a pair matches; random in-batch molecules and texts provide negative pairs. For captioning, query tokens can attend bidirectionally to one another but not to text tokens, while text tokens attend to the queries and preceding text tokens only. This causal masking makes the queries supply molecular information for autoregressive text generation. For retrieval, MolCA first uses the contrasting score to retrieve candidates and then uses the matching classifier to rerank them.

  4. Knowl 4 — PubChem324k supplies pretraining pairs and a filtered generation split

    experimental setup

    PubChem324k was assembled from molecule descriptions on the PubChem website. To reduce leakage, common or IUPAC molecule names at the start of descriptions were replaced with a generic template. Because many descriptions were uninformative, descriptions longer than 19 words were selected as a 15,000-pair downstream subset, split into 12,000 training, 1,000 validation, and 2,000 test pairs. The reported average molecule lengths for these splits were 32, 32, and 31, and average text lengths were 60, 61, and 60 words, respectively. The remaining data were used for pretraining, after filtering out molecules in the validation and test sets of CheBI-20, PCDes, and MoMu; the resulting pretraining split contains 298,083 pairs. Text length was counted by splitting on spaces.

  5. Knowl 5 — MolCA improves molecule captioning on PubChem324k and CheBI-20

    empirical result

    On molecule captioning, MolCA with Galactica 1.3B and LoRA achieved 38.7 BLEU-2, 30.3 BLEU-4, 50.2 ROUGE-1, 35.9 ROUGE-2, 44.5 ROUGE-L, and 45.6 METEOR on PubChem324k. The strongest listed baseline, MoMu-Large, scored 31.1 BLEU-2 and 22.8 BLEU-4 there, so the BLEU-2 gain was 7.6 points. On CheBI-20, the same MolCA configuration achieved 62.0 BLEU-2, 53.1 BLEU-4, 68.1 ROUGE-1, 53.7 ROUGE-2, 61.8 ROUGE-L, and 65.1 METEOR; the strongest listed baseline, MoMu-Large, scored 59.9 BLEU-2, a 2.1-point difference. MolCA with the smaller Galactica 125M also exceeded the larger listed baselines on the reported captioning metrics: it scored 31.9 BLEU-2 on PubChem324k and 61.2 on CheBI-20. Models were fine-tuned on each dataset’s training split, with test performance selected using validation performance.

  6. Knowl 6 — MolCA improves IUPAC-name generation on PubChem324k

    empirical result

    For IUPAC-name prediction on PubChem324k, MolCA with Galactica 1.3B and LoRA scored 75.0 BLEU-2, 66.6 BLEU-4, 69.6 ROUGE-1, 48.2 ROUGE-2, 63.4 ROUGE-L, and 72.1 METEOR. MolCA with Galactica 125M scored 73.9, 66.3, 69.0, 47.8, 63.2, and 71.8 on those same metrics. The strongest listed SMILES-only baseline, MolT5-Large, scored 59.4 BLEU-2, 49.7 BLEU-4, 55.9 ROUGE-1, 33.3 ROUGE-2, 49.1 ROUGE-L, and 58.5 METEOR. The models were fine-tuned on the PubChem324k training split; the task prompt asked for the molecule’s IUPAC name.

  7. Knowl 7 — MolCA improves molecule–text retrieval across three datasets

    empirical result

    MolCA’s stage-1 checkpoint was evaluated for retrieval without downstream fine-tuning. It retrieved the top 128 candidates using molecule–text contrasting and reranked them with molecule–text matching. On the PubChem324k test set, MolCA achieved 66.6% molecule-to-text accuracy and 94.6% Recall@20, and 66.0% text-to-molecule accuracy and 93.5% Recall@20. MoleculeSTM, the strongest listed baseline for accuracy in both directions, achieved 45.8% and 44.3% accuracy, respectively. On PCDes, MolCA achieved 85.6% molecule-to-text and 82.3% text-to-molecule Recall@20, compared with the best listed baseline scores of 80.4% and 79.0%. On MoMu, MolCA achieved 76.8% and 73.3% Recall@20, compared with MoleculeSTM’s 70.5% and 66.9%. On PubChem324k, removing the matching-based reranking reduced accuracy to 58.3% and 56.0%, indicating that matching contributed substantially to the final retrieval results.

  8. Knowl 8 — Combining graph and SMILES inputs benefits generation and property prediction

    empirical result

    Representation ablations show that using both 2D graphs and 1D SMILES can outperform either representation alone in molecule-to-text generation. On PubChem324k captioning, the Galactica 1.3B model scored 34.6 BLEU-2 with SMILES only, 34.5 with graphs only, and 38.7 with both. On CheBI-20 captioning, SMILES-only scored 55.3 BLEU-2, while the combined model scored 62.0; for PubChem324k IUPAC prediction, the corresponding scores were 71.0 and 75.0. In molecule property prediction, combining graphs and SMILES raised mean ROC-AUC across six MoleculeNet datasets from 72.1% to 74.0%. The combined model improved ROC-AUC on Bace (79.3±0.8% to 79.8±0.5%), ClinTox (89.0±1.7% to 89.5±0.7%), ToxCast (56.2±0.7% to 64.5±0.8%), Sider (61.1±1.2% to 63.0±1.7%), and Tox21 (76.0±0.5% to 77.2±0.5%); BBBP decreased from 70.8±0.6% to 70.0±0.5%. Property results are means and standard deviations across three random seeds using scaffold splits.

  9. Knowl 9 — Graph input improves functional-group counting

    empirical result

    MolCA was evaluated on counting 85 functional-group types in molecules. RDKit-derived counts served as targets, and the model used a separate linear regressor for each group type with mean squared error training. On the PubChem324k training split, the models were fine-tuned and evaluated on the validation split; reported curves average three random seeds, with variation shown as one standard deviation. The graph-plus-SMILES model achieved lower validation RMSE than the SMILES-only model over fine-tuning, supporting the paper’s finding that 2D graph input improves functional-group counting.

  10. Knowl 10 — Ablations support both pretraining stages and the Q-Former projector

    empirical result

    On PubChem324k captioning with Galactica 1.3B, a model with neither pretraining stage scored 35.8 BLEU-2 and 27.6 BLEU-4; adding stage 1 raised these scores to 36.7 and 28.3; adding stage 2 as well raised them to 38.7 and 30.3. In a separate projector comparison using LoRA fine-tuning, the Q-Former achieved 39.8 BLEU-2 and 31.7 BLEU-4, compared with 35.2 and 28.1 for a linear projector using graph and SMILES inputs. A SMILES-only model without a cross-modal projector scored 33.7 and 26.0. These results indicate that both alignment stages and the Q-Former were useful in the tested captioning setup.

  11. Knowl 11 — Captioning quality and data scale limit practical use

    limitation

    The authors state that MolCA’s molecule-captioning performance was not yet sufficient for practical application. They identify the scale of available paired data as a possible constraint: PubChem324k contains 324,000 pairs, substantially fewer than the approximately 10-million examples used in the cited vision–language pretraining comparison. Mining weakly supervised pairs from biochemical literature is suggested as a possible remedy. The study also leaves language-model capabilities such as in-context learning and chain-of-thought reasoning outside its scope, and does not address 3D molecular modeling or drug discovery. The authors caution that the model can generate inaccurate or biased text and should be tested carefully before real-world use.

Coverage note — The paper’s qualitative generation examples and detailed training-time measurements are omitted because they add less generalizable evidence than the included benchmark results and ablations.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan. 2022. Flamingo: a visual language model for few-shot learning. In NeurIPS.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In IEEvaluation@ACL, pages 65–72. Association for Computational Linguistics.
  3. 3.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. Scibert: A pretrained language model for scientific text. In EMNLP/IJCNLP (1), pages 3613–3618. Association for Computational Linguistics.
  4. 4.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, pages 3558–3568. Computer Vision Foundation / IEEE.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  6. 6.Dimitrios Christofidellis, Giorgio Giannone, Jannis Born, Ole Winther, Teodoro Laino, and Matteo Manica. 2023. Unifying molecular and textual representations via multi-task language modelling. In ICML.
  7. 7.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning. In NeurIPS.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, pages 4171–4186. Association for Computational Linguistics.
  9. 9.Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al. 2022. Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models. arXiv preprint arXiv:2203.06904.
  10. 10.Carl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. 2022. Translation between molecules and natural language. In EMNLP, pages 375–413. Association for Computational Linguistics.
  11. 11.Carl Edwards, ChengXiang Zhai, and Heng Ji. 2021. Text2mol: Cross-modal molecule retrieval with natural language queries. In EMNLP (1), pages 595–607. Association for Computational Linguistics.
  12. 12.Henri A Favre and Warren H Powell. 2013. Nomenclature of organic chemistry: IUPAC recommendations and preferred names 2013. Royal Society of Chemistry.
  13. 13.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, et al. 2023. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394.
  14. 14.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In ICLR. OpenReview.net.
  15. 15.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. Lora: Low-rank adaptation of large language models. In ICLR. OpenReview.net.
  16. 16.Weihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay Pande, and Jure Leskovec. 2020. Strategies for pre-training graph neural networks. In ICLR.
  17. 17.Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. 2019. Pubmedqa: A dataset for biomedical research question answering. In EMNLP/IJCNLP (1), pages 2567–2577. Association for Computational Linguistics.
  18. 18.Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858.
  19. 19.Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A. Shoemaker, Paul A. Thiessen, Bo Yu, Leonid Zaslavsky, Jian Zhang, and Evan Bolton. 2021. Pubchem in 2021: new data content and improved web interfaces. Nucleic Acids Res., 49(Database-Issue):D1388–D1395.
  20. 20.Thomas N Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In ICLR.
  21. 21.Greg Landrum. 2013. Rdkit documentation. Release, 1(1-79):4.
  22. 22.Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. CoRR, abs/2301.12597.
  23. 23.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In ACL/IJCNLP (1), pages 4582–4597. Association for Computational Linguistics.
  24. 24.Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan. 2022. Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. In ICLR. OpenReview.net.
  25. 25.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  26. 26.Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. 2022a. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. In NeurIPS.
  27. 27.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. arXiv preprint arXiv:2304.08485.
  28. 28.Shengchao Liu, Weili Nie, Chengpeng Wang, Jiarui Lu, Zhuoran Qiao, Ling Liu, Jian Tang, Chaowei Xiao, and Anima Anandkumar. 2022b. Multi-modal molecule structure-text model for text-based retrieval and editing. CoRR, abs/2212.10789.
  29. 29.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In ICLR (Poster). OpenReview.net.
  30. 30.Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, and Sayak Paul. 2022. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft.
  31. 31.Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick. 2023. Linearly mapping from image to text space. In ICLR.
  32. 32.OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  33. 33.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In ACL, pages 311–318. ACL.
  34. 34.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In ICML, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR.
  35. 35.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  36. 36.Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying Wei, Wenbing Huang, and Junzhou Huang. 2020. Self-supervised graph transformer on large-scale molecular data. In NeurIPS.
  37. 37.Philipp Seidl, Andreu Vall, Sepp Hochreiter, and Günter Klambauer. 2023. Enhancing activity prediction models in drug discovery with the ability to understand human language. arXiv preprint arXiv:2303.03363.
  38. 38.Teague Sterling and John J. Irwin. 2015. ZINC 15 - ligand discovery for everyone. J. Chem. Inf. Model., 55(11):2324–2337.
  39. 39.Bing Su, Dazhao Du, Zhao Yang, Yujie Zhou, Jiangmeng Li, Anyi Rao, Hao Sun, Zhiwu Lu, and Ji-Rong Wen. 2022. A molecular multimodal foundation model associating molecule graphs with natural language. CoRR, abs/2209.05481.
  40. 40.Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. Galactica: A large language model for science. CoRR, abs/2211.09085.
  41. 41.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  42. 42.Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. 2021. Multimodal few-shot learning with frozen language models. In NeurIPS, pages 200–212.
  43. 43.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NIPS, pages 5998–6008.
  44. 44.Alain C Vaucher, Philippe Schwaller, Joppe Geluykens, Vishnu H Nair, Anna Iuliano, and Teodoro Laino. 2021. Inferring experimental procedures from text-based representations of chemical reactions. Nature communications, 12(1):2573.
  45. 45.David Weininger. 1988. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci., 28(1):31–36.
  46. 46.Alexander Frank Wells. 2012. Structural inorganic chemistry. Oxford university press.
  47. 47.Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. 2018. Moleculenet: a benchmark for molecular machine learning. Chemical science, 9(2):513–530.
  48. 48.Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2019. How powerful are graph neural networks? In ICLR.
  49. 49.Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2022. FILIP: fine-grained interactive language-image pre-training. In ICLR. OpenReview.net.
  50. 50.Yuning You, Tianlong Chen, Yongduo Sui, Ting Chen, Zhangyang Wang, and Yang Shen. 2020. Graph contrastive learning with augmentations. In NeurIPS.
  51. 51.Zheni Zeng, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2022. A deep-learning system bridging molecule structure and biomedical text with comprehension comparable to human professionals. Nature communications, 13(1):862.
  52. 52.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068.
  53. 53.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models. CoRR, abs/2303.18223.
  54. 54.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.

Citation

MLA
Liu, Z., et al. “MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 15623–38, https://doi.org/10.18653/v1/2023.emnlp-main.966.
APA
Liu, Z., Li, S., Luo, Y., Fei, H., Cao, Y., Kawaguchi, K., Wang, X., & Chua, T.-S. (2023). MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 15623–15638. https://doi.org/10.18653/v1/2023.emnlp-main.966
Chicago
Liu, Z., S. Li, Y. Luo, et al. 2023. “MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 15623–38. https://doi.org/10.18653/v1/2023.emnlp-main.966.
Harvard
Liu, Z. et al. (2023) “MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 15623–15638. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.966.
Vancouver
1. Liu Z, Li S, Luo Y, Fei H, Cao Y, Kawaguchi K, Wang X, Chua T-S (2023) MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 15623–15638

BibTeX

@inproceedings{liu-etal-2023-molca,
    title = "{M}ol{CA}: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter",
    author = "Liu, Zhiyuan  and
      Li, Sihang  and
      Luo, Yanchen  and
      Fei, Hao  and
      Cao, Yixin  and
      Kawaguchi, Kenji  and
      Wang, Xiang  and
      Chua, Tat-Seng",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.966/",
    doi = "10.18653/v1/2023.emnlp-main.966",
    pages = "15623--15638"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/