Translation between Molecules and Natural Language

Carl EdwardsTuan Manh LaiKevin RosGarrett HonkeKyunghyun ChoHeng Ji

article2022EMNLP338 citations

Presents MolT5, a self-supervised framework pretrained on unlabeled text and chemical strings to bridge language and chemistry through bidirectional translation tasks like molecule captioning and text-conditioned molecule generation.

Listen

Designing new chemical compounds has historically relied on manual, trial-and-error laboratory development that costs billions of dollars and spans over a decade. While computational chemistry and deep learning have begun assisting in drug design, existing tools primarily target narrow numerical properties rather than functional, human-understandable requirements. Creating automated systems that can understand and generate chemical structures from text has been hindered by a severe shortage of paired molecule-and-text data, as manual chemical annotation requires deep domain expertise.

The article evaluates MolT5, a self-supervised learning framework designed to translate bidirectionally between molecular structures and natural language. Specifically, the article demonstrates two novel tasks: generating descriptive captions for molecular structures and generating new molecular structures directly from plain-text descriptions.

The researchers addressed the data bottleneck by pretraining an encoder-decoder model on massive unaligned datasets: standard English web text and 100 million molecular text representations from public chemistry databases. Using a denoising objective, the model learned the underlying rules of both modalities simultaneously without requiring initial pairwise alignments. The model was subsequently finetuned on a benchmark dataset of 33,010 expert-annotated molecule-description pairs. Evaluation was conducted using standard natural language generation metrics, chemical fingerprint similarity scores, and a specialized cross-modal retrieval model to measure how accurately generated outputs matched their corresponding inputs.

The primary finding is that joint pretraining on language and molecular structures enables effective translation in both directions, substantially outperforming traditional recurrent neural networks and standard architectures trained from scratch. In molecule captioning, the largest model variant achieved statistically significant improvements across language quality and cross-modal retrieval metrics, correctly identifying complex molecular classes and functional roles. In molecule generation from text, the framework achieved an exact chemical match rate of 31.1% on test samples—representing an approximate 11% relative increase over standard language models of equivalent scale—while maintaining over 90% chemical validity. Applying specialized diverse decoding strategies further raised the syntactic validity of generated molecules to as high as 99.6%.

These findings indicate that natural language can serve as an effective interface for chemical design, allowing scientists to specify functional requirements directly in plain text. Translating between natural language and molecular structures has the potential to shorten drug discovery timelines and reduce early-stage research costs. While larger general-purpose language models exhibit surprising chemical generation capabilities, adding explicit molecular pretraining provides critical improvements in generating valid, structurally accurate compounds.

Organizations should treat these findings as proof of concept for text-guided chemical design. Decision-makers should consider supporting pilot workflows where natural language models assist domain experts in exploring molecular candidates, provided that all model-generated molecules are rigorously validated through standard laboratory synthesis and clinical testing before real-world deployment.

A primary limitation of this work is the potential for bias inherited from public web corpora, which may influence generation outputs. Additionally, linear text representations of molecules can occasionally yield syntactically invalid structures, and the current benchmark evaluation relies on single-reference descriptions. The article's empirical comparisons provide moderate-to-high confidence in the model's computational capabilities, but decision-makers must exercise caution until prospective laboratory evaluations confirm biological efficacy.

arXiv: 2204.11817blender-nlp/MolT5
Cover for Translation between Molecules and Natural Language

Abstract

We present MolT5 – a self-supervised learning framework for pretraining models on a vast amount of unlabeled natural language text and molecule strings. MolT5 allows for new, useful, and challenging analogs of traditional vision-language tasks, such as molecule captioning and text-based de novo molecule generation (altogether: translation between molecules and language), which we explore for the first time. Since MolT5 pretrains models on single-modal data, it helps overcome the chemistry domain shortcoming of data scarcity. Furthermore, we consider several metrics, including a new cross-modal embedding-based metric, to evaluate the tasks of molecule captioning and text-based molecule generation. Our results show that MolT5-based models are able to generate outputs, both molecules and captions, which in many cases are high quality1.

Table of Contents

  • 1 Introduction
  • 2 Tasks
  • 2.1 Molecule Captioning
  • 2.2 Text-Based de Novo Molecule Generation
  • 3 Evaluation Metrics
  • 3.1 Text2Mol Metric
  • 3.2 Evaluating Molecule Captioning
  • 3.3 Evaluating Text-Based de Novo Molecule Generation
  • 4 MolT5 – Multimodal Text-Molecule Representation Model
  • 5 Experiments and Results
  • 5.1 Data
  • 5.2 Baselines
  • 5.3 Pretraining Process
  • 5.4 Molecule Captioning
  • 5.5 Text-Based de novo Molecule Generation
  • 5.6 Probing the Model
  • 6 Related Work
  • 6.1 Multimedia Representation
  • 6.2 Image Captioning and Text-Guided Image Generation
  • 6.3 Molecule Representation
  • 7 Conclusions and Future Work
  • 8 Broader Impacts
  • 8.1 Risks
  • 9 Limitations
  • Acknowledgement
  • References
  • A Baselines and Hyperparameters
  • B Reproducibility Checklist
  • C Decoding with Huggingface Model
  • D High Validity Molecule Generation
  • E Ablations
  • F More Examples
  • G Testing Model Diversity with Retrieval
  • H Statistical Significance
  • I NLP Capabilities of MolT5
  • J Model Probing Tests

Knowls

  1. Knowl 1 — MolT5 Pretraining and Fine-Tuning Architecture

    model/method

    MolT5 (Molecular T5) is a self-supervised sequence-to-sequence learning framework designed for bidirectional translation between molecular graph linearizations (SMILES strings) and natural language text.

    The model uses an encoder-decoder Transformer architecture initialized from public checkpoints of T5.1.1 across three scales: Small (60M parameters), Base (220M parameters), and Large (770M parameters).

    Pretraining is performed without cross-modal parallel data on a combined mixture of two monolingual corpora:

    1. Natural language text from the Colossal Clean Crawled Corpus (C4, ≈700 GB\approx 700\text{ GB} of English text).
    2. Molecular SMILES representations from ZINC-15 (100 million SMILES strings).

    During pretraining, the model optimizes the "replace corrupted spans" denoising objective:

    • Batches of 256 sequences are evenly divided between natural language and SMILES strings.
    • Consecutive token spans in each sequence are masked at random and substituted by sentinel tokens (e.g., [X], [Y]).
    • The decoder autoregressively predicts the corrupted token spans preceded by their sentinel identifiers.
    • Pretraining runs for 1,000,0001,000,000 steps.

    After pretraining, MolT5 is fine-tuned for 50,00050,000 steps on paired datasets (such as ChEBI-20) for either:

    • Molecule Captioning: mapping an input SMILES string to an output natural language description.
    • Text-Guided Molecule Generation: mapping an input natural language description to an output SMILES string.
  2. Knowl 2 — Tasks of Molecule Captioning and Text-Guided de Novo Molecule Generation

    definition

    Bidirectional translation between molecules and natural language encompasses two core sequence-to-sequence formulation tasks:

    1. Molecule Captioning: Given a chemical molecule represented as a Simplified Molecular Input Line Entry System (SMILES) string, the task is to generate a descriptive natural language text detailing the molecule's identity, functional groups, structural classifications, synthetic precursors, chemical/biological properties, and functional roles. To prevent models from merely memorizing named entities from the start of descriptions, molecule names in reference descriptions are normalized to the generic prefix "The molecule is [...]".

    2. Text-Guided de Novo Molecule Generation: Given a natural language prompt describing desired chemical attributes, biological functions, or structural motifs, the task is to generate a syntactically valid SMILES string representing a molecule that matches the specified natural language criteria.

  3. Knowl 3 — Cross-Modal Evaluation Metric via Text2Mol Cosine Similarity

    model/method

    Standard n-gram natural language generation metrics (e.g., BLEU, ROUGE, METEOR) and structural chemical similarity metrics struggle to evaluate cross-modal translation because molecules can be accurately described in many non-overlapping ways, and multiple distinct molecules can satisfy a text specification.

    To measure semantic alignment across modalities, a cross-modal evaluation metric is constructed using a trained Text2Mol retrieval model. A multi-layer perceptron (MLP) embeds both SMILES strings and natural language texts into a shared multimodal vector space where matching molecule-text pairs yield high positive cosine similarity and mismatched negative pairs average near zero.

    For an evaluated molecule-description candidate pair, the cross-modal similarity is calculated as: Text2Mol Score(M,T)=cos⁡(fmol(M),ftext(T))\text{Text2Mol Score}(M, T) = \cos\left(f_{\text{mol}}(M), f_{\text{text}}(T)\right) where fmolf_{\text{mol}} is the molecule embedding MLP encoder, ftextf_{\text{text}} is the natural language embedding MLP encoder, MM is the candidate or ground-truth SMILES string, and TT is the candidate or ground-truth natural language text. On the ChEBI-20 test split, the ground-truth pairs achieve an average Text2Mol score of 0.6090.609.

  4. Knowl 4 — Molecule Captioning Performance on ChEBI-20

    data/table

    Molecule captioning performance on the test split of the ChEBI-20 dataset (3,301 test pairs). Baseline models include a 4-layer bidirectional GRU RNN and a 6-layer vanilla Transformer trained from scratch, compared against T5.1.1 and MolT5 across Small, Base, and Large sizes. ROUGE scores represent F1F_1 measures.

    Model BLEU-2 BLEU-4 ROUGE-1 ROUGE-2 ROUGE-L METEOR Text2Mol
    Ground Truth - - - - - - 0.609
    RNN (4-layer GRU) 0.251 0.176 0.450 0.278 0.394 0.363 0.426
    Transformer (6-layer) 0.061 0.027 0.204 0.087 0.186 0.114 0.057
    T5-Small 0.501 0.415 0.602 0.446 0.545 0.532 0.526
    MolT5-Small 0.519 0.436 0.620 0.469 0.563 0.551 0.540
    T5-Base 0.511 0.423 0.607 0.451 0.550 0.539 0.523
    MolT5-Base 0.540 0.457 0.634 0.485 0.578 0.569 0.547
    T5-Large 0.558 0.467 0.630 0.478 0.569 0.586 0.563
    MolT5-Large 0.594 0.508 0.654 0.510 0.594 0.614 0.582

    Pretrained models substantially outperform models trained from scratch. The vanilla Transformer overfits to repeating generic property words regardless of relevance, leading to a very low Text2Mol score (0.0570.057). MolT5 consistently surpasses standard T5 across all metrics for each model size. Statistical paired t-tests between T5-Large and MolT5-Large confirm significant improvement (e.g., p=1.053×10−29p = 1.053 \times 10^{-29} for Text2Mol and p=1.53×10−22p = 1.53 \times 10^{-22} for ROUGE-1).

  5. Knowl 5 — Text-Guided Molecule Generation Performance on ChEBI-20

    data/table

    Performance of models on text-guided de novo molecule generation evaluated on the ChEBI-20 test split. BLEU, Exact match percentage, Levenshtein distance, and Validity are calculated on all model outputs. Fingerprint Tanimoto Similarity (MACCS FTS, RDK FTS, Morgan FTS), Fréchet ChemNet Distance (FCD), and Text2Mol scores are calculated only on syntactically valid generated SMILES strings.

    Model BLEU↑\uparrow Exact↑\uparrow Levenshtein↓\downarrow MACCS FTS↑\uparrow RDK FTS↑\uparrow Morgan FTS↑\uparrow FCD↓\downarrow Text2Mol↑\uparrow Validity↑\uparrow
    Ground Truth 1.000 1.000 0.0 1.000 1.000 1.000 0.0 0.609 1.000
    RNN 0.652 0.005 38.090 0.591 0.400 0.362 4.55 0.409 0.542
    Transformer 0.499 0.000 57.660 0.480 0.320 0.217 11.32 0.277 0.906
    T5-Small 0.741 0.064 27.703 0.704 0.578 0.525 2.89 0.479 0.608
    MolT5-Small 0.755 0.079 25.988 0.703 0.568 0.517 2.49 0.482 0.721
    T5-Base 0.762 0.069 24.950 0.731 0.605 0.545 2.48 0.499 0.660
    MolT5-Base 0.769 0.081 24.458 0.721 0.588 0.529 2.18 0.496 0.772
    T5-Large 0.854 0.279 16.721 0.823 0.731 0.670 1.22 0.552 0.902
    MolT5-Large 0.854 0.311 16.071 0.834 0.746 0.684 1.20 0.554 0.905

    Standard T5 models pretrained purely on English text demonstrate strong transfer capabilities to SMILES generation. Pretraining MolT5 on ZINC SMILES delivers substantial gains in syntactic validity (e.g., MolT5-Small validity is 72.1% vs 60.8% for T5-Small; MolT5-Base is 77.2% vs 66.0%) and achieves the highest exact match percentage (31.1% for MolT5-Large) and lowest Fréchet ChemNet Distance (1.20).

  6. Knowl 6 — High-Validity Decoding Strategy for Molecule Generation

    algorithm

    Because autoregressive sequence models generating SMILES strings can output chemically invalid syntax, a high-validity (HV) decoding procedure uses diverse beam search coupled with chemical validation filtering to increase output validity.

    Input: Natural language query TT, model MM, initial beam width B=30B = 30, diversity penalty λ=0.5\lambda = 0.5
    Output: A chemically valid SMILES string S∗S^*
    function HighValidityDecode(TT, MM, BB, λ\lambda):
        while B>0B > 0 do:
            try:
                Beams = DiverseBeamSearch(MM, input=TT, beam_width=BB, beam_groups=BB, diversity_penalty=λ\lambda)
                for each candidate sequence S∈BeamsS \in \text{Beams} (ordered by beam rank) do:
                    if RDKit.MolFromSmiles(SS) is not None then:
                        return SS
                return Beams[0]
            except MemoryError:
                B←B−5B \leftarrow B - 5
        return ""

    Applying this high-validity selection mechanism increases the proportion of valid generated molecules:

    • MolT5-Small-HV reaches 98.3%98.3\% validity (compared to 72.5%72.5\% under standard decoding).
    • MolT5-Base-HV reaches 97.9%97.9\% validity (compared to 78.7%78.7\% under standard decoding).
    • MolT5-Large-HV reaches 99.6%99.6\% validity (compared to 90.5%90.5\% under standard decoding).
  7. Knowl 7 — Pretraining Modality Ablation on MolT5-Small

    empirical result

    Ablation experiments comparing pretraining on text alone (C4-only), SMILES strings alone (ZINC-only), and both modalities jointly (C4+ZINC) on MolT5-Small evaluated on ChEBI-20 demonstrate the value of multimodal pretraining:

    1. Molecule Captioning:

      • C4-only: BLEU-4 = 0.4330.433, ROUGE-L = 0.5710.571, METEOR = 0.5450.545, Text2Mol = 0.5300.530.
      • ZINC-only: BLEU-4 = 0.4340.434, ROUGE-L = 0.5730.573, METEOR = 0.5480.548, Text2Mol = 0.5380.538.
      • C4+ZINC: BLEU-4 = 0.4450.445, ROUGE-L = 0.5830.583, METEOR = 0.5570.557, Text2Mol = 0.5430.543. Pretraining on both corpora outperforms single-corpus pretraining across all natural language metrics.
    2. Molecule Generation (Normalized by Validity): When valid-only metrics are normalized by multiplying higher-is-better scores by validity and dividing lower-is-better scores (FCD) by validity:

      • C4-only: Normalized Morgan FTS = 0.40700.4070, Normalized FCD = 4.714.71, Normalized Text2Mol = 0.35240.3524, Validity = 63.5%63.5\%.
      • ZINC-only: Normalized Morgan FTS = 0.42290.4229, Normalized FCD = 3.413.41, Normalized Text2Mol = 0.37360.3736, Validity = 80.7%80.7\%.
      • C4+ZINC: Normalized Morgan FTS = 0.43570.4357, Normalized FCD = 3.593.59, Normalized Text2Mol = 0.38790.3879, Validity = 72.5%72.5\%.

      Pretraining on ZINC alone yields the highest raw validity (80.7%80.7\%), but joint C4+ZINC pretraining yields the highest overall normalized Morgan FTS and Text2Mol semantic similarity to ground truth.

  8. Knowl 8 — Evaluating Output Diversity via Cross-Modal Retrieval

    model/method

    To evaluate whether generative models produce specific, diverse outputs rather than collapsing to generic mode predictions, a cross-modal retrieval assessment is applied using Text2Mol across the generated test dataset:

    1. Molecule Generation Diversity: The full set of molecules generated for test descriptions acts as a retrieval corpus. The input natural language descriptions are used as queries to retrieve their corresponding generated molecules via Text2Mol cosine ranking.

      • MolT5-Large achieves Mean Rank = 87.487.4, MRR = 0.5700.570, Hits@1 = 44.6%44.6\%, Hits@10 = 80.1%80.1\%, and Hits@100 = 91.0%91.0\%.
      • The vanilla Transformer baseline collapses, achieving Mean Rank = 426.4426.4, MRR = 0.1060.106, and Hits@1 = 5.62%5.62\%.
    2. Molecule Captioning Diversity: The full set of generated descriptions acts as the corpus, queried by the input molecules' SMILES representations.

      • MolT5-Large achieves Mean Rank = 16.116.1, MRR = 0.5580.558, Hits@1 = 40.4%40.4\%, Hits@10 = 84.2%84.2\%, and Hits@100 = 96.8%96.8\%.
      • Ground Truth reference text achieves Mean Rank = 5.65.6, MRR = 0.7030.703, Hits@1 = 56.4%56.4\%, Hits@10 = 94.3%94.3\%.
      • The vanilla Transformer baseline fails to produce retrievable outputs, achieving Mean Rank = 17501750, MRR = 0.0070.007, and Hits@1 = 0.4%0.4\%.

    These retrieval metrics confirm that MolT5 produces fine-grained outputs that uniquely reflect specific input prompts rather than invariant generic answers.

  9. Knowl 9 — Limitations of SMILES vs. SELFIES String Representations in MolT5

    limitation

    MolT5 relies on SMILES strings as the linear molecular representation, which lack syntactic validity guarantees when produced by autoregressive decoders, leading to invalid molecules that fail parsing in cheminformatics packages like RDKit.

    While robust alternative representations such as Self-Referencing Embedded Strings (SELFIES) guarantee 100%100\% chemical validity, SELFIES performed poorly when fine-tuning from public pretrained T5 checkpoints (which are critical for sample efficiency and language transfer). In addition, several molecular compounds present in the ChEBI-20 benchmark dataset encounter parsing and validity failures within standard SELFIES implementations.

Coverage note — Deliberately omitted qualitative probing output figures (Figures 10–61 in Appendix J) and the SST-2 GLUE benchmark accuracy verification (MolT5 95.6% vs T5 95.2% in Appendix I), as these serve as qualitative visual examples and standard sanity checks rather than core methodology or primary benchmark contributions.

References

  1. 1.Viraj Bagal, Rishal Aggarwal, PK Vinod, and U Deva Priyakumar. 2021. Molgpt: Molecular generation using a transformer-decoder model. Journal of Chemical Information and Modeling.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72.
  3. 3.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. Scibert: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3615–3620.
  4. 4.Daniel Campos and Heng Ji. 2021. Img2smi: Translating molecular structure images to simplified molecular-input line-entry system. arXiv preprint arXiv:2109.04202.
  5. 5.Adrià Cereto-Massagué, María José Ojeda, Cristina Valls, Miquel Mulero, Santiago Garcia-Vallvé, and Gerard Pujadas. 2015. Molecular fingerprint similarity search in virtual screening. Methods, 71:58–63.
  6. 6.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015. Microsoft coco captions: Data collection and evaluation server. ArXiv, abs/1504.00325.
  7. 7.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. In Computer Vision – ECCV 2020, pages 104–120, Cham. Springer International Publishing.
  8. 8.Seyone Chithrananda, Gabe Grand, and Bharath Ramsundar. 2020. Chemberta: Large-scale self-supervised pretraining for molecular property prediction. arXiv preprint arXiv:2010.09885.
  9. 9.Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1724–1734, Doha, Qatar. Association for Computational Linguistics.
  10. 10.Dina Demner-Fushman, Marc D Kohli, Marc B Rosenman, Sonya E Shooshan, Laritza Rodriguez, Sameer Antani, George R Thoma, and Clement J McDonald. 2016. Preparing a collection of radiology examinations for distribution and retrieval. Journal of the American Medical Informatics Association, 23(2):304–310.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  12. 12.Joseph L Durant, Burton A Leland, Douglas R Henry, and James G Nourse. 2002. Reoptimization of mdl keys for use in drug discovery. Journal of chemical information and computer sciences, 42(6):1273–1280.
  13. 13.David K Duvenaud, Dougal Maclaurin, Jorge Iparraguirre, Rafael Bombarell, Timothy Hirzel, Alán Aspuru-Guzik, and Ryan P Adams. 2015. Convolutional networks on graphs for learning molecular fingerprints. Advances in neural information processing systems, 28.
  14. 14.Carl Edwards, ChengXiang Zhai, and Heng Ji. 2021. Text2mol: Cross-modal molecule retrieval with natural language queries. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 595–607.
  15. 15.Benedek Fabian, Thomas Edlich, Héléna Gaspar, Marwin Segler, Joshua Meyers, Marco Fiscato, and Mohamed Ahmed. 2020. Molecular representation learning with language models and domain-relevant auxiliary tasks. arXiv preprint arXiv:2011.13230.
  16. 16.Thomas Gaudelet, Ben Day, Arian R Jamasb, Jyothish Soman, Cristian Regep, Gertrude Liu, Jeremy BR Hayter, Richard Vickers, Charles Roberts, Jian Tang, et al. 2021. Utilizing graph machine learning within drug discovery and development. Briefings in bioinformatics, 22(6):bbab159.
  17. 17.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30.
  18. 18.MD Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga. 2019. A comprehensive survey of deep learning for image captioning. ACM Computing Surveys (CsUR), 51(6):1–36.
  19. 19.Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang. 2022. Scaling up vision-language pre-training for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 17980–17989.
  20. 20.Ross Irwin, Spyridon Dimitriadis, Jiazhen He, and Esben Bjerrum. 2021. Chemformer: A pre-trained transformer for computational chemistry. ChemRxiv.
  21. 21.Sabrina Jaeger, Simone Fulle, and Samo Turk. 2018. Mol2vec: unsupervised machine learning approach with chemical intuition. Journal of chemical information and modeling, 58(1):27–35.
  22. 22.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  23. 23.Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. 2021. Transformers in vision: A survey. ACM Computing Surveys (CSUR).
  24. 24.Mario Krenn, Florian Häse, AkshatKumar Nigam, Pascal Friederich, and Alan Aspuru-Guzik. 2020. Self-referencing embedded strings (selfies): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1(4):045024.
  25. 25.Greg Landrum. 2021. Rdkit: Open-source cheminformatics software.
  26. 26.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, pages 121–137. Springer.
  27. 27.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  28. 28.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  29. 29.Di Lu, Spencer Whitehead, Lifu Huang, Heng Ji, and Shih-Fu Chang. 2018. Entity-aware image caption generation. In Proc. 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP2018).
  30. 30.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, pages 13–23.
  31. 31.Jieyu Lu and Yingkai Zhang. 2022. Unified deep learning model for multitask reaction predictions with explanation. Journal of Chemical Information and Modeling.
  32. 32.Tomás Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013a. Efficient estimation of word representations in vector space. In 1st International Conference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Workshop Track Proceedings.
  33. 33.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013b. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems, pages 3111–3119.
  34. 34.Frederic P Miller, Agnes F Vandome, and John McBrewster. 2009. Levenshtein distance: Information theory, computer science, string (computer science), string metric, damerau? levenshtein distance, spell checker, hamming distance.
  35. 35.Jia-Yu Pan, Hyung-Jeong Yang, Pinar Duygulu, and Christos Faloutsos. 2004. Automatic image captioning. In 2004 IEEE International Conference on Multimedia and Expo (ICME)(IEEE Cat. No. 04TH8763), volume 3, pages 1987–1990. IEEE.
  36. 36.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  37. 37.John Pavlopoulos, Vasiliki Kougia, and Ion Androutsopoulos. 2019. A survey on biomedical image captioning. In Proceedings of the second workshop on shortcomings in vision and language, pages 26–36.
  38. 38.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  39. 39.Daniil Polykovskiy, Alexander Zhebrak, Benjamin Sanchez-Lengeling, Sergey Golovanov, Oktai Tatanov, Stanislav Belyaev, Rauf Kurbanov, Aleksey Artamonov, Vladimir Aladinskiy, Mark Veselov, et al. 2020. Molecular sets (moses): a benchmarking platform for molecular generation models. Frontiers in pharmacology, 11:1931.
  40. 40.Kristina Preuer, Philipp Renz, Thomas Unterthiner, Sepp Hochreiter, and Günter Klambauer. 2018. Fréchet chemnet distance: A metric for generative models for molecules in drug discovery. Journal of chemical information and modeling, 58 9:1736–1741.
  41. 41.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR.
  42. 42.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  43. 43.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125.
  44. 44.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR.
  45. 45.Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016. Generative adversarial text to image synthesis. In International conference on machine learning, pages 1060–1069. PMLR.
  46. 46.Ahmet Sureyya Rifaioglu, Heval Atas, Maria Jesus Martin, Rengul Cetin-Atalay, Volkan Atalay, and Tunca Dogan. 2018. Recent applications of deep learning and machine intelligence on in silico drug discovery: methods, tools and databases. Briefings in Bioinformatics, 20(5):1878–1912.
  47. 47.Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Andrew Chen, Kathleen Kenealy, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, and Andrea Gesmundo. 2022. Scaling up models and data with t5x and seqio. arXiv preprint arXiv:2203.17189.
  48. 48.David Rogers and Mathew Hahn. 2010. Extended-connectivity fingerprints. Journal of chemical information and modeling, 50(5):742–754.
  49. 49.Nadine Schneider, Roger A. Sayle, and Gregory A. Landrum. 2015. Get your atoms in order - an open-source implementation of a novel and robust molecular canonicalization algorithm. Journal of chemical information and modeling, 55 10:2111–20.
  50. 50.Philippe Schwaller, Benjamin Hoover, Jean-Louis Reymond, Hendrik Strobelt, and Teodoro Laino. 2021a. Extraction of organic chemistry grammar from unsupervised learning of chemical reactions. Science Advances, 7(15):eabe4166.
  51. 51.Philippe Schwaller, Daniel Probst, Alain C Vaucher, Vishnu H Nair, David Kreutter, Teodoro Laino, and Jean-Louis Reymond. 2021b. Mapping the space of chemical reactions using attention-based neural networks. Nature Machine Intelligence, 3(2):144–152.
  52. 52.Kihyuk Sohn. 2016. Improved deep metric learning with multi-class n-pair loss objective. In Proceedings of the 30th International Conference on Neural Information Processing Systems, pages 1857–1865.
  53. 53.Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi, Silvia Cascianelli, Giuseppe Fiameni, and Rita Cucchiara. 2021. From show to tell: A survey on image captioning. arXiv preprint arXiv:2107.06912.
  54. 54.T. Sterling and John J. Irwin. 2015a. Zinc 15 – ligand discovery for everyone. Journal of Chemical Information and Modeling, 55:2324 – 2337.
  55. 55.Teague Sterling and John J Irwin. 2015b. Zinc 15–ligand discovery for everyone. Journal of chemical information and modeling, 55(11):2324–2337.
  56. 56.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. VL-BERT: pre-training of generic visual-linguistic representations. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  57. 57.Chenkai Sun, Weijiang Li, Jinfeng Xiao, Nikolaus Nova Parulian, ChengXiang Zhai, and Heng Ji. 2021. Fine-grained chemical entity typing with multi-modal knowledge representation. In 2021 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 1984–1991. IEEE.
  58. 58.Taffee T Tanimoto. 1958. Elementary mathematical theory of classification and prediction.
  59. 59.Andreu Vall, Sepp Hochreiter, and Günter Klambauer. 2021. Bioassayclr: Prediction of biological activity for novel bioassays based on rich textual descriptions. ELLIS Machine Learning for Molecule Discovery Workshop.
  60. 60.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  61. 61.Alain C Vaucher, Philippe Schwaller, Joppe Geluykens, Vishnu H Nair, Anna Iuliano, and Teodoro Laino. 2021. Inferring experimental procedures from text-based representations of chemical reactions. Nature communications, 12(1):1–11.
  62. 62.Alain C Vaucher, Federico Zipoli, Joppe Geluykens, Vishnu H Nair, Philippe Schwaller, and Teodoro Laino. 2020. Automated extraction of chemical synthesis actions from experimental procedures. Nature communications, 11(1):1–11.
  63. 63.Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2016. Diverse beam search: Decoding diverse solutions from neural sequence models. arXiv preprint arXiv:1610.02424.
  64. 64.Oriol Vinyals, Alexander Toshev, Samy Bengio, and D. Erhan. 2015. Show and tell: A neural image caption generator. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3156–3164.
  65. 65.Hongwei Wang, Weijiang Li, Xiaomeng Jin, Kyunghyun Cho, Heng Ji, Jiawei Han, and Martin Burke. 2022. Chemical-reaction-aware molecule representation learning. In Proc. The International Conference on Learning Representations (ICLR2022).
  66. 66.David Weininger. 1988. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of chemical information and computer sciences, 28(1):31–36.
  67. 67.David Weininger, Arthur Weininger, and Joseph L Weininger. 1989. Smiles. 2. algorithm for generation of unique smiles notation. Journal of chemical information and computer sciences, 29(2):97–101.
  68. 68.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019. Huggingface’s transformers: State-of-the-art natural language processing. ArXiv, abs/1910.03771.
  69. 69.Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. 2018. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1316–1324.
  70. 70.Zheni Zeng, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2022. A deep-learning system bridging molecule structure and biomedical text with comprehension comparable to human professionals. Nature communications, 13(1):1–11.
  71. 71.Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. 2017. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In Proceedings of the IEEE international conference on computer vision, pages 5907–5915.

Citation

MLA
Edwards, C., et al. “Translation Between Molecules and Natural Language”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 375–413, https://doi.org/10.18653/v1/2022.emnlp-main.26.
APA
Edwards, C., Lai, T., Ros, K., Honke, G., Cho, K., & Ji, H. (2022). Translation between Molecules and Natural Language. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 375–413. https://doi.org/10.18653/v1/2022.emnlp-main.26
Chicago
Edwards, C., T. Lai, K. Ros, G. Honke, K. Cho, and H. Ji. 2022. “Translation Between Molecules and Natural Language”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 375–413. https://doi.org/10.18653/v1/2022.emnlp-main.26.
Harvard
Edwards, C. et al. (2022) “Translation between Molecules and Natural Language”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 375–413. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.26.
Vancouver
1. Edwards C, Lai T, Ros K, Honke G, Cho K, Ji H (2022) Translation between Molecules and Natural Language. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 375–413

BibTeX

@inproceedings{edwards-etal-2022-translation,
    title = "Translation between Molecules and Natural Language",
    author = "Edwards, Carl  and
      Lai, Tuan  and
      Ros, Kevin  and
      Honke, Garrett  and
      Cho, Kyunghyun  and
      Ji, Heng",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.26/",
    doi = "10.18653/v1/2022.emnlp-main.26",
    pages = "375--413"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/