Unifying Molecular and Textual Representations via Multi-task Language Modelling

Dimitrios ChristofidellisGiorgio GiannoneJannis BornOle WintherTeodoro LainoMatteo Manica

article2023ICML142 citations

Introduces Text+Chem T5, a unified multi-task language model that bridges natural language and chemical structures to execute both single-domain and cross-domain chemistry workflows without needing task-specific architectures or expensive single-modality pre-training.

Listen

Data-driven scientific discovery increasingly relies on generative language models to design molecules, predict chemical reactions, and automate laboratory procedures. However, existing artificial intelligence approaches typically require separate specialized models for each individual task or depend on costly pre-training across separate single-domain datasets. This separation creates a barrier between natural language text and chemical representations, complicating human-machine collaboration and preventing models from sharing knowledge across related tasks.

The article demonstrates and evaluates a unified, multi-domain, multi-task framework called Multitask Text and Chemistry T5 (Text+Chem T5). The primary objective is to enable a single model to process natural language and chemical notations simultaneously, performing text-based, chemistry-based, and cross-domain translation tasks without specialized task-specific heads or expensive single-domain pre-training.

The authors implemented their approach using a standard encoder-decoder language architecture initialized from natural language models (ranging from 60 million to 220 million parameters). They trained the model on balanced multi-task datasets comprising 11.5 million and 33.5 million samples gathered from chemical reaction records, procedural reaction literature, and chemical text-structure databases. The model was evaluated across five core benchmarks: forward reaction prediction, retrosynthesis, molecular captioning, text-conditional molecule generation, and procedural action extraction from text.

The evaluation yielded several key findings. First, Text+Chem T5 outperformed existing specialized models on cross-domain tasks. For text-to-molecule generation, the base model achieved an accuracy of 32.2%—a fourfold improvement over the specialized MolT5 baseline (8.1%)—alongside superior sequence similarity scores. Second, for molecular captioning, the system consistently scored highest across standard text evaluation metrics, achieving a BLEU-4 score of 0.542 compared to 0.457 for MolT5-base. Third, architectural ablations revealed that sharing and fine-tuning a unified encoder across all domains yielded much better results than using separate domain-specific encoders with late-stage attention merging. Fourth, the model retained competitive performance on purely chemical and textual benchmarks while scaling more effectively with model and dataset size than baseline approaches. Finally, end-to-end qualitative testing showed that Text+Chem T5 successfully navigated a multi-step discovery workflow (generating a molecule from description, planning its synthesis, and extracting laboratory actions), whereas general foundation models like ChatGPT and Galactica produced invalid or incorrect outputs.

These results demonstrate that a single, unified model can replace fragmented, task-specific pipelines in automated discovery workflows. By removing the need for costly separate pre-training runs and specialized engineering for each chemical task, organizations can substantially reduce computational expenses, shorten model development timelines, and streamline human-AI interaction in digital laboratories.

Organizations developing computational chemistry and automated synthesis platforms should consider adopting multi-task, shared-encoder architectures to consolidate their workflow infrastructure. Teams planning implementation should prioritize training on balanced, high-volume datasets, as expanded data directly improves chemical string accuracy. Further pilot testing is warranted to integrate these models with laboratory automation hardware for automated experimentation.

Decision-makers should note certain limitations: chemical representations rely on string-based notations (SMILES), which carry a minor risk of generating invalid molecular structures. In addition, the training data reflects inherent biases from historical literature, and all outputs remain intended for research purposes. Any computational predictions must undergo standard experimental and clinical validation prior to practical chemical or medical deployment.

No sufficiently relevant recommendations were found.

Cover for Unifying Molecular and Textual Representations via Multi-task Language Modelling

Abstract

The recent advances in neural language models have also been successfully applied to the field of chemistry, offering generative solutions for classical problems in molecular design and synthesis planning. These new methods have the potential to fuel a new era of data-driven automation in scientific discovery. However, specialized models are still typically required for each task, leading to the need for problem-specific fine-tuning and neglecting task interrelations. The main obstacle in this field is the lack of a unified representation between natural language and chemical representations, complicating and limiting human-machine interaction. Here, we propose the first multi-domain, multi-task language model that can solve a wide range of tasks in both the chemical and natural language domains. Our model can handle chemical and natural language concurrently, without requiring expensive pre-training on single domains or task-specific models. Interestingly, sharing weights across domains remarkably improves our model when benchmarked against state-of-the-art baselines on single-domain and cross-domain tasks. In particular, sharing information across domains and tasks gives rise to large improvements in cross-domain tasks, the magnitude of which increase with scale, as measured by more than a dozen of relevant metrics. Our work suggests that such models can robustly and efficiently accelerate discovery in physical sciences by superseding problem-specific fine-tuning and enhancing human-model interactions.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Background
  • 3 Method
  • 4 Experiments
  • 4.1 Results
  • 5 Related Work
  • 6 Limitations
  • 7 Conclusion
  • Code Availability
  • References
  • A Additional Experiments
  • B Comparison with recent Language Models and Workflow example
  • C Prompt templates
  • D Experimental Details

Knowls

  1. Knowl 1 — Text+Chem T5 shares one encoder–decoder across chemical and textual tasks

    model/method

    Text+Chem T5 is a T5 encoder–decoder model trained jointly on natural-language and chemical-string tasks. It starts from a text-pretrained T5 checkpoint and uses a shared encoder for both domains, with a shared T5 decoder and no separate task-specific output heads. Natural-language task prompts specify which mapping to perform. The model is fine-tuned jointly on the mixed task distribution rather than first undergoing expensive pretraining on a large chemistry-only corpus or receiving separate task-specific fine-tuning. The authors’ central design claim is that sharing encoder weights, tasks, and domains lets the model transfer information between language and chemistry while retaining the ability to perform tasks within either domain.

  2. Knowl 2 — The model covers five tasks across three domain-mapping categories

    definition

    Text+Chem T5 is trained and evaluated on five tasks. The two chemistry-to-chemistry tasks are forward reaction prediction, which maps reaction precursors to the main product, and one-step retrosynthesis, which maps a product to its precursors. Molecule captioning maps a molecule represented as SMILES to natural-language text (mol2text); text-conditional de novo generation maps a molecular description to SMILES (text2mol). Paragraph-to-action maps a natural-language description of a chemical procedure to a stepwise action sequence. The paper groups these as mol2mol, mol2text, text2mol, and text2text tasks, respectively; the mol2mol category contains both reaction tasks.

  3. Knowl 3 — Training balances five task datasets and uses a larger reaction-data variant

    experimental setup

    The standard training corpus combines reaction data, procedure-to-action data, and molecule-description/SMILES pairs. The reaction dataset contains 2.3 million training pairs and 10,000 each for validation and testing. The procedures dataset contains 2.16 million training examples and 270,000 each for validation and testing. The CheBI-20 molecule-description dataset contains approximately 26,000 training pairs and approximately 3,000 each for validation and testing. To balance the five tasks, the authors use 2.3 million training examples per task, repeating examples for tasks with smaller datasets, for 11.5 million training examples total. An augmented variant expands the reaction data to 6.7 million pairs and balances the tasks to 33.5 million total training examples.

    The models are initialized from T5-small or T5-base checkpoints. Small and base versions have 60 million and 220 million parameters, respectively; they use 6 and 12 layers and 8 and 12 attention heads. All variants use a maximum input length of 512 tokens, one training epoch, batch size 64, and four accumulated gradient batches. The reported learning rates are 4e-4 for standard small and 6e-4 for the other three configurations (standard base, augmented small, and augmented base).

  4. Knowl 4 — A shared model performs unevenly across specialist chemistry and procedure tasks

    empirical result

    On forward reaction prediction, Text+Chem T5 scored 0.412 (small), 0.459 (base), 0.413 (augmented small), and 0.594 (augmented base) accuracy. These scores are below the RXN specialist baseline at 0.685 and the fine-tuned T5-base score of 0.629. On one-step retrosynthesis, the corresponding Text+Chem T5 scores were 0.249, 0.478, 0.405, and 0.372; the RXN specialist scored 0.733, while fine-tuned T5-small scored 0.245. The retrosynthesis metric is roundtrip accuracy.

    For paragraph-to-action generation, BLEU scores were 0.929 (small), 0.935 (base), 0.926 (augmented small), and 0.943 (augmented base), compared with 0.953 for fine-tuned T5-small and 0.850 for the RXN specialist. Thus, the shared model is competitive on the text task and improves over the specialist RXN score there, but does not uniformly match the best task-specific results across the chemical reaction tasks.

  5. Knowl 5 — Augmented Text+Chem T5-base leads the reported text-to-SMILES comparisons

    empirical result

    For text-conditional molecule generation, the paper evaluates generated SMILES using BLEU, exact accuracy, Levenshtein distance, MACCS, RDK, and Morgan fingerprint similarity (FTS), Fréchet ChemNet Distance (FCD), and validity. In that metric order, the results are:

    • Transformer: 0.499 BLEU; 0 accuracy; 57.66 Levenshtein; 0.480 MACCS FTS; 0.320 RDK FTS; 0.217 Morgan FTS; 11.32 FCD; 0.906 validity.
    • Fine-tuned T5-small: 0.741; 0.064; 27.7; 0.704; 0.578; 0.525; 2.89; 0.608.
    • MolT5-small: 0.755; 0.079; 25.99; 0.703; 0.568; 0.517; 2.49; 0.721.
    • Text+Chem T5-small: 0.739; 0.157; 28.54; 0.859; 0.736; 0.660; 0.066; 0.776.
    • Augmented Text+Chem T5-small: 0.815; 0.191; 21.78; 0.864; 0.744; 0.672; 0.060; 0.951.
    • Fine-tuned T5-base: 0.762; 0.069; 24.95; 0.731; 0.605; 0.545; 2.48; 0.660.
    • MolT5-base: 0.769; 0.081; 24.49; 0.721; 0.588; 0.529; 0.218; 0.772.
    • Text+Chem T5-base: 0.750; 0.212; 27.39; 0.874; 0.767; 0.697; 0.061; 0.792.
    • Augmented Text+Chem T5-base: 0.853; 0.322; 16.87; 0.901; 0.816; 0.757; 0.050; 0.943.

    Augmented Text+Chem T5-base has the best reported value on each listed metric among these comparisons. Its exact-match accuracy is 0.322, while the non-augmented Text+Chem T5-base has higher BLEU than the augmented small model’s baselines only selectively; the augmented base model’s advantage is clearest across the fingerprint, FCD, validity, and exact-match measures. Increasing size from small to base improves the augmented model’s reported results on all these metrics.

  6. Knowl 6 — Text+Chem T5 improves SMILES-to-caption scores across the reported metrics

    empirical result

    For molecule captioning, models receive a SMILES string and generate a natural-language description. The metrics, in order, are BLEU-2, BLEU-4, ROUGE-1, ROUGE-2, ROUGE-L, and METEOR:

    • Transformer: 0.061, 0.027, 0.188, 0.0597, 0.165, 0.126.
    • Fine-tuned T5-small: 0.501, 0.415, 0.602, 0.446, 0.545, 0.532.
    • MolT5-small: 0.519, 0.436, 0.620, 0.469, 0.563, 0.551.
    • Text+Chem T5-small: 0.553, 0.462, 0.633, 0.481, 0.574, 0.583.
    • Augmented Text+Chem T5-small: 0.560, 0.470, 0.638, 0.488, 0.580, 0.588.
    • Fine-tuned T5-base: 0.511, 0.424, 0.607, 0.451, 0.550, 0.539.
    • MolT5-base: 0.540, 0.457, 0.634, 0.485, 0.578, 0.569.
    • Text+Chem T5-base: 0.580, 0.490, 0.647, 0.498, 0.586, 0.604.
    • Augmented Text+Chem T5-base: 0.625, 0.542, 0.682, 0.543, 0.622, 0.648.

    Augmented Text+Chem T5-base has the highest reported score for all six metrics. The standard and augmented Text+Chem T5 results also improve from small to base on every listed metric.

  7. Knowl 7 — Encoder sharing and tuning outperform separate-encoder aggregation variants

    empirical result

    An architecture ablation compared separate text and chemistry encoders with mean or cross-attention aggregation against the shared, jointly tuned encoder in Text+Chem T5. Scores are BLEU for text2mol and mol2text, in that order. Separate-encoder MDe²-CLM with mean aggregation scored 0.572 and 0.123; using cross-attention raised these scores to 0.702 and 0.274. Separate-encoder, multi-task MDMTe²-CLM with cross-attention scored 0.247 and 0.119 without encoder tuning, and 0.211 and 0.075 with encoder tuning. Shared-encoder Text+Chem T5 scored 0.750 and 0.580, while its augmented version scored 0.853 and 0.625. In this comparison, sharing and training the encoder across domains gave the strongest cross-domain results; cross-attention improved over mean aggregation in the non-multi-task separate-encoder variant but did not match the shared-encoder models.

  8. Knowl 8 — Text-only T5 performs near zero on cross-domain tasks without adaptation

    empirical result

    In the reported cross-domain comparison, unadapted T5 produces almost no useful translation: the small model scores 0.000 BLEU on text2mol and 0.004 on mol2text, and the base model scores 0.000 and 0.003. By comparison, fine-tuned T5 scores 0.762/0.501 (small) and 0.762/0.511 (base), while Text+Chem T5 scores 0.815/0.560 (small) and 0.853/0.625 (base); each pair gives text2mol and mol2text BLEU, respectively. This result shows that text-only pretraining alone did not confer the cross-domain translation ability measured in these tests.

  9. Knowl 9 — A single model completes a demonstrated Monuron discovery workflow

    empirical result

    In a qualitative case study, the authors used three consecutive Text+Chem T5 calls to move from a textual description of the herbicide Monuron to a proposed synthesis procedure: text2mol generated its SMILES, retrosynthesis generated precursors, and paragraph-to-action converted a procedure into a stepwise protocol. The generated molecule matched the target SMILES, and the generated retrosynthesis matched the target reaction; the IBM RXN forward-prediction system assigned the generated reaction confidence 1.0, compared with 0.6 for the alternative reaction suggested by ChatGPT. For the procedure-extraction step, Text+Chem T5 and ChatGPT conceptually succeeded, whereas Galactica introduced substantial invented information. This is a single illustrative workflow, not a benchmark establishing general synthesis reliability.

  10. Knowl 10 — The study flags SMILES validity, data bias, and research-use constraints

    limitation

    The authors identify inherited risks from unintended biases in training data and limitations of representing molecules as SMILES, which can produce invalid sequences. They note that a representation with validity guarantees, such as SELFIES, could address the latter issue. The presented models are intended for research purposes; molecules they generate should undergo standard clinical testing before use in medical or other applications.

Coverage note — Detailed prompt wording and secondary qualitative examples are omitted because they are implementation examples rather than distinct central contributions.

References

  1. 1.Abid, A., Abdalla, A., Abid, A., Khan, D., Alfozan, A., and Zou, J. (2019). Gradio: Hassle-free sharing and testing of ml models in the wild. arXiv preprint arXiv:1906.02569.
  2. 2.Banerjee, S. and Lavie, A. (2005). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan. Association for Computational Linguistics.
  3. 3.Born, J. and Manica, M. (2023). Regression transformer enables concurrent sequence regression and generation for molecular language modelling. Nature Machine Intelligence, 5(4):432–444.
  4. 4.Born, J., Manica, M., Cadow, J., Markert, G., Mill, N. A., Filipavicius, M., Janakarajan, N., Cardinale, A., Laino, T., and Martínez, M. R. (2021a). Data-driven molecular design for discovery and synthesis of novel ligands: a case study on sars-cov-2. Machine Learning: Science and Technology, 2(2):025024.
  5. 5.Born, J., Manica, M., Oskooei, A., Cadow, J., Markert, G., and Martínez, M. R. (2021b). Paccmannrl: De novo generation of hit-like anticancer molecules from transcriptomic data via reinforcement learning. Iscience, 24(4):102269.
  6. 6.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  7. 7.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. (2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  8. 8.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. (2022). Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  9. 9.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  10. 10.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.
  11. 11.Durant, J. L., Leland, B. A., Henry, D. R., and Nourse, J. G. (2002). Reoptimization of mdl keys for use in drug discovery. Journal of chemical information and computer sciences, 42(6):1273–1280.
  12. 12.Edwards, C., Lai, T., Ros, K., Honke, G., and Ji, H. (2022). Translation between molecules and natural language. arXiv preprint arXiv:2204.11817.
  13. 13.Edwards, C., Zhai, C., and Ji, H. (2021). Text2mol: Cross-modal molecule retrieval with natural language queries. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 595–607.
  14. 14.Falcon, W. and The PyTorch Lightning team (2019). PyTorch Lightning.
  15. 15.Fu, Y., Peng, H., Sabharwal, A., Clark, P., and Khot, T. (2022). Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720.
  16. 16.Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al. (2021). Highly accurate protein structure prediction with alphafold. Nature, 596(7873):583–589.
  17. 17.Krenn, M., Häse, F., Nigam, A., Friederich, P., and Aspuru-Guzik, A. (2020). Self-referencing embedded strings (selfies): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1(4):045024.
  18. 18.Levenshtein, V. I. (1966). Binary codes capable of correcting deletions, insertions, and reversals. Soviet physics doklady, 10(8):707–710.
  19. 19.Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al. (2022). Solving quantitative reasoning problems with language models. arXiv preprint arXiv:2206.14858.
  20. 20.Liang, P. P., Wu, C., Morency, L.-P., and Salakhutdinov, R. (2021). Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning, pages 6565–6576. PMLR.
  21. 21.Liévin, V., Hother, C. E., and Winther, O. (2022). Can large language models reason about medical questions? arXiv preprint arXiv:2207.08143.
  22. 22.Lin, C.-Y. (2004). Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  23. 23.Lu, J. and Zhang, Y. (2022). Unified deep learning model for multitask reaction predictions with explanation. Journal of Chemical Information and Modeling, 62(6):1376–1387.
  24. 24.Manica, M., Born, J., Cadow, J., Christofidellis, D., Dave, A., Clarke, D., Teukam, Y. G. N., Giannone, G., Hoffman, S. C., Buchan, M., Chenthamarakshan, V., Donovan, T., Hsu, H. H., Zipoli, F., Schilter, O., Kishimoto, A., Hamada, L., Padhi, I., Wehden, K., McHugh, L., Khrabrov, A., Das, P., Takeda, S., and Smith, J. R. (2023). Accelerating material design with the generative toolkit for scientific discovery. npj Computational Materials, 9(1):69.
  25. 25.Nextmove (2023). Nextmove software pistachio.
  26. 26.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  27. 27.O’Neill, S. (2021). Ai-driven robotic laboratories show promise. Engineering, 7(1351.10):1016.
  28. 28.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002). Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  29. 29.Preuer, K., Renz, P., Unterthiner, T., Hochreiter, S., and Klambauer, G. (2018). Fréchet chemnet distance: a metric for generative models for molecules in drug discovery. Journal of chemical information and modeling, 58(9):1736–1741.
  30. 30.Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. (2018). Improving language understanding by generative pre-training.
  31. 31.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019). Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  32. 32.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  33. 33.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125.
  34. 34.Rogers, D. and Hahn, M. (2010). Extended-connectivity fingerprints. Journal of chemical information and modeling, 50(5):742–754.
  35. 35.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al. (2022). Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487.
  36. 36.Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al. (2021). Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  37. 37.Schwaller, P., Gaudin, T., Lanyi, D., Bekas, C., and Laino, T. (2018). “found in translation”: predicting outcomes of complex organic chemistry reactions using neural sequence-to-sequence models. Chemical science, 9(28):6091–6098.
  38. 38.Schwaller, P., Laino, T., Gaudin, T., Bolgar, P., Hunter, C. A., Bekas, C., and Lee, A. A. (2019). Molecular transformer: a model for uncertainty-calibrated chemical reaction prediction. ACS central science, 5(9):1572–1583.
  39. 39.Schwaller, P., Petraglia, R., Zullo, V., Nair, V. H., Haeuselmann, R. A., Pisoni, R., Bekas, C., Iuliano, A., and Laino, T. (2020). Predicting retrosynthetic pathways using transformer-based models and a hyper-graph exploration strategy. Chemical science, 11(12):3316–3325.
  40. 40.Schwaller, P., Probst, D., Vaucher, A. C., Nair, V. H., Kreutter, D., Laino, T., and Reymond, J.-L. (2021). Mapping the space of chemical reactions using attention-based neural networks. Nature Machine Intelligence, 3(2):144–152.
  41. 41.Tanimoto, T. T. (1958). Elementary mathematical theory of classification and prediction.
  42. 42.Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R. (2022). Galactica: A large language model for science. arXiv preprint arXiv:2211.09085.
  43. 43.Toniato, A., Schwaller, P., Cardinale, A., Geluykens, J., and Laino, T. (2021). Unassisted noise reduction of chemical reaction datasets. Nature Machine Intelligence, 3(6):485–494.
  44. 44.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
  45. 45.Vaucher, A. C., Zipoli, F., Geluykens, J., Nair, V. H., Schwaller, P., and Laino, T. (2020). Automated extraction of chemical synthesis actions from experimental procedures. Nature communications, 11(1):1–11.
  46. 46.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D. (2022). Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  47. 47.Weininger, D. (1988). Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of chemical information and computer sciences, 28(1):31–36.
  48. 48.Winata, G. I., Madotto, A., Lin, Z., Liu, R., Yosinski, J., and Fung, P. (2021). Language models are few-shot multilingual learners. arXiv preprint arXiv:2109.07684.
  49. 49.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M. (2020). Transformers: State-of-the-Art Natural Language Processing. pages 38–45. Association for Computational Linguistics.

Citation

MLA
Christofidellis, D., et al. “Unifying Molecular and Textual Representations via Multi-task Language Modelling”. International Conference on Machine Learning, vol. 202, 2023, pp. 6140–57, https://proceedings.mlr.press/v202/christofidellis23a.html.
APA
Christofidellis, D., Giannone, G., Born, J., Winther, O., Laino, T., & Manica, M. (2023). Unifying Molecular and Textual Representations via Multi-task Language Modelling. International Conference on Machine Learning, 202, 6140–6157. https://proceedings.mlr.press/v202/christofidellis23a.html
Chicago
Christofidellis, D., G. Giannone, J. Born, O. Winther, T. Laino, and M. Manica. 2023. “Unifying Molecular and Textual Representations via Multi-task Language Modelling”. International Conference on Machine Learning 202: 6140–57. https://proceedings.mlr.press/v202/christofidellis23a.html.
Harvard
Christofidellis, D. et al. (2023) “Unifying Molecular and Textual Representations via Multi-task Language Modelling”, International Conference on Machine Learning. PMLR, pp. 6140–6157. Available at: https://proceedings.mlr.press/v202/christofidellis23a.html.
Vancouver
1. Christofidellis D, Giannone G, Born J, Winther O, Laino T, Manica M (2023) Unifying Molecular and Textual Representations via Multi-task Language Modelling. In: International Conference on Machine Learning. PMLR, pp 6140–6157

BibTeX

@InProceedings{pmlr-v202-christofidellis23a,
  title = 	 {Unifying Molecular and Textual Representations via Multi-task Language Modelling},
  author =       {Christofidellis, Dimitrios and Giannone, Giorgio and Born, Jannis and Winther, Ole and Laino, Teodoro and Manica, Matteo},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {6140--6157},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/christofidellis23a/christofidellis23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/christofidellis23a.html},
  abstract = 	 {The recent advances in neural language models have also been successfully applied to the field of chemistry, offering generative solutions for classical problems in molecular design and synthesis planning. These new methods have the potential to fuel a new era of data-driven automation in scientific discovery. However, specialized models are still typically required for each task, leading to the need for problem-specific fine-tuning and neglecting task interrelations. The main obstacle in this field is the lack of a unified representation between natural language and chemical representations, complicating and limiting human-machine interaction. Here, we propose the first multi-domain, multi-task language model that can solve a wide range of tasks in both the chemical and natural language domains. Our model can handle chemical and natural language concurrently, without requiring expensive pre-training on single domains or task-specific models. Interestingly, sharing weights across domains remarkably improves our model when benchmarked against state-of-the-art baselines on single-domain and cross-domain tasks. In particular, sharing information across domains and tasks gives rise to large improvements in cross-domain tasks, the magnitude of which increase with scale, as measured by more than a dozen of relevant metrics. Our work suggests that such models can robustly and efficiently accelerate discovery in physical sciences by superseding problem-specific fine-tuning and enhancing human-model interactions.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/