A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery

Yu ZhangXiusi ChenBowen JinSheng WangShuiwang JiWei WangJiawei Han

article2024EMNLP108 citations

Presents a systematic review of more than 260 scientific large language models across multiple disciplines and data modalities, categorizing their architectures, pre-training strategies, and practical applications in accelerating scientific discovery.

Listen

Scientific research across medicine, chemistry, biology, physics, and climate science increasingly relies on artificial intelligence to process vast amounts of complex data. While specialized large language models have emerged within individual disciplines, earlier reviews have analyzed these tools in isolation, focusing only on single fields or single data types like plain text. This siloed perspective obscures common architectural principles and prevents researchers from sharing proven technical strategies across domains. The article addresses this fragmentation by establishing a unified framework that synthesizes how scientific language models are built, trained, and applied across diverse scientific disciplines.

The main objective of the article is to provide a comprehensive cross-field and cross-modal evaluation of specialized large language models and to demonstrate how these systems augment modern scientific discovery. To accomplish this, the authors conducted an extensive literature review covering more than 260 scientific models ranging in scale from 100 million to over 100 billion parameters. The analysis spans multiple fields—including general science, mathematics, physics, chemistry, materials science, biology, medicine, and geoscience—while examining how non-text data such as molecular graphs, biological sequences, crystal lattices, tables, images, and climate time series are adapted for model pre-training and evaluation.

The article identifies several core findings regarding the development and utility of scientific language models. First, it demonstrates that scientific pre-training converges into three universal paradigms regardless of discipline: masked sequence modeling for structured representation learning, autoregressive next-token prediction often coupled with instruction tuning for complex reasoning and generation, and contrastive multi-modal alignment to map distinct data types into shared latent spaces. Second, model development has shifted from smaller, task-specific encoder architectures toward multi-billion-parameter generative models capable of following natural language instructions across diverse tasks. Third, when applied directly to discovery workflows, these models demonstrate autonomous problem-solving capabilities, including generating viable research hypotheses, designing CRISPR gene-editing experiments, outperforming human benchmarks in mathematical Olympiad geometry, planning chemical syntheses, and predicting global weather patterns faster than traditional numerical simulations.

These findings indicate that treating complex scientific structures—such as amino acid sequences and molecular graphs—as specialized languages significantly lowers the barrier to deploying unified artificial intelligence workflows across science. Adopting standardized model architectures reduces the cost and development time needed to build domain-specific tools, while multi-modal integration improves data utilization across multiomics, clinical records, and material databases. However, because specialized scientific tasks demand absolute precision, the tendency of language models to generate plausible but factually incorrect outputs introduces safety and reliability risks, particularly in clinical and chemical decision-making.

To address these challenges, the article recommends developing targeted, theme-focused knowledge bases to prevent models from losing rare domain insights during broad training, as well as advancing multi-modal retrieval-augmented generation systems that verify outputs against verified experimental data and chemical structures. Future efforts should also incorporate invariant learning techniques to ensure models remain reliable when tested on unseen molecular scaffolds or novel scientific concepts. Because current evaluations remain centered on mathematics and natural sciences rather than social sciences, stakeholders should maintain cautious confidence in specialized model outputs until domain experts thoroughly validate them against empirical benchmarks.

Cover for A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery

Abstract

In many scientific fields, large language models (LLMs) have revolutionized the way text and other modalities of data (e.g., molecules and proteins) are handled, achieving superior performance in various applications and augmenting the scientific discovery process. Nevertheless, previous surveys on scientific LLMs often concentrate on one or two fields or a single modality. In this paper, we aim to provide a more holistic view of the research landscape by unveiling cross-field and cross-modal connections between scientific LLMs regarding their architectures and pre-training techniques. To this end, we comprehensively survey over 260 scientific LLMs, discuss their commonalities and differences, as well as summarize pre-training datasets and evaluation tasks for each field and modality. Moreover, we investigate how LLMs have been deployed to benefit scientific discovery. Resources related to this survey are available at https://github.com/yuzhimanhua/Awesome-Scientific-Language-Models.

Table of Contents

  • 1 Introduction
  • 2 LLMs in General Science (Table A1)
  • 2.1 Language
  • 2.2 Language + Graph
  • 2.3 Applications in Scientific Discovery
  • 3 LLMs in Mathematics (Table A2)
  • 3.1 Language
  • 3.2 Language + Vision
  • 3.3 Table
  • 3.4 Applications in Scientific Discovery
  • 4 LLMs in Physics (Table A3)
  • 4.1 Language
  • 4.2 Applications in Scientific Discovery
  • 5 LLMs in Chemistry and Materials Science (Table A4)
  • 5.1 Language
  • 5.2 Language + Graph
  • 5.3 Language + Vision
  • 5.4 Molecule
  • 5.5 Applications in Scientific Discovery
  • 6 LLMs in Biology and Medicine (Table A5)
  • 6.1 Language
  • 6.2 Language + Graph
  • 6.3 Language + Vision
  • 6.4 Protein, DNA, RNA, and Multiomics
  • 6.5 Applications in Scientific Discovery
  • 7 LLMs in Geography, Geology, and Environmental Science (Table A6)
  • 7.1 Language
  • 7.2 Language + Graph
  • 7.3 Language + Vision
  • 7.4 Climate Time Series
  • 7.5 Applications in Scientific Discovery
  • 8 Challenges and Future Directions
  • Limitations
  • Acknowledgments
  • References
  • A Summary Tables of Scientific LLMs

Knowls

  1. Knowl 1 — Taxonomy of Scientific Large Language Model Pre-training Strategies

    model/method

    Pre-training strategies for scientific large language models (LLMs) across diverse disciplines (e.g., mathematics, physics, chemistry, biology, medicine, and geoscience) and modalities fall into three primary paradigms based on model architecture and data representation:

    1. Encoder Models with Masked Language Modeling (MLM): Bidirectional Transformer encoders are pre-trained by predicting masked elements within sequentially formatted scientific data. Inputs include naturally sequential data (e.g., scientific papers, FASTA protein/DNA/RNA sequences) or artificially linearized structures (e.g., simplified molecular-input line-entry system (SMILES) or SELFIES strings for molecules; sequences of venue, author, and reference metadata nodes for citation networks).

    2. (Encoder-)Decoder Models with Autoregressive Next Token Prediction and Instruction Tuning: Autoregressive decoder models and sequence-to-sequence encoder-decoder models are trained via next-token prediction on large domain corpora, followed by supervised fine-tuning on domain-specific instructions and preference optimization. Complex multi-modal data is sequentialized by flattening table rows and columns, expressing 3D crystalline systems as particle coordinates, or encoding visual/diagrammatic features into visual tokens prepended to text token sequences.

    3. Dual-Encoder Models with Cross-Modal Contrastive Learning: Dual encoders project paired multi-modal scientific data points into a shared latent space by pulling positive pairs together and pushing negative pairs apart (following Dense Passage Retrieval and CLIP frameworks). When both modalities are sequential (e.g., scientific document text pairs or text paired with amino acid sequences), two LLM encoders are employed; when preserving non-sequential topologies (e.g., molecular graphs, chest X-rays, histopathology slides, aerial satellite views), specialized graph neural network (GNN) or vision Transformer (ViT) encoders are paired with an LLM text encoder.

  2. Knowl 2 — Scientific LLM Architectures and Benchmarks for Academic Literature and Graph Modeling

    model/method

    LLMs for general science primarily target scientific literature comprehension, citation analysis, and bibliographic graph structure:

    • Pre-training Data Sources: Large-scale bibliographic repositories and corpora including AMiner, Microsoft Academic Graph (MAG), Semantic Scholar, and S2ORC, providing titles, abstracts, and full texts.
    • Language-Only Models: Early models employ BERT architectures trained with masked language modeling (MLM) and next sentence prediction (NSP) on paper text (e.g., SciBERT, ScholarBERT, AcademicRoBERTa). Autoregressive models (e.g., SciGPT2, Galactica, SciGLM, DARWIN, FORGE) scale pre-training with next-token prediction and instruction tuning on scientific question-answering and multi-step reasoning data.
    • Graph-Augmented Language Models: Scientific papers form rich heterogeneous graphs defined by citations, venues, and authorship. Graph-aware architectures integrate these signals either by linearizing metadata into context sequences for MLM (e.g., OAG-BERT), adopting contrastive dual-encoder objectives over citation graph topology (e.g., SPECTER, ASPIRE, SciNCL), or incorporating graph-specific modules such as Adapters (SPECTER 2.0), GNN-nested Transformers (SciPatton), and Mixture-of-Experts (MoE) Transformers (SciMult).
    • Evaluation Tasks: Evaluated on traditional NLP tasks (named entity recognition (NER), relation extraction (RE), question answering (QA), extreme summarization) and relational graph tasks (paper classification, link prediction, document retrieval, paper recommendation, author name disambiguation, and paper-reviewer matching) using benchmarks such as SciDocs and SciRepEval.
  3. Knowl 3 — Architectures, Sequentialization, and Benchmarks for Mathematical LLMs

    model/method

    Mathematical LLMs address formal proofs, quantitative reasoning, word problems, geometric diagrams, and tabular data:

    • Language-Centric Math Models: Trained on multiple-choice datasets (MathQA, Ape210K, Math23K), generative datasets (GSM8K, MATH, MetaMathQA), and web-scale mathematical corpora (Proof-Pile-2, OpenWebMath). While early approaches employ encoder-based MLM (GenBERT, MathBERT, MWP-BERT), modern systems employ large-scale decoder architectures (Minerva, WizardMath, MAmmoTH, MetaMath, DeepSeekMath, InternLM-Math, Rho-Math) pre-trained on mathematical text/code and tuned with synthetic step-by-step reasoning instructions.
    • Language + Vision for Geometry: Geometry problem-solving models integrate diagrammatic information with text. Methods encode diagrams using object detectors (e.g., RetinaNet in Inter-GPS) or Vision Transformers (e.g., ViT in G-LLaVA, Geoformer, SCA-GPS), concatenate projected visual tokens with textual embeddings, and feed the sequence into an LLM decoder (e.g., LLaMA-2) using sequence-to-sequence or instruction tuning objectives with auxiliary losses (masked image modeling, image reconstruction, text-image matching). Evaluated on Geometry3K, GeoQA, GEOS, and MathVista.
    • Tabular Reasoning Models: Tabular math data is linearized by flattening structured grid cells into context tokens prepended to question prompts. Encoder models (TAPAS, TaBERT, TUTA) utilize masked token prediction and cell-structure classification, whereas modern generalist table LLMs (TAPEX, ReasTAP, Table-GPT, TableLlama, TableLLM) utilize sequence-to-sequence execution or instruction tuning on benchmarks such as WikiTableQuestions, WikiSQL, and TableInstruct across table QA, column type classification, formula prediction, and table fact verification.
  4. Knowl 4 — Chemistry and Materials Science LLM Modalities and Architectures

    model/method

    LLMs in chemistry and materials science operate over chemical literature, molecular graph representations, crystal coordinates, and 1D string representations:

    • Pure Molecular Sequence Models: 1D molecular representations (SMILES and SELFIES strings) are modeled with bidirectional encoders (ChemBERTa, MoLFormer, SMILES-BERT, polyBERT) via MLM for molecular property prediction (toxicity, atomization energy) and virtual screening. Autoregressive sequence-to-sequence and decoder models (BARTSmiles, ChemGPT, T5Chem, MolGen) generate molecules, predict reaction products, and perform retrosynthetic pathway design.
    • Language + Molecular Graph Models: To bridge molecular structures with natural language descriptions (from PubChem, ChEBI-20, ZINC), two main architectures are used: (1) Dual-encoder contrastive learning frameworks (Text2Mol, MoMu, MoleculeSTM) connecting a GNN molecular graph encoder with an LLM text encoder via text-graph matching objectives; (2) Projection-based multimodal LLMs (MolCA, InstructMol, 3D-MoLM, GIMLET) that map 2D GNN or 3D coordinate representations (e.g., Uni-Mol) into virtual tokens fed into LLM backbones (e.g., LLaMA, Galactica) for cross-modal retrieval, molecule captioning, and instruction-based property prediction.
    • Language + Vision Models: Full multi-modal frameworks (e.g., GIT-Mol) unify molecular graph representations, text descriptions, and 2D molecular structure images into a shared latent space via Swin Transformer vision encoders, GNNs, and T5 language backbones.
    • Materials and Crystal Modeling: Text-based LLMs (MatSciBERT, BatteryBERT, LLM-Prop) process materials literature and database records. For crystalline solids, generative autoregressive LLMs (CrystalLLM) process particle coordinates and composition strings for automated crystal structure generation.
  5. Knowl 5 — Biomedical, Multiomics, and Clinical Multi-Modal LLMs

    model/method

    LLMs designed for biomedicine encompass clinical language, knowledge graphs, biomedical imaging, and biological macro-sequences:

    • Clinical and Biomedical Text Models: Trained on PubMed, PMC, electronic health records (MIMIC-III, MIMIC-IV), and medical KGs (UMLS). Progression spans moderate-sized encoder/decoder models (BioBERT, ClinicalBERT, BioLinkBERT, BioGPT) evaluated on the BLURB benchmark, to instruction-tuned billion-parameter clinical dialogue models (Med-PaLM 1 & 2, MedAlpaca, PMC-LLaMA, HuatuoGPT, BioMistral, MEDITRON, Me LLaMA).
    • Biomedical Vision-Language Models: Pair clinical text (radiology reports from MIMIC-CXR, CheXpert; histopathology figure-captions from ROCO, MedICaT, Quilt-1M) with visual encoders. Architectures predominantly build upon contrastive dual-encoders (ConVIRT, GLoRIA, MedCLIP, BiomedCLIP, PMC-CLIP, BioViL, Mammo-CLIP) or multimodal decoders with visual token projectors (LLaVA-Med, Med-Flamingo, RadFM, Med-PaLM M, Med-Gemini) for visual question answering, chest X-ray segmentation, and automated medical report generation.
    • Macromolecular Sequence Models (Protein, DNA, RNA): Biological sequences formatted as FASTA strings are treated as discrete linguistic tokens. For proteins, encoder models (ESM-1b, ESM-2, ProtTrans, SaProt) use MLM over amino acid sequences for structure prediction and mutation effect estimation; autoregressive decoders (ProtGPT2, ProGen, ProLLaMA) generate functional sequences. Multimodal protein-text architectures (ProtST, Prot2Text) align sequences with textual functional annotations via contrastive or sequence-to-sequence loss. For genomics and transcriptomics, nucleotide LLMs (DNABERT, DNABERT-2, Nucleotide Transformer, HyenaDNA, RNABERT, RNA-FM) predict epigenetic marks, promoter sites, and secondary structures.
    • Single-Cell Multiomics Foundation Models: Single-cell RNA sequencing (scRNA-seq) expression profiles are modeled using BERT, GPT, and Performer architectures (scBERT, scGPT, Geneformer, scFoundation, CellPLM) with linear attention complexity to perform cell-type annotation, perturbation prediction, and gene network inference.
  6. Knowl 6 — Geoscience, Spatial Data, and Climate Foundation Models

    model/method

    Geoscience and atmospheric LLMs address text, geospatial entities, satellite imagery, and spatiotemporal climate grids:

    • Geoscience Language Models: Pre-trained on environmental reports, geographic corpora, OpenStreetMap, and academic databases. Encoders (ClimateBERT, SpaBERT, MGeo, GeoLM) perform masked language modeling and entity typing for geographic entity linking and query-to-Point-of-Interest (POI) matching. Large autoregressive models (K2, OceanGPT, GeoGalactica) adapt general LLMs (e.g., LLaMA, Galactica) via instruction tuning on curated geoscience data, evaluated on benchmarks like GeoBench and OceanBench.
    • Geospatial Graph and Vision Models: Heterogeneous POI graphs and spatial knowledge bases are integrated with text using graph aggregation layers in Transformer backbones (ERNIE-GeoL) or pointer-generation networks (PK-Chat). Language-vision models (UrbanCLIP) align aerial satellite imagery with descriptive regional profiles using contrastive vision-language pre-training to predict urban development indicators.
    • Climate Time Series Foundation Models: Spatiotemporal atmospheric and ocean climate datasets (ERA5 global reanalysis grids, CMIP6 climate model ensembles) are treated as foundation model pre-training inputs using Vision Transformer (ViT) and Swin Transformer architectures (e.g., FourCastNet, Pangu-Weather, ClimaX, FengWu, W-MAE, FuXi) trained via regression or masked autoencoding objectives for medium-range global weather forecasting and climate projection downscaling.
  7. Knowl 7 — Deployment of Scientific LLMs Across Stages of the Scientific Discovery Pipeline

    empirical result

    Scientific LLMs are deployed to automate, augment, and accelerate multiple stages of the scientific discovery cycle across disciplines:

    • Hypothesis Generation and Idea Brainstorming: Systems such as ResearchAgent and SciMon query and synthesize literature corpora to generate novel scientific hypotheses and research directions. Automated reviewing frameworks (e.g., MARG, ReviewerGPT) utilize LLMs to produce structured peer-review feedback on paper drafts.
    • Theorem Proving and Mathematical Discovery: Combining LLMs with symbolic reasoning engines yields breakthrough mathematical problem-solving. AlphaGeometry couples a language model (for proposing geometric auxiliary constructions) with a symbolic deduction engine, solving 25 out of 30 classical International Mathematical Olympiad (IMO) geometry problems without human demonstrations; adding Wu's method increases performance to 27 out of 30, surpassing human gold-medalist levels. FunSearch couples an LLM with evolutionary program search, discovering new mathematical solutions to the open cap set problem in combinatorial optimization.
    • Physics and Quantum Experiment Design: Transformer LLMs predict integer scattering amplitude coefficients in Planar N=4\mathcal{N}=4 Super Yang-Mills theory, learn qubit measurement outcome distributions in Rydberg atom arrays (RydbergGPT), and automatically synthesize executable Python experiment blueprints for quantum systems.
    • Autonomous Chemistry and Catalyst Discovery: Autonomous agentic systems (ChemCrow, Coscientist) integrate LLMs with external software tools, robotic execution platforms, and atomistic neural network simulators to design, plan, and execute chemical syntheses. Optimization frameworks (ChatDrug, DrugAssist, ChemReasoner) use LLM-guided heuristic search and in-context feedback for molecular property editing and catalyst selection.
    • Biological Design and Clinical Synthesis: Multiomics and protein LLMs enable de novo generation of SARS-CoV-2 antibodies with specific antigen-binding capabilities, evolutionary fitness evaluation of viral escape variants, and automated protocol design for CRISPR-based gene editing (CRISPR-GPT).
  8. Knowl 8 — Open Challenges and Research Frontiers for Scientific Large Language Models

    limitation

    The survey identifies three fundamental open challenges in the pre-training and deployment of scientific LLMs:

    1. Tail Knowledge in Fine-Grained Themes: While LLMs capture coarse-grained scientific fields (e.g., general chemistry), pre-training on general-domain or broad corpora causes high-frequency common signals to dominate the model parameter space. Highly specialized tail concepts (e.g., specific reaction mechanisms like Suzuki coupling) risk being diluted or erased. Integrating automatically curated, in-depth thematic knowledge graphs into the generation pipeline is a critical path to preserving fine-grained domain fidelity.

    2. Out-of-Distribution (OOD) Scientific Generalization: Scientific advancement inherently generates distribution shifts between training and evaluation phases, such as newly discovered concepts, molecules with unseen chemical scaffolds, or proteins with novel peptide topologies. Pre-trained scientific LLMs struggle with 이러한 OOD shifts. Incorporating principles from invariant risk minimization into scientific LLM pre-training frameworks offers a theoretical pathway to ensure robust generalization across unseen scientific spaces.

    3. Factuality and Cross-Modal Hallucination Mitigation: Hallucinations pose acute safety risks in high-stakes scientific areas like drug development and clinical medicine. Although Retrieval-Augmented Generation (RAG) mitigates factual drift in text and knowledge graphs, scientific data is inherently multi-modal. Developing cross-modal RAG systems—capable of dynamically retrieving and conditioning text generation on 3D molecular structures, crystallography coordinates, biological sequences, and observational sensor data—is essential for trustworthy scientific foundation models.

Coverage note — None was omitted; the full taxonomy, all domain modalities (general science, math, physics, chemistry/materials, biology/medicine, geoscience/climate), applications in discovery, and future challenges from the survey paper were synthesized into standalone, self-sufficient knowls.

References

  1. 1.Hisham Abdel-Aty and Ian R Gould. 2022. Large-scale distributed training of transformers for chemical fingerprinting. Journal of Chemical Information and Modeling, 62(20):4852–4862.
  2. 2.Hadi Abdine, Michail Chatzianastasis, Costas Bouyioukos, and Michalis Vazirgiannis. 2024. Prot2text: Multimodal protein’s function generation with gnns and transformers. In AAAI’24, pages 10757–10765.
  3. 3.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  4. 4.Emre Can Acikgoz, Osman Batur ˙Ince, Rayene Bench, Arda Anıl Boz, ˙Ilker Kesen, Aykut Erdem, and Erkut Erdem. 2024. Hippocrates: An open-source framework for advancing large language models in healthcare. arXiv preprint arXiv:2404.16621.
  5. 5.Manato Akiyama and Yasubumi Sakakibara. 2022. Informative rna base embedding for rna structural alignment and clustering by deep representation learning. NAR Genomics and Bioinformatics, 4(1):lqac012.
  6. 6.Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew McDermott. 2019. Publicly available clinical bert embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop, pages 72–78.
  7. 7.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. In NAACL’19, pages 2357–2367.
  8. 8.Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu Ha, et al. 2018. Construction of the literature graph in semantic scholar. In NAACL’18, pages 84–91.
  9. 9.Luis M Antunes, Keith T Butler, and Ricardo Grau-Crespo. 2023. Crystal structure generation with autoregressive large language modeling. arXiv preprint arXiv:2307.04340.
  10. 10.Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. 2019. Invariant risk minimization. arXiv preprint arXiv:1907.02893.
  11. 11.Sören Arlt, Haonan Duan, Felix Li, Sang Michael Xie, Yuhuai Wu, and Mario Krenn. 2024. Meta-designing quantum experiments with language models. arXiv preprint arXiv:2406.02470.
  12. 12.Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q Jiang, Jia Deng, Stella Biderman, and Sean Welleck. 2024. Llemma: An open language model for mathematics. In ICLR’24.
  13. 13.Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang. 2024. Researchagent: Iterative research idea generation over scientific literature with large language models. arXiv preprint arXiv:2404.07738.
  14. 14.Viraj Bagal, Rishal Aggarwal, PK Vinod, and U Deva Priyakumar. 2022. Molgpt: molecular generation using a transformer-decoder model. Journal of Chemical Information and Modeling, 62(9):2064–2076.
  15. 15.Fan Bai, Yuxin Du, Tiejun Huang, Max Q-H Meng, and Bo Zhao. 2024. M3d: Advancing 3d medical image analysis with multi-modal large language models. arXiv preprint arXiv:2404.00578.
  16. 16.Amos Bairoch and Rolf Apweiler. 2000. The swiss-prot protein sequence database and its supplement trembl in 2000. Nucleic Acids Research, 28(1):45–48.
  17. 17.Shruthi Bannur, Stephanie Hyland, Qianchu Liu, Fernando Perez-Garcia, Maximilian Ilse, Daniel C Castro, Benedikt Boecking, Harshita Sharma, Kenza Bouzid, Anja Thieme, et al. 2023. Learning to exploit temporal structure for biomedical vision–language processing. In CVPR’23, pages 15016–15027.
  18. 18.Zhijie Bao, Wei Chen, Shengze Xiao, Kuang Ren, Jiaao Wu, Cheng Zhong, Jiajie Peng, Xuanjing Huang, and Zhongyu Wei. 2023. Disc-medllm: Bridging general large language models and real-world medical consultation. arXiv preprint arXiv:2308.14346.
  19. 19.Marco Basaldella, Fangyu Liu, Ehsan Shareghi, and Nigel Collier. 2020. Cometa: A corpus for medical entity linking in the social media. In EMNLP’20, pages 3122–3137.
  20. 20.Jeff Beck and Ed Sequeira. 2003. Pubmed central (pmc): An archive for literature from life sciences journals. The NCBI Handbook.
  21. 21.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. Scibert: A pretrained language model for scientific text. In EMNLP’19, pages 3615–3620.
  22. 22.Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. 2023a. Accurate medium-range global weather forecasting with 3d neural networks. Nature, 619(7970):533–538.
  23. 23.Zhen Bi, Ningyu Zhang, Yida Xue, Yixin Ou, Daxiong Ji, Guozhou Zheng, and Huajun Chen. 2023b. Oceangpt: A large language model for ocean science tasks. In ACL’24, pages 3357–3372.
  24. 24.Olivier Bodenreider. 2004. The unified medical language system (umls): integrating biomedical terminology. Nucleic Acids Research, 32(suppl_1):D267–D270.
  25. 25.Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C Castro, Anton Schwaighofer, Stephanie Hyland, Maria Wetscherek, Tristan Naumann, Aditya Nori, Javier Alvarez-Valle, et al. 2022. Making the most of text semantics to improve biomedical vision–language processing. In ECCV’22, pages 1–21.
  26. 26.Daniil A Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. 2023. Autonomous chemical research with large language models. Nature, 624(7992):570–578.
  27. 27.Elliot Bolton, Abhinav Venigalla, Michihiro Yasunaga, David Hall, Betty Xiong, Tony Lee, Roxana Daneshjou, Jonathan Frankle, Percy Liang, Michael Carbin, et al. 2024. Biomedlm: A 2.7 b parameter language model trained on biomedical text. arXiv preprint arXiv:2403.18421.
  28. 28.Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016. A full-text learning to rank dataset for medical information retrieval. In ECIR’16, pages 716–722.
  29. 29.Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. 2024. Augmenting large language models with chemistry tools. Nature Machine Intelligence, 6(5):525–535.
  30. 30.Nadav Brandes, Dan Ofer, Yam Peleg, Nadav Rappoport, and Michal Linial. 2022. Proteinbert: a universal deep-learning model of protein sequence and function. Bioinformatics, 38(8):2102–2110.
  31. 31.Keno K Bressem, Lisa C Adams, Robert A Gaudin, Daniel Tröltzsch, Bernd Hamm, Marcus R Makowski, Chan-Yong Schüle, Janis L Vahldiek, and Stefan M Niehues. 2020. Highly accurate classification of chest radiographic reports using a deep learning natural language model pre-trained on 3.8 million text reports. Bioinformatics, 36(21):5255–5261.
  32. 32.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In NeurIPS’20, pages 1877–1901.
  33. 33.Aurelia Bustos, Antonio Pertusa, Jose-Maria Salinas, and Maria De La Iglesia-Vaya. 2020. Padchest: A large chest x-ray image dataset with multi-label annotated reports. Medical Image Analysis, 66:101797.
  34. 34.Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel S Weld. 2020. Tldr: Extreme summarization of scientific documents. In Findings of EMNLP’20, pages 4766–4777.
  35. 35.Tianji Cai, Garrett W Merz, François Charton, Niklas Nolte, Matthias Wilhelm, Kyle Cranmer, and Lance J Dixon. 2024. Transforming the bootstrap: Using transformers to compute scattering amplitudes in planar n= 4 super yang-mills theory. Machine Learning: Science and Technology, 5(3):035073.
  36. 36.He Cao, Zijing Liu, Xingyu Lu, Yuan Yao, and Yu Li. 2023. Instructmol: Multi-modal integration for building a versatile and reliable molecular assistant in drug discovery. arXiv preprint arXiv:2311.16208.
  37. 37.Souradip Chakraborty, Ekaba Bisong, Shweta Bhatt, Thomas Wagner, Riley Elliott, and Francesco Mosconi. 2020. Biomedbert: A pre-trained biomedical language model for qa and ir. In COLING’20, pages 669–679.
  38. 38.Jinho Chang and Jong Chul Ye. 2024. Bidirectional generation of structure and properties through a single molecular foundation model. Nature Communications, 15(1):2323.
  39. 39.Jiaqi Chen, Tong Li, Jinghui Qin, Pan Lu, Liang Lin, Chongyu Chen, and Xiaodan Liang. 2022a. Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression. In EMNLP’22, pages 3313–3323.
  40. 40.Jiaqi Chen, Jianheng Tang, Jinghui Qin, Xiaodan Liang, Lingbo Liu, Eric Xing, and Liang Lin. 2021. Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning. In Findings of ACL’21, pages 513–523.
  41. 41.Jiayang Chen, Zhihang Hu, Siqi Sun, Qingxiong Tan, Yixuan Wang, Qinze Yu, Licheng Zong, Liang Hong, Jin Xiao, Tao Shen, et al. 2022b. Interpretable rna foundation model from unannotated data for highly accurate rna structure and function predictions. arXiv preprint arXiv:2204.00300.
  42. 42.Junying Chen, Xidong Wang, Anningzhe Gao, Feng Jiang, Shunian Chen, Hongbo Zhang, Dingjie Song, Wenya Xie, Chuyi Kong, Jianquan Li, et al. 2023a. Huatuogpt-ii, one-stage training for medical adaption of llms. arXiv preprint arXiv:2311.09774.
  43. 43.Kang Chen, Tao Han, Junchao Gong, Lei Bai, Fenghua Ling, Jing-Jia Luo, Xi Chen, Leiming Ma, Tianning Zhang, Rui Su, et al. 2023b. Fengwu: Pushing the skillful global medium-range weather forecast beyond 10 days lead. arXiv preprint arXiv:2304.02948.
  44. 44.Ken Chen, Yue Zhou, Maolin Ding, Yu Wang, Zhixiang Ren, and Yuedong Yang. 2024. Self-supervised learning on millions of primary rna sequences from 72 vertebrates improves sequence-based rna splicing prediction. Briefings in Bioinformatics, 25(3):bbae163.
  45. 45.Lei Chen, Xiaohui Zhong, Feng Zhang, Yuan Cheng, Yinghui Xu, Yuan Qi, and Hao Li. 2023c. Fuxi: a cascade machine learning forecasting system for 15-day global weather forecast. npj Climate and Atmospheric Science, 6(1):190.
  46. 46.Yirong Chen, Zhenyu Wang, Xiaofen Xing, Zhipei Xu, Kai Fang, Junhong Wang, Sihang Li, Jieling Wu, Qi Liu, Xiangmin Xu, et al. 2023d. Bianque: Balancing the questioning and suggestion ability of health llms with multi-turn health conversations polished by chatgpt. arXiv preprint arXiv:2310.15896.
  47. 47.Zeming Chen, Alejandro Hernández Cano, Angelika Romanou, Antoine Bonnet, Kyle Matoba, Francesco Salvi, Matteo Pagliardini, Simin Fan, Andreas Köpf, Amirkeivan Mohtashami, et al. 2023e. Meditron-70b: Scaling medical pretraining for large language models. arXiv preprint arXiv:2311.16079.
  48. 48.Zhihong Chen, Yuhao Du, Jinpeng Hu, Yang Liu, Guanbin Li, Xiang Wan, and Tsung-Hui Chang. 2022c. Multi-modal masked autoencoders for medical vision-and-language pre-training. In MICCAI’22, pages 679–689.
  49. 49.Zhihong Chen, Guanbin Li, and Xiang Wan. 2022d. Align, reason and learn: Enhancing medical vision-and-language pre-training with knowledge. In ACM MM’22, pages 5152–5161.
  50. 50.Pujin Cheng, Li Lin, Junyan Lyu, Yijin Huang, Wenhan Luo, and Xiaoying Tang. 2023. Prior: Prototype representation joint learning from medical images and reports. In CVPR’23, pages 21361–21371.
  51. 51.Zhoujun Cheng, Haoyu Dong, Ran Jia, Pengfei Wu, Shi Han, Fan Cheng, and Dongmei Zhang. 2022. Fortap: Using formulas for numerical-reasoning-aware table pretraining. In ACL’22, pages 1150–1166.
  52. 52.Gayane Chilingaryan, Hovhannes Tamoyan, Ani Tevosyan, Nelly Babayan, Lusine Khondkaryan, Karen Hambardzumyan, Zaven Navoyan, Hrant Khachatrian, and Armen Aghajanyan. 2024. Bartsmiles: Generative masked language models for molecular representations. Journal of Chemical Information and Modeling, 64(15):5832–5843.
  53. 53.Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. 2020. Chemberta: large-scale self-supervised pretraining for molecular property prediction. arXiv preprint arXiv:2010.09885.
  54. 54.Shang-Ching Chou. 1988. An introduction to wu’s method for mechanical theorem proving in geometry. Journal of Automated Reasoning, 4(3):237–267.
  55. 55.Ratul Chowdhury, Nazim Bouatta, Surojit Biswas, Christina Floristean, Anant Kharkar, Koushik Roy, Charlotte Rochereau, Gustaf Ahdritz, Joanna Zhang, George M Church, et al. 2022. Single-sequence protein structure prediction using a language model and deep learning. Nature Biotechnology, 40(11):1617–1623.
  56. 56.Dimitrios Christofidellis, Giorgio Giannone, Jannis Born, Ole Winther, Teodoro Laino, and Matteo Manica. 2023. Unifying molecular and textual representations via multi-task language modelling. In ICML’23, pages 6140–6157.
  57. 57.Yanyi Chu, Dan Yu, Yupeng Li, Kaixuan Huang, Yue Shen, Le Cong, Jason Zhang, and Mengdi Wang. 2024. A 5’ utr language model for decoding untranslated regions of mrna and function predictions. Nature Machine Intelligence, 6(4):449–460.
  58. 58.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  59. 59.Arman Cohan, Waleed Ammar, Madeleine van Zuylen, and Field Cady. 2019. Structural scaffolds for citation intent classification in scientific publications. In NAACL’19, pages 3586–3596.
  60. 60.Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S Weld. 2020. Specter: Document-level representation learning using citation-informed transformers. In ACL’20, pages 2270–2282.
  61. 61.The 1000 Genomes Project Consortium. 2015. A global reference for human genetic variation. Nature, 526(7571):68–74.
  62. 62.The RNAcentral Consortium. 2019. Rnacentral: a hub of information for non-coding rna sequences. Nucleic Acids Research, 47(D1):D221–D229.
  63. 63.Haotian Cui, Chloe Wang, Hassaan Maan, Kuan Pang, Fengning Luo, Nan Duan, and Bo Wang. 2024. scgpt: toward building a foundation model for single-cell multi-omics using generative ai. Nature Methods, 21(8):1470–1480.
  64. 64.Hugo Dalla-Torre, Liam Gonzalez, Javier Mendoza-Revilla, Nicolas Lopez Carranza, Adam Henryk Grzywaczewski, Francesco Oteri, Christian Dallago, Evan Trop, Bernardo P de Almeida, Hassan Sirelkhatim, et al. 2023. The nucleotide transformer: Building and evaluating robust foundation models for human genomics. bioRxiv, pages 2023–01.
  65. 65.Mike D’Arcy, Tom Hope, Larry Birnbaum, and Doug Downey. 2024. Marg: Multi-agent review generation for scientific papers. arXiv preprint arXiv:2401.04259.
  66. 66.Cheng Deng, Yuting Jia, Hui Xu, Chong Zhang, Jingyao Tang, Luoyi Fu, Weinan Zhang, Haisong Zhang, Xinbing Wang, and Chenghu Zhou. 2021. Gakg: A multimodal geoscience academic knowledge graph. In CIKM’21, pages 4445–4454.
  67. 67.Cheng Deng, Bo Tong, Luoyi Fu, Jiaxin Ding, Dexing Cao, Xinbing Wang, and Chenghu Zhou. 2023. Pk-chat: Pointer network guided knowledge driven generative dialogue model. arXiv preprint arXiv:2304.00592.
  68. 68.Cheng Deng, Tianhang Zhang, Zhongmou He, Qiyuan Chen, Yuanyuan Shi, Yi Xu, Luoyi Fu, Weinan Zhang, Xinbing Wang, Chenghu Zhou, et al. 2024. K2: A foundation language model for geoscience knowledge understanding and utilization. In WSDM’24, pages 161–170.
  69. 69.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL’19, pages 4171–4186.
  70. 70.Ruixue Ding, Boli Chen, Pengjun Xie, Fei Huang, Xin Li, Qiang Zhang, and Yao Xu. 2023. Mgeo: Multimodal geographic language model pre-training. In SIGIR’23, pages 185–194.
  71. 71.Rezarta Islamaj Doğan, Robert Leaman, and Zhiyong Lu. 2014. Ncbi disease corpus: a resource for disease name recognition and concept normalization. Journal of Biomedical Informatics, 47:1–10.
  72. 72.Carl Edwards, Tuan Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. 2022. Translation between molecules and natural language. In EMNLP’22, pages 375–413.
  73. 73.Carl Edwards, ChengXiang Zhai, and Heng Ji. 2021. Text2mol: Cross-modal molecule retrieval with natural language queries. In EMNLP’21, pages 595–607.
  74. 74.Ahmed Elnaggar, Hazem Essam, Wafaa Salah-Eldin, Walid Moustafa, Mohamed Elkerdawy, Charlotte Rochereau, and Burkhard Rost. 2023. Ankh: Optimized protein language model unlocks general-purpose modelling. arXiv preprint arXiv:2301.06568.
  75. 75.Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, et al. 2021. Prottrans: Toward understanding the language of life through self-supervised learning. IEEE TPAMI, 44(10):7112–7127.
  76. 76.Veronika Eyring, Sandrine Bony, Gerald A Meehl, Catherine A Senior, Bjorn Stevens, Ronald J Stouffer, and Karl E Taylor. 2016. Overview of the coupled model intercomparison project phase 6 (cmip6) experimental design and organization. Geoscientific Model Development, 9(5):1937–1958.
  77. 77.Benedek Fabian, Thomas Edlich, Héléna Gaspar, Marwin Segler, Joshua Meyers, Marco Fiscato, and Mohamed Ahmed. 2020. Molecular representation learning with language models and domain-relevant auxiliary tasks. arXiv preprint arXiv:2011.13230.
  78. 78.Yin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu, Rui Huang, Zhuo Chen, Xiaohui Fan, and Huajun Chen. 2024a. Mol-instructions: A large-scale biomolecular instruction dataset for large language models. In ICLR’24.
  79. 79.Yin Fang, Ningyu Zhang, Zhuo Chen, Lingbing Guo, Xiaohui Fan, and Huajun Chen. 2024b. Domain-agnostic molecular generation with chemical feedback. In ICLR’24.
  80. 80.Noelia Ferruz and Birte Höcker. 2022. Controllable protein design with language models. Nature Machine Intelligence, 4(6):521–532.
  81. 81.Noelia Ferruz, Steffen Schmidt, and Birte Höcker. 2022. Protgpt2 is a deep unsupervised language model for protein design. Nature Communications, 13(1):4348.
  82. 82.Veniamin Fishman, Yuri Kuratov, Maxim Petrov, Aleksei Shmelev, Denis Shepelin, Nikolay Chekanov, Olga Kardymon, and Mikhail Burtsev. 2023. Gena-lm: A family of open-source foundational dna language models for long sequences. bioRxiv, pages 2023–06.
  83. 83.David Fitzek, Yi Hong Teoh, Hin Pok Fung, Gebremedhin A Dagnew, Ejaaz Merali, M Schuyler Moss, Benjamin MacLellan, and Roger G Melko. 2024. Rydberggpt. arXiv preprint arXiv:2405.21052.
  84. 84.Daniel Flam-Shepherd and Alán Aspuru-Guzik. 2023. Language models can generate molecules, materials, and protein binding sites directly in three dimensions as xyz, cif, and pdb files. arXiv preprint arXiv:2305.05708.
  85. 85.Daniel Flam-Shepherd, Kevin Zhu, and Alán Aspuru-Guzik. 2022. Language models can learn complex molecular distributions. Nature Communications, 13(1):3293.
  86. 86.Oscar Franzén, Li-Ming Gan, and Johan LM Björkegren. 2019. Panglaodb: a web server for exploration of mouse and human single-cell rna sequencing data. Database, 2019:baz046.
  87. 87.Nathan C Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi, Rafael Gomez-Bombarelli, Connor W Coley, and Vijay Gadepally. 2023. Neural scaling of deep chemical models. Nature Machine Intelligence, 5(11):1297–1305.
  88. 88.Jiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye, Wanjun Zhong, Yufei Wang, Lanqing Hong, Jianhua Han, Hang Xu, Zhenguo Li, et al. 2023. G-llava: Solving geometric problem with multi-modal large language model. arXiv preprint arXiv:2312.11370.
  89. 89.Anna Gaulton, Anne Hersey, Michał Nowotka, A Patricia Bento, Jon Chambers, David Mendez, Prudence Mutowo, Francis Atkinson, Louisa J Bellis, Elena Cibrián-Uhalte, et al. 2017. The chembl database in 2017. Nucleic Acids Research, 45(D1):D945–D954.
  90. 90.Mor Geva, Ankit Gupta, and Jonathan Berant. 2020. Injecting numerical reasoning skills into language models. In ACL’20, pages 946–958.
  91. 91.Shantanu Ghosh, Clare B Poynton, Shyam Visweswaran, and Kayhan Batmanghelich. 2024. Mammo-clip: A vision language foundation model to enhance data efficiency and robustness in mammography. arXiv preprint arXiv:2405.12255.
  92. 92.Michael Glass, Mustafa Canim, Alfio Gliozzo, Saneem Chemmengath, Vishwajeet Kumar, Rishav Chakravarti, Avirup Sil, Feifei Pan, Samarth Bharadwaj, and Nicolas Rodolfo Fauceglia. 2021. Capturing row and column semantics in transformer based question answering over tables. In NAACL’21, pages 1212–1224.
  93. 93.Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2024. Tora: A tool-integrated reasoning agent for mathematical problem solving. In ICLR’24.
  94. 94.F Grezes, S Blanco-Cuaresma, A Accomazzi, MJ Kurtz, G Shapurian, E Henneken, CS Grant, DM Thompson, R Chyla, S McDonald, et al. 2024. Building astrobert, a language model for astronomy & astrophysics. In Astronomical Society of the Pacific Conference Series, volume 535, page 119.
  95. 95.Nate Gruver, Anuroop Sriram, Andrea Madotto, Andrew Gordon Wilson, C Lawrence Zitnick, and Zachary Ulissi. 2024. Fine-tuned language models generate stable inorganic materials as text. In ICLR’24.
  96. 96.Xuemei Gu and Mario Krenn. 2024. Generation and human-expert evaluation of interesting research ideas using knowledge graphs and large language models. arXiv preprint arXiv:2405.17044.
  97. 97.Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 3(1):1–23.
  98. 98.Jiang Guo, A Santiago Ibanez-Lopez, Hanyu Gao, Victor Quach, Connor W Coley, Klavs F Jensen, and Regina Barzilay. 2022. Automated chemical reaction extraction from scientific literature. Journal of Chemical Information and Modeling, 62(9):2035–2045.
  99. 99.Taicheng Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh Chawla, Olaf Wiest, Xiangliang Zhang, et al. 2023. What can large language models do in chemistry? a comprehensive benchmark on eight tasks. In NeurIPS’23.
  100. 100.Tanishq Gupta, Mohd Zaki, NM Anoop Krishnan, and Mausam. 2022. Matscibert: A materials domain language model for text mining and information extraction. npj Computational Materials, 8(1):102.
  101. 101.Mordechai Haklay and Patrick Weber. 2008. Openstreetmap: User-generated street maps. IEEE Pervasive Computing, 7(4):12–18.
  102. 102.Tianyu Han, Lisa C Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Löser, Daniel Truhn, and Keno K Bressem. 2023. Medalpaca–an open-source collection of medical conversational ai models and training data. arXiv preprint arXiv:2304.08247.
  103. 103.Minsheng Hao, Jing Gong, Xin Zeng, Chiming Liu, Yucheng Guo, Xingyi Cheng, Taifeng Wang, Jianzhu Ma, Xuegong Zhang, and Le Song. 2024. Large-scale foundation model on single-cell transcriptomics. Nature Methods, 21(8):1481–1491.
  104. 104.Jennifer Harrow, Adam Frankish, Jose M Gonzalez, Electra Tapanari, Mark Diekhans, Felix Kokocinski, Bronwen L Aken, Daniel Barrell, Amonida Zadissa, Stephen Searle, et al. 2012. Gencode: the reference human genome annotation for the encode project. Genome Research, 22(9):1760–1774.
  105. 105.Haohuai He, Bing He, Lei Guan, Yu Zhao, Feng Jiang, Guanxing Chen, Qingge Zhu, Calvin Yu-Chian Chen, Ting Li, and Jianhua Yao. 2024a. De novo generation of sars-cov-2 antibody cdrh3 with a pre-trained generative large language model. Nature Communications, 15(1):6867.
  106. 106.Yuting He, Fuxiang Huang, Xinrui Jiang, Yuxiang Nie, Minghao Wang, Jiguang Wang, and Hao Chen. 2024b. Foundation model for advancing healthcare: Challenges, opportunities, and future directions. arXiv preprint arXiv:2404.03264.
  107. 107.Thorsten Hellert, João Montenegro, and Andrea Pollastro. 2024. Physbert: A text embedding model for physics scientific literature. arXiv preprint arXiv:2408.09574.
  108. 108.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021a. Measuring massive multitask language understanding. In ICLR’21.
  109. 109.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021b. Measuring mathematical problem solving with the math dataset. In NeurIPS’21.
  110. 110.Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, András Horányi, Joaquín Muñoz-Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, et al. 2020. The era5 global reanalysis. Quarterly Journal of the Royal Meteorological Society, 146(730):1999–2049.
  111. 111.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Mueller, Francesco Piccinno, and Julian Eisenschlos. 2020. Tapas: Weakly supervised table parsing via pre-training. In ACL’20, pages 4320–4333.
  112. 112.Brian Hie, Ellen D Zhong, Bonnie Berger, and Bryan Bryson. 2021. Learning the language of viral evolution and escape. Science, 371(6526):284–288.
  113. 113.Xanh Ho, Anh Khoa Duong Nguyen, An Tuan Dao, Junfeng Jiang, Yuki Chida, Kaito Sugimoto, Huy Quoc To, Florian Boudin, and Akiko Aizawa. 2024. A survey of pre-trained language models for processing scientific text. arXiv preprint arXiv:2401.17824.
  114. 114.Zhi Hong, Aswathy Ajith, James Pauloski, Eamon Duede, Kyle Chard, and Ian Foster. 2023. The diminishing returns of masked language models to science. In Findings of ACL’23, pages 1270–1283.
  115. 115.Tom Hope, Aida Amini, David Wadden, Madeleine van Zuylen, Sravanthi Parasa, Eric Horvitz, Daniel S Weld, Roy Schwartz, and Hannaneh Hajishirzi. 2021. Extracting a knowledge base of mechanisms from covid-19 papers. In NAACL’21, pages 4489–4503.
  116. 116.John J Horton. 2023. Large language models as simulated economic agents: What can we learn from homo silicus? arXiv preprint arXiv:2301.07543.
  117. 117.Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives. 2022. Learning inverse folding from millions of predicted structures. In ICML’22, pages 8946–8970.
  118. 118.Jizhou Huang, Haifeng Wang, Yibo Sun, Yunsheng Shi, Zhengjie Huang, An Zhuo, and Shikun Feng. 2022. Ernie-geol: A geography-and-language pre-trained model and its applications in baidu maps. In KDD’22, pages 3029–3039.
  119. 119.Kaixuan Huang, Yuanhao Qu, Henry Cousins, William A Johnson, Di Yin, Mihir Shah, Denny Zhou, Russ Altman, Mengdi Wang, and Le Cong. 2024a. Crispr-gpt: An llm agent for automated design of gene-editing experiments. arXiv preprint arXiv:2404.18021.
  120. 120.Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. 2019. Clinicalbert: Modeling clinical notes and predicting hospital readmission. arXiv preprint arXiv:1904.05342.
  121. 121.Kexin Huang, Abhishek Singh, Sitong Chen, Edward Moseley, Chih-Ying Deng, Naomi George, and Charolotta Lindvall. 2020. Clinical xlnet: Modeling sequential clinical notes and predicting prolonged mechanical ventilation. In Proceedings of the 3rd Clinical Natural Language Processing Workshop, pages 94–100.
  122. 122.Shih-Cheng Huang, Liyue Shen, Matthew P Lungren, and Serena Yeung. 2021. Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition. In ICCV’21, pages 3942–3951.
  123. 123.Shu Huang and Jacqueline M Cole. 2022. Batterybert: A pretrained language model for battery database enhancement. Journal of Chemical Information and Modeling, 62(24):6365–6377.
  124. 124.Weijian Huang, Cheng Li, Hong-Yu Zhou, Hao Yang, Jiarun Liu, Yong Liang, Hairong Zheng, Shaoting Zhang, and Shanshan Wang. 2024b. Enhancing representation in radiography-reports foundation model: A granular alignment algorithm using masked contrastive learning. Nature Communications, 15(1):7620.
  125. 125.Zhi Huang, Federico Bianchi, Mert Yuksekgonul, Thomas J Montine, and James Zou. 2023. A visual–language foundation model for pathology image analysis using medical twitter. Nature Medicine, 29(9):2307–2316.
  126. 126.Hiroshi Iida, Dung Thai, Varun Manjunatha, and Mohit Iyyer. 2021. Tabbie: Pretrained representations of tabular data. In NAACL’21, pages 3446–3456.
  127. 127.Wisdom Oluchi Ikezogwo, Mehmet Saygin Seyfioglu, Fatemeh Ghezloo, Dylan Stefan Chan Geva, Fatwir Sheikh Mohammed, Pavan Kumar Anand, Ranjay Krishna, and Linda Shapiro. 2023. Quilt-1m: One million image-text pairs for histopathology. In NeurIPS’23.
  128. 128.Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Silviana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, et al. 2019. Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. In AAAI’19, pages 590–597.
  129. 129.Ross Irwin, Spyridon Dimitriadis, Jiazhen He, and Esben Jannik Bjerrum. 2022. Chemformer: a pre-trained transformer for computational chemistry. Machine Learning: Science and Technology, 3(1):015022.
  130. 130.Kevin Maik Jablonka, Philippe Schwaller, Andres Ortega-Guerrero, and Berend Smit. 2024. Leveraging large language models for predictive chemistry. Nature Machine Intelligence, 6(2):161–169.
  131. 131.Anubhav Jain, Shyue Ping Ong, Geoffroy Hautier, Wei Chen, William Davidson Richards, Stephen Dacek, Shreyas Cholia, Dan Gunter, David Skinner, Gerbrand Ceder, et al. 2013. Commentary: The materials project: A materials genome approach to accelerating materials innovation. APL materials, 1(1).
  132. 132.Yanrong Ji, Zhihan Zhou, Han Liu, and Ramana V Davuluri. 2021. Dnabert: pre-trained bidirectional encoder representations from transformers model for dnalanguage in genome. Bioinformatics, 37(15):2112–2120.
  133. 133.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  134. 134.Zhengbao Jiang, Yi Mao, Pengcheng He, Graham Neubig, and Weizhu Chen. 2022. Omnitab: Pretraining with natural and synthetic data for few-shot tablebased question answering. In NAACL’22, pages 932–942.
  135. 135.Zhanming Jie, Jierui Li, and Wei Lu. 2022. Learning to reason deductively: Math word problem solving as complex relation extraction. In ACL’22, pages 5944–5955.
  136. 136.Bowen Jin, Gang Liu, Chi Han, Meng Jiang, Heng Ji, and Jiawei Han. 2023a. Large language models on graphs: A comprehensive survey. arXiv preprint arXiv:2312.02783.
  137. 137.Bowen Jin, Chulin Xie, Jiawei Zhang, Kashob Kumar Roy, Yu Zhang, Zheng Li, Ruirui Li, Xianfeng Tang, Suhang Wang, Yu Meng, and Jiawei Han. 2024. Graph chain-of-thought: Augmenting large language models by reasoning on graphs. In Findings of ACL’24, pages 163–184.
  138. 138.Bowen Jin, Wentao Zhang, Yu Zhang, Yu Meng, Xinyang Zhang, Qi Zhu, and Jiawei Han. 2023b. Patton: Language model pretraining on text-rich networks. In ACL’23, pages 7005–7020.
  139. 139.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421.
  140. 140.Qiao Jin, Bhuwan Dhingra, William Cohen, and Xinghua Lu. 2019. Probing biomedical embeddings from language models. In Proceedings of the 3rd Workshop on Evaluating Vector Space Representations for NLP, pages 82–89.
  141. 141.Qiao Jin, Won Kim, Qingyu Chen, Donald C Comeau, Lana Yeganova, W John Wilbur, and Zhiyong Lu. 2023c. Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics, 39(11):btad651.
  142. 142.Wengong Jin, Connor W Coley, Regina Barzilay, and Tommi Jaakkola. 2017. Predicting organic reaction outcomes with weisfeiler-lehman network. In NIPS’17, pages 2604–2613.
  143. 143.Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al. 2023. Mimic-iv, a freely accessible electronic health record dataset. Scientific Data, 10(1):1.
  144. 144.Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chihying Deng, Roger G Mark, and Steven Horng. 2019. Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data, 6(1):317.
  145. 145.Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016. Mimic-iii, a freely accessible critical care database. Scientific Data, 3(1):1–9.
  146. 146.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for opendomain question answering. In EMNLP’20, pages 6769–6781.
  147. 147.Yash Khare, Viraj Bagal, Minesh Mathew, Adithi Devi, U Deva Priyakumar, and CV Jawahar. 2021. Mmbert: Multimodal bert pretraining for improved medical vqa. In ISBI’21, pages 1033–1036.
  148. 148.Chanwoo Kim, Soham U Gadgil, Alex J DeGrave, Jesutofunmi A Omiye, Zhuo Ran Cai, Roxana Daneshjou, and Su-In Lee. 2024. Transparent medical image ai via an image–text foundation model grounded in medical literature. Nature Medicine, 30(4):1154–1165.
  149. 149.Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al. 2019. Pubchem 2019 update: improved access to chemical data. Nucleic Acids Research, 47(D1):D1102–D1109.
  150. 150.Thomas N Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In AAAI’17.
  151. 151.Mario Krenn, Florian Häse, AkshatKumar Nigam, Pascal Friederich, and Alan Aspuru-Guzik. 2020. Self-referencing embedded strings (selfies): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1(4):045024.
  152. 152.Christopher Kuenneth and Rampi Ramprasad. 2023. polybert: a chemical language model to enable fully machine-driven ultrafast polymer informatics. Nature Communications, 14(1):4099.
  153. 153.Avan Kumar, Bhavik R Bakshi, Manojkumar Ramteke, and Hariprasad Kodamana. 2023. Recycle-bert: extracting knowledge about plastic waste recycling by natural language processing. ACS Sustainable Chemistry & Engineering, 11(32):12123–12134.
  154. 154.Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024. Biomistral: A collection of open-source pretrained large language models for medical domains. In Findings of ACL’24, pages 5848–5864.
  155. 155.Dan Lahav, Jon Saad Falcon, Bailey Kuehl, Sophie Johnson, Sravanthi Parasa, Noam Shomron, Duen Horng Chau, Diyi Yang, Eric Horvitz, Daniel S Weld, et al. 2022. A search engine for discovery of scientific challenges and directions. In AAAI’22, pages 11982–11990.
  156. 156.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240.
  157. 157.Oliver Lehmberg, Dominique Ritze, Robert Meusel, and Christian Bizer. 2016. A large public corpus of web tables containing time and context metadata. In WWW’16, pages 75–76.
  158. 158.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020a. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In ACL’20, pages 7871–7880.
  159. 159.Patrick Lewis, Myle Ott, Jingfei Du, and Veselin Stoyanov. 2020b. Pretrained language models for biomedical and clinical tasks: understanding and extending the state-of-the-art. In Proceedings of the 3rd Clinical Natural Language Processing Workshop, pages 146–157.
  160. 160.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al. 2022. Solving quantitative reasoning problems with language models. In NeurIPS’22, pages 3843–3857.
  161. 161.Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2023a. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. In NeurIPS’23.
  162. 162.Fei Li, Yonghao Jin, Weisong Liu, Bhanu Pratap Singh Rawat, Pengshan Cai, Hong Yu, et al. 2019. Fine-tuning bidirectional encoder representations from transformers (bert)–based models on large-scale electronic health record notes: an empirical study. JMIR Medical Informatics, 7(3):e14830.
  163. 163.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023b. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML’23, pages 19730–19742.
  164. 164.Michael Y Li, Emily B Fox, and Noah D Goodman. 2024a. Automated statistical model discovery with language models. In ICML’24, pages 27791–27807.
  165. 165.Peng Li, Yeye He, Dror Yashar, Weiwei Cui, Song Ge, Haidong Zhang, Danielle Rifinski Fainman, Dongmei Zhang, and Surajit Chaudhuri. 2024b. Table-gpt: Table fine-tuned gpt for diverse table tasks. Proceedings of the ACM on Management of Data, 2(3):1–28.
  166. 166.Pengfei Li, Gang Liu, Jinlong He, Zixu Zhao, and Shenjun Zhong. 2023c. Masked vision and language pre-training with unimodal and multimodal contrastive losses for medical visual question answering. In MICCAI’23, pages 374–383.
  167. 167.Sihang Li, Zhiyuan Liu, Yanchen Luo, Xiang Wang, Xiangnan He, Kenji Kawaguchi, Tat-Seng Chua, and Qi Tian. 2024c. Towards 3d molecule-text interpretation in language models. In ICLR’24.
  168. 168.Sizhen Li, Saeed Moayedpour, Ruijiang Li, Michael Bailey, Saleh Riahi, Lorenzo Kogler-Anele, Milad Miladi, Jacob Miner, Dinghai Zheng, Jun Wang, et al. 2023d. Codonbert: Large language models for mrna design and optimization. bioRxiv, pages 2023–09.
  169. 169.Yikuan Li, Shishir Rao, José Roberto Ayala Solares, Abdelaali Hassaine, Rema Ramakrishnan, Dexter Canoy, Yajie Zhu, Kazem Rahimi, and Gholamreza Salimi-Khorshidi. 2020. Behrt: transformer for electronic health records. Scientific Reports, 10(1):7155.
  170. 170.Yikuan Li, Ramsey M Wehbe, Faraz S Ahmad, Hanyin Wang, and Yuan Luo. 2022a. Clinical-longformer and clinical-bigbird: Transformers for long clinical sequences. arXiv preprint arXiv:2201.11838.
  171. 171.Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang. 2023e. Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge. Cureus, 15(6).
  172. 172.Zekun Li, Jina Kim, Yao-Yi Chiang, and Muhao Chen. 2022b. Spabert: A pretrained language model from geographic data for geo-entity representation. In Findings of EMNLP’22, pages 2757–2769.
  173. 173.Zekun Li, Wenxuan Zhou, Yao-Yi Chiang, and Muhao Chen. 2023f. Geolm: Empowering language models for geospatially grounded language understanding. In EMNLP’23, pages 5227–5240.
  174. 174.Zhongli Li, Wenxuan Zhang, Chao Yan, Qingyu Zhou, Chao Li, Hongzhi Liu, and Yunbo Cao. 2022c. Seeking patterns, not just memorizing procedures: Contrastive learning for solving math word problems. In Findings of ACL’22, pages 2486–2496.
  175. 175.Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, et al. 2024a. Monitoring ai-modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews. In ICML’24.
  176. 176.Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, et al. 2024b. Mapping the increasing use of llms in scientific papers. arXiv preprint arXiv:2404.01268.
  177. 177.Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Yi Ding, Xinyu Yang, Kailas Vodrahalli, Siyu He, Daniel Scott Smith, Yian Yin, et al. 2024c. Can large language models provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI, 1(8):AIoa2400196.
  178. 178.Zhenwen Liang, Kehan Guo, Gang Liu, Taicheng Guo, Yujun Zhou, Tianyu Yang, Jiajun Jiao, Renjie Pi, Jipeng Zhang, and Xiangliang Zhang. 2024d. Scemqa: A scientific college entrance level multimodal question answering benchmark. In ACL’24, pages 109–119.
  179. 179.Zhenwen Liang, Tianyu Yang, Jipeng Zhang, and Xiangliang Zhang. 2023. Unimath: A foundational and multimodal mathematical reasoner. In EMNLP’23, pages 7126–7133.
  180. 180.Zhenwen Liang, Jipeng Zhang, Lei Wang, Wei Qin, Yunshi Lan, Jie Shao, and Xiangliang Zhang. 2022. Mwp-bert: Numeracy-augmented pre-training for math word problem solving. In Findings of NAACL’22, pages 997–1009.
  181. 181.Jiacheng Lin, Hanwen Xu, Zifeng Wang, Sheng Wang, and Jimeng Sun. 2024a. Panacea: A foundation model for clinical trial search, summarization, design, and recruitment. arXiv preprint arXiv:2407.11007.
  182. 182.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. 2017. Focal loss for dense object detection. In ICCV’17, pages 2980–2988.
  183. 183.Weixiong Lin, Ziheng Zhao, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023a. Pmc-clip: Contrastive language-image pre-training using biomedical documents. In MICCAI’23, pages 525–536.
  184. 184.Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al. 2023b. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123–1130.
  185. 185.Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, Yelong Shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, et al. 2024b. Rho-1: Not all tokens are what you need. arXiv preprint arXiv:2404.07965.
  186. 186.Zhouhan Lin, Cheng Deng, Le Zhou, Tianhang Zhang, Yi Xu, Yutong Xu, Zhongmou He, Yuanyuan Shi, Beiya Dai, Yunchong Song, et al. 2024c. Geogalactica: A scientific large language model in geoscience. arXiv preprint arXiv:2401.00434.
  187. 187.David J Lipman and William R Pearson. 1985. Rapid and sensitive protein similarity searches. Science, 227(4693):1435–1441.
  188. 188.Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. 2021a. Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In ISBI’21, pages 1650–1654.
  189. 189.Che Liu, Sibo Cheng, Chen Chen, Mengyun Qiao, Weitong Zhang, Anand Shah, Wenjia Bai, and Rossella Arcucci. 2023a. M-flag: Medical vision-language pre-training with frozen language models and latent space geometry optimization. In MICCAI’23, pages 637–647.
  190. 190.Fangyu Liu, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, and Nigel Collier. 2021b. Self-alignment pretraining for biomedical entity representations. In NAACL’21, pages 4228–4238.
  191. 191.Junling Liu, Ziming Wang, Qichen Ye, Dading Chong, Peilin Zhou, and Yining Hua. 2023b. Qilin-med-vl: Towards chinese large vision-language model for general healthcare. arXiv preprint arXiv:2310.17956.
  192. 192.Pengfei Liu, Yiming Ren, Jun Tao, and Zhixiang Ren. 2024a. Git-mol: A multi-modal large language model for molecular science with graph, image, and text. Computers in Biology and Medicine, 171:108073.
  193. 193.Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022a. Tapex: Table pre-training via learning a neural sql executor. In ICLR’22.
  194. 194.Ryan Liu and Nihar B Shah. 2023. Reviewergpt? an exploratory study on using large language models for paper reviewing. arXiv preprint arXiv:2306.00622.
  195. 195.Shengchao Liu, Yanjing Li, Zhuoxinran Li, Anthony Gitter, Yutao Zhu, Jiarui Lu, Zhao Xu, Weili Nie, Arvind Ramanathan, Chaowei Xiao, et al. 2023c. A text-guided protein design framework. arXiv preprint arXiv:2302.04611.
  196. 196.Shengchao Liu, Weili Nie, Chengpeng Wang, Jiarui Lu, Zhuoran Qiao, Ling Liu, Jian Tang, Chaowei Xiao, and Animashree Anandkumar. 2023d. Multi-modal molecule structure–text model for text-based retrieval and editing. Nature Machine Intelligence, 5(12):1447–1457.
  197. 197.Shengchao Liu, Jiongxiao Wang, Yijin Yang, Chengpeng Wang, Ling Liu, Hongyu Guo, and Chaowei Xiao. 2024b. Conversational drug editing using retrieval and domain feedback. In ICLR’24.
  198. 198.Xiao Liu, Da Yin, Jingnan Zheng, Xingjian Zhang, Peng Zhang, Hongxia Yang, Yuxiao Dong, and Jie Tang. 2022b. Oag-bert: Towards a unified backbone language model for academic knowledge services. In KDD’22, pages 3418–3428.
  199. 199.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  200. 200.Zhiyuan Liu, Sihang Li, Yanchen Luo, Hao Fei, Yixin Cao, Kenji Kawaguchi, Xiang Wang, and Tat-Seng Chua. 2023e. Molca: Molecular graph-language modeling with cross-modal projector and uni-modal adapter. In EMNLP’23, pages 15623–15638.
  201. 201.Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel S Weld. 2020. S2orc: The semantic scholar open research corpus. In ACL’20, pages 4969–4983.
  202. 202.Jieyu Lu and Yingkai Zhang. 2022. Unified deep learning model for multitask reaction predictions with explanation. Journal of Chemical Information and Modeling, 62(6):1376–1387.
  203. 203.Ming Y Lu, Bowen Chen, Andrew Zhang, Drew FK Williamson, Richard J Chen, Tong Ding, Long Phi Le, Yung-Sung Chuang, and Faisal Mahmood. 2023. Visual language pretrained multiple instance zero-shot transfer for histopathology images. In CVPR’23, pages 19764–19775.
  204. 204.Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2024. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. In ICLR’24.
  205. 205.Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-chun Zhu. 2021. Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning. In ACL’21, pages 6774–6786.
  206. 206.Zhiyong Lu. 2011. Pubmed and beyond: a survey of web tools for searching biomedical literature. Database, 2011:baq036.
  207. 207.Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018. Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction. In EMNLP’18, pages 3219–3232.
  208. 208.Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023a. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. arXiv preprint arXiv:2308.09583.
  209. 209.Ling Luo, Jinzhong Ning, Yingwen Zhao, Zhijun Wang, Zeyuan Ding, Peng Chen, Weiru Fu, Qinyu Han, Guangtao Xu, Yunzhi Qiu, et al. 2024. Taiyi: a bilingual fine-tuned large language model for diverse biomedical tasks. JAMIA, 31(9):1865–1874.
  210. 210.Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 23(6):bbac409.
  211. 211.Yizhen Luo, Kai Yang, Massimo Hong, Xingyi Liu, and Zaiqing Nie. 2023b. Molfm: A multimodal molecular foundation model. arXiv preprint arXiv:2307.09484.
  212. 212.Yizhen Luo, Jiahuan Zhang, Siqi Fan, Kai Yang, Yushuai Wu, Mu Qiao, and Zaiqing Nie. 2023c. Biomedgpt: Open multimodal generative pre-trained transformer for biomedicine. arXiv preprint arXiv:2308.09442.
  213. 213.Kelvin Luu, Xinyi Wu, Rik Koncel-Kedziorski, Kyle Lo, Isabel Cachola, and Noah A Smith. 2021. Explaining relationships between scientific documents. In ACL’21, pages 2130–2144.
  214. 214.Liuzhenghao Lv, Zongying Lin, Hao Li, Yuyang Liu, Jiaxi Cui, Calvin Yu-Chian Chen, Li Yuan, and Yonghong Tian. 2024. Prollama: A protein large language model for multi-task protein language processing. arXiv preprint arXiv:2402.16445.
  215. 215.Ali Madani, Ben Krause, Eric R Greene, Subu Subramanian, Benjamin P Mohr, James M Holton, Jose Luis Olmos, Caiming Xiong, Zachary Z Sun, Richard Socher, et al. 2023. Large language models generate functional protein sequences across diverse families. Nature Biotechnology, 41(8):1099–1106.
  216. 216.Xin Man, Chenghong Zhang, Jin Feng, Changyu Li, and Jie Shao. 2023. W-mae: Pre-trained weather model with masked autoencoder for multi-variable weather forecasting. arXiv preprint arXiv:2304.08754.
  217. 217.Łukasz Maziarka, Tomasz Danel, Sławomir Mucha, Krzysztof Rataj, Jacek Tabor, and Stanisław Jastrzębski. 2020. Molecule attention transformer. arXiv preprint arXiv:2002.08264.
  218. 218.Łukasz Maziarka, Dawid Majchrowski, Tomasz Danel, Piotr Gaiński, Jacek Tabor, Igor Podolak, Paweł Morkisz, and Stanisław Jastrzębski. 2024. Relative molecule self-attention transformer. Journal of Cheminformatics, 16(1):3.
  219. 219.Joshua Meier, Roshan Rao, Robert Verkuil, Jason Liu, Tom Sercu, and Alex Rives. 2021. Language models enable zero-shot prediction of the effects of mutations on protein function. In NeurIPS’21, pages 29287–29303.
  220. 220.Yiwen Meng, William Speier, Michael K Ong, and Corey W Arnold. 2021a. Bidirectional representation learning from transformers using multimodal electronic health record data to predict depression. IEEE Journal of Biomedical and Health Informatics, 25(8):3121–3129.
  221. 221.Zaiqiao Meng, Fangyu Liu, Thomas Clark, Ehsan Shareghi, and Nigel Collier. 2021b. Mixture-of-partitions: Infusing large biomedical knowledge graphs into bert. In EMNLP’21, pages 4672–4681.
  222. 222.Giacomo Miolo, Giulio Mantoan, and Carlotta Orsenigo. 2021. Electramed: a new pre-trained language representation model for biomedical nlp. arXiv preprint arXiv:2104.09585.
  223. 223.Adrian Mirza, Nawaf Alampara, Sreekanth Kunchapu, Benedict Emoekabu, Aswanth Krishnan, Mara Wilhelmi, Macjonathan Okereke, Juliane Eberhardt, Amir Mohammad Elahi, Maximilian Greiner, et al. 2024. Are large language models superhuman chemists? arXiv preprint arXiv:2404.01475.
  224. 224.Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, et al. 2022. Lila: A unified benchmark for mathematical reasoning. In EMNLP’22, pages 5807–5832.
  225. 225.Jong Hak Moon, Hyungyung Lee, Woncheol Shin, Young-Hak Kim, and Edward Choi. 2022. Multimodal understanding and generation for medical images and text via vision-language pre-training. IEEE JBHI, 26(12):6070–6080.
  226. 226.Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Eduardo Pontes Reis, and Pranav Rajpurkar. 2023. Med-flamingo: a multimodal medical few-shot learner. In ML4H’23, pages 353–367.
  227. 227.S Mostafa Mousavi, William L Ellsworth, Weiqiang Zhu, Lindsay Y Chuang, and Gregory C Beroza. 2020. Earthquake transformer—an attentive deep-learning model for simultaneous earthquake detection and phase picking. Nature Communications, 11(1):3952.
  228. 228.Martin Müller, Marcel Salathé, and Per E Kummervold. 2023. Covid-twitter-bert: A natural language processing model to analyse covid-19 content on twitter. Frontiers in Artificial Intelligence, 6:1023281.
  229. 229.Philip Müller, Georgios Kaissis, Congyu Zou, and Daniel Rueckert. 2022. Joint learning of localized representations from medical images and reports. In ECCV’22, pages 685–701.
  230. 230.Sheshera Mysore, Arman Cohan, and Tom Hope. 2022. Multi-vector models with textual guidance for fine-grained scientific document similarity. In NAACL’22, pages 4453–4470.
  231. 231.Usman Naseem, Adam G Dunn, Matloob Khushi, and Jinman Kim. 2022. Benchmarking for biomedical natural language processing tasks with a domain specific albert. BMC Bioinformatics, 23(1):144.
  232. 232.Eric Nguyen, Michael Poli, Marjan Faizi, Armin W Thomas, Michael Wornow, Callum Birch-Sykes, Stefano Massaroli, Aman Patel, Clayton M Rabideau, Yoshua Bengio, et al. 2023a. Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution. In NeurIPS’23.
  233. 233.Tuan Dung Nguyen, Yuan-Sen Ting, Ioana Ciuca, Charles O’Neill, Ze-Chang Sun, Maja Jabłońska, Sandor Kruk, Ernest Perkowski, Jack Miller, Jason Jason Jingsh Li, et al. 2023b. Astrollama: Towards specialized foundation models in astronomy. In Proceedings of the Second Workshop on Information Extraction from Scientific Publications, pages 49–55.
  234. 234.Tung Nguyen, Johannes Brandstetter, Ashish Kapoor, Jayesh K Gupta, and Aditya Grover. 2023c. Climax: A foundation model for weather and climate. In ICML’23, pages 25904–25938.
  235. 235.Erik Nijkamp, Jeffrey A Ruffolo, Eli N Weinstein, Nikhil Naik, and Ali Madani. 2023. Progen2: exploring the boundaries of protein language models. Cell Systems, 14(11):968–978.
  236. 236.Maizhen Ning, Qiu-Feng Wang, Kaizhu Huang, and Xiaowei Huang. 2023. A symbolic characters aware model for solving geometry problems. In ACM MM’23, pages 7767–7775.
  237. 237.Janghoon Ock, Chakradhar Guntuboina, and Amir Barati Farimani. 2023. Catalyst energy prediction with catberta: Unveiling feature exploration strategies through large language models. ACS Catalysis, 13(24):16032–16044.
  238. 238.Malte Ostendorff, Nils Rethmeier, Isabelle Augenstein, Bela Gipp, and Georg Rehm. 2022. Neighborhood contrastive learning for scientific document representations with citation embeddings. In EMNLP’22, pages 11670–11688.
  239. 239.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. In NeurIPS’22, pages 27730–27744.
  240. 240.Ibrahim Burak Ozyurt. 2020. On the effectiveness of small, discriminatively pre-trained language representation models for biomedical text mining. In Proceedings of the First Workshop on Scholarly Document Processing, pages 104–112.
  241. 241.Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In CHIL’22, pages 248–260.
  242. 242.Panupong Pasupat and Percy Liang. 2015. Compositional semantic parsing on semi-structured tables. In ACL’15, pages 1470–1480.
  243. 243.Jaideep Pathak, Shashank Subramanian, Peter Harrington, Sanjeev Raja, Ashesh Chattopadhyay, Morteza Mardani, Thorsten Kurth, David Hall, Zongyi Li, Kamyar Azizzadenesheli, et al. 2022. Fourcastnet: A global data-driven high-resolution weather model using adaptive fourier neural operators. arXiv preprint arXiv:2202.11214.
  244. 244.Qizhi Pei, Lijun Wu, Kaiyuan Gao, Jinhua Zhu, Yue Wang, Zun Wang, Tao Qin, and Rui Yan. 2024. Leveraging biomolecule and natural language through multi-modal learning: A survey. arXiv preprint arXiv:2403.01528.
  245. 245.Qizhi Pei, Wei Zhang, Jinhua Zhu, Kehan Wu, Kaiyuan Gao, Lijun Wu, Yingce Xia, and Rui Yan. 2023. Biot5: Enriching cross-modal integration in biology with chemical knowledge and natural language associations. In EMNLP’23, pages 1102–1123.
  246. 246.Obioma Pelka, Sven Koitka, Johannes Rückert, Felix Nensa, and Christoph M Friedrich. 2018. Radiology objects in context (roco): a multimodal image dataset. In 7th Joint International Workshop, CVII-STENT and 3rd International Workshop, LABELS, Held in Conjunction with MICCAI’18, pages 180–189.
  247. 247.Chantal Pellegrini, Matthias Keicher, Ege Özsoy, Petra Jiraskova, Rickmer Braren, and Nassir Navab. 2023. Xplainer: From x-ray observations to explainable zero-shot diagnosis. In MICCAI’23, pages 420–429. Springer.
  248. 248.Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019. Transfer learning in biomedical natural language processing: An evaluation of bert and elmo on ten benchmarking datasets. In Proceedings of the 18th BioNLP Workshop and Shared Task, pages 58–65.
  249. 249.Ernest Perkowski, Rui Pan, Tuan Dung Nguyen, Yuan-Sen Ting, Sandor Kruk, Tong Zhang, Charlie O’Neill, Maja Jablonska, Zechang Sun, Michael J Smith, et al. 2024. Astrollama-chat: Scaling astrollama with conversational and diverse datasets. Research Notes of the AAS, 8(1):7.
  250. 250.Long N Phan, James T Anibal, Hieu Tran, Shaurya Chanana, Erol Bahadroglu, Alec Peltekian, and Grégoire Altan-Bonnet. 2021. Scifive: a text-to-text transformer model for biomedical literature. arXiv preprint arXiv:2106.03598.
  251. 251.Sara Pieri, Sahal Shaji Mullappilly, Fahad Shahbaz Khan, Rao Muhammad Anwer, Salman Khan, Timothy Baldwin, and Hisham Cholakkal. 2024. Bimedix: Bilingual medical mixture of experts llm. arXiv preprint arXiv:2402.13253.
  252. 252.Haoke Qiu, Lunyang Liu, Xuepeng Qiu, Xuemin Dai, Xiangling Ji, and Zhao-Yan Sun. 2024a. Polync: a natural and chemical language model for the prediction of unified polymer properties. Chemical Science, 15(2):534–544.
  253. 253.Pengcheng Qiu, Chaoyi Wu, Xiaoman Zhang, Weixiong Lin, Haicheng Wang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2024b. Towards building multilingual language model for medicine. arXiv preprint arXiv:2402.13963.
  254. 254.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In ICML’21, pages 8748–8763.
  255. 255.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21(140):1–67.
  256. 256.Raghunathan Ramakrishnan, Pavlo O Dral, Matthias Rupp, and O Anatole Von Lilienfeld. 2014. Quantum chemistry structures and properties of 134 kilo molecules. Scientific Data, 1(1):1–7.
  257. 257.Mayk Caldas Ramos, Shane S Michtavy, Marc D Porosoff, and Andrew D White. 2023. Bayesian optimization of catalysts with in-context learning. arXiv preprint arXiv:2304.05341.
  258. 258.Roshan M Rao, Jason Liu, Robert Verkuil, Joshua Meier, John Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives. 2021. Msa transformer. In ICML’21, pages 8844–8856.
  259. 259.Laila Rasmy, Yang Xiang, Ziqian Xie, Cui Tao, and Degui Zhi. 2021. Med-bert: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction. npj Digital Medicine, 4(1):86.
  260. 260.Sereina Riniker and Gregory A Landrum. 2013. Open-source platform to benchmark fingerprints for ligand-based virtual screening. Journal of Cheminformatics, 5(1):26.
  261. 261.Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C Lawrence Zitnick, Jerry Ma, et al. 2021. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. PNAS, 118(15):e2016239118.
  262. 262.Alexey Romanov and Chaitanya Shivade. 2018. Lessons from natural language inference in the clinical domain. In ACL’18, pages 1586–1596.
  263. 263.Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. 2024. Mathematical discoveries from program search with large language models. Nature, 625(7995):468–475.
  264. 264.Jerret Ross, Brian Belgodere, Vijil Chenthamarakshan, Inkit Padhi, Youssef Mroueh, and Payel Das. 2022. Large-scale chemical language representations capture molecular structure and properties. Nature Machine Intelligence, 4(12):1256–1264.
  265. 265.Andre Niyongabo Rubungo, Craig Arnold, Barry P Rand, and Adji Bousso Dieng. 2023. Llm-prop: Predicting physical and electronic properties of crystalline solids from their text descriptions. arXiv preprint arXiv:2310.14029.
  266. 266.Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn, Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, et al. 2024. Capabilities of gemini models in medicine. arXiv preprint arXiv:2404.18416.
  267. 267.Tobias Schimanski, Julia Bingler, Mathias Kraus, Camilla Hyslop, and Markus Leippold. 2023. Climatebert-netzero: Detecting and assessing net zero and reduction targets. In EMNLP’23, pages 15745–15756.
  268. 268.Nadine Schneider, Nikolaus Stiefl, and Gregory A Landrum. 2016. What’s what: The (nearly) definitive guide to reaction role assignment. Journal of Chemical Information and Modeling, 56(12):2336–2346.
  269. 269.Philippe Schwaller, Benjamin Hoover, Jean-Louis Reymond, Hendrik Strobelt, and Teodoro Laino. 2021a. Extraction of organic chemistry grammar from unsupervised learning of chemical reactions. Science Advances, 7(15):eabe4166.
  270. 270.Philippe Schwaller, Daniel Probst, Alain C Vaucher, Vishnu H Nair, David Kreutter, Teodoro Laino, and Jean-Louis Reymond. 2021b. Mapping the space of chemical reactions using attention-based neural networks. Nature Machine Intelligence, 3(2):144–152.
  271. 271.Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, Oren Etzioni, and Clint Malcolm. 2015. Solving geometry problems: Combining text and diagram interpretation. In EMNLP’15, pages 1466–1476.
  272. 272.Junyuan Shang, Tengfei Ma, Cao Xiao, and Jimeng Sun. 2019. Pre-training of graph augmented transformers for medication recommendation. In IJCAI’19, pages 5953–5959.
  273. 273.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, YK Li, Y Wu, and Daya Guo. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300.
  274. 274.Jia Tracy Shen, Michiharu Yamashita, Ethan Prihar, Neil Heffernan, Xintao Wu, Ben Graff, and Dongwon Lee. 2021. Mathbert: A pre-trained language model for general nlp tasks in mathematics education. arXiv preprint arXiv:2106.07340.
  275. 275.Pranav Shetty, Arunkumar Chitteth Rajan, Chris Kuenneth, Sonakshi Gupta, Lakshmi Prerana Panchumarti, Lauren Holm, Chao Zhang, and Rampi Ramprasad. 2023. A general-purpose material property data extraction pipeline from large polymer corpora using natural language processing. npj Computational Materials, 9(1):52.
  276. 276.Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, and Raghav Mani. 2020. Biomegatron: larger biomedical domain language model. In EMNLP’20, pages 4700–4706.
  277. 277.Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. 2024. Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers. arXiv preprint arXiv:2409.04109.
  278. 278.Amanpreet Singh, Mike D’Arcy, Arman Cohan, Doug Downey, and Sergey Feldman. 2023. Scirepeval: A multi-format benchmark for scientific document representations. In EMNLP’23, pages 5548–5566.
  279. 279.Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023a. Large language models encode clinical knowledge. Nature, 620(7972):172–180.
  280. 280.Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, et al. 2023b. Towards expert-level medical question answering with large language models. arXiv preprint arXiv:2305.09617.
  281. 281.Arnab Sinha, Zhihong Shen, Yang Song, Hao Ma, Darrin Eide, Bo-June Hsu, and Kuansan Wang. 2015. An overview of microsoft academic service (mas) and applications. In WWW’15, pages 243–246.
  282. 282.Shiven Sinha, Ameya Prabhu, Ponnurangam Kumaraguru, Siddharth Bhat, and Matthias Bethge. 2024. Wu’s method can boost symbolic ai to rival silver medalists and alphageometry to outperform gold medalists at imo geometry. arXiv preprint arXiv:2404.06405.
  283. 283.Panayiotis Smeros, Carlos Castillo, and Karl Aberer. 2021. Sciclops: Detecting and contextualizing scientific claims for assisting manual fact-checking. In CIKM’21, pages 1692–1702.
  284. 284.Henry Sprueill, Carl Edwards, Mariefel Olarte, Udishnu Sanyal, Heng Ji, and Sutanay Choudhury. 2023. Monte carlo thought search: Large language model querying for complex scientific reasoning in catalyst design. In Findings of EMNLP’23, pages 8348–8365.
  285. 285.Henry W Sprueill, Carl Edwards, Khushbu Agarwal, Mariefel V Olarte, Udishnu Sanyal, Conrad Johnston, Hongbin Liu, Heng Ji, and Sutanay Choudhury. 2024. Chemreasoner: Heuristic search over a large language model’s knowledge space using quantum-chemical feedback. In ICML’24, pages 46351–46374.
  286. 286.Teague Sterling and John J Irwin. 2015. Zinc 15–ligand discovery for everyone. Journal of Chemical Information and Modeling, 55(11):2324–2337.
  287. 287.Samuel Stevens, Jiaman Wu, Matthew J Thompson, Elizabeth G Campolongo, Chan Hee Song, David Edward Carlyn, Li Dong, Wasila M Dahdul, Charles Stewart, Tanya Berger-Wolf, et al. 2024. Bioclip: A vision foundation model for the tree of life. In CVPR’24, pages 19412–19424.
  288. 288.Bing Su, Dazhao Du, Zhao Yang, Yujie Zhou, Jiangmeng Li, Anyi Rao, Hao Sun, Zhiwu Lu, and Ji-Rong Wen. 2022. A molecular multimodal foundation model associating molecule graphs with natural language. arXiv preprint arXiv:2209.05481.
  289. 289.Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, and Fajie Yuan. 2024. Saprot: protein language modeling with structure-aware vocabulary. In ICLR’24.
  290. 290.Sanjay Subramanian, Lucy Lu Wang, Sachin Mehta, Ben Bogin, Madeleine van Zuylen, Sravanthi Parasa, Sameer Singh, Matt Gardner, and Hannaneh Hajishirzi. 2020. Medicat: A dataset of medical images, captions, and textual references. In Findings of EMNLP’20, pages 2112–2120.
  291. 291.Baris E Suzek, Yuqi Wang, Hongzhan Huang, Peter B McGarvey, Cathy H Wu, and UniProt Consortium. 2015. Uniref clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics, 31(6):926–932.
  292. 292.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, et al. 2023. Challenging big-bench tasks and whether chain-of-thought can solve them. In Findings of ACL’23, pages 13003–13051.
  293. 293.Jie Tang, Jing Zhang, Limin Yao, Juanzi Li, Li Zhang, and Zhong Su. 2008. Arnetminer: extraction and mining of academic social networks. In KDD’08, pages 990–998.
  294. 294.Tim Tanida, Philip Müller, Georgios Kaissis, and Daniel Rueckert. 2023. Interactive and explainable region-guided radiology report generation. In CVPR’23, pages 7433–7442.
  295. 295.Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085.
  296. 296.Omkar Thawkar, Abdelrahman Shaker, Sahal Shaji Mullappilly, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, Jorma Laaksonen, and Fahad Shahbaz Khan. 2024. Xraygpt: Chest radiographs summarization using medical vision-language models. In Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, pages 440–448.
  297. 297.Christina V Theodoris, Ling Xiao, Anant Chopra, Mark D Chaffin, Zeina R Al Sayed, Matthew C Hill, Helene Mantineo, Elizabeth M Brydon, Zexian Zeng, X Shirley Liu, et al. 2023. Transfer learning enables predictions in network biology. Nature, 618(7965):616–624.
  298. 298.Ekin Tiu, Ellie Talius, Pujan Patel, Curtis P Langlotz, Andrew Y Ng, and Pranav Rajpurkar. 2022. Expert-level detection of pathologies from unannotated chest x-ray images via self-supervised learning. Nature Biomedical Engineering, 6(12):1399–1406.
  299. 299.Shubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, and Igor Gitman. 2024. Openmathinstruct-1: A 1.8 million math instruction tuning dataset. arXiv preprint arXiv:2402.10176.
  300. 300.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  301. 301.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  302. 302.Amalie Trewartha, Nicholas Walker, Haoyan Huo, Sanghoon Lee, Kevin Cruse, John Dagdelen, Alexander Dunn, Kristin A Persson, Gerbrand Ceder, and Anubhav Jain. 2022. Quantifying the advantage of domain-specific pre-training on named entity recognition tasks in materials science. Patterns, 3(4).
  303. 303.Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. 2024. Solving olympiad geometry without human demonstrations. Nature, 625(7995):476–482.
  304. 304.Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Charles Lau, Ryutaro Tanno, Ira Ktena, et al. 2024. Towards generalist biomedical ai. NEJM AI, 1(3):AIoa2300138.
  305. 305.Saeid Ashraf Vaghefi, Dominik Stammbach, Veruska Muccione, Julia Bingler, Jingwei Ni, Mathias Kraus, Simon Allen, Chiara Colesanti-Senni, Tobias Wekhof, Tobias Schimanski, et al. 2023. Chatclimate: Grounding conversational ai in climate science. Communications Earth & Environment, 4(1):480.
  306. 306.Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. Trec-covid: constructing a pandemic information retrieval test collection. SIGIR Forum, 54(1):1–12.
  307. 307.Shoya Wada, Toshihiro Takeda, Shiro Manabe, Shozo Konishi, Jun Kamohara, and Yasushi Matsumura. 2020. Pre-training technique to localize medical bert and enhance biomedical bert. arXiv preprint arXiv:2005.07202.
  308. 308.Zhongwei Wan, Che Liu, Mi Zhang, Jie Fu, Benyou Wang, Sibo Cheng, Lei Ma, César Quilodrán-Casas, and Rossella Arcucci. 2023. Med-unic: Unifying cross-lingual medical vision-language pre-training by diminishing bias. In NeurIPS’23.
  309. 309.Benyou Wang, Qianqian Xie, Jiahuan Pei, Zhihong Chen, Prayag Tiwari, Zhao Li, and Jie Fu. 2023a. Pretrained language models in biomedical domain: A systematic survey. ACM Computing Surveys, 56(3):1–52.
  310. 310.Dongjie Wang, Chang-Tien Lu, and Yanjie Fu. 2023b. Towards automated urban planning: When generative and chatgpt-like ai meets urban planning. arXiv preprint arXiv:2304.03892.
  311. 311.Fuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti, and Lequan Yu. 2022a. Multi-granularity cross-modal alignment for generalized medical visual representation learning. In NeurIPS’22, pages 33536–33549.
  312. 312.Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, et al. 2023c. Scientific discovery in the age of artificial intelligence. Nature, 620(7972):47–60.
  313. 313.Hanyin Wang, Chufan Gao, Christopher Dantona, Bryan Hull, and Jimeng Sun. 2024a. Drg-llama: tuning llama model to predict diagnosis-related group for hospitalized patients. npj Digital Medicine, 7(1):16.
  314. 314.Haochun Wang, Chi Liu, Nuwa Xi, Zewen Qiang, Sendong Zhao, Bing Qin, and Ting Liu. 2023d. Huatuo: Tuning llama model with chinese medical knowledge. arXiv preprint arXiv:2304.06975.
  315. 315.Haorui Wang, Marta Skreta, Cher-Tian Ser, Wenhao Gao, Lingkai Kong, Felix Streith-Kalthoff, Chenru Duan, Yuchen Zhuang, Yue Yu, Yanqiao Zhu, et al. 2024b. Efficient evolutionary search over chemical space with large language models. arXiv preprint arXiv:2406.16976.
  316. 316.Ke Wang, Houxing Ren, Aojun Zhou, Zimu Lu, Sichun Luo, Weikang Shi, Renrui Zhang, Linqi Song, Mingjie Zhan, and Hongsheng Li. 2024c. Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning. In ICLR’24.
  317. 317.Qingyun Wang, Doug Downey, Heng Ji, and Tom Hope. 2023e. Scimon: Scientific inspiration machines optimized for novelty. In ACL’24, pages 279–299.
  318. 318.Qingyun Wang, Carl Edwards, Heng Ji, and Tom Hope. 2024d. Towards a human-computer collaborative scientific paper lifecycle: A pilot study and hands-on tutorial. In COLING’24, pages 56–67.
  319. 319.Ruida Wang, Jipeng Zhang, Yizhen Jia, Rui Pan, Shizhe Diao, Renjie Pi, and Tong Zhang. 2024e. Theoremlllama: Transforming general-purpose llms into lean4 experts. arXiv preprint arXiv:2407.03203.
  320. 320.Sheng Wang, Yuzhi Guo, Yuhong Wang, Hongmao Sun, and Junzhou Huang. 2019. Smiles-bert: large scale unsupervised pre-training for molecular property prediction. In ACM BCB’19, pages 429–436.
  321. 321.Xiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu, Jieyu Zhang, Satyen Subramaniam, Arjun R Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang. 2024f. Scibench: Evaluating college-level scientific problem-solving abilities of large language models. In ICML’24, pages 50622–50649.
  322. 322.Xidong Wang, Guiming Hardy Chen, Dingjie Song, Zhiyi Zhang, Zhihong Chen, Qingying Xiao, Feng Jiang, Jianquan Li, Xiang Wan, Benyou Wang, et al. 2024g. Cmb: A comprehensive medical benchmark in chinese. In NAACL’24, pages 6184–6205.
  323. 323.Yan Wang, Xiaojiang Liu, and Shuming Shi. 2017. Deep neural solver for math word problems. In EMNLP’17, pages 845–854.
  324. 324.Zhiruo Wang, Haoyu Dong, Ran Jia, Jia Li, Zhiyi Fu, Shi Han, and Dongmei Zhang. 2021. Tuta: Tree-based transformers for generally structured table pre-training. In KDD’21, pages 1780–1790.
  325. 325.Zifeng Wang, Lang Cao, Benjamin Danek, Yichi Zhang, Qiao Jin, Zhiyong Lu, and Jimeng Sun. 2024h. Accelerating clinical evidence synthesis with large language models. arXiv preprint arXiv:2406.17755.
  326. 326.Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, and Jimeng Sun. 2022b. Medclip: Contrastive learning from unpaired medical images and text. In EMNLP’22.
  327. 327.Neha Warikoo, Yung-Chun Chang, and Wen-Lian Hsu. 2021. Lbert: Lexically aware transformer-based bidirectional encoder representation model for learning universal bio-entity relations. Bioinformatics, 37(3):404–412.
  328. 328.Nicolas Webersinke, Mathias Kraus, Julia Anna Bingler, and Markus Leippold. 2021. Climatebert: A pretrained language model for climate-related text. arXiv preprint arXiv:2110.12010.
  329. 329.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022a. Finetuned language models are zero-shot learners. In ICLR’22.
  330. 330.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022b. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS’22, pages 24824–24837.
  331. 331.David Weininger. 1988. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of Chemical Information and Computer Sciences, 28(1):31–36.
  332. 332.Johannes Welbl, Nelson F Liu, and Matt Gardner. 2017. Crowdsourcing multiple choice science questions. In Proceedings of the 3rd Workshop on Noisy Usergenerated Text, pages 94–106.
  333. 333.Sean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi, and Yejin Choi. 2022. Naturalprover: Grounded mathematical proof generation with language models. In NeurIPS’22, pages 4913–4927.
  334. 334.Hongzhi Wen, Wenzhuo Tang, Xinnan Dai, Jiayuan Ding, Wei Jin, Yuying Xie, and Jiliang Tang. 2024. Cellplm: Pre-training of cell language model beyond single cells. In ICLR’24.
  335. 335.Andrew D White. 2023. The future of chemistry is language. Nature Reviews Chemistry, 7(7):457–458.
  336. 336.Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang, Weidi Xie, and Yanfeng Wang. 2024. Pmc-llama: toward building open-source language models for medicine. JAMIA, 31(9):1833–1843.
  337. 337.Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. Towards generalist foundation model for radiology. arXiv preprint arXiv:2308.02463.
  338. 338.Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. 2018. Moleculenet: a benchmark for molecular machine learning. Chemical Science, 9(2):513–530.
  339. 339.Jun Xia, Yanqiao Zhu, Yuanqi Du, and Stan Z Li. 2023. A systematic survey of chemical pre-trained models. In IJCAI’23, pages 6787–6795.
  340. 340.Qianqian Xie, Qingyu Chen, Aokun Chen, Cheng Peng, Yan Hu, Fongci Lin, Xueqing Peng, Jimin Huang, Jeffrey Zhang, Vipina Keloth, et al. 2024. Me llama: Foundation large language models for medical applications. arXiv preprint arXiv:2402.12749.
  341. 341.Tong Xie, Yuwei Wan, Wei Huang, Zhenyu Yin, Yixuan Liu, Shaozhou Wang, Qingyuan Linghu, Chunyu Kit, Clara Grazian, Wenjie Zhang, et al. 2023. Darwin series: Domain specific large language models for natural science. arXiv preprint arXiv:2308.13565.
  342. 342.Guangzhi Xiong, Qiao Jin, Zhiyong Lu, and Aidong Zhang. 2024. Benchmarking retrieval-augmented generation for medicine. In Findings of ACL’24, pages 6233–6251.
  343. 343.Honglin Xiong, Sheng Wang, Yitao Zhu, Zihao Zhao, Yuxiao Liu, Linlin Huang, Qian Wang, and Dinggang Shen. 2023. Doctorglm: Fine-tuning your chinese doctor is not a herculean task. arXiv preprint arXiv:2304.01097.
  344. 344.Changwen Xu, Yuyang Wang, and Amir Barati Farimani. 2023a. Transpolymer: a transformer-based language model for polymer property predictions. npj Computational Materials, 9(1):64.
  345. 345.Minghao Xu, Xinyu Yuan, Santiago Miret, and Jian Tang. 2023b. Protst: Multi-modality learning of protein sequences and biomedical texts. In ICML’23, pages 38749–38767.
  346. 346.Ran Xu, Wenqi Shi, Yue Yu, Yuchen Zhuang, Yanqiao Zhu, May D Wang, Joyce C Ho, Chao Zhang, and Carl Yang. 2024. Bmretriever: Tuning large language models as better biomedical text retrievers. arXiv preprint arXiv:2404.18443.
  347. 347.Hiroki Yamauchi, Tomoyuki Kajiwara, Marie Katsurai, Ikki Ohmukai, and Takashi Ninomiya. 2022. A japanese masked language model for academic domain. In Proceedings of the Third Workshop on Scholarly Document Processing, pages 152–157.
  348. 348.Yibo Yan, Haomin Wen, Siru Zhong, Wei Chen, Haodong Chen, Qingsong Wen, Roger Zimmermann, and Yuxuan Liang. 2024. Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web. In WWW’24, pages 4006–4017.
  349. 349.Fan Yang, Wenchuan Wang, Fang Wang, Yuan Fang, Duyu Tang, Junzhou Huang, Hui Lu, and Jianhua Yao. 2022a. scbert as a large-scale pretrained deep language model for cell type annotation of singlecell rna-seq data. Nature Machine Intelligence, 4(10):852–866.
  350. 350.Lin Yang, Shawn Xu, Andrew Sellergren, Timo Kohlberger, Yuchen Zhou, Ira Ktena, Atilla Kiraly, Faruk Ahmed, Farhad Hormozdiari, Tiam Jaroensri, et al. 2024a. Advancing multimodal medical capabilities of gemini. arXiv preprint arXiv:2405.03162.
  351. 351.Songhua Yang, Hanjie Zhao, Senbin Zhu, Guangyu Zhou, Hongfei Xu, Yuxiang Jia, and Hongying Zan. 2024b. Zhongjing: Enhancing the chinese medical capabilities of large language model through expert feedback and real-world multi-turn dialogue. In AAAI’24, pages 19368–19376.
  352. 352.Xi Yang, Jiang Bian, William R Hogan, and Yonghui Wu. 2020. Clinical concept extraction using transformers. JAMIA, 27(12):1935–1942.
  353. 353.Xi Yang, Aokun Chen, Nima PourNejatian, Hoo Chang Shin, Kaleb E Smith, Christopher Parisien, Colin Compas, Cheryl Martin, Anthony B Costa, Mona G Flores, et al. 2022b. A large language model for electronic health records. npj Digital Medicine, 5(1):194.
  354. 354.Xianjun Yang, Junfeng Gao, Wenxin Xue, and Erik Alexandersson. 2024c. Pllama: An open-source large language model for plant science. arXiv preprint arXiv:2401.01600.
  355. 355.Zonglin Yang, Xinya Du, Junxian Li, Jie Zheng, Soujanya Poria, and Erik Cambria. 2024d. Large language models for automated open-domain scientific hypotheses discovery. In Findings of ACL’24, pages 13545–13565.
  356. 356.Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christopher D Manning, Percy S Liang, and Jure Leskovec. 2022a. Deep bidirectional language-knowledge graph pretraining. In NeurIPS’22, pages 37309–37323.
  357. 357.Michihiro Yasunaga, Jure Leskovec, and Percy Liang. 2022b. Linkbert: Pretraining language models with document links. In ACL’22, pages 8003–8016.
  358. 358.Geyan Ye, Xibao Cai, Houtim Lai, Xing Wang, Junhong Huang, Longyue Wang, Wei Liu, and Xiangxiang Zeng. 2023a. Drugassist: A large language model for molecule optimization. arXiv preprint arXiv:2401.10334.
  359. 359.Qichen Ye, Junling Liu, Dading Chong, Peilin Zhou, Yining Hua, and Andrew Liu. 2023b. Qilinmed: Multi-stage knowledge injection advanced medical large language model. arXiv preprint arXiv:2310.09089.
  360. 360.Junqi Yin, Sajal Dash, Feiyi Wang, and Mallikarjun Shankar. 2023. Forge: pre-training open foundation models for science. In SC’23, pages 1–13.
  361. 361.Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020. Tabert: Pretraining for joint understanding of textual and tabular data. In ACL’20, pages 8413–8426.
  362. 362.Huaiyuan Ying, Shuo Zhang, Linyang Li, Zhejian Zhou, Yunfan Shao, Zhaoye Fei, Yichuan Ma, Jiawei Hong, Kuikun Liu, Ziyi Wang, et al. 2024. Internlm-math: Open math large language models toward verifiable reasoning. arXiv preprint arXiv:2402.06332.
  363. 363.Kihyun You, Jawook Gu, Jiyeon Ham, Beomhee Park, Jiho Kim, Eun K Hong, Woonhyuk Baek, and Byungseok Roh. 2023. Cxr-clip: Toward large scale chest x-ray language-image pre-training. In MICCAI’23, pages 101–111.
  364. 364.Botao Yu, Frazier N Baker, Ziqi Chen, Xia Ning, and Huan Sun. 2024a. Llasmol: Advancing large language models for chemistry with a large-scale, comprehensive, high-quality instruction tuning dataset. arXiv preprint arXiv:2402.09391.
  365. 365.Fei Yu, Anningzhe Gao, and Benyou Wang. 2024b. Ovm, outcome-supervised value models for planning in mathematical reasoning. In Findings of NAACL’24, pages 858–875.
  366. 366.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2024c. Metamath: Bootstrap your own mathematical questions for large language models. In ICLR’24.
  367. 367.Tao Yu, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang, Yi Chern Tan, Xinyi Yang, Dragomir Radev, Richard Socher, and Caiming Xiong. 2021. Grappa: Grammar-augmented pre-training for table semantic parsing. In ICLR’21.
  368. 368.Hongyi Yuan, Zheng Yuan, Ruyi Gan, Jiaxing Zhang, Yutao Xie, and Sheng Yu. 2022a. Biobart: Pretraining and evaluation of a biomedical generative language model. In Proceedings of the 21st Workshop on Biomedical Language Processing, pages 97–109.
  369. 369.Zheng Yuan, Yijia Liu, Chuanqi Tan, Songfang Huang, and Fei Huang. 2021. Improving biomedical pretrained language models with knowledge. In Proceedings of the 20th Workshop on Biomedical Language Processing, pages 180–190.
  370. 370.Zheng Yuan, Zhengyun Zhao, Haixia Sun, Jiao Li, Fei Wang, and Sheng Yu. 2022b. Coder: Knowledgeinfused cross-lingual medical term embedding for term normalization. Journal of Biomedical Informatics, 126:103983.
  371. 371.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. 2024a. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In CVPR’24, pages 9556–9567.
  372. 372.Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2024b. Mammoth: Building math generalist models through hybrid instruction tuning. In ICLR’24.
  373. 373.Xiang Yue, Tuney Zheng, Ge Zhang, and Wenhu Chen. 2024c. Mammoth2: Scaling instructions from the web. arXiv preprint arXiv:2405.03548.
  374. 374.Atakan Yüksel, Erva Ulusoy, Atabey Ünlü, and Tunca Doğan. 2023. Selformer: molecular representation learning via selfies language models. Machine Learning: Science and Technology, 4(2):025035.
  375. 375.Zheni Zeng, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2022. A deep-learning system bridging molecule structure and biomedical text with comprehension comparable to human professionals. Nature Communications, 13(1):862.
  376. 376.Dan Zhang, Ziniu Hu, Sining Zhoubian, Zhengxiao Du, Kaiyu Yang, Zihan Wang, Yisong Yue, Yuxiao Dong, and Jie Tang. 2024a. Sciglm: Training scientific language models with self-reflective instruction annotation and tuning. arXiv preprint arXiv:2401.07950.
  377. 377.Daoan Zhang, Weitong Zhang, Bing He, Jianguo Zhang, Chenchen Qin, and Jianhua Yao. 2023a. Dnagpt: A generalized pretrained tool for multiple dna sequence analysis tasks. bioRxiv, pages 2023–07.
  378. 378.Di Zhang, Wei Liu, Qian Tan, Jingdan Chen, Hang Yan, Yuliang Yan, Jiatong Li, Weiran Huang, Xiangyu Yue, Dongzhan Zhou, et al. 2024b. Chemllm: A chemical large language model. arXiv preprint arXiv:2402.06852.
  379. 379.Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Guiming Chen, Jianquan Li, Xiangbo Wu, Zhang Zhiyi, Qingying Xiao, et al. 2023b. Huatuogpt, towards taming language model to be a doctor. In Findings of EMNLP’23, pages 10859–10885.
  380. 380.Kai Zhang, Rong Zhou, Eashan Adhikarla, Zhiling Yan, Yixin Liu, Jun Yu, Zhengliang Liu, Xun Chen, Brian D Davison, Hui Ren, et al. 2024c. A generalist vision–language foundation model for diverse biomedical tasks. Nature Medicine, pages 1–13.
  381. 381.Ningyu Zhang, Qianghuai Jia, Kangping Yin, Liang Dong, Feng Gao, and Nengwei Hua. 2020. Conceptualized representation learning for chinese biomedical text mining. arXiv preprint arXiv:2008.10813.
  382. 382.Qiang Zhang, Keyang Ding, Tianwen Lyv, Xinda Wang, Qingyu Yin, Yiwen Zhang, Jing Yu, Yuhao Wang, Xiaotong Li, Zhuoyi Xiang, et al. 2024d. Scientific large language models: A survey on biological & chemical domains. arXiv preprint arXiv:2401.14656.
  383. 383.Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, et al. 2023c. Biomedclip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs. arXiv preprint arXiv:2303.00915.
  384. 384.Tianshu Zhang, Xiang Yue, Yifei Li, and Huan Sun. 2024e. Tablellama: Towards open large generalist models for tables. In NAACL’24, pages 6024–6044.
  385. 385.Xiaokang Zhang, Jing Zhang, Zeyao Ma, Yang Li, Bohan Zhang, Guanlin Li, Zijun Yao, Kangli Xu, Jinchang Zhou, Daniel Zhang-Li, et al. 2024f. Tablellm: Enabling tabular data manipulation by llms in real office usage scenarios. arXiv preprint arXiv:2403.19318.
  386. 386.Xinlu Zhang, Chenxin Tian, Xianjun Yang, Lichang Chen, Zekun Li, and Linda Ruth Petzold. 2023d. Alpacare: Instruction-tuned large language models for medical application. arXiv preprint arXiv:2310.14558.
  387. 387.Xuan Zhang, Limei Wang, Jacob Helwig, Youzhi Luo, Cong Fu, Yaochen Xie, Meng Liu, Yuchao Lin, Zhao Xu, Keqiang Yan, et al. 2023e. Artificial intelligence for science in quantum, atomistic, and continuum systems. arXiv preprint arXiv:2307.08423.
  388. 388.Yikun Zhang, Mei Lang, Jiuhong Jiang, Zhiqiang Gao, Fan Xu, Thomas Litfin, Ke Chen, Jaswinder Singh, Xiansong Huang, Guoli Song, et al. 2024g. Multiple sequence alignment-based rna language model and its application to structural inference. Nucleic Acids Research, 52(1):e3–e3.
  389. 389.Yu Zhang, Hao Cheng, Zhihong Shen, Xiaodong Liu, Ye-Yi Wang, and Jianfeng Gao. 2023f. Pre-training multi-task contrastive learning models for scientific literature understanding. In Findings of EMNLP’23, pages 12259–12275.
  390. 390.Yu Zhang, Bowen Jin, Qi Zhu, Yu Meng, and Jiawei Han. 2023g. The effect of metadata on scientific literature tagging: A cross-field cross-model study. In WWW’23, pages 1626–1637.
  391. 391.Yu Zhang, Yanzhen Shen, SeongKu Kang, Xiusi Chen, Bowen Jin, and Jiawei Han. 2023h. Chain-of-factors paper-reviewer matching. arXiv preprint arXiv:2310.14483.
  392. 392.Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz. 2022. Contrastive learning of medical visual representations from paired images and text. In MLHC’22, pages 2–25.
  393. 393.Yunkun Zhang, Jin Gao, Mu Zhou, Xiaosong Wang, Yu Qiao, Shaoting Zhang, and Dequan Wang. 2023i. Text-guided foundation model adaptation for pathological image classification. In MICCAI’23, pages 272–282.
  394. 394.Haiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu, Jie Fu, Zhi-Hong Deng, Lingpeng Kong, and Qi Liu. 2023a. Gimlet: A unified graph-text model for instruction-based molecule zero-shot learning. In NeurIPS’23.
  395. 395.Suyuan Zhao, Jiahuan Zhang, and Zaiqing Nie. 2023b. Large-scale cell representation learning via divide-and-conquer contrastive learning. arXiv preprint arXiv:2306.04371.
  396. 396.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023c. A survey of large language models. arXiv preprint arXiv:2303.18223.
  397. 397.Wei Zhao, Mingyue Shang, Yang Liu, Liang Wang, and Jingming Liu. 2020. Ape210k: A large-scale and template-rich dataset of math word problems. arXiv preprint arXiv:2009.11506.
  398. 398.Yilun Zhao, Linyong Nan, Zhenting Qi, Rui Zhang, and Dragomir Radev. 2022. Reastap: Injecting table reasoning skills during pre-training via synthetic reasoning examples. In EMNLP’22, pages 9006–9018.
  399. 399.Zihan Zhao, Da Ma, Lu Chen, Liangtai Sun, Zihao Li, Hongshen Xu, Zichen Zhu, Su Zhu, Shuai Fan, Guodong Shen, et al. 2024. Chemdfm: Dialogue foundation model for chemistry. arXiv preprint arXiv:2401.14818.
  400. 400.Yizhen Zheng, Huan Yee Koh, Jiaxin Ju, Anh TN Nguyen, Lauren T May, Geoffrey I Webb, and Shirui Pan. 2023a. Large language models for scientific synthesis, inference and explanation. arXiv preprint arXiv:2310.07984.
  401. 401.Zaixiang Zheng, Yifan Deng, Dongyu Xue, Yi Zhou, Fei Ye, and Quanquan Gu. 2023b. Structure-informed language models are protein designers. In ICML’23, pages 42317–42338.
  402. 402.Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2sql: Generating structured queries from natural language using reinforcement learning. arXiv preprint arXiv:1709.00103.
  403. 403.Gengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng, Hongteng Xu, Zhewei Wei, Linfeng Zhang, and Guolin Ke. 2023. Uni-mol: A universal 3d molecular representation learning framework. In ICLR’23.
  404. 404.Zhihan Zhou, Yanrong Ji, Weijian Li, Pratik Dutta, Ramana Davuluri, and Han Liu. 2024a. Dnabert-2: Efficient foundation model and benchmark for multi-species genome. In ICLR’24.
  405. 405.Zhilun Zhou, Yuming Lin, Depeng Jin, and Yong Li. 2024b. Large language model for participatory urban planning. arXiv preprint arXiv:2402.17161.
  406. 406.Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024. Can large language models transform computational social science? Computational Linguistics, 50(1):237–291.
  407. 407.Maxim Zvyagin, Alexander Brace, Kyle Hippe, Yuntian Deng, Bin Zhang, Cindy Orozco Bohorquez, Austin Clyde, Bharat Kale, Danilo Perez-Rivera, Heng Ma, et al. 2023. Genslms: Genome-scale language models reveal sars-cov-2 evolutionary dynamics. The International Journal of High Performance Computing Applications, 37(6):683–705.

Citation

MLA
Zhang, Y., et al. “A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 8783–817, https://doi.org/10.18653/v1/2024.emnlp-main.498.
APA
Zhang, Y., Chen, X., Jin, B., Wang, S., Ji, S., Wang, W., & Han, J. (2024). A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 8783–8817. https://doi.org/10.18653/v1/2024.emnlp-main.498
Chicago
Zhang, Y., X. Chen, B. Jin, et al. 2024. “A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 8783–8817. https://doi.org/10.18653/v1/2024.emnlp-main.498.
Harvard
Zhang, Y. et al. (2024) “A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 8783–8817. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.498.
Vancouver
1. Zhang Y, Chen X, Jin B, Wang S, Ji S, Wang W, Han J (2024) A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 8783–8817

BibTeX

@inproceedings{zhang-etal-2024-comprehensive-survey,
    title = "A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery",
    author = "Zhang, Yu  and
      Chen, Xiusi  and
      Jin, Bowen  and
      Wang, Sheng  and
      Ji, Shuiwang  and
      Wang, Wei  and
      Han, Jiawei",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.498/",
    doi = "10.18653/v1/2024.emnlp-main.498",
    pages = "8783--8817"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/