A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery
Yu ZhangXiusi ChenBowen JinSheng WangShuiwang JiWei WangJiawei Han
Presents a systematic review of more than 260 scientific large language models across multiple disciplines and data modalities, categorizing their architectures, pre-training strategies, and practical applications in accelerating scientific discovery.
Scientific research across medicine, chemistry, biology, physics, and climate science increasingly relies on artificial intelligence to process vast amounts of complex data. While specialized large language models have emerged within individual disciplines, earlier reviews have analyzed these tools in isolation, focusing only on single fields or single data types like plain text. This siloed perspective obscures common architectural principles and prevents researchers from sharing proven technical strategies across domains. The article addresses this fragmentation by establishing a unified framework that synthesizes how scientific language models are built, trained, and applied across diverse scientific disciplines.
The main objective of the article is to provide a comprehensive cross-field and cross-modal evaluation of specialized large language models and to demonstrate how these systems augment modern scientific discovery. To accomplish this, the authors conducted an extensive literature review covering more than 260 scientific models ranging in scale from 100 million to over 100 billion parameters. The analysis spans multiple fields—including general science, mathematics, physics, chemistry, materials science, biology, medicine, and geoscience—while examining how non-text data such as molecular graphs, biological sequences, crystal lattices, tables, images, and climate time series are adapted for model pre-training and evaluation.
The article identifies several core findings regarding the development and utility of scientific language models. First, it demonstrates that scientific pre-training converges into three universal paradigms regardless of discipline: masked sequence modeling for structured representation learning, autoregressive next-token prediction often coupled with instruction tuning for complex reasoning and generation, and contrastive multi-modal alignment to map distinct data types into shared latent spaces. Second, model development has shifted from smaller, task-specific encoder architectures toward multi-billion-parameter generative models capable of following natural language instructions across diverse tasks. Third, when applied directly to discovery workflows, these models demonstrate autonomous problem-solving capabilities, including generating viable research hypotheses, designing CRISPR gene-editing experiments, outperforming human benchmarks in mathematical Olympiad geometry, planning chemical syntheses, and predicting global weather patterns faster than traditional numerical simulations.
These findings indicate that treating complex scientific structures—such as amino acid sequences and molecular graphs—as specialized languages significantly lowers the barrier to deploying unified artificial intelligence workflows across science. Adopting standardized model architectures reduces the cost and development time needed to build domain-specific tools, while multi-modal integration improves data utilization across multiomics, clinical records, and material databases. However, because specialized scientific tasks demand absolute precision, the tendency of language models to generate plausible but factually incorrect outputs introduces safety and reliability risks, particularly in clinical and chemical decision-making.
To address these challenges, the article recommends developing targeted, theme-focused knowledge bases to prevent models from losing rare domain insights during broad training, as well as advancing multi-modal retrieval-augmented generation systems that verify outputs against verified experimental data and chemical structures. Future efforts should also incorporate invariant learning techniques to ensure models remain reliable when tested on unseen molecular scaffolds or novel scientific concepts. Because current evaluations remain centered on mathematics and natural sciences rather than social sciences, stakeholders should maintain cautious confidence in specialized model outputs until domain experts thoroughly validate them against empirical benchmarks.
- Paper: A Survey of Large Language Models, Wayne Xin Zhao et al. (2023). Provides a comprehensive foundational review of large language model architectures, pre-training paradigms, and alignment techniques that are adapted for domain-specific scientific tasks.
- Paper: SciBERT: A Pretrained Language Model for Scientific Text, Iz Beltagy et al. (2019). Demonstrates the foundational methodology of pre-training language models directly on scientific literature corpora to capture domain-specific vocabulary and semantics.
- Paper: ProtTrans: Towards Cracking the Language of Life’s Code Through Self-Supervised Deep Learning and High Performance Computing, Ahmed Elnaggar et al. (2020). Establishes how self-supervised language modeling architectures can treat biological sequences as structured languages, a core paradigm surveyed in scientific discovery.
- Paper: Solving Quantitative Reasoning Problems with Language Models, Aitor Lewkowycz et al. (2022). Introduces key techniques for pre-training large language models on technical and mathematical corpora to enable quantitative reasoning across STEM domains.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). Surveys the design spaces and cross-modal alignment mechanisms of multimodal large language models that underpin multimodal scientific foundation models.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). Outlines paradigms for integrating structured knowledge bases and knowledge graphs with language models to mitigate factual errors in specialized domains.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Establishes the standard architectural taxonomy of memory, planning, and action modules for LLM-based autonomous agents executing complex problem-solving workflows.
- Paper: A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al. (2023). Categorizes the mechanisms and mitigation strategies of hallucinations in large language models, explaining the reliability challenges highlighted in scientific deployment.
- Paper: Innovator-VL: A Multimodal Large Language Model for Scientific Discovery, Zichen Wen et al. (2026). Applies multimodal scientific modeling principles to build a unified vision-language model trained specifically on complex scientific imagery such as molecular structures and micrographs.
- Paper: LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery, Pingchuan Ma et al. (2024). Extends the surveyed concepts of generative scientific reasoning by coupling large language models with differentiable physical simulations in a bilevel optimization framework for automated discovery.
- Paper: Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models, Kevin Murphy (2026). Implements an autonomous scientific discovery agent that couples LLM hypothesis generation with Bayesian experiment design to identify mechanistic world models across physics, chemistry, and neuroscience.
- Paper: OceanGPT: A Large Language Model for Ocean Science Tasks, Zhen Bi et al. (2024). Provides a concrete realization of domain-specific scientific LLMs by developing OceanGPT and multi-agent instruction synthesis for oceanographic science tasks.
- Paper: Idea2Story: An Automated Pipeline for Transforming Research Concepts into Complete Scientific Narratives, Tengyue Xu et al. (2026). Operationalizes automated scientific narrative and research plan generation by combining LLMs with structured knowledge graphs of reusable research method units.
- Paper: Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models, Fengli Xu et al. (2025). Explores advanced reinforcement learning and test-time reasoning paradigms that directly address the complex multi-step reasoning bottlenecks identified across scientific tasks.
- Paper: Position: LLMs can't jump, Tom Zahavy (2026). Presents a critical epistemological critique on whether transformer-based LLMs can achieve fundamental, creative theoretical leaps in scientific discovery beyond statistical induction.
