Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

Tianyu LiuAllen Xin WangAntonia PanescuLisa Xinyi ChenWenxin LongXinyu WeiYueqian JingZiyao ZengJihang ChenSihan Jiang

article2026arXiv0 citations

Introduces SciAgentArena, an interactive evaluation benchmark of roughly 200 multi-domain tasks with stepwise verification, revealing that while AI agents handle structured data analysis effectively, they struggle with open-ended exploration and novel scientific reasoning.

Listen

Artificial intelligence agents based on large language models are increasingly promoted as tools to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchmarks either evaluate language models on static, multiple-choice questions or assess general coding and mathematical reasoning in non-scientific settings. The article addresses this evaluation gap by introducing SciAgentArena, a systematic benchmark designed to evaluate how effectively AI agents navigate complex, multi-step scientific workflows across five key biomedical domains: drug discovery, single-cell omics, spatial omics, electronic health records modeling, and statistical genetics.

The main objective of the article is to systematically assess the practical capabilities, reliability, and failure modes of current generalist and domain-specialist AI agents on realistic, multi-step scientific research tasks. The benchmark measures agent performance across four core scientific dimensions: standard data analysis, model selection, open-ended model optimization, and scientific validity checking.

To evaluate these capabilities, the researchers constructed approximately 200 verifiable tasks curated by domain experts across the five scientific domains. They evaluated 18 AI agents, including frontier general-purpose models, specialized biomedical agents, and coding-assistant frameworks. The benchmarking environment uses an agent-agnostic architecture that separates the agent execution sandbox from the automated evaluation system. This structure enables unit-test-style verification of intermediate computational states, tool use, runtime data structures, and final scientific outputs without configuration conflicts.

The findings show that current AI agents act as promising but highly uneven scientific collaborators. Agents perform best on structured data analysis and established preprocessing pipelines where procedures and evaluation criteria are explicitly specified. However, performance drops sharply in open-ended model optimization; for example, agents routinely solve single-objective molecular designs but struggle with multi-constraint optimization tasks. In model selection, agents exhibit severe conservative convergence by defaulting to popular, documented methods (such as Harmony for batch correction or Leiden for clustering) rather than choosing optimal approaches for specific data constraints. Domain-specialist agents generally offer advantages in tasks requiring curated tools and specialized workflows, but generalist agents with robust code-execution loops frequently match or exceed specialist performance. Critically, agents perform poorly on validity checks: they often accept flawed prompts, execute calculations on contradictory data without noticing anomalies, or hallucinate non-existent programming interfaces rather than refusing invalid tasks.

These findings indicate that while AI agents can reliably automate routine, well-defined data processing pipelines, they lack autonomous scientific judgment, robust state tracking, and spontaneous error correction. Relying on agents for critical research decisions without rigorous verification poses substantial risks, including silent propagation of data errors, flawed causal conclusions, and wasted experimental resources. The results challenge the assumption that expanding domain-specific tool libraries alone will produce competent AI scientists; procedural execution ability does not equate to sound scientific reasoning.

The article recommends that organizations deploy AI agents primarily for bounded, verifiable data analysis workflows while maintaining mandatory human oversight for model selection, optimization, and scientific validation. Developers should enhance future agents with explicit runtime verification mechanisms, such as checking installed software interfaces, validating data context before execution, enforcing step-wise state persistence, and integrating built-in refusal protocols for ill-posed requests. Future benchmarking efforts should expand into additional scientific disciplines, such as physics and materials science, to further assess agent reliability across diverse research environments.

Cover for Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

Abstract

AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchmarks for AI agents rarely capture the complexity, heterogeneity, and extended reasoning required by scientific work, whereas benchmarks for scientific tasks often reduce research to static, direct problems and provide limited support for interactive evaluation. Here, we introduce SciAgentArena, a systematic benchmark for evaluating AI agents in real-world scientific research scenarios drawn from emerging needs across multiple domains. SciAgentArena comprises approximately 200 tasks with stepwise verification and an interactive, agent-agnostic environment for assessing diverse AI agents. Using this benchmark, we find that current agents can contribute effectively to well-specified data-analysis workflows, particularly when the task structure and evaluation criteria are clear. However, their performance remains uneven across scientific contexts: agents struggle to generate genuinely novel insights, sustain self-directed exploration, and formulate robust solutions for open-ended research questions. We further characterize common failure modes across agents and identify opportunities for improving their reliability, autonomy, and scientific reasoning. Together, SciAgentArena provides a practical framework for measuring progress in AI agents for science and for guiding the design of future agents capable of addressing complex scientific challenges. Full codes, tasks, and datasets can be accessed via this link: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Results
  • 2.1 Benchmark Overview.
  • 2.2 Benchmarking AI Agents for Computational Drug Discovery.
  • 2.3 Benchmarking AI Agents for Processing Cellular and Tissue-Level Signals.
  • 2.4 Benchmarking AI Agents on Processing Patient-Level Information.
  • 2.5 Benchmarking AI Agents on Cross-Domain Tasks.
  • 2.6 Error Analyses and General Suggestions Summarized from Different Scientific Fields.
  • 3 Discussion
  • 4 Methods
  • 4.1 Framework
  • 4.2 Domain: Computational Drug Discovery
  • 4.3 Domain: Single-Cell Omics
  • 4.4 Domain: Spatial Omics
  • 4.5 Domain: Electronic Health Records
  • 4.6 Domain: Genetics
  • 4.7 Cross Domain Tasks.
  • 5 Code and Data Availability
  • 6 Author Contributions
  • 7 Acknowledgments
  • 8 Conflict of Interests
  • References
  • A Categories of AI Agents
  • B Differences between our framework and other benchmark studies
  • C Prompt list
  • D Supplementary Figures

Citation

MLA
Liu, T., et al. “Benchmarking AI Agents for Addressing Scientific Challenges Across Scales”. arXiv, 2026, http://arxiv.org/abs/2606.12736v1.
APA
Liu, T., Wang, A. X., Panescu, A., Chen, L. X., Long, W., Wei, X., Jing, Y., Zeng, Z., Chen, J., Jiang, S., Wang, Z., Gu, S., Chen, S., Hu, X., Shao, H., Xu, L., Zheng, W., Cao, Z., Fang, A., … Zhao, H. (2026). Benchmarking AI Agents for Addressing Scientific Challenges Across Scales. arXiv. http://arxiv.org/abs/2606.12736v1
Chicago
Liu, T., A. X. Wang, A. Panescu, et al. 2026. “Benchmarking AI Agents for Addressing Scientific Challenges Across Scales”. arXiv. http://arxiv.org/abs/2606.12736v1.
Harvard
Liu, T. et al. (2026) “Benchmarking AI Agents for Addressing Scientific Challenges Across Scales”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2606.12736v1.
Vancouver
1. Liu T, Wang AX, Panescu A, et al (2026) Benchmarking AI Agents for Addressing Scientific Challenges Across Scales. arXiv

BibTeX

@article{liu2026benchmarking,
  title = {Benchmarking AI Agents for Addressing Scientific Challenges Across Scales},
  author = {Liu, Tianyu and Wang, Allen Xin and Panescu, Antonia and Chen, Lisa Xinyi and Long, Wenxin and Wei, Xinyu and Jing, Yueqian and Zeng, Ziyao and Chen, Jihang and Jiang, Sihan and Wang, Ziqing and Gu, Siyi and Chen, Siyu and Hu, Xinyang and Shao, Haoran and Xu, Leqi and Zheng, Wangjie and Cao, Zhiyuan and Fang, Ada and Yu, Botao and Sun, Kunyang and Ying, Rex and Cohan, Arman and Chen, Qingyu and Xue, Lingzhou and Ding, Kaize and Du, Yuanqi and Jin, Wengong and Yang, Zhuoran and Zitnik, Marinka and Zou, James and Xu, Hua and Zhao, Hongyu},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2606.12736v1},
  eprint = {2606.12736}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission