Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
Tianyu LiuAllen Xin WangAntonia PanescuLisa Xinyi ChenWenxin LongXinyu WeiYueqian JingZiyao ZengJihang ChenSihan Jiang
Introduces SciAgentArena, an interactive evaluation benchmark of roughly 200 multi-domain tasks with stepwise verification, revealing that while AI agents handle structured data analysis effectively, they struggle with open-ended exploration and novel scientific reasoning.
Artificial intelligence agents based on large language models are increasingly promoted as tools to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchmarks either evaluate language models on static, multiple-choice questions or assess general coding and mathematical reasoning in non-scientific settings. The article addresses this evaluation gap by introducing SciAgentArena, a systematic benchmark designed to evaluate how effectively AI agents navigate complex, multi-step scientific workflows across five key biomedical domains: drug discovery, single-cell omics, spatial omics, electronic health records modeling, and statistical genetics.
The main objective of the article is to systematically assess the practical capabilities, reliability, and failure modes of current generalist and domain-specialist AI agents on realistic, multi-step scientific research tasks. The benchmark measures agent performance across four core scientific dimensions: standard data analysis, model selection, open-ended model optimization, and scientific validity checking.
To evaluate these capabilities, the researchers constructed approximately 200 verifiable tasks curated by domain experts across the five scientific domains. They evaluated 18 AI agents, including frontier general-purpose models, specialized biomedical agents, and coding-assistant frameworks. The benchmarking environment uses an agent-agnostic architecture that separates the agent execution sandbox from the automated evaluation system. This structure enables unit-test-style verification of intermediate computational states, tool use, runtime data structures, and final scientific outputs without configuration conflicts.
The findings show that current AI agents act as promising but highly uneven scientific collaborators. Agents perform best on structured data analysis and established preprocessing pipelines where procedures and evaluation criteria are explicitly specified. However, performance drops sharply in open-ended model optimization; for example, agents routinely solve single-objective molecular designs but struggle with multi-constraint optimization tasks. In model selection, agents exhibit severe conservative convergence by defaulting to popular, documented methods (such as Harmony for batch correction or Leiden for clustering) rather than choosing optimal approaches for specific data constraints. Domain-specialist agents generally offer advantages in tasks requiring curated tools and specialized workflows, but generalist agents with robust code-execution loops frequently match or exceed specialist performance. Critically, agents perform poorly on validity checks: they often accept flawed prompts, execute calculations on contradictory data without noticing anomalies, or hallucinate non-existent programming interfaces rather than refusing invalid tasks.
These findings indicate that while AI agents can reliably automate routine, well-defined data processing pipelines, they lack autonomous scientific judgment, robust state tracking, and spontaneous error correction. Relying on agents for critical research decisions without rigorous verification poses substantial risks, including silent propagation of data errors, flawed causal conclusions, and wasted experimental resources. The results challenge the assumption that expanding domain-specific tool libraries alone will produce competent AI scientists; procedural execution ability does not equate to sound scientific reasoning.
The article recommends that organizations deploy AI agents primarily for bounded, verifiable data analysis workflows while maintaining mandatory human oversight for model selection, optimization, and scientific validation. Developers should enhance future agents with explicit runtime verification mechanisms, such as checking installed software interfaces, validating data context before execution, enforcing step-wise state persistence, and integrating built-in refusal protocols for ill-posed requests. Future benchmarking efforts should expand into additional scientific disciplines, such as physics and materials science, to further assess agent reliability across diverse research environments.
- Paper: MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation, Qian Huang et al. (2024). It introduces MLAgentBench, establishing foundational protocols and execution environments for evaluating autonomous agents on interactive, multi-step scientific and machine learning experimentation.
- Paper: SciAgent: Tool-augmented Language Models for Scientific Reasoning, Yubo Ma et al. (2024). It presents SciAgent and SciToolBench, providing essential background on how tool augmentation and procedural execution facilitate multi-domain scientific reasoning in language models.
- Paper: ScienceWorld: Is your Agent Smarter than a 5th Grader?, Ruoyao Wang et al. (2022). It establishes ScienceWorld, demonstrating the critical need for interactive simulation environments rather than static QA tests to evaluate procedural scientific reasoning.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). It introduces stepwise trajectory inspection and agentic judges, laying crucial groundwork for interactive, multi-step verification of agentic systems across complex workflows.
- Paper: SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models, Xiaoxuan Wang et al. (2024). It provides SciBench for evaluating complex collegiate-level scientific problem-solving, illustrating the baseline reasoning challenges that interactive agent benchmarks seek to expand upon.
- Paper: A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery, Yu Zhang et al. (2024). It offers a comprehensive survey of scientific language models and modalities across disciplines, contextualizing the core scientific domains and tasks addressed by multi-scale agent benchmarks.
- Paper: ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models, Jinheon Baek et al. (2025). It demonstrates iterative research idea generation and peer-review evaluation, highlighting the open-ended scientific reasoning workflows that SciAgentArena aims to benchmark systematically.
- Paper: LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery, Pingchuan Ma et al. (2024). It develops a bilevel optimization framework combining LLMs with scientific simulations, providing a key technical paradigm for evaluating physical scientific discovery.
- Paper: AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery, Guiyao Tie et al. (2026). It synthesizes workflow automation and defines an autonomy spectrum for AI-driven scientific discovery, providing a broad structural framework that extends beyond benchmark-level task evaluations.
- Paper: Accelerating Scientific Research with Gemini in the Real-World, Samuel Schmidgall et al. (2026). It presents Gemini Co-Scientist, demonstrating the end-to-end deployment of multi-agent scientific systems across wet-lab and computational domains in real-world research settings.
- Paper: Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models, Kevin Murphy (2026). It implements an active Bayesian experiment design framework for discovering mechanistic world models, directly addressing the open-ended hypothesis testing bottlenecks identified in scientific agent benchmarking.
- Paper: Innovator-VL: A Multimodal Large Language Model for Scientific Discovery, Zichen Wen et al. (2026). It develops Innovator-VL to tackle multimodal scientific reasoning and discovery, extending foundation model capabilities to specialized scientific visual tasks.
- Paper: Idea2Story: An Automated Pipeline for Transforming Research Concepts into Complete Scientific Narratives, Tengyue Xu et al. (2026). It introduces a structured knowledge-graph framework to transform informal research ideas into scientific narratives, operationalizing automated scientific ideation.
- Paper: PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing, Yiwen Song et al. (2026). It extends autonomous research workflows to the automated drafting and refinement of full scientific papers through multi-agent collaboration.
- Paper: Code as Agent Harness, Xuying Ning et al. (2026). It formalizes executable code harnesses to overcome reliability and long-horizon execution failures in scientific and autonomous agent workflows.
- Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). It provides a diagnostic guardrail framework to monitor, detect, and mitigate failure modes and security risks in autonomous agent tool execution.
