HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Qianchu LiuSheng ZhangGuanghui QinJeya Maria Jose ValanarasuMaximilian RokussMing LuTimothy OssowskiJuan Manuel Zambrano ChavesCliff WongPeniel Argaw
Introduces HealthAgentBench, an evaluation suite of 54 end-to-end clinical tasks across seven multimodal environments that exposes critical reasoning bottlenecks where top frontier AI agents achieve only a 42% success rate.
The rapid advancement of artificial intelligence (AI) has shifted focus from simple conversational models to autonomous agents capable of multi-step planning, tool use, and environment interaction. While existing healthcare AI evaluations predominantly rely on static question-answering or narrow simulated dialogues, real-world clinical practice requires end-to-end reasoning over massive, heterogeneous datasets—such as three-dimensional computed tomography (CT) scans, gigapixel pathology slides, and extensive electronic health records (EHR). Consequently, there is an urgent need for realistic, holistic benchmark suites to measure how frontier AI agents handle comprehensive clinical workflows.
The article introduces HealthAgentBench, a unified suite of 54 agentic healthcare tasks spanning seven distinct clinical categories. It evaluates the autonomous capabilities of frontier AI agents across critical workflow stages—including data management, diagnostics, research, and treatment planning—measuring task success against rigorous, expert-reviewed clinical baselines.
To construct this benchmark, the authors packaged each task within a containerized terminal environment using real patient artifacts, such as images, tabular EHR data, and clinical trial documents. Agents receive minimal instructions and must autonomously formulate strategies, execute code, navigate files, and produce structured outputs without human scaffolding. The evaluation measures binary success criteria across 10 frontier agents from the GPT and Claude model families using three different harnesses, evaluating performance, dollar cost, and wall-clock execution time over three repeated trials per task.
The findings reveal that frontier agents currently struggle with end-to-end healthcare tasks, leaving the benchmark far from saturation. The top-performing system, Codex GPT-5.5, achieves an overall task success rate of approximately 42%, while older or smaller models score as low as 16%. Agents perform strongly on tabular research workflows—such as building machine learning prediction pipelines and configuring data conversion scripts—frequently matching or exceeding baseline models. However, severe bottlenecks emerge in medical imaging (CT, X-ray, and pathology), where average success drops to roughly 17%, and in complex database retrieval tasks that require navigating large search spaces and multi-step reasoning. Additionally, the GPT model family consistently defines the cost-efficiency frontier, outperforming Claude models on imaging while using fewer tokens and costing less per trial.
These results demonstrate that while current AI agents are already capable of automated data engineering and predictive modeling on structured clinical data, they remain far from clinically viable for direct diagnostic imaging and complex unstructured data retrieval. Deploying existing general-purpose agents directly into multimodal diagnostic pipelines presents significant safety and performance risks. Successful models distinguish themselves by dedicating substantial effort to verifying outputs, orienting within environments, and dynamically generating intermediate visual representations.
Organizations developing or integrating healthcare agents should focus near-term automation on structured data pipelines, predictive modeling, and extract-transform-load (ETL) workflows where agent reliability is highest. For complex clinical diagnostics and large search spaces, developers should build task-specific vision backends, implement decomposition strategies, and deploy lightweight triage architectures rather than relying solely on larger general-purpose agents. Further research should expand the benchmark to include additional modalities, interactive interfaces, and multi-agent coordination environments before strong deployment decisions are made in safety-critical clinical settings.
The primary limitations of the evaluation stem from its terminal-based execution setup, which models backend technical workflows rather than clinician-facing user interfaces. Furthermore, while the 54 tasks provide diverse, high-signal coverage across major modalities, they represent a sampled subset rather than the entirety of clinical practice. Nevertheless, because all tasks utilize rigorous anti-cheating mechanisms, objective verification, and hand-reviewed gold standards, stakeholders can have high confidence in the relative capability gaps and structural bottlenecks identified across evaluated agent families.
- Paper: MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation, Qian Huang et al. (2024). Provides the foundational framework for evaluating autonomous language model agents performing multi-step iterative machine learning experimentation in code and file workspaces.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). Establishes standard methodologies for evaluating autonomous agents using realistic sandboxed environments and end-to-end task execution metrics.
- Paper: Large language models encode clinical knowledge, Karan Singhal et al. (2022). Introduces foundational medical question answering and clinical reasoning benchmarks for evaluating large language models in healthcare.
- Paper: EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images, Seongsu Bae et al. (2023). Pioneers the multimodal clinical evaluation combining tabular electronic health records with radiological imaging datasets.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Surveys the core architectural principles, tool-use mechanisms, and evaluation paradigms for LLM-based autonomous agents.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). Develops intermediate trajectory validation and agentic grading methods essential for evaluating multi-step agent execution.
- Paper: MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification, Jiancheng Yang et al. (2021). Supplies standardized 2D and 3D biomedical imaging datasets that underpin clinical vision evaluation pipelines.
- Paper: Benchmarking AI Agents for Addressing Scientific Challenges Across Scales, Tianyu Liu et al. (2026). Expands agentic evaluation across broader scientific research domains, including EHR modeling and multi-step omics pipelines.
- Paper: MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis, Akshat Sanghvi et al. (2026). Explores interactive multi-agent clinical consultation workflows to improve diagnostic reasoning beyond single-agent execution.
- Paper: 3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis, Ziyue Wang et al. (2026). Specializes agentic reasoning techniques to tackle complex 3D medical imaging analysis, a modality noted as challenging in HealthAgentBench.
- Paper: Can Revealed Preferences Clarify LLM Alignment and Steering?, Khurram Yamin et al. (2026). Investigates revealed preferences and decision-theoretic risk tradeoffs in clinical AI decision support systems.
- Paper: SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations, Shuaiqi Wang et al. (2026). Addresses the evaluation of synthetic data quality for validating multi-turn, tool-calling agents in sensitive domains.
- Paper: MemGym: a Long-Horizon Memory Environment for LLM Agents, Wujiang Xu et al. (2026). Examines long-horizon memory management environments to isolate context retention errors in extended autonomous agent workflows.
