HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

Qianchu LiuSheng ZhangGuanghui QinJeya Maria Jose ValanarasuMaximilian RokussMing LuTimothy OssowskiJuan Manuel Zambrano ChavesCliff WongPeniel Argaw

article2026arXiv7 citations

Introduces HealthAgentBench, an evaluation suite of 54 end-to-end clinical tasks across seven multimodal environments that exposes critical reasoning bottlenecks where top frontier AI agents achieve only a 42% success rate.

Listen

The rapid advancement of artificial intelligence (AI) has shifted focus from simple conversational models to autonomous agents capable of multi-step planning, tool use, and environment interaction. While existing healthcare AI evaluations predominantly rely on static question-answering or narrow simulated dialogues, real-world clinical practice requires end-to-end reasoning over massive, heterogeneous datasets—such as three-dimensional computed tomography (CT) scans, gigapixel pathology slides, and extensive electronic health records (EHR). Consequently, there is an urgent need for realistic, holistic benchmark suites to measure how frontier AI agents handle comprehensive clinical workflows.

The article introduces HealthAgentBench, a unified suite of 54 agentic healthcare tasks spanning seven distinct clinical categories. It evaluates the autonomous capabilities of frontier AI agents across critical workflow stages—including data management, diagnostics, research, and treatment planning—measuring task success against rigorous, expert-reviewed clinical baselines.

To construct this benchmark, the authors packaged each task within a containerized terminal environment using real patient artifacts, such as images, tabular EHR data, and clinical trial documents. Agents receive minimal instructions and must autonomously formulate strategies, execute code, navigate files, and produce structured outputs without human scaffolding. The evaluation measures binary success criteria across 10 frontier agents from the GPT and Claude model families using three different harnesses, evaluating performance, dollar cost, and wall-clock execution time over three repeated trials per task.

The findings reveal that frontier agents currently struggle with end-to-end healthcare tasks, leaving the benchmark far from saturation. The top-performing system, Codex GPT-5.5, achieves an overall task success rate of approximately 42%, while older or smaller models score as low as 16%. Agents perform strongly on tabular research workflows—such as building machine learning prediction pipelines and configuring data conversion scripts—frequently matching or exceeding baseline models. However, severe bottlenecks emerge in medical imaging (CT, X-ray, and pathology), where average success drops to roughly 17%, and in complex database retrieval tasks that require navigating large search spaces and multi-step reasoning. Additionally, the GPT model family consistently defines the cost-efficiency frontier, outperforming Claude models on imaging while using fewer tokens and costing less per trial.

These results demonstrate that while current AI agents are already capable of automated data engineering and predictive modeling on structured clinical data, they remain far from clinically viable for direct diagnostic imaging and complex unstructured data retrieval. Deploying existing general-purpose agents directly into multimodal diagnostic pipelines presents significant safety and performance risks. Successful models distinguish themselves by dedicating substantial effort to verifying outputs, orienting within environments, and dynamically generating intermediate visual representations.

Organizations developing or integrating healthcare agents should focus near-term automation on structured data pipelines, predictive modeling, and extract-transform-load (ETL) workflows where agent reliability is highest. For complex clinical diagnostics and large search spaces, developers should build task-specific vision backends, implement decomposition strategies, and deploy lightweight triage architectures rather than relying solely on larger general-purpose agents. Further research should expand the benchmark to include additional modalities, interactive interfaces, and multi-agent coordination environments before strong deployment decisions are made in safety-critical clinical settings.

The primary limitations of the evaluation stem from its terminal-based execution setup, which models backend technical workflows rather than clinician-facing user interfaces. Furthermore, while the 54 tasks provide diverse, high-signal coverage across major modalities, they represent a sampled subset rather than the entirety of clinical practice. Nevertheless, because all tasks utilize rigorous anti-cheating mechanisms, objective verification, and hand-reviewed gold standards, stakeholders can have high confidence in the relative capability gaps and structural bottlenecks identified across evaluated agent families.

Cover for HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

Abstract

As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment. The benchmark suite spans diverse workflows throughout the patient journey and a broad range of modalities. Each task is designed to replicate an end-to-end clinical workflow: given minimal instructions, an agent must explore raw healthcare data, operate within a complex environment, and execute multi-step solutions that go beyond naive prompting. A final task success rate is reported to provide a single, interpretable metric for HealthAgentBench overall performance for each agent. Evaluating frontier agents on HealthAgentBench, we find that overall task success rate remains low, underscoring the difficulty of the suite. The strongest and the most cost effective agent, Codex GPT-5.5, achieves only approximately 42% success rate. Beyond aggregate performance, HealthAgentBench reveals nuanced strengths and weaknesses across task categories. Frontier agents show promise in automatically developing research modeling pipelines over EHR data, but medical imaging remains especially challenging, particularly for Claude Code models, while Codex GPT-5.5 shows emerging capability. Tasks that combine large search spaces with compositional reasoning requirements remain difficult for all current agents. Together, these results suggest that HealthAgentBench provides a challenging and realistic benchmark with substantial room for future progress. We release our benchmark at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Task Creation
  • 3.1 Task selection criteria.
  • 3.2 Sourcing tasks.
  • 3.3 Task design principles.
  • 3.4 Post-creation checks.
  • 4 Baselines and Results
  • 4.1 Overall Performance
  • 4.2 Breaking Down Performance by Task Categories
  • 5 Conclusion
  • References
  • A Benchmark Comparison
  • B The HealthAgentBench Benchmark Suite
  • C Cost vs. performance trade-off
  • D Per-Category Performance Breakdowns
  • D.1 Per-category results: success rate, time, and cost
  • D.2 Task-specific scores
  • D.3 EHR Data Quality Auditing: the cost of a large search space
  • D.4 EHR Event Modelling: AUROC vs. human engineered baseline
  • E HealthAgentBench Task Categories
  • E.1 EHR Format Conversion: EHR ETL pipeline customization
  • E.2 X-ray Report Correction: chest X-ray report correction
  • E.3 Clinical Trial Matching: patient–trial eligibility matching
  • E.4 CT Abnormality Classification: chest CT interpretation
  • E.5 Pathology Tumor Area Selection: pathology whole-slide reasoning
  • E.6 ehrshot: longitudinal clinical event prediction
  • E.7 EHR Data Quality Auditing: EHR data-quality auditing
  • F Trajectory Analysis for Successful Trials
  • F.1 EHR Format Conversion: minimal scoped config override
  • F.2 Clinical Trial Matching: triage, then parallel eligibility adjudication
  • F.3 EHR Event Modelling: leakage-safe ML pipeline
  • F.4 CT Abnormality Classification: multi-window rendering and noise control
  • F.5 Pathology Tumor Area Selection: multi-resolution tiling and morphology-driven calls
  • F.6 X-ray Report Correction: read priors, let the image decide
  • F.7 EHR Data Quality Auditing: scripted detection with iterative recall recovery
  • F.8 Common themes

Knowls

  1. Knowl 1 — HealthAgentBench Benchmark Framework and Task Taxonomy

    model/method

    HealthAgentBench is an evaluation suite of 54 agentic healthcare tasks distributed across 7 task categories, designed to evaluate autonomous AI agents operating on raw, heterogeneous clinical data within interactive terminal environments.

    Each task is encapsulated in an isolated Docker container managed via the Harbor execution framework. Agents are provided with raw clinical artifacts and minimal natural language instructions, and interact with the environment via command-line execution (e.g., executing Python scripts, installing software packages, querying databases, reading and manipulating multimodal files) before writing a structured submission file to /workspace/submission/.

    The benchmark covers 5 healthcare data modalities and 4 clinical workflow stages across 7 distinct task categories:

    1. EHR Format Conversion (1 task): Data Management workflow; converts raw electronic health record (EHR) data from the MIMIC-IV demo into the standardized Medical Event Data Standard (MEDS) parquet format by configuring and executing an extract-transform-load (ETL) pipeline.
    2. X-ray Report Correction (10 tasks): Diagnosis workflow; inspects longitudinal chest radiographs (2D DICOM/JPG) alongside historical clinical reports to identify and correct clinically significant factual errors in a corrupted draft findings section.
    3. Clinical Trial Matching (9 tasks): Treatment Planning workflow; analyzes an unstructured patient admission note and triages a cohort of 300--450 candidate ClinicalTrials.gov XML protocol documents to produce a confidence-ordered list of eligible trial identifiers.
    4. CT Abnormality Classification (10 tasks): Diagnosis workflow; evaluates a 3D non-contrast chest computed tomography (CT) volume (200--700 slices in NIfTI format from CT-RATE) and produces binary presence/absence classifications for 4--12 specified clinical findings.
    5. Pathology Tumor Area Selection (10 tasks): Research and Diagnosis workflow; processes a gigapixel whole-slide lymph node pathology image (WSI from CAMELYON16) and identifies which tiles on a 256×256256 \times 256 grid (at 16×16\times downsampling, where each tile covers 4096×40964096 \times 4096 level-0 pixels) contain metastatic breast carcinoma covering ≥20%\ge 20\% of the tile area.
    6. EHR Event Modelling (6 tasks): Research workflow; learns a predictive machine learning model from labeled training and validation event logs (Stanford STARR cohort via EHRSHOT) and outputs per-patient onset probabilities on an unlabeled held-out test cohort across 6 diagnosis targets.
    7. EHR Data Quality Auditing (8 tasks): Data Management workflow; inspects 8 corrupted tabular EHR tables (>800,000 rows derived from MIMIC-IV) and flags specific row IDs exhibiting synthetic data quality violations (impossible values, demographic conflicts, or cross-table duplicate/inconsistent records).
  2. Knowl 2 — HealthAgentBench Binary Task Success and Scoring Criteria

    definition

    Every task in HealthAgentBench is scored using a binary outcome (reward R∈{0,1}R \in \{0, 1\}), where a score of 1 indicates that the agent's final submission satisfies a rigorous task-specific success threshold reflecting near-expert human performance, and 0 indicates failure. The aggregate metric for comparing agents is the task success rate, defined as the proportion of evaluated trials that earn a reward of 1.

    The criteria across the 7 task categories are defined as follows:

    • EHR Format Conversion: R=1R = 1 if and only if all verifier checks pass: the default configuration file is preserved unmodified, the custom YAML exists and parses, all 7 required code-prefix families (INSURANCE//, LANGUAGE//, MARITAL_STATUS//, RACE//, OMR//, HOSP_LAB//, ICU_CHARTEVENT//) are present in the output metadata, and per-shard row counts and canonical content hashes match the gold standard.
    • X-ray Report Correction: R=1R = 1 if and only if a 5-vote majority judge panel using the CheXprompt framework (with GPT-5.4 as judge) casts ≥3/5\ge 3/5 votes asserting zero clinically significant errors in the corrected findings section compared to expert-reviewed gold reports.
    • Clinical Trial Matching: R=1R = 1 if and only if the agent achieves a top-50 recall of Recall@50=1.0\text{Recall}@50 = 1.0, meaning that every physician-annotated gold eligible clinical trial is included within the agent's top 50 ranked predictions (selected from a candidate pool of 300--450 trials).
    • CT Abnormality Classification: R=1R = 1 if and only if per-label exact match accuracy is 1.01.0, requiring every individual clinical abnormality label (4--12 labels per volume) to be classified correctly as present or absent.
    • Pathology Tumor Area Selection: R=1R = 1 if and only if the tile-level F1F_1 score between the predicted set of tumor grid tiles and the hidden ground-truth tumor mask satisfies F1≥0.90F_1 \ge 0.90.
    • EHR Event Modelling: R=1R = 1 if and only if the Area Under the Receiver Operating Characteristic curve (AUROC) of the agent's predictions on the unlabeled test cohort meets or exceeds the AUROC of a human-engineered count-based LightGBM baseline (AUROC≥AUROCcount+GBM\text{AUROC} \ge \text{AUROC}_{\text{count+GBM}}).
    • EHR Data Quality Auditing: R=1R = 1 if and only if the agent achieves cluster-level recall of 1.01.0 (every injected synthetic error cluster has at least one correctly identified row member) while maintaining a precision floor of Precision≥0.01\text{Precision} \ge 0.01 to prevent unconstrained over-flagging.
  3. Knowl 3 — HealthAgentBench Task Creation and Anti-Leakage Design Principles

    model/method

    HealthAgentBench constructs and validates tasks using a four-stage lifecycle designed to guarantee realism, verifiability, and resistance to evaluation shortcuts:

    1. Candidate Screening: Tasks must necessitate multi-step agentic execution (preventing direct zero-shot prompt completion due to massive input dimensions like 3D volumes, gigapixel slides, or multi-table databases), represent real clinical workflows, span varied modalities, and have an expected random-guessing success rate below 10%10\%.
    2. Task Sourcing and Transformation: Tasks are either adapted from existing benchmarks by stripping human scaffolding (e.g., adapting EHRSHOT from a fixed feature baseline into an autonomous end-to-end ML modeling challenge) or constructed by applying deterministic, seeded perturbation recipes to real patient databases (e.g., injecting impossible clinical measurements, demographic mismatches, and duplicate records into MIMIC-IV).
    3. Standardized Environment Design: Tasks are deployed using five core principles:
      • Versatile Terminal Environment: Execution within Docker containers orchestrated by Harbor, providing agents full terminal access to Python, shell tools, and the internet for downloading standard libraries.
      • Anti-Cheat by Construction: Separation of agent runtime from evaluation via a two-service Docker topology. Credentials, gold ground-truth labels, and scoring scripts are mounted exclusively into a verifier container and are inaccessible to the agent. Agent-visible filenames and identifiers are anonymized (e.g., study_NN, case_NN), upstream database names are scrubbed, and web browsing is blocked during runs.
      • Minimalist Instructions: Agents receive 2--3 sentence objective descriptions and target output formats without predefined procedural recipes or tool prompts.
      • Representative Sampling: 5 to 15 representative cases are selected per category across varied disease classes, patient demographics, and difficulty tiers.
    4. Post-Creation Verification Checks: Before inclusion, candidate tasks undergo sweeps with frontier agents. Execution trajectories are manually and automatically audited to eliminate accidental shortcut leaks (such as leaked evaluation scripts) or task ambiguities with multiple conflicting valid solutions. Trivial tasks with 100%100\% zero-effort completion across all models are removed.
  4. Knowl 4 — Overall Performance and Cost Pareto Frontier of Frontier AI Agents

    data/table

    Evaluating 10 frontier LLM agent systems across 3 harnesses (Codex, Claude Code, and Copilot-CLI) with maximum reasoning effort and disabled web browsing over 3 attempts per task (162 trials total per agent) demonstrates that HealthAgentBench remains challenging and far from saturation. The top-performing system, Codex GPT-5.5, achieves an aggregate task success rate of 42%42\%.

    Agent Success Rate (%) Mean Cost / Task (USD) Mean Time / Task (min)
    Codex GPT-5.5 42 2.8 15
    Copilot Opus 4.8 36 3.1 18
    Copilot GPT-5.5 35 2.6 18
    Claude Code Opus 4.8 32 4.0 16
    Codex GPT-5.4 28 1.3 18
    Claude Code Opus 4.7 27 4.8 16
    Codex GPT-5.3 22 1.0 14
    Claude Code Opus 4.6 19 4.1 22
    Claude Code Sonnet 4.6 17 2.9 24
    Codex GPT-5.4-mini 16 0.6 16

    The cost-performance Pareto frontier is defined entirely by the GPT-5 model family (Codex GPT-5.4-mini, Codex GPT-5.3, Codex GPT-5.4, Copilot GPT-5.5, and Codex GPT-5.5). Across all performance tiers, every Claude Code configuration is either matched or outperformed by a GPT-5 model at lower financial cost. Claude Code agents incur higher costs (up to $4.8 per task for Opus 4.7) and longer execution durations (up to 24 minutes per task for Sonnet 4.6) due to verbose multi-agent subprocess spawning and emitting approximately twice as many tokens per trial (44,000 vs. 20,000 output tokens) compared to Codex runs.

  5. Knowl 5 — Per-Category Agent Success Rates and the Modality Gap

    data/table

    Agent performance across HealthAgentBench reveals a stark divergence between structured/textual tasks and medical imaging tasks. Across all evaluated Codex and Claude Code agent runs, the mean success rate on text-based tasks (EHR Format Conversion, EHR Event Modelling, EHR Data Quality Auditing, Clinical Trial Matching) is 49%49\%, whereas the mean success rate on medical imaging tasks (X-ray Report Correction, CT Abnormality Classification, Pathology Tumor Area Selection) drops to 17%17\%.

    Agent X-ray CT Tumor Trial DQ EHR MEDS
    Codex GPT-5.5 33 33 40 52 25 72 100
    Codex GPT-5.4 40 30 7 18 12 61 100
    Codex GPT-5.4-mini 10 23 0 18 4 39 100
    Codex GPT-5.3 20 10 17 18 12 61 100
    Claude Code Opus 4.8 17 17 20 48 33 67 100
    Claude Code Opus 4.7 13 13 10 26 33 78 100
    Claude Code Opus 4.6 10 0 17 22 42 33 0
    Claude Code Sonnet 4.6 10 3 10 11 12 61 100
    Copilot GPT-5.5 27 23 13 44 42 67 100
    Copilot Opus 4.8 20 17 17 67 33 72 100

    Codex GPT models significantly outperform Claude Code models on perception-heavy imaging tasks, averaging 22%22\% success compared to 12%12\% for Claude Code agents. Codex GPT-5.5 attains the highest imaging performance (35%35\% averaged across imaging tasks, leading Pathology Tumor Area Selection at 40%40\% and matching top performance on CT Abnormality Classification at 33%33\%). Conversely, text task performance between the two families is comparable (50%50\% for Codex vs. 48%48\% for Claude Code).

  6. Knowl 6 — Frontier Agent Capabilities in Autonomous EHR Pipeline Construction

    empirical result

    Frontier agents demonstrate strong capabilities in automating medical data engineering and predictive machine learning research workflows over structured electronic health records:

    • ETL Pipeline Customization (EHR Format Conversion): 9 out of 10 evaluated agents achieve a 100%100\% success rate, correctly parsing an unfamiliar open-source ETL repository (MIMIC_IV_MEDS), modifying the configuration to extract admission demographics as timestamped events and assigning specified domain code prefixes, executing the pipeline using uv, and emitting canonical parquet shards. Only Claude Code Opus 4.6 fails (0%0\% success) due to consistently dropping an admission demographic field.
    • End-to-End Clinical Event Prediction (EHR Event Modelling): When tasked with training an ML model from multi-gigabyte raw event logs (Stanford STARR cohort with 6,739 patients and 41.6 million observations) within a 1-hour time budget to predict 6 new-onset diagnoses, leading agents achieve high success rates against the human-engineered count+GBM baseline (78%78\% for Claude Code Opus 4.7, 72%72\% for Codex GPT-5.5, and 72%72\% for Copilot Opus 4.8).
    • Competitive Modeling Performance: Trajectory analysis indicates agents formulate leakage-safe data loading pipelines (using tools such as Polars lazy joins and PyArrow to filter events strictly prior to prediction time TfirstT_{\text{first}}) and train gradient-boosted trees (LightGBM/XGBoost). The resulting test AUROC values (ranging from 0.720.72 to 0.770.77 across agents) match or exceed the published deep learning CLMBR neural network foundation model baseline on the majority of disease prediction targets.
  7. Knowl 7 — Search Space and Compositionality Bottlenecks in EHR Quality Auditing

    empirical result

    EHR Data Quality Auditing (flagging synthetic errors across 8 relational tables containing over 800,000 rows) is one of the benchmark's most challenging categories, with no agent exceeding a 42%42\% success rate (Claude Code Opus 4.6 reaches 42%42\%, Copilot GPT-5.5 achieves 42%42\%, and all Codex models other than GPT-5.5 achieve ≤12.5%\le 12.5\%).

    Ablation experiments isolate two primary drivers of task failure:

    1. Search Space Size Bottleneck: Providing agents with explicit search clues (disclosing the injected error sub-type and naming the target tables to inspect) substantially improves cluster recall across all tested agents compared to base instructions without clues (e.g., Claude Code Opus 4.8 increases recall by +0.24+0.24 to 0.970.97, Codex GPT-5.3 increases by +0.25+0.25, and Codex GPT-5.5 increases by +0.15+0.15). This demonstrates that navigating unconstrained relational database schemas is a major point of failure.
    2. Task Compositionality Bottleneck: In the combined task where agents must simultaneously identify all three error categories (impossible clinical values, demographic contradictions, and duplicate/inconsistent records) in a single run, agent recall degrades significantly compared to evaluating the same error types in three isolated passes (e.g., Codex GPT-5.4-mini drops recall by 0.340.34, Claude Code Sonnet 4.6 drops by 0.290.29, and Claude Code Opus 4.6 drops by 0.290.29). Agents struggle to maintain detection sensitivity across multiple distinct error hypotheses simultaneously.
  8. Knowl 8 — Triage and Parallel Subagent Orchestration in Clinical Trial Matching

    empirical result

    In the Clinical Trial Matching task category, agents evaluate an unstructured free-text patient admission note against 300--450 candidate trial XML files to retrieve all eligible trials with perfect top-50 recall (Recall@50=1.0\text{Recall}@50 = 1.0).

    A substantial performance gap emerges between leading systems and weaker models: Copilot CLI running Opus 4.8 achieves the highest task success rate (67%67\%, 18/2718/27 trials), followed by Codex GPT-5.5 (52%52\%, 14/2714/27), Claude Code Opus 4.8 (48%48\%, 13/2713/27), and Copilot GPT-5.5 (44%44\%, 12/2712/27), while all remaining agents score ≤26%\le 26\%.

    Trajectory analysis of the top-performing Copilot CLI Opus 4.8 system reveals a successful three-stage execution strategy:

    1. Scripted Initial Triage: The primary agent writes an inline Python script to parse XML fields, enforcing hard demographic constraints (e.g., age, sex) and keyword-based cardiology relevance scoring to prune the search space into candidate batches.
    2. Parallel Subagent Adjudication: The orchestrator spawns 5 background subagents prompted as eligibility adjudicators, each assigned a disjoint batch of trials and equipped with a structured patient profile explicitly listing negative clinical findings (e.g., absent conditions or lack of previous implants) to evaluate inclusion and exclusion criteria concurrently.
    3. Central Re-Verification: The lead agent collects candidate verdicts and performs deep criterion-level verification of borderline cases before submitting the final confidence-ranked NCT-ID list.

    While strong agents attain high continuous recall (Recall@50≈0.82−−0.93\text{Recall}@50 \approx 0.82--0.93), achieving the binary pass standard of 1.01.0 recall requires exhaustive exclusion-criterion comprehension to avoid omitting edge-case trials.

  9. Knowl 9 — Tool-Driven View Synthesis and Perception Dynamics in Medical Imaging Tasks

    empirical result

    Because volumetric CT scans (200--700 slices) and gigapixel pathology slides (roughly 100,000×100,000100,000 \times 100,000 pixels) cannot be ingested directly into LLM context windows, successful agent trajectories rely on autonomous tool-mediated view synthesis and hierarchical visual inspection:

    • Pathology WSI Inspection Strategy: In Pathology Tumor Area Selection, top agent Codex GPT-5.5 (40%40\% success, tile-F1=0.845F_1 = 0.845) uses OpenSlide and tifffile to build a multi-resolution image pyramid. It inspects low-power overviews (level-6) to locate tissue fragments, generates intermediate contact sheets (level-4 and level-2), and distinguishes metastatic carcinoma from benign lymphoid tissue using morphological features rather than simple color/stain thresholds. It conducts boundary tile refinement at full resolution alongside an 'outside pass' over unflagged high-tissue regions to recover missed tumor foci.
    • CT Volumetric Slicing Strategy: In CT Abnormality Classification, Codex GPT-5.5 (33%33\% success) writes Python scripts utilizing nibabel to render orthogonal reformations (axial, coronal, sagittal) and dual-window contact sheets (lung and mediastinal windows), generating small z-slab mean projections and Hounsfield Unit (HU) density measurements on candidate fluid collections.
    • Primary Vision Failure Modes: Agents frequently fail via two symmetric error modes: under-calling (failing to identify subtle, low-contrast pathologies such as paraseptal emphysema or small nodules) and over-calling (flagging normal anatomical variations or tissue artifacts as abnormal findings, which violates the 1.01.0 exact-match accuracy requirement on CT tasks).
  10. Knowl 10 — Limitations of HealthAgentBench

    limitation

    The authors identify several limitations of the HealthAgentBench evaluation suite:

    1. Non-Exhaustive Healthcare Coverage: While the benchmark incorporates 7 task categories across 5 modalities and 4 workflow stages, it does not encompass all dimensions of healthcare agent capabilities, omitting real-time clinician-patient spoken dialogue, bedside procedural assistance, robotic control, and dynamic multi-party hospital operational coordination.
    2. Terminal vs. Clinical User Interface Gap: Tasks are executed inside text-based command-line Docker environments. While optimal for benchmarking autonomous coding and API tool use, this does not reflect the clinician-facing graphical user interfaces (GUIs), electronic medical record UI widgets, or hospital communication platforms encountered during real-world clinical deployment.
    3. Dependence on General-Purpose Vision Capabilities: Current general-purpose multimodal LLMs lack domain-specialized radiology and pathology visual backends, creating a major performance bottleneck on volumetric CT, whole-slide imaging, and subtle radiograph interpretation that general scaling alone may not resolve without specialized vision tool integration.

Coverage note — No substantial contributed material was omitted. The knowls cover the full benchmark taxonomy, formal success criteria, task creation workflow, overall and per-category empirical results, cost-performance analyses, detailed task ablations, trajectory analyses, and stated limitations.

References

  1. 1.Mohammad Almansoori, Komal Kumar, and Hisham Cholakkal. Self-evolving multi-agent simulations for realistic clinical interactions, 2025. URL https://arxiv.org/abs/2503.22678.
  2. 2.Anthropic. Claude code documentation, 2025. URL https://code.claude.com/docs.
  3. 3.Anthropic. Claude opus system card, 2026. URL https://www.anthropic.com/claude/opus.
  4. 4.Anthropic. Claude sonnet system card, 2026. URL https://www.anthropic.com/claude/sonnet.
  5. 5.Rahul K Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, et al. Healthbench: Evaluating large language models towards improved human health. arXiv preprint arXiv:2505.08775, 2025. URL https://doi.org/10.48550/arXiv.2505.08775.
  6. 6.Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, Johannes Heidecke, and Karan Singhal. HealthBench: Evaluating Large Language Models Towards Improved Human Health, 2025. URL https://doi.org/10.48550/arXiv.2505.08775.
  7. 7.Suhana Bedi, Ryan Welch, Ethan Steinberg, Michael Wornow, Taeil Matthew Kim, Haroun Ahmed, Peter Sterling, Bravim Purohit, Qurat Akram, Angelic Acosta, et al. Healthadminbench: Evaluating computer-use agents on healthcare administration tasks. arXiv preprint arXiv:2604.09937, 2026. URL https://doi.org/10.48550/arXiv.2604.09937.
  8. 8.Zhihao Fan, Jialong Tang, Wei Chen, Siyuan Wang, Zhongyu Wei, Jun Xi, Fei Huang, and Jingren Zhou. Ai hospital: Benchmarking large language models in a multi-agent medical interaction simulator, 2024. URL https://arxiv.org/abs/2402.09742.
  9. 9.Yifan Gao, Haoyue Li, Feng Yuan, Xin Gao, Weiran Huang, and Xiaosong Wang. Camyla: Scaling autonomous research in medical image segmentation, 2026. URL https://arxiv.org/abs/2604.10696.
  10. 10.GitHub. Github copilot cli, 2025. URL https://github.com/features/copilot/cli.
  11. 11.Ibrahim Ethem Hamamci, Sezgin Er, Chenyu Wang, Furkan Almas, Ayse Gulnihan Simsek, Sevval Nil Esirgun, Irem Dogan, Omer Faruk Durugol, Benjamin Hou, Suprosanna Shit, Weicheng Dai, Murong Xu, Hadrien Reynaud, Muhammed Furkan Dasdelen, Bastian Wittmann, Tamaz Amiranashvili, Enis Simsar, Mehmet Simsar, Emine Bensu Erdemir, Abdullah Alanbay, Anjany Sekuboyina, Berkan Lafci, Ahmet Kaplan, Zhiyong Lu, Malgorzata Polacin, Bernhard Kainz, Christian Bluethgen, Kayhan Batmanghelich, Mehmet Kemal Ozdemir, and Bjoern Menze. Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography. Nature Biomedical Engineering, 2026. URL https://doi.org/10.1038/s41551-025-01599-y.
  12. 12.Harbor Framework Team. Harbor: A framework for evaluating and optimizing agents and models in container environments, January 2026. URL https://github.com/laude-institute/harbor.
  13. 13.Rebecca Soskin Hicks, Mikhail Trofimov, Dominick Lim, Rahul K. Arora, Foivos Tsimpourlas, Preston Bowman, Michael Sharman, Chi Tong, Kavin Karthik, Arnav Dugar, Akshay Jagadeesh, Khaled Saab, Johannes Heidecke, Ashley Alexander, Nate Gross, and Karan Singhal. Healthbench professional: Evaluating large language models on real clinician chats, 2026. URL https://arxiv.org/abs/2604.27470.
  14. 14.Yixing Jiang, Kameron C. Black, Gloria Geng, Danny Park, James Zou, Andrew Y. Ng, and Jonathan H. Chen. MedAgentBench: A Virtual EHR Environment to Benchmark Medical LLM Agents. NEJM AI, 2, 2025. URL https://doi.org/10.1056/AIdbp2500144.
  15. 15.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421, 2021. URL https://doi.org/10.3390/app11146421.
  16. 16.Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Steven Horng, Leo Anthony Celi, and Roger Mark. MIMIC-IV Clinical Database Demo. PhysioNet, January 2023. doi: 10.13026/dp1f-ex47. URL https://doi.org/10.13026/dp1f-ex47. Version 2.2.
  17. 17.Alistair Johnson, Matthew Lungren, Yifan Peng, Zhiyong Lu, Roger Mark, Seth Berkowitz, and Steven Horng. MIMIC-CXR-JPG - chest radiographs with structured labels. PhysioNet, March 2024. doi: 10.13026/jsn5-t979. URL https://doi.org/10.13026/jsn5-t979. Version 2.1.0.
  18. 18.Alistair Johnson, Tom Pollard, Roger Mark, Seth Berkowitz, and Steven Horng. MIMIC-CXR Database. PhysioNet, July 2024. doi: 10.13026/4jqj-jw95. URL https://doi.org/10.13026/4jqj-jw95. Version 2.1.0.
  19. 19.Alistair E. W. Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J. Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, Li-wei H. Lehman, Leo A. Celi, and Roger G. Mark. MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 10(1):1, 2023. doi: 10.1038/s41597-022-01899-x.
  20. 20.Yusheng Liao, Shuyang Jiang, Yanfeng Wang, and Yu Wang. ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents. In Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL), 2025. URL https://doi.org/10.18653/v1/2025.acl-long.663.
  21. 21.Junqi Liu, Selena Song, Yuhan Wang, Jiawei Mao, Hardy Chen, Xiaoke Huang, Tianhao Qi, Pengfei Guo, Yucheng Tang, Yufan He, Can Zhao, Andriy Myronenko, Dong Yang, Daguang Xu, and Yuyin Zhou. Automedbench: Towards medical autoresearch with agentic ai models, 2026.
  22. 22.Ruoqi Liu, Imran Q. Mohiuddin, Austin J. Schoeffler, Kavita Renduchintala, Ashwin Nayak, Prasantha L. Vemu, Shivam C. Vedak, Kameron C. Black, John L. Havlik, Isaac Ogunmola, Stephen P. Ma, Roopa Dhatt, and Jonathan H. Chen. Physicianbench: Evaluating llm agents in real-world ehr environments, 2026. URL https://arxiv.org/abs/2605.02240.
  23. 23.Matthew McDermott and Justin Xu. MIMIC-IV MEDS: An ETL pipeline to extract MIMIC-IV data into the MEDS format. https://github.com/Medical-Event-Data-Standard/MIMIC_IV_MEDS, May 2025. Software, MIT license.
  24. 24.Mike A Merrill, Alexander G Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, E Kelly Buchanan, et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces. arXiv preprint arXiv:2601.11868, 2026.
  25. 25.Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. GAIA: A benchmark for General AI Assistants, 2023. URL https://doi.org/10.48550/arXiv.2311.12983.
  26. 26.Harsha Nori, Mayank Daswani, Christopher Kelly, Scott Lundberg, Marco Tulio Ribeiro, Marc Wilson, Xiaoxuan Liu, Viknesh Sounderajah, Jonathan Carlson, Matthew P Lungren, Bay Gross, Peter Hames, Mustafa Suleyman, Dominic King, and Eric Horvitz. Sequential diagnosis with language models, 2025. URL https://arxiv.org/abs/2506.22405.
  27. 27.OpenAI. Openai codex cli, 2025. URL https://github.com/openai/codex.
  28. 28.OpenAI. Gpt-5.3 codex, 2026. URL https://openai.com/index/introducing-gpt-5-3-codex/.
  29. 29.OpenAI. Gpt-5.4, 2026. URL https://openai.com/index/gpt-5-4.
  30. 30.OpenAI. Introducing gpt-5.5, 2026. URL https://openai.com/index/introducing-gpt-5-5.
  31. 31.Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor. Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments, 2025. URL https://arxiv.org/abs/2405.07960.
  32. 32.Ian Soboroff. Overview of trec 2021. In TREC, 2021.
  33. 33.Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason A. Fries, and Nigam H. Shah. EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models, 2023. URL https://doi.org/10.48550/arXiv.2307.02028.
  34. 34.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments, 2024. URL https://doi.org/10.48550/arXiv.2404.07972.
  35. 35.Ran Xu, Yuchen Zhuang, Yishan Zhong, Yue Yu, Zifeng Wang, Xiangru Tang, Hang Wu, May Dongmei Wang, Peifeng Ruan, Donghan Yang, Tao Wang, Guanghua Xiao, Xin Liu, Carl Yang, Yang Xie, and Wenqi Shi. Medagentgym: A scalable agentic training environment for code-centric reasoning in biomedical data science. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=jHDZEUgS4r.
  36. 36.Juan Manuel Zambrano Chaves, Shih-Cheng Huang, Yanbo Xu, Hanwen Xu, Naoto Usuyama, Sheng Zhang, Fei Wang, Yujia Xie, Mahmoud Khademi, Ziyi Yang, et al. A clinically accessible small multimodal radiology model and evaluation metric for chest x-ray findings. Nature Communications, 16(1):3108, 2025.
  37. 37.Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. WebArena: A Realistic Web Environment for Building Autonomous Agents, 2024. URL https://doi.org/10.48550/arXiv.2307.13854.
  38. 38.Yakun Zhu, Zhongzhen Huang, Qianhan Feng, Linjie Mu, Yannian Gu, Shaoting Zhang, Qi Dou, and Xiaofan Zhang. CP-Env: Evaluating Large Language Models on Clinical Pathways in a Controllable Hospital Environment, 2025. URL https://doi.org/10.48550/arXiv.2512.10206.
  39. 39.Yinghao Zhu, Ziyi He, Haoran Hu, Xiaochen Zheng, Xichen Zhang, Zixiang Wang, Junyi Gao, Liantao Ma, and Lequan Yu. MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks, 2025. URL https://doi.org/10.48550/arXiv.2505.12371.

Citation

MLA
Liu, Q., et al. “HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents”. arXiv, 2026, http://arxiv.org/abs/2606.31179v1.
APA
Liu, Q., Zhang, S., Qin, G., Valanarasu, J. M. J., Rokuss, M., Lu, M., Ossowski, T., Chaves, J. M. Z., Wong, C., Argaw, P., Hasija, Y., Wei, M., Yim, W.-. wai ., Liu, Q., Jing, Z., Entenmann, J., Usuyama, N., Naumann, T., & Poon, H. (2026). HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents. arXiv. http://arxiv.org/abs/2606.31179v1
Chicago
Liu, Q., S. Zhang, G. Qin, et al. 2026. “HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents”. arXiv. http://arxiv.org/abs/2606.31179v1.
Harvard
Liu, Q. et al. (2026) “HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2606.31179v1.
Vancouver
1. Liu Q, Zhang S, Qin G, et al (2026) HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents. arXiv

BibTeX

@article{liu2026healthagentbench,
  title = {HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents},
  author = {Liu, Qianchu and Zhang, Sheng and Qin, Guanghui and Valanarasu, Jeya Maria Jose and Rokuss, Maximilian and Lu, Mingyu and Ossowski, Timothy and Chaves, Juan Manuel Zambrano and Wong, Cliff and Argaw, Peniel and Hasija, Yashna and Wei, Mu and Yim, Wen-wai and Liu, Qin and Jing, Zilin and Entenmann, Jason and Usuyama, Naoto and Naumann, Tristan and Poon, Hoifung},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2606.31179v1},
  eprint = {2606.31179}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/