MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis

Akshat SanghviNaren AkashRaza ImamAmit SharmaMohit Jain

article2026arXiv5 citations

Proposes a multi-agent consultation framework and a 4,421-case benchmark that replace unrealistic single-shot evaluations with interactive, sequential clinical questioning, closing over half the diagnostic accuracy gap to full-information oracles.

Listen

Large language models are increasingly used for clinical decision support, but standard evaluations treat medical diagnosis as a static, single-shot task with complete information provided upfront. This formulation fails to reflect real-world clinical practice, which is fundamentally interactive and requires sequential hypothesis refinement through targeted questioning. The article addresses this gap by creating MeDxBench, a comprehensive benchmark of 4,421 diverse clinical cases across 20 specialties, and proposing MeDxAgent, a multi-agent consultation system designed to conduct interactive, multi-turn clinical diagnosis without leaking ground-truth diagnostic data.

The researchers systematically investigated prompt-, flow-, and agent-level design choices using simulated consultations among a doctor agent, a patient agent restricted to the source case description, and an evaluation judge. Built across five public datasets encompassing 4,113 unique conditions ranging from common to rare diseases, the evaluation assessed how structured multi-agent collaboration, hypothetico-deductive reasoning, external knowledge grounding, and missing-evidence tracking influence diagnostic accuracy across different foundational model families.

The findings show that MeDxAgent achieves a 57.4% diagnostic accuracy on MeDxBench using GPT-4o, delivering a 10.3 percentage point improvement over the single-agent baseline (47.1%) and closing 52.3% of the gap to a full-information oracle ceiling (66.8%). Three design elements proved critical: gathering demographics in the initial turn, converting dialogue transcripts into structured clinical summaries, and steering follow-up questions with candidate diagnoses. The timing of interventions was pivotal; activating differential questioning too early reduced accuracy to 34.7%, whereas delaying it to the tenth turn yielded 52.8%, representing an 18.1 percentage point swing. Furthermore, specialized components such as specialist ensembles, knowledge graph retrieval, and evidence-gap tracking degraded performance when used in isolation due to overconfidence or confirmation bias, yet became mutually reinforcing and beneficial in the unified system.

These results demonstrate that multi-agent systems must carefully balance broad information gathering with targeted hypothesis testing to prevent premature cognitive anchoring and inflated confidence. Across model families, the architecture transferred effectively, providing consistent gains and improving diagnostic performance most substantially on rare conditions (an 11.9 percentage point increase). Deploying such workflows could significantly improve the reliability and safety of clinical diagnostic aids, provided that agent coordination is structured to avoid premature closure during patient intake.

Organizations developing clinical decision support should implement structured clinical summarization, enforce early demographic collection, and carefully schedule targeted diagnostic inquiries rather than applying them from the first interaction. MeDxAgent should be positioned as a supportive tool for clinicians to surface differential diagnoses rather than an autonomous diagnostic system. Future initiatives should focus on expanding the framework to handle multimodal clinical inputs, such as medical imaging, and validating interactive agent policies within supervised, real-world clinical pilots.

arXiv: 2606.03416
  • Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). Synthesizes multi-agent collaborative workflows and dynamic reasoning frameworks, contextualizing specialized diagnostic consultation agents within broader autonomous agent architectures.
  • Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). Investigates how LLM agents handle incrementally revealed and evolving information in multi-turn interactions, extending the interactive questioning and state-tracking findings of MeDxAgent.
  • Paper: Decentralized Multi-Agent Systems with Shared Context, Yuzhen Mao et al. (2026). Explores decentralized multi-agent architectures with shared context, offering potential structural scaling paradigms beyond the structured consultation pipelines evaluated in MeDxAgent.
  • Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). Introduces diagnostic safety guardrails and trajectory evaluation for autonomous AI agents, providing essential safety auditing for interactive clinical decision-making agents.
Cover for MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis

Abstract

Large language models (LLMs) are increasingly used for health-related decision support. Yet most evaluations treat diagnosis as a single-shot task with complete information provided upfront, often as a multiple-choice selection. This diverges from clinical practice, where diagnosis is interactive and open-ended, involving sequential hypothesis refinement through targeted questioning. We address this gap. We build MeDxBench, a large-scale benchmark of 4,421 clinical cases across 20 specialties. We further propose MeDxAgent, a multi-agent consultation system for interactive diagnosis, and systematically study its prompt-, flow- and agent-level design choices. MeDxAgent achieves a 10.3% accuracy gain over the baseline on MeDxBench, closing 52.3% of the gap to a full-information oracle. We find that specific design choices: collecting demographics first, passing summarized dialogue for diagnosis, and feeding candidate diagnoses for targeted questioning, improve accuracy, mirroring how physicians reason, though their effect emerges fully only in combination. Code and dataset will be released upon publication.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 MeDxBench Benchmark Dataset
  • 4 Simulation Framework
  • 5 MeDxAgent System
  • 5.1 Prompt-Based Variations
  • 5.2 Flow-Based Variations
  • 5.3 Agent-Based Variations
  • 6 Results
  • 7 Discussion
  • 8 Conclusion
  • 9 Limitations
  • References
  • A Additional Results
  • A.1 Accuracy by Specialty
  • A.2 Accuracy by Disease Prevalence
  • A.3 Top-k Accuracy
  • A.4 Accuracy of the Patient Agent
  • A.5 Round-Level Conversation Analysis
  • B Implementation Details
  • C Prompts
  • C.1 Patient Agent
  • C.2 Doctor: Question Agents
  • C.3 Doctor: Diagnosis Agents
  • C.4 Summarizer Agents
  • C.5 Evidence Gap Agent
  • C.6 Judge Agent
  • D Example

Knowls

  1. Knowl 1 — MeDxBench Interactive Medical Diagnosis Benchmark

    definition

    MeDxBench is a large-scale diagnostic benchmark designed to evaluate large language models (LLMs) in multi-turn, interactive clinical diagnosis rather than static, single-turn question answering. It comprises 4,421 clinical cases across 20 medical specialties, consolidated and standardized from five public datasets:

    1. CRAFT-MD: 200 free-text clinical vignettes curated by clinical experts.
    2. DiagnosisArena: 897 peer-reviewed journal case reports.
    3. MedMCQA: 847 exam-style multiple-choice questions from AIIMS and NEET-PG entrance examinations.
    4. MedQA: 1,489 exam-style board examination questions with clinical reasoning chains.
    5. PubMed: 988 physician-curated PubMed case reports.

    The benchmark contains 4,113 unique disease labels categorized by global prevalence:

    • Common (> 1 in 100): 655 diseases across 822 cases.
    • Uncommon (1 in 100 to 1 in 2,000): 1,098 diseases across 1,351 cases.
    • Rare (< 1 in 2,000): 2,360 diseases across 2,248 cases.

    Case descriptions range from 212 to 9,848 characters (mean 1,785±1,9551,785 \pm 1,955), with 546 cases containing multiple ground-truth diagnosis labels. To prevent diagnostic data leakage, all explicit mentions of ground-truth diagnoses are redacted from case descriptions while retaining clinical evidence, and non-diagnostic tasks or cases requiring direct image/video interpretation are excluded.

  2. Knowl 2 — MeDxAgent Multi-Agent Consultation Architecture

    model/method

    MeDxAgent is a multi-agent consultation framework for interactive disease diagnosis that simulates clinical reasoning across a maximum budget of 20 dialogue turns. The framework operates in two distinct phases:

    1. Broad Information Gathering (Turns 1–9):

      • The Summarizer Agent distills the ongoing dialogue into a structured clinical problem representation at every turn.
      • In Turn 1, the Question Agent collects baseline demographic information (age and sex).
      • In Turns 2–9, the Evidence Gap Agent identifies general missing baseline information (vital signs, medical history, symptom timeline, lab tests), prompting the Question Agent to explore broad clinical evidence.
    2. Differential Diagnosis Consultation (Turns ≥\ge 10):

      • Specialist Ensemble: 10 physician persona agents (9 specialists in Cardiology, Pulmonology, Gastroenterology, Neurology, Endocrinology, Nephrology, Rheumatology, Hematology, Dermatology, plus 1 General Physician) independently review the clinical summary and predict their top-3 diagnoses with confidence scores and reasoning (yielding 30 candidates). A senior physician agent (Selector-SE) weights each specialist by case relevance.
      • Knowledge Graph Retrieval: Symptoms revealed by the patient are queried against the Unified Biomedical Knowledge Graph (UBKG) to score and retrieve top-3 symptom-matched diseases, which are consolidated with LLM predictions via a Selector-KG agent.
      • Selector Agent: Consolidates candidate diagnoses from the Specialist Ensemble and Knowledge Graph into a unified top-3 differential diagnosis.
      • Evidence Gap Agent: Analyzes the top-3 diagnoses alongside the clinical summary to generate actionable diagnosis-specific gaps (evidence needed to confirm or rule out competing diseases) and general gaps.
      • Evidence-Guided Differential Question Agent: Queries the patient with targeted questions aimed at differentiating competing candidate diagnoses and resolving high-utility evidence gaps.
  3. Knowl 3 — Interactive Multi-Turn Simulation and Evaluation Protocol

    experimental setup

    The simulation environment consists of three LLM-based agent roles interacting under strict anti-leakage constraints:

    • Patient Agent: Initialized solely with the clinical case text (with ground-truth diagnoses redacted). It answers doctor queries strictly from explicit case details in the first person using non-technical language. If requested information is absent from the case description, it must respond with "I don't know", preventing synthetic hallucination or diagnostic leakage.
    • Doctor Agent (System under Evaluation): Begins with an open greeting ("What brings you in today?") and conducts multi-turn questioning. After each turn, it predicts a ranked list of top-3 candidate diagnoses with rationale and confidence scores (0–100%). Consultations terminate early if a diagnosis reaches a ≥95%\ge 95\% confidence threshold; otherwise, they proceed to a hard cap of 20 turns, selecting the top prediction from turn 20.
    • Judge Agent: Uses semantic matching to compare predicted diagnoses against ground truth. A prediction is classified as correct if it is an exact match, a recognized medical synonym/paraphrase, or a more specific clinical subtype of the ground truth (e.g., malignant melanoma for melanoma). Predictions that are more general than the ground truth are marked incorrect. For multi-label cases, matching any ground-truth label is considered correct.

    All simulation agents (Patient, Judge) and core evaluated models are run at temperature T=0T = 0 with structured JSON schema outputs.

  4. Knowl 4 — Diagnostic Accuracy of MeDxAgent and Component Variants on MeDxBench

    data/table

    Diagnostic accuracy (%) across MeDxBench and its five constituent datasets evaluated with GPT-4o as the backbone model across prompt, flow, and agent variations (averaged over 3 runs, variance <0.5%< 0.5\%):

    Variant Avg. Acc. CRAFT-MD Diag. Arena MedMCQA MedQA PubMed
    Baseline 47.1 52.5 18.8 57.3 45.9 60.9
    Prompt Variations
    Turn Awareness 48.6 56.5 21.1 59.2 45.5 61.0
    Demographics First 49.8 54.0 20.4 61.0 51.4 62.4
    Combined Prompt 51.7 57.0 22.5 63.9 51.9 63.3
    Flow Variations
    No Early Stopping 51.6 54.5 24.9 62.6 50.1 66.1
    Differential Q (Turn 2) 34.7 35.0 12.3 47.7 36.9 41.6
    Differential Q (Turn 5) 50.4 59.0 18.8 63.6 52.6 57.7
    Differential Q (Turn 10) 52.8 60.5 23.5 64.5 54.2 61.1
    Combined Flow 52.8 61.0 20.9 63.8 54.1 64.4
    Agent Variations
    Summarizer: Paragraph 54.9 61.0 27.4 64.6 54.8 66.6
    Summarizer: Structured 54.6 61.0 27.0 63.2 55.1 67.0
    Specialist Ensemble 51.1 54.5 24.2 63.9 52.3 60.6
    Knowledge Graph 51.7 57.0 21.6 62.9 54.3 62.6
    Evidence Gap 50.6 59.5 23.2 61.0 48.8 60.5
    MeDxAgent (Full) 57.4 66.5 27.4 69.8 56.9 66.6
    Oracle (Full Information) 66.8 65.5 31.7 76.3 68.6 92.0

    MeDxAgent achieves a 10.3 percentage point gain over the baseline (47.1% to 57.4%), recovering 52.3% of the performance gap between the interactive baseline and the static full-information Oracle (66.8%).

  5. Knowl 5 — Timing Sensitivity and Anchoring Bias in Differential Questioning

    empirical result

    Differential questioning—where the question-asking agent is provided current diagnostic hypotheses and instructed to ask discriminative questions—exhibits extreme sensitivity to the turn at which it is activated:

    • Activating differential questioning early at Turn 2 causes diagnostic accuracy to drop to 34.7%, which is 12.4 percentage points below the single-agent baseline (47.1%) and 17.0 points below the Combined Prompt baseline (51.7%). This degradation is driven by cognitive anchoring bias: the doctor conditions inquiries on premature, poorly supported early hypotheses, leading to circular inquiry on incorrect diagnoses.
    • Delaying activation to Turn 5 allows general clinical evidence to accumulate first, recovering accuracy to 50.4%.
    • Activating differential questioning at Turn 10 achieves 52.8% accuracy (+5.7 percentage points over the baseline).

    Adjusting only the switch turn threshold while keeping the prompts and architecture identical results in an 18.1 percentage point variation in overall diagnostic accuracy (34.7% vs. 52.8%).

  6. Knowl 6 — Compositional Synergy and Isolation Failure of Multi-Agent Modules

    empirical result

    When evaluated in isolation on top of the Flow baseline (Differential Questioning Turn 10, 52.8% accuracy), individual diagnostic agents decrease performance, whereas combining all agents yields the highest overall accuracy (57.4%):

    1. Isolation Failures:

      • Summarizer Agent: The only component that helps in isolation (+2.1 points for paragraph format to 54.9%; +1.8 points for structured schema to 54.6%) by structuring dialogue into problem representations aligned with LLM training distributions.
      • Specialist Ensemble: Lowers accuracy to 51.1% (-1.7 points). Repeated predictions among 10 specialists inflate ensemble confidence scores, causing premature early stopping (23.3% early stop rate) before adequate evidence is gathered.
      • Knowledge Graph Grounding (UBKG): Lowers accuracy to 51.7% (-1.1 points) due to over-weighting symptom-matched graph entities lacking clinical test context.
      • Evidence Gap Agent: Lowers accuracy to 50.6% (-2.2 points) due to confirmation bias, focusing questions too heavily on validating early candidate hypotheses.
    2. Synergistic Combination and Leave-One-Out Ablation: When combined in MeDxAgent, the Summarizer provides clean state representations that enable the Knowledge Graph and Specialist Ensemble to propose diverse, unanchored candidates; these diverse candidates allow the Evidence Gap agent to guide inquiry objectively. Leave-one-out ablations confirm that every agent is essential to the full system (57.4%):

      • Full MeDxAgent: 57.4%
      • Without Summarizer: 54.4% (-3.0 points)
      • Without Evidence Gap: 55.5% (-1.9 points)
      • Without Specialist Ensemble: 55.7% (-1.7 points)
      • Without Knowledge Graph: 56.0% (-1.4 points)
  7. Knowl 7 — Cross-Model Generalization and External Baseline Comparison

    data/table

    The architectural components of MeDxAgent generalize across distinct LLM families without prompt re-tuning and outperform prior state-of-the-art interactive diagnostic frameworks.

    Cross-Model Accuracy (%) on MeDxBench (D1: CRAFT-MD, D2: DiagArena, D3: MedMCQA, D4: MedQA, D5: PubMed):

    Model Variant D1 D2 D3 D4 D5
    GPT-4o Baseline 52.5 18.8 57.3 45.9 60.9
    Combined Prompt 57.0 22.5 63.9 51.9 63.3
    Differential Q 60.5 23.5 64.5 54.2 61.1
    MeDxAgent 66.5 27.4 69.9 56.9 66.6
    Full Info. Oracle 65.5 31.7 76.3 68.6 92.0
    Grok-4.1 Baseline 53.5 26.6 60.7 47.4 69.4
    Combined Prompt 55.0 26.9 60.6 48.4 70.6
    Differential Q 58.0 26.8 62.6 50.0 67.3
    MeDxAgent 62.0 28.8 63.8 51.4 63.8
    Full Info. Oracle 60.5 35.5 74.5 62.2 92.6
    DeepSeek-V3.2 Baseline 56.0 26.3 61.6 48.2 65.8
    Combined Prompt 55.5 25.3 63.2 48.9 64.4
    Differential Q 61.5 27.9 67.2 51.0 65.7
    MeDxAgent 65.0 26.3 66.6 54.3 60.5
    Full Info. Oracle 73.0 36.2 78.0 68.0 91.5

    Comparison with Existing Systems:

    • Against MedAgentSim (GPT-4o backbone): MeDxAgent achieves 63.07% average accuracy vs. MedAgentSim's 56.87% (+6.20%), including NEJM (41.67% vs. 27.50%, +14.17%), MIMIC (76.04% vs. 75.30%, +0.74%), and MedQA (71.50% vs. 67.80%, +3.70%).
    • Against VivaBench (PubMed benchmark): MeDxAgent outperforms VivaBench across model backbones: GPT-4o (43.3% vs. 23.1%, +20.2%), o4-mini (37.7% vs. 32.0%, +5.7%), and Gemini-2.5-Pro (43.5% vs. 35.0%, +8.5%).
  8. Knowl 8 — Diagnostic Accuracy Stratified by Disease Prevalence and Clinical Specialty

    empirical result

    Performance analysis of MeDxAgent across disease prevalence and clinical specialties reveals distinct diagnostic behaviors:

    1. Prevalence-Dependent Performance: MeDxAgent's diagnostic gains over the single-agent baseline are inversely correlated with disease prevalence:

      • Common diseases (n=822n = 822): accuracy increases from 56.3% to 61.1% (+4.8 points).
      • Uncommon diseases (n=1,351n = 1,351): accuracy increases from 55.4% to 64.5% (+9.1 points).
      • Rare diseases (n=2,248n = 2,248): accuracy increases from 37.1% to 49.0% (+11.9 points). The Summarizer module contributes disproportionately to rare diseases (45.7% accuracy vs. 42.1% for the best flow variant) by consolidating sparse, non-obvious clinical clues.
    2. Specialty Performance and Confusions: MeDxAgent is the top-performing variant in 15 of 20 clinical specialties, with highest absolute accuracy in Endocrinology (68.7%), Orthopaedics (68.0%), and Obstetrics & Gynaecology (65.3%). Overall specialty-match accuracy (the predicted condition belonging to the correct specialty domain) is 70.4%. Distinctive specialties achieve high match rates: Ophthalmology (86.8%), Obstetrics & Gynaecology (84.0%), Neurology (83.7%), and Dermatology (83.0%). Specialties with systemic manifestations (such as Rheumatology) frequently produce confusions distributed across organ-specific specialties (Cardiology, Neurology, Nephrology).

  9. Knowl 9 — Dialogue Dynamics and Uninformative Response Rates in Medical Simulation

    empirical result

    Interaction dynamics across the 4,421 cases reveal key properties of simulated multi-turn consultations:

    1. Early Stopping Formula: The early stopping rate measures the percentage of the 20-round dialogue budget saved: Early Stop Rate (%)=1−Total Rounds#Cases×Max Rounds\text{Early Stop Rate (\%)} = 1 - \frac{\text{Total Rounds}}{\#\text{Cases} \times \text{Max Rounds}} where #Cases=4,421\#\text{Cases} = 4,421 and Max Rounds=20\text{Max Rounds} = 20. In MeDxAgent, 88.9% of cases (3,932 of 4,421) run for the full 20 rounds, yielding an average early stopping rate of 4.2% (total rounds: 84,739).

    2. Correlation with Uninformative Answers: Across system variants, overall diagnostic accuracy strongly negatively correlates with the proportion of "I don't know" (IDK) rounds: r=−0.79r = -0.79 In contrast, correlation with total conversation rounds (r=0.27r = 0.27) and the raw count of IDK rounds (r=−0.07r = -0.07) is weak. Effective questioning policies are characterized by eliciting substantive answers from the patient.

    3. Evidence-Guided Questioning Impact: In standard variants, the IDK response rate increases monotonically from turn 1 (~1%) to turn 20 (~80–85%). The Evidence Gap agent and MeDxAgent produce an abrupt drop in IDK rate around turn 10 (falling to 48–57%) by redirecting inquiries to specific, clinically relevant evidence gaps.

  10. Knowl 10 — Limitations and Scope of MeDxAgent and MeDxBench

    limitation

    The methodology and benchmark exhibit four primary limitations:

    1. Text-Only Modality: The framework is restricted to textual dialogue and textual imaging reports, omitting direct multimodal reasoning over medical imaging, waveforms (ECG), or video, which are integral to real-world diagnostics.
    2. Patient Simulation Fidelity: Constraining the Patient Agent to answer strictly from case vignettes prevents diagnostic leakage but causes frequent "I don't know" responses (occurring in ≈60%\approx 60\% of rounds and >80%> 80\% of cases), diverging from realistic patient interactions.
    3. Heuristic Prompt Engineering: Interaction timing and agent coordination rely on fixed turn thresholds (e.g., activating differential questioning at Turn 10) and rule-based prompting rather than learned questioning policies or dynamic stopping criteria.
    4. Deployment Scope: MeDxAgent is strictly designed as a clinical decision-support tool to assist physicians with differential diagnosis, not as an autonomous diagnostic system, as erroneous or overconfident predictions pose severe safety risks in clinical workflows.

Coverage note — None was omitted; all key contributions—including the benchmark statistics, multi-agent architecture, prompt/flow/agent experimental ablations, cross-model evaluations, external baselines, dialogue dynamics, and stated limitations—are fully covered.

References

  1. 1.Mohammad Almansoori, Komal Kumar, and Hisham Cholakkal. 2026. Medagentsim: Self-evolving multi-agent simulations for realistic clinical interactions. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2025, pages 362–372, Cham. Springer Nature Switzerland.
  2. 2.Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, Johannes Heidecke, and Karan Singhal. 2025. Healthbench: Evaluating large language models towards improved human health. Preprint, arXiv:2505.08775.
  3. 3.Suhana Bedi, Hejie Cui, Miguel Fuentes, Alyssa Unell, Michael Wornow, Juan M. Banda, Nikesh Kotecha, Timothy Keyes, Yifan Mai, Mert Oez, Hao Qiu, Shrey Jain, Leonardo Schettini, Mehr Kashyap, Jason Alan Fries, Akshay Swaminathan, Philip Chung, Fateme Nateghi Haredasht, Ivan Lopez, and 64 others. 2026. Holistic evaluation of large language models for medical tasks with medhelm. Nature Medicine, 32(3):943–951.
  4. 4.Judith L. Bowen. 2006. Educational strategies to promote clinical diagnostic reasoning. New England Journal of Medicine, 355(21):2217–2225.
  5. 5.Ruey-Wen Chang, Georges Bordage, and Kathleen J. Connell. 1998. The importance of early problem representation during case presentations. Academic Medicine, 73(10):S109–S111.
  6. 6.Kai Chen, Xinfeng Li, Tianpei Yang, Hewei Wang, Wei Dong, and Yang Gao. 2025. Mdteamgpt: A self-evolving llm-based multi-agent framework for multi-disciplinary team medical consultation. arXiv preprint arXiv:2503.13856.
  7. 7.Christopher Chiu, Silviu Pitis, and Mihaela van der Schaar. 2026. Simulating viva voce examinations to evaluate clinical reasoning in large language models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  8. 8.Beatriz Costa-Gomes, Sophia Chen, Connie Hsueh, Deborah Morgan, Philipp Schoenegger, Yash Shah, Sam Way, Yuki Zhu, Timothé Adeline, Michael Bhaskar, Mustafa Suleyman, and Seth Spielman. 2025. It’s about time: The temporal and modal dynamics of copilot usage. Preprint, arXiv:2512.11879.
  9. 9.Beatriz Costa-Gomes, Pavel Tolmachev, Eloise Taysom, Viknesh Sounderajah, Hannah Richardson, Philipp Schoenegger, Xiaoxuan Liu, Matthew M. Nour, Seth Spielman, Samuel F. Way, Yash Shah, Michael Bhaskar, Harsha Nori, Christopher Kelly, Peter Hames, Bay Gross, Mustafa Suleyman, and Dominic King. 2026. Public use of a generalist llm chatbot for health queries. Nature Health.
  10. 10.Arthur S. Elstein, Lee S. Shulman, and Sarah A. Sprafka. 1978. Medical Problem Solving: An Analysis of Clinical Reasoning. Harvard University Press, Cambridge, MA.
  11. 11.Zhihao Fan, Lai Wei, Jialong Tang, Wei Chen, Wang Siyuan, Zhongyu Wei, and Fei Huang. 2025. AI hospital: Benchmarking large language models in a multi-agent medical interaction simulator. In Proceedings of the 31st International Conference on Computational Linguistics, pages 10183–10213, Abu Dhabi, UAE. Association for Computational Linguistics.
  12. 12.Yanjun Gao, Ruizhe Li, Emma Croxford, John Caskey, Brian W. Patterson, Matthew Churpek, Timothy Miller, Dmitriy Dligach, and Majid Afshar. 2025. Leveraging medical knowledge graphs into large language models for diagnosis prediction: Design and application study. JMIR AI, 4:e58670.
  13. 13.Mark L. Graber, Nancy Franklin, and Ruthanna Gordon. 2005. Diagnostic error in internal medicine. Archives of Internal Medicine, 165(13):1493–1499.
  14. 14.Zhihao Jia, Mingyi Jia, Junwen Duan, and Jianxin Wang. 2025. Ddo: Dual-decision optimization for llm-based medical consultation via multi-agent collaboration. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 26380–26397.
  15. 15.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14).
  16. 16.Shreya Johri, Jaehwan Jeong, Benjamin A. Tran, Daniel I. Schlessinger, Shannon Wongvibulsin, Leandra A. Barnes, Hong-Yu Zhou, Zhuo Ran Cai, Eliezer M. Van Allen, David Kim, Roxana Daneshjou, and Pranav Rajpurkar. 2025. An evaluation framework for clinical use of large language models in patient interaction tasks. Nature Medicine, 31(1):77–86.
  17. 17.Yubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Marzyeh Ghassemi, Cynthia Breazeal, and Hae Won Park. 2024. Mdagents: An adaptive collaboration of llms for medical decision-making. In Advances in Neural Information Processing Systems, volume 37, pages 79410–79452. Curran Associates, Inc.
  18. 18.Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, and Jennifer Neville. 2026. LLMs get lost in multi-turn conversation. In The Fourteenth International Conference on Learning Representations.
  19. 19.Junkai Li, Yunghwei Lai, Weitao Li, Jingyi Ren, Meng Zhang, Xinhui Kang, Siyu Wang, Peng Li, Ya-Qin Zhang, Weizhi Ma, and Yang Liu. 2025. Agent hospital: A simulacrum of hospital with evolvable medical agents. Preprint, arXiv:2405.02957.
  20. 20.Shuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S. Ilgen, Emma Pierson, Pang Wei Koh, and Yulia Tsvetkov. 2024. Mediq: Question-asking llms and a benchmark for reliable interactive clinical reasoning. In Advances in Neural Information Processing Systems, volume 37, pages 28858–28888. Curran Associates, Inc.
  21. 21.Yusheng Liao, Yutong Meng, Yuhao Wang, Hongcheng Liu, Yanfeng Wang, and Yu Wang. 2024. Automatic interactive evaluation for large language models with state aware patient simulator. Preprint, arXiv:2403.08495.
  22. 22.Harsha Nori, Mayank Daswani, Christopher Kelly, Scott Lundberg, Marco Tulio Ribeiro, Marc Wilson, Xiaoxuan Liu, Viknesh Sounderajah, Jonathan Carlson, Matthew P Lungren, Bay Gross, Peter Hames, Mustafa Suleyman, Dominic King, and Eric Horvitz. 2025. Sequential diagnosis with language models. Preprint, arXiv:2506.22405.
  23. 23.Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Proceedings of the Conference on Health, Inference, and Learning, volume 174 of Proceedings of Machine Learning Research, pages 248–260. PMLR.
  24. 24.Rajat Rawat, Hudson McBride, Rajarshi Ghosh, Dhiyaan Nirmal, Jong Moon, Dhruv Alamuri, Sean O’Brien, and Kevin Zhu. 2024. DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models. In Proceedings of the Third Workshop on NLP for Positive Impact, pages 334–348, Miami, Florida, USA. Association for Computational Linguistics.
  25. 25.Daniel Philip Rose, Chia-Chien Hung, Marco Lepri, Israa Alqassem, Kiril Gashteovski, and Carolin Lawrence. 2025. MEDDxAgent: A unified modular agent framework for explainable automatic differential diagnosis. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13803–13826, Vienna, Austria. Association for Computational Linguistics.
  26. 26.Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor. 2025. Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments. Preprint, arXiv:2405.07960.
  27. 27.Tianqi Shang, Weiqing He, Charles Zheng, Lingyao Li, Li Shen, and Bingxin Zhao. 2025. Dynamicare: A dynamic multi-agent framework for interactive and open-ended medical decision-making. Preprint, arXiv:2507.02616.
  28. 28.Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Mansfield, Dina Demner-Fushman, and 13 others. 2023. Large language models encode clinical knowledge. Nature, 620(7972):172–180.
  29. 29.Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R. Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H. Chen, Nigam H. Shah, Sami Lachgar, Philip Andrew Mansfield, and 16 others. 2025. Toward expert-level medical question answering with large language models. Nature Medicine, 31(3):943–950.
  30. 30.Benjamin J. Stear, Taha Mohseni Ahooyi, J. Alan Simmons, Charles Kollar, Lance Hartman, Katherine Beigel, Aditya Lahiri, Shubha Vasisht, Tiffany J. Callahan, Christopher M. Nemarich, Jonathan C. Silverstein, and Deanne M. Taylor. 2024. Petagraph: A large-scale unifying knowledge graph framework for integrating biomolecular and biomedical data. Scientific Data, 11(1):1338.
  31. 31.Xiangru Tang, Anni Zou, Zhuosheng Zhang, Ziming Li, Yilun Zhao, Xingyao Zhang, Arman Cohan, and Mark Gerstein. 2024. MedAgents: Large language models as collaborators for zero-shot medical reasoning. In Findings of the Association for Computational Linguistics: ACL 2024, pages 599–621, Bangkok, Thailand. Association for Computational Linguistics.
  32. 32.Tao Tu, Mike Schaekermann, Anil Palepu, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Yong Cheng, Elahe Vedadi, Nenad Tomasev, Shekoofeh Azizi, Karan Singhal, Le Hou, Albert Webson, Kavita Kulkarni, S. Sara Mahdavi, Christopher Semturs, and 7 others. 2025. Towards conversational diagnostic artificial intelligence. Nature, 642(8067):442–450.
  33. 33.Qipeng Wang, Rui Sheng, Yafei Li, Huamin Qu, Yushi Sun, and Min Zhu. 2025a. Medkgi: Iterative differential diagnosis with medical knowledge graphs and information-guided inquiring. arXiv preprint arXiv:2512.24181.
  34. 34.Shansong Wang, Mingzhe Hu, Qiang Li, Mojtaba Safari, and Xiaofeng Yang. 2025b. Capabilities of gpt-5 on multimodal medical reasoning. Preprint, arXiv:2508.08224.
  35. 35.Cliff Wong, Sam Preston, Qianchu Liu, Zelalem Gero, Jaspreet Bagga, Sheng Zhang, Shrey Jain, Theodore Zhao, Yu Gu, Yanbo Xu, Sid Kiblawi, Srinivasan Yegnasubramanian, Taxiarchis Botsis, Marvin Borja, Luis M. Ahumada, Joseph C. Murray, Guo Hui Gan, Roshanthi Weerasinghe, Kristina Young, and 6 others. 2025. Universal abstraction: Harnessing frontier models to structure real-world data at scale. Preprint, arXiv:2502.00943.
  36. 36.Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, and Xiaofan Zhang. 2025. Diagnosisarena: Benchmarking diagnostic reasoning for large language models. Preprint, arXiv:2505.14107.

Citation

MLA
Sanghvi, A., et al. “MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis”. arXiv, 2026, https://doi.org/10.48550/arxiv.2606.03416.
APA
Sanghvi, A., Akash, N., Imam, R., Sharma, A., & Jain, M. (2026). MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis. arXiv. https://doi.org/10.48550/arxiv.2606.03416
Chicago
Sanghvi, A., N. Akash, R. Imam, A. Sharma, and M. Jain. 2026. “MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2606.03416.
Harvard
Sanghvi, A. et al. (2026) “MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis”. arXiv. Available at: https://doi.org/10.48550/arxiv.2606.03416.
Vancouver
1. Sanghvi A, Akash N, Imam R, Sharma A, Jain M (2026) MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis. https://doi.org/10.48550/arxiv.2606.03416

BibTeX

@misc{https://doi.org/10.48550/arxiv.2606.03416,
  doi = {10.48550/ARXIV.2606.03416},
  url = {https://arxiv.org/abs/2606.03416},
  author = {Sanghvi, Akshat and Akash, Naren and Imam, Raza and Sharma, Amit and Jain, Mohit},
  keywords = {Multiagent Systems (cs.MA), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/