MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis
Akshat SanghviNaren AkashRaza ImamAmit SharmaMohit Jain
Proposes a multi-agent consultation framework and a 4,421-case benchmark that replace unrealistic single-shot evaluations with interactive, sequential clinical questioning, closing over half the diagnostic accuracy gap to full-information oracles.
Large language models are increasingly used for clinical decision support, but standard evaluations treat medical diagnosis as a static, single-shot task with complete information provided upfront. This formulation fails to reflect real-world clinical practice, which is fundamentally interactive and requires sequential hypothesis refinement through targeted questioning. The article addresses this gap by creating MeDxBench, a comprehensive benchmark of 4,421 diverse clinical cases across 20 specialties, and proposing MeDxAgent, a multi-agent consultation system designed to conduct interactive, multi-turn clinical diagnosis without leaking ground-truth diagnostic data.
The researchers systematically investigated prompt-, flow-, and agent-level design choices using simulated consultations among a doctor agent, a patient agent restricted to the source case description, and an evaluation judge. Built across five public datasets encompassing 4,113 unique conditions ranging from common to rare diseases, the evaluation assessed how structured multi-agent collaboration, hypothetico-deductive reasoning, external knowledge grounding, and missing-evidence tracking influence diagnostic accuracy across different foundational model families.
The findings show that MeDxAgent achieves a 57.4% diagnostic accuracy on MeDxBench using GPT-4o, delivering a 10.3 percentage point improvement over the single-agent baseline (47.1%) and closing 52.3% of the gap to a full-information oracle ceiling (66.8%). Three design elements proved critical: gathering demographics in the initial turn, converting dialogue transcripts into structured clinical summaries, and steering follow-up questions with candidate diagnoses. The timing of interventions was pivotal; activating differential questioning too early reduced accuracy to 34.7%, whereas delaying it to the tenth turn yielded 52.8%, representing an 18.1 percentage point swing. Furthermore, specialized components such as specialist ensembles, knowledge graph retrieval, and evidence-gap tracking degraded performance when used in isolation due to overconfidence or confirmation bias, yet became mutually reinforcing and beneficial in the unified system.
These results demonstrate that multi-agent systems must carefully balance broad information gathering with targeted hypothesis testing to prevent premature cognitive anchoring and inflated confidence. Across model families, the architecture transferred effectively, providing consistent gains and improving diagnostic performance most substantially on rare conditions (an 11.9 percentage point increase). Deploying such workflows could significantly improve the reliability and safety of clinical diagnostic aids, provided that agent coordination is structured to avoid premature closure during patient intake.
Organizations developing clinical decision support should implement structured clinical summarization, enforce early demographic collection, and carefully schedule targeted diagnostic inquiries rather than applying them from the first interaction. MeDxAgent should be positioned as a supportive tool for clinicians to surface differential diagnoses rather than an autonomous diagnostic system. Future initiatives should focus on expanding the framework to handle multimodal clinical inputs, such as medical imaging, and validating interactive agent policies within supervised, real-world clinical pilots.
- Paper: Improving Factuality and Reasoning in Language Models through Multiagent Debate, Yilun Du et al. (2023). Establishes the foundational multi-agent debate and consultation paradigm that MeDxAgent adapts and specializes for multi-agent clinical diagnostic reasoning.
- Paper: Capabilities of GPT-4 on Medical Challenge Problems, Harsha Nori et al. (2023). Provides the foundational baseline evaluation of frontier LLMs on standardized medical licensing exams, illustrating the single-shot paradigm that MeDxAgent seeks to reform into interactive diagnosis.
- Paper: Large language models encode clinical knowledge, Karan Singhal et al. (2022). Introduces foundational medical benchmark tasks and clinical knowledge evaluation frameworks that motivate the creation of interactive benchmarks like MeDxBench.
- Paper: MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues, Ge Bai et al. (2024). Formalizes the multi-turn conversational evaluation and proactive questioning principles necessary for understanding interactive sequential diagnosis systems.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Presents a comprehensive architectural taxonomy for LLM-based autonomous agents and multi-agent workflows utilized in designing specialized medical agents.
- Paper: What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams, Di Jin et al. (2020). Supplies the canonical board exam dataset (MedQA) representing the static, full-information medical QA paradigm that MeDxBench expands into open-ended interactive consultations.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). Synthesizes multi-agent collaborative workflows and dynamic reasoning frameworks, contextualizing specialized diagnostic consultation agents within broader autonomous agent architectures.
- Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). Investigates how LLM agents handle incrementally revealed and evolving information in multi-turn interactions, extending the interactive questioning and state-tracking findings of MeDxAgent.
- Paper: Decentralized Multi-Agent Systems with Shared Context, Yuzhen Mao et al. (2026). Explores decentralized multi-agent architectures with shared context, offering potential structural scaling paradigms beyond the structured consultation pipelines evaluated in MeDxAgent.
- Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). Introduces diagnostic safety guardrails and trajectory evaluation for autonomous AI agents, providing essential safety auditing for interactive clinical decision-making agents.
