MedCoT: Medical Chain of Thought via Hierarchical Expert
Jiaxiang LiuYuan WangJiawei DuJoey ZhouZuozhu Liu
Proposes a hierarchical multi-expert chain-of-thought framework for medical visual question answering that validates step-by-step diagnostic rationales and outperforms models over twenty times its size without requiring manual rationale annotations.
Medical visual question answering systems aim to interpret clinical imagery and answer diagnostic questions, providing critical support for clinicians and personalized health consultations for patients. However, existing artificial intelligence systems in this domain often struggle because they rely on single-model architectures that produce simple, unverified answers without explaining their reasoning. In high-stakes medical diagnostics, decisions typically require multi-expert collaboration and transparent, step-by-step reasoning rather than unverified single-model outputs.
The article demonstrates and evaluates MedCoT, a hierarchical multi-expert reasoning framework designed to improve both diagnostic accuracy and interpretability in biomedical image analysis without requiring expensive human-annotated reasoning steps. The approach establishes a three-stage workflow: an initial specialist large language model generates a preliminary reasoning path, a follow-up specialist conducts self-reflection to validate and refine that path while adding image captions, and a locally deployed diagnostic model aggregates input from multiple specialized sub-networks through a sparse mixture-of-experts architecture to cast a final diagnostic vote. Experiments were conducted across four medical imaging benchmarks, including VQA-RAD, SLAKE-EN, PathVQA, and Med-VQA-2019.
The findings show that MedCoT consistently outperforms existing state-of-the-art models while remaining highly parameter-efficient. On closed-end diagnostic questions, MedCoT achieved an accuracy of 87.50% on VQA-RAD and 87.26% on SLAKE-EN, outperforming the single Gemini Pro baseline by 27.21% and 14.66%, respectively. Furthermore, with approximately 256 million parameters, MedCoT exceeded the 7-billion-parameter LLaVA-Med model by 5.52% on VQA-RAD and 4.09% on SLAKE-EN. Ablation experiments demonstrated that removing the follow-up specialist reduced accuracy by 6.62% on VQA-RAD, and omitting the mixture-of-experts architecture caused a 4.78% performance loss, notably degrading accuracy on complex organ-specific queries like head-related diagnostics by about 10%.
These results indicate that structured multi-stage verification and specialized expert sub-networks substantially enhance diagnostic reliability and safety while significantly lowering computational overhead. Providing explicit reasoning paths alongside diagnostic decisions allows medical professionals to audit system logic, reducing the clinical risks associated with black-box automated systems. The success of a compact 256-million-parameter diagnostic model also suggests that organizations can achieve superior clinical performance on local hardware without incurring the massive computational and infrastructure costs of multi-billion-parameter models.
Decision-makers should consider adopting collaborative, multi-tiered architectures for automated clinical diagnostics rather than deploying standalone models. Future implementations should focus on combining commercial language models with lightweight local models to balance reasoning power with data governance. However, stakeholders should note that the framework remains susceptible to language model hallucinations when both initial and follow-up specialists agree on incorrect reasoning paths, and the multi-step verification process introduces added computational latency (11.23 seconds per sample compared to 5.02 seconds for standard approaches). Further validation on domain-specific medical foundational models is recommended prior to clinical deployment.
No sufficiently relevant recommendations were found.
- Paper: VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge, Vishwesh Nath et al. (2025). VILA-M3 carries medical expert collaboration into conversational vision-language systems by routing specialist-tool findings back into the model, extending MedCoT’s hierarchical expert reasoning in a broader clinical workflow.
- Paper: MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency, Dongzhi Jiang et al. (2025). MME-CoT extends the chain-of-thought evaluation problem by measuring intermediate reasoning quality, robustness, and efficiency rather than judging only final-answer accuracy.
