MedCoT: Medical Chain of Thought via Hierarchical Expert

Jiaxiang LiuYuan WangJiawei DuJoey ZhouZuozhu Liu

article2024EMNLP60 citations

Proposes a hierarchical multi-expert chain-of-thought framework for medical visual question answering that validates step-by-step diagnostic rationales and outperforms models over twenty times its size without requiring manual rationale annotations.

Listen

Medical visual question answering systems aim to interpret clinical imagery and answer diagnostic questions, providing critical support for clinicians and personalized health consultations for patients. However, existing artificial intelligence systems in this domain often struggle because they rely on single-model architectures that produce simple, unverified answers without explaining their reasoning. In high-stakes medical diagnostics, decisions typically require multi-expert collaboration and transparent, step-by-step reasoning rather than unverified single-model outputs.

The article demonstrates and evaluates MedCoT, a hierarchical multi-expert reasoning framework designed to improve both diagnostic accuracy and interpretability in biomedical image analysis without requiring expensive human-annotated reasoning steps. The approach establishes a three-stage workflow: an initial specialist large language model generates a preliminary reasoning path, a follow-up specialist conducts self-reflection to validate and refine that path while adding image captions, and a locally deployed diagnostic model aggregates input from multiple specialized sub-networks through a sparse mixture-of-experts architecture to cast a final diagnostic vote. Experiments were conducted across four medical imaging benchmarks, including VQA-RAD, SLAKE-EN, PathVQA, and Med-VQA-2019.

The findings show that MedCoT consistently outperforms existing state-of-the-art models while remaining highly parameter-efficient. On closed-end diagnostic questions, MedCoT achieved an accuracy of 87.50% on VQA-RAD and 87.26% on SLAKE-EN, outperforming the single Gemini Pro baseline by 27.21% and 14.66%, respectively. Furthermore, with approximately 256 million parameters, MedCoT exceeded the 7-billion-parameter LLaVA-Med model by 5.52% on VQA-RAD and 4.09% on SLAKE-EN. Ablation experiments demonstrated that removing the follow-up specialist reduced accuracy by 6.62% on VQA-RAD, and omitting the mixture-of-experts architecture caused a 4.78% performance loss, notably degrading accuracy on complex organ-specific queries like head-related diagnostics by about 10%.

These results indicate that structured multi-stage verification and specialized expert sub-networks substantially enhance diagnostic reliability and safety while significantly lowering computational overhead. Providing explicit reasoning paths alongside diagnostic decisions allows medical professionals to audit system logic, reducing the clinical risks associated with black-box automated systems. The success of a compact 256-million-parameter diagnostic model also suggests that organizations can achieve superior clinical performance on local hardware without incurring the massive computational and infrastructure costs of multi-billion-parameter models.

Decision-makers should consider adopting collaborative, multi-tiered architectures for automated clinical diagnostics rather than deploying standalone models. Future implementations should focus on combining commercial language models with lightweight local models to balance reasoning power with data governance. However, stakeholders should note that the framework remains susceptible to language model hallucinations when both initial and follow-up specialists agree on incorrect reasoning paths, and the multi-step verification process introduces added computational latency (11.23 seconds per sample compared to 5.02 seconds for standard approaches). Further validation on domain-specific medical foundational models is recommended prior to clinical deployment.

No sufficiently relevant recommendations were found.

Cover for MedCoT: Medical Chain of Thought via Hierarchical Expert

Abstract

Artificial intelligence has advanced in Medical Visual Question Answering (Med-VQA), but prevalent research tends to focus on the accuracy of the answers, often overlooking the reasoning paths and interpretability, which are crucial in clinical settings. Besides, current Med-VQA algorithms, typically reliant on singular models, lack the robustness needed for real-world medical diagnostics which usually require collaborative expert evaluation. To address these shortcomings, this paper presents MedCoT, a novel hierarchical expert verification reasoning chain method designed to enhance interpretability and accuracy in biomedical imaging inquiries. MedCoT is predicated on two principles: The necessity for explicit reasoning paths in Med-VQA and the requirement for multi-expert review to formulate accurate conclusions. The methodology involves an Initial Specialist proposing diagnostic rationales, followed by a Follow-up Specialist who validates these rationales, and finally, a consensus is reached through a vote among a sparse Mixture of Experts within the locally deployed Diagnostic Specialist, which then provides the definitive diagnosis. Experimental evaluations on four standard Med-VQA datasets demonstrate that MedCoT surpasses existing state-of-the-art approaches, providing significant improvements in performance and interpretability. Code is released at https://github.com/JXLiu-AI/MedCoT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Med-VQA
  • 2.2 Multimodal CoT
  • 2.3 MoE
  • 3 Methodology
  • 3.1 Preliminaries
  • 3.2 Initial Specialist
  • 3.3 Follow-up Specialist
  • 3.4 Diagnostic Specialist
  • 3.4.1 Multimodal T5
  • 3.4.2 MoE
  • 4 Experiments
  • 4.1 Experimental Setting
  • 4.2 Main Results
  • 4.3 Ablation Study
  • 4.4 Discussion
  • 5 Conclusion
  • Limitation
  • Acknowledgements
  • References
  • Appendix
  • A MedCoT on Four Datasets
  • B Method Details
  • B.1 Gate Mechanism
  • B.2 MedCoT Method Details
  • C Datasets
  • D The Effect of Initial Specialist
  • E Self-Reflection in Follow-up Specialist
  • E.1 Effectiveness of Self-Reflection in Follow-up Specialist
  • E.2 Error-Analysis on Self-Reflection
  • F Computing Resource Costs
  • G Assessing the Impact of Different Prompts
  • H Prompt Template
  • H.1 Initial Specialist CoT prompt
  • H.2 Follow-up Specialist Prompt
  • H.3 Follow-up specialist Image Caption Prompt

Knowls

  1. Knowl 1 — MedCoT uses a three-stage, expert-verified reasoning chain

    model/method

    MedCoT answers a medical visual question by passing an image, question, and any answer options through three diagnostic stages. The Initial Specialist, implemented with an external large language model, proposes a textual rationale. The Follow-up Specialist checks and, where needed, revises that rationale; it also produces an image caption. A locally deployed Diagnostic Specialist then uses the image, question, rationale, and caption to generate the final answer. This design provides a rationale alongside the answer and does not require manually annotated rationales for training the system. The pipeline diagram on page 3 depicts this progression from preliminary rationale, through review and knowledge infusion, to the local diagnostic model.

  2. Knowl 2 — Follow-up review filters rationales and adds an image caption

    model/method

    For an image II, text input TT (including the question and options), and Initial Specialist rationale RiR_i, the Follow-up Specialist assesses whether the rationale is valid for the image and question. It retains an effective rationale and regenerates an ineffective one; the prompt also allows an effective rationale to be summarized and refined. In compact form, the paper defines the resulting rationale as Rf=RiR_f=R_i when RiR_i is judged effective, and Rf=LLM(T,I,pf)R_f=\mathrm{LLM}(T,I,p_f) otherwise, where pfp_f is the Follow-up Specialist prompt. The Follow-up Specialist separately generates a caption from the image, which is supplied with RfR_f to the Diagnostic Specialist to enrich its textual context and help bridge the image–text modality gap.

  3. Knowl 3 — Diagnostic Specialist combines text and image features before answer generation

    model/method

    The locally deployed Diagnostic Specialist is an encoder–decoder multimodal model. A textual encoder maps the input text TT—question, options, rationale, and caption—to token features FT∈Rn×dF_T\in\mathbb{R}^{n\times d}, where nn is the number of text tokens and dd is the hidden dimension. A visual encoder maps image II to patch features FI∈Rm×dF_I\in\mathbb{R}^{m\times d}, where mm is the number of image patches. Cross-attention uses query, key, and value representations Q,K,VQ,K,V derived respectively from FT,FI,FIF_T,F_I,F_I to form text-conditioned visual features HVatt∈Rn×dH_V^{\mathrm{att}}\in\mathbb{R}^{n\times d}: HVatt=softmax(QK⊤/d)VH_V^{\mathrm{att}}=\mathrm{softmax}(QK^\top/\sqrt{d})V. The fused features are passed to a textual decoder to produce the answer. The architecture diagram on page 4 shows the visual and textual encoders, cross-attention, sparse expert layer, and answer decoder.

  4. Knowl 4 — Sparse MoE routes inputs to selected experts and fuses their feature outputs

    model/method

    Instead of relying on a single learned gate to balance image and text, MedCoT’s Diagnostic Specialist uses a sparse mixture of experts (MoE). A router selects the top kk experts from NN experts for an input. For a feature position ff, let si,fs_{i,f} be the routing score of selected expert ii, and let Ei,fE_{i,f} be that expert’s output. The selected scores are normalized as wi,f=exp⁡(si,f)/∑j=1kexp⁡(sj,f)w_{i,f}=\exp(s_{i,f})/\sum_{j=1}^{k}\exp(s_{j,f}), and the feature-level aggregation is the weighted average Ef=∑i=1kwi,fEi,fE_f=\sum_{i=1}^{k}w_{i,f}E_{i,f}. The resulting image–text mixing coefficient is λf=sigmoid(Ef)\lambda_f=\mathrm{sigmoid}(E_f), and the fused feature is FF,f=(1−λf)FT,f+λfHV,fattF_{F,f}=(1-\lambda_f)F_{T,f}+\lambda_f H_{V,f}^{\mathrm{att}}. Thus, the paper’s “Feature-level Majority Vote” is implemented as a score-weighted feature aggregation, rather than a discrete vote among answer labels. Grid search selected k=2k=2 for all four datasets, with expert counts of 6 for VQA-RAD, 10 for SLAKE-EN, and 5 each for VQA-Med 2019 and PathVQA. The organ-category charts on pages 7 and 14 show input-dependent specialization: on VQA-RAD, MoE exceeded the gate by about 10 percentage points on head-related questions, with Experts 0 and 5 prominent; on SLAKE-EN, the chart reports nearly a 16-point improvement for Brain Face questions, with Experts 2 and 5 prominent.

  5. Knowl 5 — Implementation uses Flan-T5, DETR, and Gemini Pro 1.5 specialists

    experimental setup

    MedCoT’s Diagnostic Specialist uses the Base encoder–decoder architecture of Flan-T5 and DETR ResNet-101 DC5 as its visual encoder; the reported DETR visual representation has shape (100,256)(100,256). The Initial and Follow-up Specialists use Gemini Pro 1.5. The Diagnostic Specialist was trained for 100 epochs with learning rate 8×10−58\times10^{-5} and batch size 8, using PyTorch and HuggingFace on four NVIDIA GeForce RTX 3090 GPUs. The sparse-MoE configuration was searched per dataset; the selected expert counts were 6 for VQA-RAD, 10 for SLAKE-EN, and 5 for each of VQA-Med 2019 and PathVQA, with k=2k=2 throughout. The resulting model has approximately 256M parameters (257M with 6 experts and 261M with 10). Closed-ended questions were evaluated by accuracy; open-ended questions were also evaluated with Rouge and BLEU metrics.

  6. Knowl 6 — Closed-ended accuracy is strong on three datasets but below the best listed PathVQA baseline

    data/table

    The closed-ended benchmark comparison reports accuracy in percent. MedCoT achieves the highest listed score on VQA-RAD, SLAKE-EN, and VQA-Med 2019, while its PathVQA score is below the best listed baseline. The model is approximately 256M parameters, compared with 7B for the cited LLaVA-Med variants.

    Dataset MedCoT Best listed baseline Baseline model
    VQA-RAD 87.50 84.19 LLaVA-Med (From LLaVA)
    SLAKE-EN 87.26 86.78 LLaVA-Med (BioMed CLIP)
    VQA-Med 2019 82.81 81.20 WDAN
    PathVQA 90.37 91.65 LLaVA-Med (From Vicuna)

    The corresponding differences are +3.31, +0.48, and +1.61 percentage points on the first three datasets, and −1.28 points on PathVQA. These comparisons preserve the mixed result in the reported benchmark rather than implying that MedCoT leads on every dataset.

  7. Knowl 7 — MedCoT exceeds MedThink on the reported open-ended metrics

    data/table

    For open-ended questions, where answers may have multiple valid phrasings, the paper compares MedCoT with MedThink using Rouge-1, Rouge-L, Rouge-Lsum, and BLEU-1. MedCoT scores higher on all four reported metrics on both SLAKE-EN and VQA-RAD.

    Dataset Model Rouge-1 Rouge-L Rouge-Lsum BLEU-1
    SLAKE-EN MedCoT 80.86 80.14 80.12 78.33
    SLAKE-EN MedThink 80.12 79.91 79.93 77.94
    VQA-RAD MedCoT 66.30 65.78 65.98 61.29
    VQA-RAD MedThink 58.10 58.09 58.10 51.76

    The largest reported gaps are on VQA-RAD: MedCoT exceeds MedThink by 8.20 points in Rouge-1 and 9.53 points in BLEU-1.

  8. Knowl 8 — Ablations show gains from both follow-up review and sparse MoE

    data/table

    An ablation on VQA-RAD and SLAKE-EN compares combinations of the Follow-up Specialist and sparse MoE in the Diagnostic Specialist. Removing either component lowers accuracy relative to the complete system, and the configuration with both components performs best.

    Follow-up Specialist Sparse MoE VQA-RAD accuracy SLAKE-EN accuracy
    No No 77.57 83.17
    Yes No 82.72 86.05
    No Yes 80.88 83.65
    Yes Yes 87.50 87.26

    With both components present, removing follow-up review reduces VQA-RAD accuracy by 6.62 points and SLAKE-EN accuracy by 4.09 points. Removing MoE reduces them by 4.78 and 0.21 points, respectively. The table therefore supports contributions from both components, with the measured size of each effect depending on dataset and ablation.

  9. Knowl 9 — Initial Specialist rationales improve answers in a 100-question utility check

    empirical result

    To assess whether rationale prompting helps the Initial Specialist, the authors sampled 100 question sets from each of four Med-VQA datasets and an additional mixture set. Correct-answer counts were higher with Initial Specialist chain-of-thought prompting than with LLM prompting without a rationale in every group: VQA-RAD, 72 versus 47; SLAKE-EN, 70 versus 44; VQA-2019, 68 versus 37; PathVQA, 76 versus 49; and the mixture set, 73 versus 46. These results are counts correct out of 100 sampled questions per group, and support the utility of eliciting a rationale in this evaluation.

  10. Knowl 10 — LLM hallucinations remain a risk, and the hierarchical pipeline costs more time

    limitation

    The authors report that the Initial and Follow-up Specialists can hallucinate, and that self-reflection and hierarchical review reduce but do not eliminate this risk. In a documented example, both LLM specialists accepted an incorrect rationale about pneumomediastinum, and the Diagnostic Specialist consequently produced an incorrect answer; agreement between the two LLM stages therefore does not guarantee correctness. The added stages also increase latency relative to the reported baselines: on VQA-RAD, MedCoT used 1,550 tokens per sample and took 11.23 seconds, compared with 1,493 tokens and 5.02 seconds for Vanilla CoT; MedThink took 7.25 seconds, with token cost unreported. Their respective VQA-RAD accuracies were 87.50%, 60.29%, and 83.50%. The paper identifies the extra time and unresolved hallucination risk as limitations.

Coverage note — Prompt-wording robustness tests and the detailed qualitative case studies are omitted as supplementary analyses; the main pipeline, quantitative comparisons, component ablations, MoE specialization findings, and stated limitations are included.

References

  1. 1.Asma Ben Abacha, Sadid A. Hasan, Vivek Datla, Joey Liu, Dina Demner-Fushman, and Henning Müller. 2019a. Vqa-med: Overview of the medical visual question answering task at imageclef 2019. In Conference and Labs of the Evaluation Forum.
  2. 2.Asma Ben Abacha, Sadid A Hasan, Vivek V Datla, Joey Liu, Dina Demner-Fushman, and Henning Müller. 2019b. Vqa-med: Overview of the medical visual question answering task at imageclef 2019. CLEF (working notes), 2(6).
  3. 3.Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, and Chitta Baral. 2021. Weaqa: Weak supervision via captions for visual question answering. Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021.
  4. 4.Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Laila Bashmal, and Mansour Zuair. 2023. Vision–language model for visual question answering in medical imagery. Bioengineering, 10(3):380.
  5. 5.Asma Ben Abacha, Sadid A Hasan, Vivek V Datla, Dina Demner-Fushman, and Henning Müller. 2019. Vqa-med: Overview of the medical visual question answering task at imageclef 2019. In Proceedings of CLEF (Conference and Labs of the Evaluation Forum) 2019 Working Notes. 9-12 September 2019.
  6. 6.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer.
  7. 7.Soravit Changpinyo, Doron Kukliansy, Idan Szpektor, Xi Chen, Nan Ding, and Radu Soricut. 2022. All you may need for vqa are image captions. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1947–1963.
  8. 8.Zhihong Chen, Guanbin Li, and Xiang Wan. 2022. Align, reason and learn: Enhancing medical vision-and-language pre-training with knowledge. In Proceedings of the 30th ACM International Conference on Multimedia, pages 5152–5161.
  9. 9.Sedigheh Eslami, Gerard de Melo, and Christoph Meinel. 2021. Does clip benefit visual question answering in the medical domain as much as it does in the general domain? arXiv preprint arXiv:2112.13906.
  10. 10.Sedigheh Eslami, Christoph Meinel, and Gerard de Melo. 2023. PubMedCLIP: How much does CLIP benefit visual question answering in the medical domain? In Findings of the Association for Computational Linguistics: EACL 2023, pages 1181–1193, Dubrovnik, Croatia. Association for Computational Linguistics.
  11. 11.William Fedus, Jeff Dean, and Barret Zoph. 2022a. A review of sparse expert models in deep learning. arXiv preprint arXiv:2209.01667.
  12. 12.William Fedus, Barret Zoph, and Noam Shazeer. 2022b. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39.
  13. 13.Xiaotang Gai, Chenyi Zhou, Jiaxiang Liu, Yang Feng, Jian Wu, and Zuozhu Liu. 2024. Medthink: Explaining medical visual question answering via multimodal decision-making rationale. arXiv preprint arXiv:2404.12372.
  14. 14.Haifan Gong, Guanqi Chen, Sishuo Liu, Yizhou Yu, and Guanbin Li. 2021. Cross-modal self-attention with multi-task pre-training for medical visual question answering. In Proceedings of the 2021 International Conference on Multimedia Retrieval, pages 456–460.
  15. 15.Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. 2020. Pathvqa: 30000+ questions for medical visual question answering. Preprint, arXiv:2003.10286.
  16. 16.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232.
  17. 17.Xiaofei Huang and Hongfang Gong. 2023. A dual-attention learning network with word and sentence embedding for medical visual question answering. IEEE Transactions on Medical Imaging.
  18. 18.Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991. Adaptive mixtures of local experts. Neural computation, 3(1):79–87.
  19. 19.Yash Khare, Viraj Bagal, Minesh Mathew, Adithi Devi, U Deva Priyakumar, and CV Jawahar. 2021. Mmbert: multimodal bert pretraining for improved medical vqa. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pages 1033–1036. IEEE.
  20. 20.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. Unifiedqa: Crossing format boundaries with a single qa system. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1896–1907.
  21. 21.Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. 2018. Bilinear attention networks. Advances in neural information processing systems, 31.
  22. 22.Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. 2018. A dataset of clinically generated visual questions and answers about radiology images. Scientific data, 5(1):1–10.
  23. 23.Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. Gshard: Scaling giant models with conditional computation and automatic sharding. In International Conference on Learning Representations.
  24. 24.Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2024. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36.
  25. 25.Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. 2021. Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pages 1650–1654. IEEE.
  26. 26.Jiaxiang Liu, Tianxiang Hu, Yan Zhang, Yang Feng, Jin Hao, Junhui Lv, and Zuozhu Liu. 2023a. Parameter-efficient transfer learning for medical visual question answering. IEEE Transactions on Emerging Topics in Computational Intelligence.
  27. 27.Jiaxiang Liu, Tianxiang Hu, Yan Zhang, Xiaotang Gai, YANG FENG, and Zuozhu Liu. 2023b. A chatgpt aided explainable framework for zero-shot medical image diagnosis. In ICML 3rd Workshop on Interpretable Machine Learning in Healthcare (IMLH).
  28. 28.Yunyi Liu, Zhanyu Wang, Dong Xu, and Luping Zhou. 2023c. Q2atransformer: Improving medical vqa via an answer querying decoder. In International Conference on Information Processing in Medical Imaging, pages 445–456. Springer.
  29. 29.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521.
  30. 30.Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023. Chameleon: Plug-and-play compositional reasoning with large language models. arXiv preprint arXiv:2304.09842.
  31. 31.Binh D Nguyen, Thanh-Toan Do, Binh X Nguyen, Tuong Do, Erman Tjiputra, and Quang D Tran. 2019. Overcoming data limitation in medical visual question answering. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 522–530. Springer.
  32. 32.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.
  33. 33.Obioma Pelka, Sven Koitka, Johannes Rückert, Felix Nensa, and Christoph M Friedrich. 2018. Radiology objects in context (roco): a multimodal image dataset. In Intravascular Imaging and Computer Assisted Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis: 7th Joint International Workshop, CVII-STENT 2018 and Third International Workshop, LABELS 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, September 16, 2018, Proceedings 3, pages 180–189. Springer.
  34. 34.Zhangyang Qi, Ye Fang, Mengchen Zhang, Zeyi Sun, Tong Wu, Ziwei Liu, Dahua Lin, Jiaqi Wang, and Hengshuang Zhao. 2023. Gemini vs gpt-4v: A preliminary comparison and combination of vision-language models through qualitative cases. arXiv preprint arXiv:2312.15011.
  35. 35.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  36. 36.Fuji Ren and Yangyang Zhou. 2020. Cgmvqa: A new classification and generative model for medical visual question answering. IEEE Access, 8:50626–50636.
  37. 37.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2016. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In International Conference on Learning Representations.
  38. 38.Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. Preprint, arXiv:2303.11366.
  39. 39.Haoyu Song, Li Dong, Weinan Zhang, Ting Liu, and Furu Wei. 2022. Clip models are few-shot learners: Empirical studies on vqa and visual entailment. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6088–6100.
  40. 40.Anthony Meng Huat Tiong, Junnan Li, Boyang Li, Silvio Savarese, and Steven C.H. Hoi. 2022a. Plug-and-play VQA: Zero-shot VQA by conjoining large pretrained models with zero training. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 951–967, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  41. 41.Anthony Meng Huat Tiong, Junnan Li, Boyang Li, Silvio Savarese, and Steven C.H. Hoi. 2022b. Plug-and-play VQA: Zero-shot VQA by conjoining large pretrained models with zero training. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 951–967, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  42. 42.Tom Van Sonsbeek, Mohammad Mahdi Derakhshani, Ivona Najdenkoska, Cees GM Snoek, and Marcel Worring. 2023. Open-ended medical visual question answering through prefix tuning of language models. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 726–736. Springer.
  43. 43.Zhecan Wang, Bin Xiao, Noel Codella, Jianwei Yang, Yen-Chun Chen, Luowei Zhou, Shih-Fu Chang, Xiyang Dai, Haoxuan You, and Lu Yuan. 2022. Clip-td: Clip targeted distillation for vision-language tasks. In International Conference on Learning Representations.
  44. 44.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45.
  45. 45.Rongwu Xu, Zehan Qi, and Wei Xu. 2024. Preemptive answer "attacks" on chain-of-thought reasoning. Preprint, arXiv:2405.20902.
  46. 46.Li-Ming Zhan, Bo Liu, Lu Fan, Jiaxin Chen, and Xiao-Ming Wu. 2020. Medical visual question answering via conditional reasoning. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2345–2354.
  47. 47.Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. 2023a. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199.
  48. 48.Ruiyuan Zhang, Jiaxiang Liu, Zexi Li, Hao Dong, Jie Fu, and Chao Wu. 2024. Scalable geometric fracture assembly via co-creation space among assemblers. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 7269–7277.
  49. 49.Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023b. Multi-modal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923.
  50. 50.Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang. 2023. Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models. In Thirty-seventh Conference on Neural Information Processing Systems.

Citation

MLA
Liu, J., et al. “MedCoT: Medical Chain of Thought via Hierarchical Expert”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 17371–89, https://doi.org/10.18653/v1/2024.emnlp-main.962.
APA
Liu, J., Wang, Y., Du, J., Zhou, J. T., & Liu, Z. (2024). MedCoT: Medical Chain of Thought via Hierarchical Expert. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 17371–17389. https://doi.org/10.18653/v1/2024.emnlp-main.962
Chicago
Liu, J., Y. Wang, J. Du, J. T. Zhou, and Z. Liu. 2024. “MedCoT: Medical Chain of Thought via Hierarchical Expert”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 17371–89. https://doi.org/10.18653/v1/2024.emnlp-main.962.
Harvard
Liu, J. et al. (2024) “MedCoT: Medical Chain of Thought via Hierarchical Expert”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 17371–17389. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.962.
Vancouver
1. Liu J, Wang Y, Du J, Zhou JT, Liu Z (2024) MedCoT: Medical Chain of Thought via Hierarchical Expert. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 17371–17389

BibTeX

@inproceedings{liu-etal-2024-medcot,
    title = "{M}ed{C}o{T}: Medical Chain of Thought via Hierarchical Expert",
    author = "Liu, Jiaxiang  and
      Wang, Yuan  and
      Du, Jiawei  and
      Zhou, Joey Tianyi  and
      Liu, Zuozhu",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.962/",
    doi = "10.18653/v1/2024.emnlp-main.962",
    pages = "17371--17389"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/