keyword
reasoning chain
A reasoning chain is a sequential series of intermediate logical steps, rationales, or subtasks used to progress from an initial premise or question to a final conclusion. In artificial intelligence and cognitive systems, reasoning chains decompose complex problems into explicit, manageable stages of deduction, evidence evaluation, or calculation. This step-by-step progression enhances the interpretability and transparency of decision-making, facilitates the verification of intermediate claims, and improves accuracy on multi-step tasks. Because each stage builds upon preceding assertions, a reasoning chain provides a structured path for composing simpler tasks into comprehensive solutions while also linking the validity of the final outcome directly to the correctness of each intermediate step.
3 items

MedCoT: Medical Chain of Thought via Hierarchical Expert
Jiaxiang Liu, Yuan Wang, Jiawei Du, Joey Zhou, Zuozhu Liu
Why you should read this
Proposes a hierarchical multi-expert chain-of-thought framework for medical visual question answering that validates step-by-step diagnostic rationales and outperforms models over twenty times its size without requiring manual rationale annotations.
Artificial intelligence has advanced in Medical Visual Question Answering (Med-VQA), but prevalent research tends to focus on the accuracy of the answers, often overlooking the reasoning paths and interpretability, which are crucial in clinical settings. Besides, current Med-VQA algorithms, typically reliant on singular models, lack the robustness needed for real-world medical diagnostics which usually require collaborative expert evaluation. To address these shortcomings, this paper presents MedCoT, a novel hierarchical expert verification reasoning chain method designed to enhance interpretability and accuracy in biomedical imaging inquiries. MedCoT is predicated on two principles: The necessity for explicit reasoning paths in Med-VQA and the requirement for multi-expert review to formulate accurate conclusions. The methodology involves an Initial Specialist proposing diagnostic rationales, followed by a Follow-up Specialist who validates these rationales, and finally, a consensus is reached through a vote among a sparse Mixture of Experts within the locally deployed Diagnostic Specialist, which then provides the definitive diagnosis. Experimental evaluations on four standard Med-VQA datasets demonstrate that MedCoT surpasses existing state-of-the-art approaches, providing significant improvements in performance and interpretability. Code is released at https://github.com/JXLiu-AI/MedCoT.
Added
2026-10-03

Demonstrations, CoT, and Prompting: A Theoretical Analysis of ICL
Xuhan Tong, Yuchen Zeng, Jiawei Zhang
Why you should read this
Establishes theoretical generalization bounds for in-context learning under mild assumptions, explaining how demonstration selection, Chain-of-Thought task decomposition, and prompt templates directly govern model performance on unseen tasks.
In-Context Learning (ICL) enables pretrained LLMs to adapt to downstream tasks by conditioning on a small set of input-output demonstrations, without any parameter updates. Although there have been many theoretical efforts to explain how ICL works, most either rely on strong architectural or data assumptions, or fail to capture the impact of key practical factors such as demonstration selection, Chain-of-Thought (CoT) prompting, the number of demonstrations, and prompt templates. We address this gap by establishing a theoretical analysis of ICL under mild assumptions that links these design choices to generalization behavior. We derive an upper bound on the ICL test loss, showing that performance is governed by (i) the quality of selected demonstrations, quantified by Lipschitz constants of the ICL loss along paths connecting test prompts to pretraining samples, (ii) an intrinsic ICL capability of the pretrained model, and (iii) the degree of distribution shift. Within the same framework, we analyze CoT prompting as inducing a task decomposition and show that it is beneficial when demonstrations are well chosen at each substep and the resulting subtasks are easier to learn. Finally, we characterize how ICL performance sensitivity to prompt templates varies with the number of demonstrations. Together, our study shows that pretraining equips the model with the ability to generalize beyond observed tasks, while CoT enables the model to compose simpler subtasks into more complex ones, and demonstrations and instructions enable it to retrieve similar or complex tasks, including those that can be composed into more complex ones, jointly supporting generalization to unseen tasks. All theoretical insights are corroborated by experiments.
Added
2026-09-30

How Language Model Hallucinations Can Snowball
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, Noah A. Smith
Why you should read this
Reveals that large language models frequently invent false justifications to maintain consistency with their own earlier mistakes, generating secondary errors that they are otherwise capable of correctly identifying as false in isolation.
A major risk of using language models in practical applications is their tendency to hallucinate incorrect statements. Hallucinations are often attributed to knowledge gaps in LMs, but we show that LMs sometimes produce hallucinations that they can separately recognize as incorrect. To do this, we construct three question-answering datasets where LMs often state an incorrect answer which is followed by an explanation with at least one incorrect claim. Crucially, we find that GPT-3.5, GPT-4, and LLaMA2-70B-chat can identify 67%, 87%, and 94% of these incorrect claims, respectively. We show that this phenomenon doesn’t disappear under higher temperatures sampling, beam search, and zero-shot chain-of-thought prompting. These findings reveal that LM hallucinations can snowball: early mistakes by an LM can lead to more mistakes that otherwise would not be made.
Added
2026-09-28
