3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis
Ziyue WangLinghan CaiChang Han LowHao LiuJunde WuJingyu WangRui WangLei SongJiang BianJingjing Fu
Presents 3DMedAgent, an agent-based framework that enables standard 2D multimodal large language models to perform 3D CT analysis without volumetric fine-tuning by coordinating specialized tools through structured memory and multi-step reasoning.
Three-dimensional medical imaging, such as computed tomography (CT), is essential for modern diagnosis, but manually reviewing dense volumetric scans slice-by-slice places an unsustainable workload on radiologists and increases the risk of diagnostic errors. While multimodal large language models (MLLMs) offer strong potential for automated medical reasoning, existing models are fundamentally designed for two-dimensional inputs. Adapting them to volumetric 3D scans typically requires expensive 3D training that compresses fine-grained anatomical details and encourages superficial pattern matching, leading to brittle performance across diverse clinical settings.
The article demonstrates and evaluates 3DMedAgent, a modular system that enables standard 2D multimodal models to perform comprehensive 3D CT analysis—ranging from basic physical measurements to high-level clinical reasoning—without requiring 3D-specific model fine-tuning.
The approach employs a query-adaptive agent that coordinates specialized visual and textual tools, decomposing complex volumetric analysis into structured evidence. It first establishes an Organ-Aware Memory Initialization to ground major anatomical structures and spatial boundaries, followed by Coarse-to-Fine Lesion Targeting that uses dense similarity heatmaps to narrow candidate regions. For ambiguous cases, an iterative loop selects informative individual slices for focused visual verification. The resulting findings are stored in a long-term structured memory to guide multi-step clinical reasoning. To rigorously evaluate the system alongside an abdominal benchmark (DeepTumorVQA), the authors introduced DeepChestVQA, a thoracic CT benchmark spanning 1,020 visual question-answering pairs across 17 clinical capability dimensions.
The key findings show substantial performance gains. When powered by GPT-5, 3DMedAgent achieved an overall accuracy of 66% on the abdominal benchmark and 57% on the thoracic benchmark, delivering an average improvement of over 20 percentage points compared to baseline general, medical, and 3D-specific models, which frequently performed near random guess levels. On complex medical reasoning tasks, the agent improved accuracy by more than 27 percentage points over baselines. The system demonstrated strong generalizability across diverse organ systems and multi-center data sources. Furthermore, validation with medical experts confirmed that the slices autonomously selected by the system for visual verification closely matched the preferences of experienced radiologists.
These results imply that building tool-augmented reasoning agents is a more effective and scalable strategy for 3D medical AI than end-to-end 3D model fine-tuning. By grounding high-level clinical diagnoses in verified visual evidence, the framework improves diagnostic transparency and reduces the risk of model hallucinations. This evidence-based approach lowers development costs by leveraging existing 2D models while providing reliable decision support that could significantly shorten diagnostic review times in clinical workflows.
Before clinical deployment, healthcare organizations and developers should conduct prospective pilot validations in live workflows under direct human supervision. Next development steps should focus on optimizing the agent's routing policy using reinforcement learning or supervised fine-tuning, expanding the available tool suite, and improving multi-structure spatial reasoning.
The findings are subject to several limitations. The system's performance depends on the accuracy of its underlying visual tools, meaning segmentation errors can propagate into downstream reasoning. Additionally, on tasks requiring complex 3D spatial adjacency between organs, relying on selected 2D slices can occasionally underperform compared to models relying on memorized anatomical priors. Readers can place high confidence in the benchmark improvements, but automated outputs should continue to operate strictly as clinical decision support alongside human medical review.
- Paper: VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge, Vishwesh Nath et al. (2025). It introduces a framework for equipping multimodal language models with dynamic medical expert tools, laying foundational agentic principles adapted by 3DMedAgent.
- Paper: LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day, Chunyuan Li et al. (2023). It establishes foundational multimodal vision-language instruction tuning for biomedical imaging, representing the core 2D medical MLLM baseline expanded to 3D analysis.
- Paper: Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image Analysis, Yucheng Tang et al. (2022). It introduces self-supervised volumetric transformer backbones for 3D computed tomography, providing the specialized 3D visual perception tools utilized by agent workflows.
- Paper: Segment anything in medical images, Jun Ma et al. (2023). It provides the foundational prompt-driven medical segmentation tools that 3DMedAgent dynamically invokes to isolate regional anatomical visual evidence.
- Paper: Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale, Junying Chen et al. (2024). It presents methods for scaling visual-linguistic clinical reasoning in MLLMs, establishing the baseline capabilities of 2D models before agentic multi-step 3D extension.
- Paper: MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification, Jiancheng Yang et al. (2021). It provides a standardized benchmark bridging 2D and 3D biomedical image evaluation, establishing the comparative gap between direct 3D modeling and multi-view 2D methods.
- Paper: EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images, Seongsu Bae et al. (2023). It introduces executable query decomposition and tool-use interfaces for multi-modal clinical question answering, foreshadowing agent-based medical evidence aggregation.
No sufficiently relevant recommendations were found.
