3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis

Ziyue WangLinghan CaiChang Han LowHao LiuJunde WuJingyu WangRui WangLei SongJiang BianJingjing Fu

article2026arXiv5 citations

Presents 3DMedAgent, an agent-based framework that enables standard 2D multimodal large language models to perform 3D CT analysis without volumetric fine-tuning by coordinating specialized tools through structured memory and multi-step reasoning.

Listen

Three-dimensional medical imaging, such as computed tomography (CT), is essential for modern diagnosis, but manually reviewing dense volumetric scans slice-by-slice places an unsustainable workload on radiologists and increases the risk of diagnostic errors. While multimodal large language models (MLLMs) offer strong potential for automated medical reasoning, existing models are fundamentally designed for two-dimensional inputs. Adapting them to volumetric 3D scans typically requires expensive 3D training that compresses fine-grained anatomical details and encourages superficial pattern matching, leading to brittle performance across diverse clinical settings.

The article demonstrates and evaluates 3DMedAgent, a modular system that enables standard 2D multimodal models to perform comprehensive 3D CT analysis—ranging from basic physical measurements to high-level clinical reasoning—without requiring 3D-specific model fine-tuning.

The approach employs a query-adaptive agent that coordinates specialized visual and textual tools, decomposing complex volumetric analysis into structured evidence. It first establishes an Organ-Aware Memory Initialization to ground major anatomical structures and spatial boundaries, followed by Coarse-to-Fine Lesion Targeting that uses dense similarity heatmaps to narrow candidate regions. For ambiguous cases, an iterative loop selects informative individual slices for focused visual verification. The resulting findings are stored in a long-term structured memory to guide multi-step clinical reasoning. To rigorously evaluate the system alongside an abdominal benchmark (DeepTumorVQA), the authors introduced DeepChestVQA, a thoracic CT benchmark spanning 1,020 visual question-answering pairs across 17 clinical capability dimensions.

The key findings show substantial performance gains. When powered by GPT-5, 3DMedAgent achieved an overall accuracy of 66% on the abdominal benchmark and 57% on the thoracic benchmark, delivering an average improvement of over 20 percentage points compared to baseline general, medical, and 3D-specific models, which frequently performed near random guess levels. On complex medical reasoning tasks, the agent improved accuracy by more than 27 percentage points over baselines. The system demonstrated strong generalizability across diverse organ systems and multi-center data sources. Furthermore, validation with medical experts confirmed that the slices autonomously selected by the system for visual verification closely matched the preferences of experienced radiologists.

These results imply that building tool-augmented reasoning agents is a more effective and scalable strategy for 3D medical AI than end-to-end 3D model fine-tuning. By grounding high-level clinical diagnoses in verified visual evidence, the framework improves diagnostic transparency and reduces the risk of model hallucinations. This evidence-based approach lowers development costs by leveraging existing 2D models while providing reliable decision support that could significantly shorten diagnostic review times in clinical workflows.

Before clinical deployment, healthcare organizations and developers should conduct prospective pilot validations in live workflows under direct human supervision. Next development steps should focus on optimizing the agent's routing policy using reinforcement learning or supervised fine-tuning, expanding the available tool suite, and improving multi-structure spatial reasoning.

The findings are subject to several limitations. The system's performance depends on the accuracy of its underlying visual tools, meaning segmentation errors can propagate into downstream reasoning. Additionally, on tasks requiring complex 3D spatial adjacency between organs, relying on selected 2D slices can occasionally underperform compared to models relying on memorized anatomical priors. Readers can place high confidence in the benchmark improvements, but automated outputs should continue to operate strictly as clinical decision support alongside human medical review.

No sufficiently relevant recommendations were found.

Cover for 3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis

Abstract

3D CT analysis spans a continuum from low-level perception to high-level clinical understanding. Existing 3D-oriented analysis methods adopt either isolated task-specific modeling or task-agnostic end-to-end paradigms to produce one-hop outputs, impeding the systematic accumulation of perceptual evidence for downstream reasoning. In parallel, recent multimodal large language models (MLLMs) exhibit improved visual perception and can integrate visual and textual information effectively, yet their predominantly 2D-oriented designs fundamentally limit their ability to perceive and analyze volumetric medical data. To bridge this gap, we propose 3DMedAgent, a unified agent that enables 2D MLLMs to perform general 3D CT analysis without 3D-specific fine-tuning. 3DMedAgent coordinates heterogeneous visual and textual tools through a flexible MLLM agent, progressively decomposing complex 3D analysis into tractable subtasks that transition from global to regional views, from 3D volumes to informative 2D slices, and from visual evidence to structured textual representations. Central to this design, 3DMedAgent maintains a long-term structured memory that aggregates intermediate tool outputs and supports query-adaptive, evidence-driven multi-step reasoning. We further introduce the DeepChestVQA benchmark for evaluating unified perception-to-understanding capabilities in 3D thoracic imaging. Experiments across over 40 tasks demonstrate that 3DMedAgent consistently outperforms general, medical, and 3D-specific MLLMs, highlighting a scalable path toward general-purpose 3D clinical this http URL and data are available at \href{this https URL}{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Medical Vision-Language Models
  • 2.2 Medical Agentic Systems
  • 2.3 Medical Benchmarks
  • 3 DeepChestVQA Benchmark Construction
  • 4 Methodology
  • 4.1 Organ-Aware Memory Initialization
  • 4.2 Coarse-to-Fine Lesion Targeting
  • 4.3 Think-with-1-Slice Loop
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Comparison with MLLMs
  • 5.3 Generalization Ability Evaluation
  • 5.4 Ablation Study on Key Components
  • 5.5 Detailed Analysis of CFLT
  • 5.6 Detailed Analysis of T1S-Loop
  • 6 Conclusion
  • 7 Impact Statement
  • References
  • A Details of DeepChestVQA Dataset Construction
  • A.1 Recognition Question Type
  • A.2 Visual Reasoning Question Type
  • A.3 Medical Reasoning Question Type
  • B Details of DeepTumorVQA Dataset Usage
  • B.1 Source Datasets
  • B.2 The preprocessing details of DeepTumorVQA
  • C Limitations and Discussions

Knowls

  1. Knowl 1 — 3DMedAgent Framework for Volumetric CT Analysis

    model/method

    3DMedAgent is an agentic framework designed to enable 2D multimodal large language models (MLLMs) to analyze 3D computed tomography (CT) scans across both low-level perception and high-level clinical understanding without fine-tuning on 3D data. The framework decouples perception from clinical reasoning by coordinating heterogeneous specialized visual tools and maintaining a long-term structured memory that stores compact textual and spatial evidence.

    The system operates across three progressive stages:

    1. Organ-Aware Memory Initialization (OAMI): Global anatomical parsing tools segment major organs to extract size, attenuation (Hounsfield Units), and cranio-caudal span into an initial memory bank.
    2. Coarse-to-Fine Lesion Targeting (CFLT): For lesion-specific queries, visual-language 3D embeddings generate dense spatial activation heatmaps, which are constrained by organ bounds and filtered to identify high-probability regions of interest (ROIs) and key axial slices.
    3. Think-with-1-Slice Loop (T1S-Loop): If ambiguity remains, the agent iteratively selects informative 2D slices, invokes 2D visual tools (such as crop-and-zoom or mask overlays), performs slice-level verification, and updates its memory to reach a final clinical conclusion.
  2. Knowl 2 — Organ-Aware Memory Initialization

    model/method

    Organ-Aware Memory Initialization (OAMI) generates a standardized, organ-level global overview of a 3D CT volume to initialize the agent's memory bank M0\mathcal{M}_0. Given a CT volume, an anatomical grounding model (such as VISTA3D) extracts 3D segmentation masks YoY_o for a set of major anatomical organs O\mathcal{O}. For each segmented organ o∈Oo \in \mathcal{O}, predefined deterministic functions calculate its volume/size SoS_o, its mean Hounsfield Unit attenuation value HoH_o, and its cranio-caudal extent along the z-axis [zmin⁡(o),zmax⁡(o)][z_{\min}(o), z_{\max}(o)].

    The initialized memory is formulated as: M0={mo∣o∈O},mo=(So,Ho,[zmin⁡(o),zmax⁡(o)])\mathcal{M}_0 = \{ m_o \mid o \in \mathcal{O} \}, \quad m_o = \left( S_o, H_o, [z_{\min}(o), z_{\max}(o)] \right)

    Lesion segmentation is deliberately omitted from initial memory creation because lesion taxonomies and boundaries exhibit high variability across clinical datasets, risking the injection of noisy priors. In contrast, standardized organ segmentations provide robust global spatial and physical baselines to answer volumetric measurement queries and guide subsequent regional searches.

  3. Knowl 3 — Coarse-to-Fine Lesion Targeting

    model/method

    Coarse-to-Fine Lesion Targeting (CFLT) localizes candidate lesion slices and regions of interest (ROIs) using a pretrained 3D vision-language alignment model (CT-CLIP). Given a CT volume V∈RH×W×DV \in \mathbb{R}^{H \times W \times D}, the 3D vision encoder extracts a dense local feature map F=fimg(V)∈Rh×w×d×nF = f_{\text{img}}(V) \in \mathbb{R}^{h \times w \times d \times n}, while a clinical text prompt pp (e.g., describing a target lesion) is mapped by text encoder ftxtf_{\text{txt}} to text embedding t=ftxt(p)∈Rnt = f_{\text{txt}}(p) \in \mathbb{R}^n.

    Each normalized local patch embedding f^i,j,k∈Rn\hat{f}_{i,j,k} \in \mathbb{R}^n is multiplied with the normalized text embedding t^∈Rn\hat{t} \in \mathbb{R}^n via cosine similarity to construct a dense 3D similarity heatmap H∈Rh×w×d\mathcal{H} \in \mathbb{R}^{h \times w \times d}: Hi,j,k=f^i,j,k⊤t^\mathcal{H}_{i,j,k} = \hat{f}_{i,j,k}^\top \hat{t}

    Targeting proceeds via two steps:

    1. Volume Cropping: The agent infers target organ oo from the user query and uses memory [zmin⁡(o),zmax⁡(o)][z_{\min}(o), z_{\max}(o)] to prune all activations outside the organ's z-axis boundaries, producing an organ-masked heatmap H(o)\mathcal{H}^{(o)}.
    2. Lesion Targeting: Candidate ROIs RR (e.g., axial slices or anatomical sub-segments) are projected onto the grid as Π(R)\Pi(R). Discarding patches below an activation threshold τ\tau, each ROI is scored by weighting the heatmap response by the patch organ overlap ratio ρ(P)=∣P∩Yo∣∣P∣\rho(P) = \frac{|P \cap Y_o|}{|P|}: S(R)=∑P⊂Π(R)I[H(P)≥τ]ρ(P)H(P)S(R) = \sum_{P \subset \Pi(R)} \mathbb{I}[\mathcal{H}(P) \ge \tau] \rho(P) \mathcal{H}(P)

    Candidate ROIs are ranked by S(R)S(R), and top candidates are appended to memory M0\mathcal{M}_0 to produce updated lesion memory Mℓ\mathcal{M}_\ell.

  4. Knowl 4 — Think-with-1-Slice Reasoning and Verification Loop

    algorithm

    The Think-with-1-Slice Loop (T1S-Loop) performs on-demand, slice-level visual verification when the aggregated memory evidence is insufficient to finalize an answer. The loop is bounded by maximum turns Tmax⁡=5T_{\max} = 5.

    Input: Query qq, memory with lesion candidates Mℓ\mathcal{M}_\ell, max turns Tmax⁡=5T_{\max} = 5, MLLM agent AθA_\theta, router Rθ\mathcal{R}_\theta, visual tools T\mathcal{T}
    Output: Final clinical answer y^\hat{y}
    Initialize t←0t \leftarrow 0, Mℓ0←Mℓ\mathcal{M}_\ell^0 \leftarrow \mathcal{M}_\ell
    while t<Tmax⁡t < T_{\max} do
        Generate reasoning state: ut←Aθ(q,Mℓt)=(rt,y^t,Et,At)u_t \leftarrow A_\theta(q, \mathcal{M}_\ell^t) = (r_t, \hat{y}_t, \mathcal{E}_t, \mathcal{A}_t)
        Update memory: Mℓt←Mℓt∪{ut}\mathcal{M}_\ell^t \leftarrow \mathcal{M}_\ell^t \cup \{ u_t \}
        Determine action: (bt,τt)←Rθ(ut,T)(b_t, \tau_t) \leftarrow \mathcal{R}_\theta(u_t, \mathcal{T})
        if bt=0b_t = 0 then
            return y^t\hat{y}_t
        end if
        Select candidate ROI/slice rtr_t from Mℓt\mathcal{M}_\ell^t
        Apply visual tool τt∈T\tau_t \in \mathcal{T} (e.g., crop-and-zoom or mask overlay) to rtr_t
        Conduct slice multimodal reasoning: ut+1←Aθ(q,Mℓt,τt(rt))u_{t+1} \leftarrow A_\theta(q, \mathcal{M}_\ell^t, \tau_t(r_t))
        Update memory Mℓt+1←Mℓt∪{ut+1}\mathcal{M}_\ell^{t+1} \leftarrow \mathcal{M}_\ell^t \cup \{ u_{t+1} \}
        Drop used slice rtr_t from candidates
        t←t+1t \leftarrow t + 1
    end while
    return y^Tmax⁡\hat{y}_{T_{\max}}

    In each turn tt, rtr_t denotes the rationale, y^t\hat{y}_t the current answer, Et\mathcal{E}_t the supporting evidence retrieved from memory, At\mathcal{A}_t explicit assumptions made due to missing evidence, and bt∈{0,1}b_t \in \{0, 1\} indicates whether new visual tool acquisition is required.

  5. Knowl 5 — DeepChestVQA Benchmark Specification

    definition

    DeepChestVQA is a 3D thoracic computed tomography (CT) benchmark designed to evaluate multimodal perception, localization, and medical reasoning across 17 distinct capability dimensions. The benchmark comprises 892 CT volumes and 1,020 multiple-choice question-answer pairs sourced from CT-RATE, ReXGroundingCT, and the NSCLC-Radiomics dataset.

    The benchmark is structured into three primary question categories:

    • Recognition (3 subtypes, 180 QA pairs): Binary existence tasks for abnormalities in specific anatomical regions (Bronchus Lesion Existence, Lung Lesion Existence, Pleura Lesion Existence).
    • Visual Reasoning (8 subtypes, 480 QA pairs): Fine-grained spatial and physical properties derived from 3D connected components and masks (Largest Lesion Diameter, Largest Lesion Location across 5 lung lobes, Largest Lesion Slice percentile, Lesion Counting, Lesion Counting by Location, Organ Enlargement, Organ Atrophy, Lesion-Organ Hounsfield Unit Difference).
    • Medical Reasoning (6 subtypes, 360 QA pairs): Multi-step clinical interpretation tasks (Attenuation Pattern Classification, Volume-Loss Lesion Classification, Imaging Phenotype Analysis, Phenotype Mixing Identification, Emphysema Severity Grading, Pleural Effusion Grading).

    Ground-truth labels for questions are constructed deterministically using validated organ and lesion masks along with metadata rules without needing full-volume inference at annotation time, and closed-set multiple-choice options are balanced across answer classes.

  6. Knowl 6 — Empirical Performance on DeepTumorVQA and DeepChestVQA

    data/table

    Zero-shot performance comparison of 3DMedAgent against general multimodal LLMs (GPT-5, Qwen3-VL-30B), 2D medical MLLMs (MedGemma-27B, HuatuoGPT-Vision-34B), and 3D medical MLLMs (M3D, RadFM) across DeepTumorVQA (1,740 QA pairs) and DeepChestVQA (1,020 QA pairs). Values report Multiple-Choice Question (MCQ) accuracy.

    Dataset / Question Type Rand GPT-5 Qwen3-VL MedGemma HuatuoGPT M3D RadFM 3DMedAgent (GPT-5)
    DeepTumorVQA
    Measurement 0.25 0.33 0.35 0.29 0.26 0.28 0.31 0.63
    Recognition 0.50 0.49 0.51 0.51 0.54 0.53 0.47 0.76
    Visual Reasoning 0.34 0.39 0.35 0.38 0.36 0.34 0.34 0.56
    Medical Reasoning 0.40 0.43 0.45 0.35 0.36 0.42 0.37 0.70
    Total Average 0.37 0.41 0.42 0.38 0.38 0.39 0.37 0.66
    DeepChestVQA
    Recognition 0.50 0.52 0.53 0.50 0.53 0.52 0.47 0.69
    Visual Reasoning 0.31 0.38 0.32 0.31 0.29 0.32 0.30 0.49
    Medical Reasoning 0.40 0.39 0.45 0.41 0.40 0.38 0.35 0.53
    Total Average 0.40 0.43 0.43 0.41 0.41 0.41 0.37 0.57

    Direct application of 2D and 3D MLLM baselines achieves near-random performance due to 2D context loss or domain shift in compressed 3D tokenizers. 3DMedAgent (GPT-5) achieves average accuracy gains of +25% on DeepTumorVQA and +14% on DeepChestVQA over the strongest baseline models.

  7. Knowl 7 — Ablation of 3DMedAgent Key Modules

    data/table

    Incremental ablation study analyzing the contribution of Organ-Aware Memory Initialization (OAMI), Coarse-to-Fine Lesion Targeting (CFLT), and Think-with-1-Slice Loop (T1S-Loop) using GPT-5 as the core agent across DeepTumorVQA and DeepChestVQA benchmarks.

    Setting MLLM Input DeepTumorVQA DeepChestVQA
    OAMI CFLT T1S-Loop Text / Image Mea. Rec. Vis. Reason. Med. Reason. Rec. Vis. Reason. Med. Reason.
    - - - Image + Text 0.33 0.49 0.38 0.43 0.52 0.38 0.39
    ✓ - - Text 0.54 0.49 0.43 0.62 0.51 0.44 0.45
    ✓ ✓ - Text 0.60 0.73 0.54 0.66 0.67 0.46 0.50
    ✓ ✓ ✓ Image + Text 0.63 0.76 0.56 0.70 0.69 0.49 0.53

    OAMI provides initial global organ priors that drive large accuracy gains in measurement (from 0.33 to 0.54) and medical reasoning (from 0.43 to 0.62). CFLT provides precise lesion spatial localization, substantially improving recognition (from 0.49 to 0.73 on DeepTumorVQA; 0.51 to 0.67 on DeepChestVQA). T1S-Loop adds targeted 2D visual verification, providing the highest accuracy across all tasks.

  8. Knowl 8 — DeepTumorVQA Preprocessing and Curation Strategy

    experimental setup

    To eliminate systematic evaluation biases present in the raw DeepTumorVQA dataset (such as 97% class imbalance toward cysts in lesion classification and redundant linguistic paraphrases), a refined subset was constructed consisting of 1,740 VQA pairs from 1,280 CT volumes across 29 subtypes using four curation principles:

    1. Source Balancing: CT scans are sampled across more than 20 public abdominal CT source datasets in an approximately equal ratio to avoid overfitting to dominant source distributions.
    2. Paraphrase Deduplication: At most one question instance is sampled per question subtype for a specific CT case, removing artificial scale inflation from linguistic rewording.
    3. Volume Diversity: The number of VQA items per CT volume is capped to prevent information leakage across repeated questions on the same patient scan.
    4. Content-Based Option Balancing: For fixed closed-set multiple-choice questions (e.g., T-stage categories T1-T4), sample counts are balanced across the underlying semantic categories rather than superficial option letters.
  9. Knowl 9 — Radiologist Slice Selection Alignment and T1S-Loop Dynamics

    empirical result

    Empirical analysis of 3DMedAgent's intermediate stages demonstrates high clinical alignment and adaptive compute allocation:

    1. Expert Slice Alignment: Comparing CFLT-selected candidate slices against slice preferences independently annotated by two certified radiologists (R1 and R2) shows that the top-3 candidate slices chosen by CFLT achieve an agreement rate with radiologists that closely approaches the inter-radiologist agreement rate (R1 vs. R2R1 \text{ vs. } R2).
    2. Adaptive Compute Allocation: Harder cases are automatically routed to higher iteration turns kk in the T1S-Loop. While baseline accuracy without visual verification drops steeply on cases requiring more turns (reflecting case difficulty), enabling T1S-Loop provides consistent accuracy improvements across all iteration turns, with the steepest marginal gains occurring in the first 1-2 turns.
  10. Knowl 10 — Limitations of 3DMedAgent

    limitation

    3DMedAgent exhibits three primary limitations:

    1. Tool Dependency and Error Propagation: Initial reasoning stages (OAMI and CFLT) depend directly on the segmentation accuracy of VISTA3D and the localization fidelity of CT-CLIP. Errors or domain shifts in these vision modules directly propagate inaccurate evidence to the long-term memory.
    2. Spatial and Relational Reasoning Constraints: Reasoning over a restricted set of individual 2D slices limits the agent's ability to interpret complex 3D topological relationships across multiple organs (e.g., determining which organ lies immediately adjacent to a pancreatic lesion), where baseline MLLMs leveraging global prior heuristics sometimes perform comparably or better.
    3. Fixed Prompt-Based Policy: Action routing, tool invocation, and evidence distillation are governed by zero-shot prompting rather than a learned reinforcement learning policy or supervised fine-tuning.

Coverage note — None was omitted; all primary methodological formulations (OAMI, CFLT, T1S-Loop), benchmark constructions (DeepChestVQA and DeepTumorVQA preprocessing), quantitative experimental tables, ablation findings, and limitation analyses have been captured as knowls.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Aerts, H. J. W. L., Wee, L., Rios Velazquez, E., Leijenaar, R. T. H., Parmar, C., Grossmann, P., Carvalho, S., Bussink, J., Monshouwer, R., Haibe-Kains, B., Rietveld, D., Hoebers, F., Rietbergen, M. M., Leemans, C. R., Dekker, A., Quackenbush, J., Gillies, R. J., and Lambin, P. Data From NSCLC-Radiomics (version 4), 2014. URL https://doi.org/10.7937/K9/TCIA.2015.PF0M9REI. [Data set].
  3. 3.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736, 2022.
  4. 4.AN, P., XU, Y., and WU, P. Miccai flare23 challenge-attention mechanism-based deep supervision network for abdominal multi-organ segmentation. In MICCAI 2023 FLARE Challenge, 2023.
  5. 5.Antonelli, M., Reinke, A., Bakas, S., Farahani, K., Kopp-Schneider, A., Landman, B. A., Litjens, G., Menze, B., Ronneberger, O., Summers, R. M., et al. The medical segmentation decathlon. Nature communications, 13(1): 4128, 2022.
  6. 6.Austin, J., Müller, N. L., Friedman, P. J., Hansell, D. M., Naidich, D. P., Remy-Jardin, M., Webb, W. R., and Zerhouni, E. A. Glossary of terms for ct of the lungs: recommendations of the nomenclature committee of the fleischner society. Radiology, 200(2):327–331, 1996.
  7. 7.Baharoon, M., Luo, L., Moritz, M., Kumar, A., Kim, S. E., Zhang, X., Zhu, M., Alabbad, M. H., Alhazmi, M. S., Mistry, N. P., et al. Rexgroundingct: A 3d chest ct dataset for segmentation of findings from free-text reports. arXiv preprint arXiv:2507.22030, 2025.
  8. 8.Bai, F., Du, Y., Huang, T., Meng, M. Q.-H., and Zhao, B. M3d: Advancing 3d medical image analysis with multi-modal large language models. arXiv preprint arXiv:2404.00578, 2024.
  9. 9.Blankemeier, L., Cohen, J. P., Kumar, A., Van Veen, D., Gardezi, S. J. S., Paschali, M., Chen, Z., Delbrouck, J.-B., Reis, E., Truyts, C., et al. Merlin: A vision language foundation model for 3d computed tomography. Research Square, pp. rs–3, 2024.
  10. 10.Bruls, R. and Kwee, R. Workload for radiologists during on-call hours: dramatic increase in the past 15 years. Insights into imaging, 11(1):121, 2020.
  11. 11.Chen, J., Gui, C., Ouyang, R., Gao, A., Chen, S., Chen, G. H., Wang, X., Zhang, R., Cai, Z., Ji, K., et al. Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale. arXiv preprint arXiv:2406.19280, 2024a.
  12. 12.Chen, J., Cai, L., Wang, Z., Huang, Y., Jiang, S., Huang, S., Wang, H., and Zhang, Y. Pathagent: Toward interpretable analysis of whole-slide pathology images via large language model-based agentic reasoning. arXiv preprint arXiv:2511.17052, 2025a.
  13. 13.Chen, Y., Xiao, W., Bassi, P. R., Zhou, X., Er, S., Hamamci, I. E., Zhou, Z., and Yuille, A. Are vision language models ready for clinical diagnosis? a 3d medical benchmark for tumor-centric visual question answering. arXiv preprint arXiv:2505.18915, 2025b.
  14. 14.Chen, Z., Varma, M., Delbrouck, J.-B., Paschali, M., Blankemeier, L., Van Veen, D., Valanarasu, J. M. J., Youssef, A., Cohen, J. P., Reis, E. P., et al. Chexagent: Towards a foundation model for chest x-ray interpretation. arXiv preprint arXiv:2401.12208, 2024b.
  15. 15.Fallahpour, A., Ma, J., Munim, A., Lyu, H., and Wang, B. Medrax: Medical reasoning agent for chest x-ray. arXiv preprint arXiv:2502.02673, 2025.
  16. 16.Fu, Y., Zhao, Y., Zeng, Z., Chen, C., and Jin, Y. Unleashing the power of image-tabular self-supervised learning via breaking cross-tabular barriers. arXiv preprint arXiv:2512.14026, 2025.
  17. 17.Gai, X., Liu, J., Li, Y., Meng, Z., Wu, J., and Liu, Z. 3d-rad: A comprehensive 3d radiology med-vqa dataset with multi-temporal analysis and diverse diagnostic tasks. arXiv preprint arXiv:2506.11147, 2025.
  18. 18.Ghezloo, F., Seyfioglu, M. S., Soraki, R., Ikezogwo, W. O., Li, B., Vivekanandan, T., Elmore, J. G., Krishna, R., and Shapiro, L. Pathfinder: A multi-modal multi-agent system for medical diagnostic decision-making applied to histopathology. arXiv preprint arXiv:2502.08916, 2025.
  19. 19.Group, D. T. Ct or invasive coronary angiography in stable chest pain. New England Journal of Medicine, 386(17): 1591–1602, 2022.
  20. 20.Hamamci, I. E., Er, S., Almas, F., Simsek, A. G., Esirgun, S. N., Dogan, I., Dasdelen, M. F., Wittmann, B., Simsar, E., Simsar, M., et al. A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero-shot detection of abnormalities. CoRR, 2024a.
  21. 21.Hamamci, I. E., Er, S., Wang, C., Almas, F., Simsek, A. G., Esirgun, S. N., Dogan, I., Durugol, O. F., Hou, B., Shit, S., et al. Developing generalist foundation models from a multimodal dataset for 3d computed tomography. arXiv preprint arXiv:2403.17834, 2024b.
  22. 22.Hansell, D. M., Bankier, A. A., MacMahon, H., McLoud, T. C., Muller, N. L., and Remy, J. Fleischner society: glossary of terms for thoracic imaging. Radiology, 246 (3):697–722, 2008.
  23. 23.Hasbun, R., Abrahams, J., Jekel, J., and Quagliarello, V. J. Computed tomography of the head before lumbar puncture in adults with suspected meningitis. New England Journal of Medicine, 345(24):1727–1733, 2001.
  24. 24.Hatamizadeh, A., Tang, Y., Nath, V., Yang, D., Myronenko, A., Landman, B., Roth, H. R., and Xu, D. Unetr: Transformers for 3d medical image segmentation. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 574–584, 2022.
  25. 25.Haydel, M. J., Preston, C. A., Mills, T. J., Luber, S., Blaudeau, E., and DeBlieux, P. M. Indications for computed tomography in patients with minor head injury. New England Journal of Medicine, 343(2):100–105, 2000.
  26. 26.He, X., Zhang, Y., Mou, L., Xing, E., and Xie, P. Pathvqa: 30000+ questions for medical visual question answering. arXiv preprint arXiv:2003.10286, 2020.
  27. 27.He, Y., Guo, P., Tang, Y., Myronenko, A., Nath, V., Xu, Z., Yang, D., Zhao, C., Simon, B., Belue, M., et al. Vista3d: A unified segmentation foundation model for 3d medical imaging. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 20863–20873, 2025.
  28. 28.Heller, N., Sathianathen, N., Kalapara, A., Walczak, E., Moore, K., Kaluzniak, H., Rosenberg, J., Blake, P., Rengel, Z., Oestreich, M., et al. The kits19 challenge data: 300 kidney tumor cases with clinical context, ct semantic segmentations, and surgical outcomes. arXiv preprint arXiv:1904.00445, 2019.
  29. 29.Hu, X., Qian, Y., Yu, J., Liu, J., Tang, P., Ji, X., Xu, C., Liu, J., Yan, X., Yu, X., et al. The landscape of medical agents: A survey. Authorea Preprints, 2025a.
  30. 30.Hu, Y., Li, T., Lu, Q., Shao, W., He, J., Qiao, Y., and Luo, P. Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22170–22183, 2024.
  31. 31.Hu, Y., Zheng, Y., Miao, S., Zhang, X., Xia, J., Qi, Y., Zhang, Y., He, Y., Chen, Q., Ye, J., et al. Cardiac-clip: A vision-language foundation model for 3d cardiac ct images. arXiv preprint arXiv:2507.22024, 2025b.
  32. 32.Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J., and Maier-Hein, K. H. nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 18(2):203–211, 2021.
  33. 33.Ji, Y., Bai, H., Ge, C., Yang, J., Zhu, Y., Zhang, R., Li, Z., Zhanng, L., Ma, W., Wan, X., et al. Amos: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation. Advances in neural information processing systems, 35:36722–36732, 2022.
  34. 34.Jiang, S., Liu, F., Wang, Z., Cai, L., and Zhang, Y. Pathreasoner-r1: Instilling structured reasoning into pathology vision-language model via knowledge-guided policy optimization. arXiv preprint arXiv:2601.21617, 2026.
  35. 35.Kim, Y., Park, C., Jeong, H., Chan, Y. S., Xu, X., McDuff, D., Lee, H., Ghassemi, M., Breazeal, C., and Park, H. W. Mdagents: An adaptive collaboration of llms for medical decision-making. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  36. 36.Lau, J. J., Gayen, S., Ben Abacha, A., and Demner-Fushman, D. A dataset of clinically generated visual questions and answers about radiology images. Scientific data, 5(1):1–10, 2018.
  37. 37.Li, B., Yan, T., Pan, Y., Luo, J., Ji, R., Ding, J., Xu, Z., Liu, S., Dong, H., Lin, Z., et al. Mmedagent: Learning to use medical tools with multi-modal agent. In Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 8745–8760, 2024.
  38. 38.Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon, H., and Gao, J. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36:28541–28564, 2023.
  39. 39.Lin, J., Xia, Y., Zhang, J., Yan, K., Cao, K., Lu, L., Luo, J., and Zhang, L. Ct-glip: 3d grounded language-image pretraining with ct scans and radiology reports for full-body scenarios. arXiv preprint arXiv:2404.15272, 2024.
  40. 40.Lin, W., Zhao, Z., Zhang, X., Wu, C., Zhang, Y., Wang, Y., and Xie, W. Pmc-clip: Contrastive language-image pre-training using biomedical documents. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 525–536. Springer, 2023.
  41. 41.Liu, B., Zhan, L.-M., Xu, L., Ma, L., Yang, Y., and Wu, X.-M. Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In 2021 IEEE 18th international symposium on biomedical imaging (ISBI), pp. 1650–1654. IEEE, 2021.
  42. 42.Liu, F., Jiang, S., Cai, L., Wang, Z., and Zhang, Y. Pathflip: Fine-grained language-image pretraining for versatile computational pathology. arXiv preprint arXiv:2512.17621, 2025.
  43. 43.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. Advances in neural information processing systems, 36:34892–34916, 2023.
  44. 44.Low, C. H., Wang, Z., Zhang, T., Zeng, Z., Zhuo, Z., Mazomenos, E. B., and Jin, Y. Surgraw: Multi-agent workflow with chain-of-thought reasoning for surgical intelligence. arXiv preprint arXiv:2503.10265, 2025a.
  45. 45.Low, C. H., Zhuo, Z., Wang, Z., Xu, J., Liu, H., Sirajudeen, N., Boal, M., Edwards, P. J., Stoyanov, D., Francis, N., et al. Cares: Collaborative agentic reasoning for error detection in surgery. arXiv preprint arXiv:2508.08764, 2025b.
  46. 46.Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., and Zhang, Y. An empirical study of catastrophic forgetting in large language models during continual fine-tuning. IEEE Transactions on Audio, Speech and Language Processing, 2025.
  47. 47.Ma, J., Zhang, Y., Gu, S., Zhu, C., Ge, C., Zhang, Y., An, X., Wang, C., Wang, Q., Liu, X., et al. Abdomenct-1k: Is abdominal organ segmentation a solved problem? IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):6695–6714, 2021.
  48. 48.Mei, X., Liu, Z., Singh, A., Lange, M., Boddu, P., Gong, J. Q., Lee, J., DeMarco, C., Cao, C., Platt, S., et al. Interstitial lung disease diagnosis and prognosis using an ai system integrating longitudinal data. Nature communications, 14(1):2272, 2023.
  49. 49.Oren, O., Gersh, B. J., and Bhatt, D. L. Artificial intelligence in medical imaging: switching from radiographic pathological data to clinically meaningful endpoints. The Lancet Digital Health, 2(9):e486–e488, 2020.
  50. 50.Patel, A. G., Pizzitola, V. J., Johnson, C. D., Zhang, N., and Patel, M. D. Radiologists make more errors interpreting off-hours body ct studies during overnight assignments as compared with daytime assignments. Radiology, 297 (2):374–379, 2020.
  51. 51.Rister, B., Yi, D., Shivakumar, K., Nobashi, T., and Rubin, D. L. Ct-org, a new dataset for multiple organ segmentation in computed tomography. Scientific Data, 7(1):381, 2020.
  52. 52.Roth, H., Farag, A., Turkbey, E. B., Lu, L., Liu, J., and Summers, R. M. Data from pancreas-ct. (No Title), 2016.
  53. 53.Rudie, J. D., Lin, H.-M., Ball, R. L., Jalal, S., Prevedello, L. M., Nicolaou, S., Marinelli, B. S., Flanders, A. E., Magudia, K., Shih, G., et al. The rsna abdominal traumatic injury ct (ratic) dataset. Radiology: Artificial Intelligence, 6(6):e240101, 2024.
  54. 54.Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36:68539–68551, 2023.
  55. 55.Sellergren, A., Kazemzadeh, S., Jaroensri, T., Kiraly, A., Traverse, M., Kohlberger, T., Xu, S., Jamil, F., Hughes, C., Lau, C., et al. Medgemma technical report. arXiv preprint arXiv:2507.05201, 2025.
  56. 56.Sharma, D., Purushotham, S., and Reddy, C. K. Medfusenet: An attention-based multimodal deep learning model for visual question answering in the medical domain. Scientific Reports, 11(1):19826, 2021.
  57. 57.Tang, X., Zou, A., Zhang, Z., Li, Z., Zhao, Y., Zhang, X., Cohan, A., and Gerstein, M. Medagents: Large language models as collaborators for zero-shot medical reasoning. In Findings of the Association for Computational Linguistics ACL 2024, pp. 599–621, 2024.
  58. 58.Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  59. 59.Topol, E. J. High-performance medicine: the convergence of human and artificial intelligence. Nature medicine, 25 (1):44–56, 2019.
  60. 60.Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P.-C., Carroll, A., Lau, C., Tanno, R., Ktena, I., et al. Towards generalist biomedical ai. Nejm Ai, 1(3): AIoa2300138, 2024.
  61. 61.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. Transactions on Machine Learning Research.
  62. 62.Wang, X., Peng, Y., Lu, L., Lu, Z., and Summers, R. M. Tienet: Text-image embedding network for common thorax disease classification and reporting in chest x-rays. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 9049–9058, 2018.
  63. 63.Wang, Z., Wu, J., Cai, L., Low, C. H., Yang, X., Li, Q., and Jin, Y. Medagent-pro: Towards evidence-based multimodal medical diagnosis via reasoning agentic workflow. arXiv preprint arXiv:2503.18968, 2025.
  64. 64.Wasserthal, J., Breit, H.-C., Meyer, M. T., Pradella, M., Hinck, D., Sauter, A. W., Heye, T., Boll, D. T., Cyriac, J., Yang, S., et al. Totalsegmentator: robust segmentation of 104 anatomic structures in ct images. Radiology: Artificial Intelligence, 5(5):e230024, 2023.
  65. 65.Wu, C., Zhang, X., Zhang, Y., Hui, H., Wang, Y., and Xie, W. Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data. Nature Communications, 16(1):7866, 2025.
  66. 66.Xia, P., Wang, J., Peng, Y., Zeng, K., Wu, X., Tang, X., Zhu, H., Li, Y., Liu, S., Lu, Y., et al. Mmedagent-rl: Optimizing multi-agent collaboration for multimodal medical reasoning. arXiv preprint arXiv:2506.00555, 2025.
  67. 67.Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.
  68. 68.Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y. React: Synergizing reasoning and acting in language models. In The eleventh international conference on learning representations, 2022.
  69. 69.Zech, J. R., Badgeley, M. A., Liu, M., Costa, A. B., Titano, J. J., and Oermann, E. K. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS medicine, 15(11):e1002683, 2018.
  70. 70.Zhang, S., Xu, Y., Usuyama, N., Xu, H., Bagga, J., Tinn, R., Preston, S., Rao, R., Wei, M., Valluri, N., et al. Biomedclip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs. arXiv preprint arXiv:2303.00915, 2023a.
  71. 71.Zhang, X., Wu, C., Zhao, Z., Lin, W., Zhang, Y., Wang, Y., and Xie, W. Pmc-vqa: Visual instruction tuning for medical visual question answering. arXiv preprint arXiv:2305.10415, 2023b.
  72. 72.Zhang, X., Wu, C., Zhao, Z., Lei, J., Zhang, Y., Wang, Y., and Xie, W. Radgenome-chest ct: A grounded vision-language dataset for chest ct analysis. arXiv preprint arXiv:2404.16754, 2024.
  73. 73.Zhu, Y., Qi, Y., Wang, Z., Gu, L., Sui, D., Hu, H., Zhang, X., He, Z., He, J., Ma, L., et al. Healthflow: A self-evolving ai agent with meta planning for autonomous healthcare research. arXiv preprint arXiv:2508.02621, 2025.
  74. 74.Zuo, Y., Qu, S., Li, Y., Chen, Z., Zhu, X., Hua, E., Zhang, K., Ding, N., and Zhou, B. Medxpertqa: Benchmarking expert-level medical reasoning and understanding. arXiv preprint arXiv:2501.18362, 2025.

Citation

MLA
Wang, Z., et al. “3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis”. arXiv, 2026, http://arxiv.org/abs/2602.18064v2.
APA
Wang, Z., Cai, L., Low, C. H., Liu, H., Wu, J., Wang, J., Wang, R., Song, L., Bian, J., Fu, J., & Jin, Y. (2026). 3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis. arXiv. http://arxiv.org/abs/2602.18064v2
Chicago
Wang, Z., L. Cai, C. H. Low, et al. 2026. “3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis”. arXiv. http://arxiv.org/abs/2602.18064v2.
Harvard
Wang, Z. et al. (2026) “3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2602.18064v2.
Vancouver
1. Wang Z, Cai L, Low CH, et al (2026) 3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis. arXiv

BibTeX

@article{wang20263dmedagent,
  title = {3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis},
  author = {Wang, Ziyue and Cai, Linghan and Low, Chang Han and Liu, Haofeng and Wu, Junde and Wang, Jingyu and Wang, Rui and Song, Lei and Bian, Jiang and Fu, Jingjing and Jin, Yueming},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2602.18064v2},
  eprint = {2602.18064}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/