EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images
Seongsu BaeDaeun KyungJaehee RyuEunbyeol ChoGyubok LeeSunjun KweonJungwoo OhLei JiEric I-Chao ChangTackeun Kim
Presents EHRXQA, a multimodal electronic health record benchmark integrating structured patient tables with chest X-ray images alongside a NeuralSQL-based baseline to enable joint clinical reasoning across text and imaging modalities.
Modern healthcare relies heavily on Electronic Health Records (EHRs) to track patient care, combining structured tabular data such as diagnoses, procedures, and prescriptions with visual diagnostic data like chest X-ray images. While automated question-answering systems offer substantial promise for clinical decision support and medical research, current solutions remain siloed. They operate either exclusively over structured relational databases using standard database queries or exclusively over individual medical images using visual reasoning. This lack of integration prevents healthcare providers from asking complex, cross-modal questions that connect clinical timelines and interventions with radiological findings.
The article introduces and evaluates EHRXQA, the first multi-modal question answering dataset and benchmark bridging structured EHR relational tables with chest X-ray images. The main objective is to establish a standardized benchmark for multi-modal clinical question answering and demonstrate a system capable of joint reasoning across tabular and visual medical records.
To build this benchmark, the researchers integrated three established medical databases: structured clinical tables from MIMIC-IV, radiological images from MIMIC-CXR, and annotated anatomical scene graphs from Chest ImaGenome. They established an intermediate visual question answering benchmark named MIMIC-CXR-VQA comprising 377,391 unique instances across 48 clinically validated templates. Combining these resources yielded the EHRXQA dataset containing 46,152 questions spanning single-modality and multi-modality inquiries across individual and cohort patient scopes. To answer these queries, the authors proposed a framework that couples a large language model with an external visual perception programming interface, translating natural language questions into an executable query language that queries tables and triggers image analysis.
The evaluation revealed several critical findings. First, purely table-based questions achieved high overall execution accuracy of approximately 92.9%, demonstrating that large language models handle complex clinical database queries effectively when supplied with relevant examples. Second, performance dropped significantly when answering questions requiring visual perception: overall execution accuracy fell to 65.9% for combined image-and-table queries and 48.2% for image-only queries. Third, system accuracy declined sharply as the number of required images increased; accuracy dropped to 39.6% when evaluating multiple sequential images for a single patient and plunged to 1.7% when processing broader patient cohorts. Fourth, while the language model achieved high logical parsing accuracy across all categories (72.5% to 87.3%), the visual models failed to match this precision, reaching a peak standalone visual accuracy of only 69.2%.
These findings indicate that logical reasoning and query translation are not the primary bottlenecks preventing real-world deployment of multi-modal clinical intelligence tools. Instead, visual perception accuracy is the main limiting factor. The compounding of minor visual misclassifications across multiple studies or patient groups rapidly undermines final answers, posing safety and reliability risks in critical clinical settings where precise decision-making is vital.
Organizations developing healthcare artificial intelligence should prioritize advancing core medical computer vision models before deploying automated clinical query systems. Future research and development should concentrate on integrating visual uncertainty estimation, expanding datasets to include unanswerable queries to mitigate false conclusions, and developing multi-modal systems capable of dialogue. Decision-makers should treat current multi-modal question-answering systems as experimental research tools rather than autonomous clinical solutions, keeping in mind that the current dataset relies on a single institutional database and a constrained set of visual labels.
- Paper: CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison, Jeremy Irvin et al. (2019). This paper establishes foundational methodologies for extracting chest radiograph labels and handling uncertainty, which underpin the radiological reasoning and visual perception challenges addressed in EHRXQA.
- Paper: ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases, Xiaosong Wang et al. (2017). It introduces hospital-scale chest X-ray disease classification and localization benchmarks that ground the automated thoracic pathology recognition leveraged by multi-modal EHR question answering.
- Paper: Scalable and accurate deep learning with electronic health records, Alvin Rajkomar et al. (2018). It provides the foundational framework for scalable deep learning over structured electronic health records that EHRXQA integrates with visual imaging data.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). It formulates the standard Visual Question Answering task and baseline multimodal architectures that EHRXQA adapts to the clinical domain.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). It introduces a foundational unified transformer baseline for vision-and-language tasks that informs cross-modal reasoning architectures.
- Paper: Large language models encode clinical knowledge, Karan Singhal et al. (2022). It benchmarks large language models on clinical knowledge retrieval and reasoning, establishing the linguistic capabilities required for parsing medical queries in EHRXQA.
- Paper: LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day, Chunyuan Li et al. (2023). This paper develops an end-to-end multimodal biomedical conversational assistant, extending the joint vision-language clinical capabilities benchmarked in EHRXQA.
- Paper: MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding, Yuxin Zuo et al. (2025). It extends multimodal clinical evaluation by benchmarking expert-level medical reasoning across complex patient records and diagnostic imagery.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). It advances multimodal reasoning by introducing step-by-step visual chain-of-thought localization, addressing the visual perception bottlenecks identified in EHRXQA.
