EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images

Seongsu BaeDaeun KyungJaehee RyuEunbyeol ChoGyubok LeeSunjun KweonJungwoo OhLei JiEric I-Chao ChangTackeun Kim

article2023NeurIPS51 citations

Presents EHRXQA, a multimodal electronic health record benchmark integrating structured patient tables with chest X-ray images alongside a NeuralSQL-based baseline to enable joint clinical reasoning across text and imaging modalities.

Listen

Modern healthcare relies heavily on Electronic Health Records (EHRs) to track patient care, combining structured tabular data such as diagnoses, procedures, and prescriptions with visual diagnostic data like chest X-ray images. While automated question-answering systems offer substantial promise for clinical decision support and medical research, current solutions remain siloed. They operate either exclusively over structured relational databases using standard database queries or exclusively over individual medical images using visual reasoning. This lack of integration prevents healthcare providers from asking complex, cross-modal questions that connect clinical timelines and interventions with radiological findings.

The article introduces and evaluates EHRXQA, the first multi-modal question answering dataset and benchmark bridging structured EHR relational tables with chest X-ray images. The main objective is to establish a standardized benchmark for multi-modal clinical question answering and demonstrate a system capable of joint reasoning across tabular and visual medical records.

To build this benchmark, the researchers integrated three established medical databases: structured clinical tables from MIMIC-IV, radiological images from MIMIC-CXR, and annotated anatomical scene graphs from Chest ImaGenome. They established an intermediate visual question answering benchmark named MIMIC-CXR-VQA comprising 377,391 unique instances across 48 clinically validated templates. Combining these resources yielded the EHRXQA dataset containing 46,152 questions spanning single-modality and multi-modality inquiries across individual and cohort patient scopes. To answer these queries, the authors proposed a framework that couples a large language model with an external visual perception programming interface, translating natural language questions into an executable query language that queries tables and triggers image analysis.

The evaluation revealed several critical findings. First, purely table-based questions achieved high overall execution accuracy of approximately 92.9%, demonstrating that large language models handle complex clinical database queries effectively when supplied with relevant examples. Second, performance dropped significantly when answering questions requiring visual perception: overall execution accuracy fell to 65.9% for combined image-and-table queries and 48.2% for image-only queries. Third, system accuracy declined sharply as the number of required images increased; accuracy dropped to 39.6% when evaluating multiple sequential images for a single patient and plunged to 1.7% when processing broader patient cohorts. Fourth, while the language model achieved high logical parsing accuracy across all categories (72.5% to 87.3%), the visual models failed to match this precision, reaching a peak standalone visual accuracy of only 69.2%.

These findings indicate that logical reasoning and query translation are not the primary bottlenecks preventing real-world deployment of multi-modal clinical intelligence tools. Instead, visual perception accuracy is the main limiting factor. The compounding of minor visual misclassifications across multiple studies or patient groups rapidly undermines final answers, posing safety and reliability risks in critical clinical settings where precise decision-making is vital.

Organizations developing healthcare artificial intelligence should prioritize advancing core medical computer vision models before deploying automated clinical query systems. Future research and development should concentrate on integrating visual uncertainty estimation, expanding datasets to include unanswerable queries to mitigate false conclusions, and developing multi-modal systems capable of dialogue. Decision-makers should treat current multi-modal question-answering systems as experimental research tools rather than autonomous clinical solutions, keeping in mind that the current dataset relies on a single institutional database and a constrained set of visual labels.

Cover for EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images

Abstract

Electronic Health Records (EHRs), which contain patients’ medical histories in various multi-modal formats, often overlook the potential for joint reasoning across imaging and table modalities underexplored in current EHR Question Answering (QA) systems. In this paper, we introduce EHRXQA, a novel multi-modal question answering dataset combining structured EHRs and chest X-ray images. To develop our dataset, we first construct two uni-modal resources: 1) The MIMIC-CXR-VQA dataset, our newly created medical visual question answering (VQA) benchmark, specifically designed to augment the imaging modality in EHR QA, and 2) EHRSQL (MIMIC-IV), a refashioned version of a previously established table-based EHR QA dataset. By integrating these two uni-modal resources, we successfully construct a multi-modal EHR QA dataset that necessitates both uni-modal and cross-modal reasoning. To address the unique challenges of multi-modal questions within EHRs, we propose a NeuralSQL-based strategy equipped with an external VQA API. This pioneering endeavor enhances engagement with multi-modal EHR sources and we believe that our dataset can catalyze advances in real-world medical scenarios such as clinical decision-making and research. EHRXQA is available at https://github.com/baeseongsu/ehrxqa.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminary: Ingredients for Multi-Modal EHR QA
  • 3.1 Uni-Modal Data Resources
  • 3.2 Uni-Modal EHR QA datasets
  • 3.2.1 Table-based EHR QA: EHRSQL
  • 3.2.2 Image-based EHR QA: MIMIC-CXR-VQA
  • 4 EHRXQA: A Multi-Modal EHR Question Answering Dataset
  • 4.1 Dataset Construction
  • 4.1.1 Database Construction
  • 4.1.2 Question Template Construction
  • 4.1.3 SQL/NeuralSQL Annotation
  • 4.1.4 Data Generation
  • 4.2 Data Statistics and Comparisons with other EHR QA datasets
  • 5 NeuralSQL with Visual Question Answering
  • 6 Experiments
  • 6.1 MIMIC-CXR-VQA
  • 6.2 EHRXQA
  • 7 Discussion
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — EHRXQA Dataset

    definition

    EHRXQA is a multi-modal electronic health record (EHR) question answering dataset that integrates structured relational tables from MIMIC-IV with chest X-ray (CXR) images from MIMIC-CXR and scene-graph annotations from Chest ImaGenome. It contains a total of 46,152 question-query-answer samples split across training (36,174), validation (5,170), and test (4,808) sets.

    The dataset spans three modality-based query scopes:

    1. Table-related QA (16,529 samples: 12,961 train / 1,852 valid / 1,716 test): Questions answered exclusively through structured EHR tables using SQL queries.
    2. Image-related QA (16,366 samples: 12,860 train / 1,838 valid / 1,668 test): Questions requiring visual reasoning over CXR images linked to patient records using NeuralSQL queries.
    3. Image+Table-related QA (13,257 samples: 10,353 train / 1,480 valid / 1,424 test): Questions requiring joint cross-modal reasoning over structured clinical tables and CXR images using NeuralSQL queries.

    The questions also span three patient-level scopes:

    • None: General questions not specific to individual patients.
    • Single: Questions about a specific patient, querying single images (1-image), comparing consecutive studies (2-image), or aggregating across multiple studies (NN-image).
    • Group: Cohort-level questions aggregating records across multiple patients.
  2. Knowl 2 — MIMIC-CXR-VQA Benchmark Dataset

    definition

    MIMIC-CXR-VQA is a medical visual question answering (VQA) benchmark constructed by pairing frontal-view (AP/PA) chest radiographs from MIMIC-CXR with bounding-box and relationship annotations from the Chest ImaGenome silver and gold scene graph datasets. Each CXR image is characterized across 563 distinct relations between 36 anatomical objects and 68 attributes grouped into 5 categories ('anatomical finding', 'disease', 'device', 'tubes/lines', 'technical assessment').

    The dataset comprises 377,391 unique (Image I,Question Q,Answer A)(\text{Image } I, \text{Question } Q, \text{Answer } A) triples across seven content categories: presence, anatomy, attribute, abnormality, size, plane, and gender. Question generation uses 48 seed templates expanded with logical/set operations and paraphrased using GPT-4 (averaging 16.5 paraphrases per template). The VQA task is formulated as multi-label classification over 110 target labels (36 objects, 68 attributes, 'M', 'F', 'AP', 'PA', 'yes', and 'no'). The training and validation sets are constructed from 800 patients in the machine-generated silver dataset (95:5 split), while the test set is constructed from 400 patients in the clinician-validated gold dataset.

  3. Knowl 3 — EHR-Image Database Integration and Timeframe Normalization

    model/method

    To link unstructured chest radiographs with relational EHR databases, an image reference table named TB_CXR is added to the MIMIC-IV relational schema. The TB_CXR table consists of six schema columns:

    • subject_id: Unique identifier for the patient.
    • hadm_id: Unique identifier for the hospital admission.
    • study_id: Unique identifier for the radiological study.
    • image_id: Unique identifier for the specific frontal chest X-ray image.
    • studydatetime: Timestamp when the imaging study occurred.
    • viewposition: Radiographic projection (e.g., AP or PA).

    To establish consistent temporal references (e.g., 'last year', 'this month', 'since 2102') across patients whose timestamps were artificially shifted between 2100 and 2200 for de-identification:

    1. Patient medical event timelines are compressed into a standardized 2100–2105 timeframe while preserving relative intervals and study order.
    2. The current evaluation timestamp is fixed to 2105-12-31 23:59:00. Any medical record or image timestamp occurring after this cutoff is excluded.
    3. Patients with an admission timestamp prior to the cutoff but lacking a discharge timestamp due to the cutoff are designated as currently admitted.
  4. Knowl 4 — NeuralSQL Framework with External VQA API

    model/method

    NeuralSQL is an executable querying framework that extends standard SQL to process multi-modal EHR databases by delegating image perception subqueries to an external Visual Question Answering (VQA) model.

    The framework operates in two distinct stages:

    1. NeuralSQL Parsing (Q→ZQ \to Z): A Large Language Model (LLM) translates a natural language question QQ into an executable NeuralSQL query ZZ. For image-dependent sub-tasks, the query embeds a specialized API call: FUNC_VQA(qI,cI)\text{FUNC\_VQA}(q_I, c_I) where qIq_I is a natural language subquestion targeting visual findings (e.g., "does a chest x-ray study indicate lung cancer in the left lung?") and cIc_I is the database column containing image study identifiers (referencing TB_CXR.study_id).

    2. NeuralSQL Execution (Z→AZ \to A): The NeuralSQL query ZZ is parsed into an abstract syntax tree (AST). When evaluating a node containing FUNC_VQA(qI,cI)\text{FUNC\_VQA}(q_I, c_I), the interpreter retrieves the corresponding chest radiograph(s) II via an internal image loader using identifier cIc_I. The image(s) II and visual subquestion qIq_I are passed to a fine-tuned vision-language model (such as M3AE\text{M}^3\text{AE}), which executes perception inference. The boolean or categorical output of the VQA model is returned as a structured relational column object, enabling the SQL engine to continue relational filtering, aggregation, and sorting to output the final answer AA.

  5. Knowl 5 — Four-Stage Templated Data Generation Pipeline for EHRXQA

    algorithm

    The generation of question-query-answer tuples in EHRXQA follows a structured five-phase pipeline (Stage 0 to Stage 4) followed by neural execution filtering and linguistic paraphrasing:

    Input: Base question template schema TT, Database DD, Preprocessed visual attributes VV
    Output: Set of valid (Question QQ, Query ZZ, Answer AA) pairs
    Stage 0: Select a base question-query template from TT (categorized by modality and patient scope).
    Stage 1 (Visual Value Sampling):
        Fill visual placeholders object,{object}, {attribute}, and categoryusingvalidclinicaltuplesfrom{category} using valid clinical tuples from V$.
    Stage 2 (Operation Value Sampling):
        Sample operation/grouping values from predefined domain constraints (e.g., [age_group] in {20s, 30s, 40s, 50s, 60 or above}).
    Stage 3 (Time Template Sampling):
        Instantiate temporal operators with natural language templates (e.g., [time_filter_global], [time_filter_within], [time_filter_exact]).
    Stage 4 (Condition Value Sampling):
        Sample database-grounded condition placeholders (e.g., {gender}, {year}, {drug_name}, {procedure_name}).
    Construct synchronized (Question QrawQ_{\text{raw}}, Query ZZ) pair with matching sampled arguments.
    Execute ZZ on database DD to obtain Answer AA.
    if AA is non-empty and valid then
        Generate paraphrased question variations QQ using GPT-4.
        Manually review paraphrased candidates for medical and grammatical fidelity.
        Add (QQ, ZZ, AA) to dataset pool.
    end if
  6. Knowl 6 — Relative AUROC Metric for Medical VQA Perception Upper-Bound Calibration

    equation

    Because human radiologists exhibit diagnostic discrepancy and imperfect perceptual accuracy on chest radiographs, the achievable upper bound on image perception tasks is below 100%. To calibrate VQA model performance against an empirical perceptual upper bound on basic single-image verification queries (e.g., "Is there attributein{attribute} in {object}?"), a relative metric mrelm_{\text{rel}} is defined:

    mrel(o,a)=mVQA(o,a)mref(o,a)m_{\text{rel}}(o, a) = \frac{m_{\text{VQA}}(o, a)}{m_{\text{ref}}(o, a)}

    where:

    • oo denotes a specific anatomical object (from 36 candidate anatomical locations).
    • aa denotes a specific attribute or finding (from 68 candidate attributes).
    • mVQA(o,a)m_{\text{VQA}}(o, a) is the metric score achieved by the evaluated VQA model on the verification task for pair (o,a)(o, a).
    • mref(o,a)m_{\text{ref}}(o, a) is the reference metric score achieved by a dedicated multi-label visual classification reference model trained directly on the gold/silver scene graph labels for pair (o,a)(o, a).

    When Area Under the Receiver Operating Characteristic is used as the base metric mm, the overall relative metric is denoted as AUROCrel\text{AUROC}_{\text{rel}}.

  7. Knowl 7 — Semantic Parsing and Execution Evaluation Metrics for EHRXQA

    definition

    To isolate the performance of the semantic parsing stage from the neural visual execution stage in multi-modal EHR question answering, EHRXQA uses three distinct evaluation metrics:

    1. Logical Form Accuracy (AccLF\text{Acc}_{\text{LF}}): Evaluates the semantic parsing stage (Q→ZQ \to Z). It measures the exact match percentage between the predicted executable query Z^\hat{Z} (SQL or NeuralSQL) and the ground-truth query ZZ: AccLF=I(Z^≡Z)\text{Acc}_{\text{LF}} = \mathbb{I}(\hat{Z} \equiv Z)

    2. Ground-truth Execution Accuracy (AccEX∣gt\text{Acc}_{\text{EX}|\text{gt}}): Evaluates the execution stage (Z→AZ \to A) using the ground-truth query ZZ. It compares the result obtained by executing ZZ with the true answer AA: AccEX∣gt=I(Execute(Z)=A)\text{Acc}_{\text{EX}|\text{gt}} = \mathbb{I}(\text{Execute}(Z) = A) For Table-related QA, AccEX∣gt=100%\text{Acc}_{\text{EX}|\text{gt}} = 100\%, whereas for Image-related and Image+Table-related QA, this metric measures the isolated accuracy of the external VQA perception module.

    3. Prediction Execution Accuracy (AccEX∣pred\text{Acc}_{\text{EX}|\text{pred}}): Evaluates the end-to-end question answering pipeline by executing the predicted query Z^\hat{Z} and comparing the result with ground-truth answer AA: AccEX∣pred=I(Execute(Z^)=A)\text{Acc}_{\text{EX}|\text{pred}} = \mathbb{I}(\text{Execute}(\hat{Z}) = A)

  8. Knowl 8 — Performance of Vision-Language Pre-trained Baselines on MIMIC-CXR-VQA

    data/table

    Evaluation of baseline VQA models on the MIMIC-CXR-VQA benchmark demonstrates that vision-language pre-trained (VLP) models outperform language-only priors, with M3AE\text{M}^3\text{AE} achieving the highest accuracy and perceptual alignment. Models marked with an asterisk (∗*) were pre-trained on a standardized MIMIC-CXR corpus of chest radiographs and free-text reports.

    Model Validation Set Test Set
    Acc (%) F1 (micro) Acc (%) F1 (micro) AUCrel\textbf{AUC}_{\textbf{rel}}
    Prior (Most) 26.8 0.27 25.4 0.25 -
    Prior (Question) 34.3 0.34 32.4 0.32 -
    PubMedCLIP 55.1±1.755.1 \pm 1.7 0.56±0.020.56 \pm 0.02 54.9±1.354.9 \pm 1.3 0.54±0.020.54 \pm 0.02 0.82±0.090.82 \pm 0.09
    PubMedCLIP∗^* 56.6±1.956.6 \pm 1.9 0.58±0.020.58 \pm 0.02 56.5±2.156.5 \pm 2.1 0.56±0.020.56 \pm 0.02 0.83±0.090.83 \pm 0.09
    MedViLL∗^* 64.7±0.264.7 \pm 0.2 0.69±0.000.69 \pm 0.00 63.6±0.163.6 \pm 0.1 0.67±0.000.67 \pm 0.00 0.98±0.080.98 \pm 0.08
    M3AE\text{M}^3\text{AE} 68.9±0.268.9 \pm 0.2 0.73±0.000.73 \pm 0.00 68.9±0.368.9 \pm 0.3 0.72±0.000.72 \pm 0.00 1.02±0.081.02 \pm 0.08
    M3AE∗\text{M}^3\text{AE}^* 70.2±0.1\mathbf{70.2 \pm 0.1} 0.74±0.00\mathbf{0.74 \pm 0.00} 69.2±0.4\mathbf{69.2 \pm 0.4} 0.73±0.00\mathbf{0.73 \pm 0.00} 1.05±0.09\mathbf{1.05 \pm 0.09}

    The language-only baseline Prior (Question) achieves only 32.4% test accuracy, confirming low language shortcut bias in the dataset. M3AE∗\text{M}^3\text{AE}^* reaches 69.2%69.2\% test accuracy and an AUCrel\text{AUC}_{\text{rel}} of 1.05±0.091.05 \pm 0.09, indicating visual perception on par with dedicated task-specific classifiers.

  9. Knowl 9 — EHRXQA Benchmark Performance with ChatGPT and M3AE

    data/table

    Benchmark results on the EHRXQA test set using ChatGPT (gpt-3.5-turbo-0613) as the semantic parser and a fine-tuned M3AE\text{M}^3\text{AE} vision-language model as the external VQA API demonstrate the impact of few-shot in-context retrieval (Fixed 10-shot vs BM25 10-shot retrieval from the training set).

    Model Prompt Image-related Table-related Image+Table-related
    AccLF\text{Acc}_{\text{LF}} AccEX∣gt\text{Acc}_{\text{EX}|\text{gt}} AccEX∣pred\text{Acc}_{\text{EX}|\text{pred}} AccLF\text{Acc}_{\text{LF}} AccEX∣gt\text{Acc}_{\text{EX}|\text{gt}} AccEX∣pred\text{Acc}_{\text{EX}|\text{pred}} AccLF\text{Acc}_{\text{LF}} AccEX∣gt\text{Acc}_{\text{EX}|\text{gt}} AccEX∣pred\text{Acc}_{\text{EX}|\text{pred}}
    ChatGPT Fixed 1.1 49.4 17.4 4.9 100.0 30.0 4.8 68.8 35.7
    + M3AE\text{M}^3\text{AE} BM25 87.3 49.4 48.2 73.0 100.0 92.9 72.5 68.8 65.9

    Retrieving domain-specific few-shot examples via BM25 dramatically improves logical form parsing accuracy across all scopes: from 1.1% to 87.3% on Image-related QA, 4.9% to 73.0% on Table-related QA, and 4.8% to 72.5% on Image+Table-related QA.

  10. Knowl 10 — Visual Perception Bottleneck in Multi-Image and Group Queries

    empirical result

    In multi-modal EHR question answering with NeuralSQL, the primary performance bottleneck stems from perceptual error propagation across multiple images rather than logical parsing errors. When evaluating ChatGPT (gpt-3.5-turbo-0613) with BM25 in-context learning paired with an M3AE\text{M}^3\text{AE} VQA API on the EHRXQA benchmark:

    1. Discrepancy Between Parsing and Execution Accuracy: On Image-related QA, logical form parsing accuracy is high (AccLF=87.3%\text{Acc}_{\text{LF}} = 87.3\%), but predicted execution accuracy drops significantly (AccEX∣pred=48.2%\text{Acc}_{\text{EX}|\text{pred}} = 48.2\%) due to the ground-truth execution bound (AccEX∣gt=49.4%\text{Acc}_{\text{EX}|\text{gt}} = 49.4\%). In contrast, Table-related QA achieves AccEX∣pred=92.9%\text{Acc}_{\text{EX}|\text{pred}} = 92.9\% with AccLF=73.0%\text{Acc}_{\text{LF}} = 73.0\%.
    2. Impact of Multi-Image Aggregation: For Image-related queries within a single patient scope, AccEX∣pred\text{Acc}_{\text{EX}|\text{pred}} drops to 39.6%39.6\% when querying across multiple images (NN-image scope).
    3. Severe Degradation on Group-Level Queries: In cohort/group queries requiring visual inference across many patients' chest radiographs, cumulative classification errors cause AccEX∣pred\text{Acc}_{\text{EX}|\text{pred}} to plummet to 1.7%1.7\%.
    4. Modality Filtering Effect: Image+Table queries achieve higher predicted execution accuracy (65.9%65.9\%) than pure Image-related queries (48.2%48.2\%) because relational table constraints filter and restrict the candidate image set to a smaller size, mitigating cascading perception errors.

Coverage note — None was omitted. All primary dataset contributions (EHRXQA, MIMIC-CXR-VQA), the NeuralSQL method, database schemas, evaluation metrics, and empirical baseline results are covered.

References

  1. 1.Asma Ben Abacha, Sadid A Hasan, Vivek V Datla, Joey Liu, Dina Demner-Fushman, and Henning Müller. Vqa-med: Overview of the medical visual question answering task at imageclef 2019. CLEF (working notes), 2(6), 2019.
  2. 2.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433, 2015.
  3. 3.Seongsu Bae, Daeyoung Kim, Jiho Kim, and Edward Choi. Question answering for complex electronic health records database using unified encoder-decoder architecture. In Machine Learning for Health, pages 13–25. PMLR, 2021.
  4. 4.Jayetri Bardhan, Anthony Colas, Kirk Roberts, and Daisy Zhe Wang. Drugehrqa: A question answering dataset on structured and unstructured electronic health records for medicine related queries. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 1083–1097, 2022.
  5. 5.Asma Ben Abacha, Mourad Sarrouti, Dina Demner-Fushman, Sadid A Hasan, and Henning Müller. Overview of the vqa-med task at imageclef 2021: Visual question answering and generation in the medical domain. In Proceedings of the CLEF 2021 Conference and Labs of the Evaluation Forum-working notes. 21-24 September 2021, 2021.
  6. 6.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. Semantic parsing on freebase from question-answer pairs. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1533–1544, 2013.
  7. 7.Adrian Brady, Risteárd Ó Laoide, Peter McCarthy, and Ronan McDermott. Discrepancy and error in radiology: concepts, causes and consequences. The Ulster medical journal, 81(1):3, 2012.
  8. 8.Adrian P Brady. Error and discrepancy in radiology: inevitable or avoidable? Insights into imaging, 8:171–182, 2017.
  9. 9.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  10. 10.Yingshan Chang, Mridu Narang, Hisami Suzuki, Guihong Cao, Jianfeng Gao, and Yonatan Bisk. Webqa: Multihop and multimodal qa. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16495–16504, 2022.
  11. 11.Mark Chen et al. Evaluating large language models trained on code. In arXiv, 2021.
  12. 12.Wenhu Chen, Hexiang Hu, Xi Chen, Pat Verga, and William Cohen. Murag: Multimodal retrieval-augmented generator for open question answering over images and text. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5558–5570, 2022.
  13. 13.Zhihong Chen, Yuhao Du, Jinpeng Hu, Yang Liu, Guanbin Li, Xiang Wan, and Tsung-Hui Chang. Multi-modal masked autoencoders for medical vision-and-language pre-training. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2022: 25th International Conference, Singapore, September 18–22, 2022, Proceedings, Part V, pages 679–689. Springer, 2022.
  14. 14.Zui Chen, Zihui Gu, Lei Cao, Ju Fan, Sam Madden, and Nan Tang. Symphony: Towards natural language query answering over multi-modal data lakes. In Conference on Innovative Data Systems Research, CIDR, pages 8–151, 2023.
  15. 15.Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, et al. Binding language models in symbolic languages. arXiv preprint arXiv:2210.02875, 2022.
  16. 16.Philipp Christmann, Rishiraj Saha Roy, and Gerhard Weikum. Conversational question answering on heterogeneous sources. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 144–154, 2022.
  17. 17.Nicholas J Dobbins, Bin Han, Weipeng Zhou, Kristine Lan, H Nina Kim, Robert Harrington, Ozlem Uzuner, and Meliha Yetisgen. Leafai: query generator for clinical cohort discovery rivaling a human programmer. arXiv preprint arXiv:2304.06203, 2023.
  18. 18.Sedigheh Eslami, Christoph Meinel, and Gerard De Melo. Pubmedclip: How much does clip benefit visual question answering in the medical domain? In Findings of the Association for Computational Linguistics: EACL 2023, pages 1151–1163, 2023.
  19. 19.Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang. Vqa-lol: Visual question answering under the lens of logic. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI 16, pages 379–396. Springer, 2020.
  20. 20.Darryl Hannan, Akshay Jain, and Mohit Bansal. Manymodalqa: Modality disambiguation and qa over diverse inputs. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7879–7886, 2020.
  21. 21.Sadid A Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Henning Müller, and Matthew P Lungren. Overview of imageclef 2018 medical domain visual question answering task. In CLEF (Working Notes), 2018.
  22. 22.Xuehai He. Towards visual question answering on pathology images. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, volume 2, 2021.
  23. 23.Xinyue Hu, Lin Gu, Kazuma Kobayashi, Qiyuan An, Qingyu Chen, Zhiyong Lu, Chang Su, Tatsuya Harada, and Yingying Zhu. Interpretable medical image visual question answering via multi-modal relationship graph learning. arXiv preprint arXiv:2302.09636, 2023.
  24. 24.Jian Huang, Yihao Chen, Yong Li, Zhenguo Yang, Xuehao Gong, Fu Lee Wang, Xiaohong Xu, and Wenyin Liu. Medical knowledge-based network for patient-oriented visual question answering. Information Processing & Management, 2023.
  25. 25.Yefan Huang, Xiaoli Wang, Feiyan Liu, and Guofeng Huang. Ovqa: A clinically generated visual question answering dataset. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2924–2938, 2022.
  26. 26.Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6700–6709, 2019.
  27. 27.Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Benjamin Moody, Brian Gow, Li-wei H Lehman, et al. Mimic-iv, a freely accessible electronic health record dataset. Scientific data, 10(1):1, 2023.
  28. 28.Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng. Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data, 6(1):317, 2019.
  29. 29.Alistair EW Johnson, Tom J Pollard, Lu Shen, H Lehman Li-Wei, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. Mimic-iii, a freely accessible critical care database. Scientific data, 3(1):1–9, 2016.
  30. 30.Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2901–2910, 2017.
  31. 31.Olga Kovaleva, Chaitanya Shivade, Satyananda Kashyap, Karina Kanjaria, Joy Wu, Deddeh Ballah, Adam Coy, Alexandros Karargyris, Yufan Guo, David Beymer Beymer, et al. Towards visual dialog for radiology. In Proceedings of the 19th SIGBioMed Workshop on Biomedical Language Processing, pages 60–69, 2020.
  32. 32.Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images. Scientific data, 5(1):1–10, 2018.
  33. 33.Gyubok Lee, Hyeonji Hwang, Seongsu Bae, Yeonsu Kwon, Woncheol Shin, Seongjun Yang, Minjoon Seo, Jong-Yeup Kim, and Edward Choi. Ehrsql: A practical text-to-sql benchmark for electronic health records. Advances in Neural Information Processing Systems, 35:15589–15601, 2022.
  34. 34.Eric Lehman, Vladislav Lialin, Katelyn Y Legaspi, Anne Janelle R Sy, Patricia Therese S Pile, Nicole Rose I Alberto, Richard Raymund R Ragasa, Corinna Victoria M Puyat, Isabelle Rose I Alberto, Pia Gabrielle I Alfonso, et al. Learning to ask like a physician. In Workshop on Clinical Natural Language Processing, 2022.
  35. 35.Jinyang Li, Binyuan Hui, Ge Qu, Binhua Li, Jiaxi Yang, Bowen Li, Bailin Wang, Bowen Qin, Rongyu Cao, Ruiying Geng, et al. Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls. arXiv preprint arXiv:2305.03111, 2023.
  36. 36.Yongqi Li, Wenjie Li, and Liqiang Nie. Mmcoqa: Conversational question answering over text, tables, and images. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4220–4231, 2022.
  37. 37.Zhihong Lin, Donghao Zhang, Qingyi Tac, Danli Shi, Gholamreza Haffari, Qi Wu, Mingguang He, and Zongyuan Ge. Medical visual question answering: A survey. arXiv preprint arXiv:2111.10056, 2021.
  38. 38.Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pages 1650–1654. IEEE, 2021.
  39. 39.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521, 2022.
  40. 40.Jong Hak Moon, Hyungyung Lee, Woncheol Shin, Young-Hak Kim, and Edward Choi. Multi-modal understanding and generation for medical images and text via vision-language pre-training. IEEE Journal of Biomedical and Health Informatics, 26(12):6070–6080, 2022.
  41. 41.Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz. Capabilities of gpt-4 on medical challenge problems. 2023.
  42. 42.OpenAI. Introducing chatgpt, 2022.
  43. 43.OpenAI. Gpt-4 technical report. In arXiv, 2023.
  44. 44.Anusri Pampari, Preethi Raghavan, Jennifer Liang, and Jian Peng. emrqa: A large corpus for question answering on electronic medical records. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2357–2368, 2018.
  45. 45.Tom J Pollard, Alistair EW Johnson, Jesse D Raffa, Leo A Celi, Roger G Mark, and Omar Badawi. The eicu collaborative research database, a freely available multi-center database for critical care research. Scientific data, 5(1):1–13, 2018.
  46. 46.Preethi Raghavan, Jennifer J Liang, Diwakar Mahajan, Rachita Chandra, and Peter Szolovits. emrkbqa: A clinical knowledge-base question answering dataset. In Proceedings of the 20th Workshop on Biomedical Language Processing, pages 64–73, 2021.
  47. 47.Stephen Robertson, Hugo Zaragoza, et al. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389, 2009.
  48. 48.Hrituraj Singh, Anshul Nasery, Denil Mehta, Aishwarya Agarwal, Jatin Lamba, and Balaji Vasan Srinivasan. Mimoqa: Multimodal input multimodal output question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5317–5332, 2021.
  49. 49.Sarvesh Soni, Surabhi Datta, and Kirk Roberts. quehry: a question answering system to query electronic health records. Journal of the American Medical Informatics Association, 30(6):1091–1102, 2023.
  50. 50.Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant. Multimodalqa: Complex question answering over text, tables and images. arXiv preprint arXiv:2104.06039, 2021.
  51. 51.Richard Tarbell, Kim-Kwang Raymond Choo, Glenn Dietrich, and Anthony Rios. Towards understanding the generalization of medical text-to-sql models and datasets. arXiv preprint arXiv:2303.12898, 2023.
  52. 52.Matthias Urban and Carsten Binnig. Towards multi-modal dbmss for seamless querying of texts and tables. arXiv preprint arXiv:2304.13559, 2023.
  53. 53.Ping Wang, Tian Shi, and Chandan K Reddy. Text-to-sql generation for question answering on electronic medical records. In Proceedings of The Web Conference 2020, pages 350–361, 2020.
  54. 54.Joy T Wu, Nkechinyere Nneka Agu, Ismini Lourentzou, Arjun Sharma, Joseph Alexander Paguio, Jasper Seth Yao, Edward Christopher Dee, William G Mitchell, Satyananda Kashyap, Andrea Giovannini, et al. Chest imagenome dataset for clinical reasoning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  55. 55.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3911–3921, 2018.
  56. 56.Yilun Zhao, Yunxiang Li, Chenying Li, and Rui Zhang. Multihiertt: Numerical reasoning over multi hierarchical tabular and textual data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6588–6600, 2022.
  57. 57.Victor Zhong, Caiming Xiong, and Richard Socher. Seq2sql: Generating structured queries from natural language using reinforcement learning. arXiv preprint arXiv:1709.00103, 2017.

Citation

MLA
Bae, S., et al. “EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images”. arXiv, 2023, http://arxiv.org/abs/2310.18652v2.
APA
Bae, S., Kyung, D., Ryu, J., Cho, E., Lee, G., Kweon, S., Oh, J., Ji, L., Chang, E. I.-C., Kim, T., & Choi, E. (2023). EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images. arXiv. http://arxiv.org/abs/2310.18652v2
Chicago
Bae, S., D. Kyung, J. Ryu, et al. 2023. “EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images”. arXiv. http://arxiv.org/abs/2310.18652v2.
Harvard
Bae, S. et al. (2023) “EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.18652v2.
Vancouver
1. Bae S, Kyung D, Ryu J, et al (2023) EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images. arXiv

BibTeX

@article{bae2023ehrxqa,
  title = {EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images},
  author = {Bae, Seongsu and Kyung, Daeun and Ryu, Jaehee and Cho, Eunbyeol and Lee, Gyubok and Kweon, Sunjun and Oh, Jungwoo and Ji, Lei and Chang, Eric I-Chao and Kim, Tackeun and Choi, Edward},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.18652v2},
  eprint = {2310.18652}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission