Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling

Xinlu ZhangShiyang LiZhiyu ChenXifeng YanLinda Ruth Petzold

article2023ICML67 citations

Presents a multimodal framework for electronic health records that explicitly captures temporal irregularities across both numerical time series and clinical text through dynamic gating, time-attention mechanisms, and interleaved cross-attention fusion to achieve superior ICU outcome predictions.

Listen

Care provided during the initial hours of an intensive care unit stay is critical to patient survival, yet early clinical decisions are frequently vulnerable to error. While automated deep learning tools can assist clinicians by forecasting outcomes from electronic health records, existing systems struggle with the complex, irregular nature of medical data. Real-world hospital records combine numerical vital signs and laboratory measurements alongside unstructured clinical notes, both collected at irregular, misaligned intervals. The article addresses the challenge of systematically handling this temporal irregularity within individual data sources and across combined multimodal representations.

The main objective of the article is to demonstrate a unified deep learning system that explicitly models irregularity in both time-series data and clinical note sequences, and evaluates whether integrating this temporal information during multimodal fusion improves early clinical predictions. Specifically, the framework introduces a gating mechanism to blend complementary time-series methods, treats encoded text sequences as multivariate irregular time series via a time-attention mechanism, and applies an interleaved fusion architecture to combine data across temporal steps.

To evaluate this approach, the authors conducted experiments using the public MIMIC-III database across two early-stage intensive care unit prediction benchmarks: 48-hour in-hospital mortality prediction and 24-hour phenotype classification across 25 acute conditions. The datasets comprised over 16,000 and 22,000 patient records, respectively. The framework extracted up to the last five clinical notes prior to prediction time using a domain-specific long-context language model, paired them with numerical time series, and benchmarked the architecture against competitive baseline models across unimodal and multimodal setups.

The results demonstrate substantial performance improvements across all evaluation settings. First, in the multimodal setting, the proposed interleaved attention fusion outperformed all baseline fusion strategies, achieving relative F1 score improvements of 4.3% on mortality prediction and setting top marks across precision-recall and area-under-the-curve metrics. Second, dynamically unifying hand-crafted imputation with learned multi-time attention embeddings for time series generated a 6.5% relative F1 gain over the best standalone baseline for phenotype classification. Third, treating clinical notes as irregular time series through the time-attention module improved text-only F1 scores by up to 7.8% relative to baseline text sequence models. Finally, ablation studies showed that capturing irregularity within individual data types directly boosted downstream fusion quality, and scaling text sequence capacity up to 1,024 tokens steadily enhanced predictive accuracy.

These findings indicate that treating timing and irregularity as core features—rather than discarding them through simple averaging or naive concatenation—is essential for accurate medical predictions. Incorporating time-aware multimodal architectures into hospital clinical decision support systems can improve risk stratification for high-stakes conditions, potentially reducing preventable medical errors and improving patient safety without requiring prohibitive computing resources.

Healthcare technology leaders and practitioners should adopt time-aware interpolation and synchronous cross-attention architectures when developing predictive clinical systems from health records. Organizations should also prioritize natural language processing backbones that support longer text sequences to preserve vital context from clinical narratives. Before clinical deployment, future work should evaluate this architecture on additional external hospital databases, test the pipeline in real-time operational workflows, and expand its capacity beyond the recent five-note constraint to encompass complete patient stays.

No sufficiently relevant recommendations were found.

Cover for Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling

Abstract

Health conditions among patients in intensive care units (ICUs) are monitored via electronic health records (EHRs), composed of numerical time series and lengthy clinical note sequences, both taken at irregular time intervals. Dealing with such irregularity in every modality, and integrating irregularity into multimodal representations to improve medical predictions, is a challenging problem. Our method first addresses irregularity in each single modality by (1) modeling irregular time series by dynamically incorporating handcrafted imputation embeddings into learned interpolation embeddings via a gating mechanism, and (2) casting a series of clinical note representations as multivariate irregular time series and tackling irregularity via a time attention mechanism. We further integrate irregularity in multimodal fusion with an interleaved attention mechanism across temporal steps. To the best of our knowledge, this is the first work to thoroughly model irregularity in multimodalities for improving medical predictions. Our proposed methods for two medical prediction tasks consistently outperforms state-of-the-art (SOTA) baselines in each single modality and multimodal fusion scenarios. Specifically, we observe relative improvements of 6.5%, 3.6%, and 4.3% in F1 for time series, clinical notes, and multimodal fusion, respectively. These results demonstrate the effectiveness of our methods and the importance of considering irregularity in multimodal EHRs.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Problem setup
  • 3.2. MISTS
  • 3.2.1. TDE Methods
  • 3.2.2. Unifying TDE Methods
  • 3.3. Irregular clinical notes
  • 3.4. Multimodal fusion
  • 4. Experiments
  • 4.1. Experimental setup
  • 4.2. Main results
  • 4.3. Ablation study
  • 5. Conclusion
  • Acknowledgments
  • References
  • Appendix
  • A. Computation resource of UTDE
  • B. Data preprocessing
  • C. Baselines
  • C.1. MISTS baselines
  • C.2. Irregular clinical notes baselines
  • C.3. Multimodal fusion baselines
  • D. Hyperparameters and training details
  • D.1. MISTS
  • D.2. Irregular clinical notes
  • D.3. Multimodal fusion

Knowls

  1. Knowl 1 — UTDE gates imputation and time-attention representations of irregular measurements

    model/method

    Unified TDE (UTDE) produces a regular-grid representation of multivariate irregularly sampled time series (MISTS) by combining two complementary encodings. For a patient with dmd_m measured variables, first discretize each variable onto an hourly grid of HH time steps. Within a bin, retain the last observation if there are several; fill an empty bin with the most recent earlier observation, or with that variable’s global mean across patients if no earlier observation exists. A causal one-dimensional convolution maps the imputed sequence to eimp∈RH×dhe^{\mathrm{imp}}\in\mathbb{R}^{H\times d_h}, where dhd_h is the embedding width.

    The second encoding, eattn∈RH×dhe^{\mathrm{attn}}\in\mathbb{R}^{H\times d_h}, is produced by discretized multi-time attention (mTAND). For each measured variable, mTAND uses grid times as queries, that variable’s observation times as keys, and its observed values as values; attention therefore interpolates each variable onto the regular grid while using the original irregular timestamps. It uses VV learned Time2Vec time embeddings, whose iith component for time τ\tau is ωiτ+ϕi\omega_i\tau+\phi_i for i=1i=1 and sin⁡(ωiτ+ϕi)\sin(\omega_i\tau+\phi_i) for 1<i≤dv1<i\leq d_v, with learned parameters ωi,ϕi\omega_i,\phi_i and time-embedding width dvd_v. The VV interpolated representations are concatenated and linearly projected to width dhd_h.

    A multilayer perceptron computes a gate from the two encodings, and UTDE combines them as

    zts=g⊙eimp+(1−g)⊙eattn,g=f(eimp⊕eattn).z^{\mathrm{ts}}=g\odot e^{\mathrm{imp}}+(1-g)\odot e^{\mathrm{attn}},\qquad g=f(e^{\mathrm{imp}}\oplus e^{\mathrm{attn}}).

    Here zts∈RH×dhz^{\mathrm{ts}}\in\mathbb{R}^{H\times d_h} is the resulting MISTS representation; ⊕\oplus denotes concatenation, ⊙\odot elementwise multiplication with broadcasting as needed, and ff is the gate MLP. The gate can be patient-level (one scalar), temporal (one value per grid step), or hidden-space-level (one value per time step and embedding dimension). The authors use imputation and mTAND as the two submodules and select the gate granularity on validation data.

  2. Knowl 2 — Irregular clinical notes are represented as time-indexed feature sequences

    model/method

    The note encoder maps each clinical note to its pretrained language model’s [CLS] representation, orders the representations by note-taking time, and treats each embedding dimension as a time series observed at the notes’ irregular timestamps. In the experiments, Clinical-Longformer is the text encoder; its output sequence is truncated to a maximum of 1024 tokens per note. Applying mTAND to this time-indexed sequence uses the regular prediction-grid times as queries, note times as keys, and note representations as values. The result is ztxt∈RH×dhz^{\mathrm{txt}}\in\mathbb{R}^{H\times d_h}, a note representation at each of the HH regular grid steps, with the same width dhd_h as the MISTS representation.

    The time-embedding functions used for MISTS and notes share learned Time2Vec parameters because both modalities’ timestamps lie in the same time space. Their other mTAND components are learned separately to accommodate the different measurement and text representation spaces.

  3. Knowl 3 — Interleaved self- and cross-attention fuses modalities at each temporal step

    model/method

    The fusion module takes the aligned sequences zts,ztxt∈RH×dhz^{\mathrm{ts}},z^{\mathrm{txt}}\in\mathbb{R}^{H\times d_h} and applies JJ identical layers. In each layer, each modality first performs multi-head self-attention over its own temporal sequence. Then, in each direction, multi-head cross-attention uses one modality’s self-attended sequence as queries and the other modality’s sequence as keys and values. Thus the time-series stream receives note information, while the note stream receives time-series information; the next layer again performs self-attention before cross-attention. Each attention and position-wise feed-forward sublayer uses pre-layer normalization and residual connections.

    After the final layer, the model extracts the last hidden state from each modality stream, concatenates the two states, and feeds the result to a fully connected classifier. The experimental model uses J=3J=3 layers. Alternating temporal self-attention and cross-modal exchange is the proposed way to combine temporal context and cross-modal information, rather than learning those components in separate blocks.

  4. Knowl 4 — Evaluation uses early prediction on MIMIC-III ICU records

    experimental setup

    The experiments use MIMIC-III for two tasks: binary 48-hour in-hospital mortality prediction (48-IHM) and 24-hour classification of 25 acute-care phenotypes (24-PHE). Models receive time-series measurements and clinical notes recorded within the relevant prediction window. Patients without clinical notes before the prediction time are excluded. When a patient has more than five eligible notes, only the last five before the prediction time are used, a restriction the authors attribute to computational-resource limits.

    The 48-IHM train, validation, and test sets contain 11,181, 2,473, and 2,488 patients; the corresponding 24-PHE sets contain 15,561, 3,410, and 3,379. Evaluation metrics are F1 and AUPR for mortality, and macro-F1 and AUROC for phenotype classification. Each reported result is the mean and standard deviation over three runs on fixed task splits. The mortality data have an approximately 1:7 death-to-discharge ratio. Time-series inputs are rescaled to [0,1], and the clinical-note encoder is Clinical-Longformer with a maximum sequence length of 1024 tokens.

  5. Knowl 5 — UTDE improves over MISTS baselines on both prediction tasks

    empirical result

    On MIMIC-III, UTDE is compared with imputation, IP-Net, mTAND, GRU-D, SeFT, RAINDROP, DGM²-O, and MTGNN. The table reports mean ± standard deviation across three runs; scores are on the reported 0–100 scale. UTDE achieves the best result for every listed metric. Relative to the strongest non-UTDE baseline, its reported gains are 4.4% in AUPR for 48-IHM and 6.5% in macro-F1 for 24-PHE.

    Task Metric Imputation IP-Net mTAND GRU-D SeFT RAINDROP DGM2^2-O MTGNN UTDE
    48-IHM F1 39.73±\pm1.39 37.22±\pm2.75 43.87±\pm0.54 42.82±\pm0.57 16.46±\pm8.61 39.46±\pm3.70 39.08±\pm1.53 38.60±\pm2.50 45.26±\pm0.70
    48-IHM AUPR 44.36±\pm1.36 39.36±\pm1.10 47.54±\pm1.28 45.90±\pm0.40 23.89±\pm0.46 36.23±\pm0.37 37.79±\pm1.54 36.49±\pm2.10 49.64±\pm1.00
    24-PHE Macro-F1 23.36±\pm0.45 17.90±\pm0.66 19.90±\pm0.38 18.96±\pm0.99 6.10±\pm0.15 21.81±\pm1.71 18.40±\pm0.18 14.48±\pm1.69 24.89±\pm0.43
    24-PHE AUROC 74.93±\pm0.22 73.45±\pm0.10 73.48±\pm0.11 73.33±\pm0.10 65.66±\pm0.11 73.95±\pm0.89 71.71±\pm0.16 70.56±\pm0.68 75.56±\pm0.17
  6. Knowl 6 — Time-attention modeling improves clinical-note predictions

    empirical result

    For the clinical-note modality, mTANDtxt is compared with Flat note averaging, HierTrans, T-LSTM, FT-LSTM, and GRU-D. Results are mean ± standard deviation over three runs on the reported 0–100 scale. mTANDtxt has the highest scores across all four task metrics. Its F1 is 52.57 on 48-IHM and 52.95 on 24-PHE; compared with HierTrans, these are relative improvements of 7.8% and 5.3%, respectively. On 24-PHE, mTANDtxt also exceeds the strongest other irregularity-modeling baseline, GRU-D, by 3.6% in F1. The results distinguish temporal sequence modeling from simply averaging notes and support explicitly modeling note-taking irregularity.

    Method 48-IHM F1 48-IHM AUPR 24-PHE Macro-F1 24-PHE AUROC
    Flat 39.78±\pm1.14 51.69±\pm0.79 18.14±\pm1.36 74.81±\pm0.22
    HierTrans 48.76±\pm2.44 52.98±\pm1.69 50.25±\pm1.21 84.90±\pm0.25
    T-LSTM 50.32±\pm0.89 52.57±\pm3.25 39.13±\pm1.35 82.03±\pm0.07
    FT-LSTM 48.51±\pm1.67 54.39±\pm1.38 38.24±\pm0.61 81.07±\pm0.27
    GRU-D 51.01±\pm1.50 54.34±\pm0.75 51.09±\pm1.02 84.19±\pm0.20
    mTANDtxt 52.57±\pm1.30 56.05±\pm1.09 52.95±\pm0.06 85.43±\pm0.07
  7. Knowl 7 — Interleaved attention gives the strongest multimodal fusion results

    empirical result

    The fusion comparison uses UTDE for MISTS and mTANDtxt for notes, then compares concatenation, Tensor Fusion (TF), MAG, MulT, and the proposed interleaved attention. Results are mean ± standard deviation over three runs on the reported 0–100 scale. Interleaved attention ranks first on all four metrics. On 48-IHM it reaches F1 56.45 and AUPR 60.23; its F1 is a 4.3% relative improvement over MulT (54.13). On 24-PHE it reaches macro-F1 54.84 and AUROC 86.06. Most fusion approaches beat at least one single modality; the synchronous attention approaches perform better overall than the asynchronous late-fusion baselines in these experiments.

    Method 48-IHM F1 48-IHM AUPR 24-PHE Macro-F1 24-PHE AUROC
    Time series only 45.26±\pm0.70 49.64±\pm1.00 24.89±\pm0.43 75.56±\pm0.17
    Notes only 52.57±\pm1.30 56.05±\pm1.09 52.95±\pm0.06 85.43±\pm0.07
    Concat 52.77±\pm0.70 57.13±\pm0.7 53.30±\pm0.35 85.94±\pm0.21
    TF 51.44±\pm0.66 57.07±\pm0.82 49.84±\pm0.83 84.74±\pm0.16
    MAG 53.20±\pm2.13 57.86±\pm1.07 53.73±\pm0.37 85.94±\pm0.07
    MulT 54.13±\pm1.20 58.94±\pm1.94 54.20±\pm0.33 85.96±\pm0.07
    Interleaved attention 56.45±\pm1.30 60.23±\pm1.54 54.84±\pm0.31 86.06±\pm0.06
  8. Knowl 8 — UTDE gains persist across submodule choices and time-series backbones

    empirical result

    An ablation replaces mTAND in UTDE with IP-Net, retaining imputation as the other submodule. The gated combination of imputation and IP-Net outperforms either submodule alone on both tasks, although it is weaker than the imputation–mTAND UTDE variant. In addition, comparisons using CNN, LSTM, and Transformer time-series encoders show that the relative performance of imputation and mTAND varies with the backbone, while UTDE outperforms its individual submodules across the tested backbones. This supports the paper’s claim that the gate can combine complementary temporal encodings rather than relying on one TDE method or one particular sequence encoder.

    Task Metric Imputation IP-Net UTDE (Imputation + IP-Net) UTDE (Imputation + mTAND)
    48-IHM F1 39.73±\pm1.39 37.22±\pm2.75 44.88±\pm1.96 45.26±\pm0.70
    48-IHM AUPR 44.36±\pm1.36 39.36±\pm1.10 45.49±\pm3.45 49.64±\pm1.00
    24-PHE Macro-F1 23.36±\pm0.45 17.90±\pm0.66 24.06±\pm0.51 24.89±\pm0.43
    24-PHE AUROC 74.93±\pm0.22 73.45±\pm0.10 75.17±\pm0.07 75.56±\pm0.17
  9. Knowl 9 — Ablations show that both time-series and note irregularity matter in fusion

    empirical result

    The multimodal ablation compares the full model with variants that replace UTDE by imputation alone or mTAND alone, and with a variant that omits mTANDtxt. Both single-TDE replacements score below the full model on all four reported metrics. Omitting mTANDtxt and directly fusing the note representations also reduces performance, particularly 48-IHM F1, which falls from 56.45 to 51.14. The results indicate that handling irregularity in each modality contributes to multimodal prediction, rather than only improving its corresponding unimodal encoder. Values are mean ± standard deviation over three runs on the reported 0–100 scale.

    Fusion variant 48-IHM F1 48-IHM AUPR 24-PHE Macro-F1 24-PHE AUROC
    Full model 56.45±\pm1.30 60.23±\pm1.54 54.84±\pm0.31 86.06±\pm0.06
    Without UTDE; imputation used 54.59±\pm0.91 56.80±\pm0.54 54.46±\pm0.17 85.98±\pm0.02
    Without UTDE; mTAND used 54.89±\pm1.09 59.11±\pm1.21 54.07±\pm0.51 85.92±\pm0.12
    Without mTANDtxt 51.14±\pm1.79 57.81±\pm0.76 53.33±\pm0.62 85.60±\pm0.06
  10. Knowl 10 — Longer clinical-note inputs improve multimodal performance

    empirical result

    The authors vary the maximum token sequence length used to encode each note, comparing Bio-ClinicalBERT at lengths 128, 256, and 512 with Clinical-Longformer at length 1024. In the multimodal model, the plotted F1 and secondary metrics improve as the maximum sequence length increases for both 48-IHM and 24-PHE. This result supports retaining longer note context in multimodal prediction. The 1024-token Clinical-Longformer setting covers more than 98% of notes in the two tasks. Separately, because of computational limits, the preprocessing retains at most the five notes closest to the prediction time for patients with more than five eligible notes; the paper presents the assumption that later notes may be more influential as a hypothesis, not as a separately tested finding.

Coverage note — The exact score breakdown for every time-series backbone and the full training and hyperparameter grids are omitted as supporting experimental detail; the cross-backbone UTDE finding is retained.

References

  1. 1.Adler-Milstein, J., DesRoches, C. M., Kralovec, P., Foster, G., Worzala, C., Charles, D., Searcy, T., and Jha, A. K. Electronic health record adoption in us hospitals: progress continues, but challenges persist. Health affairs, 34(12):2174–2180, 2015.
  2. 2.Afessa, B., Gajic, O., and Keegan, M. T. Severity of illness and organ failure assessment in adult intensive care units. Critical care clinics, 23(3):639–658, 2007.
  3. 3.Alberti, C., Brun-Buisson, C., Burchardi, H., Martin, C., Goodman, S., Artigas, A., Sicignano, A., Palazzo, M., Moreno, R., Boulme, R., et al. Epidemiology of sepsis and infection in icu patients from an international multicentre cohort study. Intensive care medicine, 28(2):108–121, 2002.
  4. 4.Alsentzer, E., Murphy, J. R., Boag, W., Weng, W.-H., Jin, D., Naumann, T., and McDermott, M. Publicly available clinical bert embeddings. arXiv preprint arXiv:1904.03323, 2019.
  5. 5.Arbabi, A., Adams, D. R., Fidler, S., Brudno, M., et al. Identifying clinical terms in medical text using ontology-guided machine learning. JMIR medical informatics, 7(2):e12596, 2019.
  6. 6.Baytas, I. M., Xiao, C., Zhang, X., Wang, F., Jain, A. K., and Zhou, J. Patient subtyping via time-aware lstm networks. In Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining, pp. 65–74, 2017.
  7. 7.Che, Z., Purushotham, S., Cho, K., Sontag, D., and Liu, Y. Recurrent neural networks for multivariate time series with missing values. Scientific reports, 8(1):1–12, 2018.
  8. 8.Chen, R. T., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K. Neural ordinary differential equations. Advances in neural information processing systems, 31, 2018.
  9. 9.Choi, E., Bahadori, M. T., Schuetz, A., Stewart, W. F., and Sun, J. Doctor ai: Predicting clinical events via recurrent neural networks. In Machine learning for healthcare conference, pp. 301–318. PMLR, 2016.
  10. 10.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  11. 11.Deznabi, I., Iyyer, M., and Fiterau, M. Predicting in-hospital mortality by combining clinical notes with time-series data. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 4026–4031, 2021.
  12. 12.Golmaei, S. N. and Luo, X. Deepnote-gnn: predicting hospital readmission using clinical notes and patient network. In Proceedings of the 12th ACM Conference on Bioinformatics, Computational Biology, and Health Informatics, pp. 1–9, 2021.
  13. 13.Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., and Poon, H. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), 3(1):1–23, 2021.
  14. 14.Gupta, P., Malhotra, P., Vig, L., and Shroff, G. Transfer learning for clinical time series analysis using recurrent neural networks. arXiv preprint arXiv:1807.01705, 2018.
  15. 15.Harutyunyan, H., Khachatrian, H., Kale, D. C., Ver Steeg, G., and Galstyan, A. Multitask learning and benchmarking with clinical time series data. Scientific Data, 6(1):96, 2019. ISSN 2052-4463. doi: 10.1038/s41597-019-0103-9. URL https://doi.org/10.1038/s41597-019-0103-9.
  16. 16.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  17. 17.Horn, M., Moor, M., Bock, C., Rieck, B., and Borgwardt, K. Set functions for time series. In III, H. D. and Singh, A. (eds.), Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pp. 4353–4363. PMLR, 13–18 Jul 2020.
  18. 18.Huang, K., Altosaar, J., and Ranganath, R. Clinicalbert: Modeling clinical notes and predicting hospital readmission. arXiv preprint arXiv:1904.05342, 2019.
  19. 19.Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E. Adaptive mixtures of local experts. Neural computation, 3(1):79–87, 1991.
  20. 20.Johnson, A. E., Pollard, T. J., Shen, L., Lehman, L.-w. H., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Anthony Celi, L., and Mark, R. G. Mimic-iii, a freely accessible critical care database. Scientific data, 3(1):1–9, 2016.
  21. 21.Kazemi, S. M., Goel, R., Eghbali, S., Ramanan, J., Sahota, J., Thakur, S., Wu, S., Smyth, C., Poupart, P., and Brubaker, M. Time2vec: Learning a vector representation of time. arXiv preprint arXiv:1907.05321, 2019.
  22. 22.Khadanga, S., Aggarwal, K., Joty, S., and Srivastava, J. Using clinical notes with time series data for icu management. arXiv preprint arXiv:1909.09702, 2019.
  23. 23.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  24. 24.LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  25. 25.Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alche-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/6775a0635c302542da2c32aa19d86be0-Paper.pdf.
  26. 26.Li, Y., Wehbe, R. M., Ahmad, F. S., Wang, H., and Luo, Y. Clinical-longformer and clinical-bigbird: Transformers for long clinical sequences. arXiv preprint arXiv:2201.11838, 2022.
  27. 27.Lim, B. and Zohren, S. Time-series forecasting with deep learning: a survey. Philosophical Transactions of the Royal Society A, 379(2194):20200209, 2021.
  28. 28.Lin, K., Hu, Y., and Kong, G. Predicting in-hospital mortality of patients with acute kidney injury in the icu using random forest model. International journal of medical informatics, 125:55–61, 2019.
  29. 29.Lipton, Z. C., Kale, D., and Wetzel, R. Directly modeling missing data in sequences with rnns: Improved classification of clinical time series. In Machine learning for healthcare conference, pp. 253–270. PMLR, 2016.
  30. 30.Liu, Z., Shen, Y., Lakshminarasimhan, V. B., Liang, P. P., Zadeh, A., and Morency, L.-P. Efficient low-rank multimodal fusion with modality-specific factors. arXiv preprint arXiv:1806.00064, 2018.
  31. 31.Liu, Z., Zhang, J., Hou, Y., Zhang, X., Li, G., and Xiang, Y. Machine learning for multimodal electronic health records-based research: Challenges and perspectives. arXiv preprint arXiv:2111.04898, 2021.
  32. 32.Mahbub, M., Srinivasan, S., Danciu, I., Peluso, A., Begoli, E., Tamang, S., and Peterson, G. D. Unstructured clinical notes within the 24 hours since admission predict short, mid & long-term mortality in adult icu patients. Plos one, 17(1):e0262182, 2022.
  33. 33.McDermott, M., Nestor, B., Kim, E., Zhang, W., Goldenberg, A., Szolovits, P., and Ghassemi, M. A comprehensive ehr timeseries pre-training benchmark. In Proceedings of the Conference on Health, Inference, and Learning, pp. 257–278, 2021.
  34. 34.Otero-Lopez, M. J., Alonso-Hernández, P., Maderuelo-Fernández, J. A., Garrido-Corro, B., Domínguez-Gil, A., and Sánchez-Rodríguez, A. Preventable adverse drug events in hospitalized patients. Medicina clinica, 126(3):81–87, 2006.
  35. 35.Pappagari, R., Zelasko, P., Villalba, J., Carmiel, Y., and Dehak, N. Hierarchical transformers for long document classification. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pp. 838–844. IEEE, 2019.
  36. 36.Rahman, W., Hasan, M. K., Lee, S., Bagher Zadeh, A., Mao, C., Morency, L.-P., and Hoque, E. Integrating multimodal information in large pretrained transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 2359–2369, Online, July 2020a. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.214. URL https://www.aclweb.org/anthology/2020.acl-main.214.
  37. 37.Rahman, W., Hasan, M. K., Lee, S., Zadeh, A., Mao, C., Morency, L.-P., and Hoque, E. Integrating multimodal information in large pretrained transformers. In Proceedings of the conference. Association for Computational Linguistics. Meeting, volume 2020, pp. 2359. NIH Public Access, 2020b.
  38. 38.Rubanova, Y., Chen, R. T., and Duvenaud, D. K. Latent ordinary differential equations for irregularly-sampled time series. Advances in neural information processing systems, 32, 2019.
  39. 39.Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.
  40. 40.Shickel, B., Tighe, P. J., Bihorac, A., and Rashidi, P. Deep ehr: a survey of recent advances in deep learning techniques for electronic health record (ehr) analysis. IEEE journal of biomedical and health informatics, 22(5):1589–1604, 2017.
  41. 41.Shukla, S. N. and Marlin, B. Interpolation-prediction networks for irregularly sampled time series. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=r1efr3C9Ym.
  42. 42.Shukla, S. N. and Marlin, B. M. Multi-time attention networks for irregularly sampled time series. arXiv preprint arXiv:2101.10318, 2021.
  43. 43.Tisherman, S. A. and Stein, D. M. Icu management of trauma patients. Critical Care Medicine, 46(12):1991–1997, 2018.
  44. 44.Tsai, Y.-H. H., Bai, S., Liang, P. P., Kolter, J. Z., Morency, L.-P., and Salakhutdinov, R. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the conference. Association for Computational Linguistics. Meeting, volume 2019, pp. 6558. NIH Public Access, 2019.
  45. 45.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  46. 46.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, October 2020. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/2020.emnlp-demos.6.
  47. 47.Wu, Y., Ni, J., Cheng, W., Zong, B., Song, D., Chen, Z., Liu, Y., Zhang, X., Chen, H., and Davidson, S. B. Dynamic gaussian mixture based deep generative model for robust forecasting on sparse multivariate time series. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 651–659, 2021.
  48. 48.Wu, Z., Pan, S., Long, G., Jiang, J., Chang, X., and Zhang, C. Connecting the dots: Multivariate time series forecasting with graph neural networks. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, pp. 753–763, 2020.
  49. 49.Xiao, C., Choi, E., and Sun, J. Opportunities and challenges in developing deep learning models using electronic health records data: a systematic review. Journal of the American Medical Informatics Association, 25(10):1419–1428, 2018.
  50. 50.Xu, Z., So, D. R., and Dai, A. M. Mufasa: Multimodal fusion architecture search for electronic health records. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 10532–10540, 2021.
  51. 51.Yang, B. and Wu, L. How to leverage multimodal ehr data for better medical predictions? arXiv preprint arXiv:2110.15763, 2021.
  52. 52.Yang, H., Kuang, L., and Xia, F. Multimodal temporal-clinical note network for mortality prediction. Journal of Biomedical Semantics, 12(1):1–14, 2021.
  53. 53.Zadeh, A., Chen, M., Poria, S., Cambria, E., and Morency, L.-P. Tensor fusion network for multimodal sentiment analysis. arXiv preprint arXiv:1707.07250, 2017.
  54. 54.Zerveas, G., Jayaraman, S., Patel, D., Bhamidipaty, A., and Eickhoff, C. A transformer-based framework for multivariate time series representation learning. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp. 2114–2124, 2021.
  55. 55.Zhang, D., Thadajarassiri, J., Sen, C., and Rundensteiner, E. Time-aware transformer-based network for clinical notes series prediction. In Machine Learning for Healthcare Conference, pp. 566–588. PMLR, 2020.
  56. 56.Zhang, X., Li, S., Cheng, Z., Callcut, R., and Petzold, L. Domain adaptation for trauma mortality prediction in ehrs with feature disparity. In 2021 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pp. 1145–1152, 2021a. doi: 10.1109/BIBM52615.2021.9669798.
  57. 57.Zhang, X., Zeman, M., Tsiligkaridis, T., and Zitnik, M. Graph-guided network for irregularly sampled multivariate time series. arXiv preprint arXiv:2110.05357, 2021b.

Citation

MLA
Zhang, X., et al. “Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling”. International Conference on Machine Learning, vol. 202, 2023, pp. 41300–13, https://proceedings.mlr.press/v202/zhang23v.html.
APA
Zhang, X., Li, S., Chen, Z., Yan, X., & Petzold, L. R. (2023). Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling. International Conference on Machine Learning, 202, 41300–41313. https://proceedings.mlr.press/v202/zhang23v.html
Chicago
Zhang, X., S. Li, Z. Chen, X. Yan, and L. R. Petzold. 2023. “Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling”. International Conference on Machine Learning 202: 41300–41313. https://proceedings.mlr.press/v202/zhang23v.html.
Harvard
Zhang, X. et al. (2023) “Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling”, International Conference on Machine Learning. PMLR, pp. 41300–41313. Available at: https://proceedings.mlr.press/v202/zhang23v.html.
Vancouver
1. Zhang X, Li S, Chen Z, Yan X, Petzold LR (2023) Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling. In: International Conference on Machine Learning. PMLR, pp 41300–41313

BibTeX

@InProceedings{pmlr-v202-zhang23v,
  title = 	 {Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling},
  author =       {Zhang, Xinlu and Li, Shiyang and Chen, Zhiyu and Yan, Xifeng and Petzold, Linda Ruth},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {41300--41313},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/zhang23v/zhang23v.pdf},
  url = 	 {https://proceedings.mlr.press/v202/zhang23v.html},
  abstract = 	 {Health conditions among patients in intensive care units (ICUs) are monitored via electronic health records (EHRs), composed of numerical time series and lengthy clinical note sequences, both taken at $\textit{irregular}$ time intervals. Dealing with such irregularity in every modality, and integrating irregularity into multimodal representations to improve medical predictions, is a challenging problem. Our method first addresses irregularity in each single modality by (1) modeling irregular time series by dynamically incorporating hand-crafted imputation embeddings into learned interpolation embeddings via a gating mechanism, and (2) casting a series of clinical note representations as multivariate irregular time series and tackling irregularity via a time attention mechanism. We further integrate irregularity in multimodal fusion with an interleaved attention mechanism across temporal steps. To the best of our knowledge, this is the first work to thoroughly model irregularity in multimodalities for improving medical predictions. Our proposed methods for two medical prediction tasks consistently outperforms state-of-the-art (SOTA) baselines in each single modality and multimodal fusion scenarios. Specifically, we observe relative improvements of 6.5%, 3.6%, and 4.3% in F1 for time series, clinical notes, and multimodal fusion, respectively. These results demonstrate the effectiveness of our methods and the importance of considering irregularity in multimodal EHRs.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/