Scalable and accurate deep learning with electronic health records

Alvin RajkomarEyal OrenKai ChenAndrew M. DaiNissan HajajMichaela HardtPeter J. LiuXiaobing LiuJake MarcusMimi Sun

article2018npj Digital Medicine2,898 citations

Demonstrates that deep learning models trained on raw electronic health records structured with the FHIR standard outperform traditional clinical risk scores across multiple hospital systems for predicting mortality, readmission, length of stay, and discharge diagnoses without manual feature engineering.

Listen

A deep learning system was developed to predict key clinical events directly from the full, uncurated electronic health records of hospitalized adults, rather than from a small set of hand-selected variables. The work addressed the practical barriers that have long limited predictive modeling in health care: the high cost of creating task-specific datasets, the loss of most available information when records are reduced to a few dozen variables, and the difficulty of deploying models across different hospitals without extensive manual harmonization.

Researchers converted raw records from two large academic medical centers into a uniform Fast Healthcare Interoperability Resources format and trained neural-network models on more than 46 billion discrete data points drawn from 216,221 admissions. The same modeling approach was applied without site-specific tuning to four clinically distinct tasks: inpatient mortality, 30-day unplanned readmission, length of stay of seven days or longer, and the full set of discharge diagnoses. Performance was compared with established clinical scores and logistic-regression baselines that used far fewer inputs.

The deep-learning models achieved AUROCs of 0.930.94 for mortality, 0.750.76 for readmission, 0.850.86 for prolonged stay, and 0.90 for diagnoseseach materially higher than the corresponding traditional models. The improvement translated into roughly half as many false alerts for mortality risk and allowed accurate predictions 2448 hours earlier than baseline methods. Attribution techniques showed that the networks identified clinically meaningful elements such as malignant effusions, antibiotic orders, and nursing assessments within individual charts.

These results indicate that accurate, scalable forecasts for safety, quality, and resource-use outcomes can be generated from existing hospital data without laborious feature engineering. Because the approach works on the entire record, including free-text notes, it reduces reliance on incomplete or noisy structured fields and may lower alert fatigue for clinicians.

Further prospective trials are required to determine whether the improved predictions actually change care and outcomes. Additional work is also needed to test transfer of models to new sites, to quantify the incremental value of narrative notes, and to confirm that interpretability methods remain reliable at scale. The current evidence rests on retrospective data from two U.S. academic centers; performance in community hospitals or outside the United States remains untested.

arXiv: 1801.07860
  • Paper: Representation Learning: A Review and New Perspectives, Yoshua Bengio et al. (2012). Reading this review of representation learning provides the necessary foundational concepts on unsupervised feature learning and deep architectures before exploring their application to raw electronic health records.
  • Paper: Long Short-Term Memory, Sepp Hochreiter et al. (1997). Understanding the mechanics of Long Short-Term Memory networks in this foundational paper prepares the reader for how sequential clinical events are modeled over time in raw patient records.
  • Paper: Model Cards for Model Reporting, Margaret Mitchell et al. (2019). Following the source study on scalable deep learning for electronic health records, this paper introduces model cards to rigorously document performance disaggregation and safety limitations for deployed healthcare models.
Cover for Scalable and accurate deep learning with electronic health records

Abstract

Predictive modeling with electronic health record (EHR) data is anticipated to drive personalized medicine and improve healthcare quality. Constructing predictive statistical models typically requires extraction of curated predictor variables from normalized EHR data, a labor-intensive process that discards the vast majority of information in each patient’s record. We propose a representation of patientsentire raw EHR records based on the Fast Healthcare Interoperability Resources (FHIR) format. We demonstrate that deep learning methods using this representation are capable of accurately predicting multiple medical events from multiple centers without site-specific data harmonization. We validated our approach using de-identified EHR data from two US academic medical centers with 216,221 adult patients hospitalized for at least 24 h. In the sequential format we propose, this volume of EHR data unrolled into a total of 46,864,534,945 data points, including clinical notes. Deep learning models achieved high accuracy for tasks such as predicting: in-hospital mortality (area under the receiver operator curve [AUROC] across sites 0.930.94), 30-day unplanned readmission (AUROC 0.750.76), prolonged length of stay (AUROC 0.850.86), and all of a patient’s final discharge diagnoses (frequency-weighted AUROC 0.90). These models outperformed traditional, clinically-used predictive models in all cases. We believe that this approach can be used to create accurate and scalable predictions for a variety of clinical scenarios. In a case study of a particular prediction, we demonstrate that neural networks can be used to identify relevant information from the patient’s chart.

Table of Contents

  • INTRODUCTION
  • Related work
  • RESULTS
  • Mortality
  • Readmissions
  • Long length of stay
  • Inferring discharge diagnoses
  • Case study of model interpretation
  • DISCUSSION
  • METHODS
  • Datasets
  • Data representation and processing
  • Outcomes
  • Prediction timing
  • Study cohort
  • Algorithm development and analysis
  • Comparison to previously published algorithms
  • Explanation of predictions
  • Model evaluation and statistical analysis
  • Data availability
  • Code availability
  • ACKNOWLEDGEMENTS
  • AUTHOR CONTRIBUTIONS
  • ADDITIONAL INFORMATION
  • REFERENCES

Knowls

  1. Knowl 1 — FHIR-Based Raw EHR Sequence Representation

    model/method

    To represent electronic health record (EHR) data without site-specific manual variable curation or data harmonization, all clinical records for a patient are structured as a chronological sequence of Fast Healthcare Interoperability Resources (FHIR) data objects.

    Raw EHR data—encompassing patient demographics, clinician orders, diagnoses, procedures, medications, laboratory values, vital signs, nursing flowsheets, and unstructured free-text clinical notes—are mapped directly into standard FHIR resource containers. Within each resource, data attributes are decomposed into discrete categorical tokens (e.g., individual medication ingredients, generic and brand names, order types, and individual words from clinical text) and normalized numerical values.

    The input to predictive neural networks is the full temporal sequence of tokens ordered by timestamp from the start of a patient's historical medical record up to the designated prediction time-point, allowing direct feature learning across tens of thousands of raw variables.

  2. Knowl 2 — Ensemble Deep Learning Architecture for Longitudinal EHR Modeling

    model/method

    To process longitudinal sequences of variable-length EHR tokens, an ensemble of three distinct model architectures is deployed for each clinical prediction task:

    1. Long Short-Term Memory (LSTM) Recurrent Neural Networks: Processes sequential token streams to capture long-range temporal dependencies and order of clinical events.
    2. Time-Aware Neural Networks (TANN): Employs attention mechanisms that explicitly encode continuous time intervals between clinical events, allowing the network to weight relevant historical data points.
    3. Boosted Time-Based Decision Stumps: Uses ensemble boosting over temporal decision stumps to model sparse and non-linear interactions across time-indexed variables.

    Each model architecture is trained independently on the full sequence of patient tokens available up to a given prediction time point, and their predicted probabilities are combined via ensembling to yield final predictions.

  3. Knowl 3 — Clinical Prediction Performance Across Tasks and Time Points

    data/table

    The deep learning models outperform traditional clinical baseline scoring algorithms across four distinct predictive tasks evaluated at various time points during hospital admission on test cohorts from Hospital A (UCSF) and Hospital B (UCM):

    Task and Prediction Time Hospital A AUROC (95% CI) Hospital B AUROC (95% CI)
    Inpatient mortality
    24 h before admission 0.87 (0.85–0.89) 0.81 (0.79–0.83)
    At admission 0.90 (0.88–0.92) 0.90 (0.86–0.91)
    24 h after admission 0.95 (0.94–0.96) 0.93 (0.92–0.94)
    Baseline (aEWS) at 24 h after admission 0.85 (0.81–0.89) 0.86 (0.83–0.88)
    30-day readmission
    At admission 0.73 (0.71–0.74) 0.72 (0.71–0.73)
    At 24 h after admission 0.74 (0.72–0.75) 0.73 (0.72–0.74)
    At discharge 0.77 (0.75–0.78) 0.76 (0.75–0.77)
    Baseline (mHOSPITAL) at discharge 0.70 (0.68–0.72) 0.68 (0.67–0.69)
    Length of stay \ge 7 days
    At admission 0.81 (0.80–0.82) 0.80 (0.80–0.81)
    At 24 h after admission 0.86 (0.86–0.87) 0.85 (0.85–0.86)
    Baseline (mLiu) at 24 h after admission 0.76 (0.75–0.77) 0.74 (0.73–0.75)
    Discharge diagnoses (weighted AUROC)
    At admission 0.87 0.86
    At 24 h after admission 0.89 0.88
    At discharge 0.90 0.90

    The deep learning ensemble achieved higher discrimination (AUROC) than standard clinical baselines at every evaluated time point for in-hospital mortality (compared to the 28-variable augmented Early Warning Score, aEWS), 30-day unplanned readmissions (compared to modified HOSPITAL score), and long length of stay 7\ge 7 days (compared to the modified Liu logistic regression model).

  4. Knowl 4 — Inpatient Mortality Discrimination and Alert Reduction

    empirical result

    For predicting inpatient mortality 24 hours after hospital admission, the deep learning model achieved an AUROC of 0.95 (95% CI 0.94–0.96) for Hospital A and 0.93 (95% CI 0.92–0.94) for Hospital B, compared to 0.85 (95% CI 0.81–0.89) and 0.86 (95% CI 0.83–0.88) for the augmented Early Warning Score (aEWS) baseline.

    At an alert threshold corresponding to 80% sensitivity, the deep learning model roughly halved the work-up-to-detection ratio (number of patients needed to evaluate per true positive case) compared to aEWS:

    • Hospital A: 7.4 needed to evaluate for the deep learning model vs. 14.3 for aEWS.
    • Hospital B: 8.0 needed to evaluate for the deep learning model vs. 15.4 for aEWS.

    Furthermore, the deep learning model attained a level of predictive discrimination comparable to baseline 24-hour scores at 24 to 48 hours earlier in the patient's stay (achieving an AUROC of 0.87–0.81 at 24 hours prior to inpatient admission).

  5. Knowl 5 — 30-Day Unplanned Readmission and Prolonged Length of Stay Accuracy

    empirical result

    When evaluated on unplanned readmissions within 30 days of discharge, the deep learning model at discharge achieved an AUROC of 0.77 (95% CI 0.75–0.78) for Hospital A and 0.76 (95% CI 0.75–0.77) for Hospital B, significantly outperforming the baseline modified HOSPITAL score of 0.70 (95% CI 0.68–0.72) and 0.68 (95% CI 0.67–0.69).

    For predicting prolonged hospitalization (length of stay 7\ge 7 days, representing the ~75th percentile of stays), the deep learning model at 24 hours after admission achieved an AUROC of 0.86 (95% CI 0.86–0.87) for Hospital A and 0.85 (95% CI 0.84–0.86) for Hospital B, compared to 0.76 (95% CI 0.75–0.77) and 0.74 (95% CI 0.73–0.75) for the modified Liu baseline model.

  6. Knowl 6 — Multi-Label Discharge Diagnosis Code Assignment

    empirical result

    The deep learning system predicted the entire set of primary and secondary billing diagnoses simultaneously from a vocabulary of 14,025 ICD-9 codes without requiring manual feature selection. Correctness required full-length ICD-9 code agreement (e.g., code 250.4 was evaluated as distinct from 250.42).

    • Macro-weighted AUROC: Reached 0.86–0.87 at hospital admission, 0.88–0.89 at 24 hours post-admission, and 0.90 for both Hospital A and Hospital B at discharge.
    • Micro-weighted F1 Score: Reached 0.41 for Hospital A and 0.40 for Hospital B at discharge using a single global threshold selected on the validation set.
  7. Knowl 7 — Token-Level Attribution and Clinical Explainability

    model/method

    To explain individual predictions and mitigate the 'black box' opacity of deep neural networks, token-level attribution mechanisms (derived from attention weights in Time-Aware Neural Networks) highlight the specific raw EHR tokens that contributed most strongly to a prediction.

    When evaluated on specific patient cases, the attribution mechanism isolated clinically relevant multi-modal concepts without manual rules, including:

    1. Mentions of critical findings in free-text clinical notes and radiology reports (e.g., 'malignant pleural effusions', 'empyema', trade names for chest drainage catheters like 'pleurx').
    2. Nursing flowsheet scores indicating high functional risk (e.g., elevated Braden scale scores for pressure ulcer risk).
    3. Timelines of specific intravenous antibiotic and medication administrations.
  8. Knowl 8 — Experimental Cohort and Multi-Center Evaluation Setup

    experimental setup

    The evaluation cohort comprised 216,221 inpatient hospitalizations representing 114,003 unique adult patients (ge18\\ge 18 years of age) with hospital stays ge24\\ge 24 hours from two geographically distinct US academic medical centers:

    • Hospital A (University of California, San Francisco): 2012–2016; 85,522 training encounters, 9,624 test encounters.
    • Hospital B (University of Chicago Medicine): 2009–2016; 108,948 training encounters, 12,127 test encounters (including de-identified clinical free-text notes).

    Datasets were split at the patient level into 80% development/training, 10% validation, and 10% held-out test sets. Model discrimination was evaluated on the test set using AUROC with 95% confidence intervals derived from 1,000 bootstrap iterations. Across both centers, the unrolled raw sequence data at discharge totaled 46,864,534,945 discrete tokens.

  9. Knowl 9 — Methodological and Practical Limitations of Raw EHR Deep Learning

    limitation

    The deep learning approach on raw FHIR records exhibits several key limitations:

    1. Retrospective Evaluation: The findings are derived entirely from retrospective observational EHR data; prospective clinical trials are necessary to verify whether alerting on these predictions improves clinical outcomes.
    2. Lack of Cross-Site Data Harmonization: Because variables and local site codes are not mapped to a shared semantic ontology, the models do not support direct zero-shot transfer learning between hospitals and require retraining on site-specific historical records.
    3. Computational Complexity: Training deep sequence models over billions of tokens is computationally intensive and demands distributed infrastructure, although per-patient inference execution takes only a few milliseconds.
    4. Unmeasured Component Contributions: The study evaluated aggregate predictive accuracy rather than quantifying the isolated incremental contribution of individual data modalities (such as unstructured notes versus structured laboratory and flowsheet elements) on identical patient cohorts.

Coverage note — None was omitted; all primary architectural methods, predictive tasks, empirical metrics, attribution mechanisms, experimental cohorts, and stated limitations were fully extracted into self-contained knowls.

References

  1. 1.The Digital Universe: Driving Data Growth in Healthcare. Available at: https://www.emc.com/analyst-report/digital-universe-healthcare-vertical-report-ar.pdf (Accessed 23 Feb 2017).
  2. 2.Parikh, R. B., Schwartz, J. S. & Navathe, A. S. Beyond genes and molecules - a precision delivery initiative for precision medicine. N. Engl. J. Med. 376, 1609–1612 (2017).
  3. 3.Parikh, R. B., Kakad, M. & Bates, D. W. Integrating predictive analytics into high-value care: the dawn of precision delivery. JAMA 315, 651–652 (2016).
  4. 4.Bates, D. W., Saria, S., Ohno-Machado, L., Shah, A. & Escobar, G. Big data in health care: using analytics to identify and manage high-risk and high-cost patients. Health Aff. 33, 1123–1131 (2014).
  5. 5.Krumholz, H. M. Big data and new knowledge in medicine: the thinking, training, and tools needed for a learning health system. Health Aff. 33, 1163–1170 (2014).
  6. 6.Jameson, J. L. & Longo, D. L. Precision medicine--personalized, problematic, and promising. N. Engl. J. Med. 372, 2229–2234 (2015).
  7. 7.Goldstein, B. A., Navar, A. M., Pencina, M. J. & Ioannidis, J. P. A. Opportunities and challenges in developing risk prediction models with electronic health records data: a systematic review. J. Am. Med. Inform. Assoc. 24, 198–208 (2017).
  8. 8.Press, G. Cleaning big data: most time-consuming, least enjoyable data science task, survey says. Forbes (2016). Available at: https://www.forbes.com/sites/gilpress/2016/03/23/data-preparation-most-time-consuming-least-enjoyable-data-science-task-survey-says/ (Accessed 22 Oct 2017).
  9. 9.Lohr, S. For Big-Data Scientists, ‘Janitor Work’ Is Key Hurdle to Insights. (NY Times, 2014).
  10. 10.Drew, B. J. et al. Insights into the problem of alarm fatigue with physiologic monitor devices: a comprehensive observational study of consecutive intensive care unit patients. PLoS ONE 9, e110274 (2014).
  11. 11.Chopra, V. & McMahon, L. F. Jr. Redesigning hospital alarms for patient safety: alarmed and potentially dangerous. JAMA 311, 1199–1200 (2014).
  12. 12.Kaukonen, K.-M., Bailey, M., Pilcher, D., Cooper, D. J. & Bellomo, R. Systemic inflammatory response syndrome criteria in defining severe sepsis. N. Engl. J. Med. 372, 1629–1638 (2015).
  13. 13.LeCun, Y., Bengio, Y. & Hinton, G. Deep learning. Nature 521, 436–444 (2015).
  14. 14.Frome, A. et al. DeViSE: a deep visual-semantic embedding model. In Advances in Neural Information Processing Systems 26 (eds Burges, C. J. C., Bottou, L., Welling, M., Ghahramani, Z. & Weinberger, K. Q.), pp 2121–2129 (Curran Associates, Inc. Red Hook, NY, 2013).
  15. 15.Gulshan, V. et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA 316, 2402–2410 (2016).
  16. 16.Wu, Y. et al. Google’s neural machine translation system: bridging the gap between human and machine translation. arXiv [cs.CL] (2016).
  17. 17.Dai, A. M. & Le, Q. V. Semi-supervised sequence learning. In Advances in Neural Information Processing Systems 28 (eds Cortes, C., Lawrence, N. D., Lee, D. D., Sugiyama, M. & Garnett, R.), pp 3079–3087 (Curran Associates, Inc. Red Hook, NY, 2015).
  18. 18.Bengio, Y., Courville, A. & Vincent, P. Representation learning: a review and new perspectives. IEEE. Trans. Pattern Anal. Mach. Intell. 35, 1798–1828 (2013).
  19. 19.Weed, L. L. Medical records that guide and teach. N. Engl. J. Med. 278, 652–657 (1968). concl.
  20. 20.Adler-Milstein, J. et al. Electronic health record adoption In US hospitals: progress continues, but challenges persist. Health Aff. 34, 2174–2180 (2015).
  21. 21.Mandell, L. A. et al. Infectious Diseases Society of America/American Thoracic Society consensus guidelines on the management of community-acquired pneumonia in adults. Clin. Infect. Dis. 44, S27–S72 (2007). Suppl 2.
  22. 22.Lim, W. S., Smith, D. L., Wise, M. P. & Welham, S. A. British Thoracic Society community acquired pneumonia guideline and the NICE pneumonia guideline: how they fit together. BMJ Open Respir. Res. 2, e000091 (2015).
  23. 23.Churpek, M. M. et al. Multicenter comparison of machine learning methods and conventional regression for predicting clinical deterioration on the wards. Crit. Care. Med. 44, 368–374 (2016).
  24. 24.Howell, M. D. et al. Sustained effectiveness of a primary-team-based rapid response system. Crit. Care. Med. 40, 2562–2568 (2012).
  25. 25.Sun, H. et al. Semantic processing of EHR data for clinical research. J. Biomed. Inform. 58, 247–259 (2015).
  26. 26.Newton, K. M. et al. Validation of electronic medical record-based phenotyping algorithms: results and lessons learned from the eMERGE network. J. Am. Med. Inform. Assoc. 20, e147–e154 (2013).
  27. 27.OHDSI. OMOP common data model. Observational health data sciences and informatics. Available at: https://www.ohdsi.org/data-standardization/the-common-data-model/ (Accessed 23 Jan 2018).
  28. 28.Mandel, J. C., Kreda, D. A., Mandl, K. D., Kohane, I. S. & Ramoni, R. B. SMART on FHIR: a standards-based, interoperable apps platform for electronic health records. J. Am. Med. Inform. Assoc. 23, 899–908 (2016).
  29. 29.Miotto, R., Li, L., Kidd, B. A. & Dudley, J. T. Deep patient: an unsupervised representation to predict the future of patients from the electronic health records. Sci. Rep. 6, 26094 (2016).
  30. 30.Lipton, Z. C., Kale, D. C., Elkan, C. & Wetzel, R. Learning to diagnose with LSTM recurrent neural networks. arXiv [cs.LG] (2015).
  31. 31.Aczon, M. et al. Dynamic mortality risk predictions in pediatric critical care using recurrent neural networks. arXiv [stat.ML] (2017).
  32. 32.Choi, E., Bahadori, M. T., Schuetz, A., Stewart, W. F. & Sun, J. Doctor AI: predicting clinical events via recurrent neural networks. In Proceedings of the 1st Machine Learning for Healthcare Conference, vol 56 (eds F. Doshi-Velez, J. Fackler, D. Kale and B. Wallace, J. Wiens) 301–318 (PMLR, Los Angeles, CA, 2016).
  33. 33.Suresh, H. et al. Clinical intervention prediction and understanding using deep networks. arXiv [cs.LG] (PMLR, Los Angeles, CA, USA, 2017).
  34. 34.Razavian, N., Marcus, J. & Sontag, D. Multi-task prediction of disease onsets from longitudinal laboratory tests. In Proceedings of the 1st Machine Learning for Healthcare Conference, (eds F. Doshi-Velez, J. Fackler, D. Kale and B. Wallace, J. Wiens) Vol. 56, pp 73–100 (PMLR, Los Angeles, CA, 2016).
  35. 35.Che, Z., Purushotham, S., Cho, K., Sontag, D. & Liu, Y. Recurrent neural networks for multivariate time series with missing values. arXiv [cs.LG] (2016).
  36. 36.Johnson, A. E. W. et al. MIMIC-III, a freely accessible critical care database. Sci. Data 3, 160035 (2016).
  37. 37.Harutyunyan, H., Khachatrian, H., Kale, D. C. & Galstyan, A. Multitask learning and benchmarking with clinical time series data. arXiv [stat.ML] (2017).
  38. 38.Society of Critical Care Medicine. Critical care statistics. Available at: http://www.sccm.org/Communications/Pages/CriticalCareStats.aspx (Accessed 25 Jan 2018).
  39. 39.American Hospital Association. Fast facts on U.S. Hospitals, 2018. Available at: https://www.aha.org/statistics/fast-facts-us-hospitals (Accessed 25 Jan 2018).
  40. 40.Shickel, B., Tighe, P., Bihorac, A. & Rashidi, P. Deep EHR: a survey of recent advances in deep learning techniques for electronic health record (EHR) analysis. arXiv [cs.LG] (2017).
  41. 41.Bergstrom, N., Braden, B. J., Laguzza, A. & Holman, V. The braden scale for predicting pressure sore risk. Nurs. Res. 36, 205–210 (1987).
  42. 42.Tabak, Y. P., Sun, X., Nunez, C. M. & Johannes, R. S. Using electronic health record data to develop inpatient mortality predictive model: Acute Laboratory Risk of Mortality Score (ALaRMS). J. Am. Med. Inform. Assoc. 21, 455–463 (2014).
  43. 43.Nguyen, O. K. et al. Predicting all-cause readmissions using electronic health record data from the entire hospitalization: Model development and comparison. J. Hosp. Med. 11, 473–480 (2016).
  44. 44.Liu, V., Kipnis, P., Gould, M. K. & Escobar, G. J. Length of stay predictions: improvements through the use of automated laboratory and comorbidity variables. Med. Care 48, 739–744 (2010).
  45. 45.Walsh, C. & Hripcsak, G. The effects of data sources, cohort selection, and outcome definition on a predictive model of risk of thirty-day hospital readmissions. J. Biomed. Inform. 52, 418–426 (2014).
  46. 46.Kellett, J. & Kim, A. Validation of an abbreviated VitalpacTM Early Warning Score (ViEWS) in 75,419 consecutive admissions to a Canadian regional hospital. Resuscitation 83, 297–302 (2012).
  47. 47.Escobar, G. J. et al. Risk-adjusting hospital inpatient mortality using automated inpatient, outpatient, and laboratory databases. Med. Care. 46, 232–239 (2008).
  48. 48.van Walraven, C. et al. Derivation and validation of an index to predict early death or unplanned readmission after discharge from hospital to the community. CMAJ 182, 551–557 (2010).
  49. 49.Yamana, H., Matsui, H., Fushimi, K. & Yasunaga, H. Procedure-based severity index for inpatients: development and validation using administrative database. BMC Health Serv. Res. 15, 261 (2015).
  50. 50.Pine, M. et al. Modifying ICD-9-CM coding of secondary diagnoses to improve risk-adjustment of inpatient mortality rates. Med. Decis. Making 29, 69–81 (2009).
  51. 51.Smith, G. B., Prytherch, D. R., Meredith, P., Schmidt, P. E. & Featherstone, P. I. The ability of the National Early Warning Score (NEWS) to discriminate patients at risk of early cardiac arrest, unanticipated intensive care unit admission, and death. Resuscitation 84, 465–470 (2013).
  52. 52.Khurana, H. S. et al. Real-time automated sampling of electronic medical records predicts hospital mortality. Am. J. Med. 129, 688–698.e2 (2016).
  53. 53.Rothman, M. J., Rothman, S. I. & Beals, J. 4th Development and validation of a continuous measure of patient condition using the electronic medical record. J. Biomed. Inform. 46, 837–848 (2013).
  54. 54.Finlay, G. D., Rothman, M. J. & Smith, R. A. Measuring the modified early warning score and the Rothman index: advantages of utilizing the electronic medical record in an early warning system. J. Hosp. Med. 9, 116–119 (2014).
  55. 55.Zapatero, A. et al. Predictive model of readmission to internal medicine wards. Eur. J. Intern. Med. 23, 451–456 (2012).
  56. 56.Shams, I., Ajorlou, S. & Yang, K. A predictive analytics approach to reducing 30-day avoidable readmissions among patients with heart failure, acute myocardial infarction, pneumonia, or COPD. Health Care. Manag. Sci. 18, 19–34 (2015).
  57. 57.Tsui, E., Au, S. Y., Wong, C. P., Cheung, A. & Lam, P. Development of an automated model to predict the risk of elderly emergency medical admissions within a month following an index hospital visit: A Hong Kong experience. Health Inform. J. 21, 46–56 (2013).
  58. 58.Choudhry, S. A. et al. A public-private partnership develops and externally validates a 30-day hospital readmission risk prediction model. Online J. Public Health Inform. 5, 219 (2013).
  59. 59.Caruana, R. et al. Intelligible models for healthcare: predicting pneumonia risk and hospital 30-day readmission. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp 1721–1730. http://doi.acm.org/10.1145/2783258.2788613 (ACM, Sydney, NSW, Australia, 2015).
  60. 60.Tonkikh, O. et al. Functional status before and during acute hospitalization and readmission risk identification. J. Hosp. Med. 11, 636–641 (2016).
  61. 61.Betihavas, V. et al. An absolute risk prediction model to determine unplanned cardiovascular readmissions for adults with chronic heart failure. Heart Lung Circ. 24, 1068–1073 (2015).
  62. 62.Whitlock, T. L. et al. A scoring system to predict readmission of patients with acute pancreatitis to the hospital within thirty days of discharge. Clin. Gastroenterol. Hepatol. 9, 175–180 (2011). quiz e18.
  63. 63.Coleman, E. A., Min, S.-J., Chomiak, A. & Kramer, A. M. Posthospital care transitions: patterns, complications, and risk identification. Health Serv. Res. 39, 1449–1465 (2004).
  64. 64.Graboyes, E. M., Liou, T.-N., Kallogjeri, D., Nussenbaum, B. & Diaz, J. A. Risk factors for unplanned hospital readmission in otolaryngology patients. Otolaryngol. Head Neck Surg. 149, 562–571 (2013).
  65. 65.He, D., Mathews, S. C., Kalloo, A. N. & Hutfless, S. Mining high-dimensional administrative claims data to predict early hospital readmissions. J. Am. Med. Inform. Assoc. 21, 272–279 (2014).
  66. 66.Futoma, J., Morris, J. & Lucas, J. A comparison of models for predicting early hospital readmissions. J. Biomed. Inform. 56, 229–238 (2015).
  67. 67.Donzé, J., Aujesky, D., Williams, D. & Schnipper, J. L. Potentially avoidable 30-day hospital readmissions in medical patients: derivation and validation of a prediction model. JAMA Intern. Med. 173, 632–638 (2013).
  68. 68.Perotte, A. et al. Diagnosis code assignment: models and evaluation metrics. J. Am. Med. Inform. Assoc. 21, 231–237 (2014).
  69. 69.Krumholz, H. M., Terry, S. F. & Waldstreicher, J. Data acquisition, curation, and use for a continuously learning health system. JAMA 316, 1669–1670 (2016).
  70. 70.Grumbach, K., Lucey, C. R. & Claiborne Johnston, S. Transforming from centers of learning to learning health systems: the challenge for academic health centers. JAMA 311, 1109–1110 (2014).
  71. 71.Halamka, J. D. & Tripathi, M. The HITECH era in retrospect. N. Engl. J. Med. 377, 907–909 (2017).
  72. 72.Bates, D. W. et al. Ten commandments for effective clinical decision support: making the practice of evidence-based medicine a reality. J. Am. Med. Inform. Assoc. 10, 523–530 (2003).
  73. 73.Obermeyer, Z. & Emanuel, E. J. Predicting the future --- big data, machine learning, and clinical medicine. N. Engl. J. Med. 375, 1216–1219 (2016).
  74. 74.Avati, A. et al. Improving palliative care with deep learning. arXiv [cs.CY] (2017).
  75. 75.Health Level 7. FHIR Specification Home Page (2017). Available at: http://hl7.org/fhir/ (Accessed 3 Aug 2017).
  76. 76.Escobar, G. J. et al. Nonelective rehospitalizations and postdischarge mortality: predictive models suitable for use in real time. Med. Care. 53, 916–923 (2015).
  77. 77.2016 Measure updates and specifications report: hospital-wide all-cause unplanned readmission --- version 5.0. Yale--New Haven Health Services Corporation/Center for Outcomes Research & Evaluation (New Haven, CT, 2016).
  78. 78.Knaus, W. A., Draper, E. A., Wagner, D. P. & Zimmerman, J. E. APACHE II: a severity of disease classification system. Crit. Care Med. 13, 818–829 (1985).
  79. 79.Kansagara, D. et al. Risk prediction models for hospital readmission: a systematic review. JAMA 306, 1688–1698 (2011).
  80. 80.Hochreiter, S. & Schmidhuber, J. Long short-term memory. Neural Comput. 9, 1735–1780 (1997).
  81. 81.Rokach, L. Ensemble-based Classifiers. Artif. Intell. Rev. 33,1–39 (2010).
  82. 82.Cabitza, F., Rasoini, R. & Gensini, G. F. Unintended consequences of machine learning in medicine. JAMA 18, 517–518 (2017).
  83. 83.Bahdanau, D., Cho, K. & Bengio, Y. Neural machine translation by jointly learning to align and translate. arXiv [cs.CL] (2014).
  84. 84.Pedregosa, F. et al. Scikit-learn: machine learning in python. J. Mach. Learn. Res. 12, 2825–2830 (2011).
  85. 85.Pencina, M. J. & D’Agostino, R. B. Sr. Evaluating discrimination of risk prediction models: the C statistic. JAMA 314, 1063–1064 (2015).
  86. 86.Kramer, A. A. & Zimmerman, J. E. Assessing the calibration of mortality benchmarks in critical care: the Hosmer-Lemeshow test revisited. Crit. Care Med. 35, 2052–2056 (2007).
  87. 87.Romero-Brufau, S., Huddleston, J. M., Escobar, G. J. & Liebow, M. Why the C-statistic is not informative to evaluate early warning scores and what metrics to use. Crit. Care 19, 285 (2015).
  88. 88.SciKit Learn. SciKit learn documentation on area under the curve scores. Available at: http://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html (Accessed 3 Aug 2017).
  89. 89.SciKit Learn. SciKit learn documentation on F1 score. Available at: http://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html (Accessed 3 Aug 2017).

Citation

MLA
Rajkomar, A., et al. “Scalable and Accurate Deep Learning with Electronic Health Records”. Npj Digital Medicine, vol. 1, no. 1, 2018, https://doi.org/10.1038/s41746-018-0029-1.
APA
Rajkomar, A., Oren, E., Chen, K., Dai, A. M., Hajaj, N., Hardt, M., Liu, P. J., Liu, X., Marcus, J., Sun, M., Sundberg, P., Yee, H., Zhang, K., Zhang, Y., Flores, G., Duggan, G. E., Irvine, J., Le, Q., Litsch, K., … Dean, J. (2018). Scalable and accurate deep learning with electronic health records. Npj Digital Medicine, 1(1). https://doi.org/10.1038/s41746-018-0029-1
Chicago
Rajkomar, A., E. Oren, K. Chen, et al. 2018. “Scalable and Accurate Deep Learning with Electronic Health Records”. Npj Digital Medicine 1 (1). https://doi.org/10.1038/s41746-018-0029-1.
Harvard
Rajkomar, A. et al. (2018) “Scalable and accurate deep learning with electronic health records”, npj Digital Medicine, 1(1). Available at: https://doi.org/10.1038/s41746-018-0029-1.
Vancouver
1. Rajkomar A, Oren E, Chen K, et al (2018) Scalable and accurate deep learning with electronic health records. npj Digital Medicine. https://doi.org/10.1038/s41746-018-0029-1

BibTeX

@article{Rajkomar_2018, title={Scalable and accurate deep learning with electronic health records}, volume={1}, ISSN={2398-6352}, url={http://dx.doi.org/10.1038/s41746-018-0029-1}, DOI={10.1038/s41746-018-0029-1}, number={1}, journal={npj Digital Medicine}, publisher={Springer Science and Business Media LLC}, author={Rajkomar, Alvin and Oren, Eyal and Chen, Kai and Dai, Andrew M. and Hajaj, Nissan and Hardt, Michaela and Liu, Peter J. and Liu, Xiaobing and Marcus, Jake and Sun, Mimi and Sundberg, Patrik and Yee, Hector and Zhang, Kun and Zhang, Yi and Flores, Gerardo and Duggan, Gavin E. and Irvine, Jamie and Le, Quoc and Litsch, Kurt and Mossin, Alexander and Tansuwan, Justin and Wang, De and Wexler, James and Wilson, Jimbo and Ludwig, Dana and Volchenboum, Samuel L. and Chou, Katherine and Pearson, Michael and Madabushi, Srinivasan and Shah, Nigam H. and Butte, Atul J. and Howell, Michael D. and Cui, Claire and Corrado, Greg S. and Dean, Jeffrey}, year={2018}, month=May }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF