ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission

Kexin HuangJaan AltosaarRajesh Ranganath

article2019arXiv1,380 citations

Introduces ClinicalBERT, a domain-specific transformer that learns representations from clinical notes to uncover medical relationships and improve 30-day hospital readmission predictions.

Listen

Hospital readmissions place a heavy financial burden on healthcare systems—estimated at $17.9 billion annually in the United States, with roughly 76% considered preventable—while also reducing patient quality of life. Electronic health records contain rich narrative clinical notes that capture critical context such as symptom progression and diagnostic rationale. However, because these unstructured notes are sparse, complex, and full of specialized clinical jargon, existing hospital predictive models rely almost entirely on structured tabular data, missing valuable insights.

The article evaluates whether adapting modern language models to clinical text can produce deep language representations that improve the prediction of 30-day hospital readmissions dynamically during a patient's stay and at discharge.

To address this, the authors developed ClinicalBERT, a specialized deep learning language model based on the transformer architecture. They adapted the model to medical text by pre-training it on over 2 million unstructured clinical notes from the MIMIC-III database, covering more than 34,000 intensive care unit patients between 2001 and 2012. The model was then fine-tuned on the specific task of predicting 30-day hospital readmissions. To make the evaluation clinically actionable and address hospital alarm fatigue, the authors evaluated performance at a strict fixed positive predictive value of 80% (precision), alongside standard discrimination metrics, comparing ClinicalBERT against established baselines such as standard language models, bag-of-words models, and recurrent neural networks.

The analysis yielded several key findings. First, ClinicalBERT achieved substantially higher clinical language understanding, reaching 85.7% accuracy on masked token prediction compared to 49.5% for standard baseline models. Second, its learned medical concept relationships correlated more closely with human physician judgments (0.670 correlation) than standard biomedical word embedding baselines. Third, for 30-day readmission prediction at discharge, ClinicalBERT outperformed all baselines, achieving a recall of 24.2% at 80% precision compared to 17.2% for standard base models. Fourth, the model demonstrated superior predictive capability early in a patient's stay, showing higher recall within the first 24 to 72 hours of admission than all comparative baselines. Finally, the model's internal attention mechanisms successfully highlighted clinically relevant keywords (such as "chronic" or "acute"), offering interpretable explanations for its risk scores.

These findings indicate that specialized language models can unlock the predictive value hidden in unstructured clinical narratives. Implementing such models allows care teams to dynamically assess readmission risk within the first few days of admission rather than waiting for post-discharge summaries. This shift enables early clinical interventions, potentially reducing costly preventable readmissions, optimizing intensive care workflows, and improving patient outcomes without overwhelming clinicians with false alarms.

For practical adoption, organizations should retrain or fine-tune ClinicalBERT using their own internal institutional records, as language patterns and documentation styles vary across hospital systems and outpatient settings. Model outputs can then be integrated into clinical decision support tools to monitor readmission, mortality, or length of stay.

Readers should note certain limitations. The primary study utilized data from intensive care units within a single medical center, and long notes required heuristic splitting strategies that may not capture all complex interactions across very lengthy patient histories. Nevertheless, the strong empirical improvements across standard and clinically focused metrics provide solid confidence in the framework's effectiveness for clinical text modeling.

Cover for ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission

Abstract

Clinical notes contain information about patients that goes beyond structured data like lab values and medications. However, clinical notes have been underused relative to structured data, because notes are high-dimensional and sparse. This work develops and evaluates representations of clinical notes using bidirectional transformers (ClinicalBERT). ClinicalBERT uncovers high-quality relationships between medical concepts as judged by humans. ClinicalBert outperforms baselines on 30-day hospital readmission prediction using both discharge summaries and the first few days of notes in the intensive care unit. Code and model parameters are available.

Table of Contents

  • 1 Introduction
  • 1.1 Background
  • 1.2 Significance
  • 2 Methods
  • 2.1 BERT Model
  • 2.2 Clinical Text Embedding
  • 2.3 Self-Attention Mechanism
  • 2.4 Pre-training ClinicalBERT
  • 2.5 Fine-tuning ClinicalBERT
  • 3 Empirical Study
  • 3.1 Data
  • 3.2 Empirical Study I: Language Modeling and Clinical Word Similarity
  • 3.2.1 Clinical Language Modeling.
  • 3.2.2 Qualitative Analysis.
  • 3.2.3 Quantitative Analysis.
  • 3.3 Empirical Study II: 30-Day Hospital Readmission Prediction
  • 3.3.1 Cohort.
  • 3.3.2 Scalable Readmission Prediction.
  • 3.3.3 Evaluation.
  • 3.3.4 Models.
  • 3.3.5 Readmission Prediction with Discharge Summaries.
  • 3.3.6 Readmission Prediction with Early Clinical Notes.
  • 3.3.7 Interpretability.
  • 4 Guidelines on using ClinicalBERT in Practice
  • 5 Discussion
  • 6 Acknowledgements
  • References
  • A Hyperparameters and training details
  • B Preprocessing Notes for Pretraining ClinicalBERT

Knowls

  1. Knowl 1 — Subsequence Probability Aggregation for Patient Readmission Prediction

    model/method

    Because transformer architectures have a fixed maximum sequence length (e.g., 512 tokens), a patient's concatenated clinical notes are divided into nn distinct subsequences. Given model-predicted readmission probabilities for each subsequence, the aggregate patient-level probability P(readmit=1∣hpatient)P(\text{readmit} = 1 \mid h_{\text{patient}}) is computed as a weighted trade-off between the maximum subsequence probability and the mean subsequence probability:

    P(readmit=1∣hpatient)=Pmax⁡n+Pmeann(n/c)1+n/cP(\text{readmit} = 1 \mid h_{\text{patient}}) = \frac{P_{\max}^n + P_{\text{mean}}^n (n / c)}{1 + n / c}

    where:

    • nn is the total number of subsequences for the patient.
    • Pmax⁡n=max⁡i∈{1,…,n}PiP_{\max}^n = \max_{i \in \{1, \dots, n\}} P_i is the maximum predicted readmission probability across all nn subsequences.
    • Pmeann=1n∑i=1nPiP_{\text{mean}}^n = \frac{1}{n} \sum_{i=1}^n P_i is the average predicted readmission probability across all nn subsequences.
    • cc is a positive scaling parameter chosen on validation data (c=2c = 2).
    • hpatienth_{\text{patient}} denotes the implicit joint representation of all notes for the patient.

    This formulation leverages Pmax⁡nP_{\max}^n to capture localized, high-risk signals present in only a few critical subsequences (such as acute deterioration) while using PmeannP_{\text{mean}}^n and the scaling weight n/cn/c to prevent false alarms caused by noisy isolated subsequences, scaling up the mean's weight for patients with large numbers of notes. This aggregation outperforms a pure mean prediction by 3% to 8%.

  2. Knowl 2 — 30-Day Hospital Readmission Prediction Using Discharge Summaries

    empirical result

    On the MIMIC-III dataset (filtered to 34,560 non-newborn patients with 2,963 positive 30-day readmissions), ClinicalBERT was evaluated against three baseline models using patient discharge summaries under 5-fold cross-validation.

    Model AUROC AUPRC RP80
    ClinicalBERT 0.714±0.0180.714 \pm 0.018 0.701±0.0210.701 \pm 0.021 0.242±0.1110.242 \pm 0.111
    Bag-of-words 0.684±0.0250.684 \pm 0.025 0.674±0.0270.674 \pm 0.027 0.217±0.1190.217 \pm 0.119
    Bi-LSTM 0.694±0.0250.694 \pm 0.025 0.686±0.0290.686 \pm 0.029 0.223±0.1030.223 \pm 0.103
    BERT (Standard) 0.692±0.0190.692 \pm 0.019 0.678±0.0160.678 \pm 0.016 0.172±0.1010.172 \pm 0.101

    ClinicalBERT outperforms the Bag-of-words logistic regression baseline, a Bidirectional LSTM with Word2Vec embeddings, and standard BERT pre-trained on Wikipedia and BookCorpus across Area Under the Receiver Operating Characteristic (AUROC), Area Under the Precision-Recall Curve (AUPRC), and Recall at 80% Precision (RP80).

  3. Knowl 3 — Dynamic Early-Stage Readmission Prediction Using 24-48h and 48-72h ICU Notes

    empirical result

    To evaluate actionable readmission risk before patient discharge, models were tested using notes collected within the first 24–48 hours and 48–72 hours of intensive care unit admission (excluding patients discharged before the cutoff time). Results are reported as the mean and standard deviation over 5 independent runs on MIMIC-III:

    Model Cutoff time AUROC AUPRC RP80
    ClinicalBERT 24–48h 0.674±0.0380.674 \pm 0.038 0.674±0.0390.674 \pm 0.039 0.154±0.0990.154 \pm 0.099
    ClinicalBERT 48–72h 0.672±0.0390.672 \pm 0.039 0.677±0.0360.677 \pm 0.036 0.170±0.1140.170 \pm 0.114
    Bag-of-words 24–48h 0.648±0.0290.648 \pm 0.029 0.650±0.0270.650 \pm 0.027 0.144±0.0940.144 \pm 0.094
    Bag-of-words 48–72h 0.654±0.0350.654 \pm 0.035 0.657±0.0260.657 \pm 0.026 0.122±0.1060.122 \pm 0.106
    Bi-LSTM 24–48h 0.649±0.0440.649 \pm 0.044 0.660±0.0360.660 \pm 0.036 0.143±0.0800.143 \pm 0.080
    Bi-LSTM 48–72h 0.656±0.0350.656 \pm 0.035 0.668±0.0280.668 \pm 0.028 0.150±0.0810.150 \pm 0.081
    BERT (Standard) 24–48h 0.659±0.0340.659 \pm 0.034 0.656±0.0210.656 \pm 0.021 0.141±0.0800.141 \pm 0.080
    BERT (Standard) 48–72h 0.661±0.0280.661 \pm 0.028 0.668±0.0210.668 \pm 0.021 0.167±0.0880.167 \pm 0.088

    ClinicalBERT consistently outperforms all baselines at early timepoints, demonstrating that domain-specific transformer pre-training enables earlier and more accurate identification of patients at high risk of 30-day readmission.

  4. Knowl 4 — Two-Stage Pre-Training Protocol and Cross-Validation for ClinicalBERT

    experimental setup

    ClinicalBERT is initialized with the BERT-Base architecture (hidden dimension d=768d = 768, 12 layers, 12 attention heads per layer). Pre-training uses the Adam optimizer with a learning rate of 2×10−52 \times 10^{-5} on MIMIC-III clinical notes in two sequential stages:

    1. Sequence packing up to a maximum length of 128 tokens for 100,000 steps with a batch size of 64.
    2. Sequence packing up to a maximum length of 512 tokens for an additional 100,000 steps with a batch size of 8.

    To prevent data leakage between pre-training and downstream evaluation, all hospital admissions are partitioned into 5 patient-level folds. In each independent cross-validation run, four folds are used for both pre-training and fine-tuning training, while the remaining fifth fold is held out strictly for validation and testing.

  5. Knowl 5 — Rule-Based Clinical Text Preprocessing and Sentence Segmentation

    algorithm

    To prepare unstructured clinical notes for masked language modeling and next sentence prediction without sentence boundary errors caused by medical abbreviations and formatting, notes are preprocessed using the following procedure:

    Input: Raw clinical note text TT
    Output: List of clean segmented sentences SS
    T←T \leftarrow convert TT to lowercase
    T←T \leftarrow remove line breaks and carriage returns from TT
    T←T \leftarrow remove de-identification bracket tags and special character sequences (e.g., '==', '--') from TT
    T←T \leftarrow remove rule-based numbering patterns that disrupt segmentation (e.g., '1.2.')
    T←T \leftarrow replace abbreviations containing periods (e.g., replace 'M.D.' and 'dr.' with 'MD' and 'Dr')
    RawSentences←RawSentences \leftarrow apply SpaCy rule-based sentence segmentation to TT
    S←S \leftarrow empty list
    for each segment ss in RawSentencesRawSentences do
        if word_count(ss) < 20 and SS is not empty then
            S[last]←S[\text{last}] \leftarrow concatenate(S[last]S[\text{last}], ' ', ss)
        else
            append ss to SS
        end if
    end for
    return SS
  6. Knowl 6 — Recall at 80% Precision (RP80) for Clinical Alarm Fatigue Reduction

    definition

    Recall at Precision of 80% (RP80) is an evaluation metric designed for clinical decision-support systems where alarm fatigue is a primary operational bottleneck. In RP80, the model's decision threshold is set to the point on its precision-recall curve where the positive predictive value (precision) is exactly 0.800.80 (i.e., at most a 20% false positive rate among flagged cases):

    Precision=True PositivesTrue Positives+False Positives=0.80\text{Precision} = \frac{\text{True Positives}}{\text{True Positives} + \text{False Positives}} = 0.80

    RP80 is the recall (sensitivity) achieved by the model at this fixed precision threshold. It measures the fraction of true positive events detected while guaranteeing a bounded, clinically acceptable rate of false alarms.

  7. Knowl 7 — Fine-Tuning Architecture and Classification Head for Readmission

    model/method

    For downstream 30-day hospital readmission prediction, ClinicalBERT processes an input note sequence and extracts the final hidden state vector h[CLS]∈R768h_{[\text{CLS}]} \in \mathbb{R}^{768} corresponding to the classification token [CLS]. The probability of readmission is computed via a multi-layer binary classification head:

    P(readmit=1∣h[CLS])=σ(Wh[CLS])P(\text{readmit} = 1 \mid h_{[\text{CLS}]}) = \sigma(W h_{[\text{CLS}]})

    where σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} is the logistic sigmoid function, and the classifier is parameterized as a three-layer neural network with layer dimensions 768×2048768 \times 2048, 2048×7682048 \times 768, and 768×1768 \times 1.

    Fine-tuning is conducted for 3 epochs with a batch size of 56, an Adam learning rate of 2×10−52 \times 10^{-5}, and early stopping monitored on validation loss to maximize the binary classification log-likelihood.

  8. Knowl 8 — Language Modeling Accuracy on MIMIC-III Corpus

    empirical result

    ClinicalBERT was evaluated on unsupervised masked language modeling (predicting held-out tokens) and next sentence prediction (binary classification of whether two text spans are consecutive) across 5 folds of the MIMIC-III clinical note corpus.

    Model Language Modeling Next Sentence Prediction
    ClinicalBERT 0.857±0.0020.857 \pm 0.002 0.994±0.0030.994 \pm 0.003
    BERT (Standard) 0.495±0.0070.495 \pm 0.007 0.539±0.0060.539 \pm 0.006

    Standard BERT fails to model clinical notes effectively without domain-specific pre-training due to specialized clinical jargon, non-standard grammar, and abbreviations.

  9. Knowl 9 — Correlation with Physician-Assessed Medical Concept Similarity

    empirical result

    To evaluate semantic representation quality, embedding representations were evaluated on a benchmark of 30 medical concept pairs with physician-assigned similarity ratings ranging from 1.0 (least similar) to 4.0 (most similar). Concept representations in ClinicalBERT were obtained by computing the 768-dimensional mean subword embedding from the sum of the last four transformer encoder layers, and concept-pair similarity was measured via cosine similarity.

    Model Pearson Correlation (rr)
    ClinicalBERT 0.670
    Word2Vec 0.553
    FastText 0.487

    ClinicalBERT achieves a Pearson correlation of r=0.670r = 0.670 with physician ratings, outperforming Word2Vec (r=0.553r = 0.553, evaluated on the 27 in-vocabulary pairs) and FastText (r=0.487r = 0.487), each trained on the full MIMIC-III corpus (2.8B words).

  10. Knowl 10 — Self-Attention Mechanism and Interpretability in ClinicalBERT

    model/method

    ClinicalBERT contains 144 self-attention mechanisms (12 attention heads across 12 transformer encoder layers). For an input query token vector q∈Rdq \in \mathbb{R}^d and a key matrix K∈RN×dK \in \mathbb{R}^{N \times d} (where d=64d = 64 is the per-head key/query dimension and NN is the sequence length), the attention distribution over key tokens is:

    AttentionWeight(q,K)=softmax(qK⊤d)\text{AttentionWeight}(q, K) = \text{softmax}\left(\frac{q K^\top}{\sqrt{d}}\right)

    Inspecting the attention distributions across heads provides interpretable explanations for predictions. Specific attention heads specialize into bag-of-words-like feature detectors by assigning uniformly high attention weights to clinically predictive terms (such as 'acute' and 'chronic') across all query tokens in the sentence.

Coverage note — Foundational transformer equations from prior work (Vaswani et al. / Devlin et al.) that were not modified specifically for ClinicalBERT were omitted or integrated directly into self-contained knowls.

References

  1. 1.E. Alsentzer, J. R. Murphy, W. Boag, W.-H. Weng, D. Jin, T. Naumann, and M. B. A. McDermott. “Publicly Available Clinical BERT Embeddings”. In: arXiv:1904.03323 (2019).
  2. 2.G. F. Anderson and E. P. Steinberg. “Hospital readmissions in the Medicare population”. In: New England Journal of Medicine 21 (1984).
  3. 3.D. Banerjee, C. Thompson, C. Kell, R. Shetty, Y. Vetteth, H. Grossman, A. DiBiase, and M. Fowler. “An informatics-based approach to reducing heart failure all-cause readmissions: the Stanford heart failure dashboard”. In: Journal of the American Medical Informatics Association 3 (2016).
  4. 4.S. Basu Roy, A. Teredesai, K. Zolfaghar, R. Liu, D. Hazel, S. Newman, and A. Marinez. “Dynamic Hierarchical Classification for Patient Risk-of-Readmission”. In: Knowledge Discovery and Data Mining (2015).
  5. 5.W. Boag, D. Doss, T. Naumann, and P. Szolovits. “What’s in a Note? Unpacking Predictive Value in Clinical Note Representations”. In: AMIA Joint Summits on Translational Science (2018).
  6. 6.P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov. “Enriching word vectors with subword information”. In: Transactions of the Association for Computational Linguistics (2017).
  7. 7.X. Cai, O. Perez-Concha, E. Coiera, F. Martin-Sanchez, R. Day, D. Roffe, and B. Gallego. “Real-time prediction of mortality, readmission, and length of stay using electronic health record data”. In: Journal of the American Medical Informatics Association 3 (2015).
  8. 8.R. Caruana, Y. Lou, J. Gehrke, P. Koch, M. Sturm, and N. Elhadad. “Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission”. In: Knowledge Discovery and Data Mining. 2015.
  9. 9.W. W. Chapman, P. M. Nadkarni, L. Hirschman, L. W. D’Avolio, G. K. Savova, and O. Uzuner. “Overcoming barriers to NLP for clinical text: the role of shared tasks and the need for additional creative solutions”. In: Journal of the American Medical Informatics Association 5 (2011).
  10. 10.B. Chiu, G. Crichton, A. Korhonen, and S. Pyysalo. “How to Train good Word Embeddings for Biomedical NLP”. In: Proceedings of the 15th Workshop on Biomedical Natural Language Processing, ACL 2016 ().
  11. 11.J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”. In: arXiv:1810.04805 (2018).
  12. 12.J. Futoma, J. Morris, and J. Lucas. “A comparison of models for predicting early hospital readmissions”. In: Journal of Biomedical Informatics (2015).
  13. 13.B. A. Goldstein, A. M. Navar, M. J. Pencina, and J. P. A. Ioannidis. “Opportunities and challenges in developing risk prediction models with electronic health records data: a systematic review”. In: Journal of the American Medical Informatics Association (2017).
  14. 14.S. Hochreiter and J. Schmidhuber. “Long Short-Term Memory”. In: Neural Computation 8 (1997).
  15. 15.A. E. W. Johnson, T. J. Pollard, L. Shen, L.-w. H. Lehman, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. Anthony Celi, and R. G. Mark. “MIMIC-III, a freely accessible critical care database”. In: Scientific Data (2016).
  16. 16.A. J. Kind and M. A. Smith. “Documentation of mandated discharge summary components in transitions from acute to subacute care”. In: Agency for Healthcare Research and Quality (2008).
  17. 17.J. Lee, W. Yoon, S. Kim, D. Kim, S. Kim, C. H. So, and J. Kang. “BioBERT: a pre-trained biomedical language representation model for biomedical text mining”. In: arXiv:1901.08746 (2019).
  18. 18.J. Liu, Z. Zhang, and N. Razavian. “Deep EHR: Chronic Disease Prediction Using Medical Notes”. In: Proceedings of the 3rd Machine Learning for Healthcare Conference. 2018.
  19. 19.L. van der Maaten and G. Hinton. “Visualizing data using t-SNE”. In: Journal of Machine Learning Research (2008).
  20. 20.T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. “Distributed representations of words and phrases and their compositionality”. In: Advances in Neural Information Processing Systems. 2013.
  21. 21.C. A. Pedersen, P. J. Schneider, and D. J. Scheckelhoff. “ASHP national survey of pharmacy practice in hospital settings: Prescribing and transcribing—2016”. In: American Journal of Health-System Pharmacy 17 (2017).
  22. 22.T. Pedersen, S. V. Pakhomov, S. Patwardhan, and C. G. Chute. “Measures of semantic similarity and relatedness in the biomedical domain”. In: Journal of Biomedical Informatics 3 (2007).
  23. 23.J. Pennington, R. Socher, and C. Manning. “Glove: Global Vectors for Word Representation”. In: EMNLP (2014).
  24. 24.M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer. “Deep contextualized word representations”. In: arXiv:1802.05365 (2018).
  25. 25.A. Radford. “Improving Language Understanding by Generative Pre-Training”. https:// s3 - us - west - 2. amazonaws. com/openai-assets/research-covers/language-unsupervised/ language_understanding_paper.pdf. 2018.
  26. 26.A. Rajkomar, E. Oren, K. Chen, A. M. Dai, N. Hajaj, M. Hardt, P. J. Liu, X. Liu, J. Marcus, M. Sun, P. Sundberg, H. Yee, K. Zhang, Y. Zhang, G. Flores, G. E. Duggan, J. Irvine, Q. Le, K. Litsch, A. Mossin, J. Tansuwan, D. Wang, J. Wexler, J. Wilson, D. Ludwig, S. L. Volchenboum, K. Chou, M. Pearson, S. Madabushi, N. H. Shah, A. J. Butte, M. D. Howell, C. Cui, G. S. Corrado, and J. Dean. “Scalable and accurate deep learning with electronic health records”. In: NPJ Digital Medicine 1 (2018).
  27. 27.M. Schuster and K. K. Paliwal. “Bidirectional recurrent neural networks”. In: IEEE Trans. Signal Processing (1997).
  28. 28.S. Sendelbach and M. Funk. “Alarm fatigue: a patient safety concern”. In: AACN Advanced Critical Care 4 (2013).
  29. 29.R. Sennrich, B. Haddow, and A. Birch. “Neural Machine Translation of Rare Words with Subword Units”. In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics. 2016.
  30. 30.B. Shickel, P. J. Tighe, A. Bihorac, and P. Rashidi. “Deep EHR: A survey of recent advances in deep learning techniques for electronic health record (EHR) analysis”. In: IEEE Journal of Biomedical and Health Informatics 5 (2018).
  31. 31.Y. Si, J. Wang, H. Xu, and K. Roberts. “Enhancing clinical concept extraction with contextual embeddings”. In: Journal of the American Medical Informatics Association 11 (2019).
  32. 32.C. Van Walraven, R. Seth, P. C. Austin, and A. Laupacis. “Effect of discharge summary availability during post-discharge visits on hospital readmission”. In: Journal of General Internal Medicine 3 (2002).
  33. 33.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. “Attention is all you need”. In: Advances in Neural Information Processing Systems. 2017.
  34. 34.Y. Wang, S. Liu, N. Afzal, M. Rastegar-Mojarad, L. Wang, F. Shen, P. Kingsbury, and H. Liu. “A comparison of word embeddings for the biomedical natural language processing”. In: Journal of Biomedical Informatics (2018).
  35. 35.W.-H. Weng, K. B. Wagholikar, A. T. McCray, P. Szolovits, and H. C. Chueh. “Medical Subdomain Classification of Clinical Notes Using a Machine Learning-Based Natural Language Processing Approach”. In: BMC Medical Informatics and Decision Making 1 (2017).
  36. 36.C. Xiao, E. Choi, and J. Sun. “Opportunities and challenges in developing deep learning models using electronic health records data: a systematic review”. In: Journal of the American Medical Informatics Association 10 (2018).
  37. 37.K.-H. Yu, A. L. Beam, and I. S. Kohane. “Artificial intelligence in healthcare”. In: Nature Biomedical Engineering 10 (2018).
  38. 38.Y. Zhang, R. Jin, and Z.-H. Zhou. “Understanding bag-of-words model: a statistical framework”. In: International Journal of Machine Learning and Cybernetics 1 (2010).
  39. 39.Y. Zhang, R. Henao, Z. Gan, Y. Li, and L. Carin. “Multi-Label Learning from Medical Plain Text with Convolutional Residual Models”. In: Proceedings of the 3rd Machine Learning for Healthcare Conference. 2018.
  40. 40.R. B. Zuckerman, S. H. Sheingold, E. J. Orav, J. Ruhter, and A. M. Epstein. “Readmissions, observation, and the hospital readmissions reduction program”. In: New England Journal of Medicine 16 (2016).

Citation

MLA
Huang, K., et al. “ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission”. arXiv, 2019, http://arxiv.org/abs/1904.05342v3.
APA
Huang, K., Altosaar, J., & Ranganath, R. (2019). ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission. arXiv. http://arxiv.org/abs/1904.05342v3
Chicago
Huang, K., J. Altosaar, and R. Ranganath. 2019. “ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission”. arXiv. http://arxiv.org/abs/1904.05342v3.
Harvard
Huang, K., Altosaar, J. and Ranganath, R. (2019) “ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1904.05342v3.
Vancouver
1. Huang K, Altosaar J, Ranganath R (2019) ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission. arXiv

BibTeX

@article{huang2019clinicalbert,
  title = {ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission},
  author = {Huang, Kexin and Altosaar, Jaan and Ranganath, Rajesh},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1904.05342v3},
  eprint = {1904.05342}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF