Code Synonyms Do Matter: Multiple Synonyms Matching Network for Automatic ICD Coding
Zheng YuanChuanqi TanSongfang Huang
Proposes a multiple synonyms matching network that incorporates UMLS terminology via a specialized attention mechanism to capture varied clinical expressions in electronic medical records, achieving state-of-the-art ICD coding performance on MIMIC-III.
Assigning standard disease classification codes to electronic medical records is essential for clinical decision support, patient tracking, and healthcare billing. However, manual coding by specialized personnel is labor-intensive, costly, and prone to human error. Existing automated systems often struggle because clinical notes contain varied medical jargon, abbreviations, and shorthand that differ markedly from standard official disease descriptions.
The article develops and evaluates a deep learning framework called the Multiple Synonyms Matching Network to demonstrate whether incorporating diverse medical synonyms improves automated disease code assignment. By linking official diagnostic codes to a comprehensive biomedical terminology repository, the method incorporates alternate phrasing directly into the classification process.
The framework processes medical discharge notes and synonymous terms using recurrent neural network encoders. It applies a multi-synonyms attention mechanism that uses each synonym to identify relevant snippets across the text, then evaluates similarities to assign appropriate diagnostic codes without relying on code-specific training parameters for rare conditions. The approach was tested on the benchmark MIMIC-III clinical database across both a full set of thousands of diagnostic codes and a focused subset of the fifty most frequent codes.
The evaluation produced several key findings. Incorporating multiple synonyms outperformed existing state-of-the-art models across standard evaluation metrics. On the full code dataset, the model improved the area under the ROC curve to 95.0% (a 2.0 percentage point gain) and achieved top precision scores of 75.2% and 59.9% across top-8 and top-15 predictions. On the top-50 code dataset, the model improved overall precision and balance, raising the macro F1 score by 1.7 percentage points to 68.3%. Analysis confirmed that performance steadily improved when expanding from a single standard description up to four or eight synonyms per code, aligning with the natural synonym density in biomedical databases.
These findings indicate that addressing linguistic variation directly in automated medical coding systems substantially enhances classification accuracy. For healthcare organizations, adopting synonym-aware models can reduce manual review workloads, lower administrative billing costs, and minimize revenue risks from coding errors. The system also mitigates data scarcity issues for rare diseases by using code-independent similarity scoring.
Healthcare technology leaders should consider integrating synonym enrichment into clinical documentation and automated coding pipelines. When implementing such models, organizations must balance computing resources, as memory usage scales linearly with the number of synonyms utilized. While the framework demonstrates strong benchmark performance, potential adopters should exercise caution regarding performance volatility on extremely rare disease codes in the long tail and should validate the system on diverse institutional datasets before full-scale deployment.
- Paper: Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis, Benjamin Shickel et al. (2017). This survey establishes how deep learning is used to analyze EHR data, including clinical text and coding-related tasks, framing the source’s automatic coding problem.
- Paper: Publicly Available Clinical BERT Embeddings, Emily Alsentzer et al. (2019). Its clinical-BERT models and MIMIC-III evaluations provide useful foundations for understanding the source’s clinical-text representations and dataset setting.
- Paper: BERTMap: A BERT-Based Ontology Alignment System, Yuan He et al. (2022). Because it uses ontology labels and synonyms to align biomedical concepts, this work clarifies the knowledge-alignment strategy underlying the source’s code-synonym approach.
No sufficiently relevant recommendations were found.
