CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

Ningyu ZhangMosha ChenZhen BiXiaozhuan LiangLei LiXin ShangKangping YinChuanqi TanJian XuFei Huang

article2022ACL239 citations

Introduces the first comprehensive Chinese biomedical language understanding benchmark spanning eight diverse clinical and medical NLP tasks, establishing baseline evaluations across eleven pre-trained models to expose significant performance gaps compared to human capability.

Listen

Artificial intelligence applications are expanding rapidly across healthcare and clinical research, yet most standardized evaluation benchmarks remain centered entirely on English. Because Chinese is spoken by a quarter of the global population and exhibits unique grammatical structures, pervasive colloquial phrasing, and complex domain terminology, English-centric benchmarks cannot effectively validate Chinese medical language models. The lack of standardized datasets has hindered reliable evaluation, clinical implementation, and cross-model comparison for Chinese biomedical text processing.

To address this critical gap, the article introduces the Chinese Biomedical Language Understanding Evaluation (CBLUE) benchmark. The objective is to establish the first comprehensive, open-access evaluation suite and public leaderboard specifically designed to assess and advance artificial intelligence models on Chinese biomedical text understanding.

To construct CBLUE, the researchers collected authentic, anonymized real-world data across eight distinct tasks covering clinical trial criteria, electronic health records, medical forums, textbooks, and search engine logs. Tasks span named entity recognition, information extraction, diagnosis normalization, eligibility criteria classification, and search intent or relevance matching. Domain specialists annotated the data with high inter-rater agreement. The authors then benchmarked eleven leading Chinese pre-trained language models against human baseline performance established by trained non-specialists.

Key findings show that artificial intelligence models still lag significantly behind humans across Chinese medical language understanding. While trained human baselines achieved an overall average score of 77.1%, the top-performing model reached only 70.0%, falling short across all eight evaluated tasks. In complex extraction and normalization tasks, model accuracy was notably low, with top scores ranging from 55.9% to 59.3%. Domain-specific pre-training provided isolated advantages for technical medical terminology, but specialized medical models still underperformed expectations overall. Error analyses revealed that overlapping entity boundaries, syntactic ambiguity, multiple trigger phrases, and informal query phrasing were the primary causes of model failures.

These findings indicate that directly deploying general or existing domain-adapted models into high-stakes clinical workflows poses substantial operational and diagnostic risks. In contrast to English medical benchmarks where top models approach or exceed human parity, Chinese biomedical language models require further structural advancements to reliably interpret specialized clinical terms and everyday patient language.

Stakeholders and developers should utilize CBLUE as a standardized testbed to rigorously assess model capabilities before real-world deployment, while focusing future research on resolving colloquial ambiguities, complex entity structures, and cross-disease domain shifts. Decision-makers should also support the benchmark's expansion into interactive medical dialogues and dynamic evaluations to ensure continuous improvements in patient safety and clinical artificial intelligence reliability.

arXiv: 2106.08087
Cover for CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

Abstract

With the development of biomedical language understanding benchmarks, Artificial Intelligence applications are widely used in the medical field. However, most benchmarks are limited to English, which makes it challenging to replicate many of the successes in English for other languages. To facilitate research in this direction, we collect real-world biomedical data and present the first Chinese Biomedical Language Understanding Evaluation (CBLUE) benchmark: a collection of natural language understanding tasks including named entity recognition, information extraction, clinical diagnosis normalization, and an associated online platform for model evaluation, comparison, and analysis. To establish evaluation on these tasks, we report empirical results with the current 11 pre-trained Chinese models, and experimental results show that state-of-the-art neural models perform far worse than the human ceiling†. Our benchmark is released at https://tianchi.aliyun.com/dataset/dataDetail?dataId=95414&lang=en-us.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 CBLUE Overview
  • 3.1 Design Principle
  • 3.2 Tasks
  • 3.3 Data Collection
  • 3.4 Annotation
  • 3.5 Characteristics
  • 3.6 Leaderboard
  • 3.7 Distribution and Maintenance
  • 3.8 Reproducibility
  • 4 Experiments
  • 4.1 Benchmark Results
  • 4.2 Human Performance
  • 4.3 Case studies
  • 5 Conclusion
  • Acknowledgments
  • Ethical Considerations
  • References
  • A Broader Impact
  • B Negative Impact
  • C Limitations
  • D CBLUE Background
  • E Detailed Task Introduction
  • E.1 Chinese Medical Named Entity Recognition Dataset (CMeEE)
  • E.2 Chinese Medical Information Extraction Dataset (CMeIE)
  • E.3 CHIP - Clinical Diagnosis Normalization Dataset (CHIP-CDN)
  • E.4 Clinical Trial Criterion Dataset (CHIP-CTC)
  • E.5 Semantic Textual Similarity Dataset (CHIP-STS)
  • E.6 KUAKE-Query Intent Classification Dataset (KUAKE-QIC)
  • E.7 KUAKE- Query Title Relevance Dataset (KUAKE-QTR)
  • E.8 KUAKE - Query Query Relevance Dataset (KUAKE-QQR)
  • F Experiments Details
  • G Error Analysis for Other Tasks
  • Contributions

Knowls

  1. Knowl 1 — CBLUE Benchmark Tasks and Dataset Specifications

    data/table

    The Chinese Biomedical Language Understanding Evaluation (CBLUE) benchmark comprises eight natural language understanding tasks across token-level, sequence-level, and sequence-pair understanding in the Chinese biomedical domain:

    • CMeEE (Chinese Medical Named Entity Recognition): Sequence labeling to identify and classify medical entity spans into 9 categories: disease (dis), clinical manifestations (sym), drugs (dru), medical equipment (equ), medical procedures (pro), body (bod), medical examination items (ite), microorganisms (mic), and department (dep).
    • CMeIE (Chinese Medical Information Extraction): Relation extraction requiring the extraction of Subject-Predicate-Object (SPO) triples across 53 predefined schema relations (10 genus-level relations and 43 fine-grained sub-relations).
    • CHIP-CDN (Clinical Diagnosis Normalization): Mapping non-standard colloquial or clinical diagnostic terms from electronic medical records to standard International Classification of Diseases (ICD-10, Beijing Clinical Edition v601) terminology codes and phrases.
    • CHIP-STS (Semantic Textual Similarity): Binary sentence-pair classification to determine semantic equivalence between patient question pairs across 5 disease domains (diabetes, hypertension, hepatitis, AIDS, breast cancer) evaluated in a non-i.i.d. transfer learning setup where training and test disease types are disjoint.
    • CHIP-CTC (Clinical Trial Criterion Classification): Multi-class sentence classification categorizing eligibility criteria statements from the Chinese Clinical Trial Registry (ChiCTR) into 44 pre-defined semantic classes.
    • KUAKE-QIC (Query Intent Classification): Multi-class classification assigning Chinese search engine health queries to 11 intent categories (diagnosis, etiology analysis, treatment plan, medical advice, test result analysis, disease description, consequence prediction, precautions, intended effects, treatment fees, and others).
    • KUAKE-QTR (Query-Document Title Relevance): Ranking relevance between a search query and a document title across 4 ordinal levels (00 = unrelated, 11 = poorly related, 22 = related, 33 = strongly related).
    • KUAKE-QQR (Query-Query Relevance): Assessing semantic equivalence and relevance between pairs of user queries across 3 ordinal levels (00 = unrelated/poorly related, 11 = related, 22 = strongly related).
    Dataset Task Train Dev Test Metric
    CMeEE Named Entity Recognition 15,000 5,000 3,000 Micro F1F_1
    CMeIE Information Extraction 14,339 3,585 4,482 Micro F1F_1
    CHIP-CDN Diagnosis Normalization 6,000 2,000 10,192 Micro F1F_1
    CHIP-STS Sentence Similarity (Non-i.i.d.) 16,000 4,000 10,000 Macro F1F_1
    CHIP-CTC Eligibility Criteria Classification 22,962 7,682 10,000 Macro F1F_1
    KUAKE-QIC Intent Classification 6,931 1,955 1,994 Accuracy
    KUAKE-QTR Query-Document Relevance 24,174 2,913 5,465 Accuracy
    KUAKE-QQR Query-Query Relevance 15,000 1,600 1,596 Accuracy
  2. Knowl 2 — Baseline Pre-Trained Model Performance on CBLUE Benchmark

    data/table

    A systematic evaluation of 11 Chinese pre-trained language models and a human baseline across all eight CBLUE benchmark tasks demonstrates that general large pre-trained models outperform smaller architectures, while domain-specific pre-trained medical models do not consistently outperform general-purpose models.

    Model CMeEE CMeIE CDN CTC STS QIC QTR QQR Avg.
    BERT-base 62.1 54.0 55.4 69.2 83.0 84.3 60.0 84.7 69.1
    BERT-wwm-ext-base 61.7 54.0 55.4 70.1 83.9 84.5 60.9 84.4 69.4
    RoBERTa-large 62.1 54.4 56.5 70.9 84.7 84.2 60.9 82.9 69.6
    RoBERTa-wwm-ext-base 62.4 53.7 56.4 69.4 83.7 85.5 60.3 82.7 69.3
    RoBERTa-wwm-ext-large 61.8 55.9 55.7 69.0 85.2 85.3 62.8 84.4 70.0
    ALBERT-tiny 50.5 35.9 50.2 61.0 79.7 75.8 55.5 79.8 61.1
    ALBERT-xxlarge 61.8 47.6 37.5 66.9 84.8 84.8 62.2 83.1 66.1
    ZEN 61.0 50.1 57.8 68.6 83.5 83.2 60.3 83.0 68.4
    MacBERT-base 60.7 53.2 57.7 67.7 84.4 84.9 59.7 84.0 69.0
    MacBERT-large 62.4 51.6 59.3 68.6 85.6 82.7 62.9 83.5 69.6
    PCL-MedBERT 60.6 49.1 55.8 67.8 83.8 84.3 59.3 82.5 67.9
    Human 67.0 66.0 65.0 78.0 93.0 88.0 71.0 89.0 77.1

    Key observations include:

    • RoBERTa-wwm-ext-large achieved the highest average baseline score (70.070.0), followed by MacBERT-large (69.669.6) and RoBERTa-large (69.669.6).
    • The medical-specific language model PCL-MedBERT scored 67.967.9 on average, falling behind several general-domain pre-trained models.
    • Whole-word masking (wwm) does not consistently improve performance on sentence classification and relevance tasks (e.g., CHIP-CTC, KUAKE-QIC, KUAKE-QTR, KUAKE-QQR).
    • ALBERT-tiny achieved competitive scores on several tasks (e.g., 50.250.2 on CHIP-CDN and 79.779.7 on CHIP-STS), whereas ALBERT-xxlarge experienced catastrophic degradation on CHIP-CDN (37.537.5).
    • All models perform substantially below human performance across every single task, exhibiting an average gap of 7.17.1 points between the best model and human majority vote.
  3. Knowl 3 — Human Performance Evaluation Protocol and Gap Against Neural Baselines

    empirical result

    Human ceiling performance on the CBLUE benchmark was evaluated using trained amateur annotators without medical backgrounds. Annotators completed training on development set instances with validation and error correction against gold expert annotations until mastering task guidelines, after which they annotated test instances.

    Evaluator CMeEE CMeIE CDN CTC STS QIC QTR QQR Avg.
    Annotator 1 69.0 62.0 60.0 73.0 94.0 87.0 75.0 80.0 75.0
    Annotator 2 62.0 65.0 69.0 75.0 93.0 91.0 62.0 88.0 75.6
    Annotator 3 69.0 67.0 62.0 80.0 88.0 83.0 71.0 90.0 76.2
    Human Average 66.7 64.7 63.7 76.0 91.7 87.0 69.3 86.0 75.6
    Human Majority Vote 67.0 66.0 65.0 78.0 93.0 88.0 71.0 89.0 77.1
    Best Baseline Model 62.4 55.9 59.3 70.9 85.6 85.5 62.9 84.7 70.0

    The majority vote of amateur annotators exceeds the highest-performing neural baseline across all eight tasks, with the largest performance deficits observed in information extraction (CMeIE: +10.1+10.1 points), query-document relevance (KUAKE-QTR: +8.1+8.1 points), non-i.i.d. sentence similarity (CHIP-STS: +7.4+7.4 points), and clinical criteria classification (CHIP-CTC: +7.1+7.1 points).

  4. Knowl 4 — Evaluation Metric for Clinical Diagnosis Normalization (CHIP-CDN)

    equation

    In the CHIP-CDN diagnosis normalization task, performance is measured by comparing the predicted set of (original clinical term, normalized standard phrase) pairs against the gold-standard reference pairs.

    Let m∈Nm \in \mathbb{N} be the total number of gold pairs in the evaluation set, n∈Nn \in \mathbb{N} be the total number of pairs predicted by the model, and k∈Nk \in \mathbb{N} (where 0≤k≤min⁡(m,n)0 \le k \le \min(m, n)) be the number of correctly predicted pairs matching the ground truth. Precision PP, Recall RR, and Micro-F1F_1 are defined as:

    P=knP = \frac{k}{n}

    R=kmR = \frac{k}{m}

    F1=2⋅P⋅RP+R=2km+nF_1 = \frac{2 \cdot P \cdot R}{P + R} = \frac{2k}{m + n}

  5. Knowl 5 — Macro-Averaged F1 Metric for Multi-Class Evaluation in CBLUE

    equation

    For multi-class sentence classification benchmarks in CBLUE, specifically CHIP-CTC (comprising n=44n = 44 semantic eligibility criteria categories) and CHIP-STS (binary sentence pair classification under disease distribution shift), performance is evaluated via the unweighted macro-averaged F1F_1 score across all classes.

    For nn classes C1,C2,…,CnC_1, C_2, \dots, C_n, let PiP_i and RiR_i denote the precision and recall for class CiC_i:

    Pi=number of instances correctly predicted as class Citotal number of instances predicted as class CiP_i = \frac{\text{number of instances correctly predicted as class } C_i}{\text{total number of instances predicted as class } C_i}

    Ri=number of instances correctly predicted as class Citotal number of true instances in class CiR_i = \frac{\text{number of instances correctly predicted as class } C_i}{\text{total number of true instances in class } C_i}

    The macro-averaged metric Average-F1\text{Average-}F_1 is computed as:

    Average-F1=1n∑i=1n2⋅Pi⋅RiPi+Ri\text{Average-}F_1 = \frac{1}{n} \sum_{i=1}^{n} \frac{2 \cdot P_i \cdot R_i}{P_i + R_i}

  6. Knowl 6 — Error Distribution and Linguistic Challenges in Chinese Medical NER (CMeEE)

    empirical result

    Error analysis on the Chinese Medical Named Entity Recognition (CMeEE) dataset reveals seven primary error categories responsible for baseline model failures:

    • Ambiguity (24%24\%): Contextual ambiguity where entities appear in similar phrasing but differ in intended semantic meaning.
    • Entity Overlap (22%22\%): Overlapping and nested entity spans (e.g., overlapping disease names, body locations, and clinical procedures), confusing standard sequential BIO/BMES token classification models.
    • Need Domain Knowledge (18%18\%): Complex biomedical terminologies (e.g., genetic abnormalities such as deletions, translocations, and inversions, or specific antibody formulations) requiring specialized medical semantics.
    • Annotation Error (12%12\%): Inconsistent or incorrect ground-truth annotations in benchmark instances.
    • Wrong Entity Boundary (9%9\%): Model predictions that identify the entity concept but shift, truncate, or over-extend span boundaries.
    • Others (8%8\%): Input sequences with long sentence length or out-of-vocabulary rare words.
    • Need Syntactic Knowledge (7%7\%): Complex syntactic structures and multi-clause Chinese grammatical patterns obscuring entity scope.
  7. Knowl 7 — Error Distribution and Linguistic Challenges in Medical Query Intent Classification (KUAKE-QIC)

    empirical result

    Error analysis on the search engine query intent classification dataset (KUAKE-QIC) categorizes prediction failures into seven distinct sources:

    • Multiple Triggers (26.2%26.2\%): The presence of multiple conflicting indicative keyword cues in a single query representing different intents (e.g., simultaneously asking about disease causes and requesting treatment plans).
    • Colloquialism (22.5%22.5\%): Informal spoken Chinese, severe phrasing simplifications, non-standard abbreviations, and irregular syntax characteristic of consumer search engine queries.
    • Ambiguity (13.7%13.7\%): Query expressions with multiple plausible intent interpretations in the absence of broader conversational or clinical context.
    • Rare Words (11.2%11.2\%): Low-frequency medical expressions and rare illness descriptions.
    • Annotation Error (10.0%10.0\%): Mislabeled intent categories in gold labels.
    • Need Domain Knowledge (8.7%8.7\%): Specialized laboratory and diagnostic terminology (e.g., interpreting neutrophil and lymphocyte ratio variations) requiring medical domain reasoning to distinguish test result analysis from general illness diagnosis.
    • Irrelevant Description (7.5%7.5\%): Incidental narrative descriptions in user search queries misleading classification attention.
  8. Knowl 8 — Two-Stage Retrieval and Ranking Architecture for CHIP-CDN

    model/method

    The baseline system for the CHIP-CDN clinical diagnosis normalization task models terminology mapping through a two-stage recall and ranking framework:

    1. Candidate Recall Stage: For each query diagnostic term from electronic medical records, an initial candidate retrieval step selects the top k=200k = 200 standard ICD-10 candidate terms (recall_k = 200).
    2. Ranking and Multi-Label Normalization Stage:
      • A pre-trained language model sequence classifier encodes the concatenated pair consisting of the original query term and a candidate standard phrase, predicting a pairwise relevance score. Ranking models are trained with 10 negative samples sampled per query instance (num_negative_sample = 10).
      • A concurrent sequence classification model takes the original term as input to predict the total integer count of corresponding standard phrases (accommodating composite diagnoses mapping to multiple ICD-10 standard phrases separated by delimiter tokens ##).
      • The top candidate phrases selected by the ranking model matching the predicted count are returned as the final standardized outputs.
  9. Knowl 9 — Design Characteristics and Benchmark Distinctions of CBLUE

    definition

    CBLUE differs from general Chinese NLP benchmarks (such as CLUE) and English biomedical NLP benchmarks (such as BLURB) along four structural dimensions:

    • Diverse Data Sources: Aggregates Chinese medical texts across clinical trial eligibility criteria (ChiCTR), hospital electronic health records (EHR final diagnoses), authoritative medical textbooks (pediatrics and clinical practice), online patient consultation forums, and search engine query logs (Alibaba Quark).
    • Real-World Zipfian Distribution: Retains authentic long-tailed frequency distributions without artificial upsampling or downsampling, and incorporates hierarchical label schemas with coarse- and fine-grained relation structures in information extraction.
    • Explicit Transfer Learning (Non-i.i.d.) Setting: In CHIP-STS, question pairs in the evaluation set represent disease categories completely disjoint from the disease categories present in the training set to directly test cross-disease transferability.
    • Privacy and Quality Control: All data undergoes utility-preserving de-identification and institutional review board (IRB) approval to remove protected health information (PHI) and personally identifiable information (PII). Annotations are performed and validated by medical professionals from Class A tertiary hospitals, achieving high inter-annotator agreement (Fleiss' κ≈0.9\kappa \approx 0.9).
  10. Knowl 10 — Limitations of the CBLUE Benchmark

    limitation

    The authors identify three primary limitations of the CBLUE benchmark:

    1. Incomplete Task Coverage: The benchmark focuses on natural language understanding (classification, sequence labeling, retrieval ranking, and normalization), omitting generative medical tasks such as multi-turn medical dialogue generation and end-to-end automated clinical diagnostic generation.
    2. Static Benchmark Vulnerability: As a static evaluation dataset, models may exploit superficial dataset cues or spurious correlations, achieving competitive benchmark scores while remaining brittle to out-of-distribution real-world clinical shifts and adversarial challenge cases.
    3. Residual Label Noise: Despite multi-stage expert curation, a small percentage of incorrect annotations persists across the datasets (accounting for 10%10\% to 12%12\% of error cases in analyzed tasks), which could lead to suboptimal real-world model deployment decisions if models are selected purely based on benchmark leaderboard rankings.

Coverage note — Omitted fine-tuning hyperparameter search tables for individual baseline runs (Tables 15-26) and specific raw Chinese text examples (Tables 5-14, 27-32) as they represent secondary implementation configurations and qualitative samples rather than standalone conceptual or empirical contributions.

References

  1. 1.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. Scibert: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 3613–3618. Association for Computational Linguistics.
  2. 2.Alexis Conneau and Douwe Kiela. 2018. Senteval: An evaluation toolkit for universal sentence representations. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation, LREC 2018, Miyazaki, Japan, May 7-12, 2018. European Language Resources Association (ELRA).
  3. 3.Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Shijin Wang, and Guoping Hu. 2020. Revisiting pretrained models for chinese natural language processing. arXiv preprint arXiv:2004.13922.
  4. 4.Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Ziqing Yang, Shijin Wang, and Guoping Hu. 2019. Pre-training with whole word masking for chinese bert. arXiv preprint arXiv:1906.08101.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT.
  6. 6.Shizhe Diao, Jiaxin Bai, Yan Song, Tong Zhang, and Yonggang Wang. 2019. Zen: pre-training chinese text encoder enhanced by n-gram representations. arXiv preprint arXiv:1911.00720.
  7. 7.Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin, 76(5):378.
  8. 8.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna M. Wallach, Hal Daumé III, and Kate Crawford. 2018. Datasheets for datasets. CoRR, abs/1803.09010.
  9. 9.Pieter Gijsbers, Erin LeDell, Janek Thomas, Sébastien Poirier, Bernd Bischl, and Joaquin Vanschoren. 2019. An open source automl benchmark. arXiv preprint arXiv:1907.00909.
  10. 10.Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2020. Domain-specific language model pretraining for biomedical natural language processing. CoRR, abs/2007.15779.
  11. 11.T. Guan, H. Zan, X. Zhou, H. Xu, and K Zhang. 2020. CMeIE: Construction and Evaluation of Chinese Medical Information Extraction Dataset. Natural Language Processing and Chinese Computing, 9th CCF International Conference, NLPCC 2020, Zhengzhou, China, October 14–18, 2020, Proceedings, Part I.
  12. 12.Zan Hongying, Li Wenxin, Zhang Kunli, Ye Yajuan, Chang Baobao, and Sui Zhifang. 2020. Building a pediatric medical corpus: Word segmentation and named entity annotation. In Workshop on Chinese Lexical Semantics, pages 652–664.
  13. 13.Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. 2019. Pubmedqa: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 2567–2577. Association for Computational Linguistics.
  14. 14.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942.
  15. 15.Hyukki Lee, Soohyung Kim, Jong Wook Kim, and Yon Dohn Chung. 2017. Utility-preserving anonymization for health data publishing. BMC Medical Informatics Decis. Mak., 17(1):104:1–104:12.
  16. 16.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240.
  17. 17.Patrick Lewis, Myle Ott, Jingfei Du, and Veselin Stoyanov. 2020. Pretrained language models for biomedical and clinical tasks: Understanding and extending the state-of-the-art. In Proceedings of the 3rd Clinical Natural Language Processing Workshop, pages 146–157, Online. Association for Computational Linguistics.
  18. 18.J. Li, Yueping Sun, Robin J. Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, A. P. Davis, C. Mattingly, Thomas C. Wiegers, and Zhiyong Lu. 2016. Biocreative v cdr task corpus: a resource for chemical disease relation extraction. Database: The Journal of Biological Databases and Curation, 2016.
  19. 19.Shuai Lin, Pan Zhou, Xiaodan Liang, Jianheng Tang, Ruihui Zhao, Ziliang Chen, and Liang Lin. 2020. Graph-evolving meta-learning for low-resource medical dialogue generation. CoRR, abs/2012.11988.
  20. 20.Wenge Liu, Jianheng Tang, Jinghui Qin, Lin Xu, Zhen Li, and Xiaodan Liang. 2020. Meddg: A large-scale medical consultation dataset for building medical dialogue system. CoRR, abs/2010.07497.
  21. 21.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  22. 22.Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018. The natural language decathlon: Multitask learning as question answering. CoRR, abs/1806.08730.
  23. 23.Dimitris Pappas, Ion Androutsopoulos, and Haris Papageorgiou. 2018. Bioread: A new dataset for biomedical reading comprehension. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation, LREC 2018, Miyazaki, Japan, May 7-12, 2018. European Language Resources Association (ELRA).
  24. 24.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems, pages 8024–8035.
  25. 25.Tatiana Shavrina, Alena Fenogenova, Anton A. Emelyanov, Denis Shevelev, Ekaterina Artemova, Valentin Malykh, Vladislav Mikhailov, Maria Tikhonova, Andrey Chertok, and Andrey Evlampiev. 2020. Russiansuperglue: A russian language understanding evaluation benchmark. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 4717–4726. Association for Computational Linguistics.
  26. 26.Xiaoming Shen and Yonghao Gui. 2013. Clinical Pediatrics 2nd edn. People’s Medical Publishing House.
  27. 27.George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R. Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, Yannis Almirantis, John Pavlopoulos, Nicolas Baskiotis, Patrick Gallinari, Thierry Artières, Axel-Cyrille Ngonga Ngomo, Norman Heino, Éric Gaussier, Liliana Barrio-Alvers, Michael Schroeder, Ion Androutsopoulos, and Georgios Paliouras. 2015. An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition. BMC Bioinform., 16:138:1–138:28.
  28. 28.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a. Superglue: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 3261–3275.
  29. 29.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019b. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  30. 30.Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Darrin Eide, Kathryn Funk, Rodney Kinney, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuansan Wang, Chris Wilhelm, Boya Xie, Douglas Raymond, Daniel S. Weld, Oren Etzioni, and Sebastian Kohlmeier. 2020. CORD-19: the covid-19 open research dataset. CoRR, abs/2004.10706.
  31. 31.Weiping Wang, Kun Song, and Liwen Chang. 2018. Pediatrics 9th edn. People’s Medical Publishing House.
  32. 32.Zhongyu Wei, Qianlong Liu, Baolin Peng, Huaixiao Tou, Ting Chen, Xuanjing Huang, Kam-Fai Wong, and Xiangying Dai. 2018. Task-oriented dialogue system for automatic diagnosis. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 2: Short Papers, pages 201–207. Association for Computational Linguistics.
  33. 33.Y. Wu, Ruibang Luo, H. Leung, H. Ting, and T. Lam. 2019. Renet: A deep learning approach for extracting gene-disease associations from literature. In RECOMB.
  34. 34.Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, and Zhenzhong Lan. 2020. CLUE: A chinese language understanding evaluation benchmark. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pages 4762–4772. International Committee on Computational Linguistics.
  35. 35.Kun-Hsing Yu, Andrew L Beam, and Isaac S Kohane. 2018. Artificial intelligence in healthcare. Nature biomedical engineering, 2(10):719–731.
  36. 36.Guangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang, Sicheng Wang, Ruisi Zhang, Meng Zhou, Jiaqi Zeng, Xiangyu Dong, Ruoyu Zhang, Hongchao Fang, Penghui Zhu, Shu Chen, and Pengtao Xie. 2020. Meddialog: Large-scale medical dialogue datasets. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 9241–9250. Association for Computational Linguistics.
  37. 37.Hui Zong, Jinxuan Yang, Zeyu Zhang, Zuofeng Li, and Xiaoyan Zhang. 2021. Semantic categorization of chinese eligibility criteria in clinical trials using machine learning methods. BMC Medical Informatics Decis. Mak., 21(1):128.

Citation

MLA
Zhang, N., et al. “CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7888–915, https://doi.org/10.18653/v1/2022.acl-long.544.
APA
Zhang, N., Chen, M., Bi, Z., Liang, X., Li, L., Shang, X., Yin, K., Tan, C., Xu, J., Huang, F., Si, L., Ni, Y., Xie, G., Sui, Z., (常宝宝), B. C., Zong, H., Yuan, Z., Li, L., Yan, J., … Chen, Q. (2022). CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7888–7915. https://doi.org/10.18653/v1/2022.acl-long.544
Chicago
Zhang, N., M. Chen, Z. Bi, et al. 2022. “CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7888–7915. https://doi.org/10.18653/v1/2022.acl-long.544.
Harvard
Zhang, N. et al. (2022) “CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7888–7915. Available at: https://doi.org/10.18653/v1/2022.acl-long.544.
Vancouver
1. Zhang N, Chen M, Bi Z, et al (2022) CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7888–7915

BibTeX

@inproceedings{zhang-etal-2022-cblue,
    title = "{CBLUE}: A {C}hinese Biomedical Language Understanding Evaluation Benchmark",
    author = "Zhang, Ningyu  and
      Chen, Mosha  and
      Bi, Zhen  and
      Liang, Xiaozhuan  and
      Li, Lei  and
      Shang, Xin  and
      Yin, Kangping  and
      Tan, Chuanqi  and
      Xu, Jian  and
      Huang, Fei  and
      Si, Luo  and
      Ni, Yuan  and
      Xie, Guotong  and
      Sui, Zhifang  and
      Chang, Baobao  and
      Zong, Hui  and
      Yuan, Zheng  and
      Li, Linfeng  and
      Yan, Jun  and
      Zan, Hongying  and
      Zhang, Kunli  and
      Tang, Buzhou  and
      Chen, Qingcai",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.544/",
    doi = "10.18653/v1/2022.acl-long.544",
    pages = "7888--7915"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/