CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
Ningyu ZhangMosha ChenZhen BiXiaozhuan LiangLei LiXin ShangKangping YinChuanqi TanJian XuFei Huang
Introduces the first comprehensive Chinese biomedical language understanding benchmark spanning eight diverse clinical and medical NLP tasks, establishing baseline evaluations across eleven pre-trained models to expose significant performance gaps compared to human capability.
Artificial intelligence applications are expanding rapidly across healthcare and clinical research, yet most standardized evaluation benchmarks remain centered entirely on English. Because Chinese is spoken by a quarter of the global population and exhibits unique grammatical structures, pervasive colloquial phrasing, and complex domain terminology, English-centric benchmarks cannot effectively validate Chinese medical language models. The lack of standardized datasets has hindered reliable evaluation, clinical implementation, and cross-model comparison for Chinese biomedical text processing.
To address this critical gap, the article introduces the Chinese Biomedical Language Understanding Evaluation (CBLUE) benchmark. The objective is to establish the first comprehensive, open-access evaluation suite and public leaderboard specifically designed to assess and advance artificial intelligence models on Chinese biomedical text understanding.
To construct CBLUE, the researchers collected authentic, anonymized real-world data across eight distinct tasks covering clinical trial criteria, electronic health records, medical forums, textbooks, and search engine logs. Tasks span named entity recognition, information extraction, diagnosis normalization, eligibility criteria classification, and search intent or relevance matching. Domain specialists annotated the data with high inter-rater agreement. The authors then benchmarked eleven leading Chinese pre-trained language models against human baseline performance established by trained non-specialists.
Key findings show that artificial intelligence models still lag significantly behind humans across Chinese medical language understanding. While trained human baselines achieved an overall average score of 77.1%, the top-performing model reached only 70.0%, falling short across all eight evaluated tasks. In complex extraction and normalization tasks, model accuracy was notably low, with top scores ranging from 55.9% to 59.3%. Domain-specific pre-training provided isolated advantages for technical medical terminology, but specialized medical models still underperformed expectations overall. Error analyses revealed that overlapping entity boundaries, syntactic ambiguity, multiple trigger phrases, and informal query phrasing were the primary causes of model failures.
These findings indicate that directly deploying general or existing domain-adapted models into high-stakes clinical workflows poses substantial operational and diagnostic risks. In contrast to English medical benchmarks where top models approach or exceed human parity, Chinese biomedical language models require further structural advancements to reliably interpret specialized clinical terms and everyday patient language.
Stakeholders and developers should utilize CBLUE as a standardized testbed to rigorously assess model capabilities before real-world deployment, while focusing future research on resolving colloquial ambiguities, complex entity structures, and cross-disease domain shifts. Decision-makers should also support the benchmark's expansion into interactive medical dialogues and dynamic evaluations to ensure continuous improvements in patient safety and clinical artificial intelligence reliability.
- Paper: GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, Alex Wang et al. (2018). It introduces the foundational multi-task benchmark and leaderboard paradigm for natural language understanding upon which Chinese language benchmarks like CBLUE are directly structured.
- Paper: Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing, Yu Gu et al. (2020). It presents BLURB, establishing the standardized biomedical language understanding benchmark design and evaluation protocol that CBLUE adapts for Chinese clinical and biomedical NLP.
- Paper: What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams, Di Jin et al. (2020). It introduces MedQA, providing foundational methodology and multi-lingual dataset resources—including Chinese medical board exam questions—essential to understanding Chinese clinical QA evaluation.
- Paper: BioBERT: a pre-trained biomedical language representation model for biomedical text mining, Jinhyuk Lee et al. (2019). It establishes domain-specific pre-training for biomedical text representations, motivating the baseline evaluation of specialized biomedical language models in CBLUE.
- Paper: SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems, Alex Wang et al. (2019). It extends multi-task benchmark methodologies to more demanding NLU reasoning tasks and diagnostic setups that inform CBLUE's task curation.
- Paper: PubMedQA: A Dataset for Biomedical Research Question Answering, Qiao Jin et al. (2019). It illustrates biomedical natural language understanding and reasoning benchmarks built from clinical abstracts, exemplifying task design in biomedical NLP.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). It introduces the pre-trained Transformer architecture and fine-tuning framework evaluated across all baseline models in CBLUE.
- Paper: Large language models encode clinical knowledge, Karan Singhal et al. (2022). It advances clinical NLP evaluation beyond classical NLU tasks to large language models and generative clinical QA, building upon the need for robust medical benchmarks shown in CBLUE.
- Paper: MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding, Yuxin Zuo et al. (2025). It extends medical reasoning benchmarks to expert-level multimodal diagnostics and specialized reasoning challenges where earlier medical NLU benchmarks plateau.
- Paper: A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery, Yu Zhang et al. (2024). It synthesizes pre-training and evaluation paradigms for scientific and medical language models, contextualizing the benchmark results seen in domain-specific evaluations like CBLUE.
- Paper: DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains, Yanis Labrak et al. (2023). It develops and evaluates non-English clinical and biomedical language models (for French), paralleling and continuing CBLUE's mission to address the dominance of English-centric medical NLP.
- Paper: AlignBench: Benchmarking Chinese Alignment of Large Language Models, Xiao Liu et al. (2024). It broadens Chinese language model evaluation from specialized NLU benchmarks to comprehensive instruction-following and safety alignment in Chinese.
- Paper: ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools, Team Glm Aohan Zeng et al. (2024). It develops advanced bilingual foundation models capable of handling broader Chinese domain understanding, directly extending the pre-trained Chinese model landscape assessed in CBLUE.
