ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences
Yuanhe TianRuyi GanYan SongJiaxing ZhangYongdong Zhang
Develops ChiMed-GPT, an open-source Chinese medical large language model trained through continual pre-training, supervised fine-tuning, and rejection-sampling-based reinforcement learning from human feedback to generate accurate, safe clinical responses with an extended context window.
Growing healthcare demands and a constrained medical workforce have created a major strain on medical services. Natural language processing and large language models offer strong potential to automate routine healthcare communication and assist clinical workflows. However, general-purpose models often lack specialized medical knowledge, while existing medical language models predominantly rely solely on supervised fine-tuning. This limited training approach restricts their ability to master deep medical knowledge, align with human preferences, handle longer clinical contexts, and mitigate harmful biases.
The article demonstrates the development and evaluation of CHIMED-GPT, a 13-billion-parameter open-source benchmark language model built specifically for Chinese medical text processing. The main objective was to evaluate whether executing a comprehensive training regime—incorporating continuous pre-training, supervised fine-tuning, and reinforcement learning from human feedback—along with an expanded context window, delivers superior diagnostic accuracy, conversational quality, and safety compared to existing general and medical models.
To achieve this, the authors built CHIMED-GPT upon the Ziya-13B-v2 foundation model and extended its context capacity to 4,096 tokens, doubling the 2,048-token limit typical of existing open-source medical models. The team trained the model across three distinct stages using diverse data sources: continuous pre-training on 214 million tokens from medical encyclopedias and textbooks; supervised fine-tuning across over 1.2 million doctor-patient interaction and dialogue pairs alongside 100,000 safety prompts; and reinforcement learning via rejection sampling fine-tuning. For the feedback stage, the authors constructed a fine-grained reward model by ranking outputs across clinical doctors, commercial models, and existing domain models. The resulting model was evaluated across information extraction, multiple-choice and open-ended question answering, and multi-turn dialogue generation, alongside standardized bias assessments measuring public and clinical attitudes toward mental illness.
The findings show that CHIMED-GPT consistently outperforms both general and medical baseline models. In medical information extraction, the model achieved entity recognition F1 scores of 40.82 and 41.04 across two benchmark datasets, substantially outperforming existing Chinese medical models (which scored between 23.80 and 30.90) and matching GPT-4. In multi-choice medical exams, CHIMED-GPT reached 68.29% accuracy on the C-Eval benchmark, surpassing all open-source models by 19 to 32 percentage points and closely trailing GPT-4 (71.29%). In multi-turn dialogue generation, the model achieved a ROUGE-L score of 42.16, roughly doubling the performance of GPT-4 (17.14) and other Chinese medical models (under 15.0). Human evaluations confirmed that CHIMED-GPT generated significantly more fluent, complete, and precise clinical advice, while bias tests demonstrated the lowest prejudice scores among all tested models when responding to sensitive mental health prompts.
These results demonstrate that a full training lifecycle combining domain-specific pre-training with preference-aligned reinforcement learning is essential for deploying language models in high-stakes fields. For healthcare organizations and technology leaders, specialized models trained in this manner reduce clinical risks, minimize misinformation, and provide safer interactions than general-purpose commercial interfaces. The model's expanded context capacity also allows it to digest lengthy medical records and patient histories without losing coherence.
Organizations developing or deploying medical artificial intelligence should adopt end-to-end training pipelines that include human preference alignment rather than relying solely on supervised fine-tuning. Before integrating such models into patient-facing or clinical decision-support systems, stakeholders should conduct pilot studies in controlled clinical environments to evaluate real-time diagnostic performance and workflow integration.
The article's primary limitations include reliance on existing benchmark datasets and simulated dialogue interactions, with human evaluations restricted to sample sizes of 50 instances per task. While confidence in the model's textual processing and safety alignment is high based on empirical benchmarks, readers should exercise appropriate caution regarding real-world clinical deployment until broader clinical trials and regulatory reviews are conducted.
- Paper: Large language models encode clinical knowledge, Karan Singhal et al. (2022). Its Med-PaLM experiments establish an earlier medical-LLM baseline and clinician-oriented alignment approach that clarify what ChiMed-GPT’s fuller training pipeline seeks to improve.
- Paper: CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark, Ningyu Zhang et al. (2022). CBLUE’s Chinese biomedical tasks provide the benchmark context for understanding ChiMed-GPT’s evaluation of Chinese medical information extraction.
- Paper: Rationale-Guided Retrieval Augmented Generation for Medical Question Answering, Jiwoong Sohn et al. (2025). RAG2 carries forward the medical-reliability challenge by adding retrieval and confidence-based evidence filtering to improve medical question answering without retraining.
