ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences

Yuanhe TianRuyi GanYan SongJiaxing ZhangYongdong Zhang

article2024ACL86 citations

Develops ChiMed-GPT, an open-source Chinese medical large language model trained through continual pre-training, supervised fine-tuning, and rejection-sampling-based reinforcement learning from human feedback to generate accurate, safe clinical responses with an extended context window.

Listen

Growing healthcare demands and a constrained medical workforce have created a major strain on medical services. Natural language processing and large language models offer strong potential to automate routine healthcare communication and assist clinical workflows. However, general-purpose models often lack specialized medical knowledge, while existing medical language models predominantly rely solely on supervised fine-tuning. This limited training approach restricts their ability to master deep medical knowledge, align with human preferences, handle longer clinical contexts, and mitigate harmful biases.

The article demonstrates the development and evaluation of CHIMED-GPT, a 13-billion-parameter open-source benchmark language model built specifically for Chinese medical text processing. The main objective was to evaluate whether executing a comprehensive training regime—incorporating continuous pre-training, supervised fine-tuning, and reinforcement learning from human feedback—along with an expanded context window, delivers superior diagnostic accuracy, conversational quality, and safety compared to existing general and medical models.

To achieve this, the authors built CHIMED-GPT upon the Ziya-13B-v2 foundation model and extended its context capacity to 4,096 tokens, doubling the 2,048-token limit typical of existing open-source medical models. The team trained the model across three distinct stages using diverse data sources: continuous pre-training on 214 million tokens from medical encyclopedias and textbooks; supervised fine-tuning across over 1.2 million doctor-patient interaction and dialogue pairs alongside 100,000 safety prompts; and reinforcement learning via rejection sampling fine-tuning. For the feedback stage, the authors constructed a fine-grained reward model by ranking outputs across clinical doctors, commercial models, and existing domain models. The resulting model was evaluated across information extraction, multiple-choice and open-ended question answering, and multi-turn dialogue generation, alongside standardized bias assessments measuring public and clinical attitudes toward mental illness.

The findings show that CHIMED-GPT consistently outperforms both general and medical baseline models. In medical information extraction, the model achieved entity recognition F1 scores of 40.82 and 41.04 across two benchmark datasets, substantially outperforming existing Chinese medical models (which scored between 23.80 and 30.90) and matching GPT-4. In multi-choice medical exams, CHIMED-GPT reached 68.29% accuracy on the C-Eval benchmark, surpassing all open-source models by 19 to 32 percentage points and closely trailing GPT-4 (71.29%). In multi-turn dialogue generation, the model achieved a ROUGE-L score of 42.16, roughly doubling the performance of GPT-4 (17.14) and other Chinese medical models (under 15.0). Human evaluations confirmed that CHIMED-GPT generated significantly more fluent, complete, and precise clinical advice, while bias tests demonstrated the lowest prejudice scores among all tested models when responding to sensitive mental health prompts.

These results demonstrate that a full training lifecycle combining domain-specific pre-training with preference-aligned reinforcement learning is essential for deploying language models in high-stakes fields. For healthcare organizations and technology leaders, specialized models trained in this manner reduce clinical risks, minimize misinformation, and provide safer interactions than general-purpose commercial interfaces. The model's expanded context capacity also allows it to digest lengthy medical records and patient histories without losing coherence.

Organizations developing or deploying medical artificial intelligence should adopt end-to-end training pipelines that include human preference alignment rather than relying solely on supervised fine-tuning. Before integrating such models into patient-facing or clinical decision-support systems, stakeholders should conduct pilot studies in controlled clinical environments to evaluate real-time diagnostic performance and workflow integration.

The article's primary limitations include reliance on existing benchmark datasets and simulated dialogue interactions, with human evaluations restricted to sample sizes of 50 instances per task. While confidence in the model's textual processing and safety alignment is high based on empirical benchmarks, readers should exercise appropriate caution regarding real-world clinical deployment until broader clinical trials and regulatory reviews are conducted.

Tian et al (2024).pdf
Cover for ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences

Table of Contents

  • Introduction
  • 2 The CHIMED-GPT
  • 2.1 Pre-training
  • 2.2 Supervised Fine-tuning
  • 2.3 RLHF
  • 3 Results and Analysis
  • 3.1 Information Extraction
  • 3.2 Question Answering
  • 3.3 Multi-turn Dialogue
  • 4 Bias Analysis
  • 5 Related Work
  • 6 Conclusion
  • References
  • Appendix A: The Effect of Data Augmentation for Reward Model Training
  • Appendix B: Full Examples in Case Study for QA
  • Appendix C: More Comparisons between CHIMED-GPT and Other LLMs
  • Appendix D: Full Example in Case Study for Dialogue
  • Appendix E: More Showcases for Bias Analysis

Knowls

  1. Knowl 1 — CHIMED-GPT’s foundation and full training regime

    model/method

    CHIMED-GPT is a Chinese medical language model built by continuing training Ziya-13B-v2, a 13-billion-parameter decoder-only Transformer pretrained on 600 billion Chinese and English tokens. It inherits the base model’s 4,096-token context length. Its training has three stages: continued medical-domain pretraining, supervised fine-tuning (SFT) on medical instructions and dialogues, and reinforcement learning from human feedback (RLHF) using reward-model-guided rejection sampling.

  2. Knowl 2 — Continued pretraining on Chinese medical texts

    model/method

    CHIMED-GPT’s continued pretraining uses the pretraining subset of the Chinese Medical Dataset (CMD): medical encyclopedia documents and medical-textbook articles. The paper reports 369,800 instances, 214 million tokens, and 603 MB; it also describes the collection as including 8,475 textbook articles. Training uses next-token prediction, byte-pair encoding with the Ziya-13B-v2 vocabulary, and AdamW with β1=0.9\beta_1=0.9, β2=0.95\beta_2=0.95, initial learning rate 5×10−55\times10^{-5}, weight decay 0.1, and gradient clipping at 1.0. Distributed training uses Megatron-LM with tensor parallelism of 2, alongside bf16 mixed precision, ZeRO, and flash attention.

  3. Knowl 3 — Medical instruction and dialogue supervised fine-tuning

    model/method

    SFT trains CHIMED-GPT on prompt–response pairs from four Chinese medical datasets: ChiMed (200,744 instances; 84 million tokens), CMD (SFT) (1,015,000; 460 million tokens), MC (44,983; 17 million tokens), and MedDialog (9,060; 3 million tokens). For question–answer data, the question is the prompt and its answer is the target; for doctor–patient dialogues, the dialogue history and latest patient utterance form the prompt, and the doctor’s reply is the target. The training mixture also includes 100,000 Safety-Prompts examples with desired responses to toxic prompts, including refusals. The authors remove redundant material and personal information, optimize full model parameters with cross-entropy, and concatenate short prompt–response examples to use the available sequence length. SFT uses learning rate 2×10−52\times10^{-5}, weight decay 0.1, and batch size 16.

  4. Knowl 4 — Augmented preference data for reward-model training

    model/method

    The RLHF reward model is trained using 4,000 CMD (Reward) instances, split into 3,800 training, 100 validation, and 100 test instances. Each instance contains a question, a doctor-written accepted answer, and a rejected answer generated by the Chinese medical model BenTsao. To create finer-grained preferences, the authors add answers generated by GPT-4 and GPT-3.5-Turbo and rank the four responses as doctor answer > GPT-4 > GPT-3.5-Turbo > BenTsao. Each adjacent pair in that ranking becomes a positive-versus-negative training example. Reward-model training runs for two epochs with batch size 8, a cosine learning-rate schedule from 5×10−65\times10^{-6} down to 10% of that value, and warm-up over 3% of steps (at least 5 steps). The authors report that training on the original binary pairs reached 98% validation accuracy after about 200 steps, which they regarded as rapid convergence with overfitting risk; the augmented, more fine-grained data produced lower overall validation accuracy, which they interpret as a more challenging task and less overfitting.

  5. Knowl 5 — Reward-guided rejection-sampling fine-tuning

    algorithm

    After training the reward model, the authors randomly draw 10,000 prompts from the SFT data and generate CHIMED-GPT responses to them. The reward model scores the generated responses; responses are ranked by score, and the top-kk are used as target answers for further fine-tuning. The paper does not specify the value of kk. This rejection-sampling stage runs for 400 iterations with batch size 64 and AdamW (β1=0.9\beta_1=0.9, β2=0.95\beta_2=0.95, ϵ=10−5\epsilon=10^{-5}), learning rate 10−510^{-5}, and weight decay 0.1.

  6. Knowl 6 — Evaluation design for medical language tasks

    experimental setup

    CHIMED-GPT is evaluated on Chinese medical named-entity recognition (NER), question answering (QA), and multi-turn dialogue generation against general-purpose and medical LLMs, including GPT-3.5-Turbo, GPT-4, Ziya variants, Baichuan, Taiyi, and MedicalGPT variants. NER uses five-shot prompts on CCKS-2019 and ChiMST and reports F1. Multi-choice QA uses five-shot prompts and accuracy on the medical subsets of C-Eval and CMMLU and the Chinese subset of MedQA; open-ended QA uses zero-shot prompts on ChiMed and reports BLEU-1/2 and ROUGE-1/2/L. Dialogue generation uses zero-shot prompts on the MC dataset and reports BLEU and ROUGE. For the automatic evaluations, each experiment is run five times and the mean is reported. Human evaluation samples 50 QA answers or dialogues and has two annotators score fluency, completeness, and precision from 1 to 3, with higher scores indicating better quality.

  7. Knowl 7 — NER results: near GPT-4 and strongest among open models

    empirical result

    In five-shot NER on CCKS-2019 and ChiMST, CHIMED-GPT achieves F1 scores of 40.82 and 41.04, respectively. For comparison, the CCKS-2019 / ChiMST scores are: GPT-3.5-Turbo 31.42 / 32.15; GPT-4 41.37 / 41.25; Ziya-v1 25.31 / 22.26; Ziya-v2 27.84 / 25.76; Baichuan 24.14 / 21.20; Taiyi 30.90 / 30.55; MedicalGPT (Z) 29.59 / 28.12; and MedicalGPT (B) 23.80 / 26.16. Thus CHIMED-GPT has the highest score among the compared open-source LLMs on both datasets, while GPT-4 is slightly higher on both.

  8. Knowl 8 — Question-answering benchmark and human-evaluation results

    empirical result

    On five-shot multiple-choice QA, the metric order below is C-Eval accuracy / CMMLU accuracy / MedQA accuracy. Scores are: GPT-3.5-Turbo 56.58 / 49.91 / 44.50; GPT-4 71.29 / 69.55 / 67.99; Ziya-v1 36.59 / 29.07 / 12.50; Ziya-v2 39.02 / 49.06 / 13.00; Baichuan 41.46 / 45.28 / 13.00; Taiyi 48.78 / 45.20 / 39.20; MedicalGPT (Z) 48.78 / 34.56 / 25.99; MedicalGPT (B) 39.02 / 43.82 / 18.50; and CHIMED-GPT 68.29 / 52.92 / 44.50. GPT-4 leads the multiple-choice results; CHIMED-GPT ties GPT-3.5-Turbo on MedQA. On zero-shot open-ended QA on ChiMed, the metric order is BLEU-1 / BLEU-2 / ROUGE-1 / ROUGE-2 / ROUGE-L: GPT-3.5-Turbo 33.61 / 28.27 / 26.51 / 7.13 / 16.63; GPT-4 39.15 / 32.85 / 26.61 / 7.31 / 16.84; Ziya-v1 6.18 / 5.77 / 18.59 / 3.94 / 12.66; Ziya-v2 38.41 / 31.90 / 26.91 / 7.90 / 18.67; Baichuan 5.81 / 5.25 / 16.91 / 3.01 / 11.30; Taiyi 11.73 / 9.96 / 21.76 / 5.26 / 15.46; MedicalGPT (Z) 39.02 / 32.35 / 26.76 / 8.10 / 18.16; MedicalGPT (B) 5.82 / 5.26 / 16.61 / 2.94 / 11.11; and CHIMED-GPT 44.58 / 37.22 / 27.11 / 8.89 / 19.86. CHIMED-GPT leads all five open-ended metrics. In human QA ratings (fluency / completeness / precision), it scores 2.57 / 2.45 / 2.57, compared with Taiyi at 2.17 / 2.02 / 2.01, MedicalGPT (Z) at 2.30 / 2.10 / 2.13, and MedicalGPT (B) at 2.27 / 2.17 / 2.22.

  9. Knowl 9 — Multi-turn dialogue results

    empirical result

    On zero-shot multi-turn dialogue generation using MC, scores are reported in BLEU-1 / BLEU-2 / ROUGE-1 / ROUGE-2 / ROUGE-L order. CHIMED-GPT scores 33.14 / 30.86 / 43.43 / 34.91 / 42.16. The comparison scores are: GPT-3.5-Turbo 18.58 / 15.76 / 18.92 / 6.62 / 14.55; GPT-4 24.29 / 20.17 / 20.64 / 8.39 / 17.14; Ziya-v1 15.85 / 11.75 / 9.92 / 3.04 / 9.02; Ziya-v2 14.21 / 10.99 / 12.20 / 4.45 / 10.61; Baichuan 3.44 / 1.61 / 3.87 / 0.34 / 3.49; Taiyi 5.81 / 4.67 / 14.23 / 4.55 / 11.99; MedicalGPT (Z) 20.26 / 16.42 / 17.51 / 5.42 / 14.21; and MedicalGPT (B) 3.94 / 2.19 / 4.34 / 0.13 / 3.50. In human ratings of fluency / completeness / precision, CHIMED-GPT scores 2.44 / 2.38 / 2.50, versus Taiyi at 1.96 / 2.01 / 2.02, MedicalGPT (Z) at 2.09 / 2.05 / 2.11, and MedicalGPT (B) at 2.15 / 2.23 / 2.20. CHIMED-GPT therefore leads the reported automatic dialogue metrics and human-scored comparisons.

  10. Knowl 10 — Prompt-based bias assessment on mental-illness attitudes

    experimental setup

    The authors assess attitudes expressed by CHIMED-GPT and comparison LLMs using two established statement scales: CAMI, with 40 statements about public attitudes toward people with mental illnesses, and MICA, with 16 statements about clinicians’ attitudes toward such patients. They manually translate the English statements into Chinese, prompt models to select an agreement response, and map answers to bias scores using each scale’s scoring rules; higher scores indicate stronger bias. CAMI scores range from 1 to 5 and MICA scores from 1 to 6. The reported comparison finds that CHIMED-GPT has the lowest average bias score among the tested models on both scales. The paper presents these averages graphically rather than reporting exact numerical values in the text.

Coverage note — The appendix’s individual qualitative examples of medicine-description generation, medical-record generation, and responses to toxic prompts are omitted because they are illustrative cases rather than separate systematic evaluations; the main benchmark results and scale-based bias analysis are included.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language Models are Few-shot Learners. Advances in neural information processing systems, 33:1877–1901.
  2. 2.Yang Trista Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta, Varun Kumar, Jwala Dhamala, and Aram Galstyan. 2022. On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations. arXiv preprint arXiv:2203.13928.
  3. 3.Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023. Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1504–1532, Toronto, Canada.
  4. 4.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality.
  5. 5.Jiaxi Cui, Zongjian Li, Yang Yan, Bohua Chen, and Li Yuan. 2023. ChatLaw: Open-source Legal Large Language Model with Integrated External Knowledge Bases. arXiv preprint arXiv:2306.16092.
  6. 6.Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. Flashattention: Fast and Memory-efficient Exact Attention with Io-awareness. Advances in Neural Information Processing Systems, 35:16344–16359.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  8. 8.Shizhe Diao, Jiaxin Bai, Yan Song, Tong Zhang, and Yonggang Wang. 2020. ZEN: Pre-training Chinese Text Encoder Enhanced by N-gram Representations. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4729–4740.
  9. 9.Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11737–11762, Toronto, Canada.
  10. 10.Jheanell Gabbidon, Sarah Clement, Adrienne van Nieuwenhuizen, Aliya Kassam, Elaine Brohan, Ian Norman, and Graham Thornicroft. 2013. Mental Illness: Clinicians’ Attitudes (mica) Scale—Psychometric Properties of a Version for Healthcare Students and Professionals. Psychiatry research, 206(1):81–87.
  11. 11.Ruyi Gan, Ziwei Wu, Renliang Sun, Junyu Lu, Xiaojun Wu, Dixiang Zhang, Kunhao Pan, Ping Yang, Qi Yang, Jiaxing Zhang, and Yan Song. 2023. Ziya2: Data-centric Learning is All LLMs Need. arXiv preprint arXiv:2311.03301.
  12. 12.Patrick Haller, Ansar Aynetdinov, and Alan Akbik. 2023. OpinionGPT: Modelling Explicit Biases in Instruction-Tuned LLMs. arXiv preprint arXiv:2309.03876.
  13. 13.Jialong Han, Yan Song, Wayne Xin Zhao, Shuming Shi, and Haisong Zhang. 2018. Hyperdoc2vec: Distributed Representations of Hypertext Documents. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2384–2394, Melbourne, Australia.
  14. 14.Tianyu Han, Lisa C Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Löser, Daniel Truhn, and Keno B Bressem. 2023. MedAlpaca–An Open-Source Collection of Medical Conversational AI Models and Training Data. arXiv preprint arXiv:2304.08247.
  15. 15.Xianpei Han, Zhichun Wang, Jiangtao Zhang, Qinghua Wen, Wenqi Li, Buzhou Tang, Qi Wang, Zhifan Feng, Yang Zhang, Yajuan Lu, et al. 2020. Overview of the CCKS 2019 Knowledge Graph Evaluation Track: Entity, Relation, Event and QA. arXiv preprint arXiv:2003.03875.
  16. 16.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring Massive Multitask Language Understanding. arXiv preprint arXiv:2009.03300.
  17. 17.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. 2023. C-Eval: A Multi-level Multi-Discipline Chinese Evaluation Suite for Foundation Models. arXiv preprint arXiv:2305.08322.
  18. 18.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What Disease does This Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams. Applied Sciences, 11(14):6421.
  19. 19.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online.
  20. 20.Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2023a. CMMLU: Measuring Massive Multitask Language Understanding in Chinese. arXiv preprint arXiv:2306.09212.
  21. 21.Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang. 2023b. ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge. Cureus, 15(6).
  22. 22.Ilya Loshchilov and Frank Hutter. 2017. Decoupled Weight Decay Regularization. arXiv preprint arXiv:1711.05101.
  23. 23.Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining. Briefings in Bioinformatics, 23(6):bbac409.
  24. 24.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. 2017. Mixed Precision Training. arXiv preprint arXiv:1710.03740.
  25. 25.OpenAI. 2023. GPT-4 Technical Report. ArXiv, abs/2303.08774.
  26. 26.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training Language Models to Follow Instructions with Human Feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  27. 27.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language Models are Unsupervised Multitask Learners. OpenAI blog, 1(8):9.
  28. 28.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-text Transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  29. 29.Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. Zero: Memory Optimizations toward Training Trillion Parameter Models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16. IEEE.
  30. 30.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015. Neural Machine Translation of Rare Words with Subword Units. arXiv preprint arXiv:1508.07909.
  31. 31.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training Multi-billion Parameter Language Models using Model Parallelism. arXiv preprint arXiv:1909.08053.
  32. 32.Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, et al. 2023. Towards Expert-level Medical Question Answering with Large Language Models. arXiv preprint arXiv:2305.09617.
  33. 33.Yan Song and Shuming Shi. 2018. Complementary Learning of Word Embeddings. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18, pages 4368–4374.
  34. 34.Yan Song, Shuming Shi, and Jing Li. 2018. Joint Learning Embeddings for Chinese Words and Their Components via Ladder Structured Networks. In Proceedings of the 27th International Joint Conference on Artificial Intelligence, pages 4375–4381.
  35. 35.Yan Song, Yuanhe Tian, Nan Wang, and Fei Xia. 2020. Summarizing Medical Conversations via Identifying Important Utterances. In Proceedings of the 28th International Conference on Computational Linguistics, pages 717–729.
  36. 36.Yan Song, Tong Zhang, Yonggang Wang, and Kai-Fu Lee. 2021. ZEN 2.0: Continue Training and Adaption for N-gram Enhanced Text Encoders. arXiv preprint arXiv:2105.01279.
  37. 37.Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang. 2023. Safety Assessment of Chinese Large Language Models. arXiv preprint arXiv:2304.10436.
  38. 38.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford Alpaca: An Instruction-following LLaMA Model. GitHub repository.
  39. 39.S Martin Taylor and Michael J Dear. 1981. Scaling Community Attitudes toward the Mentally Ill. Schizophrenia bulletin, 7(2):225–240.
  40. 40.Yuanhe Tian, Weicheng Ma, Fei Xia, and Yan Song. 2019. ChiMed: A Chinese Medical Corpus for Question Answering. In Proceedings of the 18th BioNLP Workshop and Shared Task, pages 250–260, Florence, Italy.
  41. 41.Yuanhe Tian, Han Qin, Fei Xia, and Yan Song. 2022. ChiMST: A Chinese Medical Corpus for Word Segmentation and Medical Term Recognition. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 5654–5664, Marseille, France.
  42. 42.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and Efficient Foundation Language Models. arXiv preprint arXiv:2302.13971.
  43. 43.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open Foundation and Fine-tuned Chat Models. arXiv preprint arXiv:2307.09288.
  44. 44.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is All You Need. Advances in neural information processing systems, 30.
  45. 45.A Venigalla, J Frankle, and M Carbin. 2022. BioMedLM: a Domain-specific Large Language Model for Biomedical Text. MosaicML. Accessed: Dec, 23(3):2.
  46. 46.Haochun Wang, Chi Liu, Nuwa Xi, Zewen Qiang, Sendong Zhao, Bing Qin, and Ting Liu. 2023. Huatuo: Tuning Llama Model with Chinese Medical Knowledge. arXiv preprint arXiv:2304.06975.
  47. 47.Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023. BloombergGPT: A Large Language Model for Finance. arXiv preprint arXiv:2303.17564.
  48. 48.Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2023a. Baize: An Open-source Chat Model with Parameter-efficient Tuning on Self-chat Data. arXiv preprint arXiv:2304.01196.
  49. 49.Guohai Xu, Jiayi Liu, Ming Yan, Haotian Xu, Jinghui Si, Zhuoran Zhou, Peng Yi, Xing Gao, Jitao Sang, Rong Zhang, et al. 2023b. CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility. arXiv preprint arXiv:2307.09705.
  50. 50.Ming Xu. 2023. MedicalGPT: Training Medical GPT Model.
  51. 51.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019. XLNet: Generalized Autoregressive Pretraining for Language Understanding. In Advances in Neural Information Processing Systems 32, pages 5753–5763.
  52. 52.Jiaxing Zhang, Ruyi Gan, Junjie Wang, Yuxiang Zhang, Lin Zhang, Ping Yang, Xinyu Gao, Ziwei Wu, Xiaoqun Dong, Junqing He, Jianheng Zhuo, Qi Yang, Yongfeng Huang, Xiayu Li, Yanghan Wu, Junyu Lu, Xinyu Zhu, Weifeng Chen, Ting Han, Kunhao Pan, Rui Wang, Hao Wang, Xiaojun Wu, Zhongshen Zeng, and Chongpei Chen. 2022. Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence. CoRR, abs/2209.02970.

Citation

MLA
Tian, Y., et al. “ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7156–73, https://doi.org/10.18653/v1/2024.acl-long.386.
APA
Tian, Y., Gan, R., Song, Y., Zhang, J., & Zhang, Y. (2024). ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7156–7173. https://doi.org/10.18653/v1/2024.acl-long.386
Chicago
Tian, Y., R. Gan, Y. Song, J. Zhang, and Y. Zhang. 2024. “ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7156–73. https://doi.org/10.18653/v1/2024.acl-long.386.
Harvard
Tian, Y. et al. (2024) “ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7156–7173. Available at: https://doi.org/10.18653/v1/2024.acl-long.386.
Vancouver
1. Tian Y, Gan R, Song Y, Zhang J, Zhang Y (2024) ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7156–7173

BibTeX

@inproceedings{tian-etal-2024-chimed,
    title = "{C}hi{M}ed-{GPT}: A {C}hinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences",
    author = "Tian, Yuanhe  and
      Gan, Ruyi  and
      Song, Yan  and
      Zhang, Jiaxing  and
      Zhang, Yongdong",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.386/",
    doi = "10.18653/v1/2024.acl-long.386",
    pages = "7156--7173"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/