AlignBench: Benchmarking Chinese Alignment of Large Language Models

Xiao LiuXuanyu LeiShengyuan WangYue HuangAndrew FengBosi WenJiale ChengPei KeYifan XuWeng Lam Tam

article2024ACL75 citationsArea Chair Award

Presents ALIGNBENCH, a comprehensive Chinese alignment benchmark spanning eight real-world task categories that pairs human-verified reference evidence with a rule-calibrated, multi-dimensional LLM-as-judge evaluation method to reliably grade open-ended model responses.

Listen

As large language models become central to practical digital workflows, aligning them to follow human instructions and satisfy user preferences is essential. While English evaluation frameworks have matured, there has been a critical lack of standardized, challenging, and automated benchmarks specifically designed to measure alignment in Chinese-language applications. To resolve this problem, the article presents ALIGNBENCH, a comprehensive evaluation suite designed to assess how effectively language models fulfill authentic, open-ended Chinese user queries.

The benchmark establishes a curated dataset of 683 challenging tasks derived from real-world usage and researcher-designed problems across eight core categories, including reasoning, mathematics, writing, role play, and advanced Chinese understanding. To evaluate open-ended outputs accurately without incurring prohibitive manual evaluation costs, the authors developed a rule-calibrated, multidimensional evaluation methodology using GPT-4 as an automated judge. Each query is paired with a human-verified reference answer—supported by web citations for knowledge-intensive topics—which acts as an objective baseline to anchor automated grading.

Evaluating 17 prominent open-source and proprietary models on this framework revealed several pivotal findings. First, top-tier frontier models like GPT-4 lead overall performance with an average score of 8.01 out of 10, maintaining substantial leads in factual correctness, complex logic, and mathematical reasoning. Second, leading Chinese commercial models achieved competitive scores near or exceeding 6.0, performing on par with or slightly above GPT-3.5 and demonstrating equal or superior mastery in culturally specific Chinese language tasks. Third, open-source Chinese models showed rapid maturation, with models such as Qwen-14B and Baichuan2-13B rivaling several closed commercial APIs. Finally, the proposed evaluation method demonstrated superior alignment with human judgments, outperforming standard automated judging techniques by 12% to 20% in pairwise quality comparisons.

These findings suggest that while regional developers have closed the gap in linguistic fluency and cultural nuances, severe weaknesses persist in complex logical reasoning and mathematical problem-solving. This divide poses operational risks for organizations deploying models in specialized domains that require absolute factual and computational accuracy. Consequently, developers should prioritize high-quality reasoning and instruction-tuning datasets rather than pure fluency optimization. Meanwhile, benchmark maintainers must expand evaluation coverage to include long-context tasks and integrate dynamic fact-checking mechanisms, as automated judges can still be misled when reference data is incomplete or flawed.

arXiv: 2311.18743
Cover for AlignBench: Benchmarking Chinese Alignment of Large Language Models

Abstract

Alignment has become a critical step for instruction-tuned Large Language Models (LLMs) to become helpful assistants. However, the effective evaluation of alignment for emerging Chinese LLMs is still significantly lacking, calling for real-scenario grounded, open-ended, challenging and automatic evaluations tailored for alignment. To fill in this gap, we introduce ALIGNBENCH, a comprehensive multi-dimensional benchmark for evaluating LLMs' alignment in Chinese. We tailor a human-in-the-loop data curation pipeline, containing 8 main categories, 683 real-scenario rooted queries and corresponding human verified references. To ensure the correctness of references, each knowledge-intensive query is accompanied with evidences collected from reliable web sources (including URLs and quotations) by our annotators. For automatic evaluation, our benchmark employs a rule-calibrated multi-dimensional LLM-as-Judge (Zheng et al., 2023) approach with Chain-of-Thought to generate explanations and final ratings, ensuring high reliability and interpretability. All evaluation codes, data, and LLM generations are available at https://github.com/THUDM/AlignBench.

Table of Contents

  • 1 Introduction
  • 2 Dataset
  • 2.1 Dataset Composition
  • 2.2 Dataset Construction
  • 3 Methods
  • 4 Human Evaluation on ALIGNBENCH
  • 4.1 Agreement Evaluation
  • 4.2 Quality Evaluation
  • 5 ALIGNBENCH: Benchmarking Results
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Detailed Description of the Dataset
  • A.1.1 Fundamental Language Ability
  • A.1.2 Advanced Chinese Understanding
  • A.1.3 Open-ended Questions
  • A.1.4 Writing Ability
  • A.1.5 Logical Reasoning
  • A.1.6 Mathematics
  • A.1.7 Task-oriented Role Play
  • A.1.8 Professional Knowledge
  • A.2 Prompts and Details of Methods
  • A.3 Dimensional Performance
  • A.4 Case Study
  • A.4.1 Misleading.
  • A.4.2 Reference-free Judgements

Knowls

  1. Knowl 1 — ALIGNBENCH dataset composition and taxonomy

    data/table

    ALIGNBENCH is a Chinese alignment benchmark containing 683 real-scenario, open-ended user queries with human-verified reference answers. The dataset is organized into eight categories, with the following sample counts and intended capabilities:

    • Fundamental Language Ability: 68 samples covering commonsense knowledge, reading comprehension, translation, text classification, and information extraction.
    • Advanced Chinese Understanding: 58 samples covering Chinese characters, idioms, culture, history, implicit meaning, humor, and sarcasm.
    • Open-ended Questions: 38 samples requiring opinions, advice, recommendations, or tutoring for everyday and professional situations.
    • Writing Ability: 75 samples covering practical, creative, professional, and customized writing with requirements on style, format, consistency, and creativity.
    • Logical Reasoning: 92 samples requiring deductive, abductive, multi-hop, commonsense, proof-oriented, or puzzle-based reasoning.
    • Mathematics: 112 samples ranging from elementary to advanced and applied mathematics, including calculation, problem solving, concept explanation, and theorem-proof tasks.
    • Task-oriented Role Play: 116 samples requiring the model to adopt an identity and perform a corresponding task, including celebrity, emotional, entertainment, functional, and daily-life roles.
    • Professional Knowledge: 124 samples covering domain-specific questions in areas such as physics, chemistry, computer science, medicine, economics, astronomy, sociology, history, music, law, sports, geography, and literature.

    The eight category counts sum to 683. Unlike closed-form knowledge tests, the benchmark is designed to measure whether a Chinese LLM can satisfy realistic user intentions through detailed open-ended responses.

  2. Knowl 2 — Human-in-the-loop query and reference curation pipeline

    algorithm

    ALIGNBENCH constructs its evaluation set through a semi-automatic human-in-the-loop pipeline.

    Input: Candidate Chinese user queries from an online chat service and researcher-written challenging problems
    Output: Filtered benchmark items containing a query, a category, and a verified reference answer
    1. Retain queries that express a concrete user task or intention.
    2. Remove queries that are unclear, linguistically awkward, insufficiently difficult, or unsafe and insensitive.
    3. Use GPT-4 to draft an answer for each retained query.
    4. Have human annotators inspect, revise, and refine each draft into a reference answer.
    5. For knowledge-intensive queries, especially in Professional Knowledge, Mathematics, and Logical Reasoning,
       have annotators search reliable web sources and record URLs and quoted evidence before finalizing the reference.
    6. Ask GPT-3.5-turbo, ChatGLM, and Sparkdesk to answer the candidate queries.
    7. Use GPT-4 to score the three model responses and calculate their average score for each query.
    8. Discard the 50% of queries with the highest average scores, because those queries are less difficult.
    9. Classify the remaining queries into the eight-category taxonomy.
    10. Return the resulting query, human-verified reference, and category.

    The pipeline combines real-user authenticity with researcher-controlled difficulty and human verification. Web evidence is explicitly retained for knowledge-intensive items; such items account for 66.5% of ALIGNBENCH.

  3. Knowl 3 — Rule-calibrated multi-dimensional point-wise judging

    model/method

    ALIGNBENCH evaluates each model response independently with GPT-4 rather than comparing pairs of responses. The judge receives a user query, a human-curated reference answer, and one candidate response. It first produces a chain-of-thought-style analytical explanation, evaluates the response along task-specific dimensions, and then outputs an overall integer score from 1 to 10.

    The reference answer is anchored at score 8 as a scoring pivot. The grading rules map score ranges to reference-relative quality: scores 1–2 indicate irrelevance, severe errors, or harmful content; 3–4 indicate low quality with minor errors; 5–6 indicate medium quality; 7–8 indicate quality comparable to the reference; and 9–10 indicate quality exceeding the reference. The judge is instructed to prioritize factual correctness and satisfaction of the user’s request when aggregating dimensions.

    This design combines three elements: point-wise scoring, which avoids the quadratic cost and position bias of pairwise comparisons; chain-of-thought explanations, which improve interpretability; and explicit reference-based rules, which reduce score variance and overproduction of very high scores.

  4. Knowl 4 — Task-conditioned evaluation dimensions and generation temperatures

    model/method

    ALIGNBENCH does not apply one universal rubric to all query types. It assigns evaluation dimensions according to the nature of the requested response:

    • Factual and explanatory questions, used for Fundamental Language Ability, Advanced Chinese Understanding, and Professional Knowledge, are judged on factual correctness, user satisfaction, clarity, and completeness.
    • Logical reasoning questions, used for Logical Reasoning and Mathematics, are judged on factual correctness, user satisfaction, logical coherence, and completeness.
    • Generative questions, used for Writing Ability and Task-oriented Role Play, are judged on factual correctness, user satisfaction, logical coherence, creativity, and richness.
    • Recommendation questions, used for Open-ended Questions, are judged on factual correctness, user satisfaction, fairness and responsibility, and creativity.

    The dimensions are defined operationally: factual correctness concerns accuracy and reliable evidence; user satisfaction concerns fulfilling the user’s purpose; logical coherence concerns consistency and absence of contradiction; completeness concerns coverage of needed information; clarity concerns ease of understanding; richness concerns depth, context, detail, and examples; creativity concerns novelty; and fairness and responsibility concern feasibility, risks, and consequences.

    The benchmark also conditions response generation temperature on task type. Mathematics, Logical Reasoning, Fundamental Language Ability, Advanced Chinese Understanding, and Professional Knowledge use temperature 0.1 for deterministic answers, whereas Open-ended Questions, Writing Ability, and Task-oriented Role Play use temperature 0.7 to encourage diverse and creative responses.

  5. Knowl 5 — Agreement with human evaluation

    empirical result

    The rule-calibrated multi-dimensional judge was evaluated on 400 sampled queries and 3,200 question-answer pairings generated by eight Chinese-capable LLMs. Human annotators rated responses from 1 to 5, with 1 denoting an irrelevant, incorrect, or potentially harmful response. GPT-4 judges were compared using sample-level Pearson correlation, system-level Pearson correlation, and pairwise agreement excluding ties.

    Across all categories, the general judge obtained sample-level Pearson 0.618, system-level Pearson 0.998, and pairwise agreement 0.751. The rule-calibrated judge obtained 0.628, 0.999, and 0.724, respectively. ALIGNBENCH’s combined rule-calibrated multi-dimensional judge obtained 0.638, 0.998, and 0.753, respectively.

    Thus, the combined method had the strongest overall sample-level correlation and the strongest overall pairwise agreement with human judgments, while all three methods showed very high system-level correlation. Category-level results were mixed rather than uniformly dominant: for example, the combined method achieved sample-level Pearson correlations of 0.739 for Advanced Chinese Understanding, 0.589 for Fundamental Language Ability, 0.677 for Writing Ability, 0.544 for Open-ended Questions, 0.653 for Task-oriented Role Play, and 0.622 for Logical Reasoning.

  6. Knowl 6 — Human preference for explanation quality

    empirical result

    ALIGNBENCH evaluated the quality of judge explanations separately from score agreement. The experiment sampled 500 question-answer pairs, generated explanations with the general, rule-calibrated, and combined rule-calibrated multi-dimensional judges, and formed 1,500 pairwise comparisons for human annotators. Annotators assessed rationality, readability, and consistency between the explanation and its final rating.

    The combined judge defeated the general judge in 217 comparisons, tied in 94, and lost in 155; its win rate excluding ties was 58.3%, an advantage of 12.4 percentage points. Against the rule-calibrated judge, the combined judge won 241 comparisons, tied in 102, and lost in 139; its win rate was 63.4%, an advantage of 20.4 percentage points. The rule-calibrated judge defeated the general judge with 186 wins, 167 ties, and 147 losses, for a 55.9% win rate and a 7.8-point advantage.

    These comparisons indicate that task-specific dimensions combined with explicit reference-based rules produced explanations that human annotators considered more rational, readable, consistent, and useful as feedback.

  7. Knowl 7 — Benchmark scores for 17 Chinese-capable LLMs

    data/table

    The benchmark evaluation used gpt-4-0613 as the judge and reported an overall score plus reasoning and language aggregates. In the list below, R is the reasoning average, M is Mathematics, L is Logical Reasoning, Lang is the language average, F is Fundamental Language Ability, C is Advanced Chinese Understanding, O is Open-ended Questions, W is Writing Ability, Role is Task-oriented Role Play, and Pro is Professional Knowledge. Scores are on a 1–10 scale and are reported in the order Overall; R, M, L; Lang, F, C, O, W, Role, Pro.

    • gpt-4-1106-preview: 8.01; 7.73, 7.80, 7.66; 8.29, 7.99, 7.33, 8.61, 8.67, 8.47, 8.65.
    • gpt-4-0613: 7.53; 7.47, 7.56, 7.37; 7.59, 7.81, 6.93, 7.42, 7.93, 7.51, 7.94.
    • chatglm-turbo: 6.24; 5.00, 4.74, 5.26; 7.49, 6.82, 7.17, 8.16, 7.77, 7.76, 7.24.
    • erniebot-3.5: 6.14; 5.15, 5.03, 5.27; 7.13, 6.62, 7.60, 7.26, 7.56, 6.83, 6.90.
    • gpt-3.5-turbo-0613: 6.08; 5.35, 5.68, 5.02; 6.82, 6.71, 5.81, 7.29, 7.03, 7.28, 6.77.
    • chatglm-pro: 5.83; 4.65, 4.54, 4.75; 7.01, 6.51, 6.76, 7.47, 7.07, 7.34, 6.89.
    • spark_desk_v2: 5.74; 4.73, 4.71, 4.74; 6.76, 5.84, 6.97, 7.29, 7.18, 6.92, 6.34.
    • Qwen-14B-Chat: 5.72; 4.81, 4.91, 4.71; 6.63, 6.90, 6.36, 6.74, 6.64, 6.59, 6.56.
    • Baichuan2-13B-Chat: 5.25; 3.92, 3.76, 4.07; 6.59, 6.22, 6.05, 7.11, 6.97, 6.75, 6.43.
    • ChatGLM3-6B: 4.97; 3.85, 3.55, 4.14; 6.10, 5.75, 5.29, 6.71, 6.83, 6.28, 5.73.
    • Baichuan2-7B-Chat: 4.97; 3.66, 3.56, 3.75; 6.28, 5.81, 5.50, 7.13, 6.84, 6.53, 5.84.
    • InternLM-20B: 4.96; 3.66, 3.39, 3.92; 6.26, 5.96, 5.50, 7.18, 6.19, 6.49, 6.22.
    • Qwen-7B-Chat: 4.91; 3.73, 3.62, 3.83; 6.09, 6.40, 5.74, 6.26, 6.31, 6.19, 5.66.
    • ChatGLM2-6B: 4.48; 3.39, 3.16, 3.61; 5.58, 4.91, 4.52, 6.66, 6.25, 6.08, 5.08.
    • InternLM-Chat-7B: 3.65; 2.56, 2.45, 2.66; 4.75, 4.34, 4.09, 5.82, 4.89, 5.32, 4.06.
    • Chinese-LLaMA-2-7B-Chat: 3.57; 2.68, 2.29, 3.07; 4.46, 4.31, 4.26, 4.50, 4.63, 4.91, 4.13.
    • LLaMA-2-13B-Chinese-Chat: 3.35; 2.47, 2.21, 2.73; 4.23, 4.13, 3.31, 4.79, 3.93, 4.53, 4.71.
  8. Knowl 8 — Observed capability differences among evaluated models

    empirical result

    The 17-model evaluation shows a pronounced separation between the strongest GPT-4 variants and the remaining systems. gpt-4-1106-preview achieved the highest overall score, 8.01, followed by gpt-4-0613 at 7.53. The strongest non-GPT-4 systems scored near or above 6 overall, while the strongest listed open-source Chinese models, Qwen-14B-Chat and Baichuan2-13B-Chat, scored 5.72 and 5.25.

    Reasoning was a major weakness for most Chinese-oriented models. gpt-4-1106-preview scored 7.80 in Mathematics and 7.66 in Logical Reasoning, whereas chatglm-turbo scored 4.74 and 5.26, and Qwen-14B-Chat scored 4.91 and 4.71. In contrast, several Chinese-developed models performed relatively strongly on Advanced Chinese Understanding: erniebot-3.5 scored 7.60 and chatglm-turbo scored 7.17, compared with 7.33 for gpt-4-1106-preview. The paper attributes this possible advantage to Chinese-specific instruction-tuning data, while treating that explanation as an interpretation rather than a causal demonstration.

    The dimensional evaluation further found that gpt-4-1106-preview led in every reported dimension, with scores of 8.33 for correctness, 8.03 for user satisfaction, 9.25 for logical coherence, 8.40 for completeness, 9.12 for clarity, 8.04 for creativity, 8.59 for richness, and 9.58 for fairness and responsibility.

  9. Knowl 9 — Failure modes of reference-based and reference-free judging

    limitation

    The paper documents two concrete weaknesses of LLM-as-Judge evaluation. First, a wrong reference can mislead the evaluator. For the question asking which country bordering China has the highest GDP, the correct answer is India, but the supplied reference incorrectly stated Japan. GPT-4 answered India when solving the question directly, yet judged an answer of India as factually wrong when instructed to evaluate it against the erroneous reference.

    Second, insufficient references make factual verification difficult. For a question about the beginning of the modern Olympic Games, a response correctly stated 1896 but added that more than 30 countries participated. The available reference contained only the date, so the evaluator accepted the additional claim despite the paper identifying the actual number as 13. The authors suggest that autonomous factual verification backed by a dynamically updated information database would be needed to address this weakness.

  10. Knowl 10 — Scope limitations of ALIGNBENCH and its evaluator

    limitation

    ALIGNBENCH’s GPT-4 judge correlates relatively well with human ratings, but the paper reports substantial room to improve correlation and score granularity. LLM-as-Judge evaluations remain vulnerable to position bias, verbosity bias, and self-enhancement bias, which can distort comparisons for particular models.

    The benchmark’s 683 queries also do not exhaust the space of Chinese user needs. The authors state that more queries and topics are required for more stable evaluation, and that the current benchmark does not evaluate long-text queries. These limitations constrain how broadly ALIGNBENCH scores should be generalized beyond the sampled real-scenario tasks.

Coverage note — No substantial contributed material was omitted; related work, background, acknowledgements, and proof-only details were excluded as non-contributory.

References

  1. 1.Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, et al. 2023. PaLM 2 technical report.
  2. 2.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023a. Qwen technical report. arXiv preprint arXiv:2309.16609.
  3. 3.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  4. 4.Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. 2023b. Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20, Red Hook, NY, USA. Curran Associates Inc.
  6. 6.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2023. A survey on evaluation of large language models. arXiv preprint arXiv:2307.03109.
  7. 7.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  9. 9.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems.
  10. 10.Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023. Ultrafeedback: Boosting language models with high-quality feedback. arXiv preprint arXiv:2310.01377.
  11. 11.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335.
  12. 12.Oren Etzioni, Michele Banko, Stephen Soderland, and Daniel S Weld. 2008. Open information extraction from the web. Communications of the ACM, 51(12):68–74.
  13. 13.Amelia Glaese, Nat McAleese, Maja Tr˛ebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Sona Mokrá, Nicholas Fernando, Boxi Wu, Rachel ˇ Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving. 2022. Improving alignment of dialogue agents via targeted human judgements.
  14. 14.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In International Conference on Learning Representations.
  15. 15.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. 2023a. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. arXiv preprint arXiv:2305.08322.
  16. 16.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, et al. 2023b. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. arXiv preprint arXiv:2305.08322.
  17. 17.Pei Ke, Bosi Wen, Zhuoer Feng, Xiao Liu, Xuanyu Lei, Jaile Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, Jie Tang, and Minlie Huang. 2023. Critiquellm: Scaling llm-as-critic for effective and explainable evaluation of large language model generation. arXiv preprint arXiv:2311.18702.
  18. 18.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval.
  19. 19.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
  20. 20.Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. 2023a. Agentbench: Evaluating llms as agents. arXiv preprint arXiv:2308.03688.
  21. 21.Xiao Liu, Fanjin Zhang, Zhenyu Hou, Li Mian, Zhaoyu Wang, Jing Zhang, and Jie Tang. 2021. Self-supervised learning: Generative or contrastive. IEEE transactions on knowledge and data engineering, 35(1):857–876.
  22. 22.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b. G-eval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  23. 23.Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. 2023c. Trustworthy llms: a survey and guideline for evaluating large language models’ alignment.
  24. 24.OpenAI. 2022. Introducing chatgpt.
  25. 25.OpenAI. 2023. GPT-4 technical report. arXiv preprint arXiv:2303.08774.
  26. 26.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  27. 27.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290.
  28. 28.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research.
  29. 29.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  30. 30.Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. 2023. Principle-driven self-alignment of language models from scratch with minimal human supervision. arXiv preprint arXiv:2305.03047.
  31. 31.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158.
  32. 32.InternLM Team. 2023. Internlm: A multilingual language model with progressively enhanced capabilities. https://github.com/InternLM/InternLM-techreport.
  33. 33.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  34. 34.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems, 32.
  35. 35.Jiaan Wang, Yunlong Liang, Fandong Meng, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023a. Is chatgpt a good nlg evaluator? a preliminary study. arXiv preprint arXiv:2303.04048.
  36. 36.Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023b. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926.
  37. 37.Sida I Wang and Christopher D Manning. 2012. Baselines and bigrams: Simple, good sentiment and topic classification. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 90–94.
  38. 38.Xingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng, and Heng Ji. 2023c. Mint: Evaluating llms in multi-turn interaction with tools and language feedback.
  39. 39.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023d. Self-instruct: Aligning language models with self-generated instructions.
  40. 40.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  41. 41.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023a. Wizardlm: Empowering large language models to follow complex instructions.
  42. 42.Liang Xu, Anqi Li, Lei Zhu, Hang Xue, Changtai Zhu, Kangkang Zhao, Haonan He, Xuanwei Zhang, Qiyue Kang, and Zhenzhong Lan. 2023b. Superclue: A comprehensive chinese large language model benchmark. arXiv preprint arXiv:2307.15020.
  43. 43.Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, et al. 2023. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305.
  44. 44.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. 2022. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414.
  45. 45.Hui Zeng. 2023. Measuring massive multitask chinese understanding. arXiv preprint arXiv:2304.12986.
  46. 46.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  47. 47.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.
  48. 48.Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2023. Safetybench: Evaluating the safety of large language models with multiple choice questions. arXiv preprint arXiv:2309.07045.
  49. 49.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoic. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685.
  50. 50.Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023. Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364.
  51. 51.Ziyu Zhuang, Qiguang Chen, Longxuan Ma, Mingda Li, Yi Han, Yushan Qian, Haopeng Bai, Zixian Feng, Weinan Zhang, and Ting Liu. 2023. Through the lens of core competency: Survey on evaluation of large language models.

Citation

MLA
Liu, X., et al. “AlignBench: Benchmarking Chinese Alignment of Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 11621–40, https://doi.org/10.18653/v1/2024.acl-long.624.
APA
Liu, X., Lei, X., Wang, S., Huang, Y., Feng, A., Wen, B., Cheng, J., Ke, P., Xu, Y., Tam, W. L., Zhang, X., Sun, L., Gu, X., Wang, H., Zhang, J., Huang, M., Dong, Y., & Tang, J. (2024). AlignBench: Benchmarking Chinese Alignment of Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11621–11640. https://doi.org/10.18653/v1/2024.acl-long.624
Chicago
Liu, X., X. Lei, S. Wang, et al. 2024. “AlignBench: Benchmarking Chinese Alignment of Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11621–40. https://doi.org/10.18653/v1/2024.acl-long.624.
Harvard
Liu, X. et al. (2024) “AlignBench: Benchmarking Chinese Alignment of Large Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11621–11640. Available at: https://doi.org/10.18653/v1/2024.acl-long.624.
Vancouver
1. Liu X, Lei X, Wang S, et al (2024) AlignBench: Benchmarking Chinese Alignment of Large Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 11621–11640

BibTeX

@inproceedings{liu-etal-2024-alignbench,
    title = "{A}lign{B}ench: Benchmarking {C}hinese Alignment of Large Language Models",
    author = "Liu, Xiao  and
      Lei, Xuanyu  and
      Wang, Shengyuan  and
      Huang, Yue  and
      Feng, Andrew  and
      Wen, Bosi  and
      Cheng, Jiale  and
      Ke, Pei  and
      Xu, Yifan  and
      Tam, Weng Lam  and
      Zhang, Xiaohan  and
      Sun, Lichao  and
      Gu, Xiaotao  and
      Wang, Hongning  and
      Zhang, Jing  and
      Huang, Minlie  and
      Dong, Yuxiao  and
      Tang, Jie",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.624/",
    doi = "10.18653/v1/2024.acl-long.624",
    pages = "11621--11640"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/