AlignBench: Benchmarking Chinese Alignment of Large Language Models
Xiao LiuXuanyu LeiShengyuan WangYue HuangAndrew FengBosi WenJiale ChengPei KeYifan XuWeng Lam Tam
Presents ALIGNBENCH, a comprehensive Chinese alignment benchmark spanning eight real-world task categories that pairs human-verified reference evidence with a rule-calibrated, multi-dimensional LLM-as-judge evaluation method to reliably grade open-ended model responses.
As large language models become central to practical digital workflows, aligning them to follow human instructions and satisfy user preferences is essential. While English evaluation frameworks have matured, there has been a critical lack of standardized, challenging, and automated benchmarks specifically designed to measure alignment in Chinese-language applications. To resolve this problem, the article presents ALIGNBENCH, a comprehensive evaluation suite designed to assess how effectively language models fulfill authentic, open-ended Chinese user queries.
The benchmark establishes a curated dataset of 683 challenging tasks derived from real-world usage and researcher-designed problems across eight core categories, including reasoning, mathematics, writing, role play, and advanced Chinese understanding. To evaluate open-ended outputs accurately without incurring prohibitive manual evaluation costs, the authors developed a rule-calibrated, multidimensional evaluation methodology using GPT-4 as an automated judge. Each query is paired with a human-verified reference answer—supported by web citations for knowledge-intensive topics—which acts as an objective baseline to anchor automated grading.
Evaluating 17 prominent open-source and proprietary models on this framework revealed several pivotal findings. First, top-tier frontier models like GPT-4 lead overall performance with an average score of 8.01 out of 10, maintaining substantial leads in factual correctness, complex logic, and mathematical reasoning. Second, leading Chinese commercial models achieved competitive scores near or exceeding 6.0, performing on par with or slightly above GPT-3.5 and demonstrating equal or superior mastery in culturally specific Chinese language tasks. Third, open-source Chinese models showed rapid maturation, with models such as Qwen-14B and Baichuan2-13B rivaling several closed commercial APIs. Finally, the proposed evaluation method demonstrated superior alignment with human judgments, outperforming standard automated judging techniques by 12% to 20% in pairwise quality comparisons.
These findings suggest that while regional developers have closed the gap in linguistic fluency and cultural nuances, severe weaknesses persist in complex logical reasoning and mathematical problem-solving. This divide poses operational risks for organizations deploying models in specialized domains that require absolute factual and computational accuracy. Consequently, developers should prioritize high-quality reasoning and instruction-tuning datasets rather than pure fluency optimization. Meanwhile, benchmark maintainers must expand evaluation coverage to include long-context tasks and integrate dynamic fact-checking mechanisms, as automated judges can still be misled when reference data is incomplete or flawed.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This foundational paper establishes the LLM-as-a-judge evaluation paradigm and examines judgment biases, which AlignBench directly builds upon and adapts for Chinese alignment evaluation.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). This work introduces the framework of using large language models with chain-of-thought prompting for automated text generation evaluation, providing the methodological precursor used in AlignBench.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). This paper presents the ALCE benchmark for evaluating knowledge generation with retrieved web evidence and citations, providing foundational techniques for AlignBench's evidence-supported factual evaluation.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). LongBench established standardized bilingual (English-Chinese) open-ended multi-task evaluation methodologies from the same research group, setting essential precedent for AlignBench.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). This survey provides a comprehensive taxonomy of large language model evaluation tasks and automated judging procedures that framed the gap AlignBench targets.
- Paper: ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools, Team Glm Aohan Zeng et al. (2024). This technical report evaluates the ChatGLM/GLM-4 model family, utilizing AlignBench as a core benchmark to measure and demonstrate frontier Chinese alignment performance.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). JudgeBench critically investigates and benchmarks the reasoning capabilities and reliability of the LLM judges that benchmarks like AlignBench rely upon for automated scoring.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). This study advances the meta-evaluation of automated LLM evaluators on instruction-following adherence, diagnosing vulnerabilities in the judge architectures deployed in AlignBench.
- Paper: Split and Merge: Aligning Position Biases in LLM-based Evaluators, Zongjie Li et al. (2024). This paper develops segment-and-merge prompting strategies to systematically calibrate and mitigate the position biases present in LLM-based evaluators like those used in AlignBench.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This comprehensive survey categorizes the broader landscape and emerging techniques of the LLM-as-a-judge paradigm exemplified by AlignBench's rule-calibrated evaluation.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). BiGGen Bench extends multi-dimensional LLM-based evaluation across broad model capabilities using instance-specific rubrics similar to the calibrated judging pipeline in AlignBench.
