MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Ge BaiJie LiuXingyuan BuYancheng HeJiaheng LiuZhanhui ZhouZhuoran LinWenbo SuTiezheng GeBo Zheng

article2024ACL241 citations

Introduces MT-Bench-101, a three-tier hierarchical benchmark spanning 4,208 dialogue turns across 13 tasks, and demonstrates through testing 21 large language models that standard alignment and chat-specific techniques provide surprisingly limited benefits for multi-turn conversations.

Listen

Real-world conversational artificial intelligence relies heavily on multi-turn dialogue, where systems must track history, adapt to feedback, and proactively interact over several exchanges. However, existing evaluation benchmarks primarily focus on single-turn tasks or provide coarse-grained, two-turn assessments that fail to capture the nuances and failure modes of extended human interaction. Without rigorous multi-turn evaluations, organizations risk deploying dialogue systems that appear capable in short tests but degrade rapidly during complex, ongoing conversations.

The article introduces MT-Bench-101, a fine-grained evaluation benchmark designed to systematically assess the capabilities of large language models across multi-turn interactions. By combining real-world dialogue data with frameworks from educational psychology, the article establishes a three-tier hierarchical taxonomy spanning three overarching core competencies (Perceptivity, Adaptability, and Interactivity), seven detailed sub-abilities, and 13 distinct tasks.

To construct the benchmark, researchers generated and manually verified a dataset comprising 1,388 dialogues and 4,208 conversational turns across 30 diverse topic domains. The evaluation evaluated 21 popular language models, including two proprietary and 19 open-source systems. Credibility was ensured by using human-verified historical context to maintain conversational coherence and employing automated evaluation protocols using standardized scoring rubrics. The benchmark applies a minimum-score metric across dialogue rounds, reflecting the reality that a single failed turn compromises the entire conversational flow.

The analysis produced several key findings. First, existing models show widespread proficiency in basic context recall and rephrasing but suffer sharp performance deficits in complex reasoning and proactive questioning, with multi-turn mathematical reasoning proving to be the most challenging task overall. Second, proprietary models consistently outperform open-source alternatives, led by GPT-4 with a top score of 8.86 out of 10, followed by Yi-34B at 8.10. Third, increasing model size reliably enhances conversational ability—especially in proactive questioning—whereas common human preference alignment methods and chat-specific fine-tuning do not yield significant performance gains in multi-turn settings. Finally, the automated evaluation framework achieved an 87% agreement rate with expert human evaluators, exceeding the 80% baseline agreement observed among human experts themselves.

These findings indicate that success on single-turn benchmarks does not translate to robust multi-turn conversational performance. Standard post-training alignment techniques appear to overfit single-turn data while neglecting extended interaction dynamics. For decision-makers, this highlights operational risks in deploying automated agents for interactive, multi-step problem solving without dedicated multi-turn testing.

Organizations developing or deploying conversational agents should prioritize larger model architectures and design alignment pipelines specifically around multi-turn interaction data rather than relying solely on single-turn preference optimization. Evaluators should also adopt fine-grained rubrics and minimum-turn scoring metrics to prevent conversational failures from going unnoticed.

The benchmark’s primary limitation is that conversational capabilities evolve rapidly alongside new model architectures, meaning the current 13 tasks may not capture every emerging interaction pattern. Nevertheless, the high rate of agreement with human raters and the consistent performance rankings across different evaluator models provide strong confidence in the findings.

Cover for MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Abstract

The advent of Large Language Models (LLMs) has drastically enhanced dialogue systems. However, comprehensively evaluating the dialogue abilities of LLMs remains a challenge. Previous benchmarks have primarily focused on single-turn dialogues or provided coarse-grained and incomplete assessments of multi-turn dialogues, overlooking the complexity and fine-grained nuances of real-life dialogues. To address this issue, we introduce MT-Bench-101, specifically designed to evaluate the fine-grained abilities of LLMs in multi-turn dialogues. By conducting a detailed analysis of real multi-turn dialogue data, we construct a three-tier hierarchical ability taxonomy comprising 4208 turns across 1388 multi-turn dialogues in 13 distinct tasks. We then evaluate 21 popular LLMs based on MT-Bench-101, conducting comprehensive analyses from both ability and task perspectives and observing differing trends in LLMs performance across dialogue turns within various tasks. Further analysis indicates that neither utilizing common alignment techniques nor chat-specific designs has led to obvious enhancements in the multi-turn abilities of LLMs. Extensive case studies suggest that our designed tasks accurately assess the corresponding multi-turn abilities. The data and code are available at \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 MT-Bench-101
  • 3.1 Hierarchical Ability Taxonomy
  • 3.1.1 Perceptivity
  • 3.1.2 Adaptability
  • 3.1.3 Interactivity
  • 3.2 Data Collection
  • 3.3 Data Statistics
  • 3.4 Evaluation
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Main Results
  • 4.3 Further Analysis
  • 4.4 Case Study
  • 4.5 Human Evaluation
  • 5 Conclusion
  • 6 Limitations
  • 7 Ethics Statement
  • References
  • A Details on the Data Generation
  • B Detials on Data Statistics
  • C Details on Evaluation
  • D Additional results utilizing Qwen-72B-Chat
  • E Model Details
  • F More Cases
  • G Fleiss’ Kappa between GPT-4 and Humans

Knowls

  1. Knowl 1 — MT-Bench-101 Three-Tier Ability Taxonomy for Multi-Turn Dialogues

    definition

    MT-Bench-101 categorizes multi-turn dialogue capabilities of Large Language Models (LLMs) into a three-tier hierarchical taxonomy combining empirical human-LLM conversational data (from ShareGPT and RealChat) with dialogic teaching principles from educational psychology:

    1. Tier 1: Overarching Abilities

      • Perceptivity: The fundamental ability of a model to accurately track, comprehend, and utilize dialogue history to deliver consistent and logical responses.
      • Adaptability: The model's capacity to adjust its responses according to newly introduced user requirements, hypotheses, or critical feedback.
      • Interactivity: The proactive conversational capability where the model guides the interaction, solicits necessary details, and sustains user engagement.
    2. Tier 2: Detailed Abilities & Tier 3: Specific Tasks (13 Distinct Tasks)

      • Under Perceptivity:
        • Context Memory: Context Memory (CM) – Recalling and integrating specific information provided in earlier turns to answer the user's current query without repeating previous context.
        • Context Understanding: Anaphora Resolution (AR) – Correctly resolving demonstrative pronouns and referential expressions back to earlier dialogue content; Separate Input (SI) – Receiving high-level task instructions in the initial turn and processing specific inputs provided in subsequent turns without repeating instructions.
        • Context Interference: Topic Shift (TS) – Seamlessly transitioning when a user abruptly changes topics without being confused by previous dialogue; Content Confusion (CC) – Discerning queries that share structural or lexical similarity with prior questions but require completely distinct responses.
      • Under Adaptability:
        • Rephrasing: Content Rephrasing (CR) – Rephrasing prior response content according to new user constraints while preserving original facts; Format Rephrasing (FR) – Altering structural formatting (e.g., converting prose into a bulleted list) while preserving content.
        • Reflection: Self-Correction (SC) – Correcting erroneous or imprecise previous responses when valid user critique is received; Self-Affirmation (SA) – Maintaining and justifying correct previous responses when challenged by incorrect user critique.
        • Reasoning: Mathematical Reasoning (MR) – Incrementally solving complex mathematical problems as users introduce new conditions across turns; General Reasoning (GR) – Collaborative multi-turn logical deduction, induction, and puzzle-solving.
      • Under Interactivity:
        • Questioning: Instruction Clarification (IC) – Asking targeted counter-questions when initial user prompts are ambiguous or incomplete; Proactive Interaction (PI) – Proposing relevant follow-up questions or conversational prompts to encourage continued dialogue.
  2. Knowl 2 — Golden Context and Minimum-Score Evaluation Protocol for Multi-Turn Dialogues

    model/method

    Evaluating multi-turn dialogue models requires addressing two major methodological pitfalls: error propagation from model-generated conversational history and score inflation caused by learning in-context response styles across turns. MT-Bench-101 establishes a two-pronged evaluation protocol:

    1. Golden Context Conditioning: When evaluating turn tt of a dialogue, the model is fed the curated ground-truth ("golden") dialogue history up to turn t−1t-1 along with the current user prompt, rather than the model's own self-generated past responses. This ensures that response quality at turn tt is evaluated under clean, coherent dialogue conditions without cumulative compounding errors, and provides fair comparisons across models.
    2. Minimum-Score Aggregation (min\\min-Score Metric): Each turn's response is scored on a 1-to-10 scale by an automated LLM judge (such as GPT-4) using task-specific rubric guidelines. The dialogue-level overall score SdialogueS_{\text{dialogue}} is defined as the minimum score across all TT turns of the dialogue: Sdialogue=min⁡t=1,…,TstS_{\text{dialogue}} = \min_{t=1,\dots,T} s_t where st∈[1,10]s_t \in [1, 10] is the evaluation score at turn tt. Taking the minimum reflects the conversational principle that a single critical breakdown (such as yielding to incorrect criticism or failing to resolve an anaphor) invalidates the coherence and utility of the entire multi-turn interaction.
  3. Knowl 3 — MT-Bench-101 Dataset Structure and Task Statistics

    data/table

    The MT-Bench-101 benchmark contains 1,388 multi-turn dialogues comprising 4,208 turns across 13 fine-grained tasks and 30 topic domains (such as medicine, law, history, science, digital technology, finance, humanities, and arts). Each candidate dialogue generated by GPT-4 underwent screening by five human annotators, retaining only data with unanimous human approval.

    Task Name Abbr. Dialogues Turns Avg Turns/Dial Avg Words/Dial Avg Words/Turn
    Context Memory CM 80 319 3.99 235.9 59.15
    Separate Input SI 149 620 4.16 214.0 84.76
    Anaphora Resolution AR 153 560 3.66 214.1 58.49
    Topic Shift TS 83 249 3.00 145.7 48.56
    Content Confusion CC 147 352 2.39 353.9 147.80
    Content Rephrasing CR 136 389 2.86 321.4 112.40
    Format Rephrasing FR 74 197 2.66 191.4 71.91
    Self-Correction SC 77 154 2.00 79.2 39.60
    Self-Affirmation SA 73 146 2.00 85.4 42.71
    Mathematical Reasoning MR 108 224 2.07 175.8 51.44
    General Reasoning GR 71 218 3.07 157.3 51.22
    Instruction Clarification IC 150 426 2.84 142.5 50.18
    Proactive Interaction PI 87 354 4.07 127.1 31.22
    Overall – 1388 4208 3.03 202.0 66.64
  4. Knowl 4 — LLM Performance Benchmarking Across MT-Bench-101 Tasks

    data/table

    Evaluation scores for 21 LLMs evaluated at temperature 0 on MT-Bench-101 using GPT-4 judge scoring (1–10 scale) under the minimum-score dialogue aggregation metric. Closed-source models consistently outperform open-source counterparts, with GPT-4 achieving the highest score across all tasks (8.86), followed by Yi-34B-Chat (8.10).

    Model Avg. CM SI AR TS CC CR FR SC SA MR GR IC PI
    Llama2-7B-Chat 6.53 7.64 6.21 7.92 8.23 8.50 8.32 8.56 8.45 4.97 1.88 3.83 5.23 5.11
    Qwen-7B-Chat 7.12 7.65 7.75 8.73 8.42 8.76 8.89 9.16 8.49 7.28 2.25 3.57 5.41 6.24
    ChatGLM2-6B 5.56 6.14 4.69 7.27 6.13 6.26 7.47 7.98 6.97 4.19 2.11 3.00 5.16 4.90
    ChatGLM3-6B 6.47 7.16 5.42 8.21 7.43 8.03 8.38 8.81 7.40 5.63 2.60 3.21 6.19 5.61
    InternLM2-Chat-7B-SFT 6.69 7.51 6.26 8.01 8.06 8.70 8.50 8.50 7.68 6.16 3.47 4.48 4.92 4.76
    Yi-6B-Chat 6.93 7.57 5.27 8.69 8.37 8.76 8.43 8.44 7.49 7.85 2.18 3.80 7.30 6.00
    Mistral-7B-Instruct-v0.2 6.95 7.66 5.64 8.09 8.30 9.35 8.69 8.59 8.16 7.33 2.58 4.52 5.80 5.66
    Vicuna-13B-v1.5 6.37 7.06 5.62 7.81 7.45 8.79 7.96 7.72 7.47 6.70 2.31 4.03 5.05 4.80
    Baize-13B-v2 6.12 6.78 5.15 7.86 7.40 8.07 7.96 8.15 7.24 6.32 1.67 3.69 4.35 4.95
    UltraLM-13B-v2.0 4.61 4.66 4.89 5.99 6.49 8.48 2.87 2.53 6.70 5.27 1.46 2.34 4.13 4.11
    Llama2-13B-Chat 7.15 8.03 7.11 9.00 9.39 8.81 9.07 9.11 7.63 7.60 1.75 3.16 6.07 6.23
    Qwen-14B-Chat 7.82 8.33 8.36 9.04 9.22 9.50 9.12 9.39 8.41 7.97 3.50 4.55 8.21 6.12
    Baichuan2-13B-Chat 7.00 7.71 6.38 8.92 8.36 9.07 9.10 8.95 7.75 6.57 2.50 3.65 6.95 5.15
    InternLM2-Chat-20B-SFT 6.95 7.35 6.44 8.08 8.05 9.10 8.59 8.55 7.62 7.36 4.05 5.24 4.99 4.99
    Yi-34B-Chat 8.10 8.55 6.79 9.34 9.84 9.34 9.08 9.38 9.01 9.04 4.07 5.90 8.51 6.39
    Mixtral-8x7B-Instruct-v0.1 7.38 7.86 5.94 8.49 9.01 9.52 8.91 9.01 8.69 7.78 4.19 5.14 6.03 5.36
    GPT-3.5 7.99 8.77 7.67 7.67 9.68 9.87 9.56 9.51 9.18 7.23 4.48 5.31 8.57 6.32
    GPT-4 8.86 8.88 8.99 9.58 9.83 9.98 9.54 9.57 9.36 9.52 7.15 7.17 9.00 6.64
    Task Average 6.92 7.52 6.37 8.26 7.72 8.24 8.36 8.44 7.98 6.93 3.61 4.84 6.22 5.52
  5. Knowl 5 — Cross-Ability Strengths and Weaknesses of LLMs in Multi-Turn Dialogues

    empirical result

    Aggregating MT-Bench-101 task scores into the seven ability dimensions highlights systemic model capabilities and deficiencies:

    1. Core Strengths: LLMs achieve their highest average scores in Rephrasing (Content Rephrasing at 8.36, Format Rephrasing at 8.44) and Context Interference resistance (Topic Shift at 7.72, Content Confusion at 8.24).
    2. Major Deficiencies: LLMs struggle most severely on Reasoning (Mathematical Reasoning at 3.61, General Reasoning at 4.84) and Questioning/Interactivity (Instruction Clarification at 6.22, Proactive Interaction at 5.52).
    3. Memory vs. Understanding Asymmetry: Models perform significantly better in Context Memory (average 7.52) than in Context Understanding (Separate Input at 6.37). This occurs because Context Memory requires surface-level retrieval of previous dialogue statements, whereas Context Understanding demands semantic parsing across decoupled instructions and inputs.
  6. Knowl 6 — Effect of Human Preference Alignment (RLHF/DPO) on Multi-Turn Dialogue Abilities

    empirical result

    Comparing pairs of open-source models trained with Supervised Fine-Tuning (SFT) versus those further aligned using Reinforcement Learning from Human Feedback (RLHF) or Direct Preference Optimization (DPO) demonstrates that current preference alignment yields marginal or inconsistent effects on multi-turn dialogue abilities:

    Model SFT RLHF / DPO MT-Bench-101 Avg. Δ\Delta
    InternLM2-Chat-7B ✓ 6.69 –
    InternLM2-Chat-7B ✓ 6.85 +0.16
    InternLM2-Chat-20B ✓ 6.95 –
    InternLM2-Chat-20B ✓ 7.05 +0.10
    Mistral-7B ✓ 6.95 –
    Mistral-7B ✓ 6.89 -0.06

    Unlike single-turn benchmarks where RLHF and DPO yield substantial gains, multi-turn performance changes negligibly (+0.16 for InternLM2-Chat-7B, +0.10 for InternLM2-Chat-20B, and -0.06 for Mistral-7B). This indicates that existing preference tuning sets, which are dominated by single-turn prompt-response pairs, fail to transfer to multi-turn reflection, tracking, and proactive interaction.

  7. Knowl 7 — Performance of Conversational-Specialized LLMs on Multi-Turn Tasks

    empirical result

    LLMs specifically developed for chat and multi-turn conversations through synthetic self-chat fine-tuning (such as Baize-v2-13B and UltraLM-13B-v2.0) underperform general-purpose instruction-tuned models of comparable parameter scale:

    • UltraLM-13B-v2.0 achieves an overall average score of 4.61, failing severely on Content Rephrasing (2.87) and Format Rephrasing (2.53).
    • Baize-v2-13B achieves an overall average score of 6.12.
    • In contrast, standard general instruction-tuned models in the 13B–14B class achieve significantly higher overall scores: Qwen-14B-Chat achieves 7.82, Llama2-13B-Chat achieves 7.15, Baichuan2-13B-Chat achieves 7.00, and Vicuna-13B-v1.5 achieves 6.37.

    This shows that synthetic conversational self-chat preconditioning does not provide general advantages across fine-grained conversational tasks like strict formatting, reflection, or multi-turn reasoning.

  8. Knowl 8 — Per-Turn Score Trajectories and Impact of Golden vs. Self-Predicted Context

    empirical result

    Tracking turn-by-turn performance across tasks reveals distinct operational dynamics:

    1. Performance Decay over Turns: In Context Memory, Anaphora Resolution, and Rephrasing tasks, average scores drop between Turn 1 and subsequent turns as models exhibit context forgetting and comprehension drift. In Topic Shift and Content Confusion tasks, scores experience a sharp drop from Turn 1 to Turn 2 due to the sudden introduction of contextual interference.
    2. In-Context Style Learning vs. Error Cascade:
      • Providing Golden Context in tasks with strong procedural structure (Separate Input, Instruction Clarification, Proactive Interaction, and Mathematical Reasoning) produces an upward score trajectory across turns because the golden dialogue history serves as in-context demonstrations of the expected communication style and step-by-step reasoning formats.
      • Providing Self-Predicted Context (where the model converses with its own prior outputs) causes scores to steadily decrease across turns due to compounding errors and loss of dialogue coherence.
  9. Knowl 9 — Inter-Rater Reliability and Human Agreement of MT-Bench-101 Auto-Evaluation

    empirical result

    Automated evaluation using GPT-4 with task-specific prompts, grading rubrics, and the minimum-turn scoring rule aligns closely with human expert ratings evaluated on a sample of 100 multi-turn dialogues across 5 human experts:

    Evaluation Method Human Agreement Δ\Delta from Baseline
    Human Experts (Internal Consensus) 80% 0%
    MT-Bench-101 (GPT-4 Judge + Rubric + Min-Score) 87% +7%
    w/o scoring guidelines 77% -3%
    w/o minimum values metric (using average score) 75% -5%

    Inter-rater reliability measured via Fleiss' Kappa (κ\kappa) further confirms consistency:

    • Fleiss' Kappa among 5 human annotators: κ=0.672\kappa = 0.672
    • Fleiss' Kappa of GPT-4 vs. humans (over all individual annotations): κ=0.676\kappa = 0.676
    • Fleiss' Kappa of GPT-4 vs. individual human raters (average): κ=0.681\kappa = 0.681
    • Fleiss' Kappa of GPT-4 vs. human majority vote: κ=0.699\kappa = 0.699

    Removing the detailed scoring guidelines drops agreement to 77%, and using average-turn scores instead of minimum-turn scores drops agreement to 75%.

Coverage note — Detailed per-task prompt templates and qualitative conversation transcripts shown in the appendices were summarized into their respective task definitions, data generation rules, and evaluation rubrics rather than extracted as individual prompt knowls.

References

  1. 1.Robin Alexander. 2018. Developing dialogic teaching: Genesis, process, trial. Research papers in education, 33(5):561–598.
  2. 2.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  3. 3.Baichuan. 2023. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Xingyuan Bu, Junran Peng, Junjie Yan, Tieniu Tan, and Zhaoxiang Zhang. 2021. Gaia: A transfer learning system of object detection that fits your needs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 274–283.
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An opensource chatbot impressing gpt-4 with 90%* chatgpt quality.
  7. 7.Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023. Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:2305.14233.
  8. 8.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335.
  9. 9.Haodong Duan, Jueqi Wei, Chonghua Wang, Hongwei Liu, Yixiao Fang, Songyang Zhang, Dahua Lin, and Kai Chen. 2023. Botchat: Evaluating llms’ capabilities of having multi-turn dialogues. arXiv preprint arXiv:2310.13650.
  10. 10.Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. 2024. Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems, 36.
  11. 11.Weixin Feng, Xingyuan Bu, Chenchen Zhang, and Xubin Li. 2022. Beyond bounding box: Multi-modal knowledge learning for object detection. arXiv preprint arXiv:2205.04072.
  12. 12.Sarah E Finch, James D Finch, and Jinho D Choi. 2022. Don’t forget your abc’s: Evaluating the state-of-the-art in chat-oriented dialogue systems. arXiv preprint arXiv:2212.09180.
  13. 13.Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata. 2023. Improving language model negotiation with self-play and in-context learning from ai feedback. arXiv preprint arXiv:2305.10142.
  14. 14.Yuan Gao, Xingyuan Bu, Yang Hu, Hui Shen, Ti Bai, Xubin Li, and Shilei Wen. 2018. Solution for largescale hierarchical object detection datasets with incomplete annotation and data imbalance. arXiv preprint arXiv:1810.06208.
  15. 15.Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. The false promise of imitating proprietary llms. arXiv preprint arXiv:2305.15717.
  16. 16.Hongcheng Guo, Jian Yang, Jiaheng Liu, Liqun Yang, Linzheng Chai, Jiaqi Bai, Junran Peng, Xiaorong Hu, Chao Chen, Dongfeng Zhang, et al. 2023. Owl: A large language model for it operations. arXiv preprint arXiv:2309.09298.
  17. 17.Jiawei Guo, Ziming Li, Xueling Liu, Kaijing Ma, Tianyu Zheng, Zhouliang Yu, Ding Pan, Yizhi Li, Ruibo Liu, Yue Wang, et al. 2024a. Codeeditorbench: Evaluating code editing capability of large language models. arXiv preprint arXiv:2404.03543.
  18. 18.Jinyang Guo, Jianyu Wu, Zining Wang, Jiaheng Liu, Ge Yang, Yifu Ding, Ruihao Gong, Haotong Qin, and Xianglong Liu. 2024b. Compressing large language models by joint sparsification and quantization. ICML.
  19. 19.Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, and Yulia Tsvetkov. 2022. On the blind spots of model-based evaluation metrics for text generation. arXiv preprint arXiv:2212.10020.
  20. 20.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
  21. 21.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Yao Fu, et al. 2024. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. Advances in Neural Information Processing Systems, 36.
  22. 22.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023a. Mistral 7b. arXiv preprint arXiv:2310.06825.
  23. 23.Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, and Wei Wang. 2023b. Followbench: A multi-level fine-grained constraints following benchmark for large language models. arXiv preprint arXiv:2310.20410.
  24. 24.Tian Lan, Xian-Ling Mao, Wei Wei, and Heyan Huang. 2020. Which kind is better in open-domain multi-turn dialog, hierarchical or non-hierarchical models? an empirical study. arXiv preprint arXiv:2008.02964.
  25. 25.Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2023a. Cmmlu: Measuring massive multitask language understanding in chinese. arXiv preprint arXiv:2306.09212.
  26. 26.Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024a. From live data to high-quality benchmarks: The arena-hard pipeline.
  27. 27.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023b. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval.
  28. 28.Yizhi Li, Ge Zhang, Xingwei Qu, Jiali Li, Zhaoqun Li, Zekun Wang, Hao Li, Ruibin Yuan, Yinghao Ma, Kai Zhang, et al. 2024b. Cif-bench: A chinese instruction-following benchmark for evaluating the generalizability of large language models. arXiv preprint arXiv:2402.13109.
  29. 29.Jiaheng Liu, Zhiqi Bai, Yuanxing Zhang, Chenchen Zhang, Yu Zhang, Ge Zhang, Jiakai Wang, Haoran Que, Yukang Chen, Wenbo Su, et al. 2024a. E2-llm: Efficient and extreme length extension of large language models. arXiv preprint arXiv:2401.06951.
  30. 30.Jie Liu, Zhanhui Zhou, Chao Yang, Han-Sen Zhong, and Wanli Ouyang. 2024b. Storm-7b: An empirical study of iterative direct preference optimization.
  31. 31.Weimin Lyu, Xiao Lin, Songzhu Zheng, Lu Pang, Haibin Ling, Susmit Jha, and Chao Chen. 2024. Task-agnostic detector for insertion-based backdoor attacks. arXiv preprint arXiv:2403.17155.
  32. 32.Carol A Marchel. 2007. Learning to talk/talking to learn: Teaching critical dialogue. Teaching Educational Psychology, 2(1):1–15.
  33. 33.OpenAI. 2022. Introducing chatgpt.
  34. 34.OpenAI. 2023. Gpt-4 technical report. PREPRINT.
  35. 35.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  36. 36.Junran Peng, Xingyuan Bu, Ming Sun, Zhaoxiang Zhang, Tieniu Tan, and Junjie Yan. 2020. Large-scale object detection in the wild from imbalanced multi-labels. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9709–9718.
  37. 37.Junran Peng, Qing Chang, Haoran Yin, Xingyuan Bu, Jiajun Sun, Lingxi Xie, Xiaopeng Zhang, Qi Tian, and Zhaoxiang Zhang. 2023. Gaia-universe: Everything is super-netify. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10):11856–11868.
  38. 38.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36.
  39. 39.William A Scott. 1955. Reliability of content analysis: The case of nominal scale coding. Public opinion quarterly, pages 321–325.
  40. 40.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615.
  41. 41.Tao Sun, Linzheng Chai, Yuwei Yin Jian Yang, Hongcheng Guo, Jiaheng Liu, Bing Wang, Liqun Yang, and Zhoujun Li. 2024. Unicoder: Scaling code large language model via universal code. ACL.
  42. 42.Yuchong Sun, Che Liu, Jinwen Huang, Ruihua Song, Fuzheng Zhang, Di Zhang, Zhongyuan Wang, and Kun Gai. 2023. Parrot: Enhancing multi-turn chat models by learning to ask questions. arXiv preprint arXiv:2310.07301.
  43. 43.InternLM Team. 2023. Internlm: A multilingual language model with progressively enhanced capabilities. https://github.com/InternLM/InternLM.
  44. 44.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  45. 45.Hongru Wang, Rui Wang, Fei Mi, Yang Deng, Zezhong Wang, Bin Liang, Ruifeng Xu, and Kam-Fai Wong. 2023a. Cue-cot: Chain-of-thought prompting for responding to in-depth dialogue questions with llms. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12047–12064.
  46. 46.Xingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng, and Heng Ji. 2023b. Mint: Evaluating llms in multi-turn interaction with tools and language feedback. arXiv preprint arXiv:2309.10691.
  47. 47.Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, et al. 2023c. Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization. arXiv preprint arXiv:2306.05087.
  48. 48.Zekun Moore Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wenhu Chen, Jie Fu, and Junran Peng. 2023d. Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models. arXiv preprint arXiv: 2310.00746.
  49. 49.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  50. 50.Yanan Wu, Jie Liu, Xingyuan Bu, Jiaheng Liu, Zhanhui Zhou, Yuanxing Zhang, Chenchen Zhang, Zhiqi Bai, Haibin Chen, Tiezheng Ge, et al. 2024. Conceptmath: A bilingual concept-wise benchmark for measuring mathematical reasoning of large language models. arXiv preprint arXiv:2402.14660.
  51. 51.Canwen Xu, Daya Guo, Nan Duan, and Julian McAuley. 2023. Baize: An open-source chat model with parameter-efficient tuning on self-chat data. arXiv preprint arXiv:2304.01196.
  52. 52.Yi. 2023. Yi: Building the next generation of open-source and bilingual llms. https://github.com/01-ai/Yi.
  53. 53.Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-Li, Xin Lv, Hao Peng, Zijun Yao, Xiaohan Zhang, Hanming Li, et al. 2023. Kola: Carefully benchmarking world knowledge of large language models. arXiv preprint arXiv:2306.09296.
  54. 54.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric Xing, et al. 2023. Lmsys-chat-1m: A large-scale real-world llm conversation dataset. arXiv preprint arXiv:2309.11998.
  55. 55.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36.
  56. 56.Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023. Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364.
  57. 57.Zhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu, Chao Yang, Wanli Ouyang, and Yu Qiao. 2024a. Emulated disalignment: Safety alignment for large language models may backfire! arXiv preprint arXiv:2402.12343.
  58. 58.Zhanhui Zhou, Jie Liu, Chao Yang, Jing Shao, Yu Liu, Xiangyu Yue, Wanli Ouyang, and Yu Qiao. 2023. Beyond one-preference-for-all: Multi-objective direct preference optimization. arXiv preprint arXiv:2310.03708.
  59. 59.Zhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong, Chao Yang, and Yu Qiao. 2024b. Weak-to-strong search: Align large language models via searching over small language models. arXiv preprint arXiv:2405.19262.

Citation

MLA
Bai, G., et al. “MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7421–54, https://doi.org/10.18653/v1/2024.acl-long.401.
APA
Bai, G., Liu, J., Bu, X., He, Y., Liu, J., Zhou, Z., Lin, Z., Su, W., Ge, T., Zheng, B., & Ouyang, W. (2024). MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7421–7454. https://doi.org/10.18653/v1/2024.acl-long.401
Chicago
Bai, G., J. Liu, X. Bu, et al. 2024. “MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7421–54. https://doi.org/10.18653/v1/2024.acl-long.401.
Harvard
Bai, G. et al. (2024) “MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7421–7454. Available at: https://doi.org/10.18653/v1/2024.acl-long.401.
Vancouver
1. Bai G, Liu J, Bu X, et al (2024) MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7421–7454

BibTeX

@inproceedings{bai-etal-2024-mt,
    title = "{MT}-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues",
    author = "Bai, Ge  and
      Liu, Jie  and
      Bu, Xingyuan  and
      He, Yancheng  and
      Liu, Jiaheng  and
      Zhou, Zhanhui  and
      Lin, Zhuoran  and
      Su, Wenbo  and
      Ge, Tiezheng  and
      Zheng, Bo  and
      Ouyang, Wanli",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.401/",
    doi = "10.18653/v1/2024.acl-long.401",
    pages = "7421--7454"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/