MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
Ge BaiJie LiuXingyuan BuYancheng HeJiaheng LiuZhanhui ZhouZhuoran LinWenbo SuTiezheng GeBo Zheng
Introduces MT-Bench-101, a three-tier hierarchical benchmark spanning 4,208 dialogue turns across 13 tasks, and demonstrates through testing 21 large language models that standard alignment and chat-specific techniques provide surprisingly limited benefits for multi-turn conversations.
Real-world conversational artificial intelligence relies heavily on multi-turn dialogue, where systems must track history, adapt to feedback, and proactively interact over several exchanges. However, existing evaluation benchmarks primarily focus on single-turn tasks or provide coarse-grained, two-turn assessments that fail to capture the nuances and failure modes of extended human interaction. Without rigorous multi-turn evaluations, organizations risk deploying dialogue systems that appear capable in short tests but degrade rapidly during complex, ongoing conversations.
The article introduces MT-Bench-101, a fine-grained evaluation benchmark designed to systematically assess the capabilities of large language models across multi-turn interactions. By combining real-world dialogue data with frameworks from educational psychology, the article establishes a three-tier hierarchical taxonomy spanning three overarching core competencies (Perceptivity, Adaptability, and Interactivity), seven detailed sub-abilities, and 13 distinct tasks.
To construct the benchmark, researchers generated and manually verified a dataset comprising 1,388 dialogues and 4,208 conversational turns across 30 diverse topic domains. The evaluation evaluated 21 popular language models, including two proprietary and 19 open-source systems. Credibility was ensured by using human-verified historical context to maintain conversational coherence and employing automated evaluation protocols using standardized scoring rubrics. The benchmark applies a minimum-score metric across dialogue rounds, reflecting the reality that a single failed turn compromises the entire conversational flow.
The analysis produced several key findings. First, existing models show widespread proficiency in basic context recall and rephrasing but suffer sharp performance deficits in complex reasoning and proactive questioning, with multi-turn mathematical reasoning proving to be the most challenging task overall. Second, proprietary models consistently outperform open-source alternatives, led by GPT-4 with a top score of 8.86 out of 10, followed by Yi-34B at 8.10. Third, increasing model size reliably enhances conversational ability—especially in proactive questioning—whereas common human preference alignment methods and chat-specific fine-tuning do not yield significant performance gains in multi-turn settings. Finally, the automated evaluation framework achieved an 87% agreement rate with expert human evaluators, exceeding the 80% baseline agreement observed among human experts themselves.
These findings indicate that success on single-turn benchmarks does not translate to robust multi-turn conversational performance. Standard post-training alignment techniques appear to overfit single-turn data while neglecting extended interaction dynamics. For decision-makers, this highlights operational risks in deploying automated agents for interactive, multi-step problem solving without dedicated multi-turn testing.
Organizations developing or deploying conversational agents should prioritize larger model architectures and design alignment pipelines specifically around multi-turn interaction data rather than relying solely on single-turn preference optimization. Evaluators should also adopt fine-grained rubrics and minimum-turn scoring metrics to prevent conversational failures from going unnoticed.
The benchmark’s primary limitation is that conversational capabilities evolve rapidly alongside new model architectures, meaning the current 13 tasks may not capture every emerging interaction pattern. Nevertheless, the high rate of agreement with human raters and the consistent performance rankings across different evaluator models provide strong confidence in the findings.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). MT-Bench-101 directly builds upon and refines the multi-turn dialogue evaluation methodology established by the original MT-Bench.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). Understanding G-Eval provides essential context on using large language models as automated evaluators for natural language generation and dialogue responses.
- Paper: Can Large Language Models Be an Alternative to Human Evaluations?, David Cheng-Han Chiang et al. (2023). This paper investigates whether large language models can reliably replace human judgments, forming the foundation for automated dialogue evaluation frameworks.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). This survey details the taxonomy and landscape of LLM evaluation benchmarks that MT-Bench-101 aims to expand into fine-grained multi-turn dialogue.
- Paper: CoQA: A Conversational Question Answering Challenge, Siva Reddy et al. (2018). CoQA establishes fundamental concepts and challenges in conversational question answering across multi-turn interactions.
- Paper: MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling, Paweł Budzianowski et al. (2018). MultiWOZ introduces standard task definitions and state tracking for multi-domain, multi-turn dialogue systems.
- Paper: How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation, Chia-Wei Liu et al. (2016). This foundational study explores the pitfalls of automated dialogue evaluation metrics, motivating modern LLM-driven benchmarking approaches.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). Extends multi-turn conversational evaluation beyond standard short-horizon benchmarks to very long-term conversational memory spanning hundreds of turns.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). Generalizes fine-grained, instance-specific evaluation using model judges across nine core generation capabilities beyond dialogue.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). Provides a meta-benchmark to evaluate the reliability and precision of LLM judges on complex reasoning and factual tasks.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). Investigates the vulnerabilities and robustness of LLM evaluators when judging complex instruction adherence in generative tasks.
- Paper: From Crowdsourced Data to High-quality Benchmarks: Arena-Hard and Benchbuilder Pipeline, Tianle Li et al. (2025). Builds automated pipelines to extract challenging, high-quality benchmarks from crowdsourced conversational data for robust LLM evaluation.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). Synthesizes the broader paradigm of LLM-as-a-judge frameworks, detailing evaluation structures, biases, and future directions.
- Paper: Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation, Se-eun Yoon et al. (2024). Applies multi-turn conversational simulation using LLMs to interactive recommendation scenarios.
