MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
Wai-Chung KwanXingshan ZengYuxin JiangYufei WangLiangyou LiLifeng ShangXin JiangQun LiuKam-Fai Wong
Presents MT-Eval, a benchmark categorizing multi-turn conversations into recollection, expansion, refinement, and follow-up patterns to systematically evaluate and diagnose why large language models degrade in extended dialogues compared to single-turn interactions.
Large language models are increasingly deployed as interactive conversational assistants to handle multi-step tasks such as drafting text, document processing, and strategic analysis. While real-world applications rely heavily on extended dialogues where user needs evolve, standard benchmarks predominantly test models on isolated single-turn prompts or brief two-turn interactions. This evaluation gap creates a major operational blind spot, leaving organizations unable to reliably measure how well systems maintain instructions and context throughout continuous engagements.
The article introduces MT-Eval, a comprehensive evaluation benchmark designed to assess the multi-turn conversational capabilities of large language models across realistic interaction patterns. It systematically measures model response quality during extended conversations and directly quantifies performance degradation by comparing multi-turn results against identical single-turn baselines.
The researchers analyzed human-AI interactions to establish four distinct dialogue categories: recollection of earlier instructions, topic expansion, iterative instruction refinement, and contextual follow-ups. Using GPT-4 with extensive human-in-the-loop verification to avoid data contamination, the authors created 168 dialogue sessions comprising 1,170 turns with an average of roughly seven turns per conversation. They evaluated 10 prominent open-source and proprietary models, utilizing GPT-4 with chain-of-thought prompting alongside rule-based scoring and human validation to grade responses.
The findings show that current language models experience severe performance degradation in multi-turn settings compared to single-turn environments. On a 10-point scale, proprietary models led the overall multi-turn rankings with GPT-3.5-Turbo achieving an average score of 7.72, though leading open-source models like Mistral-Instruct-7B (7.46) and Mixtral-Instruct-8x7B (7.47) demonstrated competitive or superior results on specific subtasks such as follow-up queries. Crucially, a model's single-turn capability does not predict multi-turn reliability; for example, Llama-2-chat models performed well in single-turn tests but suffered drops of roughly 1.9 to 2.1 points in multi-turn dialogues. In-depth analysis revealed that roughly half of multi-turn failures stemmed from noncompliance as the distance to initial instructions increased, while another 48.8% resulted from error propagation, where an early mistake compounded across subsequent turns.
These results demonstrate that standard single-turn benchmark scores misrepresent actual operational readiness for conversational deployments. In production environments, degraded instruction-following and compounding errors increase operational risk, raise compliance concerns, and necessitate manual oversight. Deploying conversational agents under the assumption that high base capabilities guarantee reliable multi-turn adherence introduces significant functional vulnerabilities.
Organizations evaluating language models should mandate multi-turn benchmark assessments rather than relying on single-turn metrics. To mitigate performance decay, system designers should implement architectural safeguards, such as context management strategies that bring critical global instructions into nearer conversational context and validation mechanisms that detect and correct early errors before they propagate through the dialogue history.
The conclusions should be interpreted with awareness that the evaluation relied primarily on automated GPT-4 scoring, which human validation confirmed as well-aligned (0.71 Pearson correlation) but exhibiting a measurable leniency bias toward its own generated text. Additionally, computational constraints limited the study to models up to 14 billion parameters alongside select mixture-of-experts architectures, leaving the exact multi-turn degradation curve of larger open-source variants as an area for further validation.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). MT-Bench establishes an influential multi-turn dialogue benchmark and LLM-as-a-judge approach, providing useful context for MT-Eval’s evaluation design and scoring choices.
- Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). This later benchmark extends multi-turn evaluation to conversations where users progressively reveal, revise, or switch their intent, testing a dynamic challenge beyond MT-Eval’s dialogue categories.
- Paper: Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction, Ryo Kamoi et al. (2026). COCOEVAL continues turn-level conversation assessment by measuring whether generated dialogues reproduce realistic inconsistencies, misunderstandings, and interruptions.
