MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Wai-Chung KwanXingshan ZengYuxin JiangYufei WangLiangyou LiLifeng ShangXin JiangQun LiuKam-Fai Wong

article2024EMNLP98 citations

Presents MT-Eval, a benchmark categorizing multi-turn conversations into recollection, expansion, refinement, and follow-up patterns to systematically evaluate and diagnose why large language models degrade in extended dialogues compared to single-turn interactions.

Listen

Large language models are increasingly deployed as interactive conversational assistants to handle multi-step tasks such as drafting text, document processing, and strategic analysis. While real-world applications rely heavily on extended dialogues where user needs evolve, standard benchmarks predominantly test models on isolated single-turn prompts or brief two-turn interactions. This evaluation gap creates a major operational blind spot, leaving organizations unable to reliably measure how well systems maintain instructions and context throughout continuous engagements.

The article introduces MT-Eval, a comprehensive evaluation benchmark designed to assess the multi-turn conversational capabilities of large language models across realistic interaction patterns. It systematically measures model response quality during extended conversations and directly quantifies performance degradation by comparing multi-turn results against identical single-turn baselines.

The researchers analyzed human-AI interactions to establish four distinct dialogue categories: recollection of earlier instructions, topic expansion, iterative instruction refinement, and contextual follow-ups. Using GPT-4 with extensive human-in-the-loop verification to avoid data contamination, the authors created 168 dialogue sessions comprising 1,170 turns with an average of roughly seven turns per conversation. They evaluated 10 prominent open-source and proprietary models, utilizing GPT-4 with chain-of-thought prompting alongside rule-based scoring and human validation to grade responses.

The findings show that current language models experience severe performance degradation in multi-turn settings compared to single-turn environments. On a 10-point scale, proprietary models led the overall multi-turn rankings with GPT-3.5-Turbo achieving an average score of 7.72, though leading open-source models like Mistral-Instruct-7B (7.46) and Mixtral-Instruct-8x7B (7.47) demonstrated competitive or superior results on specific subtasks such as follow-up queries. Crucially, a model's single-turn capability does not predict multi-turn reliability; for example, Llama-2-chat models performed well in single-turn tests but suffered drops of roughly 1.9 to 2.1 points in multi-turn dialogues. In-depth analysis revealed that roughly half of multi-turn failures stemmed from noncompliance as the distance to initial instructions increased, while another 48.8% resulted from error propagation, where an early mistake compounded across subsequent turns.

These results demonstrate that standard single-turn benchmark scores misrepresent actual operational readiness for conversational deployments. In production environments, degraded instruction-following and compounding errors increase operational risk, raise compliance concerns, and necessitate manual oversight. Deploying conversational agents under the assumption that high base capabilities guarantee reliable multi-turn adherence introduces significant functional vulnerabilities.

Organizations evaluating language models should mandate multi-turn benchmark assessments rather than relying on single-turn metrics. To mitigate performance decay, system designers should implement architectural safeguards, such as context management strategies that bring critical global instructions into nearer conversational context and validation mechanisms that detect and correct early errors before they propagate through the dialogue history.

The conclusions should be interpreted with awareness that the evaluation relied primarily on automated GPT-4 scoring, which human validation confirmed as well-aligned (0.71 Pearson correlation) but exhibiting a measurable leniency bias toward its own generated text. Additionally, computational constraints limited the study to models up to 14 billion parameters alongside select mixture-of-experts architectures, leaving the exact multi-turn degradation curve of larger open-source variants as an area for further validation.

Kwan et al (2024).pdf
Cover for MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Table of Contents

  • MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
  • Abstract

Knowls

  1. Knowl 1 — Four interaction capabilities define MT-Eval

    definition

    MT-Eval evaluates four recurring patterns in user–assistant dialogue. Recollection tests whether an assistant retains and applies information or instructions from earlier turns. Expansion tests responses to different questions about the same topic or document, without requiring reference to a specific earlier answer. Refinement tests whether the assistant incorporates successive clarifications or added constraints on a task. Follow-up tests whether it can answer a question that depends on details in its preceding response. The benchmark contains 168 dialogue sessions and 1,170 turns, averaging 6.96 turns per session.

  2. Knowl 2 — Task construction and benchmark composition

    model/method

    MT-Eval combines four task sets. Recollection has 38 dialogues and 380 turns: a document-classification task over 13 categories, and a global-instruction task in which an instruction must persist across later requests. Expansion has 10 dialogues and 70 turns; each dialogue uses one document as context for seven tasks—question answering, summarization, named-entity recognition, part-of-speech tagging, relation extraction, translation, and classification. Refinement has 40 dialogues and 480 turns; each dialogue contains two document-based NLP tasks, each with six progressively constrained instructions. Follow-up has 80 dialogues and 240 added turns, extending two-turn MT-Bench conversations with three further questions.

    To reduce leakage from existing evaluation data, the authors created new instances with GPT-4 and human review, or reformatted existing datasets. GPT-4 generated context documents for Expansion and Refinement. Human annotators reviewed instances for quality, relevance, difficulty, and originality; they changed or replaced 82% of GPT-4-generated queries. Follow-up questions were written by humans, with GPT-4 producing initial answers that were then human-checked. Corresponding single-turn instances were created for all turns in Recollection, Expansion, and Refinement; the authors did not construct a matched single-turn version for the main Follow-up comparison.

  3. Knowl 3 — Scoring protocol and evaluation checks

    experimental setup

    The experiments used greedy decoding and each model’s corresponding chat format, with the same helpful-assistant system prompt during inference. GPT-4 evaluated open-ended responses using a zero-shot chain-of-thought prompt and integer scores from 1 to 10. The Recollection document-classification task was scored by accuracy and normalized to 10; global-instruction following was scored by the average number of turns complying with the instruction, likewise normalized to 10. The authors compared single-turn and multi-turn results for the matched Recollection, Expansion, and Refinement instances; Follow-up was excluded from that main paired comparison.

    For human validation, five annotators rated 60 sampled responses in each of Expansion, Refinement, and Follow-up (180 responses total). Human and GPT-4 scores had average Pearson correlation 0.71 and Spearman correlation 0.64; by task, the Pearson/Spearman correlations were 0.74/0.58 for Refinement, 0.67/0.65 for Expansion, and 0.72/0.70 for Follow-up. A separate 200-response check found that GPT-4’s rating exceeded human ratings in 82.5% of cases when judging GPT-4 outputs, versus 47.5% for Mixtral-Instruct-8x7B, 60.0% for Qwen-chat-14B, 55.0% for Llama-2-chat-7B, and 52.5% for ChatGLM3-6B. The authors therefore excluded GPT-4 from their main model comparison because of potential judge bias. They also note that compute limits prevented testing larger open-source models such as Llama-2-chat-70B.

  4. Knowl 4 — Multi-turn scores across the four dialogue tasks

    empirical result

    In the main comparison, GPT-3.5-Turbo had the highest average score among the evaluated models, but open-source models led on some tasks: Mixtral-Instruct-8x7B scored 9.52 on Follow-up, above GPT-3.5-Turbo’s 9.21, and Mistral-Instruct-7B led Recollection with 7.22. No model’s four-task average reached 8. The scores below are on a 10-point scale; task columns are Recollection, Expansion, Refinement, and Follow-up.

    Model Average Recollection Expansion Refinement Follow-up
    GPT-3.5-Turbo 7.72 6.90 7.87 6.92 9.21
    ChatGLM3-6B 5.49 2.92 5.90 4.73 8.39
    Qwen-chat-7B 6.55 5.25 7.02 5.47 8.49
    Vicuna-7B-v1.5 6.44 5.45 6.70 5.31 8.31
    Llama-2-chat-7B 6.11 3.86 5.87 6.20 8.53
    Mistral-Instruct-7B 7.46 7.22 6.98 6.58 9.05
    Vicuna-13B-v1.5 7.01 6.27 6.70 6.37 8.68
    Llama-2-chat-13B 6.31 3.66 6.37 6.37 8.82
    Qwen-chat-14B 7.26 6.21 7.58 6.11 9.12
    Mixtral-Instruct-8x7B 7.47 6.17 7.42 6.77 9.52
  5. Knowl 5 — Matched single-turn tests expose multi-turn performance losses

    empirical result

    For Recollection, Expansion, and Refinement, the authors built single-turn counterparts to compare response quality when each request is separated from its dialogue history. Every model had a lower average score in the multi-turn setting. The size of this average decline did not simply track single-turn performance: for example, Llama-2-chat models outscored Vicuna models in single-turn evaluation but fell behind them in multi-turn evaluation. Follow-up is omitted because the paper did not create an equivalent single-turn version for this main comparison. Scores are on a 10-point scale; the difference is multi-turn average minus single-turn average.

    Model Single-turn average Multi-turn average Difference
    GPT-3.5-Turbo 8.07 7.23 -0.84
    ChatGLM3-6B 5.71 4.52 -1.19
    Vicuna-7B-v1.5 6.31 5.82 -0.49
    Llama-2-chat-7B 7.21 5.31 -1.90
    Qwen-chat-7B 6.86 5.91 -0.95
    Mistral-Instruct-7B 7.69 6.93 -0.76
    Vicuna-13B-v1.5 7.10 6.45 -0.65
    Llama-2-chat-13B 7.55 5.47 -2.08
    Qwen-chat-14B 7.62 6.64 -0.98
    Mixtral-Instruct-8x7B 8.28 6.78 -1.50
  6. Knowl 6 — Separating relevant context from a query reduces Refinement scores

    empirical result

    To test sensitivity to distance from relevant information, the authors inserted 1, 3, or 6 irrelevant dialogue turns sampled from LMSYS-Chat-1M into Refinement conversations, either between the context document and its queries or before the document. Scores below are on a 10-point scale. Inserting turns between the document and queries lowered all four models’ scores as the number of inserted turns increased. Inserting turns at the front produced mixed effects, including small improvements for some models. This supports the finding that separating a query from the context it needs can impair performance. The authors also observed that most models scored lower on the second six-turn Refinement task than on the first, whose instructions were closer to the initial document.

    Model No insertions 1 between 3 between 6 between 1 front 3 front 6 front
    Mistral-Instruct-7B 6.53 6.44 6.25 6.08 6.66 6.68 6.83
    Vicuna-13B-v1.5 6.62 5.91 5.47 5.56 6.25 6.16 5.89
    Qwen-chat-14B 6.30 5.89 5.76 5.17 6.22 6.01 6.18
    Mixtral-Instruct-8x7B 6.90 6.47 6.57 6.33 7.01 6.58 6.89
  7. Knowl 7 — Gold dialogue history improves Recollection and Refinement

    empirical result

    The authors compared normal dialogue histories containing models’ own earlier responses with histories containing gold responses from human-verified GPT-4 outputs. Gold history substantially improved Recollection and Refinement scores, consistent with reduced propagation of earlier mistakes and with correct prior responses serving as useful examples. Expansion changed little, where turns largely posed distinct NLP tasks. Scores are on a 10-point scale.

    Recollection Expansion Refinement
    Model Self-generated Gold Self-generated Gold Self-generated Gold
    Mistral-Instruct-7B 5.25 7.29 6.98 7.02 6.58 7.38
    Vicuna-13B-v1.5 4.64 7.32 6.70 6.87 6.37 7.15
    Qwen-chat-14B 4.43 7.00 7.58 7.63 6.11 6.95
    Mixtral-Instruct-8x7B 3.21 7.11 7.42 7.47 6.77 7.17
  8. Knowl 8 — Classification history helps when examples are correct, but noisy labels hurt

    empirical result

    In the Recollection document-classification task, the authors varied the dialogue history to test whether earlier examples help classify a new document. Gold history contains randomly selected documents with correct labels; DGC uses correct labels but excludes examples with the current document’s label; SGC contains examples from one randomly selected class, excluding the current label; RC assigns random labels to selected documents. RC (5) and RC (10) report RC performance at turns 5 and 10; ST is single-turn classification without dialogue history. The scores are classification accuracy in percent. Correctly labeled examples generally helped relative to ST, whereas random labels sharply reduced accuracy, with turn-10 accuracy lower than turn-5 accuracy for every listed model. Restricting examples to one class could also hurt weaker models.

    Model Gold DGC SGC RC RC (5) RC (10) ST
    Vicuna-13B-v1.5 81.00 84.00 70.00 45.00 62.00 28.00 75.00
    Qwen-chat-14B 94.00 95.00 86.00 69.00 68.00 60.00 94.00
    Mistral-Instruct-7B 96.00 95.00 95.00 75.00 80.00 70.00 94.00
    Mixtral-Instruct-8x7B 95.00 95.00 94.00 57.00 60.00 54.00 88.00
  9. Knowl 9 — Large score-loss cases mainly reflect forgotten instructions and propagated errors

    empirical result

    The authors examined 160 responses: for each of four high-performing models, they selected the ten largest single-turn versus multi-turn score differences in each of the four dialogue tasks. Of these cases, 80 (50%) were attributed to noncompliance with earlier instructions, 77 (48.1%) to errors accumulated in prior turns that misdirected later responses, and 3 (1.9%) to evaluation errors. Examples included forgetting a global formatting constraint and continuing to follow a Refinement instruction using the wrong paragraph identified earlier. The error analysis therefore identifies both long-range instruction retention and dependence on the accuracy of dialogue history as major sources of multi-turn failure.

  10. Knowl 10 — Follow-up dialogue history can improve performance

    empirical result

    Unlike the other task categories, Follow-up showed higher multi-turn than single-turn scores for every model tested in a separate analysis. For this comparison, GPT-4 condensed the relevant dialogue history into a single query for all 240 instances, and humans reviewed and revised the condensed queries where needed. The authors found that earlier responses and explanations could provide useful reasoning or examples for answering the current question. Because the single-turn version was constructed by condensing dialogue rather than being a directly matched original query, this result applies to that specific comparison. Scores are on a 10-point scale.

    Model Single-turn average Multi-turn average
    GPT-3.5-Turbo 9.19 9.21
    GPT-4 9.24 9.60
    ChatGLM3-6B 7.60 8.39
    Vicuna-7B-v1.5 7.88 8.31
    Llama-2-chat-7B 7.97 8.53
    Qwen-chat-7B 7.98 8.49
    Mistral-Instruct-7B 8.78 9.05
    Vicuna-13B-v1.5 8.37 8.68
    Llama-2-chat-13B 8.30 8.82
    Qwen-chat-14B 8.60 9.12
    Mixtral-Instruct-8x7B 9.02 9.52

Coverage note — The per-instruction Recollection breakdown and token/cost estimates are omitted as supplementary diagnostics and accounting; the benchmark design, primary comparisons, and principal ablation findings are included.

References

  1. 1.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. 2023. Qwen Technical Report.
  2. 2.Yonatan Bitton, Hritik Bansal, Jack Hessel, Rulin Shao, Wanrong Zhu, Anas Awadalla, Josh Gardner, Rohan Taori, and Ludwig Schmidt. 2023. VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  4. 4.Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. Chateval: Towards better llm-based evaluators through multi-agent debate. In The Twelfth International Conference on Learning Representations.
  5. 5.Hailin Chen, Fangkai Jiao, Xingxuan Li, Chengwei Qin, Mathieu Ravaut, Ruochen Zhao, Caiming Xiong, and Shafiq Joty. 2023. ChatGPT’s One-year Anniversary: Are Open-Source Large Language Models Catching up?
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  7. 7.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. GLM: General Language Model Pre-training with Autoregressive Blank Infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335, Dublin, Ireland. Association for Computational Linguistics.
  8. 8.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring Massive Multitask Language Understanding. In International Conference on Learning Representations.
  9. 9.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7B.
  10. 10.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024. Mixtral of Experts.
  11. 11.Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, Rose E. Wang, Minae Kwon, Joon Sung Park, Hancheng Cao, Tony Lee, Rishi Bommasani, Michael S. Bernstein, and Percy Liang. 2023. Evaluating Human-Language Model Interaction. Transactions on Machine Learning Research.
  12. 12.Yusheng Liao, Yutong Meng, Hongcheng Liu, Yanfeng Wang, and Yu Wang. 2023. An automatic evaluation framework for multi-turn medical consultations capabilities of large language models. arXiv preprint arXiv:2309.02077.
  13. 13.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.
  14. 14.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
  15. 15.OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Moss, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, C. J. Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2023. GPT-4 Technical Report.
  16. 16.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  17. 17.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  18. 18.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open Foundation and Fine-Tuned Chat Models.
  19. 19.Xingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng, and Heng Ji. 2023. MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback.
  20. 20.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems.
  21. 21.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2023. MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.
  22. 22.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223.
  23. 23.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P. Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang. 2023a. LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.
  24. 24.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023b. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.
  25. 25.Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023. Instruction-Following Evaluation for Large Language Models.

Citation

MLA
Kwan, W. C., et al. “MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 20153–77, https://doi.org/10.18653/v1/2024.emnlp-main.1124.
APA
Kwan, W. C., Zeng, X., Jiang, Y., Wang, Y., Li, L., Shang, L., Jiang, X., Liu, Q., & Wong, K.-F. (2024). MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 20153–20177. https://doi.org/10.18653/v1/2024.emnlp-main.1124
Chicago
Kwan, W. C., X. Zeng, Y. Jiang, et al. 2024. “MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 20153–77. https://doi.org/10.18653/v1/2024.emnlp-main.1124.
Harvard
Kwan, W.C. et al. (2024) “MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 20153–20177. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.1124.
Vancouver
1. Kwan WC, Zeng X, Jiang Y, Wang Y, Li L, Shang L, Jiang X, Liu Q, Wong K-F (2024) MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 20153–20177

BibTeX

@inproceedings{kwan-etal-2024-mt,
    title = "{MT}-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models",
    author = "Kwan, Wai-Chung  and
      Zeng, Xingshan  and
      Jiang, Yuxin  and
      Wang, Yufei  and
      Li, Liangyou  and
      Shang, Lifeng  and
      Jiang, Xin  and
      Liu, Qun  and
      Wong, Kam-Fai",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1124/",
    doi = "10.18653/v1/2024.emnlp-main.1124",
    pages = "20153--20177"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/