A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
Yejin BangSamuel CahyawijayaNayeon LeeWenliang DaiDan SuBryan WilieHoly LoveniaZiwei JiTiezheng YuWilly Chung
Reveals critical limits in ChatGPT's reasoning, multilingual generation, and factual hallucination across twenty-three datasets while demonstrating how multi-turn interactive prompting significantly improves task performance.
The rapid adoption of conversational artificial intelligence has highlighted the need for rigorous, independent evaluations to understand what these models can reliably do. The article presents an extensive, zero-shot quantitative benchmark of the December 2022 release of ChatGPT. The objective of the study was to evaluate its multitask capabilities across eight core language tasks, assess its multilingual and multimodal competencies, and diagnose its reliability in terms of reasoning, factuality, hallucination, and multi-turn interactivity.
To establish these benchmarks without application programming interface access, the authors conducted single-run experiments across 23 datasets spanning tasks such as question answering, summarization, machine translation, sentiment analysis, dialogue tracking, and fact-checking. Credibility is supported through the use of standardized test collections and manual verification of responses. A specialized national flag-drawing dataset was also created to assess whether the model could generate visual output using scalable vector graphics code as an intermediate medium. For most tasks, sample sizes ranged from 30 to 200 instances.
The key findings reveal notable strengths alongside clear vulnerabilities. First, in zero-shot performance across standard language datasets, the model outperformed earlier foundational systems on 9 out of 13 benchmarks and matched or exceeded fully specialized, fine-tuned systems on 4 tasks. Second, the system demonstrated significant performance disparities across languages: while it performed strongly on high-resource, Latin-script languages, it degraded sharply on low-resource and non-Latin languages, showing better capability in understanding non-Latin text than generating it. Third, the model proved to be an inconsistent reasoner, scoring an average accuracy of only 63.41% across 10 reasoning categories. It performed well on commonsense, deductive, and analogical tasks but struggled severely on inductive, mathematical (scoring 23.33%), multi-hop (scoring 26.67%), and spatial tasks. Fourth, the system consistently exhibited extrinsic hallucinations by generating unverified information not contained in the source input. Finally, multi-turn interactivity served as an effective mechanism for improvement, where follow-up instructions increased summarization quality by about 8% on standard lexical overlap metrics and improved translation accuracy.
These results demonstrate that while interactive language models offer strong general-purpose capabilities and intuitive collaborative refinement, their propensity to hallucinate and fail at multi-step reasoning introduces operational, legal, and compliance risks. Deploying the system in high-stakes environments—such as task-oriented customer service, complex analytical workflows, or low-resource linguistic regions—without human oversight could lead to factual inaccuracies and flawed decision-making.
Organizations should treat conversational models as interactive collaborators rather than fully autonomous agents. Stakeholders should implement multi-turn prompting frameworks to refine and verify outputs, integrate external knowledge retrieval systems to ground factuality, and deploy automated hallucination detection safeguards. Strong caution is advised for tasks requiring complex logic, mathematical precision, or low-resource translation until targeted fine-tuning and further domain-specific validation are completed.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). Provides the foundational Massive Multitask Language Understanding (MMLU) benchmark used to evaluate core knowledge and reasoning capabilities in large language models.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). Introduces the standard methodology and benchmark for quantifying model hallucinations and factual falsehoods.
- Paper: Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench authors (2022). Establishes a broad multi-task and multi-capability evaluation suite that directly informs the evaluation framing for complex language tasks.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Establishes chain-of-thought prompting paradigms necessary to interpret and benchmark LLM multi-step reasoning capabilities.
- Paper: CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge, Alon Talmor et al. (2019). Supplies the standard benchmark dataset used for assessing commonsense reasoning in natural language models.
- Paper: PIQA: Reasoning about Physical Commonsense in Natural Language, Yonatan Bisk et al. (2019). Establishes the physical commonsense reasoning evaluation dataset relevant to comprehensive reasoning benchmarks.
- Paper: Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering, Pan Lu et al. (2022). Introduces multimodal chain-of-thought reasoning evaluation across visual science question answering tasks.
- Paper: SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems, Alex Wang et al. (2019). Provides the core benchmark framework for multi-task general-purpose language understanding evaluation.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). Synthesizes empirical findings from early ChatGPT evaluation studies into an overarching, comprehensive taxonomy of LLM evaluation practices.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). Builds on empirical interactive chatbot assessments to formalize automated multi-turn evaluation frameworks via MT-Bench and Chatbot Arena.
- Paper: MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, Yubo Wang et al. (2024). Upgrades saturated multitask evaluation suites into a more discriminative reasoning benchmark in light of early ChatGPT baseline results.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). Delves deeper into ChatGPT's code generation performance by introducing rigorous, automated test-case generation to detect subtle functional errors.
- Paper: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, Xiang Yue et al. (2023). Extends multimodal evaluation beyond intermediate code-driven generation to rigorous expert-level multimodal understanding across college disciplines.
- Paper: MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts, Pan Lu et al. (2023). Expands on the study's visual and logical reasoning observations by systematically benchmarking mathematical reasoning situated within visual contexts.
- Paper: A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT, Jules White et al. (2023). Formalizes the conversational and iterative prompt engineering strategies observed in interactive LLM evaluations into structured design patterns.
- Paper: Toolformer: Language Models Can Teach Themselves to Use Tools, Timo Schick et al. (2023). Directly tackles the intrinsic hallucination and reasoning weaknesses of parametric LLMs by enabling models to self-invoke external computational tools.
