Is ChatGPT a General-Purpose Natural Language Processing Task Solver?
Chengwei QinAston ZhangZhuosheng ZhangJiaao ChenMichihiro YasunagaDiyi Yang
Evaluates ChatGPT's zero-shot performance across twenty standard benchmarks spanning seven task categories, identifying clear strengths in reasoning alongside persistent weaknesses in structured predictions like sequence tagging.
Recent advancements in large language models have shown strong conversational abilities, most notably with the debut of ChatGPT. However, organizations and researchers face uncertainty regarding whether ChatGPT can function as a reliable, general-purpose solver for diverse language tasks without requiring task-specific training. Understanding its real-world baseline capabilities and limitations across standard workflows is essential for leaders evaluating where generative artificial intelligence can be safely and effectively deployed.
The article evaluates the zero-shot performance of ChatGPT—meaning its ability to execute tasks without prior fine-tuning on downstream training data—across 20 benchmark datasets spanning seven core categories: reasoning, natural language inference, question answering, dialogue, summarization, named entity recognition, and sentiment analysis. The study benchmarks ChatGPT against its predecessor (GPT-3.5) as well as specialized, fine-tuned models, employing both standard zero-shot prompting and step-by-step reasoning prompts.
The findings show that ChatGPT is an effective tool for selected workflows but falls short of being a universal task solver. First, ChatGPT outperforms previous models on arithmetic reasoning (achieving up to 95.8% accuracy with step-by-step reasoning prompts), dialogue reasoning (76.2% accuracy), and sentiment analysis (93.7% accuracy). Second, it demonstrates strong capabilities on natural language inference and reading comprehension tasks, showing a distinct strength in confirming factual entailments over non-entailments. Third, ChatGPT frequently underperforms GPT-3.5 on commonsense, symbolic, and logical reasoning benchmarks. Fourth, it struggles significantly with structured sequence tagging, achieving an overall score of only 53.2% on named entity recognition compared to around 94% achieved by fine-tuned models. Finally, ChatGPT exhibits verbosity in summarization, producing longer texts that yield lower standard overlap scores, with explicit word-limit constraints further degrading output quality.
These results indicate that while ChatGPT provides immediate value for dialogue generation, sentiment classification, and mathematical problem-solving, relying on it for specialized structured data extraction or strictly constrained summarization introduces operational risks and lower accuracy. Organizations should note that task-specific fine-tuned models consistently outperform zero-shot ChatGPT across nearly all benchmarks. When planning system architectures, decision-makers should deploy specialized models for high-precision information extraction and reserve ChatGPT for interactive dialogue and open-ended analysis. Further evaluations on larger datasets and few-shot in-context learning configurations should be conducted before replacing dedicated machine learning pipelines in mission-critical applications.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Establishes the foundation of evaluating large autoregressive language models as zero-shot and few-shot multitask learners across diverse NLP benchmarks, which the source directly assesses with ChatGPT.
- Paper: Language Models are Unsupervised Multitask Learners, Alec Radford et al. (2019). Introduces the core concept that scaling unsupervised language models enables general-purpose zero-shot performance across standard natural language processing tasks.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). Demonstrates how instruction tuning transforms pretrained language models into capable zero-shot generalist task solvers, the foundational paradigm underlying conversational models like ChatGPT.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Provides the background on unlocking zero-shot multi-step and arithmetic reasoning in large language models through prompting, a key strength analyzed in the source.
- Paper: Multitask Prompted Training Enables Zero-Shot Task Generalization, Victor Sanh et al. (2021). Presents essential groundwork on multitask prompted training and standardized evaluation setups for testing zero-shot generalization across NLP task categories.
- Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Colin Raffel et al. (2020). Establishes the unified text-to-text framing for mapping diverse NLP tasks into natural language inputs and outputs that the source relies upon.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). Investigates how zero-shot natural language prompt framing locates task capabilities in large language models without task-specific updates.
- Paper: GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding, Alex Wang et al. (2018). Introduces the standard multi-task evaluation framework and benchmark dataset concepts used to systematically evaluate general-purpose natural language understanding.
- Paper: A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity, Yejin Bang et al. (2023). Extends the empirical zero-shot evaluation of ChatGPT into broader dimensions including multilingual capabilities, multimodal tasks, factuality, and hallucination.
- Paper: Empirical Study of Zero-Shot NER with ChatGPT, Tingyu Xie et al. (2023). Directly tackles ChatGPT's documented weaknesses on sequence tagging tasks by introducing structured reasoning and decomposition strategies for zero-shot named entity recognition.
- Paper: Consistency Analysis of ChatGPT, Myeongjun Jang et al. (2023). Probes the logical and semantic consistency of ChatGPT's zero-shot NLP performance under input perturbations across standard benchmark datasets.
- Paper: ChatGPT outperforms crowd workers for text-annotation tasks, Fabrizio Gilardi et al. (2023). Applies ChatGPT's zero-shot text processing capabilities to practical text-annotation and classification benchmarks, comparing performance against human crowd workers.
- Paper: Exploring the Potential of Large Language Models in Computational Argumentation, Guizhen Chen et al. (2024). Deepens the assessment of ChatGPT on specialized, complex reasoning tasks by evaluating its zero-shot and few-shot capabilities across computational argumentation benchmarks.
- Paper: Can Large Language Models Be an Alternative to Human Evaluations?, David Cheng-Han Chiang et al. (2023). Investigates whether the general-purpose natural language capabilities of ChatGPT allow it to serve as a reliable evaluator for open-ended text generation tasks.
- Paper: UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation, Daixuan Cheng et al. (2023). Develops universal prompt retrieval to systematically improve zero-shot evaluation performance across unseen task types in models like ChatGPT.
- Paper: Prompting Language Models for Linguistic Structure, Terra Blevins et al. (2023). Addresses large language models' difficulties with sequence tagging by introducing structured prompting methods to extract linguistic representations without fine-tuning.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). Builds upon ChatGPT's reasoning strengths and factual limitations by developing an iterative reading-and-reasoning framework over structured external data.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). Leverages conversational multi-turn capabilities to detect factual errors and inconsistencies in zero-shot language model outputs.
