Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations
Zhuoyan LiHangxiao ZhuZhuoran LuMing Yin
Demonstrates through empirical evaluation across ten datasets that the effectiveness of LLM-generated synthetic training data for text classifiers degrades significantly as task and instance subjectivity increase.
Building high-performing text classification models traditionally requires large volumes of carefully annotated real-world data, which is expensive and time-consuming to curate. As artificial intelligence advances, organizations are exploring whether large language models can automatically generate synthetic training data to replace or augment manual data collection. However, previous attempts have yielded inconsistent results, leaving decision-makers uncertain about when synthetic data generation is a viable strategy.
The article evaluates the effectiveness of synthetic text data generated by language models for training classification systems and demonstrates how task and instance subjectivity moderate this success. Specifically, it examines how model accuracy shifts across tasks requiring objective categorization versus those involving subjective human interpretation.
The authors conducted a series of empirical experiments evaluating 10 distinct text classification tasks spanning varying levels of subjectivity, such as news topic labeling, spam detection, sentiment analysis, sarcasm detection, and humor identification. Using OpenAI's GPT-3.5-Turbo, they generated synthetic training datasets under two settings: zero-shot generation (generating text purely from prompts) and few-shot generation (providing a small set of real-world examples to guide generation). They then trained standard classification models (BERT and RoBERTa) on these synthetic datasets and compared their performance against models trained on genuine, human-annotated data. Crowdsourced evaluations established the baseline subjectivity levels across both entire tasks and individual text instances.
The investigation produced four central findings. First, models trained on real-world data consistently outperformed those trained purely on synthetic data across nearly all tasks, with real data delivering an average performance advantage of approximately 7% to 17% over synthetic baselines. Second, few-shot generation significantly improved results compared to zero-shot prompting, boosting average classification scores by roughly 9% to 11%. Third, task subjectivity is the primary driver of performance degradation: for objective tasks like news topic tagging or spam filtering, synthetic data achieved performance very close to real data (often within a 2% to 6% gap), whereas highly subjective tasks like humor or sarcasm detection suffered massive accuracy drops exceeding 25% to 40%. Fourth, even within a single task, models trained on synthetic data performed noticeably better on individual instances where human annotators strongly agreed, struggling heavily on ambiguous or subjective examples where human consensus was low.
These findings indicate that large language models currently struggle to capture the diversity, nuance, and cultural context required for subjective human language. When training datasets lack variety, models converge too quickly and fail to generalize to complex real-world edge cases. For organizations, this means that while synthetic data generation can drastically cut data curation costs and development timelines for straightforward, factual tasks, relying entirely on synthetic data for nuanced or subjective tasks introduces severe performance and compliance risks.
Decision-makers should adopt a targeted, hybrid approach. For highly objective tasks, teams can confidently deploy synthetic data generation pipelines to accelerate deployment and reduce annotation expenses. For subjective or domain-specific tasks, organizations should avoid pure zero-shot generation and instead use few-shot hybrid strategies—collecting a modest set of curated human examples to guide the generation model and combining synthetic outputs with real-world data. Future operational workflows should also incorporate human feedback to deliberately enrich synthetic data diversity.
These conclusions are primarily bounded by experiments conducted using GPT-3.5-Turbo and crowd-sourced subjectivity ratings. While newer models such as GPT-4 may generate higher-quality text, practitioners should maintain caution when applying synthetic data strategies to complex, subjective language applications without preliminary validation pilots.
- Paper: EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks, Jason Wei et al. (2019). Its foundational study of text augmentation for classification provides a baseline for understanding how LLM-generated examples fit into the broader synthetic-data tradition.
- Paper: Is GPT-3 a Good Data Annotator?, Bosheng Ding et al. (2023). Its direct comparison of GPT-3-generated training data with human-labeled data establishes the central question of when LLM-produced examples help text classifiers.
- Paper: Quantifying the Persona Effect in LLM Simulations, Tiancheng Hu et al. (2024). It extends the source’s focus on subjectivity by testing whether persona prompting can account for human variation in subjective labeling—and how much it actually improves model predictions.
