Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions
John Joon Young ChungEce KamarSaleema Amershi
Demonstrates how combining large language model diversification techniques with targeted human label replacement allows smaller downstream classifiers to outperform few-shot large language models by balancing synthetic text variety and annotation accuracy.
Building custom text classification models typically requires large volumes of labeled data, which can be expensive and time-consuming to collect manually. While generative large language models can automatically synthesize training text at a low cost, simply generating data often yields repetitive or narrow examples that fail to prepare models for real-world variations. However, adjusting generation settings to increase variety frequently causes errors in class alignment or produces irrelevant text. The article evaluates methods to balance diversity and correctness in synthetic text generation by combining automated diversification techniques with targeted human supervision.
To evaluate this approach, the authors conducted experiments across eight diverse text classification tasks using GPT-3 to generate training sets of several thousand instances, which were then used to train smaller, cost-effective BERT classification models. The study examined two automated diversification techniques: logit suppression, which penalizes previously generated words to avoid repetitive phrasing, and high sampling temperature, which flattens word-selection probabilities to produce less common text. The authors also tested two human-in-the-loop intervention strategies: correcting misaligned labels (label replacement) and removing irrelevant or unclassifiable text (out-of-scope filtering), scaling human review using lightweight classifier proxy models.
The findings show that while automated diversification techniques increase text variety, they significantly reduce generation accuracy and alignment with the target task. When used without correction, combining logit suppression and high temperature yields diminishing returns. However, introducing human label replacement resolves this trade-off: fully correcting misaligned labels improved downstream model accuracy by an absolute 14.4% on datasets generated with both diversification methods. Furthermore, correcting as few as 180 samples via proxy models enabled the smaller trained models to outperform direct few-shot classification using GPT-3. In contrast, filtering out-of-scope data provided no consistent benefit to downstream accuracy, primarily because discarding samples reduced the overall training volume.
These results demonstrate that organizations can achieve superior model accuracy and lower operational costs by generating diverse synthetic datasets paired with targeted human label correction, rather than relying on expensive direct querying of large language models for ongoing inference. Practitioners aiming to create synthetic datasets should combine high-temperature sampling and token suppression with active label review, potentially utilizing proxy models to scale human oversight efficiently. Decision-makers should note that the study focused on short-text classification benchmarks and an oracle simulation based on GPT-3, meaning additional validation on longer-form tasks, complex domain taxonomies, and other generative models is recommended prior to production deployment.
- Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). It introduces foundational concepts of decoding strategies and temperature sampling in neural text generation that the source directly explores to increase generation diversity.
- Paper: ChatGPT outperforms crowd workers for text-annotation tasks, Fabrizio Gilardi et al. (2023). It provides foundational empirical evidence for using large language models as automated data annotators and generators for text classification tasks prior to human-in-the-loop intervention.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). It examines how label correctness and demonstrations influence downstream model learning, which directly informs the source's investigation into label replacement interventions for synthetic datasets.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). It establishes key principles of aligning large language model outputs with human intent and task requirements, foundational to the human-AI partnership framing used in the source.
- Paper: Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs, Wei Zhou et al. (2026). It synthesizes broad methodologies for data preparation, cleaning, and label curation using LLMs, extending the source's focused analysis on human-in-the-loop synthetic data generation.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). It scales up data curation and filtering methodologies across large language model datasets, providing a comprehensive framework for dataset quality interventions.
- Paper: RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, Harrison Lee et al. (2024). It investigates substituting human intervention with AI feedback for alignment and data generation, directly continuing the question of scalability raised in the source's human-in-the-loop study.
