ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness
Ján CeginJakub SimkoPeter Brusilovsky
Demonstrates that ChatGPT can replace human crowd workers in generating intent classification paraphrases at a fraction of the cost while yielding higher lexical and syntactic diversity without sacrificing classifier performance.
Building accurate natural language processing systems, such as intent classifiers that understand user commands in conversational applications, traditionally depends on crowdsourcing human workers to generate paraphrased training data. However, crowdsourcing is costly, time-consuming, and difficult to manage for consistent quality and linguistic diversity. With the rise of generative artificial intelligence, organizations face the strategic question of whether automated language models can effectively replace human labor in specialized dataset creation pipelines.
The main objective of the article is to determine whether ChatGPT can successfully substitute for crowd workers in generating paraphrased training examples for intent classification, specifically evaluating output validity, linguistic diversity, downstream classifier robustness, and overall cost efficiency.
To evaluate this, the researchers replicated a benchmark crowdsourcing study by generating synthetic paraphrases from seed sentences across multiple intent classes using ChatGPT and an open-source model, Falcon-40B. The data collection incorporated standard prompting as well as a constraint technique using taboo words that the system was instructed to avoid in order to encourage creative variations. The team evaluated lexical and structural diversity across thousands of generated samples, validated semantic accuracy, and trained intent classification models across five benchmark datasets to test accuracy on unseen, out-of-distribution evaluation data.
The investigation produced four central findings. First, ChatGPT proved highly reliable, producing semantically valid and intent-aligned paraphrases for all reviewed samples, whereas the open-source Falcon-40B model struggled with instructional adherence and produced invalid responses over 23% to 26% of the time. Second, ChatGPT demonstrated superior diversity compared to human workers, yielding an 11% to 29% larger vocabulary alongside significantly greater syntactic sentence variation. Third, machine learning models trained on ChatGPT data achieved equal or superior classification robustness on unseen benchmark tests compared to models trained on crowdsourced data. Finally, data generation using ChatGPT proved drastically more cost-effective, reducing acquisition costs by a ratio of roughly 1:600 relative to human crowdsourcing.
These results demonstrate that large language models provide a high-performing, economically viable mechanism for automating training data generation. Adopting model-generated paraphrasing can dramatically cut operational budgets and compress dataset development timelines from weeks to minutes without degrading model accuracy or robustness. However, automated generation requires targeted oversight because ChatGPT exhibits specific behavioral blind spots, such as avoiding informal slang, producing overly formal phrasings approximately 5% of the time, and failing to substitute acronyms or alternative names for named entities like locations or people.
Organizations should consider adopting generative models as a primary data augmentation tool for natural language pipelines while implementing hybrid human-in-the-loop oversight. Practitioners should deploy automated quality filters to remove rare constraint violations and duplicate outputs, while reserving human effort for high-value tasks such as validating edge cases, injecting localized slang, and diversifying named entities. Next steps should include pilot studies exploring prompt engineering techniques to address named entity handling, investigating diminishing returns during high-volume generation, and testing non-English language datasets.
Confidence in these findings is high for English-language intent classification under standard prompting frameworks. Nevertheless, readers should exercise caution regarding long-term reproducibility, as commercial language model versions evolve continuously over time, and potential training data overlap with public benchmark sets represents an inherent boundary condition in evaluating foundation models.
No sufficiently relevant recommendations were found.
- Paper: Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI, Nick Pangakis et al. (2025). This later study tests how to retain human oversight when generative models automate annotation, extending the source’s findings into a broader framework for responsible deployment.
