Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations

Zhuoyan LiHangxiao ZhuZhuoran LuMing Yin

article2023EMNLP202 citations

Demonstrates through empirical evaluation across ten datasets that the effectiveness of LLM-generated synthetic training data for text classifiers degrades significantly as task and instance subjectivity increase.

Listen

Building high-performing text classification models traditionally requires large volumes of carefully annotated real-world data, which is expensive and time-consuming to curate. As artificial intelligence advances, organizations are exploring whether large language models can automatically generate synthetic training data to replace or augment manual data collection. However, previous attempts have yielded inconsistent results, leaving decision-makers uncertain about when synthetic data generation is a viable strategy.

The article evaluates the effectiveness of synthetic text data generated by language models for training classification systems and demonstrates how task and instance subjectivity moderate this success. Specifically, it examines how model accuracy shifts across tasks requiring objective categorization versus those involving subjective human interpretation.

The authors conducted a series of empirical experiments evaluating 10 distinct text classification tasks spanning varying levels of subjectivity, such as news topic labeling, spam detection, sentiment analysis, sarcasm detection, and humor identification. Using OpenAI's GPT-3.5-Turbo, they generated synthetic training datasets under two settings: zero-shot generation (generating text purely from prompts) and few-shot generation (providing a small set of real-world examples to guide generation). They then trained standard classification models (BERT and RoBERTa) on these synthetic datasets and compared their performance against models trained on genuine, human-annotated data. Crowdsourced evaluations established the baseline subjectivity levels across both entire tasks and individual text instances.

The investigation produced four central findings. First, models trained on real-world data consistently outperformed those trained purely on synthetic data across nearly all tasks, with real data delivering an average performance advantage of approximately 7% to 17% over synthetic baselines. Second, few-shot generation significantly improved results compared to zero-shot prompting, boosting average classification scores by roughly 9% to 11%. Third, task subjectivity is the primary driver of performance degradation: for objective tasks like news topic tagging or spam filtering, synthetic data achieved performance very close to real data (often within a 2% to 6% gap), whereas highly subjective tasks like humor or sarcasm detection suffered massive accuracy drops exceeding 25% to 40%. Fourth, even within a single task, models trained on synthetic data performed noticeably better on individual instances where human annotators strongly agreed, struggling heavily on ambiguous or subjective examples where human consensus was low.

These findings indicate that large language models currently struggle to capture the diversity, nuance, and cultural context required for subjective human language. When training datasets lack variety, models converge too quickly and fail to generalize to complex real-world edge cases. For organizations, this means that while synthetic data generation can drastically cut data curation costs and development timelines for straightforward, factual tasks, relying entirely on synthetic data for nuanced or subjective tasks introduces severe performance and compliance risks.

Decision-makers should adopt a targeted, hybrid approach. For highly objective tasks, teams can confidently deploy synthetic data generation pipelines to accelerate deployment and reduce annotation expenses. For subjective or domain-specific tasks, organizations should avoid pure zero-shot generation and instead use few-shot hybrid strategies—collecting a modest set of curated human examples to guide the generation model and combining synthetic outputs with real-world data. Future operational workflows should also incorporate human feedback to deliberately enrich synthetic data diversity.

These conclusions are primarily bounded by experiments conducted using GPT-3.5-Turbo and crowd-sourced subjectivity ratings. While newer models such as GPT-4 may generate higher-quality text, practitioners should maintain caution when applying synthetic data strategies to complex, subjective language applications without preliminary validation pilots.

arXiv: 2310.07849
  • Paper: Quantifying the Persona Effect in LLM Simulations, Tiancheng Hu et al. (2024). It extends the source’s focus on subjectivity by testing whether persona prompting can account for human variation in subjective labeling—and how much it actually improves model predictions.
Cover for Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations

Abstract

The collection and curation of high-quality training data is crucial for developing text classification models with superior performance, but it is often associated with significant costs and time investment. Researchers have recently explored using large language models (LLMs) to generate synthetic datasets as an alternative approach. However, the effectiveness of the LLM-generated synthetic data in supporting model training is inconsistent across different classification tasks. To better understand factors that moderate the effectiveness of the LLM-generated synthetic data, in this study, we look into how the performance of models trained on these synthetic data may vary with the subjectivity of classification. Our results indicate that subjectivity, at both the task level and instance level, is negatively associated with the performance of the model trained on synthetic data. We conclude by discussing the implications of our work on the potential and limitations of leveraging LLM for synthetic data generation¹.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Zero-shot Synthetic Data Generation
  • 3.2 Few-shot Synthetic Data Generation
  • 4 Evaluation I: Comparison Across Different Types of Tasks
  • 4.1 Datasets and Tasks
  • 4.2 Task-level Subjectivity Determination
  • 4.3 Model Training
  • 4.4 Evaluation Results
  • 4.5 Exploratory Analysis: Data Diversity
  • 5 Evaluation II: Comparison Across Different Task Instances
  • 5.1 Instance-level Subjectivity Determination
  • 5.2 Evaluation Results
  • 6 Conclusions and Discussions
  • 6.1 Why subjectivity adversely impacts the effectiveness of the synthetic data?
  • 6.2 Explaining a few exceptions
  • 6.3 Limitations and future work
  • References
  • A Appendices
  • A.1 Descriptions of tasks in Main Study
  • A.2 Descriptions of tasks in Robustness Check
  • B Evaluation I: Comparison Across Different Types of Tasks (Additional Results)
  • B.1 Convergence Analysis
  • B.2 Potential of Few-shot Synthetic Data for Data Augmentation
  • B.3 Similarity between the Synthetic Data and the Real Data
  • B.4 Additional Results on Additional Tasks with Low Subjectivity
  • B.5 Additional Results of Two Post-2022 Datasets
  • B.6 Additional Results When Changing the Volume of Synthetic Data
  • B.7 Additional Results When Using Other LLMs
  • B.8 Additional Results of Improved Data Generation Pipeline
  • B.9 Additional Results of Using Synthetic Data as Examples in Few-shot Setting
  • B.10 Additional Results for Directly Prompting LLMs for Text Classification
  • C Evaluation II: Comparison Across Different Task Instances (Additional Results)
  • D Additional Details on the Generation of Synthetic Data

Knowls

  1. Knowl 1 — Task subjectivity predicts the effectiveness of synthetic training data

    empirical result

    Across ten text-classification tasks—news topics, relation classification, movie-review sentiment, SMS spam, Reddit emotions, tweet irony, tweet emotions, news sarcasm, financial sentiment, and humor—models trained on GPT-3.5-Turbo-generated data generally performed worse than models trained on real data, and the gap was larger for tasks judged more subjective. With RoBERTa, the average advantage of real-data training over zero-shot synthetic-data training was 16.9% in Macro-F1 and 14.9% in accuracy; over few-shot synthetic-data training it was 6.7% and 6.1%, respectively. For BERT on the six tasks ranked most subjective, zero-shot synthetic-data training fell short of real-data training by an average of 27.4% in Macro-F1 and 24.2% in accuracy. The gap was described as small on the four less-subjective tasks. Tweet-irony detection was an exception: few-shot synthetic-data-trained BERT reached 81.5% Macro-F1 and 81.9% accuracy, versus 72.2% and 73.9% with real-data training; the corresponding RoBERTa scores were 83.3% and 83.7%, versus 74.0% and 75.5%.

  2. Knowl 2 — Synthetic-data models perform better on instances with stronger human agreement

    empirical result

    For each of ten classification tasks, the authors evaluated a BERT classifier trained on zero-shot GPT-3.5-Turbo synthetic data on subsets of test instances filtered by increasing thresholds on human annotation agreement. Accuracy generally increased as the agreement threshold rose, meaning performance tended to be better on instances with less disagreement and thus lower measured instance-level subjectivity. This increasing relationship appeared in eight of the ten tasks; Sarcasm News and Financial Phrasebank were the exceptions, and Spearman rank correlations often exceeded 0.85 for the other tasks. BERT models trained on real-world data showed a similar but usually weaker relationship, with Sarcasm News and Financial Phrasebank again notable exceptions. These are associations, not evidence that agreement itself causes accuracy to increase.

  3. Knowl 3 — Zero-shot and few-shot procedures for generating labeled text

    model/method

    The study used GPT-3.5-Turbo to generate labeled synthetic text in two settings. In zero-shot generation, the prompt sequence first established a task-specific context—for example, asking the model to act as a movie reviewer—and then requested text in a specified style, class, and length. A diversity instruction was added after each batch of generated examples: every 10 examples for short texts of at most 30 words, and after every example for longer paragraphs. In few-shot generation, the context was set in the same way, but each generation request was preceded by randomly sampled real text-label examples from the task. The prompt instructed the model to imitate their patterns without rewriting or merely modifying them. The authors generated 3,000 synthetic examples per class in each setting.

  4. Knowl 4 — Evaluation tasks and classifier-training protocol

    experimental setup

    The ten evaluated datasets covered AG’s News (three news topics), relation classification (four relations), IMDB reviews (positive/negative sentiment), SMS Spam (ham/spam), Reddit Emotion (joy, sadness, surprise), Tweet Irony (irony/non-irony), Tweet Emotion (anger, joy, optimism, sadness), Sarcasm News (sarcastic/non-sarcastic headlines), Financial Phrasebank (positive/negative/neutral sentiment), and Humor Speech (humorous/non-humorous text). For few-shot generation, the example pool was a random 10% of the original training data; only generated synthetic data, not that example pool, was used to train the few-shot classifiers. The comparison also included classifiers trained on real training data, with training-set sizes matched by downsampling the larger source. BERT and RoBERTa encoders supplied last-layer representations to a classifier with a 768-unit hidden layer and an output layer; training used a learning rate of 5×10−55\times10^{-5} and batch size 64. Official test partitions were used when available; otherwise data were split 70%/5%/25% into train/validation/test. Results were averaged over three runs and measured by Macro-F1 and accuracy.

  5. Knowl 5 — Few-shot examples improve synthetic-data performance, but are not uniformly sufficient

    empirical result

    Across the ten tasks, few-shot synthetic data—generated with a pool of real examples—almost always produced better classifiers than zero-shot synthetic data. The reported mean improvement from zero-shot to few-shot was 10.6% Macro-F1 and 8.8% accuracy for BERT, and 10.3% Macro-F1 and 8.9% accuracy for RoBERTa. In a separate augmentation comparison, the few-shot synthetic data combined with the small real example pool outperformed training on that limited real pool alone on many tasks; the comparison between synthetic-only training and limited-real-only training depended on the task. Thus the results support few-shot generation as a useful way to improve synthetic training data, but do not show that it reliably replaces real data or that combining the two always improves performance.

  6. Knowl 6 — Crowd judgments establish task-level subjectivity rankings

    experimental setup

    Task-level subjectivity was measured by asking U.S. MTurk workers to compare pairs of classification tasks, shown with task descriptions, label descriptions, and examples. Workers judged which task was more objective, where objective meant that classification could rely on clear, identifiable textual features without personal interpretation driven by biases, emotions, or beliefs. After filtering inattentive workers, the study retained 540 pairwise comparisons from 54 workers. Aggregated comparisons formed a directed graph; topological sorting produced a ranking, and tasks involved in a cycle were merged as a tied group before reranking. The resulting order ran from AG’s News (one star), Relation (two), IMDB (three), and SMS Spam (four) to six tasks assigned five stars: Reddit Emotion, Humor Speech, Tweet Irony, Sarcasm News, Tweet Emotion, and Financial Phrasebank. More stars indicate greater perceived subjectivity.

  7. Knowl 7 — Human agreement varies across tasks and aligns with the subjectivity ranking

    data/table

    The study also reported average instance-level agreement and Krippendorff’s α\alpha for the task evaluation samples. In the list below, each entry gives task: average agreement (mean annotations per instance); Krippendorff’s α\alpha; task-level subjectivity stars. Higher agreement measures indicate stronger inter-annotator agreement.

    AG’s News: 0.80 (4.2); 0.51; one star. Relation: 0.78 (4.5); 0.43; two stars. IMDB: 0.76 (7.3); 0.19; three stars. SMS Spam: 0.73 (8.5); 0.27; four stars. Reddit Emotion: 0.69 (6.6); 0.30; five stars. Humor Speech: 0.68 (7.1); 0.06; five stars. Tweet Irony: 0.68 (6.7); 0.03; five stars. Sarcasm News: 0.64 (7.7); 0.01; five stars. Tweet Emotion: 0.64 (4.6); 0.17; five stars. Financial Phrasebank: 0.57 (7.6); -0.03; five stars.

    The authors reported that average agreement broadly aligned with Krippendorff’s α\alpha, and that tasks ranked as more subjective generally had lower agreement. These measurements support—but do not independently validate—the use of annotator disagreement as a proxy for subjectivity.

  8. Knowl 8 — Real data are generally more diverse than generated data

    empirical result

    An exploratory analysis compared training-data diversity using Remote Clique Score, defined in the study as the average mean distance from an instance to other instances, and Chamfer Distance Score, defined as the average minimum distance from an instance to other instances; higher scores indicate greater diversity. Across tasks, real-world data generally scored as more diverse than few-shot synthetic data, which in turn generally scored as more diverse than zero-shot synthetic data. The diversity difference between real and generated data appeared more pronounced on the six higher-subjectivity tasks, particularly for Chamfer Distance. A t-test found that the zero-shot-versus-real decrease in Chamfer Distance was larger for the high-subjectivity group than for the low-subjectivity group (p<0.01p<0.01). The authors proposed limited coverage of real-life language scenarios as a possible explanation for poorer synthetic-data performance on subjective tasks; the diversity analysis is exploratory and does not establish that diversity caused the performance gap.

  9. Knowl 9 — Instance-level subjectivity is operationalized as majority-label agreement

    definition

    For a text instance ii, the study defined its agreement score aia_i as the fraction of annotators whose label matches the most frequent label. If YY is the set of possible labels, KiK_i is the number of annotators who labeled instance ii, and rik∈Yr_i^k\in Y is annotator kk’s label, then

    ai=max⁡y∈Y∑k=1Ki1(rik=y)Ki.a_i=\frac{\max_{y\in Y}\sum_{k=1}^{K_i}\mathbf{1}(r_i^k=y)}{K_i}.

    A lower aia_i indicates more disagreement and is treated as a proxy for greater instance-level subjectivity. To collect the labels, the authors sampled 50 test instances per class for each of ten task types and required at least three annotations from distinct MTurk workers per instance. Workers were assigned a task type and labeled 20 instances after receiving task instructions and passing attention checks.

  10. Knowl 10 — Additional evaluations support the subjectivity pattern, with bounded scope

    empirical result

    Several supplementary evaluations were consistent with the main pattern. On four additional, relatively low-subjectivity datasets—BBC News, Amazon Reviews, SST-2, and Yelp—the authors reported an average real-versus-zero-shot-synthetic performance difference of 4.2% for BERT. On two post-2022 datasets, BERT Macro-F1 was 79.4% for real data, 73.3% for zero-shot synthetic data, and 76.5% for few-shot synthetic data on ChatGPT App Reviews; on 2022 Tweet Emotion the corresponding scores were 68.9%, 53.5%, and 58.8%. Evaluations using GPT2-large and Llama 2, and a SunGen generation pipeline, retained the general association between greater task subjectivity and weaker synthetic-data effectiveness, although SunGen improved on direct zero-shot prompting. These checks broaden the evidence but do not remove the study’s stated scope limitations: the main evaluation relied on GPT-3.5-Turbo, crowd workers may lack linguistic expertise when judging subjectivity, and other moderators such as formality or domain knowledge may matter. The authors’ explanations involving diversity and the difficulty of reproducing majority labels are proposed interpretations, not experimentally established causes.

Coverage note — Detailed prompt wording, training-curve plots, direct-prompt classification comparisons, and synthetic-data volume sweeps are omitted because they are supplementary analyses rather than load-bearing results for the paper’s central subjectivity findings.

References

  1. 1.
    1. Chatgpt review. Accessed on Kaggle. Available from: https://www.kaggle.com/datasets/saloni1712/chatgpt-app-reviews.
  2. 2.
    1. Tweet sentiment emotions. Accessed on Kaggle. Available from: https://www.kaggle.com/datasets/ankitkumar2635/sentiment-and-emotions-of-tweets.
  3. 3.all MiniLM-L6-v2. 2023. sentence-transformers/all-minilm-l6-v2. Accessed on Hugging Face Model Hub. Available from: https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2.
  4. 4.Tiago A. Almeida, Jose Maria Gomez Hidalgo, and Akebo Yamakami. 2011. Contributions to the study of sms spam filtering: New collection and results. In Proceedings of the 2011 ACM Symposium on Document Engineering (DOCENG’11).
  5. 5.Issa Annamoradnejad and Gohar Zoghi. 2020. Colbert: Using bert sentence embedding for humor detection. arXiv preprint arXiv:2004.12765.
  6. 6.BBC. 2022. Accessed on Hugging Face. Available from: https://huggingface.co/datasets/SetFit/bbc-news.
  7. 7.Emile Benveniste. 1971. Subjectivity in language. Problems in general linguistics, 1:223–30.
  8. 8.Victor Besnier, Himalaya Jain, Andrei Bursuc, Matthieu Cord, and Patrick Pérez. 2020. This dataset does not exist: training models from generated images. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE.
  9. 9.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  10. 10.John Joon Young Chung, Ece Kamar, and Saleema Amershi. 2023. Increasing diversity while maintaining accuracy: Text data generation with large language models and human interventions. arXiv preprint arXiv:2306.04140.
  11. 11.Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. 2021. All that’s’ human’is not gold: Evaluating human evaluation of generated text. arXiv preprint arXiv:2107.00061.
  12. 12.Thomas H Cormen, Charles E Leiserson, Ronald L Rivest, and Clifford Stein. 2022. Introduction to algorithms. MIT press.
  13. 13.Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions: A Dataset of Fine-Grained Emotions. In 58th Annual Meeting of the Association for Computational Linguistics (ACL).
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  15. 15.Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A Smith, and Yejin Choi. 2021. Is gpt-3 text indistinguishable from human text? scarecrow: A framework for scrutinizing machine text. arXiv preprint arXiv:2107.01294.
  16. 16.Paul Ekman et al. 1999. Basic emotions. Handbook of cognition and emotion, 98(45-60):16.
  17. 17.Giorgio Franceschelli and Mirco Musolesi. 2023. On the creativity of large language models. arXiv preprint arXiv:2304.00008.
  18. 18.Jiahui Gao, Renjie Pi, LIN Yong, Hang Xu, Jiacheng Ye, Zhiyong Wu, WEIZHONG ZHANG, Xiaodan Liang, Zhenguo Li, and Lingpeng Kong. 2023. Self-guided noise-free data generation for efficient zero-shot learning. In The Eleventh International Conference on Learning Representations.
  19. 19.Tianyu Gao, Xu Han, Hao Zhu, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2019. FewRel 2.0: Towards more challenging few-shot relation classification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6251–6256, Hong Kong, China. Association for Computational Linguistics.
  20. 20.Mitchell L Gordon, Michelle S Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S Bernstein. 2022. Jury learning: Integrating dissenting voices into machine learning models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1–19.
  21. 21.Nitesh Goyal, Ian D Kivlichan, Rachel Rosen, and Lucy Vasserman. 2022. Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation. Proceedings of the ACM on Human-Computer Interaction, 6(CSCW2):1–28.
  22. 22.Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating large language models in generating synthetic hci research data: A case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA. Association for Computing Machinery.
  23. 23.Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. arXiv preprint arXiv:2203.09509.
  24. 24.Ruifei He, Shuyang Sun, Xin Yu, Chuhui Xue, Wenqing Zhang, Philip Torr, Song Bai, and Xiaojuan Qi. 2022. Is synthetic data from generative models ready for image recognition? arXiv preprint arXiv:2210.07574.
  25. 25.Nitin Jindal and Bing Liu. 2007. Review spam detection. In Proceedings of the 16th international conference on World Wide Web, pages 1189–1190.
  26. 26.Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4401–4410.
  27. 27.Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020. Data augmentation using pre-trained transformer models. arXiv preprint arXiv:2003.02245.
  28. 28.Zhuoyan Li, Zhuoran Lu, and Ming Yin. 2022. Towards better detection of biased language with scarce, noisy, and biased annotations. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, pages 411–423.
  29. 29.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  30. 30.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  31. 31.P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala. 2014. Good debt or bad debt: Detecting semantic orientations in economic texts. Journal of the Association for Information Science and Technology, 65.
  32. 32.Andrew M Mcnutt, Chenglong Wang, Robert A Deline, and Steven M. Drucker. 2023. On the design of ai-powered code assistants for notebooks. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA. Association for Computing Machinery.
  33. 33.Rishabh Misra and Prahal Arora. 2023. Sarcasm detection using news headlines dataset. AI Open, 4:13–18.
  34. 34.Rishabh Misra and Jigyasa Grover. 2021. Sculpting Data for ML: The first act of Machine Learning.
  35. 35.Saif Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018. Semeval-2018 task 1: Affect in tweets. In Proceedings of the 12th international workshop on semantic evaluation, pages 1–17.
  36. 36.Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pages 188–197.
  37. 37.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741.
  38. 38.OpenAI. 2023. Gpt-4 technical report.
  39. 39.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  40. 40.Samuel Rhys Cox, Yunlong Wang, Ashraf Abdul, Christian von der Weth, and Brian Y. Lim. 2021. Directed diversity: Leveraging language embedding distances for collective creativity in crowd ideation. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, New York, NY, USA. Association for Computing Machinery.
  41. 41.Gaurav Sahu, Pau Rodriguez, Issam H. Laradji, Parmida Atighehchian, David Vazquez, and Dzmitry Bahdanau. 2022. Data augmentation for intent classification with off-the-shelf large language models.
  42. 42.Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A Smith. 2021. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. arXiv preprint arXiv:2111.07997.
  43. 43.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
  44. 44.Ruixiang Tang, Xiaotian Han, Xiaoqian Jiang, and Xia Hu. 2023. Does synthetic data generation of llms help clinical text mining? arXiv preprint arXiv:2303.04360.
  45. 45.Cynthia Van Hee, Els Lefever, and Véronique Hoste. 2018. Semeval-2018 task 3: Irony detection in english tweets. In Proceedings of The 12th International Workshop on Semantic Evaluation, pages 39–50.
  46. 46.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  47. 47.Thomas C Veatch. 1998. A theory of humor.
  48. 48.Zirui Wang, Adams Wei Yu, Orhan Firat, and Yuan Cao. 2021. Towards zero-label language learning. arXiv preprint arXiv:2109.09193.
  49. 49.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  50. 50.Janyce Wiebe, Theresa Wilson, Rebecca Bruce, Matthew Bell, and Melanie Martin. 2004. Learning subjective language. Computational linguistics, 30(3):277–308.
  51. 51.Michael Wiegand, Josef Ruppenhofer, and Thomas Kleinbauer. 2019. Detection of abusive language: the problem of biased datasets. In Proceedings of the 2019 conference of the North American Chapter of the Association for Computational Linguistics: human language technologies, volume 1 (long and short papers), pages 602–608.
  52. 52.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  53. 53.Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong. 2022. Zerogen: Efficient zero-shot learning via dataset generation. arXiv preprint arXiv:2202.07922.
  54. 54.Kang Min Yoo, Dongju Park, Jaewook Kang, Sang-Woo Lee, and Woomyeong Park. 2021. Gpt3mix: Leveraging large-scale language models for text augmentation. arXiv preprint arXiv:2104.08826.
  55. 55.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015a. Character-level convolutional networks for text classification. Advances in neural information processing systems, 28.
  56. 56.Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015b. Character-level convolutional networks for text classification. In NIPS.
  57. 57.Yuxuan Zhang, Huan Ling, Jun Gao, Kangxue Yin, Jean-Francois Lafleche, Adela Barriuso, Antonio Torralba, and Sanja Fidler. 2021. Datasetgan: Efficient labeled data factory with minimal human effort.
  58. 58.Jiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G Parker, and Munmun De Choudhury. 2023. Synthetic lies: Understanding ai-generated misinformation and evaluating algorithmic and human solutions. CHI ’23, New York, NY, USA. Association for Computing Machinery.

Citation

MLA
Li, Z., et al. “Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations”. arXiv, 2023, http://arxiv.org/abs/2310.07849v2.
APA
Li, Z., Zhu, H., Lu, Z., & Yin, M. (2023). Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations. arXiv. http://arxiv.org/abs/2310.07849v2
Chicago
Li, Z., H. Zhu, Z. Lu, and M. Yin. 2023. “Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations”. arXiv. http://arxiv.org/abs/2310.07849v2.
Harvard
Li, Z. et al. (2023) “Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.07849v2.
Vancouver
1. Li Z, Zhu H, Lu Z, Yin M (2023) Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations. arXiv

BibTeX

@article{li2023synthetic,
  title = {Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations},
  author = {Li, Zhuoyan and Zhu, Hangxiao and Lu, Zhuoran and Yin, Ming},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.07849v2},
  eprint = {2310.07849}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/