Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions

John Joon Young ChungEce KamarSaleema Amershi

article2023ACL195 citations

Demonstrates how combining large language model diversification techniques with targeted human label replacement allows smaller downstream classifiers to outperform few-shot large language models by balancing synthetic text variety and annotation accuracy.

Listen

Building custom text classification models typically requires large volumes of labeled data, which can be expensive and time-consuming to collect manually. While generative large language models can automatically synthesize training text at a low cost, simply generating data often yields repetitive or narrow examples that fail to prepare models for real-world variations. However, adjusting generation settings to increase variety frequently causes errors in class alignment or produces irrelevant text. The article evaluates methods to balance diversity and correctness in synthetic text generation by combining automated diversification techniques with targeted human supervision.

To evaluate this approach, the authors conducted experiments across eight diverse text classification tasks using GPT-3 to generate training sets of several thousand instances, which were then used to train smaller, cost-effective BERT classification models. The study examined two automated diversification techniques: logit suppression, which penalizes previously generated words to avoid repetitive phrasing, and high sampling temperature, which flattens word-selection probabilities to produce less common text. The authors also tested two human-in-the-loop intervention strategies: correcting misaligned labels (label replacement) and removing irrelevant or unclassifiable text (out-of-scope filtering), scaling human review using lightweight classifier proxy models.

The findings show that while automated diversification techniques increase text variety, they significantly reduce generation accuracy and alignment with the target task. When used without correction, combining logit suppression and high temperature yields diminishing returns. However, introducing human label replacement resolves this trade-off: fully correcting misaligned labels improved downstream model accuracy by an absolute 14.4% on datasets generated with both diversification methods. Furthermore, correcting as few as 180 samples via proxy models enabled the smaller trained models to outperform direct few-shot classification using GPT-3. In contrast, filtering out-of-scope data provided no consistent benefit to downstream accuracy, primarily because discarding samples reduced the overall training volume.

These results demonstrate that organizations can achieve superior model accuracy and lower operational costs by generating diverse synthetic datasets paired with targeted human label correction, rather than relying on expensive direct querying of large language models for ongoing inference. Practitioners aiming to create synthetic datasets should combine high-temperature sampling and token suppression with active label review, potentially utilizing proxy models to scale human oversight efficiently. Decision-makers should note that the study focused on short-text classification benchmarks and an oracle simulation based on GPT-3, meaning additional validation on longer-form tasks, complex domain taxonomies, and other generative models is recommended prior to production deployment.

arXiv: 2306.04140
Cover for Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions

Abstract

Large language models (LLMs) can be used to generate text data for training and evaluating other models. However, creating high-quality datasets with LLMs can be challenging. In this work, we explore human-AI partnerships to facilitate high diversity and accuracy in LLM-based text data generation. We first examine two approaches to diversify text generation: 1) logit suppression, which minimizes the generation of languages that have already been frequently generated, and 2) temperature sampling, which flattens the token sampling probability. We found that diversification approaches can increase data diversity but often at the cost of data accuracy (i.e., text and labels being appropriate for the target domain). To address this issue, we examined two human interventions, 1) label replacement (LR), correcting misaligned labels, and 2) out-of-scope filtering (OOSF), removing instances that are out of the user’s domain of interest or to which no considered label applies. With oracle studies, we found that LR increases the absolute accuracy of models trained with diversified datasets by 14.4%. Moreover, we found that some models trained with data generated with LR interventions outperformed LLM-based few-shot classification. In contrast, OOSF was not effective in increasing model accuracy, implying the need for future work in human-in-the-loop text data generation.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Text Data Generation for Model Training
  • 2.2 Text Generation with LLMs
  • 2.3 Human-In-The-Loop
  • 3 Diversified Text Data Generation
  • 3.1 Goals
  • 3.2 Diversifying Approaches
  • 3.2.1 Settings for Data Generation
  • 3.2.2 Logit Suppression
  • 3.2.3 High Temperature
  • 4 Experiment 1: Diversified Text Data Generation
  • 4.1 Experiment Settings
  • 4.1.1 Tasks
  • 4.1.2 Generation Method
  • 4.1.3 Training Method
  • 4.2 Metrics
  • 4.3 Results
  • 5 Human Interventions to Fix Inaccurate Text Generation
  • 6 Experiment 2: Human Interventions For Diversified Text Generation
  • 6.1 Experiment Settings
  • 6.1.1 Label Replacement
  • 6.1.2 Out-of-Scope Filtering
  • 6.2 Results
  • 6.2.1 Label Replacement
  • 6.2.2 Out-of-Scope Instances Filtering
  • 7 Conclusion
  • 8 Limitations
  • 9 Ethics Statement
  • Acknowledgements
  • References
  • A Equation for Temperature Sampling
  • B Experiment 1 Details
  • B.1 Prompts Used in LLM Generation
  • B.2 Sampling Oracle Dataset
  • C Results of the Experiment 1 on Individual Dataset
  • D Experiment 2 Details
  • D.1 Examples of OOS Instances.
  • E Results of the Experiment 2 on Varying Tasks
  • E.1 Label Replacement
  • E.2 Out-of-Scope Filtering
  • F Results on Prompt C
  • ACL 2023 Responsible NLP Checklist

Knowls

  1. Knowl 1 — Label Replacement for Correcting Misaligned Labels in LLM Data Generation

    model/method

    Label Replacement (LR) is a human-in-the-loop intervention designed to correct label misalignment in synthetic text classification datasets generated by large language models (LLMs). When LLMs are prompted to generate text corresponding to a particular target class (especially under diversification constraints), the generated text often aligns with a different category or contains mislabeled content. Label replacement resolves this by re-annotating the generated instances with their true semantic category while retaining the generated text.

    To minimize human annotation effort, LR can be deployed either via full human inspection or through a sample-efficient proxy model approach:

    1. Full Inspection: Human annotators inspect all generated text instances and re-assign labels to reflect their true class.
    2. Proxy Model Labeling: Human annotators inspect only a small initial subset of generated samples (e.g., 90, 180, or 270 instances). For each target category ii, a binary linear Support Vector Classifier (SVC) is trained on the BERT embeddings of the annotated subset, treating instances belonging to class ii as positives and all other instances as negatives. For each unannotated generated instance, the proxy models output confidence scores Sp,i∈[0,1]S_{p,i} \in [0, 1] for each candidate class. The final assigned label is determined by maximizing a blended scoring function across classes ii: Sf,i=Ss,i⋅w+Sp,i⋅(1−w)S_{f,i} = S_{s,i} \cdot w + S_{p,i} \cdot (1 - w) where Ss,i∈{0,1}S_{s,i} \in \{0, 1\} indicates whether class ii was the prompt-specified class during generation, and w∈[0,1]w \in [0, 1] is a weighting parameter (set empirically to w=0.3w = 0.3) that balances the prompt prior against the proxy classifier's prediction.
  2. Knowl 2 — Empirical Effectiveness of Label Replacement on Model Accuracy

    empirical result

    Applying Label Replacement (LR) to synthetic datasets generated by GPT-3 (text-davinci-002) substantially improves both label alignment accuracy and the classification performance of downstream models (base BERT with 109M parameters) across eight natural language classification benchmarks.

    Key empirical findings include:

    • Unlocking Diversification Benefits: Without LR, applying both logit suppression and high sampling temperature (T=1.3T = 1.3) provided little to no cumulative gain over high temperature alone. With full LR, combining logit suppression and high temperature increased absolute downstream BERT model accuracy by 14.4%14.4\% compared to no LR, and by 17.8%17.8\% compared to baseline generation (low temperature T=0.3T = 0.3, no logit suppression, no LR).
    • Outperforming LLM Few-Shot Baselines: Downstream BERT classifiers trained on synthetic datasets without LR achieved lower average classification accuracy than direct few-shot classification using GPT-3. However, when proxy model LR was applied after inspecting just 180 generated instances under high temperature and logit suppression, the downstream BERT models consistently surpassed GPT-3 few-shot accuracy.
    • Dependence on Dataset Diversity: On datasets generated with low temperature (T=0.3T = 0.3) and no logit suppression, LR substantially increased label accuracy but produced only minimal gains in downstream model accuracy, demonstrating that correcting labels yields the greatest benefit only when the underlying text pool possesses sufficient linguistic and semantic diversity.
  3. Knowl 3 — Logit Suppression Algorithm for Diversified Text Generation

    algorithm

    Logit suppression is a diversification technique for iterative LLM-based text data generation that dynamically penalizes tokens that appear frequently across prior generation iterations, preventing the LLM from repeatedly outputting identical phrasing, templates, or error patterns.

    Input: LLM API with logit bias support, target class labels CC, desired total instances NN, batch size B=20B=20, maximum token penalty cap M=7.5M=7.5, multiplier α=7.5\alpha=7.5, top token count K=100K=100
    Output: Generated dataset D={(x1,y1),…,(xN,yN)}D = \{(x_1, y_1), \dots, (x_N, y_N)\}
    D←∅D \leftarrow \emptyset
    token_counts ←\leftarrow empty frequency map
    total_generated_tokens ←0\leftarrow 0
    while ∣D∣<N|D| < N do
        for each label c∈Cc \in C do
            logit_bias ←\leftarrow empty map
            if total_generated_tokens >0> 0 then
                sorted_tokens ←\leftarrow tokens in token_counts sorted descending by count
                for each token tt in the top min⁡(K,∣sortedtokens∣)\min(K, |sorted_tokens|) of sorted_tokens do
                    appearance_ratio ←tokencounts[t]/totalgeneratedtokens\leftarrow token_counts[t] / total_generated_tokens
                    penalty ←min⁡(M,α×appearanceratio×100)\leftarrow \min(M, \alpha \times appearance_ratio \times 100)
                    logit_bias[t]←−penalty[t] \leftarrow -penalty
            prompt ←\leftarrow format_prompt(label=cc, in_context_examples=sample_examples(DD, CC))
            new_instances ←\leftarrow LLM_Generate(prompt, logit_bias=logit_bias, batch_size=BB)
            for each text x∈newinstancesx \in new_instances do
                D←D∪{(x,c)}D \leftarrow D \cup \{(x, c)\}
                for each token tt in tokenize(xx) do
                    token_counts[t]←tokencounts[t]+1[t] \leftarrow token_counts[t] + 1
                    total_generated_tokens ←totalgeneratedtokens+1\leftarrow total_generated_tokens + 1
    return DD
  4. Knowl 4 — Trade-offs of Diversification Techniques in LLM Text Generation

    empirical result

    Across eight text classification benchmarks evaluated using GPT-3 (text-davinci-002), text diversification techniques—specifically logit suppression and elevated sampling temperature (T∈{0.7,0.9,1.3}T \in \{0.7, 0.9, 1.3\})—demonstrate consistent trade-offs:

    1. Diversity vs. Label Accuracy: Both logit suppression and higher sampling temperatures increase dataset diversity (measured via the Remote-Clique metric of mean pairwise BERT embedding distances), but simultaneously degrade label accuracy (the semantic alignment between the prompt-specified label and the generated text) and reduce semantic similarity to the ground-truth target domain.
    2. Relative Severity: Logit suppression induces a larger drop in label accuracy and target domain similarity than temperature scaling alone.
    3. Diminishing Returns in Isolation: When applied simultaneously without human label intervention, high temperature (T=1.3T = 1.3) and logit suppression do not produce additive accuracy improvements in downstream trained models; the benefit of logit suppression diminishes when combined with high temperature unless label correction is introduced.
    4. In-Context Example Seeding: Providing balanced seed examples in generation prompts improves both label accuracy and downstream model accuracy, and marginally increases diversity when used in the absence of logit suppression.
  5. Knowl 5 — Predictors of Downstream Classifier Accuracy via Linear Regression

    empirical result

    A linear regression analysis evaluating the relationship between dataset quality metrics and the downstream classification accuracy of base BERT models trained on LLM-generated text revealed that label accuracy, text diversity, and similarity to the oracle domain are all statistically significant positive predictors of downstream model accuracy:

    • Label Accuracy: Regression coefficient β=0.4797\beta = 0.4797 (p<0.001p < 0.001).
    • Dataset Diversity (measured via mean pairwise BERT embedding distances using Remote-Clique): Regression coefficient β=0.2260\beta = 0.2260 (p<0.001p < 0.001).
    • Domain Similarity (measured via average pairwise BERT embedding distances to the oracle training dataset): Regression coefficient β=0.1980\beta = 0.1980 (p<0.005p < 0.005).

    These results establish that optimizing downstream model performance requires jointly maximizing diversity and domain fidelity while maintaining strict label alignment.

  6. Knowl 6 — Proxy Model Score Blending Equation for Label Replacement

    equation

    In human-in-the-loop label replacement using proxy classifiers, the decision score Sf,iS_{f,i} for assigning candidate class ii to an uninspected generated instance is given by:

    Sf,i=Ss,i⋅w+Sp,i⋅(1−w)S_{f,i} = S_{s,i} \cdot w + S_{p,i} \cdot (1 - w)

    where:

    • i∈{1,…,C}i \in \{1, \dots, C\} indexes the candidate classification categories.
    • Sf,i∈RS_{f,i} \in \mathbb{R} is the blended final selection score for class ii.
    • Ss,i∈{0,1}S_{s,i} \in \{0, 1\} is a binary indicator set to 11 if class ii was the prompt-specified target label during LLM generation, and 00 otherwise.
    • Sp,i∈[0,1]S_{p,i} \in [0, 1] is the class probability or confidence score output by a binary linear Support Vector Classifier trained on BERT embeddings of a small annotated subset of generated text.
    • w∈[0,1]w \in [0, 1] is a fixed weighting constant, set to w=0.3w = 0.3 to account for potential proxy model classification errors while retaining information from the generation prompt.

    The label assigned to the text instance is arg⁡max⁡iSf,i\arg\max_i S_{f,i}.

  7. Knowl 7 — Temperature-Scaled Token Sampling Probability

    equation

    During autoregressive LLM text generation, the temperature-adjusted probability distribution over candidate vocabulary tokens is defined by:

    fT(p)i=pi1/T∑jpj1/Tf_T(p)_i = \frac{p_i^{1/T}}{\sum_j p_j^{1/T}}

    where:

    • pi∈[0,1]p_i \in [0, 1] is the original model probability assigned to token ii such that ∑jpj=1\sum_j p_j = 1.
    • T∈(0,∞)T \in (0, \infty) is the sampling temperature parameter.
    • fT(p)i∈[0,1]f_T(p)_i \in [0, 1] is the resulting temperature-scaled probability of token ii.

    When T>1T > 1, the sampling distribution becomes flatter, increasing the selection likelihood of lower-probability tokens and enhancing vocabulary diversity. When T<1T < 1, the distribution becomes sharper, concentrating probability mass on the highest-likelihood tokens.

  8. Knowl 8 — Out-of-Scope Filtering and Its Impact on Synthetic Text Datasets

    empirical result

    Out-of-Scope Filtering (OOSF) identifies and discards generated instances that do not belong to the target domain or do not fit any of the task's valid class definitions (e.g., standalone movie titles in sentiment classification or motivational advice in emotion classification).

    Empirical evaluation of OOSF across synthetic datasets revealed:

    • OOS Prevalence Across Tasks: In an oracle annotation sample of 360 generated instances per task, out-of-scope proportions varied widely: CARER (20.56%20.56\%), SST-2 (3.61%3.61\%), SUBJ (3.06%3.06\%), Clickbait (1.39%1.39\%), PubMed (1.11%1.11\%), Formality (0.56%0.56\%), HWU64 (0.28%0.28\%), and COLA (0.00%0.00\%).
    • Proxy Model Filtering Accuracy: Binary linear Support Vector Classifiers trained on BERT embeddings of annotated samples achieved high accuracy in detecting OOS instances (94.93%94.93\% on CARER, 97.18%97.18\% on SST-2, 97.50%97.50\% on SUBJ, 100%100\% on CB).
    • Impact on Downstream Accuracy: While OOSF modestly improved or maintained label accuracy and domain similarity while reducing diversity, it did not consistently improve downstream BERT classifier performance across tasks. Removing out-of-scope instances often reduced the total training set size or eliminated informative decision-boundary examples without yielding a net accuracy benefit.
  9. Knowl 9 — Experimental Setup for LLM Synthetic Dataset Generation and Downstream Evaluation

    experimental setup

    The methodology was evaluated across eight diverse NLP classification benchmarks: SST-2 (binary sentiment), Clickbait (CB; binary headline classification), CARER (6-class emotion recognition), PubMed 200k RCT (5-class sentence role classification), Subjectivity (SUBJ; binary subjective/objective), Formality (FO; binary formal/informal), HWU64 (18 chatbot intent domains), and COLA (binary linguistic acceptability).

    • Data Generation: GPT-3 (text-davinci-002) generated batches of 20 instances per API call with maximum token length 100, frequency penalty 0.02, and top-p=1.0p = 1.0. Datasets comprised 5,600 generated instances (6,922 for SST-2), matching the size of oracle training subsets.
    • Diversification Conditions: Evaluated temperatures T∈{0.3,0.7,0.9,1.3}T \in \{0.3, 0.7, 0.9, 1.3\}, with and without logit suppression on the top 100 tokens, and with or without 15–18 balanced few-shot seed examples.
    • Downstream Model Training: For each generated dataset, base BERT classifiers (109M parameters) were fine-tuned with a randomly initialized linear classification head using the Adam optimizer (learning rate 3×10−53 \times 10^{-5}, 3 warm-up epochs, early stopping patience of 5 epochs) across 5 random runs per dataset condition on NVIDIA RTX A6000 GPUs.
  10. Knowl 10 — Limitations of Post-Hoc Proxy Labeling and Out-of-Scope Instance Removal

    limitation

    The human-in-the-loop data curation framework exhibits several methodological limitations:

    1. Post-Hoc Class Imbalance: Applying proxy models after full dataset generation can skew the class distribution, causing certain categories to be underrepresented or overrepresented.
    2. Linguistic Style Skew from Misaligned Prompts: If many instances generated under a prompt for label AA are re-labeled by LR as label BB, the resulting data for label BB may inherit linguistic framing and lexical artifacts from prompt AA, distorting intra-class language patterns.
    3. Degraded Proxy Performance in High-Class Regimes: Training binary proxy models on small human-annotated sample pools (90–270 instances) yields diminishing returns on datasets with numerous fine-grained classes (e.g., CARER, HWU64, PubMed) compared to full manual inspection.
    4. Ineffectiveness of OOS Filtering: Binary removal of out-of-scope instances reduces training set volume and can inadvertently discard hard boundary examples, failing to yield consistent downstream accuracy improvements.

Coverage note — Omitted specific prompt templates (Prompt A, B, C variants) and detailed per-task sub-plots from Appendices C, D, E, and F, as they provide supporting raw examples and individual dataset variations that are fully encapsulated in the generalizable findings and aggregate results.

References

  1. 1.Saleema Amershi, James Fogarty, Ashish Kapoor, and Desney Tan. 2009. Overview based example selection in end user interactive concept learning. In Proceedings of the 22nd Annual ACM Symposium on User Interface Software and Technology, UIST '09, page 247–256, New York, NY, USA. Association for Computing Machinery.
  2. 2.Saleema Amershi, James Fogarty, and Daniel Weld. 2012. Regroup: Interactive machine learning for on-demand group creation in social networks. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI '12, page 21–30, New York, NY, USA. Association for Computing Machinery.
  3. 3.Steve Branson, Catherine Wah, Florian Schroff, Boris Babenko, Peter Welinder, Pietro Perona, and Serge Belongie. 2010. Visual recognition with humans in the loop. In Proceedings of the 11th European Conference on Computer Vision: Part IV, ECCV'10, page 438–451, Berlin, Heidelberg. Springer-Verlag.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Ángel Alexander Cabrera, Abraham J. Druck, Jason I. Hong, and Adam Perer. 2021. Discovering and validating ai errors with crowdsourced failure reports. Proc. ACM Hum.-Comput. Interact., 5(CSCW2).
  6. 6.George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A. Efros, and Jun-Yan Zhu. 2022. Dataset distillation by matching training trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  7. 7.Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, and Niloy Ganguly. 2016. Stop clickbait: Detecting and preventing clickbaits in online news media. In 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), pages 9–16.
  8. 8.Jiaao Chen, Zichao Yang, and Diyi Yang. 2020. MixText: Linguistically-informed interpolation of hidden space for semi-supervised text classification. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2147–2157, Online. Association for Computational Linguistics.
  9. 9.Justin Cheng and Michael S. Bernstein. 2015. Flock: Hybrid crowd-machine learning classifiers. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing, CSCW '15, page 600–611, New York, NY, USA. Association for Computing Machinery.
  10. 10.Franck Dernoncourt and Ji Young Lee. 2017. PubMed 200k RCT: a dataset for sequential sentence classification in medical abstracts. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 308–313, Taipei, Taiwan. Asian Federation of Natural Language Processing.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  12. 12.Marzieh Fadaee, Arianna Bisazza, and Christof Monz. 2017. Data augmentation for low-resource neural machine translation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 567–573, Vancouver, Canada. Association for Computational Linguistics.
  13. 13.James Fogarty, Desney Tan, Ashish Kapoor, and Simon Winder. 2008. Cueflik: Interactive concept learning in image search. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI '08, page 29–38, New York, NY, USA. Association for Computing Machinery.
  14. 14.Ian J. Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. MIT Press, Cambridge, MA, USA. http://www.deeplearningbook.org.
  15. 15.Demi Guo, Yoon Kim, and Alexander Rush. 2020. Sequence-level mixed sample data augmentation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5547–5552, Online. Association for Computational Linguistics.
  16. 16.Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3309–3326, Dublin, Ireland. Association for Computational Linguistics.
  17. 17.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network.
  18. 18.Yutai Hou, Yijia Liu, Wanxiang Che, and Ting Liu. 2018. Sequence-to-sequence data augmentation for dialogue language understanding. In Proceedings of the 27th International Conference on Computational Linguistics, pages 1234–1245, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  19. 19.Ashish Kapoor, Bongshin Lee, Desney Tan, and Eric Horvitz. 2010. Interactive optimization for steering machine classification. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI '10, page 1343–1352, New York, NY, USA. Association for Computing Machinery.
  20. 20.Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych. 2023. Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future. Computational Linguistics, 49(1):157–198.
  21. 21.Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020. Data augmentation using pre-trained transformer models. In Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems, pages 18–26, Suzhou, China. Association for Computational Linguistics.
  22. 22.Shibamouli Lahiri. 2015. Squinky! a corpus of sentence-level formality, informativeness, and implicature.
  23. 23.Zachary Levonian, Chia-Jung Lee, Vanessa Murdock, and F. Maxwell Harper. 2022. Trade-offs in sampling and search for early-stage interactive text classification. In 27th International Conference on Intelligent User Interfaces, IUI '22, page 566–583, New York, NY, USA. Association for Computing Machinery.
  24. 24.Ruibo Liu, Guangxuan Xu, Chenyan Jia, Weicheng Ma, Lili Wang, and Soroush Vosoughi. 2020. Data boost: Text data augmentation through reinforcement learning guided conditional generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9031–9041, Online. Association for Computational Linguistics.
  25. 25.Xingkun Liu, Arash Eshghi, Pawel Swietojanski, and Verena Rieser. 2021. Benchmarking Natural Language Understanding Services for Building Conversational Agents, pages 165–183. Springer Singapore, Singapore.
  26. 26.Nathan Ng, Kyunghyun Cho, and Marzyeh Ghassemi. 2020. SSMBA: Self-supervised manifold based data augmentation for improving out-of-domain robustness. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1268–1283, Online. Association for Computational Linguistics.
  27. 27.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04), pages 271–278, Barcelona, Spain.
  28. 28.Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red teaming language models with language models.
  29. 29.Mary Phuong and Christoph Lampert. 2019. Towards understanding knowledge distillation. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 5142–5151. PMLR.
  30. 30.Samuel Rhys Cox, Yunlong Wang, Ashraf Abdul, Christian von der Weth, and Brian Y. Lim. 2021. Directed diversity: Leveraging language embedding distances for collective creativity in crowd ideation. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI '21, New York, NY, USA. Association for Computing Machinery.
  31. 31.Marco Tulio Ribeiro and Scott Lundberg. 2022. Adaptive testing and debugging of NLP models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3253–3267, Dublin, Ireland. Association for Computational Linguistics.
  32. 32.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online. Association for Computational Linguistics.
  33. 33.Gaurav Sahu, Pau Rodriguez, Issam Laradji, Parmida Atighehchian, David Vazquez, and Dzmitry Bahdanau. 2022. Data augmentation for intent classification with off-the-shelf large language models. In Proceedings of the 4th Workshop on NLP for Conversational AI, pages 47–57, Dublin, Ireland. Association for Computational Linguistics.
  34. 34.Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. 2018. CARER: Contextualized affect representations for emotion recognition. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3687–3697, Brussels, Belgium. Association for Computational Linguistics.
  35. 35.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
  36. 36.Jina Suh, Soroush Ghorashi, Gonzalo Ramos, Nan-Chen Chen, Steven Drucker, Johan Verwey, and Patrice Simard. 2019. Anchorviz: Facilitating semantic data exploration and concept discovery for interactive machine learning. ACM Trans. Interact. Intell. Syst., 10(1).
  37. 37.Lichao Sun, Congying Xia, Wenpeng Yin, Tingting Liang, Philip Yu, and Lifang He. 2020. Mixuptransformer: Dynamic data augmentation for NLP tasks. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3436–3440, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  38. 38.Justin Talbot, Bongshin Lee, Ashish Kapoor, and Desney S. Tan. 2009. Ensemblematrix: Interactive visualization to support machine learning with multiple classifiers. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI '09, page 1283–1292, New York, NY, USA. Association for Computing Machinery.
  39. 39.Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros. 2018. Dataset distillation. arXiv preprint arXiv:1811.10959.
  40. 40.Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. Neural network acceptability judgments. Transactions of the Association for Computational Linguistics, 7:625–641.
  41. 41.Jason Wei and Kai Zou. 2019. EDA: Easy data augmentation techniques for boosting performance on text classification tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6382–6388, Hong Kong, China. Association for Computational Linguistics.
  42. 42.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  43. 43.Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2019. Errudite: Scalable, reproducible, and testable error analysis. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 747–763, Florence, Italy. Association for Computational Linguistics.
  44. 44.Congying Xia, Chenwei Zhang, Hoang Nguyen, Jiawei Zhang, and Philip Yu. 2020. Cg-bert: Conditional text generation with bert for generalized few-shot intent detection.
  45. 45.Kang Min Yoo, Dongju Park, Jaewook Kang, Sang-Woo Lee, and Woomyoung Park. 2021. GPT3Mix: Leveraging large-scale language models for text augmentation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2225–2239, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  46. 46.Kang Min Yoo, Youhyun Shin, and Sang-goo Lee. 2019. Data augmentation for spoken language understanding via joint variational generation. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, AAAI'19/IAAI'19/EAAI'19. AAAI Press.
  47. 47.Ann Yuan, Daphne Ippolito, Vitaly Nikolaev, Chris Callison-Burch, Andy Coenen, and Sebastian Gehrmann. 2021. Synthbio: A case study in faster curation of text datasets. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  48. 48.Jun Yuan, Jesse Vig, and Nazneen Rajani. 2022. Isea: An interactive pipeline for semantic error analysis of nlp models. In 27th International Conference on Intelligent User Interfaces, IUI '22, page 878–888, New York, NY, USA. Association for Computing Machinery.
  49. 49.Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018. mixup: Beyond empirical risk minimization. In International Conference on Learning Representations.
  50. 50.Le Zhang, Zichao Yang, and Diyi Yang. 2022. TreeMix: Compositional constituency-based data augmentation for natural language understanding. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5243–5258, Seattle, United States. Association for Computational Linguistics.
  51. 51.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS'15, page 649–657, Cambridge, MA, USA. MIT Press.
  52. 52.Jing Zhou, Yanan Zheng, Jie Tang, Li Jian, and Zhilin Yang. 2022. FlipDA: Effective and robust data augmentation for few-shot learning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8646–8665, Dublin, Ireland. Association for Computational Linguistics.

Citation

MLA
Chung, J. J. Y., et al. “Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 575–93, https://doi.org/10.18653/v1/2023.acl-long.34.
APA
Chung, J. J. Y., Kamar, E., & Amershi, S. (2023). Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 575–593. https://doi.org/10.18653/v1/2023.acl-long.34
Chicago
Chung, J. J. Y., E. Kamar, and S. Amershi. 2023. “Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 575–93. https://doi.org/10.18653/v1/2023.acl-long.34.
Harvard
Chung, J.J.Y., Kamar, E. and Amershi, S. (2023) “Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 575–593. Available at: https://doi.org/10.18653/v1/2023.acl-long.34.
Vancouver
1. Chung JJY, Kamar E, Amershi S (2023) Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 575–593

BibTeX

@inproceedings{chung-etal-2023-increasing,
    title = "Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions",
    author = "Chung, John  and
      Kamar, Ece  and
      Amershi, Saleema",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.34/",
    doi = "10.18653/v1/2023.acl-long.34",
    pages = "575--593"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/