ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness

Ján CeginJakub SimkoPeter Brusilovsky

article2023EMNLP63 citations

Demonstrates that ChatGPT can replace human crowd workers in generating intent classification paraphrases at a fraction of the cost while yielding higher lexical and syntactic diversity without sacrificing classifier performance.

Listen

Building accurate natural language processing systems, such as intent classifiers that understand user commands in conversational applications, traditionally depends on crowdsourcing human workers to generate paraphrased training data. However, crowdsourcing is costly, time-consuming, and difficult to manage for consistent quality and linguistic diversity. With the rise of generative artificial intelligence, organizations face the strategic question of whether automated language models can effectively replace human labor in specialized dataset creation pipelines.

The main objective of the article is to determine whether ChatGPT can successfully substitute for crowd workers in generating paraphrased training examples for intent classification, specifically evaluating output validity, linguistic diversity, downstream classifier robustness, and overall cost efficiency.

To evaluate this, the researchers replicated a benchmark crowdsourcing study by generating synthetic paraphrases from seed sentences across multiple intent classes using ChatGPT and an open-source model, Falcon-40B. The data collection incorporated standard prompting as well as a constraint technique using taboo words that the system was instructed to avoid in order to encourage creative variations. The team evaluated lexical and structural diversity across thousands of generated samples, validated semantic accuracy, and trained intent classification models across five benchmark datasets to test accuracy on unseen, out-of-distribution evaluation data.

The investigation produced four central findings. First, ChatGPT proved highly reliable, producing semantically valid and intent-aligned paraphrases for all reviewed samples, whereas the open-source Falcon-40B model struggled with instructional adherence and produced invalid responses over 23% to 26% of the time. Second, ChatGPT demonstrated superior diversity compared to human workers, yielding an 11% to 29% larger vocabulary alongside significantly greater syntactic sentence variation. Third, machine learning models trained on ChatGPT data achieved equal or superior classification robustness on unseen benchmark tests compared to models trained on crowdsourced data. Finally, data generation using ChatGPT proved drastically more cost-effective, reducing acquisition costs by a ratio of roughly 1:600 relative to human crowdsourcing.

These results demonstrate that large language models provide a high-performing, economically viable mechanism for automating training data generation. Adopting model-generated paraphrasing can dramatically cut operational budgets and compress dataset development timelines from weeks to minutes without degrading model accuracy or robustness. However, automated generation requires targeted oversight because ChatGPT exhibits specific behavioral blind spots, such as avoiding informal slang, producing overly formal phrasings approximately 5% of the time, and failing to substitute acronyms or alternative names for named entities like locations or people.

Organizations should consider adopting generative models as a primary data augmentation tool for natural language pipelines while implementing hybrid human-in-the-loop oversight. Practitioners should deploy automated quality filters to remove rare constraint violations and duplicate outputs, while reserving human effort for high-value tasks such as validating edge cases, injecting localized slang, and diversifying named entities. Next steps should include pilot studies exploring prompt engineering techniques to address named entity handling, investigating diminishing returns during high-volume generation, and testing non-English language datasets.

Confidence in these findings is high for English-language intent classification under standard prompting frameworks. Nevertheless, readers should exercise caution regarding long-term reproducibility, as commercial language model versions evolve continuously over time, and potential training data overlap with public benchmark sets represents an inherent boundary condition in evaluating foundation models.

Cegin et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness

Table of Contents

  • 1 Introduction
  • 2 Related work: Collecting paraphrases and using ChatGPT
  • 3 ChatGPT paraphrase validity and diversity
  • 3.1 Data collection using ChatGPT
  • 3.2 ChatGPT data characteristics and validity
  • 3.3 Diversity of ChatGPT paraphrases
  • 3.4 Comparison of ChatGPT paraphrases with Falcon
  • 4 Model robustness
  • 4.1 Data and models used in the experiment
  • 4.2 Accuracy on out-of-distribution data
  • 5 Cost comparison: ChatGPT vs crowdsourcing?
  • 6 Discussion
  • 7 Conclusion
  • 8 Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Example code for sending requests to the ChatGPT API
  • B Further visualization of the collected data
  • C Datasets used in OOD experiments
  • C.1 Dataset statistics for OOD experiments
  • D Model robustness on OOD data with the inclusion of taboo samples and model training details
  • E Overlaps between data on the 5 different datasets
  • F Full results of model robustness for original, GPT and human data on 5 different dataset

Knowls

  1. Knowl 1 — ChatGPT paraphrase training data gives comparable or better OOD intent accuracy

    empirical result

    The study compared BERT-large intent classifiers trained on human-crowdsourced paraphrases with classifiers trained on ChatGPT paraphrases. Evaluation used non-seed examples from the same five benchmark datasets as out-of-distribution (OOD) test data; the values below are mean accuracy over 10 runs, with standard deviations and reported 95% confidence intervals. GPT training sets marked with an asterisk were downsampled. ChatGPT-trained models scored higher on Facebook, ATIS, and CLINC150, and slightly lower on Liu and Snips; the authors characterize the latter comparisons as comparable robustness.

    Facebook ATIS Liu CLINC150 Snips
    Human training samples 3092 2302 1140 3019 4216
    Human accuracy (SD) 72.65 (1.24) 79.46 (1.83) 93.81 (0.98) 95.42 (1.04) 98.89 (0.23)
    Human 95% CI [71.41–73.89] [78.33–80.59] [93.02–94.59] [94.59–96.26] [98.72–99.08]
    GPT training samples 2976* 2210* 1133* 3194* 3961*
    GPT accuracy (SD) 76.53 (2.71) 87.64 (3.26) 93.55 (0.41) 98.06 (0.30) 99.13 (0.18)
    GPT 95% CI [74.85–78.21] [85.62–89.67] [93.21–93.88] [97.82–98.29] [98.98–99.27]

    BERT-large was fine-tuned for five epochs. Accuracy on held-out, non-seed examples indicates that models trained on ChatGPT paraphrases were not less robust than those trained on human paraphrases in these experiments.

  2. Knowl 2 — ChatGPT paraphrases have higher lexical and syntactic diversity than human paraphrases

    empirical result

    The study compared human- and ChatGPT-generated paraphrases in Prompt and Taboo collection modes. Lexical diversity is the number of unique words in a dataset. Syntactic diversity is the mean tree edit distance (TED) across pairs of paraphrases sharing an intent; higher mean TED indicates greater structural variation. ChatGPT datasets had more unique words and higher mean TED than human datasets in both matched modes. A Mann–Whitney U comparison of TED values was reported as p=0.001p=0.001. Taboo prompting increased GPT lexical diversity, but did not increase its syntactic diversity relative to GPT Prompt data. The paper also reports that GPT paraphrases tend to be longer than human paraphrases.

    Dataset Collected samples After filtering Unique words Mean TED
    Prompt human 6091 5649 946 13.686
    Taboo human 5999 5941 1487 15.483
    Prompt Falcon 5850 2897 810 14.382
    Taboo Falcon 5850 1646 643 25.852
    Prompt GPT 5850 5170 1218 19.001
    Taboo GPT 5850 5143 1656 18.442
    Taboo GPT + taboo samples 5850 5608 1871 18.661

    The table compares matched collection modes as well as Falcon outputs. The largest direct GPT–human lexical differences are 1218 versus 946 unique words in Prompt mode and 1656 versus 1487 in Taboo mode; the corresponding mean TED values also favor GPT. The final row includes taboo-containing GPT samples that were excluded from the Taboo GPT dataset.

  3. Knowl 3 — Three-round ChatGPT collection replicated the crowd paraphrase protocol

    experimental setup

    The paraphrase-generation experiment followed the scale and general procedure of an earlier crowdsourcing protocol, using 10 personal-finance intent classes and three seed sentences per class. ChatGPT was prompted across three rounds. The Prompt condition asked for paraphrases of the seed; the Taboo condition additionally prohibited words supplied for that seed. The resulting Prompt GPT dataset combined Prompt-mode rounds, while Taboo GPT combined first-round Prompt data with Taboo-mode data from the later rounds. The same taboo words as the earlier human study were used, with three taboo words in the second round and six in the third.

    Data were collected on 5 May 2023 using the gpt-3.5-turbo-0301 checkpoint. The user prompt requested five rephrasings of the original phrase; the system message cast the model as a crowdsourcing worker who earns a living by creating paraphrases. API requests used temperature 1, 13 returned responses (n=13), and presence penalty 1.5; other parameters were left at their defaults. The study reports 5,850 generated records for each resulting GPT collection mode before cleaning.

  4. Knowl 4 — Manual review found valid intent-preserving ChatGPT paraphrases, with style and taboo caveats

    empirical result

    The authors manually assessed ChatGPT outputs against their seed sentences for semantic equivalence and preservation of intent, and reported that the paraphrases were valid and intent-adhering. They nevertheless judged approximately 5% stylistically odd, usually because of unusually formal wording; they also observed little slang or nonstandard grammar. The Prompt GPT dataset contained 5,170 samples after cleaning, and the Taboo GPT dataset contained 5,143 after cleaning and taboo filtering.

    Taboo compliance was imperfect: ChatGPT used at least one prohibited word in 167 of 1,950 samples in the round with three taboo words, and in 331 of 1,950 samples in the round with six. Those violations were removed from the final Taboo GPT dataset, following the earlier study’s filtering protocol.

  5. Knowl 5 — OOD robustness experiment used five benchmarks and seed-excluded test examples

    experimental setup

    For robustness evaluation, the study used ATIS, Liu, Facebook, Snips, and CLINC150 intent-classification datasets. Seed examples selected for paraphrase generation were excluded from the original-data evaluation set; the remaining examples served as OOD test data. Only intents represented by seed examples were included. For each benchmark’s seed set, ChatGPT paraphrases were collected in four rounds using 0, 2, 4, or 6 taboo words; empty responses, duplicates, and responses containing taboo words were filtered, and the collected paraphrases were manually checked. The authors checked for exact data overlaps and removed the small number detected, reporting that overlaps were under 1% of collected data.

    The compared classifiers were BERT-large, fine-tuned for five epochs, and a multiclass SVM using TF-IDF features. The evaluation was repeated over 10 runs. The study also evaluated GPT training sets with taboo-containing samples retained, separately from the taboo-filtered GPT sets.

  6. Knowl 6 — Retaining taboo-containing GPT samples generally improved OOD accuracy

    empirical result

    The study compared classifiers trained on human paraphrases, taboo-filtered GPT paraphrases, and GPT data with taboo-containing samples included, then evaluated them on original non-seed OOD examples from Facebook (FB), ATIS, Liu, CLINC150, and Snips. Values are mean accuracy over 10 runs. Including taboo-containing GPT samples generally raised accuracy relative to taboo-filtered GPT training, though not for every dataset or classifier; the effect was modest for BERT and more pronounced for some SVM results.

    BERT training data FB ATIS Liu CLINC150 Snips
    Human 72.65 79.46 93.81 95.42 98.89
    GPT 76.53 87.64 93.55 98.06 99.13
    GPT + taboo samples 79.64 87.74 93.13 98.42 99.07
    SVM training data FB ATIS Liu CLINC150 Snips
    Human 66.72 81.79 80.93 87.07 98.62
    GPT 60.39 81.12 81.54 94.76 97.98
    GPT + taboo samples 64.59 81.81 89.93 95.40 97.85

    The comparison shows that the effect depends on the dataset and classifier: for example, BERT accuracy on Liu and Snips is slightly lower with taboo samples included, while SVM accuracy on Snips also decreases slightly.

  7. Knowl 7 — Falcon-40B produced substantially fewer usable paraphrases than ChatGPT

    empirical result

    The open-model comparison used Falcon-40B-instruct with the same collection scale and essentially the same prompts and parameter values as the ChatGPT experiment, without specific parameter tuning. After deduplication and taboo filtering, manual review removed 887 Prompt samples (23.44% of the pre-review Prompt split) and 607 Taboo samples (26.94% of the pre-review Taboo split) as invalid. The resulting valid splits contained 2,897 Prompt samples and 1,646 Taboo samples.

    Reviewers found that Falcon often failed to perform the requested paraphrase task: outputs included meta-comments such as identifying itself as an AI, failed to preserve the seed’s intent, or answered the seed question instead of rephrasing it. Its usable datasets also had fewer unique words than the corresponding ChatGPT datasets. The authors therefore did not use Falcon-generated data in the robustness experiments.

  8. Knowl 8 — ChatGPT paraphrase collection cost about 1/600 of crowdsourcing

    empirical result

    Using the original study’s estimated cost of 0.05perhuman−producedparaphraseandtheAPIpriceatcollectiontimeof0.05 per human-produced paraphrase and the API price at collection time of 0.002 per 1,000 tokens, the authors estimated a large cost advantage for ChatGPT. In the diversity experiment, 10,050 human samples were estimated to cost about 500,versus9,750ChatGPTsamplesforabout500, versus 9,750 ChatGPT samples for about 0.50, a reported ratio of roughly 1:1,000. In the five-benchmark robustness experiment, 13,680 human samples were estimated at about 680,comparedwith26,273ChatGPTsamplesforabout680, compared with 26,273 ChatGPT samples for about 2.50, or roughly 1:525. Across both experiments, the paper reports an overall cost ratio of about 1:600 in favor of ChatGPT.

  9. Knowl 9 — GPT paraphrases overlapped less with original benchmark examples than human paraphrases

    empirical result

    To assess whether generated paraphrases duplicated examples already present in the benchmark data, the authors lowercased texts, lemmatized them, removed punctuation, and counted overlaps between GPT, human, and original data. GPT–original overlaps were lower than human–original overlaps in all five datasets. These overlaps were checked and removed before the robustness experiments.

    Dataset GPT–original overlaps Human–original overlaps
    Facebook 0 15
    ATIS 1 24
    Liu 3 8
    CLINC150 13 24
    Snips 0 12

    The comparison is relevant to interpreting the OOD results: direct overlap with the original evaluation corpus was uncommon and, under this matching procedure, less frequent for GPT paraphrases than for human paraphrases.

  10. Knowl 10 — ChatGPT paraphrases have limitations involving named entities and study scope

    limitation

    The authors observed that ChatGPT generally retained named entities rather than substituting acronyms or alternative names—for example, it did not vary a location name such as “NY” to “New York.” This limits variation when seed phrases contain names of places, people, or other entities. They also cautioned that the unusually formal or otherwise odd style of some outputs may diverge from typical human language use and could affect downstream applications.

    The experiments covered English only, used no extensive prompt engineering or systematic API-parameter study, and did not compare ChatGPT with several other contemporary language models. The authors did not determine when repeated generation from the same seed would reach diminishing returns through duplication. Because the ChatGPT service could change its underlying model without public versioning, exact reproduction may not remain possible; the authors also noted that possible overlap between ChatGPT’s pretraining data and evaluation benchmarks could have contributed to the observed robustness.

Coverage note — No substantial contributed material was deliberately omitted; related work, background, acknowledgements, and illustrative-only examples were excluded.

References

  1. 1.Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. Falcon-40B: an open large language model with state-of-the-art performance.
  2. 2.Mingda Chen, Qingming Tang, Sam Wiseman, and Kevin Gimpel. 2019. Controllable paraphrase generation with a syntactic exemplar. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5972–5984, Florence, Italy. Association for Computational Linguistics.
  3. 3.Wenqing Chen, Jidong Tian, Liqiang Xiao, Hao He, and Yaohui Jin. 2020. A semantically consistent and syntactically variational encoder-decoder framework for paraphrase generation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 1186–1198, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  4. 4.Jishnu Ray Chowdhury, Yong Zhuang, and Shuyi Wang. 2022. Novelty controlled paraphrase generation with retrieval augmented conditional prompt tuning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 10535–10544.
  5. 5.Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, Maël Primet, and Joseph Dureau. 2018. Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces. CoRR, abs/1805.10190.
  6. 6.Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Christian Petersen, Alexis Chevalier, and Julius Berner. 2023. Mathematical Capabilities of ChatGPT. ArXiv:2301.13867 [cs].
  7. 7.Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks. ArXiv:2303.15056 [cs].
  8. 8.Tanya Goyal and Greg Durrett. 2020. Neural syntactic preordering for controlled paraphrase generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 238–252, Online. Association for Computational Linguistics.
  9. 9.Sonal Gupta, Rushin Shah, Mrinal Mohit, Anuj Kumar, and Mike Lewis. 2018. Semantic parsing for task oriented dialog using hierarchical representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2787–2792, Brussels, Belgium. Association for Computational Linguistics.
  10. 10.Charles T. Hemphill, John J. Godfrey, and George R. Doddington. 1990. The ATIS spoken language systems pilot corpus. In Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27,1990.
  11. 11.Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu. 2023. Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine. ArXiv:2301.08745 [cs].
  12. 12.Nitish Joshi and He He. 2022. An investigation of the (in)effectiveness of counterfactually augmented data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3668–3681, Dublin, Ireland. Association for Computational Linguistics.
  13. 13.Kalpesh Krishna, John Wieting, and Mohit Iyyer. 2020. Reformulating unsupervised style transfer as paraphrase generation. EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, pages 737–762.
  14. 14.Stefan Larson, Anish Mahendran, Andrew Lee, Jonathan K. Kummerfeld, Parker Hill, Michael A. Laurenzano, Johann Hauswald, Lingjia Tang, and Jason Mars. 2019a. Outlier detection for improved data quality and diversity in dialog systems. NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 1:517–527.
  15. 15.Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, and Jason Mars. 2019b. An evaluation dataset for intent classification and out-of-scope prediction. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1311–1316, Hong Kong, China. Association for Computational Linguistics.
  16. 16.Stefan Larson, Anthony Zheng, Anish Mahendran, Rishi Tekriwal, Adrian Cheung, Eric Guldan, Kevin Leach, and Jonathan K. Kummerfeld. 2020. Iterative feature mining for constraint-based data collection to increase data diversity and model robustness. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8097–8106, Online. Association for Computational Linguistics.
  17. 17.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  18. 18.Xingkun Liu, Arash Eshghi, Pawel Swietojanski, and Verena Rieser. 2021. Benchmarking natural language understanding services for building conversational agents. In Increasing Naturalness and Flexibility in Spoken Dialogue Interaction: 10th International Workshop on Spoken Dialogue Systems, pages 165–183. Springer.
  19. 19.Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023. Is ChatGPT a General-Purpose Natural Language Processing Task Solver? ArXiv:2302.06476 [cs].
  20. 20.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  21. 21.Jorge Ramírez, Marcos Baez, Auday Berro, Boualem Benatallah, and Fabio Casati. 2022. Crowdsourcing Syntactically Diverse Paraphrases with Diversity-Aware Prompts and Workflows. In International Conference on Advanced Information Systems Engineering (CAiSE), pages 253–269.
  22. 22.Abhilasha Ravichander, Thomas Manzini, Matthias Grabmair, Graham Neubig, Jonathan Francis, and Eric Nyberg. 2017. How would you say it? eliciting lexically diverse dialogue for supervised semantic parsing. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 374–383, Saarbrücken, Germany. Association for Computational Linguistics.
  23. 23.Samuel Rhys Cox, Yunlong Wang, Ashraf Abdul, Christian von der Weth, and Brian Y. Lim. 2021. Directed diversity: Leveraging language embedding distances for collective creativity in crowd ideation. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, New York, NY, USA. Association for Computing Machinery.
  24. 24.Brian Thompson and Matt Post. 2020. Paraphrase generation as zero-shot multilingual translation: Disentangling semantic similarity from lexical and syntactic diversity. In Proceedings of the Fifth Conference on Machine Translation, pages 561–570, Online. Association for Computational Linguistics.
  25. 25.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models.
  26. 26.Petter Törnberg. 2023. ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning. ArXiv:2304.06588 [cs].
  27. 27.Veniamin Veselovsky, Manoel Horta Ribeiro, and Robert West. 2023. Artificial artificial artificial intelligence: Crowd workers widely use large language models for text production tasks.
  28. 28.Haohan Wang, Zeyi Huang, Hanlin Zhang, Yong Jae Lee, and Eric P Xing. 2022. Toward learning human-aligned cross-domain robust models by countering misaligned features. In Uncertainty in Artificial Intelligence, pages 2075–2084. PMLR.
  29. 29.Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023. Is ChatGPT a Good NLG Evaluator? A Preliminary Study. ArXiv:2303.04048 [cs].
  30. 30.William Yang Wang, Dan Bohus, Ece Kamar, and Eric Horvitz. 2012. Crowdsourcing the acquisition of natural language corpora: Methods and observations. In 2012 IEEE Spoken Language Technology Workshop (SLT), pages 73–78.
  31. 31.Mohammad Ali Yaghoub-Zadeh-Fard, Boualem Benatallah, Moshe Chai Barukh, and Shayan Zamanirad. 2019. A study of incorrect paraphrases in crowdsourced user utterances. NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 1:295–306.
  32. 32.Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and Dacheng Tao. 2023. Can ChatGPT Understand Too? A Comparative Study on ChatGPT and Fine-tuned BERT. ArXiv:2302.10198 [cs].
  33. 33.Yiming Zhu, Peixian Zhang, Ehsan-Ul Haq, Pan Hui, and Gareth Tyson. 2023. Can ChatGPT Reproduce Human-Generated Labels? A Study of Social Computing Tasks. ArXiv:2304.10145 [cs].

Citation

MLA
Cegin, J., et al. “ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1889–905, https://doi.org/10.18653/v1/2023.emnlp-main.117.
APA
Cegin, J., Simko, J., & Brusilovsky, P. (2023). ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1889–1905. https://doi.org/10.18653/v1/2023.emnlp-main.117
Chicago
Cegin, J., J. Simko, and P. Brusilovsky. 2023. “ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1889–1905. https://doi.org/10.18653/v1/2023.emnlp-main.117.
Harvard
Cegin, J., Simko, J. and Brusilovsky, P. (2023) “ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1889–1905. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.117.
Vancouver
1. Cegin J, Simko J, Brusilovsky P (2023) ChatGPT to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1889–1905

BibTeX

@inproceedings{cegin-etal-2023-chatgpt,
    title = "{C}hat{GPT} to Replace Crowdsourcing of Paraphrases for Intent Classification: Higher Diversity and Comparable Model Robustness",
    author = "Cegin, Jan  and
      Simko, Jakub  and
      Brusilovsky, Peter",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.117/",
    doi = "10.18653/v1/2023.emnlp-main.117",
    pages = "1889--1905"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/