Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text

Yao DouMaxwell ForbesRik Koncel-KedziorskiNoah A. SmithYejin Choi

article2022ACL163 citations

Presents SCARECROW, a fine-grained span-level error annotation framework that enables crowd workers to detect and categorize subtle generation errors across ten distinct types, exposing persistent measurable gaps between human writing and large language models like GPT-3.

Listen

Recent advances in artificial intelligence have produced language models capable of generating highly fluent and seemingly natural text. This fluency creates a significant operational and governance challenge: standard human reviews often fail to distinguish machine-generated writing from human-authored content, making subtle factual, logical, and structural errors difficult to detect and evaluate reliably.

The article develops and demonstrates SCARECROW, a fine-grained evaluation framework designed to systematically identify, categorize, and explain errors in machine-generated text using trained crowd reviewers.

To measure text quality beyond holistic ratings, the authors established a schema of ten distinct issue types divided into language errors (such as redundancy, self-contradiction, and incoherence), factual errors (including basic math mistakes and commonsense violations), and reader comprehension issues (such as obscure jargon and claims requiring search engine verification). Using this framework, crowdsourced annotators reviewed 1,300 English news paragraphs generated by humans and various model configurations—including GPT-2, Grover, and fourteen decoding variants of GPT-3—yielding an annotated dataset of more than 41,000 specific error spans.

The analysis reveals several critical findings. First, increasing model scale significantly reduces incoherence and commonsense violations, but error reductions plateau for basic arithmetic, prompt adherence, and grammar. Second, larger models exhibit complex failure modes; rather than repeating simple phrases, larger systems often generate extensive, topically redundant blocks of text or self-contradictions. Third, decoding hyperparameters exert an enormous impact on output quality: depending on configuration, GPT-3's performance ranged from worse than older, smaller models to an apparent parity with human text. Finally, deeper audit of the annotations showed that perceived parity is misleading; crowd annotators generated high rates of false positives on fluent text, whereas human-written news articles actually contained far fewer genuine errors than the best machine outputs.

These findings have direct operational implications for organizations deploying large language models. Holistic human evaluations and automated surface metrics risk missing severe, subtle failures in reasoning, factual accuracy, and narrative consistency. System performance depends heavily on decoding configurations—such as repetition penalties and sampling parameters—meaning that poor hyperparameter choices can completely negate the benefits of larger model scale. Automated error detection models trained on this data achieved high recall in flagging unverifiable claims and inconsistencies, demonstrating the feasibility of targeted quality assurance pipelines.

Organizations evaluating or deploying generative text systems should avoid relying on superficial fluency checks or standard holistic rating scales. Instead, teams should implement span-level error audits and systematically tune decoding parameters, particularly frequency penalties, to control redundancy without causing topical drift. Automated verification pipelines should be explored to pre-screen claims before public or high-stakes release.

The findings are subject to specific boundary conditions. The evaluation focused primarily on single-paragraph news continuations, and individual crowd annotators operated with high precision but low recall, requiring ten annotators per text to achieve robust coverage. While findings regarding model scaling and decoding sensitivity are robust within this context, readers should exercise caution when extrapolating these exact error distributions to longer documents, creative writing, or non-news domains.

  • Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). Its analysis of decoding strategies, including nucleus sampling, provides the foundation for understanding why SCARECROW tests decoding configurations as a source of perceived text quality differences.
Cover for Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text

Abstract

Modern neural language models can produce remarkably fluent and grammatical text. So much, in fact, that recent work by Clark et al. (2021) has reported that conventional crowdsourcing can no longer reliably distinguish between machine-authored (GPT-3) and human-authored writing. As errors in machine generations become ever subtler and harder to spot, it poses a new challenge to the research community for robust machine text evaluation.

We propose a new framework called SCARECROW for scrutinizing machine text via crowd annotation. To support the broad range of real machine errors that can be identified by laypeople, the ten error categories of SCARECROW—such as redundancy, commonsense errors, and incoherence—are identified through several rounds of crowd annotation experiments without a predefined ontology.

We then use SCARECROW to collect over 41k error spans in human-written and machine-generated paragraphs of English language news text. We isolate factors for detailed analysis, including parameter count, training data, and various decoding-time configurations. Our approach successfully quantifies measurable gaps between human authored text and generations from models of several sizes, including fourteen configurations of GPT-3. In addition, our analysis unveils new insights, with detailed rationales provided by laypeople, e.g., that the commonsense capabilities have been improving with larger models while math capabilities have not, and that the choices of simple decoding hyperparameters can make remarkable differences on the perceived quality of machine text. We release our training material, annotation toolkit and dataset at https://yao-dou.github.io/scarecrow/.

Table of Contents

  • 1 Introduction
  • 2 Key Findings
  • 3 Evaluation of Natural Language Generation
  • 4 SCARECROW Annotation Methodology
  • 4.1 Prompt and Generation
  • 4.2 Span Labeling
  • 4.3 Span Selection
  • 4.4 Error Types
  • 4.5 Severity
  • 4.6 Explanation
  • 4.7 Annotation Process
  • 5 Data Collection
  • 5.1 Models
  • 5.2 Decoding strategies
  • 5.3 Prompt Selection
  • 5.4 Generation
  • 5.5 Annotation
  • 6 Error Prediction
  • 7 Related Work
  • 8 Conclusion
  • Acknowledgments
  • References
  • A SCARECROW Annotation Schema
  • A.1 Language Errors
  • A.1.1 Grammar and Usage
  • A.1.2 Redundant
  • A.1.3 Off-Prompt
  • A.1.4 Self-Contradiction
  • A.1.5 Incoherent
  • A.2 Factual Errors
  • A.2.1 Bad Math
  • A.2.2 Commonsense
  • A.2.3 Encyclopedic
  • A.3 Reader Issues
  • A.3.1 Technical Jargon
  • A.3.2 Needs Google
  • B Annotation Details
  • B.1 Error Severity
  • B.2 Grading Details
  • C Data Quality
  • D Dataset Statistics
  • E Detailed Analysis
  • E.1 Off-Prompt
  • E.2 Self-Contradiction
  • E.3 Redundant
  • E.4 Reader Issues
  • E.5 Decoding Hyperparameters
  • E.6 Best GPT-3 vs. Humans
  • E.7 Topics
  • E.8 Error explanations
  • F Future Work
  • F.1 SCARECROW Studies: Simple
  • F.2 SCARECROW Studies: Complex
  • F.3 Broadening SCARECROW
  • F.4 Applications

Knowls

  1. Knowl 1 — SCARECROW elicits localized explanations of text problems

    model/method

    SCARECROW is a crowd-annotation framework for scrutinizing open-ended generated text through specific, explained problem spans rather than a single holistic quality score. An annotator sees a human-written one-sentence prompt and a continuation of 80–145 tokens; the prompt is identified as human-written, but the continuation’s source is concealed. Annotators select the smallest span that contains an issue (at least one word; boundaries are snapped to word boundaries), assign exactly one error type and a severity, and explain their judgment in natural language. Different, partially or fully overlapping spans can represent multiple problems in the same text. Each paragraph is annotated by 10 workers. Workers receive training and qualification, and are required to score at least 90/100 to pass. The framework is intended to elicit problems recognizable to lay readers, not only expert judgments.

  2. Knowl 2 — The ten SCARECROW error types separate text errors from reader obstacles

    definition

    SCARECROW groups ten span labels into three classes. Language errors concern how ideas are expressed: Grammar and Usage marks missing, extra, incorrect, or misplaced words; Off-Prompt marks continuation text unrelated to or contradictory to the prompt; Redundant marks repeated words or ideas, with both an antecedent and the repetition identified; Self-Contradiction marks mutually inconsistent statements within the continuation, likewise identifying the antecedent; and Incoherent marks confusing text not captured by the other language labels. Factual errors are claims judged incorrect: Bad Math covers arithmetic and unit or currency conversions; Encyclopedic covers known, externally verifiable facts the annotator knows to be wrong; and Commonsense covers violations of everyday knowledge or basic reasoning. Reader issues are not necessarily errors: Needs Google marks claims that an ordinary reader would have to look up, and Technical Jargon marks language requiring specialist expertise. Annotators are instructed not to look up Needs Google claims, so that label indicates a need for verification, not a verdict that a claim is false.

  3. Knowl 3 — The news-generation study compares model scale, training domain, and decoding

    experimental setup

    The study collected annotations for continuations of English news prompts. Prompts were the first sentences of Common Crawl news articles from January 2020, selected from articles with topic metadata. Each model generated 80–145 tokens, with generation stopped at the first detected sentence boundary after 80 tokens; human continuations used the remainder of the article under the same length rule. The compared sources were GPT-2 Small (117M parameters, WebText, no fine-tuning), GPT-2 XL (1.5B, WebText, no fine-tuning), Grover-Mega (1.5B, news-trained), GPT-3 DaVinci (175B), and human-written continuations. For the primary model comparisons, decoding used top-p =0.96=0.96, temperature =1.0=1.0, and frequency penalty =0=0. GPT-3 was additionally evaluated with 14 configurations: top-p in {0.4,0.7,0.9,0.96}\{0.4,0.7,0.9,0.96\} or temperature in {0.0,0.4,0.7,1.0}\{0.0,0.4,0.7,1.0\} varied independently, with frequency penalty either 0 or 1. The full collection comprised 1,308 paragraphs, 13,056 annotations, and 41,862 labeled spans.

  4. Knowl 4 — Scaling reduces some error types but leaves others at a plateau

    empirical result

    In comparisons using the shared top-p =0.96=0.96, temperature =1.0=1.0, no-frequency-penalty decoding setup, the plotted span-coverage trends show different relationships with model scale and training domain. Encyclopedic, Commonsense, and Incoherent errors decline with in-domain news training and larger models, with human text showing the fewest of these errors. Off-Prompt, Bad Math, and Grammar and Usage show a plateau in improvement by GPT-3: human text has fewer Off-Prompt and Grammar and Usage errors, while Bad Math appears saturated in this news setting. Self-Contradiction and Redundant errors have more complex trends: they rise for some medium- or large-scale models and fall for human text. GPT-2 Small’s low Self-Contradiction rate is interpreted cautiously because its generations are often so Incoherent or Off-Prompt that they offer fewer comprehensible claims that could contradict one another. Human-written text receives the most Needs Google and Technical Jargon spans; these labels denote reader obstacles rather than necessarily incorrect text.

  5. Knowl 5 — GPT-3 decoding settings substantially change annotated error rates

    empirical result

    For GPT-3, top-p and temperature affect error types in opposite directions when no frequency penalty is used: sampling from a larger set of words (higher top-p or temperature) is associated with more Off-Prompt spans but fewer Redundant spans. Adding frequency penalty 1 lowers overall error-span coverage at every tested top-p and temperature, and reverses the no-penalty trend: less diverse sampling then produces fewer errors. Across the tested settings, argmax sampling without frequency penalty is the worst GPT-3 configuration and performs worse than GPT-2 XL. Argmax sampling with frequency penalty 1 has the fewest worker-annotated error spans and appears comparable to human text on that unadjusted measure; the comparison does not establish that the two sources contain equally many genuine errors.

  6. Knowl 6 — Worker labels make best-decoding GPT-3 look closer to human text than manual review does

    empirical result

    The paper manually reviewed 160 randomly sampled spans: 10 from each of eight error types for GPT-3 using argmax with frequency penalty 1, and the corresponding samples for human-written text. Some human labels were triggered by Common Crawl scraping artifacts, such as inserted links or advertisements, rather than writing problems; analogous, rarer formatting artifacts in GPT-3 text were also excluded. Using the sampled legitimacy rates to adjust error-span counts, the authors estimated that 48% of GPT-3 worker-labeled errors were genuine, compared with 9% of worker-labeled errors in human-written articles. They conclude that human-written news paragraphs contain many times fewer genuine issues than the best tested GPT-3 configuration, despite similar unadjusted worker-label rates. The estimates are based on limited manual review by one author; the authors caution that annotation noise can be very high for high-quality text.

  7. Knowl 7 — A RoBERTa span classifier provides a baseline for automatic error detection

    model/method

    The paper formulates error detection as classifying a text span into one of the error types or No Error. Positive spans come from SCARECROW labels; negative spans are randomly sampled unlabeled spans, with three negatives for each error-span length represented in a generated text. The data split is by text: 1,063 training texts (28,029 error spans), 100 development texts (2,538 spans), and 100 test texts (2,677 spans). The classifier encodes each text with RoBERTa-large, represents a span using its endpoint encodings plus a learned span-length embedding, and feeds the representation to a feedforward classifier. It has 357M trainable parameters and is trained with cross-entropy for up to 15 epochs using AdamW at learning rate 10−610^{-6}; the lowest-validation-loss checkpoint was epoch 8. Evaluation considers every span up to 30 tokens and reports per-token precision, recall, and F1 against the union of 10 annotators’ spans. Compared with one annotator evaluated against the other nine, the model has higher precision only for Commonsense, but higher recall for several categories and higher F1 for five of ten categories. Model-versus-human F1 values are: Bad Math, not computable versus 0.24; Commonsense, 0.10 versus 0.04; Encyclopedic, not computable versus 0.05; Grammar and Usage, 0.26 versus 0.08; Incoherent, 0.43 versus 0.24; Off-Prompt, 0.41 versus 0.46; Redundant, 0.36 versus 0.50; Self-Contradiction, 0.12 versus 0.16; Technical Jargon, 0.29 versus 0.20; and Needs Google, 0.73 versus 0.32. For example, Needs Google reaches model precision 0.59 and recall 0.96, while individual-human precision and recall are 0.78 and 0.20.

  8. Knowl 8 — Span coverage measures the share of text marked by each error type

    definition

    The study’s main error quantity, span coverage, is the average amount of generated text covered by annotations of a given error type: span lengths are measured in tokens and normalized by generation length. Overlapping spans can contribute more than once, so coverage is not bounded by 1. The model-comparison analyses omit severity-1 Grammar and Usage spans because including them lowers agreement and increases variability. The authors also compare this measure with severity-weighted coverage and unweighted span counts; the latter can change model orderings for some error categories, whereas severity weighting mainly widens uncertainty intervals.

  9. Knowl 9 — Annotators agree well on many tokens, but overlap on positive spans is lower

    empirical result

    Token-level agreement statistics show Krippendorff’s α\alpha of 0.99 for Bad Math, 0.88 for Commonsense, 0.98 for Encyclopedic, 0.72 for Grammar and Usage, 0.73 for Incoherent, 0.71 for Off-Prompt, 0.88 for Redundant, and 0.87 for Self-Contradiction. The corresponding Two Agree rates—the percentage of tokens labeled by at least one annotator that were also labeled by at least one other—are 30%, 20%, 12%, 30%, 49%, 61%, 38%, and 26%, respectively. High α\alpha can be misleading for rare categories because most tokens are negative; Encyclopedic, for example, has α=0.98\alpha=0.98 but Two Agree of 12%. The paper treats individual annotators as relatively high-precision, low-recall judges and uses aggregate annotations to capture more labeled spans. Excluding severity-1 Grammar and Usage labels is important: including them lowers that category’s α\alpha to 0.56.

  10. Knowl 10 — Bootstrap analysis supports roughly 50 generations but flags rare-error uncertainty

    empirical result

    To test how annotation variability affects estimates from smaller samples, the authors repeatedly sampled 50 GPT-3 generations from the largest apples-to-apples condition, with 10 annotations per generation, and ran 1,000 bootstrap resamples. The total error-count estimate had a coefficient of variation of 19.72%. Category variability was higher for rare errors: Bad Math had a 44.5% coefficient of variation and Encyclopedic 29.1%, compared with 8.7% for Needs Google. Variability decreased as the number of generations increased, with the decline slowing after about 30 generations. The authors infer that analyses using around 50 generations can be relatively robust overall, while rare categories need aggregation and larger annotation budgets; they collected at least 500 annotations per studied condition.

Coverage note — The exploratory topic-specific error distributions and qualitative analysis of annotator explanations are omitted because they are secondary analyses rather than load-bearing parts of the framework or main model comparisons; proposed future applications and extensions are not empirical contributions of this study.

References

  1. 1.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72.
  2. 2.Gwern Branwen. 2020. Gpt-3 creative fiction.
  3. 3.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  4. 4.Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau, and Laurent Charlin. 2020. Language gans falling short.
  5. 5.Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. 2020. With little power comes great responsibility. In Proceedings of EMNLP.
  6. 6.Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Xiaodong Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2021. Extracting training data from large language models. In USENIX Security Symposium.
  7. 7.Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2021. Evaluation of text generation: A survey.
  8. 8.Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021. All that’s ‘human’ is not gold: Evaluating human evaluation of generated text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7282–7296, Online. Association for Computational Linguistics.
  9. 9.Elizabeth Clark and Noah A. Smith. 2021. Choose your own adventure: Paired suggestions in collaborative writing for evaluating story generation models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3566–3575, Online. Association for Computational Linguistics.
  10. 10.Liam Dugan, Daphne Ippolito, Arun Kirubarajan, and Chris Callison-Burch. 2020. Roft: A tool for evaluating human detection of machine-generated text. arXiv preprint arXiv:2010.03070.
  11. 11.Tianyu Gao, Adam Fisch, and Danqi Chen. 2020. Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723.
  12. 12.Herbert P Grice. 1975. Logic and conversation. In Speech acts, pages 41–58. Brill.
  13. 13.Jing Gu, Qing yang Wu, and Zhou Yu. 2021. Perception score: A learned metric for open-ended text generation evaluation. In AAAI.
  14. 14.Jian Guan and Minlie Huang. 2020. Union: An unreferenced metric for evaluating open-ended story generation. In EMNLP.
  15. 15.Tatsunori Hashimoto, Hugh Zhang, and Percy Liang. 2019. Unifying human and statistical evaluation for natural language generation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1689–1701, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi. 2018. Learning to write with cooperative discriminators. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1638–1649, Melbourne, Australia. Association for Computational Linguistics.
  17. 17.Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. International Conference on Learning Representations.
  18. 18.David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020. Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions. In Proceedings of the 13th International Conference on Natural Language Generation, pages 169–182, Dublin, Ireland. Association for Computational Linguistics.
  19. 19.Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, and Daniel S. Weld. 2021. Genie: A leaderboard for human-in-the-loop evaluation of text generation.
  20. 20.Klaus Krippendorff. 2018. Content analysis: An introduction to its methodology. Sage publications.
  21. 21.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  22. 22.Hugo Liu and Push Singh. 2004. Conceptnet—a practical commonsense reasoning tool-kit. BT technology journal, 22(4):211–226.
  23. 23.Yann Mathet, Antoine Widlöcher, and Jean-Philippe Métivier. 2015. The unified and holistic method gamma (γ) for inter-annotator agreement measure and alignment. Computational Linguistics, 41(3):437–479.
  24. 24.Christof Monz and Maarten de Rijke. 2001. Lightweight entailment checking for computational semantics. In Proc. of the third workshop on inference in computational semantics (ICoS-3).
  25. 25.Jekaterina Novikova, Ondrej Dusek, and Verena Rieser. 2018. Rankme: Reliable human ratings for natural language generation. In NAACL.
  26. 26.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318.
  27. 27.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Yejin Choi, and Zaid Harchaoui. 2021. Mauve: Human-machine divergence curves for evaluating open-ended text generation.
  28. 28.Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020. Stanza: A python natural language processing toolkit for many human languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 101–108, Online. Association for Computational Linguistics.
  29. 29.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  30. 30.Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–7.
  31. 31.Matthew Richardson, Christopher JC Burges, and Erin Renshaw. 2013. Mctest: A challenge dataset for the open-domain machine comprehension of text. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 193–203.
  32. 32.Ananya B Sai, Akash Kumar Mohankumar, and Mitesh M Khapra. 2020. A survey of evaluation metrics used for nlg systems. arXiv preprint arXiv:2008.12009.
  33. 33.Roger C Schank and Robert P Abelson. 1977. Scripts, plans, goals, and understanding: An inquiry into human knowledge structures. Psychology Press.
  34. 34.Hadrien Titeux and Rachid Riad. 2021. pygamma-agreement: Gamma γ measure for inter/intra-annotator agreement in python.
  35. 35.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762.
  36. 36.David Wadden, Ulme Wennberg, Yi Luan, and Hannaneh Hajishirzi. 2019. Entity, relation, and event extraction with contextualized span representations. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5784–5789, Hong Kong, China. Association for Computational Linguistics.
  37. 37.Gavin Wood, Kiel Long, Tom Feltwell, Scarlett Rowland, Phillip Brooker, Jamie Mahoney, John Vines, Julie Barnett, and Shaun Lawson. 2018. Rethinking engagement with online news through social and visual co-annotation. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, pages 1–12.
  38. 38.Rowan Zellers, Ari Holtzman, Elizabeth Clark, Lianhui Qin, Ali Farhadi, and Yejin Choi. 2021. TuringAdvice: A generative and dynamic evaluation of language use. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4856–4880, Online. Association for Computational Linguistics.
  39. 39.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending against neural fake news. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advancesin Neural Information Processing Systems 32, pages 9054–9065. Curran Associates, Inc.
  40. 40.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.

Citation

MLA
Dou, Y., et al. “Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7250–74, https://doi.org/10.18653/v1/2022.acl-long.501.
APA
Dou, Y., Forbes, M., Koncel-Kedziorski, R., Smith, N. A., & Choi, Y. (2022). Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7250–7274. https://doi.org/10.18653/v1/2022.acl-long.501
Chicago
Dou, Y., M. Forbes, R. Koncel-Kedziorski, N. A. Smith, and Y. Choi. 2022. “Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7250–74. https://doi.org/10.18653/v1/2022.acl-long.501.
Harvard
Dou, Y. et al. (2022) “Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7250–7274. Available at: https://doi.org/10.18653/v1/2022.acl-long.501.
Vancouver
1. Dou Y, Forbes M, Koncel-Kedziorski R, Smith NA, Choi Y (2022) Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7250–7274

BibTeX

@inproceedings{dou-etal-2022-gpt,
    title = "Is {GPT}-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text",
    author = "Dou, Yao  and
      Forbes, Maxwell  and
      Koncel-Kedziorski, Rik  and
      Smith, Noah A.  and
      Choi, Yejin",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.501/",
    doi = "10.18653/v1/2022.acl-long.501",
    pages = "7250--7274"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/