Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods

Kathleen C. FraserHillary DawkinsSvetlana Kiritchenko

article2025JAIR87 citations

Presents a comprehensive review of state-of-the-art AI-generated text detection methods, datasets, and practical factors that govern how reliably machine-written content can be identified across real-world scenarios.

Listen

The rapid advancement of large language models has made artificially generated text virtually indistinguishable from human writing to the human eye. This development poses critical risks across multiple sectors, including the automated spread of disinformation, academic dishonesty, fraud, and broader information ecosystem pollution. Distinguishing between human-authored and computer-generated content is now essential for establishing digital trust and security. The article comprehensively evaluates the state of AI-generated text detection, analyzing current technical methodologies, available benchmark datasets, and the key operational factors that influence how detectable artificial text is in real-world scenarios.

The article conducts an extensive, high-level synthesis of natural language processing literature, focusing on studies published through mid-2024. It categorizes text generation across a spectrum of human involvement ranging from arbitrary generation to collaborative writing, and assesses three core detection paradigms: embedded watermarking, statistical and stylistic feature analysis, and fine-tuned language model classifiers. The evaluation reviews numerous benchmark datasets spanning news, academic publications, and social media across multiple languages, while examining detector robustness against real-world constraints such as unknown source models, varying document lengths, and deliberate evasion tactics.

The analysis yields several crucial findings regarding detection efficacy. First, human evaluators perform near chance levels, frequently achieving around 50% to 59% accuracy, demonstrating that automated detection systems are mandatory. Second, detector accuracy drops significantly as the generating model becomes larger and when sophisticated decoding strategies like nucleus sampling are used. Third, text length represents a major performance boundary: while statistical and machine learning classifiers require at least 100 to 200 words to reach full capability, watermark detection can succeed with approximately 10 words. Fourth, current detectors exhibit poor generalizability, suffering accuracy drops of 10% to 25% or more when exposed to out-of-distribution domains, unseen prompts, or newer generating models. Finally, human-AI collaboration and adversarial attacks—such as machine paraphrasing, text polishing, and minor fact edits—severely degrade detection rates, often reducing statistical detector true positive rates to below 5%.

These findings indicate that relying on a single detection tool introduces severe operational, legal, and reputational risks. The widespread brittleness of current systems means organizations cannot treat automated binary classifications as absolute ground truth. Crucially, studies reveal that detectors exhibit systematic biases, such as falsely classifying essays by non-native English speakers as AI-generated at rates near 60%, raising urgent fairness and compliance concerns. Consequently, automated detection cannot be solely depended upon for high-stakes punitive decisions without rigorous governance.

Decision-makers should deploy multi-layered ensemble strategies rather than individual standalone tools, combining fine-tuned language model classifiers with precisely calibrated statistical filters set to maintain low false positive rates. When training internal detectors, organizations must assemble diverse, balanced datasets containing mixed text lengths, varied prompts, and hybrid human-AI examples. Furthermore, institutions should integrate human-in-the-loop workflows where automated detectors flag uncertainty rather than make unilateral determinations. Future technical and policy efforts must prioritize robust cross-lingual benchmarks, standardized watermark adoption across industry providers, and non-linguistic signals such as social network metadata to effectively counter evolving evasion techniques.

The conclusions of the article are constrained by the fact that most evaluated datasets were generated in artificial research settings rather than collected directly from in-the-wild web deployments. Additionally, as open-source models proliferate and techniques like low-rank adaptation make model personalization cheaper, the underlying statistical signatures of AI text will continue to shift. Leaders should maintain moderate confidence in existing detection capabilities for standard, long-form content from known model families, but exercise extreme caution when evaluating short texts, mixed-authorship documents, or adversarial content.

  • Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). Building upon the survey's detection taxonomy, this work evaluates the fundamental robustness limits and evasion attacks against state-of-the-art AI text detectors.
Cover for Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods

Abstract

Large language models (LLMs) have advanced to a point that even humans have difficulty discerning whether a text was generated by another human, or by a computer. However, knowing whether a text was produced by human or artificial intelligence (AI) is important to determining its trustworthiness, and has applications in many domains including detecting fraud and academic dishonesty, as well as combating the spread of misinformation and political propaganda. The task of AI-generated text (AIGT) detection is therefore both very challenging, and highly critical. In this survey, we summarize state-of-the art approaches to AIGT detection, including watermarking, statistical and stylistic analysis, and machine learning classification. We also provide information about existing datasets for this task. Synthesizing the research findings, we aim to provide insight into the salient factors that combine to determine how “detectable” AIGT text is under different scenarios, and to make practical recommendations for future work towards this significant technical and societal challenge.

Table of Contents

  • 1. Introduction
  • 2. The Task of AIGT Detection
  • 2.1 Text Classification and Generative AI
  • 2.2 Taxonomy of AI-Generated Text
  • 2.3 Detection Scenarios
  • 3. Current Approaches to AIGT Detection
  • 3.1 Watermarking
  • 3.2 Statistical and Stylistic Analysis
  • 3.2.1 Statistical Analysis
  • 3.2.2 Stylistic Analysis
  • 3.3 Language Model-Based Classification
  • 3.4 Off-the-Shelf Detection Tools
  • 3.5 Humans as Detectors
  • 4. AIGT Datasets
  • 5. Factors that Influence Detectability
  • 5.1 Properties of the Generating LLM
  • 5.2 Language
  • 5.3 Document Length
  • 5.4 Out-of-Distribution Domains and Generating Models
  • 5.5 Degree of Human Influence
  • 5.6 Adversarial Attacks
  • 6. Discussion and Summary
  • 7. Challenges and Future Directions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — A taxonomy of AI-generated text by human influence

    definition

    AI-generated text (AIGT) is not a homogeneous category. The survey organizes it into four classes ordered from least to greatest human control, with progressively greater expected detection difficulty:

    • Arbitrary generation: The AI determines both content and structure, as in “Please write a story.”
    • Guided generation: A human specifies a topic or intended message, while the AI produces the text, as in generating an article from a headline or asking for a particular claim about that headline.
    • Controlled generation: A human supplies the underlying content and the AI transforms it through paraphrasing, style transfer, summarization, translation, or polishing. Polishing makes only small changes to a human-written text, such as improving readability.
    • Collaborative generation: Human and AI contributions are combined through post-editing, documents containing separate human- and AI-written sections (“mixcase”), or partly automated “cyborg” accounts.

    The boundaries between these classes are fuzzy. Human intervention generally weakens detectable AI regularities, but it also increases the human time and effort required to produce the text.

  2. Knowl 2 — Detection scenarios and the information available to a detector

    definition

    AIGT detection depends on whether the generating language model is known and whether its internal probability information is accessible. The survey distinguishes four scenarios:

    • Known model: The detector knows the specific generating model, or knows that the text came from one model in a specified candidate set.
    • Unknown model: The detector has no reliable information about which model generated the text.
    • White-box access: The detector can access the generating model’s parameters or token-level output probabilities. This is possible only in the known-model setting.
    • Black-box access: The detector can query a known model and observe its outputs but cannot access its parameters or output probabilities.

    In the known-model/white-box setting, watermark extraction and probability-based measures such as token likelihood, rank, perplexity, entropy, and probability curvature are available. In the known-model/black-box setting, regeneration tests, stylistic classifiers, and supervised classifiers can be used. In the unknown-model setting, detectors must use proxy open-source models, classifiers trained on other generators, or commercial tools; performance is generally lower because the detector relies on properties expected to transfer across models. The scenario map reproduced in Figure 5 on page 21 summarizes this allocation of methods.

  3. Knowl 3 — Watermarking approaches and their robustness trade-offs

    model/method

    Watermarking embeds a secret signal during generation so that possession of the extraction procedure can identify AI-produced text without requiring ordinary linguistic classification. In the red-green-list method, a hash of the preceding token seeds a partition of the vocabulary into green and red lists at every position. The generator biases the next-token distribution toward green-list tokens, and a statistical test on the observed green-token frequency detects the watermark. Longer texts make false-positive probabilities extremely small.

    The sequential red-green construction is vulnerable to word deletion, insertion, synonym replacement, and paraphrasing because any change can alter the subsequent vocabulary partitions. Fixed red-green lists improve robustness to local perturbations, but their statistical bias makes the list easier to infer, remove, or forge through repeated black-box queries. Distortion-free schemes avoid changing the token probability distribution by using a secret key to select among otherwise valid stochastic generations; they improve imperceptibility but are intrinsically fragile to text modification unless a soft matching function is added. The survey notes that soft matching has been experimentally shown to improve robustness to paraphrasing.

    Black-box, post-generation watermarks based on formatting, Unicode substitutions, syntax, or synonym replacement can be applied without access to the generating model, but formatting signals are removable by canonicalization and lexical replacement may damage quality. Combining lexical and syntactic signals provides some complementary robustness, although paraphrasing can alter both. Neural regeneration systems such as Remark-LLM improve coherence and can train on malicious examples for robustness, but require more text to encode the same amount of information than rule-based methods. Overall, robust black-box watermarking remains unresolved, and watermark detection is usable only when the text was watermarked and the extraction method is known.

  4. Knowl 4 — Statistical and stylistic feature-based detection

    model/method

    Feature-based detectors use measurable properties of a text rather than directly fine-tuning a large classifier. Statistical methods typically exploit the tendency of generated text to contain high-probability, low-rank tokens and therefore lower perplexity, entropy, or surprisal than human writing under the relevant language model. A threshold on such a feature can be calibrated to prioritize a low false-positive rate, but practical threshold selection still requires domain- and model-specific calibration data even when the method is called “zero-shot.”

    More structural statistical methods include DetectGPT, which averages the change in a model’s log probability after perturbing or paraphrasing the input, and NPR, which uses average perturbed token rank. Both exploit the reported tendency of AIGT to lie in negative-curvature regions of the model’s probability function. Regeneration methods instead retain part of the text as a prompt, regenerate continuations, and compare the new outputs with the original; they can work with black-box access but depend strongly on the number of regenerations and the truncation ratio. Intrinsic-dimension estimation reported approximate dimensions of 9–10 for human text and 8 for AIGT across several genres and models, with apparent robustness to paraphrasing, although the survey states that evidence is insufficient to establish universality.

    When the generating model is unknown, proxy models can provide comparative features. Sniffer uses contrastive probabilities from pairs of open models, SeqXGPT uses probability lists for sentence-level mixed-authorship detection, and Ghostbuster combines token-probability features from weak models ranging from unigram models to GPT-3 variants before fitting logistic regression. Comparative proxy features are generally more informative than a single model-dependent score and can extend detection to unseen generators.

    Stylistic detectors operate without model access and measure lexical and syntactic diversity, repetition, coherence, discourse structure, punctuation, phraseology, readability, and other general-purpose properties. The most discriminative features reported in earlier studies were syntactic and lexical diversity together with general-purpose features, but transfer across decoding strategies, models, and domains was poor. Domain expectations matter: AI may appear unusually formal and correct in casual writing, whereas in professional news it may instead be detected through restricted vocabulary, weaker coherence, or other deviations from the expected style.

  5. Knowl 5 — Pre-trained language-model classifiers and human detection limits

    model/method

    Pre-trained language-model classifiers receive the raw text and learn discriminative features during fine-tuning rather than relying on a manually specified feature set. Typical systems attach a classification head to BERT or RoBERTa and fine-tune it on paired human and AI text; other systems use sequence-to-sequence models such as T5. Their representations commonly contain 256–768 dimensions, so they usually require substantially more labeled data than low-dimensional statistical or stylistic classifiers.

    The survey reports that early RoBERTa classifiers detected GPT-2 text with 95% accuracy, and that training on multiple decoding strategies and text lengths improved robustness. A T5 classifier generally outperformed comparable RoBERTa-based classification and both surpassed an OpenAI detector and GPT-2 used directly as a detector. For short messages, a multiscale positive-unlabelled formulation treats definitely AI-generated examples as positive and ambiguous short texts as unlabelled; this improved BERT and RoBERTa detection in English and Chinese. Adversarial-training systems such as RADAR improve robustness by jointly training a paraphraser to evade and a detector to recognize the resulting text.

    Despite these computational gains, human readers usually perform near random guessing when classifying isolated AI- and human-written passages. Reported exceptions are task-dependent: linguistics students averaged 59% accuracy in one study, English-language teachers reached 61% on AI-generated student essays and 67% after limited self-training, and paired human-versus-ChatGPT responses can be easier for experts. These results support the survey’s conclusion that computational detectors can exploit statistical properties that are not reliably perceptible to human observers.

  6. Knowl 6 — Generating-model size, decoding strategy, and sample length control detectability

    empirical result

    The survey synthesizes evidence that larger and better-generating language models are harder to detect. An AI Detectability Index assigns higher detectability-related scores to larger models whose perplexity and probability-curvature statistics approach human distributions. In experiments where detector and generator capacity were matched, detection accuracy decreased approximately linearly as the number of generator parameters increased exponentially.

    Decoding mismatch is also a major source of failure. When a detector trained on top-kk output was tested on nucleus-sampled output, accuracy decreased by 21% relative to in-distribution evaluation. Nucleus sampling was generally the hardest decoding strategy to detect and produced detectors with the best transfer to other sampling strategies. Even within one strategy, changing parameters reduced recall: for nucleus sampling, changing the probability threshold from 0.96 to 0.8 reduced recall by 13%; for top-kk sampling, changing kk from 40 to 160 reduced recall by 56.4%.

    Text length supplies the evidence required by most non-watermark detectors. Watermarks can carry a binary signal with roughly 10 words, whereas statistical, stylistic, and fine-tuned classifiers generally need at least 100 words. Approximately 120 words were sufficient for several statistical and fine-tuned detectors to reach their full potential across story, news, and scientific-writing domains, while approximately 200 words were sufficient in one study for ChatGPT-turbo and GPT-4. A theoretical analysis reported that around 500 words can suffice even when human and AI distributions are very close, provided the distributions are not identical. Concatenating ten short tweets improved a fine-tuned detector from 80% to nearly 100% accuracy. Length effects interact with decoding: shortening nucleus-sampled text from 256 to 64 words caused a 10% accuracy drop, compared with less than 1% for top-kk text. Training must also include short examples when sentence-level detection is required; in one benchmark, sentence-level accuracy rose from 50% to 93% after such examples were added.

  7. Knowl 7 — Language and distribution shift limit cross-domain detection

    empirical result

    AIGT detectors generalize poorly when the language, domain, generating model, or prompt differs between training and testing. Cross-lingual detection with XLM-RoBERTa was generally difficult, with only limited transfer between English and Chinese; consequently, AIGT in low-resource languages remains comparatively difficult to detect. Detectors also exhibit a fairness problem for human writers: publicly available tools achieved almost perfect accuracy on U.S. eighth-grade essays but produced a false-positive rate close to 60% on essays written in English by Chinese learners. Asking ChatGPT to make word choices sound more native reduced the probability that both human and AI texts were classified as AI, illustrating that the same linguistic properties can affect detection and evasion.

    On multi-domain data, statistical and stylistic methods lost 10–25 percentage points in accuracy even when all domains had been represented during training, with an additional 1–15 point loss on unseen domains. A fine-tuned Longformer lost more than 20% in the out-of-distribution setting. Fine-tuned RoBERTa often had the strongest in-domain performance but the weakest cross-domain transfer, whereas ELECTRA showed greater robustness in one comparison. Statistical detectors generally transferred poorly to unseen generating models; language-model classifiers transferred somewhat better but remained unsatisfactory in some settings, with transfer depending on text length and generator size. Prompt wording is another distribution shift: even semantically similar unseen prompts can reduce performance, especially in creative writing and question answering. Small amounts of new-domain data can enable transfer learning, but broad generalization is not automatic.

  8. Knowl 8 — Human editing and hybrid authorship sharply reduce detectability

    empirical result

    The degree of human influence is a primary determinant of detectability. In a comparison of human-written, fully GPT-written, GPT-completed, and GPT-polished scientific abstracts, GPT-polished text—human writing rewritten by ChatGPT for clarity—was the hardest category to detect; several commercial detectors performed worse than random guessing. Human-written scientific papers paraphrased by ChatGPT could be detected at 75% accuracy when a RoBERTa detector was trained with human–AI paraphrase examples, showing that training conditions strongly affect this result.

    Experiments on mixed human–AI text found that statistical and fine-tuned detectors performed poorly on binary detection when mixcase examples were absent from training. AI paraphrasing and simulated humanizing were the most difficult mixcase categories, while human adaptation for fluency was less damaging to detection. Fine-tuned detectors were especially sensitive to typos, and threshold-based detectors were especially sensitive to AI paraphrasing; sentence-level AI polishing weakened detection more than word-level polishing. When mixed-case examples were included in training, RADAR achieved approximately 88% accuracy across the mixcase categories.

    Hybrid misinformation is especially difficult when an AI changes only a few facts in an otherwise trustworthy human article. Machine translation also shifts detector behavior: in one small study, about 95% of original human texts were classified as human, compared with about 70% after translation into English. The survey cautions that the translation sample was too small for a conclusive estimate.

  9. Knowl 9 — Adversarial perturbation and paraphrasing attacks expose detector brittleness

    limitation

    AIGT detectors can be evaded by modifying text while attempting to preserve its meaning. A model-aware word attack identifies high-confidence, non-stop-word tokens under the generating model, proposes embedding-similar synonyms with matching part of speech, checks sentence-level semantic similarity, and selects a candidate with low probability under the generator. This attack readily fools probability-, rank-, and GLTR-style detectors, although a detector based on global factual structure was reported to be more robust. Character typos and homoglyph substitutions also damage statistical detectors but can often be removed by inference-time preprocessing.

    Full-document paraphrasing is more powerful. With a semantically faithful paraphrasing model, detection at a fixed 1% false-positive rate fell below 5% true-positive rate for statistical detectors, to 13% for a fine-tuned classifier, and to 50% for watermarking. Iterating paraphrasing five or six times further increased evasion, with only slight degradation in measured text quality. These attacks do not need queries to the detector when the paraphraser is an unseen, unprotected model; they primarily exploit poor transfer across generators.

    More difficult attacks assume that auxiliary models are also protected. Query-free word substitutions can be random, while query-based substitutions can use evolutionary search against a particular detector. Prompt tuning can produce an equally valid continuation of the original prompt that evades a detector, even when the text need not preserve the original wording or semantics. Finally, watermarking can be forged rather than erased: approximately one million black-box queries were reported as sufficient to infer red-green labels for the 181 most common English words and spoof a red-green watermark.

  10. Knowl 10 — Dataset design and deployment recommendations for robust detection

    experimental setup

    The survey finds that detector quality depends heavily on the match between training data and the intended deployment setting. Available datasets are predominantly English and are concentrated in news, academic writing, essays, misinformation, and multi-domain benchmarks; most are research-generated from a human “anchor” corpus rather than collected from the open internet. Examples in the survey include RAID with 6200K instances across eight domains, 11 models, and four decoding strategies; GPABench2 with 2800K academic-publication instances; OpenOrca with 4200K question-answering instances; CT2 with 1600K news instances; In-the-wild with 447K multi-domain instances; M4 with 150K multilingual multi-domain instances; HC3-SI with 215K instances; and Ghostbuster with 21K instances. Reported sizes combine human and AI examples and may hide class imbalance or concentration in one model or language.

    For a practical detector, the survey recommends generating data in each target language, using separate monolingual detectors when cross-lingual transfer is unreliable; matching the target domain when it is known and mixing domains otherwise; balancing human and AI length distributions while including sentence-level examples; using many human authors with varied styles and language proficiency; and generating AI text from many prompts rather than one prompt. Prompt tuning or filtering should make synthetic text resemble the human target distribution, for example using MAUVE, and training should include mixed human–AI examples whenever such cases matter. Nucleus sampling over a range of parameter values is recommended because it is difficult to detect and transfers relatively well.

    For deployment, no single method is reliable across all conditions. Threshold-based statistical detectors are valuable in narrowly calibrated known-model settings because their false-positive rate can be controlled, while language-model classifiers and proxy-model ensembles provide broader but less certain coverage. The survey therefore recommends combining differently calibrated statistical detectors for specialized cases with language-model-based detectors for unknown-model cases, and aggregating multiple tools rather than treating any one detector as authoritative.

Coverage note — The detailed inventory of off-the-shelf tools, the full dataset catalogue, and the longer regulatory, fairness, explainability, and multimedia future-work discussion were omitted because they are mainly catalogues or prospective directions rather than load-bearing findings for reconstructing the survey’s detection framework and empirical synthesis.

References

  1. 1.Aaditya Bhat (2023). Gpt-wiki-intro (revision 0e458f5). Hugging Face, https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.
  2. 2.Abdalla, M. H. I., Malberg, S., Dementieva, D., Mosca, E., & Groh, G. (2023). A benchmark dataset to distinguish human-written and machine-generated scientific papers. Information, 14 (10), 522.
  3. 3.Akram, A. (2023). An empirical study of ai-generated text detection tools. Advances in Machine Learning & Artificial Intelligence, 4 (2), 44–55.
  4. 4.Ardito, C. G. (2024). Contra generative AI detection in higher education assessments. In New Directions for Teaching and Learning special issue: Integrating Generative AI in the Design of Assessment.
  5. 5.Bao, G., Zhao, Y., Teng, Z., Yang, L., & Zhang, Y. (2024). Fast-detectGPT: Efficient zero-shot detection of machine-generated text via conditional probability curvature. In Proceedings of the Twelfth International Conference on Learning Representations (ICRL).
  6. 6.Beltagy, I., Peters, M. E., & Cohan, A. (2020). Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150.
  7. 7.Brassil, J., Low, S., Maxemchuk, N., & O’Gorman, L. (1995). Electronic marking and identification techniques to discourage document copying. IEEE Journal on Selected Areas in Communications, 13 (8), 1495–1504.
  8. 8.Chakraborty, M., Tonmoy, S. T. I., Zaman, S. M., Gautam, S., Kumar, T., Sharma, K., Barman, N., Gupta, C., Jain, V., Chadha, A., et al. (2023). Counter Turing Test (CT2): AI-generated text detection is not as easy as you may think-introducing AI detectability index (ADI). In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 2206–2239.
  9. 9.Chen, C., & Shu, K. (2023). Can LLM-generated misinformation be detected?. In Proceedings of the NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following.
  10. 10.Chen, Y., Kang, H., Zhai, V., Li, L., Singh, R., & Raj, B. (2023a). Token prediction as implicit classification to identify LLM-generated text. In Bouamor, H., Pino, J., & Bali, K. (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 13112–13120, Singapore. Association for Computational Linguistics.
  11. 11.Chen, Y., Kang, H., Zhai, V., Li, L., Singh, R., & Ramakrishnan, B. (2023b). GPT-Sentinel: Distinguishing human and ChatGPT generated content. arXiv preprint arXiv:2305.07969.
  12. 12.Choi, E. C., & Ferrara, E. (2024). FACT-GPT: Fact-checking augmentation via claim matching with LLMs. In Companion Proceedings of the ACM on Web Conference 2024, p. 883–886, New York, NY, USA. Association for Computing Machinery.
  13. 13.Christ, M., Gunn, S., & Zamir, O. (2024). Undetectable watermarks for language models. In The Thirty Seventh Annual Conference on Learning Theory, pp. 1125–1139. PMLR.
  14. 14.Cresci, S. (2020). A decade of social bot detection. Communications of the ACM, 63 (10), 72–83.
  15. 15.Crothers, E., Japkowicz, N., Viktor, H., & Branco, P. (2022). Adversarial robustness of neural-statistical features in detection of generative transformers. In Proceedings of the 2022 International Joint Conference on Neural Networks (IJCNN), pp. 1–8. IEEE.
  16. 16.Crothers, E., Japkowicz, N., & Viktor, H. L. (2023). Machine-generated text: A comprehensive survey of threat models and detection methods. IEEE Access, 11.
  17. 17.Cui, W., Zhang, L., Wang, Q., & Cai, S. (2023). Who said that? benchmarking social media AI detection. arXiv preprint arXiv:2310.08240.
  18. 18.Cutler, J., Dugan, L., Havaldar, S., & Stein, A. (2021). Automatic detection of hybrid human-machine text boundaries. Tech. rep., University of Pennsylvania.
  19. 19.Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Burstein, J., Doran, C., & Solorio, T. (Eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  20. 20.Dugan, L., Hwang, A., Trhl´ık, F., Zhu, A., Ludan, J. M., Xu, H., Ippolito, D., & Callison-Burch, C. (2024). RAID: A shared benchmark for robust evaluation of machine-generated text detectors. In Ku, L.-W., Martins, A., & Srikumar, V. (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12463–12492, Bangkok, Thailand. Association for Computational Linguistics.
  21. 21.Fagni, T., Falchi, F., Gambini, M., Martella, A., & Tesconi, M. (2021). TweepFake: About detecting deepfake tweets. PLOS ONE, 16 (5), e0251415.
  22. 22.Fr¨ohling, L., & Zubiaga, A. (2021). Feature-based detection of automated language models: tackling gpt-2, gpt-3 and grover. PeerJ Computer Science, 7, e443.
  23. 23.Gehrmann, S., Strobelt, H., & Rush, A. M. (2019). GLTR: Statistical detection and visualization of generated text. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL). Association for Computational Linguistics.
  24. 24.Ghosal, S. S., Chakraborty, S., Geiping, J., Huang, F., Manocha, D., & Bedi, A. S. (2023). Towards possibilities & impossibilities of AI-generated text detection: A survey. Transactions on Machine Learning Research.
  25. 25.Guo, B., Zhang, X., Wang, Z., Jiang, M., Nie, J., Ding, Y., Yue, J., & Wu, Y. (2023). How close is ChatGPT to human experts? Comparison corpus, evaluation, and detection. arXiv preprint arXiv:2301.07597.
  26. 26.Hajian-Tilaki, K. (2013). Receiver operating characteristic (ROC) curve analysis for medical diagnostic test evaluation. Caspian Journal of Internal Medicine, 4 (2), 627.
  27. 27.Hans, A., Schwarzschild, A., Cherepanova, V., Kazemi, H., Saha, A., Goldblum, M., Geiping, J., & Goldstein, T. (2024). Spotting llms with binoculars: zero-shot detection of machine-generated text. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org.
  28. 28.He, X., Shen, X., Chen, Z., Backes, M., & Zhang, Y. (2024). MGTBench: Benchmarking machine-generated text detection. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pp. 2251–2265.
  29. 29.Henrique, D. S. G., Kucharavy, A., & Guerraoui, R. (2023). Stochastic parrots looking for stochastic parrots: LLMs are easy to fine-tune and hard to detect with other LLMs. arXiv preprint arXiv:2304.08968.
  30. 30.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations.
  31. 31.Hu, X., Chen, P.-Y., & Ho, T.-y. (2023). RADAR: Robust AI-text detection via adversarial learning. In Proceedings of the Annual Conference on Neural Information Processing Systems.
  32. 32.Huang, K.-H., McKeown, K., Nakov, P., Choi, Y., & Ji, H. (2023). Faking fake news for real fake news detection: Propaganda-loaded training data generation. In Rogers, A., Boyd-Graber, J., & Okazaki, N. (Eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 14571–14589, Toronto, Canada. Association for Computational Linguistics.
  33. 33.Huang, Y., & Sun, L. (2023). FakeGPT: Fake news generation, explanation and detection of large language models. arXiv preprint arXiv:2310.05046.
  34. 34.Jawahar, G., Sagot, B., & Seddah, D. (2019). What does BERT learn about the structure of language?. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL).
  35. 35.Jiang, B., Tan, Z., Nirmal, A., & Liu, H. (2024). Disinformation detection: An evolving challenge in the age of LLMs. In Proceedings of the 2024 SIAM International Conference on Data Mining (SDM).
  36. 36.Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., & Goldstein, T. (2023). A watermark for large language models. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., & Scarlett, J. (Eds.), Proceedings of the 40th International Conference on Machine Learning (ICML), Vol. 202 of Proceedings of Machine Learning Research, pp. 17061–17084. PMLR.
  37. 37.Knott, A., Pedreschi, D., Chatila, R., Chakraborti, T., Leavy, S., Baeza-Yates, R., Eyers, D., Trotman, A., Teal, P. D., Biecek, P., et al. (2023). Generative AI models should include detection mechanisms as a condition for public release. Ethics and Information Technology, 25 (4), 55.
  38. 38.Koike, R., Kaneko, M., & Okazaki, N. (2024). OUTFOX: LLM-generated essay detection through in-context learning with adversarially generated examples. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence (AAAI-24), pp. 21258–21266.
  39. 39.Kong, L., Jiang, H., Zhuang, Y., Lyu, J., Zhao, T., & Zhang, C. (2020). Calibrated language model fine-tuning for in-and out-of-distribution data. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1326–1340.
  40. 40.Kreps, S., McCain, R. M., & Brundage, M. (2022). All the news that’s fit to fabricate: AI-generated text as a tool of media misinformation. Journal of Experimental Political Science, 9 (1), 104–117.
  41. 41.Krishna, K., Song, Y., Karpinska, M., Wieting, J., & Iyyer, M. (2023). Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS).
  42. 42.Kuditipudi, R., Thickstun, J., Hashimoto, T., & Liang, P. (2024). Robust distortion-free watermarks for language models. Transactions on Machine Learning Research.
  43. 43.Kumarage, T., Bhattacharjee, A., Padejski, D., Roschke, K., Gillmor, D., Ruston, S., Liu, H., & Garland, J. (2023a). J-guard: Journalism guided adversarially robust detection of AI-generated news. In Park, J. C., Arase, Y., Hu, B., Lu, W., Wijaya, D., Purwarianti, A., & Krisnadhi, A. A. (Eds.), Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 484–497, Nusa Dua, Bali. Association for Computational Linguistics.
  44. 44.Kumarage, T., Garland, J., Bhattacharjee, A., Trapeznikov, K., Ruston, S., & Liu, H. (2023b). Stylometric detection of AI-generated text in Twitter timelines. arXiv preprint arXiv:2303.03697.
  45. 45.Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., K¨uttler, H., Lewis, M., Yih, W.-t., Rockt¨aschel, T., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.
  46. 46.Li, L., Wang, P., Ren, K., Sun, T., & Qiu, X. (2023). Origin tracing and detecting of LLMs. arXiv preprint arXiv:2304.14072.
  47. 47.Li, Y., Li, Q., Cui, L., Bi, W., Wang, Z., Wang, L., Yang, L., Shi, S., & Zhang, Y. (2024). MAGE: Machine-generated text detection in the wild. In Ku, L.-W., Martins, A., & Srikumar, V. (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 36–53, Bangkok, Thailand. Association for Computational Linguistics.
  48. 48.Lian, W., Goodson, B., Pentland, E., Cook, A., Vong, C., & ”Teknium” (2023). OpenOrca: An open dataset of GPT augmented FLAN reasoning traces. https://huggingface.co/Open-Orca/OpenOrca.
  49. 49.Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns (N Y), 4 (7), 100779.
  50. 50.Lin, L., Gupta, N., Zhang, Y., Ren, H., Liu, C.-H., Ding, F., Wang, X., Li, X., Verdoliva, L., & Hu, S. (2024). Detecting multimedia generated by large AI models: A survey. arXiv preprint arXiv:2402.00045.
  51. 51.Liu, A., Pan, L., Lu, Y., Li, J., Hu, X., Zhang, X., Wen, L., King, I., Xiong, H., & Yu, P. (2024). A survey of text watermarking in the era of large language models. ACM Computing Surveys, 57 (2), 1–36.
  52. 52.Liu, X., Zhang, Z., Wang, Y., Pu, H., Lan, Y., & Shen, C. (2023a). CoCo: Coherence-enhanced machine-generated text detection under low resource with contrastive learning. In Bouamor, H., Pino, J., & Bali, K. (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 16167–16188, Singapore. Association for Computational Linguistics.
  53. 53.Liu, Y., Zhang, Z., Zhang, W., Yue, S., Zhao, X., Cheng, X., Zhang, Y., & Hu, H. (2023b). ArguGPT: evaluating, understanding and identifying argumentative essays generated by GPT models. arXiv preprint arXiv:2304.07666.
  54. 54.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  55. 55.Liu, Z., Yao, Z., Li, F., & Luo, B. (2024). On the detectability of ChatGPT content: Benchmarking, methodology, and evaluation through the lens of academic writing. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pp. 2236–2250.
  56. 56.Lucas, J., Uchendu, A., Yamashita, M., Lee, J., Rohatgi, S., & Lee, D. (2023). Fighting fire with fire: The dual role of LLMs in crafting and detecting elusive disinformation. In Bouamor, H., Pino, J., & Bali, K. (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 14279–14305, Singapore. Association for Computational Linguistics.
  57. 57.Macko, D., Moro, R., Uchendu, A., Lucas, J., Yamashita, M., Pikuliak, M., Srba, I., Le, T., Lee, D., Simko, J., & Bielikova, M. (2023). MULTITuDE: Large-scale multilingual machine-generated text detection benchmark. In Bouamor, H., Pino, J., & Bali, K. (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 9960–9987, Singapore. Association for Computational Linguistics.
  58. 58.Mazar, N., Amir, O., & Ariely, D. (2008). The dishonesty of honest people: A theory of self-concept maintenance. Journal of Marketing Research, 45 (6), 633–644.
  59. 59.Mireshghallah, N., Mattern, J., Gao, S., Shokri, R., & Berg-Kirkpatrick, T. (2024). Smaller language models are better zero-shot machine-generated text detectors. In Graham, Y., & Purver, M. (Eds.), Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 278–293, St. Julian’s, Malta. Association for Computational Linguistics.
  60. 60.Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., & Finn, C. (2023). DetectGPT: Zero-shot machine-generated text detection using probability curvature. In Proceedings of the 40th International Conference on Machine Learning (ICML).
  61. 61.Mu˜noz-Ortiz, A., G´omez-Rodr´ıguez, C., & Vilares, D. (2023). Contrasting linguistic patterns in human and LLM-generated text. arXiv preprint arXiv:2308.09067.
  62. 62.Nguyen-Son, H.-Q., Tieu, N.-D.-T., Nguyen, H. H., Yamagishi, J., & Zen, I. E. (2017). Identifying computer-generated text using statistical analysis. In Proceedings of the Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), pp. 1504–1511.
  63. 63.OpenAI (2022). Introducing ChatGPT. OpenAI blog.
  64. 64.Orenstrakh, M. S., Karnalim, O., Suarez, C. A., & Liut, M. (2024). Detecting LLM-generated text in computing education: Comparative study for ChatGPT cases. In 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC), pp. 121–126. IEEE.
  65. 65.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., & Oh, A. (Eds.), Advances in Neural Information Processing Systems, Vol. 35, pp. 27730–27744. Curran Associates, Inc.
  66. 66.Pagnoni, A., Graciarena, M., & Tsvetkov, Y. (2022). Threat scenarios and best practices to detect neural fake news. In Proceedings of the 29th International Conference on Computational Linguistics, pp. 1233–1249.
  67. 67.Pan, Y., Pan, L., Chen, W., Nakov, P., Kan, M.-Y., & Wang, W. (2023). On the risk of misinformation pollution with large language models. In Bouamor, H., Pino, J., & Bali, K. (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 1389–1403, Singapore. Association for Computational Linguistics.
  68. 68.Pillutla, K., Swayamdipta, S., Zellers, R., Thickstun, J., Welleck, S., Choi, Y., & Harchaoui, Z. (2021). MAUVE: Measuring the gap between neural text and human text using divergence frontiers. In Proceedings of the Annual Conference on Neural Information Processing Systems.
  69. 69.Pu, J., Sarwar, Z., Abdullah, S. M., Rehman, A., Kim, Y., Bhattacharya, P., Javed, M., & Viswanath, B. (2023). Deepfake text detection: Limitations and opportunities. In Proceedings of the IEEE Symposium on Security and Privacy (SP), pp. 1613–1630. IEEE.
  70. 70.Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. OpenAI Blog.
  71. 71.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21 (140), 1–67.
  72. 72.Rizzo, S. G., Bertini, F., & Montesi, D. (2016). Content-preserving text watermarking through unicode homoglyph substitution. In Proceedings of the 20th International Database Engineering & Applications Symposium, IDEAS ’16, p. 97–104, New York, NY, USA. Association for Computing Machinery.
  73. 73.Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., & Feizi, S. (2023). Can AI-generated text be reliably detected?. arXiv preprint arXiv:2303.11156.
  74. 74.Sarvazyan, A. M., Gonz´alez, J. A., Franco-Salvador, M., Rangel, F., Chulvi, B., & Rosso, ´ P. (2023). Overview of AuTexTification at IberLEF 2023: Detection and attribution of machine-generated text in multiple domains. Procesamiento del Lenguaje Natural, 71, 275–288.
  75. 75.Schuster, T., Schuster, R., Shah, D. J., & Barzilay, R. (2020). The limitations of stylometry for detecting machine-generated fake news. Computational Linguistics, 46 (2), 499–510.
  76. 76.Shi, Z., Wang, Y., Yin, F., Chen, X., Chang, K.-W., & Hsieh, C.-J. (2024). Red teaming language model detectors with language models. Transactions of the Association for Computational Linguistics, 12, 174—-189.
  77. 77.Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., Krueger, G., Kim, J. W., Kreps, S., et al. (2019). Release strategies and the social impacts of language models. arXiv preprint arXiv:1908.09203.
  78. 78.Soto, R. R., Koch, K., Khan, A., Chen, B., Bishop, M., & Andrews, N. (2024). Few-shot detection of machine-generated text using style representations. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR).
  79. 79.Spitale, G., Biller-Andorno, N., & Germani, F. (2023). AI model GPT-3 (dis)informs us better than humans. Science Advances, 9 (26).
  80. 80.Stieglitz, S., Brachten, F., Berthel´e, D., Schlaus, M., Venetopoulou, C., & Veutgen, D. (2017). Do social bots (still) act different to humans? – comparing metrics of social bots with those of humans. In Meiselwitz, G. (Ed.), Social Computing and Social Media. Human Behavior, pp. 379–395, Cham. Springer International Publishing.
  81. 81.Stiff, H., & Johansson, F. (2022). Detecting computer-generated disinformation. International Journal of Data Science and Analytics, 13 (4), 363–383.
  82. 82.Su, J., Zhuo, T., Wang, D., & Nakov, P. (2023a). DetectLLM: Leveraging log rank information for zero-shot detection of machine-generated text. In Bouamor, H., Pino, J., & Bali, K. (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 12395–12412, Singapore. Association for Computational Linguistics.
  83. 83.Su, J., Zhuo, T. Y., Mansurov, J., Wang, D., & Nakov, P. (2023b). Fake news detectors are biased against texts generated by large language models. arXiv preprint arXiv:2309.08674.
  84. 84.Su, Z., Wu, X., Zhou, W., Ma, G., & Hu, S. (2023c). HC3 Plus: A semantic-invariant human ChatGPT comparison corpus. In Proceedings of the Fifth Workshop on Knowledge-driven Analytics and Systems impacting Human Quality of Life (KDAH-CIKM-2023).
  85. 85.Tang, R., Chuang, Y.-N., & Hu, X. (2024). The science of detecting LLM-generated texts. Communications of the ACM, 67, 50––59.
  86. 86.Tian, E. (2023). Identifying GPT: First principles for generative AI detection. M.Sc. Thesis, Princeton University.
  87. 87.Tian, Y., Chen, H., Wang, X., Bai, Z., Zhang, Q., Li, R., Xu, C., & Wang, Y. (2024). Multiscale positive-unlabeled detection of AI-generated texts. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR).
  88. 88.Topkara, M., Topkara, U., & Atallah, M. J. (2006a). Words are not enough: sentence level natural language watermarking. In Proceedings of the 4th ACM International Workshop on Contents Protection and Security, MCPS ’06, p. 37–46, New York, NY, USA. Association for Computing Machinery.
  89. 89.Topkara, U., Topkara, M., & Atallah, M. J. (2006b). The hiding virtues of ambiguity: quantifiably resilient watermarking of natural language text through synonym substitutions. In Proceedings of the 8th Workshop on Multimedia and Security, MM&Sec ’06, p. 164–174, New York, NY, USA. Association for Computing Machinery.
  90. 90.Tulchinskii, E., Kuznetsov, K., Kushnareva, L., Cherniavskii, D., Nikolenko, S., Burnaev, E., Barannikov, S., & Piontkovskaya, I. (2023). Intrinsic dimension estimation for robust detection of ai-generated texts. In Proceedings of the 37th International Conference on Neural Information Processing Systems, pp. 39257—-39276.
  91. 91.Uchendu, A., Le, T., & Lee, D. (2023). Attribution and obfuscation of neural text authorship: A data mining perspective. ACM SIGKDD Explorations Newsletter, 25 (1), 1–18.
  92. 92.Uchendu, A., Ma, Z., Le, T., Zhang, R., & Lee, D. (2021). TURINGBENCH: A benchmark environment for Turing test in the age of neural text generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 2001–2016.
  93. 93.Venkatraman, S., Uchendu, A., & Lee, D. (2024). GPT-who: An information density-based machine-generated text detector. In Duh, K., Gomez, H., & Bethard, S. (Eds.), Findings of the Association for Computational Linguistics: NAACL 2024, pp. 103–115, Mexico City, Mexico. Association for Computational Linguistics.
  94. 94.Verma, V., Fleisig, E., Tomlin, N., & Klein, D. (2024). Ghostbuster: Detecting text ghostwritten by large language models. In Duh, K., Gomez, H., & Bethard, S. (Eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 1702–1717, Mexico City, Mexico. Association for Computational Linguistics.
  95. 95.Wang, P., Li, L., Ren, K., Jiang, B., Zhang, D., & Qiu, X. (2023). SeqXGPT: Sentence-level AI-generated text detection. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 1144–1156, Singapore. Association for Computational Linguistics.
  96. 96.Wang, Y., Feng, S., Hou, A., Pu, X., Shen, C., Liu, X., Tsvetkov, Y., & He, T. (2024a). Stumbling blocks: Stress testing the robustness of machine-generated text detectors under attacks. In Ku, L.-W., Martins, A., & Srikumar, V. (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2894–2925, Bangkok, Thailand. Association for Computational Linguistics.
  97. 97.Wang, Y., Mansurov, J., Ivanov, P., Su, J., Shelmanov, A., Tsvigun, A., Whitehouse, C., Mohammed Afzal, O., Mahmoud, T., Sasaki, T., Arnold, T., Aji, A., Habash, N., Gurevych, I., & Nakov, P. (2024b). M4: Multi-generator, multi-domain, and multilingual black-box machine-generated text detection. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1369–1407, St. Julian’s, Malta. Association for Computational Linguistics.
  98. 98.Wardle, C., & Derakhshan, H. (2017). Information disorder: Toward an interdisciplinary framework for research and policymaking, Vol. 27. Council of Europe.
  99. 99.Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Folt`ynek, T., Guerrero-Dib, J., Popoola, O., Sigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19 (1), 26.
  100. 100.Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., & Le, Q. V. (2021). Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  101. 101.Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W., Haas, J., Legassick, S., Irving, G., & Gabriel, I. (2022). Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’22, p. 214–229, New York, NY, USA. Association for Computing Machinery.
  102. 102.Wu, J., Guo, J., & Hooi, B. (2024). Fake news in sheep’s clothing: Robust fake news detection against LLM-empowered style attacks. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 3367–3378.
  103. 103.Wu, J., Yang, S., Zhan, R., Yuan, Y., Chao, L. S., & Wong, D. F. (2025). A survey on llm-generated text detection: Necessity, methods, and future directions. Computational Linguistics, pp. 1–65.
  104. 104.Xu, H., Ren, J., He, P., Zeng, S., Cui, Y., Liu, A., Liu, H., & Tang, J. (2023). On the generalization of training-based ChatGPT detection methods. arXiv preprint arXiv:2310.01307.
  105. 105.Yang, X., Chen, K., Zhang, W., Liu, C., Qi, Y., Zhang, J., Fang, H., & Yu, N. (2023). Watermarking text generated by black-box language models. arXiv preprint arXiv:2305.08883.
  106. 106.Yang, X., Cheng, W., Petzold, L., Wang, W. Y., & Chen, H. (2024). DNA-GPT: Divergent n-gram analysis for training-free detection of GPT-generated text. In The Twelfth International Conference on Learning Representations (ICLR).
  107. 107.Yang, X., Pan, L., Zhao, X., Chen, H., Petzold, L., Wang, W. Y., & Cheng, W. (2023). A survey on detection of LLMs-generated content. arXiv preprint arXiv:2310.15654.
  108. 108.Yoo, K., Ahn, W., Jang, J., & Kwak, N. (2023). Robust multi-bit natural language watermarking through invariant features. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2092–2115.
  109. 109.Yoo, K., Ahn, W., & Kwak, N. (2024). Advancing beyond identification: Multi-bit watermark for large language models. In Duh, K., Gomez, H., & Bethard, S. (Eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 4031–4055, Mexico City, Mexico. Association for Computational Linguistics.
  110. 110.Yu, P., Chen, J., Feng, X., & Xia, Z. (2025). CHEAT: A large-scale dataset for detecting ChatGPT-written abstracts. IEEE Transactions on Big Data.
  111. 111.Yu, X., Qi, Y., Chen, K., Chen, G., Yang, X., Zhu, P., Zhang, W., & Yu, N. (2023). GPT paternity test: GPT generated text detection with GPT genetic inheritance. arXiv preprint arXiv:2305.12519.
  112. 112.Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., & Choi, Y. (2019). Defending against neural fake news. Advances in Neural Information Processing Systems, 32.
  113. 113.Zhang, H., Edelman, B. L., Francati, D., Venturi, D., Ateniese, G., & Barak, B. (2024a). Watermarks in the sand: Impossibility of strong watermarking for generative models. In The Forty-first International Conference on Machine Learning (ICML).
  114. 114.Zhang, Q., Gao, C., Chen, D., Huang, Y., Huang, Y., Sun, Z., Zhang, S., Li, W., Fu, Z., Wan, Y., & Sun, L. (2024b). LLM-as-a-coauthor: Can mixed human-written and machine-generated text be detected?. In Duh, K., Gomez, H., & Bethard, S. (Eds.), Findings of the Association for Computational Linguistics: NAACL 2024, pp. 409–436, Mexico City, Mexico. Association for Computational Linguistics.
  115. 115.Zhang, R., Hussain, S. S., Neekhara, P., & Koushanfar, F. (2024c). {REMARK-LLM}: A robust and efficient watermarking framework for generative large language models. In 33rd USENIX Security Symposium (USENIX Security 24), pp. 1813–1830.
  116. 116.Zhao, X., Ananth, P. V., Li, L., & Wang, Y.-X. (2024). Provable robust watermarking for AI-generated text. In The Twelfth International Conference on Learning Representations (ICLR).
  117. 117.Zhong, W., Tang, D., Xu, Z., Wang, R., Duan, N., Zhou, M., Wang, J., & Yin, J. (2020). Neural deepfake detection with factual structure of text. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 2461–2470.
  118. 118.Zhou, J., Zhang, Y., Luo, Q., Parker, A. G., & De Choudhury, M. (2023). Synthetic lies: Understanding AI-generated misinformation and evaluating algorithmic and human solutions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pp. 1–20.

Citation

MLA
Fraser, K. C., et al. “Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods”. Journal of Artificial Intelligence Research, vol. 82, 2025, pp. 2233–78, https://doi.org/10.1613/jair.1.16665.
APA
Fraser, K. C., Dawkins, H., & Kiritchenko, S. (2025). Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods. Journal of Artificial Intelligence Research, 82, 2233–2278. https://doi.org/10.1613/jair.1.16665
Chicago
Fraser, K. C., H. Dawkins, and S. Kiritchenko. 2025. “Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods”. Journal of Artificial Intelligence Research 82: 2233–78. https://doi.org/10.1613/jair.1.16665.
Harvard
Fraser, K.C., Dawkins, H. and Kiritchenko, S. (2025) “Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods”, Journal of Artificial Intelligence Research, 82, pp. 2233–2278. Available at: https://doi.org/10.1613/jair.1.16665.
Vancouver
1. Fraser KC, Dawkins H, Kiritchenko S (2025) Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods. Journal of Artificial Intelligence Research 82:2233–2278

BibTeX

@article{Fraser_2025, title={Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods}, volume={82}, ISSN={1076-9757}, url={http://dx.doi.org/10.1613/jair.1.16665}, DOI={10.1613/jair.1.16665}, journal={Journal of Artificial Intelligence Research}, publisher={AI Access Foundation}, author={Fraser, Kathleen C. and Dawkins, Hillary and Kiritchenko, Svetlana}, year={2025}, month=Apr, pages={2233–2278} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/