DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature

Eric MitchellYoonho LeeAlexander KhazatskyChristopher D. ManningChelsea Finn

article2023ICML962 citations

Introduces DetectGPT, a zero-shot method that accurately identifies machine-generated text by exploiting negative curvature regions in language model log probability spaces without requiring separate classifier training or watermarks.

Listen

The rapid growth and fluency of large language models have created pressing challenges across journalism, education, and content authenticity, as automated text generation risks proliferating factual inaccuracies and complicating fair academic assessment. Because humans perform only slightly better than chance at identifying machine-generated text, dependable automated detection tools are urgently required. The article introduces and evaluates DetectGPT, a new zero-shot detection method designed to determine whether a given passage was generated by a specific language model without requiring model retraining, custom dataset collection, or explicit watermarking.

DetectGPT operates on the foundational finding that machine-generated text tends to occupy regions of negative curvature within a model's log probability function, meaning minor rewrites of model-generated text systematically lower the model's assigned probability. To measure this, the approach generates minor rephrasings of a candidate passage using an off-the-shelf mask-filling language model and evaluates the drop in log probability under the candidate source model. The article assesses DetectGPT across six diverse datasets and a wide spectrum of models ranging from 1.5 billion parameters up to 175-billion-parameter systems such as GPT-3 and Jurassic-2 Jumbo, evaluating detection accuracy via the area under the receiver operating characteristic curve.

Across the experiments, DetectGPT consistently outperformed existing zero-shot baselines that rely solely on raw probabilities, token ranks, or predictive entropy, delivering an average improvement of 0.06 in detection score. For 20-billion-parameter GPT-NeoX news generations, DetectGPT improved detection accuracy from 0.81 to 0.95 compared to the strongest zero-shot alternative. Furthermore, while supervised detectors trained on millions of text examples degraded significantly when faced with new domains such as biomedical literature or German news, DetectGPT generalized robustly without adaptation. The method also retained solid performance, scoring above 0.80, even when nearly a quarter of the text had been revised through simulated human edits.

These findings demonstrate that language models inherently expose their artificial origin through localized probability structures, enabling reliable text verification without expensive, dataset-specific detector models. From an operational perspective, zero-shot detection minimizes retraining costs and adaptation risks across shifting domains. However, DetectGPT is computationally demanding because it requires generating and scoring up to 100 perturbations per candidate passage, and it requires white-box access to evaluate model token probabilities, which may introduce API costs or operational friction when scoring proprietary models.

Organizations evaluating text detection tools should consider DetectGPT as an accurate, adaptable zero-shot baseline, particularly in settings where white-box model probabilities are accessible. Next steps supported by the article include investigating model ensembling to improve detection when the generating model is unknown, optimizing perturbation pipelines to lower computation overhead, and exploring whether this probability curvature property extends to generative models in visual and audio domains.

arXiv: 2301.11305
  • Paper: Cross-Domain Detection of GPT-2-Generated Technical Text, Juan Diego Rodriguez et al. (2022). Establishes foundational supervised cross-domain detection paradigms for transformer-generated text, highlighting the generalization challenges that motivate zero-shot, curvature-based detection methods.
  • Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). Introduces key statistical dynamics of autoregressive decoding and probability distribution landscapes in neural text generation that underpin DetectGPT's curvature hypothesis.
Cover for DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature

Abstract

The increasing fluency and widespread usage of large language models (LLMs) highlight the desirability of corresponding tools aiding detection of LLM-generated text. In this paper, we identify a property of the structure of an LLM's probability function that is useful for such detection. Specifically, we demonstrate that text sampled from an LLM tends to occupy negative curvature regions of the model's log probability function. Leveraging this observation, we then define a new curvature-based criterion for judging if a passage is generated from a given LLM. This approach, which we call DetectGPT, does not require training a separate classifier, collecting a dataset of real or generated passages, or explicitly watermarking generated text. It uses only log probabilities computed by the model of interest and random perturbations of the passage from another generic pre-trained language model (e.g., T5). We find DetectGPT is more discriminative than existing zero-shot methods for model sample detection, notably improving detection of fake news articles generated by 20B parameter GPT-NeoX from 0.81 AUROC for the strongest zero-shot baseline to 0.95 AUROC for DetectGPT. See this https URL for code, data, and other project information.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The Zero-Shot Machine-Generated Text Detection Problem
  • 4 DetectGPT: Zero-shot Machine-Generated Text Detection with Random Perturbations
  • 5 Experiments
  • 5.1 Main Results
  • 5.2 Variants of Machine-Generated Text Detection
  • 5.3 Other factors impacting performance of DetectGPT
  • 6 Discussion
  • References
  • A Complete Results for Top-pp and Top-kk Decoding

Knowls

  1. Knowl 1 — Perturbation Discrepancy and the Local Perturbation Discrepancy Gap Hypothesis

    definition

    Let xx denote a candidate text passage (such as a paragraph-length sequence of tokens), pθp_\theta denote the probability distribution defined by a generative autoregressive language model with parameters θ\theta, and q(⋅∣x)q(\cdot \mid x) denote a perturbation function that generates semantically similar variations x~\tilde{x} of xx along the natural language data manifold.

    The perturbation discrepancy d(x,pθ,q)d(x, p_\theta, q) is defined as the difference between the model log probability of the original passage xx and the expected log probability of its perturbations under pθp_\theta:

    d(x,pθ,q)≜log⁡pθ(x)−Ex~∼q(⋅∣x)[log⁡pθ(x~)]d(x, p_\theta, q) \triangleq \log p_\theta(x) - \mathbb{E}_{\tilde{x} \sim q(\cdot \mid x)} \left[ \log p_\theta(\tilde{x}) \right]

    The Local Perturbation Discrepancy Gap Hypothesis asserts that:

    1. If qq produces samples on the data manifold, d(x,pθ,q)d(x, p_\theta, q) is positive and large with high probability for passages sampled from the model, x∼pθx \sim p_\theta.
    2. For human-written passages, x∼phumanx \sim p_{\text{human}}, d(x,pθ,q)d(x, p_\theta, q) tends toward zero for all xx.

    This discrepancy arises because machine-generated passages tend to occupy negative curvature regions (such as local maxima) of the model's log-probability surface, whereas human-written passages do not systematically occupy local probability maxima of pθp_\theta.

  2. Knowl 2 — Approximation of the Negative Hessian Trace via Perturbation Discrepancy

    theoretical result

    In a continuous latent semantic space where local displacements correspond to meaning-preserving edits, let f(x)=log⁡pθ(x)f(x) = \log p_\theta(x) denote the log probability function of a language model pθp_\theta at a passage representation xx. Let Hf(x)H_f(x) denote the Hessian matrix of second-order partial derivatives of ff with respect to xx.

    By Hutchinson's trace estimator, for any random perturbation vector z∼qzz \sim q_z whose components satisfy E[zi]=0\mathbb{E}[z_i] = 0 and Var(zi)=1\text{Var}(z_i) = 1, the trace of Hf(x)H_f(x) satisfies:

    tr(Hf(x))=Ez[z⊤Hf(x)z]\text{tr}(H_f(x)) = \mathbb{E}_z \left[ z^\top H_f(x) z \right]

    Approximating the directional second derivative z⊤Hf(x)zz^\top H_f(x) z via symmetric finite differences with step size h=1h = 1 yields:

    z⊤Hf(x)z≈f(x+z)+f(x−z)−2f(x)z^\top H_f(x) z \approx f(x + z) + f(x - z) - 2f(x)

    Taking the expectation over zz and assuming that the perturbation distribution is symmetric (p(z)=p(−z)p(z) = p(-z) for all zz), the negative trace of the Hessian relates directly to the difference between the central point value and the perturbed expectation:

    −12tr(Hf(x))≈f(x)−Ez[f(x+z)]-\frac{1}{2} \text{tr}(H_f(x)) \approx f(x) - \mathbb{E}_z [f(x + z)]

    When discrete text perturbations x~∼q(⋅∣x)\tilde{x} \sim q(\cdot \mid x) generated by a masked language model approximate random directions zz on the semantic data manifold, the perturbation discrepancy d(x,pθ,q)=log⁡pθ(x)−Ex~∼q(⋅∣x)[log⁡pθ(x~)]d(x, p_\theta, q) = \log p_\theta(x) - \mathbb{E}_{\tilde{x} \sim q(\cdot \mid x)}[\log p_\theta(\tilde{x})] is approximately proportional to the negative trace of the log-probability function's Hessian (−tr(Hlog⁡pθ(x))-\text{tr}(H_{\log p_\theta}(x))) restricted to the data manifold.

  3. Knowl 3 — DetectGPT Zero-Shot Detection Algorithm

    algorithm

    DetectGPT is a zero-shot algorithm that determines whether a candidate passage xx was generated by a specified source language model pθp_\theta. The algorithm requires white-box scoring access to compute log⁡pθ(⋅)\log p_\theta(\cdot), an off-the-shelf mask-filling language model qq (such as T5-3B) to generate perturbations, a sample count hyperparameter kk, and a decision threshold ϵ\epsilon.

    To generate perturbations, DetectGPT masks randomly selected 2-word spans until 15% of the passage words are masked, and samples replacements from qq at temperature 1.0. To enhance detection stability across samples with varying variance, the raw perturbation discrepancy is normalized by the sample standard deviation of the perturbed log probabilities.

    Input: candidate passage xx, source model pθp_\theta, perturbation distribution qq, perturbation count kk, threshold ϵ\epsilon
    Output: boolean decision (true if generated by pθp_\theta, false otherwise)
    for i←1i \leftarrow 1 to kk do
        Sample perturbed passage x~i∼q(⋅∣x)\tilde{x}_i \sim q(\cdot \mid x)
        Compute log probability log⁡pθ(x~i)\log p_\theta(\tilde{x}_i)
    μ~←1k∑i=1klog⁡pθ(x~i)\tilde{\mu} \leftarrow \frac{1}{k} \sum_{i=1}^k \log p_\theta(\tilde{x}_i)
    \hat{d}_x \leftarrow \log p_\theta(x) - \tilde{\mu}
    \tilde{\sigma}_x^2 \leftarrow \frac{1}{k-1} \sum_{i=1}^k (\log p_\theta(\tilde{x}_i) - \tilde{\mu})^2
    \tilde{\sigma}_x \leftarrow \sqrt{\tilde{\sigma}_x^2}
    s \leftarrow \frac{\hat{d}_x}{\tilde{\sigma}_x}
    if s>ϵs > \epsilon then
        return true
    else
        return false

    In standard evaluations, setting k=100k = 100 perturbations provides near-maximal discriminative performance.

  4. Knowl 4 — Zero-Shot Machine-Generated Text Detection Performance Across Models and Domains

    data/table

    The table below shows the area under the receiver operating characteristic curve (AUROC) for DetectGPT compared to four zero-shot statistical baseline criteria (token log probability thresholding, average token rank, average token log-rank, and predictive entropy). Evaluations use 500 samples per dataset-model combination, sampled with temperature 1.0 conditional generation from the first 30 tokens of human text. Datasets evaluated are XSum (news articles), SQuAD (Wikipedia contexts), and WritingPrompts (Reddit creative stories) across source models ranging from 1.5B to 20B parameters: GPT-2 XL (1.5B), OPT-2.7B, GPT-Neo-2.7B, GPT-J (6B), and GPT-NeoX (20B).

    Dataset / Method GPT-2 OPT-2.7 Neo-2.7 GPT-J NeoX Average
    XSum
    log⁡p(x)\log p(x) 0.86 0.86 0.86 0.82 0.77 0.83
    Rank 0.79 0.76 0.77 0.75 0.73 0.76
    LogRank 0.89 0.88 0.90 0.86 0.81 0.87
    Entropy 0.60 0.50 0.58 0.58 0.61 0.57
    DetectGPT 0.99 0.97 0.99 0.97 0.95 0.97
    SQuAD
    log⁡p(x)\log p(x) 0.91 0.88 0.84 0.78 0.71 0.82
    Rank 0.83 0.82 0.80 0.79 0.74 0.80
    LogRank 0.94 0.92 0.90 0.83 0.76 0.87
    Entropy 0.58 0.53 0.58 0.58 0.59 0.57
    DetectGPT 0.99 0.97 0.97 0.90 0.79 0.92
    WritingPrompts
    log⁡p(x)\log p(x) 0.97 0.95 0.95 0.94 0.93 0.95
    Rank 0.87 0.83 0.82 0.83 0.81 0.83
    LogRank 0.98 0.96 0.97 0.96 0.95 0.96
    Entropy 0.37 0.42 0.34 0.36 0.39 0.38
    DetectGPT 0.99 0.99 0.99 0.97 0.93 0.97

    DetectGPT outperforms all zero-shot baselines on 14 of 15 combinations, delivering an average AUROC gain of 0.10 on XSum, 0.05 on SQuAD, and 0.01 on WritingPrompts over the strongest baseline (LogRank). On the largest open-source model evaluated, GPT-NeoX (20B), DetectGPT improves fake news detection AUROC on XSum from 0.81 (LogRank) to 0.95.

  5. Knowl 5 — Generalization of Zero-Shot DetectGPT Versus Supervised Detectors Under Distribution Shift

    empirical result

    Supervised classifiers fine-tuned for machine-generated text detection (such as OpenAI's RoBERTa-base and RoBERTa-large GPT-2 detectors) achieve high accuracy on in-distribution data but suffer severe performance drops when tested out-of-distribution, whereas zero-shot DetectGPT maintains robust detection.

    In empirical evaluations across 200 samples per dataset:

    • On in-distribution English news (XSum with GPT-2): Supervised RoBERTa-large achieves 0.997 AUROC, RoBERTa-base achieves 0.981 AUROC, and DetectGPT achieves 0.991 AUROC.
    • On out-of-domain English scientific writing (PubMedQA with PubMedGPT): RoBERTa-base drops to 0.604 AUROC and RoBERTa-large drops to 0.713 AUROC, while DetectGPT achieves 0.836 AUROC.
    • On out-of-distribution German news (WMT16 German with mGPT and mT5-3B perturbations): Supervised detectors fail completely, with RoBERTa-base reaching 0.394 AUROC (worse than random guessing) and RoBERTa-large reaching 0.537 AUROC. In contrast, DetectGPT achieves 0.962 AUROC on German news, matching its English performance.
  6. Knowl 6 — Detection Performance on 175-Billion-Parameter Models

    data/table

    Zero-shot and supervised detectors were evaluated on 150 samples per dataset generated by two 175-billion-parameter language models: OpenAI's GPT-3 and AI21 Labs' Jurassic-2 Jumbo. Because commercial APIs for these models do not return full vocabulary logit distributions, rank- and entropy-based zero-shot metrics cannot be computed; only log-probability thresholding, supervised RoBERTa detectors, and DetectGPT are compared.

    Detection Method PubMedQA XSum WritingPrompts Average
    RoBERTa-base 0.64 / 0.58 0.92 / 0.74 0.92 / 0.81 0.77
    RoBERTa-large 0.71 / 0.64 0.92 / 0.88 0.91 / 0.88 0.82
    log⁡p(x)\log p(x) thresholding 0.64 / 0.55 0.76 / 0.61 0.88 / 0.67 0.69
    DetectGPT 0.84 / 0.77 0.84 / 0.84 0.87 / 0.84 0.83

    Note: Table entries show AUROC values formatted as (GPT-3 AUROC / Jurassic-2 AUROC).

    DetectGPT achieves an overall average AUROC of 0.83 across both 175B models, outperforming raw log-probability thresholding (0.69 average AUROC) and matching or exceeding supervised detectors fine-tuned on millions of samples (RoBERTa-large at 0.82 average AUROC).

  7. Knowl 7 — Robustness of DetectGPT to Text Paraphrasing and Span Revisions

    empirical result

    To evaluate resilience against human post-editing or automated paraphrasing of machine-generated text, human revisions were simulated on GPT-J-generated XSum news articles by replacing randomly selected 5-word spans with samples from T5-3B until a fraction r∈[0,0.24]r \in [0, 0.24] of the text was replaced.

    As the replacement fraction rr increases from 0% to 24%:

    • All statistical detection methods (DetectGPT, log likelihood, token rank, token log-rank) experience monotonic degradation in AUROC as more text is modified.
    • DetectGPT consistently maintains the highest AUROC across all revision percentages, retaining an AUROC above 0.80 even when nearly a quarter (r=24%r = 24\%) of the machine-generated text has been rewritten.
    • By comparison, the baseline rank, log-rank, and likelihood methods degrade to lower AUROC values (ranging roughly from 0.65 to 0.75 at r=24%r = 24\%).
  8. Knowl 8 — Effect of Nucleus and Top-k Decoding on Machine Text Detectability

    empirical result

    Evaluating detection methods on machine text generated with truncated decoding strategies—nucleus sampling (top-pp with p=0.96p = 0.96) and top-kk sampling (k=40k = 40)—demonstrates that restricting decoding makes model outputs easier to detect on average across datasets.

    Because top-kk and top-pp sampling truncate lower-probability tokens in the tail distribution, the generated sequences have higher average per-token log likelihood and lower average token rank under the generative model compared to unconstrained temperature 1.0 direct sampling.

    Across five models (GPT-2, OPT-2.7B, GPT-Neo-2.7B, GPT-J, GPT-NeoX) on XSum, SQuAD, and WritingPrompts:

    • Under nucleus sampling (p=0.96p = 0.96), DetectGPT achieves average AUROCs of 0.98 on XSum, 0.94 on SQuAD, and 0.98 on WritingPrompts.
    • Under top-kk sampling (k=40k = 40), DetectGPT achieves average AUROCs of 0.98 on XSum, 0.93 on SQuAD, and 0.97 on WritingPrompts.
    • DetectGPT provides the highest or tied-highest average AUROC across all three datasets under both decoding constraints.
  9. Knowl 9 — Cross-Model Scoring Behavior in Black-Box Text Detection

    empirical result

    When the exact source model that generated a text candidate is unknown, DetectGPT can be evaluated in a black-box surrogate setting where a separate model BB is used to compute log probabilities for text generated by source model AA. Evaluated across 200 samples from XSum, SQuAD, and WritingPrompts over all pairwise combinations of GPT-J (6B), GPT-Neo (2.7B), and GPT-2 XL (1.5B):

    1. White-box vs. Black-box Gap: Detection accuracy is highest along the diagonal where the scoring model matches the generator (mean AUROCs: GPT-J = 0.92, GPT-Neo = 0.97, GPT-2 = 0.99). Off-diagonal black-box detection degrades to mean AUROCs between 0.60 and 0.85.
    2. Scorer Asymmetry: Models vary in their intrinsic capacity as surrogate scoring models. When averaging across all source models (column averages), GPT-Neo (mean AUROC 0.88) and GPT-2 (mean AUROC 0.87) perform substantially better as general scoring surrogates than GPT-J (mean AUROC 0.72).
  10. Knowl 10 — Scaling Effects of Mask-Filling Model Size and Perturbation Sample Budget

    empirical result

    The detection efficacy of DetectGPT depends critically on the capacity of the perturbation model and the number of perturbation samples kk used to estimate the expectation Ex~∼q(⋅∣x)[log⁡pθ(x~)]\mathbb{E}_{\tilde{x} \sim q(\cdot \mid x)}[\log p_\theta(\tilde{x})]:

    1. Perturbation Model Capacity: When evaluated on SQuAD contexts with T5 mask-filling models ranging from 60M (T5-small) to 2.7B/3B (T5-3B) parameters, detection AUROC increases monotonically with mask-filling model size for all source model scales (GPT-2 small, medium, large, XL). Perturbing tokens via uniform random replacement from the vocabulary yields near-random detection AUROC (~0.55), confirming that perturbations must lie on the natural language manifold to capture semantic curvature.
    2. Perturbation Budget (kk): As the number of sampled perturbations kk increases from 1 to 1000 on GPT-2 and GPT-J, AUROC rises steeply from k=1k=1 to k=100k=100, beyond which detection performance converges. Thus, k=100k = 100 provides an optimal trade-off between computational cost and classification accuracy.
  11. Knowl 11 — Practical Limitations of DetectGPT

    limitation

    DetectGPT has four primary limitations identified by the authors:

    1. White-Box Model Access Requirement: DetectGPT requires token-level log probabilities from the candidate generative model pθp_\theta. For models accessible only via restricted APIs that do not return token probabilities (such as ChatGPT), direct scoring is impossible without relying on surrogate proxy models.
    2. Computational Overhead: Because DetectGPT must generate kk perturbations (typically k=100k=100) and evaluate the log probability of all kk perturbed passages under pθp_\theta for every single candidate passage, it is roughly two orders of magnitude more compute-intensive than single-pass zero-shot baselines.
    3. Dependence on Perturbation Quality: The curvature estimate assumes perturbations remain on the semantically valid data manifold. For specialized domains (such as biomedical texts in PubMedQA) or low-resource languages where general-purpose mask-filling models (like T5) perform poorly, perturbation fidelity drops, leading to lower detection accuracy.
    4. Context Length Tracking Degradation: When candidate passages are long, masking 15% of tokens in a single pass introduces 20+ mask tokens simultaneously, causing T5 mask-filling models to occasionally fail at tracking and filling masks consistently.

Coverage note — None was omitted; all key theoretical definitions, derivations, algorithmic specifications, empirical benchmarks (across sizes 1.5B to 175B, multiple domains, and languages), decoding analyses, paraphrasing experiments, surrogate model evaluations, hyperparameter sweeps, and stated limitations were fully captured.

References

  1. 1.Aaronson, S. My Projects at OpenAI, Nov 2022. URL https://scottaaronson.blog/?p=6823.
  2. 2.Bakhtin, A., Gross, S., Ott, M., Deng, Y., Ranzato, M., and Szlam, A. Real or fake? Learning to discriminate machine from human generated text. arXiv, 2019. URL http://arxiv.org/abs/1906.03351.
  3. 3.Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021. URL https://doi.org/10.5281/zenodo.5297715.
  4. 4.Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., and Weinbach, S. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of the ACL Workshop on Challenges & Perspectives in Creating Large Language Models, 2022. URL https://arxiv.org/abs/2204.06745.
  5. 5.Bojar, O. r., Chatterjee, R., Federmann, C., Graham, Y., Haddow, B., Huck, M., Jimeno Yepes, A., Koehn, P., Logacheva, V., Monz, C., Negri, M., Neveol, A., Neves, M., Popel, M., Post, M., Rubino, R., Scarton, C., Specia, L., Turchi, M., Verspoor, K., and Zampieri, M. Findings of the 2016 conference on machine translation. In Proceedings of the First Conference on Machine Translation, pp. 131–198, Berlin, Germany, August 2016. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/W/W16/W16-2301.
  6. 6.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  7. 7.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. PaLM: Scaling language modeling with pathways, 2022. URL https://arxiv.org/abs/2204.02311.
  8. 8.Christian, J. CNET secretly used AI on articles that didn’t disclose that fact, staff say. https://web.archive.org/web/20230124063916/https://futurism.com/cnet-ai-articles-label, 2023. Accessed: 2023-01-25.
  9. 9.Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf.
  10. 10.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  11. 11.Dolhansky, B., Bitton, J., Pflaum, B., Lu, J., Howes, R., Wang, M., and Ferrer, C. C. The deepfake detection challenge dataset, 2020. URL https://ai.facebook.com/datasets/dfdc/.
  12. 12.Fagni, T., Falchi, F., Gambini, M., Martella, A., and Tesconi, M. Tweepfake: About detecting deepfake tweets. PLOS ONE, 16(5):1–16, 05 2021. doi: 10.1371/journal.pone.0251415. URL https://doi.org/10.1371/journal.pone.0251415.
  13. 13.Fan, A., Lewis, M., and Dauphin, Y. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 889–898, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1082. URL https://aclanthology.org/P18-1082.
  14. 14.Gehrmann, S., Strobelt, H., and Rush, A. GLTR: Statistical detection and visualization of generated text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pp. 111–116, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-3019. URL https://aclanthology.org/P19-3019.
  15. 15.Guarnera, L., Giudice, O., and Battiato, S. Deepfake detection by analyzing convolutional traces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2020.
  16. 16.Güera, D. and Delp, E. J. Deepfake video detection using recurrent neural networks. In 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), pp. 1–6, 2018. doi: 10.1109/AVSS.2018.8639163.
  17. 17.Hutchinson, M. A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines. Communications in Statistics - Simulation and Computation, 19(2):433–450, 1990. doi: 10.1080/03610919008812866. URL https://doi.org/10.1080/03610919008812866.
  18. 18.Ippolito, D., Duckworth, D., Callison-Burch, C., and Eck, D. Automatic detection of generated text is easiest when humans are fooled. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 1808–1822, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.164. URL https://www.aclweb.org/anthology/2020.acl-main.164.
  19. 19.Jawahar, G., Abdul-Mageed, M., and Lakshmanan, L. V. S. Automatic detection of machine generated text: A critical survey. In International Conference on Computational Linguistics, 2020.
  20. 20.Jin, Q., Dhingra, B., Liu, Z., Cohen, W., and Lu, X. PubMedQA: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 2567–2577, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1259. URL https://aclanthology.org/D19-1259.
  21. 21.Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T. A watermark for large language models, 2023. URL https://arxiv.org/abs/2301.10226.
  22. 22.Krishna, K., Song, Y., Karpinska, M., Wieting, J., and Iyyer, M. Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense. arXiv preprint arXiv:2303.13408, 2023.
  23. 23.Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., and Zou, J. Gpt detectors are biased against non-native english writers. arXiv preprint arXiv:2304.02819, 2023.
  24. 24.Lin, S., Hilton, J., and Evans, O. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3214–3252, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URL https://aclanthology.org/2022.acl-long.229.
  25. 25.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  26. 26.Mireshghallah, F., Mattern, J., Gao, S., Shokri, R., and Berg-Kirkpatrick, T. Smaller language models are better black-box machine-generated text detectors. arXiv preprint arXiv:2305.09859, 2023.
  27. 27.Narayan, S., Cohen, S. B., and Lapata, M. Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, 2018.
  28. 28.OpenAI. Chatgpt: Optimizing language models for dialogue. http://web.archive.org/web/20230109000707/https://openai.com/blog/chatgpt/, 2022. Accessed: 2023-01-10.
  29. 29.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners, 2019. URL https://d4mucfpksywv.cloudfront.net/better-language-models/language_models_are_unsupervised_multitask_learners.pdf.
  30. 30.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  31. 31.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383–2392, Austin, Texas, November 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1264. URL https://aclanthology.org/D16-1264.
  32. 32.Roose, K. and Newton, C. ChatGPT transforms a classroom and is ‘M3GAN’ real? Hard Fork, a New York Times Podcast, 2022. URL https://www.nytimes.com/2023/01/13/podcasts/hard-fork-chatgpt-teachers-gen-z-cameras-m3gan.html.
  33. 33.Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., and Feizi, S. Can ai-generated text be reliably detected? arXiv preprint arXiv:2303.11156, 2023.
  34. 34.Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., and Wang, J. Release strategies and the social impacts of language models, 2019. URL https://arxiv.org/ftp/arxiv/papers/1908/1908.09203.pdf.
  35. 35.Uchendu, A., Le, T., Shu, K., and Lee, D. Authorship attribution for neural text generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 8384–8395, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.673. URL https://aclanthology.org/2020.emnlp-main.673.
  36. 36.Wang, B. and Komatsuzaki, A. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021.
  37. 37.Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., and Choi, Y. Defending against neural fake news. In Neural Information Processing Systems, 2019.
  38. 38.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L. Opt: Open pre-trained transformer language models, 2022. URL https://arxiv.org/abs/2205.01068.
  39. 39.Zhao, H., Zhou, W., Chen, D., Wei, T., Zhang, W., and Yu, N. Multi-attentional deepfake detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2185–2194, 2021.
  40. 40.Zi, B., Chang, M., Chen, J., Ma, X., and Jiang, Y.-G. Wilddeepfake: A challenging real-world dataset for deepfake detection. In Proceedings of the 28th ACM international conference on multimedia, pp. 2382–2390, 2020.
  41. 41.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences, 2020.

Citation

MLA
Mitchell, E., et al. “DetectGPT: Zero-Shot Machine-Generated Text Detection Using Probability Curvature”. arXiv, 2023, http://arxiv.org/abs/2301.11305v2.
APA
Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., & Finn, C. (2023). DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature. arXiv. http://arxiv.org/abs/2301.11305v2
Chicago
Mitchell, E., Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn. 2023. “DetectGPT: Zero-Shot Machine-Generated Text Detection Using Probability Curvature”. arXiv. http://arxiv.org/abs/2301.11305v2.
Harvard
Mitchell, E. et al. (2023) “DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.11305v2.
Vancouver
1. Mitchell E, Lee Y, Khazatsky A, Manning CD, Finn C (2023) DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature. arXiv

BibTeX

@article{mitchell2023detectgpt,
  title = {DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature},
  author = {Mitchell, Eric and Lee, Yoonho and Khazatsky, Alexander and Manning, Christopher D. and Finn, Chelsea},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.11305v2},
  eprint = {2301.11305}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/