DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
Eric MitchellYoonho LeeAlexander KhazatskyChristopher D. ManningChelsea Finn
Introduces DetectGPT, a zero-shot method that accurately identifies machine-generated text by exploiting negative curvature regions in language model log probability spaces without requiring separate classifier training or watermarks.
The rapid growth and fluency of large language models have created pressing challenges across journalism, education, and content authenticity, as automated text generation risks proliferating factual inaccuracies and complicating fair academic assessment. Because humans perform only slightly better than chance at identifying machine-generated text, dependable automated detection tools are urgently required. The article introduces and evaluates DetectGPT, a new zero-shot detection method designed to determine whether a given passage was generated by a specific language model without requiring model retraining, custom dataset collection, or explicit watermarking.
DetectGPT operates on the foundational finding that machine-generated text tends to occupy regions of negative curvature within a model's log probability function, meaning minor rewrites of model-generated text systematically lower the model's assigned probability. To measure this, the approach generates minor rephrasings of a candidate passage using an off-the-shelf mask-filling language model and evaluates the drop in log probability under the candidate source model. The article assesses DetectGPT across six diverse datasets and a wide spectrum of models ranging from 1.5 billion parameters up to 175-billion-parameter systems such as GPT-3 and Jurassic-2 Jumbo, evaluating detection accuracy via the area under the receiver operating characteristic curve.
Across the experiments, DetectGPT consistently outperformed existing zero-shot baselines that rely solely on raw probabilities, token ranks, or predictive entropy, delivering an average improvement of 0.06 in detection score. For 20-billion-parameter GPT-NeoX news generations, DetectGPT improved detection accuracy from 0.81 to 0.95 compared to the strongest zero-shot alternative. Furthermore, while supervised detectors trained on millions of text examples degraded significantly when faced with new domains such as biomedical literature or German news, DetectGPT generalized robustly without adaptation. The method also retained solid performance, scoring above 0.80, even when nearly a quarter of the text had been revised through simulated human edits.
These findings demonstrate that language models inherently expose their artificial origin through localized probability structures, enabling reliable text verification without expensive, dataset-specific detector models. From an operational perspective, zero-shot detection minimizes retraining costs and adaptation risks across shifting domains. However, DetectGPT is computationally demanding because it requires generating and scoring up to 100 perturbations per candidate passage, and it requires white-box access to evaluate model token probabilities, which may introduce API costs or operational friction when scoring proprietary models.
Organizations evaluating text detection tools should consider DetectGPT as an accurate, adaptable zero-shot baseline, particularly in settings where white-box model probabilities are accessible. Next steps supported by the article include investigating model ensembling to improve detection when the generating model is unknown, optimizing perturbation pipelines to lower computation overhead, and exploring whether this probability curvature property extends to generative models in visual and audio domains.
- Paper: Cross-Domain Detection of GPT-2-Generated Technical Text, Juan Diego Rodriguez et al. (2022). Establishes foundational supervised cross-domain detection paradigms for transformer-generated text, highlighting the generalization challenges that motivate zero-shot, curvature-based detection methods.
- Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). Introduces key statistical dynamics of autoregressive decoding and probability distribution landscapes in neural text generation that underpin DetectGPT's curvature hypothesis.
- Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). Directly evaluates the vulnerability of zero-shot probability-based detectors like DetectGPT against deliberate evasion tactics such as recursive paraphrasing.
- Paper: RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors, Liam Dugan et al. (2024). Benchmarks leading machine-generated text detectors under extensive adversarial perturbations, varying decoding strategies, and cross-domain distribution shifts.
- Paper: M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection, Yuxia Wang et al. (2024). Extends machine-generated text detection benchmarking into multilingual settings, source model attribution, and mixed-document boundary detection.
- Paper: Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods, Kathleen C. Fraser et al. (2025). Provides a comprehensive meta-analysis of the practical factors, length constraints, and sampling methods that dictate the performance of statistical AI text detectors.
- Paper: Who Wrote this Code? Watermarking for Code Generation, Taehyun Lee et al. (2024). Evaluates the limits of post-hoc detectors including DetectGPT on structured programming code and proposes entropy-targeted watermarking alternatives.
- Paper: Scalable watermarking for identifying large language model outputs, Sumanth Dathathri et al. (2024). Presents production-scale generative watermarking as an alternative detection paradigm to overcome the computational perturbation overhead of zero-shot methods.
