Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Fast-DetectGPT

Fast-DetectGPT is a zero-shot algorithm designed to identify machine-generated text by analyzing statistical patterns in language model probabilities without requiring task-specific training data or fine-tuned classifiers. It operates on the principle of conditional probability curvature, exploiting the tendency of large language models to select tokens with consistently higher model-assigned probabilities compared to human authors. While earlier curvature-based methods like DetectGPT rely on computationally expensive passage perturbations from external models, Fast-DetectGPT evaluates probability curvature through direct conditional token sampling. This structural optimization significantly accelerates detection speed, lowers computational overhead, and maintains high classification accuracy across diverse models and writing domains.

1 item

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

Liam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu, Josh Magnus Ludan, Hainiu Xu, Daphne Ippolito, Chris Callison-Burch

OrganizationsCarnegie Mellon UniversityKing's College LondonUniversity College LondonUniversity of Pennsylvania

Why you should read this

Introduces a six-million-generation benchmark across eleven language models, eight domains, and eleven adversarial attacks, revealing that top AI text detectors easily fail when faced with minor sampling changes, repetition penalties, or unseen models.

Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they are, the datasets used for evaluation are insufficiently challenging—lacking variations in sampling strategy, adversarial attacks, and open-source generative models. In this work we present RAID: the largest and most challenging benchmark dataset for machine-generated text detection. RAID includes over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies. Using RAID, we evaluate the out-of-domain and adversarial robustness of 8 open- and 4 closed-source detectors and find that current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models. We release our data1 along with a leaderboard2 to encourage future research.

Added

2026-09-26