Built independently by an author, for readers. Read the story and support ChapterPal

keyword

metric-based detectors

Metric-based detectors are automated systems designed to determine whether a text was generated by an artificial intelligence model or composed by a human by evaluating intrinsic statistical and probabilistic characteristics of the writing. Rather than training a dedicated neural classifier on labeled examples of human and machine text, these detectors compute quantitative metrics using a reference language model, such as perplexity, token log-likelihood, entropy, rank distributions, or curvature under small text perturbations. They operate on the premise that machine-generated content typically exhibits distinct statistical signatures, such as lower uncertainty and higher token predictability, compared to human prose. Because metric-based detectors classify text by comparing these calculated values against predetermined thresholds, they can function in a zero-shot manner without extensive supervised retraining, though their effectiveness often depends on the alignment between the reference model and the generator of the target text.

1 item

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

Liam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu, Josh Magnus Ludan, Hainiu Xu, Daphne Ippolito, Chris Callison-Burch

OrganizationsCarnegie Mellon UniversityKing's College LondonUniversity College LondonUniversity of Pennsylvania

Why you should read this

Introduces a six-million-generation benchmark across eleven language models, eight domains, and eleven adversarial attacks, revealing that top AI text detectors easily fail when faced with minor sampling changes, repetition penalties, or unseen models.

Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they are, the datasets used for evaluation are insufficiently challenging—lacking variations in sampling strategy, adversarial attacks, and open-source generative models. In this work we present RAID: the largest and most challenging benchmark dataset for machine-generated text detection. RAID includes over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies. Using RAID, we evaluate the out-of-domain and adversarial robustness of 8 open- and 4 closed-source detectors and find that current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models. We release our data1 along with a leaderboard2 to encourage future research.

Added

2026-09-26