Built independently by an author, for readers. Read the story and support ChapterPal

keyword

unseen generative models

Unseen generative models are artificial intelligence systems whose generated content was not included in the training data, tuning pipeline, or prior exposure of a downstream evaluation framework or detection tool. In machine learning and synthetic media analysis, these models function as novel, out-of-distribution generation sources used to test whether a detector or classifier can generalize effectively beyond the specific architectures, parameter scales, and decoding methods on which it was trained. Because detection systems often memorize subtle statistical artifacts unique to their training data, assessing performance on outputs from unseen generative models serves as a standard benchmark for evaluating real-world robustness against newly developed, unfamiliar, or alternative generative technologies.

1 item

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

Liam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu, Josh Magnus Ludan, Hainiu Xu, Daphne Ippolito, Chris Callison-Burch

OrganizationsCarnegie Mellon UniversityKing's College LondonUniversity College LondonUniversity of Pennsylvania

Why you should read this

Introduces a six-million-generation benchmark across eleven language models, eight domains, and eleven adversarial attacks, revealing that top AI text detectors easily fail when faced with minor sampling changes, repetition penalties, or unseen models.

Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they are, the datasets used for evaluation are insufficiently challenging—lacking variations in sampling strategy, adversarial attacks, and open-source generative models. In this work we present RAID: the largest and most challenging benchmark dataset for machine-generated text detection. RAID includes over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies. Using RAID, we evaluate the out-of-domain and adversarial robustness of 8 open- and 4 closed-source detectors and find that current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models. We release our data1 along with a leaderboard2 to encourage future research.

Added

2026-09-26