Built independently by an author, for readers. Read the story and support ChapterPal

keyword

closed-source commercial detectors

Closed-source commercial detectors are proprietary software tools and services developed by businesses to identify machine-generated content, such as artificial intelligence text, without publicly disclosing their underlying code, training datasets, model architecture, or internal algorithms. Typically accessed through web interfaces or paid application programming interfaces, these tools analyze submitted samples and generate scores or classifications indicating the likelihood that the content was produced by an automated language model. Because their core mechanisms are protected as commercial trade secrets, users and independent researchers must evaluate these systems as black-box services, relying on input-output behavior rather than direct structural inspection to assess their performance, reliability, and vulnerabilities.

1 item

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

Liam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu, Josh Magnus Ludan, Hainiu Xu, Daphne Ippolito, Chris Callison-Burch

OrganizationsCarnegie Mellon UniversityKing's College LondonUniversity College LondonUniversity of Pennsylvania

Why you should read this

Introduces a six-million-generation benchmark across eleven language models, eight domains, and eleven adversarial attacks, revealing that top AI text detectors easily fail when faced with minor sampling changes, repetition penalties, or unseen models.

Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they are, the datasets used for evaluation are insufficiently challenging—lacking variations in sampling strategy, adversarial attacks, and open-source generative models. In this work we present RAID: the largest and most challenging benchmark dataset for machine-generated text detection. RAID includes over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies. Using RAID, we evaluate the out-of-domain and adversarial robustness of 8 open- and 4 closed-source detectors and find that current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models. We release our data1 along with a leaderboard2 to encourage future research.

Added

2026-09-26