RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
Liam DuganAlyssa HwangFilip TrhlíkAndrew ZhuJosh Magnus LudanHainiu XuDaphne IppolitoChris Callison-Burch
Introduces a six-million-generation benchmark across eleven language models, eight domains, and eleven adversarial attacks, revealing that top AI text detectors easily fail when faced with minor sampling changes, repetition penalties, or unseen models.
As large language models become increasingly capable of generating human-like text, organizations face heightened risks involving phishing attacks, disinformation campaigns, spam, and spurious scientific publications. Although numerous commercial and open-source tools claim high accuracy—frequently exceeding 99%—in detecting machine-generated content, these claims are rarely evaluated against shared, rigorous standards. Existing evaluation datasets typically lack the adversarial modifications, varied decoding strategies, and modern generative models needed to test real-world effectiveness, creating an unverified sense of security among decision-makers.
The main objective of the article is to systematically assess the out-of-domain and adversarial robustness of current machine-generated text detectors using a comprehensive, standardized benchmark. To achieve this, the authors created the Robust AI Detection (RAID) dataset, which contains over 6 million generations constructed from roughly 15,000 human-written documents. The evaluation spans 11 generative models, 8 diverse topical domains, 4 decoding strategies, and 11 query-free adversarial attacks, benchmarking 12 leading open-source, metric-based, and commercial detectors while controlling for false positive rates.
The article establishes several key findings regarding detector reliability. First, open-source detectors using default classification thresholds suffer from dangerously high false positive rates, frequently misclassifying 20% to over 90% of human-written text as artificial. Second, even when calibrated to a strict 5% false positive rate, detector performance degrades dramatically under realistic generation settings: introducing standard repetition penalties and random sampling reduces accuracy by up to 32 percentage points compared to default greedy generation. Third, detectors exhibit severe domain and model bias, performing well on text produced by models they encountered during training (often reaching 95% accuracy) but plummeting to below 60% accuracy on unseen models or new domains. Finally, simple adversarial perturbations substantially undermine detection; subtle synonym swaps and homoglyph substitutions degrade accuracy across several detectors by 35 to over 70 percentage points.
These findings indicate that current text detectors are not sufficiently reliable for automated compliance, content moderation, or punitive actions such as academic disciplinary proceedings. Relying on uncalibrated or off-the-shelf detection tools exposes organizations to severe legal, reputational, and operational risks due to high false accusation rates. Furthermore, the vulnerability of detectors to basic adversarial edits means malicious actors can easily bypass automated filters, rendering legal or platform mandates for AI labeling largely unenforceable with current technology.
Decision-makers should refrain from deploying automatic text detectors in high-stakes or punitive contexts. Organizations that must use detection systems should explicitly calibrate classification thresholds on domain-specific human data rather than relying on default settings, and they should prioritize identifying direct harms—such as fraud, hate speech, and factual inaccuracy—over general machine authorship. Benchmark developers should continue maintaining dynamic evaluation datasets across multiple languages and evolving models, while practitioners should rely on continuous adversarial testing rather than static accuracy claims.
The analysis is bounded by the ongoing evolution of language models, which will require periodic updates to the dataset to reflect newer model releases, as well as limited multilingual and domain coverage outside the core English text domains. Nevertheless, the scale and methodological rigor of the benchmark support high confidence in the finding that contemporary detection tools are fragile and easily circumvented.
- Paper: Cross-Domain Detection of GPT-2-Generated Technical Text, Juan Diego Rodriguez et al. (2022). Establishes foundational findings on how machine-generated text detectors generalize or degrade across out-of-domain technical datasets, providing essential context for RAID's large-scale benchmark.
- Paper: CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks, Xuanli He et al. (2022). Introduces key concepts of linguistic watermarking and perturbation attacks on generated text, which underpin the adversarial evaluation settings tested in RAID.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). Examines systemic flaws and lack of robustness in generated-text evaluation methodologies, motivating the need for comprehensive benchmarks like RAID.
- Paper: M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection, Yuxia Wang et al. (2024). Extends benchmark evaluation of black-box text detectors into multilingual domains, fine-grained generator attribution, and human-machine transition boundary detection.
- Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). Investigates the theoretical and practical limits of machine-generated text detection by evaluating recursive paraphrasing and spoofing attacks across diverse detector classes.
- Paper: Scalable watermarking for identifying large language model outputs, Sumanth Dathathri et al. (2024). Presents a scalable generative watermarking solution designed to overcome post-hoc text detection vulnerabilities highlighted in comprehensive benchmarks like RAID.
- Paper: Who Wrote this Code? Watermarking for Code Generation, Taehyun Lee et al. (2024). Applies the challenges of synthetic text detection and evasion to the specialized domain of machine-generated source code.
