M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection
Yuxia WangJonibek MansurovPetar IvanovJinyan SuArtem ShelmanovAkim TsvigunOsama Mohammed AfzalTarek MahmoudGiovanni PuccettiThomas Arnold
Introduces a multilingual, multi-domain benchmark covering binary, multi-generator, and boundary-level mixed-text detection to evaluate automated detectors and human ability in identifying black-box machine-generated text.
The rapid advancement and adoption of large language models have enabled the generation of human-like text at scale, introducing significant risks around disinformation, academic dishonesty, intellectual property infringement, and the erosion of trust in digital media. Addressing these risks requires effective mechanisms to differentiate machine-generated content from genuine human writing. However, existing detection frameworks often oversimplify real-world usage by focusing predominantly on English, assuming full-text machine generation, and relying strictly on binary classification without identifying specific source models.
The article introduces a comprehensive benchmark, M4GT-Bench, designed to evaluate black-box detection methods across realistic conditions. The primary objective is to evaluate how well both automated detection models and human readers perform across three practical tasks: binary classification distinguishing human from machine-generated text in monolingual and multilingual contexts; multi-way classification attributing machine-generated text to specific language models; and boundary detection locating the exact change point where human writing transitions into machine-generated text within mixed documents.
To establish this benchmark, the researchers compiled a corpus spanning nine languages, six distinct domains (including news, Wikipedia, student essays, and scientific paper abstracts/reviews), and nine modern language models (including ChatGPT, GPT-4, and the LLaMA-2 series). The evaluation tested multiple supervised detection approaches, including standard Transformer encoders (such as RoBERTa, XLM-R, DeBERTa-v3, and Longformer) alongside classifiers built on statistical word rankings, stylometric signals, and content-based features. In parallel, a controlled human study was conducted to establish a baseline for human capability in identifying specific generating models.
The findings demonstrate four critical outcomes. First, human evaluators struggle significantly with source attribution: even after reviewing reference examples, human accuracy in identifying the specific generator among four options was only 21.2%, which is below random guess performance (25%). Second, automated models achieve strong performance in binary and multi-way detection when trained and tested on the same distributions—reaching accuracies above 96–97% using Transformer architectures—but their accuracy drops severely when confronted with unseen generator models or unfamiliar domains (for example, falling from 97% to as low as 36.5% on certain out-of-domain data). Third, in boundary detection for mixed human-machine text, sequence taggers accurately identify transitions when evaluated on familiar generators (achieving Mean Absolute Error under 3 words), but error rates increase dramatically to deviations of 15 to over 50 words when generalizing to unseen models and domains. Fourth, multilingual detection showed substantial performance degradation when testing on low-resource or distantly related languages, with macro F1-scores dropping below 50% for several non-Latin languages.
These results imply that current automated text detectors are brittle and cannot be reliably deployed as out-of-the-box safeguards against new or unobserved generative models. For organizational leadership and policymakers, relying heavily on existing detection software introduces operational compliance risks and high rates of false positives or missed detections, particularly in high-stakes domains such as legal reviews, academic integrity monitoring, and cybersecurity. The observed high recall but low precision in many multilingual settings also creates a risk of disproportionately misclassifying genuine human text as machine-generated.
Based on these findings, decision-makers should avoid using standalone black-box detection tools for definitive disciplinary or legal decisions without human-in-the-loop verification. Organizations seeking to deploy automated text detection should maintain dynamic training pipelines that continually incorporate data from newly released model families and target operational domains. Further research and development should explore alternative detection paradigms, such as watermarking and few-shot in-context learning, while expanding mixed-text evaluation to encompass complex editing scenarios with multiple change points.
The conclusions should be interpreted within the scope of the article's limitations. The boundary detection task assumed a single transition from human to machine text, which does not capture all nuances of iterative co-authoring or human rewriting. Additionally, the multi-way attribution models remain vulnerable to style shifts and text obfuscation techniques, warranting caution when interpreting attribution results in adversarial environments.
- Paper: Cross-Domain Detection of GPT-2-Generated Technical Text, Juan Diego Rodriguez et al. (2022). It establishes the foundational domain transfer challenges and RoBERTa-based detection techniques for technical synthetic text that M4GT-Bench systematically scales up to multi-model and multilingual settings.
- Paper: CNN-Generated Images Are Surprisingly Easy to Spot… for Now, Sheng-Yu Wang et al. (2019). It provides foundational principles on cross-model generalization in synthetic artifact detection, motivating the cross-generator brittleness evaluated in M4GT-Bench.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). It diagnoses critical flaws in standard natural language generation evaluation and human distinguishing capabilities, providing the methodological rationale for the benchmark's controlled human baselines.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). It documents the severe performance drop across non-Latin and low-resource languages in generative AI, directly informing M4GT-Bench's multilingual detection setup.
- Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). It introduces core sampling dynamics and statistical properties of neural text generation that underpin the word-ranking and stylometric classifiers evaluated in M4GT-Bench.
- Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). It extends the evaluation of black-box text detectors by analyzing their fundamental vulnerabilities against deliberate evasion attacks such as recursive paraphrasing and spoofing.
- Paper: Scalable watermarking for identifying large language model outputs, Sumanth Dathathri et al. (2024). It pursues generative watermarking as a robust, scalable alternative to post-hoc black-box text classification, directly addressing the detection brittleness demonstrated in M4GT-Bench.
