M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection

Yuxia WangJonibek MansurovPetar IvanovJinyan SuArtem ShelmanovAkim TsvigunOsama Mohammed AfzalTarek MahmoudGiovanni PuccettiThomas Arnold

article2024ACL68 citations

Introduces a multilingual, multi-domain benchmark covering binary, multi-generator, and boundary-level mixed-text detection to evaluate automated detectors and human ability in identifying black-box machine-generated text.

Listen

The rapid advancement and adoption of large language models have enabled the generation of human-like text at scale, introducing significant risks around disinformation, academic dishonesty, intellectual property infringement, and the erosion of trust in digital media. Addressing these risks requires effective mechanisms to differentiate machine-generated content from genuine human writing. However, existing detection frameworks often oversimplify real-world usage by focusing predominantly on English, assuming full-text machine generation, and relying strictly on binary classification without identifying specific source models.

The article introduces a comprehensive benchmark, M4GT-Bench, designed to evaluate black-box detection methods across realistic conditions. The primary objective is to evaluate how well both automated detection models and human readers perform across three practical tasks: binary classification distinguishing human from machine-generated text in monolingual and multilingual contexts; multi-way classification attributing machine-generated text to specific language models; and boundary detection locating the exact change point where human writing transitions into machine-generated text within mixed documents.

To establish this benchmark, the researchers compiled a corpus spanning nine languages, six distinct domains (including news, Wikipedia, student essays, and scientific paper abstracts/reviews), and nine modern language models (including ChatGPT, GPT-4, and the LLaMA-2 series). The evaluation tested multiple supervised detection approaches, including standard Transformer encoders (such as RoBERTa, XLM-R, DeBERTa-v3, and Longformer) alongside classifiers built on statistical word rankings, stylometric signals, and content-based features. In parallel, a controlled human study was conducted to establish a baseline for human capability in identifying specific generating models.

The findings demonstrate four critical outcomes. First, human evaluators struggle significantly with source attribution: even after reviewing reference examples, human accuracy in identifying the specific generator among four options was only 21.2%, which is below random guess performance (25%). Second, automated models achieve strong performance in binary and multi-way detection when trained and tested on the same distributions—reaching accuracies above 96–97% using Transformer architectures—but their accuracy drops severely when confronted with unseen generator models or unfamiliar domains (for example, falling from 97% to as low as 36.5% on certain out-of-domain data). Third, in boundary detection for mixed human-machine text, sequence taggers accurately identify transitions when evaluated on familiar generators (achieving Mean Absolute Error under 3 words), but error rates increase dramatically to deviations of 15 to over 50 words when generalizing to unseen models and domains. Fourth, multilingual detection showed substantial performance degradation when testing on low-resource or distantly related languages, with macro F1-scores dropping below 50% for several non-Latin languages.

These results imply that current automated text detectors are brittle and cannot be reliably deployed as out-of-the-box safeguards against new or unobserved generative models. For organizational leadership and policymakers, relying heavily on existing detection software introduces operational compliance risks and high rates of false positives or missed detections, particularly in high-stakes domains such as legal reviews, academic integrity monitoring, and cybersecurity. The observed high recall but low precision in many multilingual settings also creates a risk of disproportionately misclassifying genuine human text as machine-generated.

Based on these findings, decision-makers should avoid using standalone black-box detection tools for definitive disciplinary or legal decisions without human-in-the-loop verification. Organizations seeking to deploy automated text detection should maintain dynamic training pipelines that continually incorporate data from newly released model families and target operational domains. Further research and development should explore alternative detection paradigms, such as watermarking and few-shot in-context learning, while expanding mixed-text evaluation to encompass complex editing scenarios with multiple change points.

The conclusions should be interpreted within the scope of the article's limitations. The boundary detection task assumed a single transition from human to machine text, which does not capture all nuances of iterative co-authoring or human rewriting. Additionally, the multi-way attribution models remain vulnerable to style shifts and text obfuscation techniques, warranting caution when interpreting attribution results in adversarial environments.

  • Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). It extends the evaluation of black-box text detectors by analyzing their fundamental vulnerabilities against deliberate evasion attacks such as recursive paraphrasing and spoofing.
  • Paper: Scalable watermarking for identifying large language model outputs, Sumanth Dathathri et al. (2024). It pursues generative watermarking as a robust, scalable alternative to post-hoc black-box text classification, directly addressing the detection brittleness demonstrated in M4GT-Bench.
Cover for M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection

Abstract

The advent of Large Language Models (LLMs) has brought an unprecedented surge in machine-generated text (MGT) across diverse channels. This raises legitimate concerns about its potential misuse and societal implications. The need to identify and differentiate such content from genuine human-generated text is critical in combating disinformation, preserving the integrity of education and scientific fields, and maintaining trust in communication. In this work, we address this problem by introducing a new benchmark based on a multilingual, multi-domain, and multi-generator corpus of MGTs — M4GT-Bench. The benchmark is compiled of three tasks: (1) mono-lingual and multi-lingual binary MGT detection; (2) multi-way detection where one need to identify, which particular model generated the text; and (3) mixed human-machine text detection, where a word boundary delimiting MGT from human-written content should be determined. On the developed benchmark, we have tested several MGT detection baselines and also conducted an evaluation of human performance. We see that obtaining good performance in MGT detection usually requires an access to the training data from the same domain and generators. The benchmark is available at https://github.com/mbzuai-nlp/M4GT-Bench.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Datasets and Metrics
  • 3.1 Human vs. Machine
  • 3.2 Multi-way Detection
  • 3.3 Boundary Identification
  • 4 Human Evaluation
  • 5 Experiments
  • 5.1 Monolingual Human vs. Machine
  • 5.2 Multilingual Human vs. Machine
  • 5.3 Multi-way Detection
  • 5.4 Boundary Identification
  • 6 Conclusion and Future Direction
  • Limitations
  • Ethics and Broader Impact
  • References
  • Appendix
  • A Task 3 Template Prompt
  • B Human Evaluation Results
  • C Results of Multi-way Classification
  • D Examples of Task 1 and 2
  • E Examples of Boundary Identification

Knowls

  1. Knowl 1 — M4GT-Bench defines three black-box machine-generated text detection tasks

    definition

    M4GT-Bench is a public benchmark for detecting machine-generated text (MGT) across nine languages, six domains, and nine large-language-model generators. It defines three tasks:

    • Task 1, binary detection: classify a text as human-written or machine-generated. It has a monolingual English track and a multilingual track.
    • Task 2, generator attribution: assign a text to its specific source among human writing and six LLM generators: davinci-003, ChatGPT, GPT-4, Cohere, Dolly-v2, and BLOOMz.
    • Task 3, boundary identification: identify the word position at which a text changes from an initial human-written segment to a machine-generated continuation.

    The first two tasks evaluate document-level classification, whereas the third evaluates localization of a human–machine transition within a document.

  2. Knowl 2 — Task 1 combines balanced monolingual data with multilingual generalization data

    data/table

    The binary human-versus-machine portion of M4GT-Bench contains six domains: OUTFOX student essays, Wikipedia, WikiHow, Reddit ELI5, arXiv abstracts, and PeerRead paper reviews. For the monolingual English data, human-written texts were upsampled to reduce class imbalance, producing 65,177 human texts and 73,288 machine-generated texts. The machine data use davinci-003, ChatGPT, Cohere, Dolly-v2, and BLOOMz; GPT-4 continuations were additionally generated as an unseen-generator test set containing 14,344 texts.

    The multilingual extension includes English together with Arabic, Bulgarian, Chinese, Indonesian, Russian, Urdu, Italian, and German. Its underlying human source corpora contain 5,119,196 documents, while the parallel human–machine evaluation data contain 65,100 examples: 28,000 human texts, 9,000 davinci-003 texts, 25,000 ChatGPT texts, 100 Jais-30B texts, and 3,000 LLaMA-2 texts. The authors minimally clean human data by removing conversion and crawling artifacts such as repeated newlines, URLs, references, short Wikipedia paragraphs, and PDF-induced line breaks. Binary performance is measured with accuracy, precision, recall, and F1 with respect to the MGT class.

  3. Knowl 3 — Task 2 evaluates seven-way generator attribution on unseen domains

    experimental setup

    Task 2 treats generator identification as a seven-label classification problem with the labels human, davinci-003, ChatGPT, GPT-4, Cohere, Dolly-v2, and BLOOMz. Unlike Task 1, it uses the parallel data without upsampling human texts. The dataset includes the domains Wikipedia, WikiHow, Reddit ELI5, arXiv abstracts, and PeerRead, plus the newly collected OUTFOX student-essay domain for testing domain generalization.

    For each experiment, the classifier is trained on all domains except the domain used for testing. A separate randomly split setting, called the all-domain setting, includes every domain in training, validation, and test partitions. The evaluation reports accuracy, macro-F1, and class-wise precision, recall, and F1. This setup tests whether textual signatures learned from known domains remain useful when the subject matter and writing style change.

  4. Knowl 4 — Task 3 constructs human-to-machine continuations with one annotated boundary

    experimental setup

    Task 3 models a common mixed-authorship scenario in which a person begins a document and an LLM continues it. Each example consists of one human-written prefix followed by one machine-generated continuation, with exactly one change point. The human prefix contains between 0% and 50% of the words; a human proportion of 0% therefore denotes a fully machine-generated example.

    The continuations concern PeerRead academic-paper reviews and OUTFOX student essays. The prompts provide a paper title and abstract or an essay problem statement and partial essay, then request a continuation from ChatGPT, GPT-4, or LLaMA-2-7B, LLaMA-2-13B, or LLaMA-2-70B. PeerRead contains 5,676 examples per generator for ChatGPT, LLaMA-2-7B, LLaMA-2-13B, and LLaMA-2-70B, plus 5,189 examples for an alternative LLaMA-2-7B prompt; OUTFOX contains 1,000 test examples for each of GPT-4, LLaMA-2-7B, LLaMA-2-13B, and LLaMA-2-70B. The gold annotation is the word index where machine generation begins. Performance is measured by mean absolute error, defined as the average absolute difference between the predicted and gold boundary indices.

  5. Knowl 5 — The benchmark compares Transformer, ranking-feature, stylistic, and semantic boundary detectors

    model/method

    For Tasks 1 and 2, M4GT-Bench evaluates fine-tuned RoBERTa and XLM-R classifiers, logistic regression or SVM classifiers using GLTR features, and SVM or logistic-regression classifiers using stylistic and NELA features. The GLTR representation has 14 features: counts of tokens ranked in the top 10, top 100, top 1,000, or below rank 1,000 by a language model, plus a 10-bin distribution of the actual token probability divided by the maximum token probability at its position. The GLTR logistic-regression classifier uses a maximum of 1,000 training iterations.

    Stylistic features cover character, syntactic, structural, and word-level properties. NELA features describe style, complexity, bias, affect, moral-foundation dimensions, and events. For Task 3, Longformer and DeBERTa-v3 are used as token-level sequence taggers. Human-written tokens receive label 0 and machine-generated tokens label 1; the first token predicted as machine-generated is mapped to a word-level boundary. Classification experiments use five random seeds, while boundary experiments use three seeds and report means with standard deviations.

  6. Knowl 6 — Monolingual binary detectors generalize poorly to unseen generators

    empirical result

    In English human-versus-machine detection, detectors perform well when training, validation, and test splits contain all generators but lose substantial accuracy when the test generator is excluded from training. XLM-R is the strongest overall detector in the reported experiments: its all-generator setting achieves precision 95.08%, recall 98.80%, F1 96.87%, and accuracy 96.31%.

    When the test generator is unseen, XLM-R obtains accuracy 84.32% on davinci-003, 85.62% on ChatGPT, 77.95% on GPT-4, 86.23% on Cohere, 79.43% on Dolly-v2, and 73.07% on BLOOMz. RoBERTa, XLM-R, and the stylistic-feature classifier provide the best results in the all-generator setting; GLTR performs best for the unseen Cohere condition, while NELA obtains the highest accuracy for unseen GPT-4. BLOOMz is especially difficult because its output distribution differs markedly from the other generators. Overall, the results show that high binary accuracy generally depends on exposure to the relevant generator family.

  7. Knowl 7 — Cross-language binary detection is strongest for related or well-resourced languages

    empirical result

    The multilingual XLM-R detector is trained on all languages except the language being tested. Its all-language random split reaches precision 91.49%, recall 98.52%, F1 94.86%, accuracy 94.52%, and macro-F1 94.49%. The unseen-language results are:

    • Arabic: precision 88.91%, recall 97.35%, F1 92.64%, accuracy 92.18%, macro-F1 92.12%.
    • Bulgarian: 51.41%, 99.97%, 67.90%, 52.74%, and 39.18%.
    • Chinese: 76.85%, 95.18%, 84.73%, 82.51%, and 81.93%.
    • English: 63.72%, 84.64%, 72.42%, 66.51%, and 64.55%.
    • German: 66.21%, 98.30%, 79.00%, 73.62%, and 71.58%.
    • Indonesian: 53.13%, 100.0%, 69.39%, 55.83%, and 45.00%.
    • Italian: 75.20%, 100.0%, 85.84%, 83.51%, and 83.04%.
    • Russian: 52.55%, 86.38%, 65.26%, 53.70%, and 47.60%.
    • Urdu: 91.29%, 97.93%, 94.49%, 94.39%, and 94.39%.

    The five metrics in each language entry are precision, recall, F1, accuracy, and macro-F1, respectively. Bulgarian, Russian, and Indonesian form the weakest group, while Arabic and Urdu form the strongest. The detector frequently labels texts as machine-generated, yielding high MGT recall but lower precision and poorer recognition of human-written text, especially for low-resource languages or languages unrelated to those used for training.

  8. Knowl 8 — Generator attribution collapses on unfamiliar domains despite strong in-domain performance

    empirical result

    RoBERTa is the strongest Task 2 baseline, but its performance drops sharply when the test domain is excluded from training. In the all-domain split, RoBERTa reaches precision 96.96%, recall 97.01%, macro-F1 96.94%, and accuracy 97.00%. When evaluated on unseen domains, its accuracy and macro-F1 are 36.55% and 32.29% for arXiv, 69.47% and 66.89% for PeerRead, 74.49% and 71.66% for Reddit, 68.36% and 68.85% for WikiHow, 51.91% and 39.37% for Wikipedia, and 78.25% and 66.40% for OUTFOX.

    The arXiv domain is the hardest for most detectors. Among generator classes, davinci-003 is usually the most difficult to distinguish, followed by ChatGPT, because their output distributions are relatively similar. BLOOMz is the easiest: even the weakest NELA-SVM baseline achieves more than 90% class-wise F1 for BLOOMz. RoBERTa outperforms XLM-R and the feature-based methods overall, while one-versus-one SVM is better than one-versus-rest logistic regression for GLTR features.

  9. Knowl 9 — Boundary localization is highly sensitive to both generator and domain shift

    empirical result

    Longformer and DeBERTa-v3 are trained to tag human and machine tokens and are evaluated with mean absolute error in word-position units. When training and testing involve the same PeerRead generator, errors can be small, but cross-generator testing in the same domain is substantially harder. With training on both PeerRead generator types, Longformer obtains MAE 1.89 ± 0.79 on PeerRead LLaMA-2-7B* and 4.36 ± 0.36 on PeerRead ChatGPT; DeBERTa-v3 obtains 0.57 ± 0.23 and 2.63 ± 0.20, respectively.

    Training only on PeerRead ChatGPT yields Longformer MAE 31.43 ± 6.15 on PeerRead LLaMA-2-7B* and 4.55 ± 0.36 on PeerRead ChatGPT, while training only on PeerRead LLaMA-2-7B* yields 1.94 ± 0.07 on LLaMA-2-7B* and 51.379 ± 0.72 on ChatGPT. On the unseen OUTFOX domain containing multiple generators, errors increase to 21.54 ± 0.25 for Longformer and 15.55 ± 2.60 for DeBERTa-v3 when trained on both PeerRead generators. With LLaMA-2-7B* training alone, the OUTFOX errors rise to 53.62 ± 1.60 and 32.35 ± 0.78. Thus, locating a transition is even less generator- and domain-robust than document-level detection.

  10. Knowl 10 — Humans cannot reliably distinguish the source generators from five demonstrations

    empirical result

    The human evaluation samples 140 Task 2 examples from 35 domain–generator combinations: five domains crossed with six LLM generators and human writing. Annotators are divided into four groups, each seeing three domains, four classes, and five randomly selected demonstration examples that cover all four classes. Groups 1 through 4 consist of native Italian, Chinese, English, and Russian speakers, respectively, who are NLP postdoctoral researchers or PhD students.

    The four-class accuracies are 27.42% for Group 1, 10.45% for Group 2, 15.47% for Group 3, and 24.82% for Group 4. Combining Groups 1 and 2 gives 20.27%, combining Groups 3 and 4 gives 21.13%, and the overall accuracy is 21.20%, with overall precision 22.67%, recall 22.11%, and macro-F1 21.06%. Because random guessing among four classes yields 25% accuracy, the participants perform at or below chance. Annotators report relying on superficial cues such as repetitive output, formatting artifacts, perceived quality differences among GPT models, typos, and URLs, but five demonstrations are insufficient for learning reliable generator-specific signatures.

Coverage note — The paper’s explicit limitation analysis—vulnerability to paraphrasing, back-translation, and other adversarial attacks; inability of supervised multi-way classifiers to recognize unseen generators; and the simplifying assumption of exactly one human-to-machine boundary—was omitted as a separate knowl to keep the output to the ten most contribution-critical units.

References

  1. 1.Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang. 2023. Fast-detectgpt: Efficient zero-shot detection of machine-generated text via conditional probability curvature. arXiv preprint arXiv:2310.05130.
  2. 2.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. CoRR, abs/2004.05150.
  3. 3.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116.
  4. 4.Evan Crothers, Nathalie Japkowicz, Herna Viktor, and Paula Branco. 2022. Adversarial robustness of neural-statistical features in detection of generative transformers. In 2022 International Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE.
  5. 5.Liam Dugan, Daphne Ippolito, Arun Kirubarajan, and Chris Callison-Burch. 2020. RoFT: A tool for evaluating human detection of machine-generated text. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 189–196, Online. Association for Computational Linguistics.
  6. 6.Liam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi, and Chris Callison-Burch. 2023. Real or fake text?: Investigating human ability to detect boundaries between human-written and machine-generated text. In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7-14, 2023, pages 12763–12771. AAAI Press.
  7. 7.Chujie Gao, Dongping Chen, Qihui Zhang, Yue Huang, Yao Wan, and Lichao Sun. 2024. Llm-as-a-coauthor: The challenges of detecting llm-human mixcase. arXiv preprint arXiv:2401.05952.
  8. 8.Sebastian Gehrmann, Hendrik Strobelt, and Alexander Rush. 2019a. GLTR: Statistical detection and visualization of generated text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 111–116, Florence, Italy. Association for Computational Linguistics.
  9. 9.Sebastian Gehrmann, Hendrik Strobelt, and Alexander M Rush. 2019b. Gltr: Statistical detection and visualization of generated text. arXiv preprint arXiv:1906.04043.
  10. 10.Katy Ilonka Gero, Vivian Liu, and Lydia Chilton. 2022. Sparks: Inspiration for science writing using language models. In Designing interactive systems conference, pages 1002–1019.
  11. 11.Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean Wojcik, and Peter Ditto. 2012. Moral foundations theory: The pragmatic validity of moral pluralism. Advances in Experimental Social Psychology, 47.
  12. 12.Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023. How close is chatgpt to human experts? comparison corpus, evaluation, and detection. CoRR, abs/2301.07597.
  13. 13.Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2024. Spotting llms with binoculars: Zero-shot detection of machine-generated text. arXiv preprint arXiv:2401.12070.
  14. 14.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: decoding-enhanced bert with disentangled attention. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  15. 15.Xinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023. Mgtbench: Benchmarking machine-generated text detection. CoRR, abs/2303.14822.
  16. 16.Benjamin D Horne, Jeppe Nørregaard, and Sibel Adali. 2019. Robust fake news detection over time and attack. ACM Transactions on Intelligent Systems and Technology (TIST), 11(1):1–23.
  17. 17.Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Ho. 2023. Radar: Robust ai-text detection via adversarial learning. arXiv preprint arXiv:2307.03838.
  18. 18.Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2019. Automatic detection of generated text is easiest when humans are fooled. arXiv preprint arXiv:1911.00650.
  19. 19.John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023. A watermark for large language models. CoRR, abs/2301.10226.
  20. 20.Ryuto Koike, Masahiro Kaneko, and Naoaki Okazaki. 2023. Outfox: Llm-generated essay detection through in-context learning with adversarially generated examples. arXiv preprint arXiv:2307.11729.
  21. 21.Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer. 2023. Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense. arXiv preprint arXiv:2303.13408.
  22. 22.Tharindu Kumarage, Joshua Garland, Amrita Bhattacharjee, Kirill Trapeznikov, Scott Ruston, and Huan Liu. 2023. Stylometric detection of ai-generated text in twitter timelines. arXiv preprint arXiv:2303.03697.
  23. 23.Jenny S Li, John V Monaco, Li-Chiou Chen, and Charles C Tappert. 2014. Authorship authentication using short messages from social networking sites. In 2014 IEEE 11th International Conference on e-Business Engineering, pages 314–319. IEEE.
  24. 24.Xiaoming Liu, Zhaohan Zhang, Yichen Wang, Hang Pu, Yu Lan, and Chao Shen. 2022. Coco: Coherence-enhanced machine-generated text detection under data limitation with contrastive learning. arXiv preprint arXiv:2212.10341.
  25. 25.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  26. 26.Dominik Macko, Robert Moro, Adaku Uchendu, Ivan Srba, Jason Samuel Lucas, Michiharu Yamashita, Nafis Irtiza Tripto, Dongwon Lee, Jakub Simko, and Maria Bielikova. 2024. Authorship obfuscation in multilingual machine-generated text detection. arXiv preprint arXiv:2401.07867.
  27. 27.Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn. 2023. Detectgpt: Zero-shot machine-generated text detection using probability curvature. CoRR, abs/2301.11305.
  28. 28.Shaoor Munir, Brishna Batool, Zubair Shafiq, Padmini Srinivasan, and Fareed Zaffar. 2021. Through the looking glass: Learning to attribute synthetic text generated by language models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1811–1822.
  29. 29.Rafael Rivera Soto, Kailin Koch, Aleem Khan, Barry Chen, Marcus Bishop, and Nicholas Andrews. 2024. Few-shot detection of machine-generated text using style representations. arXiv e-prints, pages arXiv–2401.
  30. 30.Zhouxing Shi, Yihan Wang, Fan Yin, Xiangning Chen, Kai-Wei Chang, and Cho-Jui Hsieh. 2023. Red teaming language model detectors with language models. arXiv preprint arXiv:2305.19713.
  31. 31.Lei Shu, Liangchen Luo, Jayakumar Hoskere, Yun Zhu, Canoee Liu, Simon Tong, Jindong Chen, and Lei Meng. 2023. Rewritelm: An instruction-tuned large language model for text rewriting. arXiv preprint arXiv:2305.15685.
  32. 32.Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. 2019. Release strategies and the social impacts of language models. arXiv preprint arXiv:1908.09203.
  33. 33.Jinyan Su, Terry Yue Zhuo, Di Wang, and Preslav Nakov. 2023. Detectllm: Leveraging log rank information for zero-shot detection of machine-generated text. arXiv preprint arXiv:2306.05540.
  34. 34.Adaku Uchendu, Thai Le, Kai Shu, and Dongwon Lee. 2020. Authorship attribution for neural text generation. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), pages 8384–8395.
  35. 35.Adaku Uchendu, Zeyu Ma, Thai Le, Rui Zhang, and Dongwon Lee. 2021. Turingbench: A benchmark environment for turing test in the age of neural text generation. arXiv preprint arXiv:2109.13296.
  36. 36.Saranya Venkatraman, Adaku Uchendu, and Dongwon Lee. 2023. Gpt-who: An information density-based machine-generated text detector. arXiv preprint arXiv:2310.06202.
  37. 37.Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. 2021. Adversarial glue: A multi-task benchmark for robustness evaluation of language models. arXiv preprint arXiv:2111.02840.
  38. 38.Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Chenxi Whitehouse, Osama Mohammed Afzal, Tarek Mahmoud, Alham Fikri Aji, et al. 2023. M4: Multi-generator, multi-domain, and multi-lingual black-box machine-generated text detection. arXiv preprint arXiv:2305.14902.
  39. 39.Zhuohan Xie, Trevor Cohn, and Jey Han Lau. 2023. The next chapter: A study of large language models in storytelling. In Proceedings of the 16th International Natural Language Generation Conference, pages 323–351.
  40. 40.Feng Xiong, Thanet Markchom, Ziwei Zheng, Subin Jung, Varun Ojha, and Huizhi Liang. 2024. Fine-tuning large language models for multigenerator, multidomain, and multilingual machine-generated text detection. arXiv preprint arXiv:2401.12326.
  41. 41.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending against neural fake news. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 9051–9062.
  42. 42.Xuandong Zhao, Prabhanjan Ananth, Lei Li, and Yu-Xiang Wang. 2023a. Provable robust watermarking for ai-generated text. arXiv preprint arXiv:2306.17439.
  43. 43.Xuandong Zhao, Yu-Xiang Wang, and Lei Li. 2023b. Protecting language generation models via invisible watermarking. CoRR, abs/2302.03162.
  44. 44.Wanjun Zhong, Duyu Tang, Zenan Xu, Ruize Wang, Nan Duan, Ming Zhou, Jiahai Wang, and Jian Yin. 2020. Neural deepfake detection with factual structure of text. arXiv preprint arXiv:2010.07475.

Citation

MLA
Wang, Y., et al. “M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 3964–92, https://doi.org/10.18653/v1/2024.acl-long.218.
APA
Wang, Y., Mansurov, J., Ivanov, P., Su, J., Shelmanov, A., Tsvigun, A., Afzal, O. M., Mahmoud, T., Puccetti, G., Arnold, T., Aji, A., Habash, N., Gurevych, I., & Nakov, P. (2024). M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3964–3992. https://doi.org/10.18653/v1/2024.acl-long.218
Chicago
Wang, Y., J. Mansurov, P. Ivanov, et al. 2024. “M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3964–92. https://doi.org/10.18653/v1/2024.acl-long.218.
Harvard
Wang, Y. et al. (2024) “M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3964–3992. Available at: https://doi.org/10.18653/v1/2024.acl-long.218.
Vancouver
1. Wang Y, Mansurov J, Ivanov P, et al (2024) M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3964–3992

BibTeX

@inproceedings{wang-etal-2024-m4gt,
    title = "{M}4{GT}-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection",
    author = "Wang, Yuxia  and
      Mansurov, Jonibek  and
      Ivanov, Petar  and
      Su, Jinyan  and
      Shelmanov, Artem  and
      Tsvigun, Akim  and
      Mohammed Afzal, Osama  and
      Mahmoud, Tarek  and
      Puccetti, Giovanni  and
      Arnold, Thomas  and
      Aji, Alham  and
      Habash, Nizar  and
      Gurevych, Iryna  and
      Nakov, Preslav",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.218/",
    doi = "10.18653/v1/2024.acl-long.218",
    pages = "3964--3992"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/