RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

Liam DuganAlyssa HwangFilip TrhlíkAndrew ZhuJosh Magnus LudanHainiu XuDaphne IppolitoChris Callison-Burch

article2024ACL124 citations

Introduces a six-million-generation benchmark across eleven language models, eight domains, and eleven adversarial attacks, revealing that top AI text detectors easily fail when faced with minor sampling changes, repetition penalties, or unseen models.

Listen

As large language models become increasingly capable of generating human-like text, organizations face heightened risks involving phishing attacks, disinformation campaigns, spam, and spurious scientific publications. Although numerous commercial and open-source tools claim high accuracy—frequently exceeding 99%—in detecting machine-generated content, these claims are rarely evaluated against shared, rigorous standards. Existing evaluation datasets typically lack the adversarial modifications, varied decoding strategies, and modern generative models needed to test real-world effectiveness, creating an unverified sense of security among decision-makers.

The main objective of the article is to systematically assess the out-of-domain and adversarial robustness of current machine-generated text detectors using a comprehensive, standardized benchmark. To achieve this, the authors created the Robust AI Detection (RAID) dataset, which contains over 6 million generations constructed from roughly 15,000 human-written documents. The evaluation spans 11 generative models, 8 diverse topical domains, 4 decoding strategies, and 11 query-free adversarial attacks, benchmarking 12 leading open-source, metric-based, and commercial detectors while controlling for false positive rates.

The article establishes several key findings regarding detector reliability. First, open-source detectors using default classification thresholds suffer from dangerously high false positive rates, frequently misclassifying 20% to over 90% of human-written text as artificial. Second, even when calibrated to a strict 5% false positive rate, detector performance degrades dramatically under realistic generation settings: introducing standard repetition penalties and random sampling reduces accuracy by up to 32 percentage points compared to default greedy generation. Third, detectors exhibit severe domain and model bias, performing well on text produced by models they encountered during training (often reaching 95% accuracy) but plummeting to below 60% accuracy on unseen models or new domains. Finally, simple adversarial perturbations substantially undermine detection; subtle synonym swaps and homoglyph substitutions degrade accuracy across several detectors by 35 to over 70 percentage points.

These findings indicate that current text detectors are not sufficiently reliable for automated compliance, content moderation, or punitive actions such as academic disciplinary proceedings. Relying on uncalibrated or off-the-shelf detection tools exposes organizations to severe legal, reputational, and operational risks due to high false accusation rates. Furthermore, the vulnerability of detectors to basic adversarial edits means malicious actors can easily bypass automated filters, rendering legal or platform mandates for AI labeling largely unenforceable with current technology.

Decision-makers should refrain from deploying automatic text detectors in high-stakes or punitive contexts. Organizations that must use detection systems should explicitly calibrate classification thresholds on domain-specific human data rather than relying on default settings, and they should prioritize identifying direct harms—such as fraud, hate speech, and factual inaccuracy—over general machine authorship. Benchmark developers should continue maintaining dynamic evaluation datasets across multiple languages and evolving models, while practitioners should rely on continuous adversarial testing rather than static accuracy claims.

The analysis is bounded by the ongoing evolution of language models, which will require periodic updates to the dataset to reflect newer model releases, as well as limited multilingual and domain coverage outside the core English text domains. Nevertheless, the scale and methodological rigor of the benchmark support high confidence in the finding that contemporary detection tools are fragile and easily circumvented.

arXiv: 2405.07940
Cover for RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

Abstract

Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they are, the datasets used for evaluation are insufficiently challenging—lacking variations in sampling strategy, adversarial attacks, and open-source generative models. In this work we present RAID: the largest and most challenging benchmark dataset for machine-generated text detection. RAID includes over 6 million generations spanning 11 models, 8 domains, 11 adversarial attacks and 4 decoding strategies. Using RAID, we evaluate the out-of-domain and adversarial robustness of 8 open- and 4 closed-source detectors and find that current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models. We release our data1 along with a leaderboard2 to encourage future research.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3.1 Overview
  • 3.2 Domains
  • 3.3 Prompts
  • 3.4 Models
  • 3.5 Decoding Strategies
  • 3.6 Adversarial Attacks
  • 3.7 Post-Processing
  • 4 Dataset
  • 4.1 Statistics
  • 4.2 Release Structure
  • 4.3 RAID-Extra
  • 5 Detectors
  • 5.1 Detector Selection
  • 5.2 Detector Evaluation
  • 6 Findings
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Experiments on Multilingual and Code Generations (RAID-Extra)
  • A.1 Data Generation
  • A.2 Results
  • B Fixed FPR Accuracy vs. F1 Score
  • C Per-Domain Threshold Tuning
  • D Leaderboard and Pypi Package
  • E Dataset Details
  • E.1 Domains
  • E.2 Generative Models
  • E.3 Prompts
  • E.4 Adversarial Attacks
  • E.5 Repetition Penalty vs. Frequency Penalty vs. Presence Penalty
  • E.6 Hardware
  • F Detector Details
  • F.1 Detectors
  • F.2 Thresholds
  • G Extended Figures and Tables
  • G.1 Dataset Statistics and Evaluations
  • G.2 Extended Heatmaps
  • G.3 Model Performance vs. Domain
  • G.4 Detector Accuracy vs. Decoding Strategy
  • G.5 Detector Accuracy vs. Adversarial Attack
  • H Example Generations

Knowls

  1. Knowl 1 — RAID benchmark composition and public evaluation infrastructure

    model/method

    RAID (Robust AI Detection) is a shared benchmark for detecting machine-generated text under variation in generators, domains, decoding procedures, and adversarial perturbations. The core benchmark contains 14,971 human-written source documents, 509,014 non-adversarial generations, and 6,287,820 texts after including adversarially modified generations. It covers 11 generators—GPT-2 XL, GPT-3, ChatGPT, GPT-4, Mistral 7B and Mistral 7B Chat, MPT-30B and MPT-30B Chat, LLaMA 2 70B Chat, and Cohere Command and Command Chat—across eight domains: abstracts, books, news, poetry, recipes, Reddit, movie reviews, and Wikipedia.

    The authors release the data, evaluation scripts, and a public leaderboard. Ten percent of the core benchmark is released without labels as a hidden test set. The leaderboard separates detectors that report training on RAID from detectors that do not, allowing out-of-domain evaluation to be distinguished from benchmark-specific training.

  2. Knowl 2 — Controlled generation procedure for the core benchmark

    experimental setup

    RAID begins with approximately 2,000 human-written documents per core domain; the exact counts are 1,966 abstracts, 1,981 books, 1,980 news articles, 1,971 poems, 1,972 recipes, 1,979 Reddit posts, 1,143 movie reviews, and 1,979 Wikipedia articles. Each document supplies a title for a zero-shot generation prompt.

    Continuation-style generators receive prompts such as “The following is the full text of a news article titled ‘{title}’ from bbc.com,” while chat-style generators receive task instructions such as “Write the body of a BBC news article titled ‘{title}’.” Prompts avoid specifying a target length and use no source-document text beyond the title. After generation, prompts are removed, failed generations are filtered, and the data are balanced so that every retained human document has exactly one corresponding output for each applicable generator, decoding setting, repetition-penalty setting, and adversarial attack.

  3. Knowl 3 — Query-free adversarial attack suite

    model/method

    RAID models an adversary that has one query, no detector knowledge, and no access to gradients. Eleven black-box, query-free attacks are applied to generated text: alternative American-to-British spelling, article deletion, paragraph insertion between sentences, upper/lower-case swapping, zero-width-space insertion, whitespace addition, visually similar homoglyph substitution, number swapping, common misspellings, neural paraphrasing, and synonym substitution.

    The attack rate θ\theta is the manually reviewed fraction of available mutations applied to a passage. RAID uses the following rates: alternative spelling 100%, article deletion 50%, paragraph insertion 50%, upper/lower swap 5%, zero-width space 100%, whitespace addition 20%, homoglyph substitution 100%, number swap 50%, misspelling 20%, paraphrase 100%, and synonym substitution 50%.

    The paraphrase attack uses DIPPER-11B with sentence interval 3, lexical diversity 60, and order diversity 0. The synonym attack masks each token, obtains up to 20 BERT candidates in a context window of 20 tokens on either side, filters candidates by part of speech and FastText cosine similarity of at least 0.5, and selects the θL\theta L highest-likelihood substitutions, where LL is the passage length in tokens. The authors manually reviewed mutations so that attacks remained relatively inconspicuous to human readers.

  4. Knowl 4 — Detector suite and fixed-false-positive evaluation protocol

    experimental setup

    The evaluation uses 12 off-the-shelf detectors in three categories. Neural detectors are RoBERTa-Base trained on GPT-2 outputs, RoBERTa-Large trained on GPT-2 outputs, RoBERTa-Base trained on ChatGPT outputs, and RADAR. Metric-based detectors are GLTR, FastDetectGPT, LLMDet, and Binoculars. Commercial detectors are GPTZero, Originality, Winston, and ZeroGPT. The authors do not train neural detectors on RAID; metric-based detectors use the default generative models supplied by their implementations.

    Each detector maps an input text xx to a scalar score s(x)s(x). A text is classified as machine-generated when s(x)≥τs(x) \geq \tau, where τ\tau is a detector threshold. For each detector and each domain, the authors choose τ\tau using human-written RAID texts so that the false-positive rate is 5%; accuracy is then the fraction of machine-generated texts correctly identified at that fixed false-positive rate. Threshold search starts from the mean human score, changes the threshold iteratively toward the target, halves and reverses the step after overshooting, and stops when the false-positive rate is within 0.0005 of the target or after 50 iterations.

  5. Knowl 5 — Naive thresholds produce unsafe false-positive rates

    empirical result

    Using uncalibrated threshold values can make open-source detectors falsely label large portions of human-written RAID text as machine-generated. The false-positive rates at naive thresholds τ=0.25\tau=0.25, 0.50.5, 0.750.75, and 0.950.95 are, respectively:

    • RoBERTa-Base GPT-2: 8.71%, 6.59%, 5.18%, 3.38%.
    • RoBERTa-Large GPT-2: 6.14%, 2.91%, 1.46%, 0.25%.
    • RoBERTa-Base ChatGPT: 21.6%, 15.8%, 15.1%, 10.4%.
    • RADAR: 7.48%, 3.48%, 2.17%, 1.23%.
    • GLTR: 100%, 99.3%, 21.0%, 0.05%.
    • FastDetectGPT: 47.3%, 23.2%, 13.1%, 1.70%.
    • LLMDet: 97.9%, 96.0%, 92.0%, 75.3%.
    • Binoculars: 0.07%, 0.00%, 0.00%, 0.00%.
    • GPTZero: 0.03%, 0.00%, 0.00%, 0.00%.
    • Originality: 0.47%, 0.25%, 0.17%, 0.07%.
    • Winston: 0.75%, 0.55%, 0.38%, 0.21%.
    • ZeroGPT: 1.71%, 1.42%, 1.21%, 0.90%.

    Closed-source detectors remain below 1.7% false-positive rate at these thresholds, whereas several open-source metric-based detectors are unusable without calibration. Detector accuracy also depends sharply on the chosen false-positive rate: detectors can reach the very high accuracies reported in informal claims only at similarly high false-positive rates. Binoculars is particularly strong at low false-positive rates, while some detectors cannot reach very low rates and plateau at 16.9% for ZeroGPT, 0.88% for FastDetectGPT, and 0.62% for Originality. The results establish fixed and explicitly reported false-positive rates as necessary for meaningful detector comparisons.

  6. Knowl 6 — Sampling and repetition penalties substantially reduce detectability

    empirical result

    Across non-adversarial RAID text, detectors perform better on greedy decoding than on random sampling, and a repetition penalty further reduces detection accuracy. The mean accuracies at a 5% false-positive rate for the three detector categories are:

    • Commercial detectors: 90% with greedy decoding, 66% with sampling, 69% with greedy decoding plus repetition penalty, and 44% with sampling plus repetition penalty.
    • Neural detectors: 84%, 63%, 51%, and 34% under the same four conditions.
    • Metric-based detectors: 89%, 70%, 58%, and 20% under the same four conditions.

    The repetition penalty is multiplicative with factor 1.2 for the open-source generators that expose this parameter. Adding it decreases accuracy by as much as 32 percentage points across detectors, generator families, and domains. Random sampling is also consistently harder to detect than greedy decoding, including when the penalty is present. Thus, evaluation only on default greedy outputs substantially overestimates detector robustness.

  7. Knowl 7 — Detectors generalize poorly across generators and domains

    empirical result

    Detector performance is strongly associated with the generators and domains represented in detector training data. RoBERTa-Large trained on GPT-2 outputs exceeds 95% accuracy on five domains when the text is generated by GPT-2, but rarely exceeds 60% on the same domains when another generator is used. RADAR performs unusually poorly on movie reviews regardless of the generator. These patterns indicate that changing only the generator or domain can invalidate apparently strong detection results.

    On non-adversarial outputs averaged across generators, domains, and applicable decoding conditions, the total accuracies at 5% false-positive rate are 59.1% for RoBERTa-Base GPT-2, 56.7% for RoBERTa-Large GPT-2, 44.8% for RoBERTa-Base ChatGPT, 70.9% for RADAR, 62.6% for GLTR, 73.6% for FastDetectGPT, 35.0% for LLMDet, 79.6% for Binoculars, 66.5% for GPTZero, 85.0% for Originality, 71.0% for Winston, and 65.5% for ZeroGPT. Metric-based detectors, especially Binoculars and FastDetectGPT, show the strongest cross-generator generalization, but even the strongest detectors can experience error rates above 95% under a changed generator, decoding setting, or penalty.

  8. Knowl 8 — Adversarial vulnerabilities are detector-specific

    empirical result

    At a 5% false-positive rate, different detectors respond very differently to small, black-box text modifications. For six representative detectors, the following accuracies were measured on unmodified text and after six attacks; values in parentheses are changes from the unmodified baseline, in percentage points.

    • RoBERTa-Large GPT-2: 56.7%; paraphrase 72.9% (+16.2), synonym swap 79.4% (+22.7), misspelling 39.5% (-17.2), homoglyph 21.3% (-35.4), whitespace 40.1% (-16.6), article deletion 33.2% (-23.5).
    • RADAR: 70.9%; 67.3% (-3.6), 67.5% (-3.4), 69.5% (-1.4), 59.3% (-11.6), 66.1% (-4.8), 67.9% (-3.0).
    • GLTR: 62.6%; 47.2% (-15.4), 31.2% (-31.4), 59.8% (-2.8), 24.3% (-38.3), 45.8% (-16.8), 52.1% (-10.5).
    • Binoculars: 79.6%; 80.3% (+0.7), 43.5% (-36.1), 78.0% (-1.6), 37.7% (-41.9), 70.1% (-9.5), 74.3% (-5.3).
    • GPTZero: 66.5%; 64.0% (-2.5), 61.0% (-5.5), 65.1% (-1.4), 66.2% (-0.3), 66.2% (-0.3), 61.0% (-5.5).
    • Originality: 85.0%; 96.7% (+11.7), 96.5% (+11.5), 78.6% (-6.4), 9.3% (-75.7), 84.9% (-0.1), 71.4% (-13.6).

    The attack order in every row is paraphrase, synonym swap, misspelling, homoglyph, whitespace, and article deletion. Synonym replacement reduces Binoculars accuracy by 36.1 points, while homoglyphs reduce five of the six listed detectors by an average of 40.6 points; GPTZero loses only 0.3 points under homoglyphs. RADAR, which uses adversarial training, is comparatively stable. Some attacks improve accuracy, as paraphrasing and synonym replacement do for RoBERTa-Large GPT-2 and paraphrasing does for Binoculars. This shows that attack effectiveness depends on the detector's training distribution rather than being uniform across detectors.

  9. Knowl 9 — RAID generations differ from human text in repetitiveness, length, and perplexity

    data/table

    In the non-adversarial portion of RAID, human text contains 14,971 examples with a mean length of 378.5 tokens, SelfBLEU 7.64, LLaMA-7B perplexity 9.09, and GPT-2-XL perplexity 21.2. Across all 509,000 non-adversarial generations, the corresponding means are 323.4 tokens, SelfBLEU 13.7, LLaMA-7B perplexity 6.61, and GPT-2-XL perplexity 23.8.

    The generator-level statistics are given as number of generations, mean tokens, SelfBLEU, LLaMA-7B perplexity, and GPT-2-XL perplexity: GPT-2 has 59,884, 384.7, 23.9, 8.33, and 8.10; GPT-3 has 29,942, 185.6, 13.6, 3.90, and 8.12; ChatGPT has 29,942, 329.4, 10.3, 3.39, and 9.31; GPT-4 has 29,942, 350.8, 9.42, 5.01, and 13.4; Cohere has 29,942, 301.9, 11.0, 5.67, and 23.7; Cohere Chat has 29,942, 239.0, 11.0, 4.93, and 11.6; Mistral has 59,884, 370.2, 19.1, 7.74, and 17.9; Mistral Chat has 59,884, 287.7, 9.16, 4.31, and 10.3; MPT has 59,884, 379.2, 22.1, 14.0, and 66.9; MPT Chat has 59,884, 219.2, 5.39, 7.06, and 56.3; and LLaMA Chat has 59,884, 404.4, 10.6, 3.33, and 9.76.

    The aggregate statistics show that machine-generated text is typically shorter and more repetitive than human text, while its mean perplexity is lower under LLaMA 7B. Chat models are generally less repetitive than continuation models, but the distributions remain sufficiently distinct that detectors can exploit generator-specific regularities.

  10. Knowl 10 — Benchmark scope, maintenance, and deployment limitations

    limitation

    RAID does not provide comprehensive multilingual coverage: multilingual data are concentrated in the additional RAID-extra release, which contains 2.3 million generations in Python code, Czech news, and German news rather than many languages across many domains. The benchmark can also become obsolete as new language models and generation settings appear, requiring repeated dataset and metric updates.

    A public robustness benchmark can itself induce overfitting. Detector developers may specialize in the particular domains, generators, decoding strategies, and attacks represented by RAID without explicitly training on its examples, recreating the overreported-accuracy problem that the benchmark is intended to address. Hidden test data and periodically updated releases mitigate but do not eliminate this limitation. The authors consequently caution against using current detectors in disciplinary or punitive settings, where false positives can cause harm.

Coverage note — The full RAID-extra detector matrix, detailed implementations of every evaluated detector, and auxiliary threshold tables were omitted because they are ancillary to the core benchmark construction and main robustness findings.

References

  1. 1.Aaditya Bhat. 2023. Gpt-wiki-intro.
  2. 2.Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. The falcon series of open language models.
  3. 3.Micael Arman. 2020. Poems dataset (nlp). https://www.kaggle.com/datasets/michaelarman/poemsdataset.
  4. 4.Shahryar Baki, Rakesh Verma, Arjun Mukherjee, and Omprakash Gnawali. 2017. Scaling and effectiveness of email masquerade attacks: Exploiting natural language generation. In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security, ASIA CCS ’17, page 469–482, New York, NY, USA. Association for Computing Machinery.
  5. 5.David Bamman and Noah A. Smith. 2013. New alignment methods for discriminative book summarization.
  6. 6.Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang. 2023. Fast-detectgpt: Efficient zero-shot detection of machine-generated text via conditional probability curvature.
  7. 7.Meghana Moorthy Bhat and Srinivasan Parthasarathy. 2020. How effectively can machines defend against machine-generated fake news? an empirical study. In Proceedings of the First Workshop on Insights from Negative Results in NLP, pages 48–53, Online. Association for Computational Linguistics.
  8. 8.Michał Bien, Michał Gilski, Martyna Maciejewska, Wojciech Taisner, Dawid Wisniewski, and Agnieszka Lawrynowicz. 2020. RecipeNLG: A cooking recipes dataset for semi-structured text generation. In Proceedings of the 13th International Conference on Natural Language Generation, pages 22–28, Dublin, Ireland. Association for Computational Linguistics.
  9. 9.Steven Bird and Edward Loper. 2004. NLTK: The natural language toolkit. In Proceedings of the ACL Interactive Poster and Demonstration Sessions, pages 214–217, Barcelona, Spain. Association for Computational Linguistics.
  10. 10.Matyáš Bohácek, Michal Bravanský, Filip Trhlík, and Václav Moravec. 2022. Fine-grained czech news article dataset: An interdisciplinary approach to trustworthiness analysis.
  11. 11.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146.
  12. 12.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  13. 13.Shuyang Cai and Wanyun Cui. 2023. Evade chatgpt detectors via a single space.
  14. 14.Megha Chakraborty, S.M Towhidul Islam Tonmoy, S M Mehedi Zaman, Shreya Gautam, Tanay Kumar, Krish Sharma, Niyar Barman, Chandan Gupta, Vinija Jain, Aman Chadha, Amit Sheth, and Amitava Das. 2023. Counter Turing test (CT2): AI-generated text detection is not as easy as you may think - introducing AI detectability index (ADI). In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2206–2239, Singapore. Association for Computational Linguistics.
  15. 15.Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021. All that’s ‘human’ is not gold: Evaluating human evaluation of generated text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7282–7296, Online. Association for Computational Linguistics.
  16. 16.NLP Team Cohere. 2024. World-class ai, at your command. Accessed: 2024-02-02.
  17. 17.Evan N. Crothers, Nathalie Japkowicz, and Herna L. Viktor. 2023. Machine-generated text: A comprehensive survey of threat models and detection methods. IEEE Access, 11:70977–71002.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  19. 19.Liam Dugan, Daphne Ippolito, Arun Kirubarajan, and Chris Callison-Burch. 2020. RoFT: A tool for evaluating human detection of machine-generated text. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 189–196, Online. Association for Computational Linguistics.
  20. 20.Liam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi, and Chris Callison-Burch. 2023. Real or fake text? investigating human ability to detect boundaries between human-written and machine-generated text. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence, AAAI’23/IAAI’23/EAAI’23. AAAI Press.
  21. 21.Salijona Dyrmishi, Salah Ghamizi, and Maxime Cordy. 2023. How do humans perceive adversarial text? a reality check on the validity and naturalness of word-based adversarial attacks. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8822–8836, Toronto, Canada. Association for Computational Linguistics.
  22. 22.Vitalii Fishchuk and Daniel Braun. 2023. Efficient black-box adversarial attacks on neural text detectors.
  23. 23.Rinaldo Gagiano, Maria Myung-Hee Kim, Xiuzhen Zhang, and Jennifer Biggs. 2021. Robustness analysis of grover for machine-generated news detection. In Proceedings of the The 19th Annual Workshop of the Australasian Language Technology Association, pages 119–127, Online. Australasian Language Technology Association.
  24. 24.Chujie Gao, Dongping Chen, Qihui Zhang, Yue Huang, Yao Wan, and Lichao Sun. 2024. Llm-as-a-coauthor: The challenges of detecting llm-human mixcase.
  25. 25.Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. 2018. Black-box generation of adversarial text sequences to evade deep learning classifiers.
  26. 26.Sebastian Gehrmann, Hendrik Strobelt, and Alexander Rush. 2019. GLTR: Statistical detection and visualization of generated text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 111–116, Florence, Italy. Association for Computational Linguistics.
  27. 27.Derek Greene and Pádraig Cunningham. 2006. Practical solutions to the problem of diagonal dominance in kernel document clustering. In Proc. 23rd International Conference on Machine learning (ICML’06), pages 377–384. ACM Press.
  28. 28.Jesus Guerrero, Gongbo Liang, and Izzat Alsmadi. 2022. A mutation-based text generation for adversarial machine learning applications.
  29. 29.Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023. How close is chatgpt to human experts? comparison corpus, evaluation, and detection.
  30. 30.Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2024. Spotting llms with binoculars: Zero-shot detection of machine-generated text.
  31. 31.Julian Hazell. 2023. Large language models can be used to effectively scale spear phishing campaigns. arXiv preprint arXiv:2305.06972.
  32. 32.Xinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023. Mgtbench: Benchmarking machine-generated text detection.
  33. 33.John Hewitt, Christopher Manning, and Percy Liang. 2022. Truncation sampling as language model desmoothing. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 3414–3427, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  34. 34.Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Ho. 2023. Radar: Robust ai-text detection via adversarial learning. Advances in Neural Information Processing Systems.
  35. 35.Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2020. Automatic detection of generated text is easiest when humans are fooled. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1808–1822, Online. Association for Computational Linguistics.
  36. 36.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b.
  37. 37.Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation.
  38. 38.Tibor Kiss and Jan Strunk. 2006. Unsupervised multilingual sentence boundary detection. Computational Linguistics, 32(4):485–525.
  39. 39.Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries. 2022. The stack: 3 tb of permissively licensed source code. Preprint.
  40. 40.Ryuto Koike, Masahiro Kaneko, and Naoaki Okazaki. 2023. How you prompt matters! even task-oriented constraints in instructions affect llm-generated text detection.
  41. 41.Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer. 2023. Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense.
  42. 42.Pranav Kulkarni, Ziqing Ji, Yan Xu, Marko Neskovic, and Kevin Nolan. 2023. Exploring semantic perturbations on grover.
  43. 43.Tharindu Kumarage, Paras Sheth, Raha Moraffah, Joshua Garland, and Huan Liu. 2023. How reliable are ai-generated-text detectors? an assessment framework using evasive soft prompts.
  44. 44.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424–8445, Dublin, Ireland. Association for Computational Linguistics.
  45. 45.Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. Contrastive decoding: Open-ended text generation as optimization. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12286–12312, Toronto, Canada. Association for Computational Linguistics.
  46. 46.Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, and Yue Zhang. 2024. Mage: Machine-generated text detection in the wild.
  47. 47.Gongbo Liang, Jesus Guerrero, and Izzat Alsmadi. 2023a. Mutation-based adversarial attacks on neural text detectors.
  48. 48.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023b. Holistic evaluation of language models.
  49. 49.Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Zou. 2023c. Gpt detectors are biased against non-native english writers.
  50. 50.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  51. 51.Ning Lu, Shengcai Liu, Rui He, Qi Wang, Yew-Soon Ong, and Ke Tang. 2023. Large language models can be guided to evade ai-generated text detection.
  52. 52.Brady D Lund, Ting Wang, Nishith Reddy Mannuru, Bing Nie, Somipam Shimray, and Ziang Wang. 2023. Chatgpt and a new academic reality: Artificial intelligence-written research papers and the ethics of the large language models in scholarly publishing. Journal of the Association for Information Science and Technology, 74(5):570–581.
  53. 53.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  54. 54.Dominik Macko, Robert Moro, Adaku Uchendu, Jason Lucas, Michiharu Yamashita, Matúš Pikuliak, Ivan Srba, Thai Le, Dongwon Lee, Jakub Simko, and Maria Bielikova. 2023. MULTITuDE: Large-scale multilingual machine-generated text detection benchmark. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9960–9987, Singapore. Association for Computational Linguistics.
  55. 55.Dominik Macko, Robert Moro, Adaku Uchendu, Ivan Srba, Jason Samuel Lucas, Michiharu Yamashita, Nafis Irtiza Tripto, Dongwon Lee, Jakub Simko, and Maria Bielikova. 2024. Authorship obfuscation in multilingual machine-generated text detection.
  56. 56.Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell. 2023. Locally typical sampling. Transactions of the Association for Computational Linguistics, 11:102–121.
  57. 57.Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn. 2023. Detectgpt: zero-shot machine-generated text detection using probability curvature. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org.
  58. 58.NLP Team MosaicML. 2023. Introducing mpt-30b: Raising the bar for open-source foundation models. Accessed: 2023-06-22.
  59. 59.Edoardo Mosca, Mohamed Hesham Ibrahim Abdalla, Paolo Basso, Margherita Musumeci, and Georg Groh. 2023. Distinguishing fact from fiction: A benchmark dataset for identifying machine-generated scientific papers in the LLM era. In Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023), pages 190–207, Toronto, Canada. Association for Computational Linguistics.
  60. 60.OpenAI. 2022. ChatGPT: Optimizing Language Models for Dialogue.
  61. 61.OpenAI. 2023. Gpt-4 technical report.
  62. 62.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  63. 63.Artidoro Pagnoni, Martin Graciarena, and Yulia Tsvetkov. 2022. Threat scenarios and best practices to detect neural fake news. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1233–1249, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  64. 64.Sayak Paul and Soumik Rakshit. 2021. arxiv paper abstracts. https://www.kaggle.com/datasets/spsayakpaul/arxiv-paper-abstracts.
  65. 65.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only.
  66. 66.J. Pu, Z. Sarwar, S. Abdullah, A. Rehman, Y. Kim, P. Bhattacharya, M. Javed, and B. Viswanath. 2023a. Deepfake text detection: Limitations and opportunities. In 2023 IEEE Symposium on Security and Privacy (SP), pages 1613–1630, Los Alamitos, CA, USA. IEEE Computer Society.
  67. 67.Xiao Pu, Jingyu Zhang, Xiaochuang Han, Yulia Tsvetkov, and Tianxing He. 2023b. On the zero-shot generalization of machine-generated text detectors.
  68. 68.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  69. 69.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  70. 70.Veselin Raychev, Pavol Bielik, and Martin Vechev. 2016. Probabilistic model for code with decision trees. SIGPLAN Not., 51(10):731–747.
  71. 71.Juan Diego Rodriguez, Todd Hay, David Gros, Zain Shamsi, and Ravi Srinivasan. 2022. Cross-domain detection of GPT-2-generated technical text. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1213–1233, Seattle, United States. Association for Computational Linguistics.
  72. 72.Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi. 2023. Can ai-generated text be reliably detected?
  73. 73.Areg Mikael Sarvazyan, José Ángel González, Paolo Rosso, and Marc Franco-Salvador. 2023a. Supervised machine-generated text detectors: Family and scale matters. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 14th International Conference of the CLEF Association, CLEF 2023, Thessaloniki, Greece, September 18–21, 2023, Proceedings, page 121–132, Berlin, Heidelberg. Springer-Verlag.
  74. 74.Areg Mikael Sarvazyan, José Ángel González, Marc Franco-Salvador, Francisco Rangel, Berta Chulvi, and Paolo Rosso. 2023b. Overview of autextification at iberlef 2023: Detection and attribution of machine-generated text in multiple domains.
  75. 75.Dietmar Schabus, Marcin Skowron, and Martin Trapp. 2017. One million posts: A data set of german online discussions. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’17, page 1241–1244, New York, NY, USA. Association for Computing Machinery.
  76. 76.Tal Schuster, Roei Schuster, Darsh J. Shah, and Regina Barzilay. 2020. The limitations of stylometry for detecting machine-generated fake news. Computational Linguistics, 46(2):499–510.
  77. 77.Tatiana Shamardina, Vladislav Mikhailov, Daniil Chernianskii, Alena Fenogenova, Marat Saidov, Anastasiya Valeeva, Tatiana Shavrina, Ivan Smurov, Elena Tutubalina, and Ekaterina Artemova. 2022. Findings of the the ruatd shared task 2022 on artificial text detection in russian. In Computational Linguistics and Intellectual Technologies. RSUH.
  78. 78.Filipo Sharevski, Jennifer Vander Loop, Peter Jachim, Amy Devine, and Emma Pieroni. 2023. Talking abortion (mis) information with chatgpt on tiktok. In 2023 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW), pages 594–608. IEEE.
  79. 79.Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, Miles McCain, Jason Newhouse, Jason Blazakis, Kris McGuffie, and Jasmine Wang. 2019. Release strategies and the social impacts of language models.
  80. 80.Rafael Rivera Soto, Kailin Koch, Aleem Khan, Barry Chen, Marcus Bishop, and Nicholas Andrews. 2024. Few-shot detection of machine-generated text using style representations.
  81. 81.Giovanni Spitale, Nikola Biller-Andorno, and Federico Germani. 2023. Ai model gpt-3 (dis)informs us better than humans. Science Advances, 9(26):eadh1850.
  82. 82.Harald Stiff and Fredrik Johansson. 2022. Detecting computer-generated disinformation. International Journal of Data Science and Analytics, 13:363–383.
  83. 83.Zhenpeng Su, Xing Wu, Wei Zhou, Guangyuan Ma, and Songlin Hu. 2024. Hc3 plus: A semantic-invariant human chatgpt comparison corpus.
  84. 84.Edward Tian and Alexander Cui. 2023. Gptzero: Towards detection of ai-generated text using zero-shot and supervised methods.
  85. 85.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  86. 86.Adaku Uchendu, Zeyu Ma, Thai Le, Rui Zhang, and Dongwon Lee. 2021. Turingbench: A benchmark environment for turing test in the age of neural text generation.
  87. 87.Vivek Verma, Eve Fleisig, Nicholas Tomlin, and Dan Klein. 2023. Ghostbuster: Detecting text ghostwritten by large language models.
  88. 88.Michael Völske, Martin Potthast, Shahbaz Syed, and Benno Stein. 2017. TL;DR: Mining Reddit to learn automatic summarization. In Proceedings of the Workshop on New Frontiers in Summarization, pages 59–63, Copenhagen, Denmark. Association for Computational Linguistics.
  89. 89.Jian Wang, Shangqing Liu, Xiaofei Xie, and Yi Li. 2023a. Evaluating aigc detectors on code content.
  90. 90.Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Chenxi Whitehouse, Osama Mohammed Afzal, Tarek Mahmoud, Alham Fikri Aji, and Preslav Nakov. 2023b. M4: Multi-generator, multi-domain, and multi-lingual black-box machine-generated text detection.
  91. 91.Max Weiss. 2019. Deepfake bot submissions to federal public comment websites cannot be distinguished from human submissions. Technology Science.
  92. 92.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Huggingface’s transformers: State-of-the-art natural language processing.
  93. 93.Max Wolff. 2020. Attacking neural text detectors.
  94. 94.Kangxi Wu, Liang Pang, Huawei Shen, Xueqi Cheng, and Tat-Seng Chua. 2023. LLMDet: A third party large language models generated text detection tool. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 2113–2133, Singapore. Association for Computational Linguistics.
  95. 95.Han Xu, Jie Ren, Pengfei He, Shenglai Zeng, Yingqian Cui, Amy Liu, Hui Liu, and Jiliang Tang. 2023. On the generalization of training-based chatgpt detection methods.
  96. 96.Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018. Texygen: A benchmarking platform for text generation models.

Citation

MLA
Dugan, L., et al. “RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 12463–92, https://doi.org/10.18653/v1/2024.acl-long.674.
APA
Dugan, L., Hwang, A., Trhlík, F., Zhu, A., Ludan, J. M., Xu, H., Ippolito, D., & Callison-Burch, C. (2024). RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12463–12492. https://doi.org/10.18653/v1/2024.acl-long.674
Chicago
Dugan, L., A. Hwang, F. Trhlík, et al. 2024. “RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12463–92. https://doi.org/10.18653/v1/2024.acl-long.674.
Harvard
Dugan, L. et al. (2024) “RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12463–12492. Available at: https://doi.org/10.18653/v1/2024.acl-long.674.
Vancouver
1. Dugan L, Hwang A, Trhlík F, Zhu A, Ludan JM, Xu H, Ippolito D, Callison-Burch C (2024) RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 12463–12492

BibTeX

@inproceedings{dugan-etal-2024-raid,
    title = "{RAID}: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors",
    author = "Dugan, Liam  and
      Hwang, Alyssa  and
      Trhl{\'i}k, Filip  and
      Zhu, Andrew  and
      Ludan, Josh Magnus  and
      Xu, Hainiu  and
      Ippolito, Daphne  and
      Callison-Burch, Chris",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.674/",
    doi = "10.18653/v1/2024.acl-long.674",
    pages = "12463--12492"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/