Built independently by an author, for readers. Read the story and support ChapterPal

keyword

paraphrasing attacks

Paraphrasing attacks are adversarial techniques used to modify text, particularly artificial intelligence-generated content, by rewording sentences and altering vocabulary or structure while preserving the original semantic meaning. In natural language processing and content provenance, these attacks are deployed to evade automated AI-text detectors or remove embedded statistical watermarks. By disrupting specific token sequences, n-gram patterns, and stylistic signatures that detection algorithms rely upon, paraphrasing attacks obscure the synthetic origin of the text without compromising its readability or underlying message.

5 items

Ghostbuster: Detecting Text Ghostwritten by Large Language Models

Ghostbuster: Detecting Text Ghostwritten by Large Language Models

Vivek Verma, Eve Fleisig, Nicholas Tomlin, Dan Klein

OrganizationsUniversity of California Berkeley

Why you should read this

Introduces Ghostbuster, a black-box AI text detector that achieves state-of-the-art accuracy across diverse writing domains and unseen generation models by combining token probabilities from smaller reference language models through structured feature search.

We introduce Ghostbuster, a state-of-the-art system for detecting AI-generated text. Our method works by passing documents through a series of weaker language models, running a structured search over possible combinations of their features, and then training a classifier on the selected features to predict whether documents are AI-generated. Crucially, Ghostbuster does not require access to token probabilities from the target model, making it useful for detecting text generated by black-box or unknown models. In conjunction with our model, we release three new datasets of human- and AI-generated text as detection benchmarks in the domains of student essays, creative writing, and news articles. We compare Ghostbuster to several existing detectors, including DetectGPT and GPTZero, as well as a new RoBERTa baseline. Ghostbuster achieves 99.0 F1 when evaluated across domains, which is 5.9 F1 higher than the best preexisting model. It also outperforms all previous approaches in generalization across writing domains (+7.5 F1), prompting strategies (+2.1 F1), and language models (+4.4 F1). We also analyze our system’s robustness to a variety of perturbations and paraphrasing attacks, and evaluate its performance on documents by non-native English speakers.

Added

2026-10-01

SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation

SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation

Abe Bohan Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov

OrganizationsJohns Hopkins UniversityMassachusetts Institute of TechnologyTencentUniversity of WashingtonXi'an Jiaotong University

Why you should read this

Proposes a sentence-level text watermarking framework that embeds signals into semantic embedding spaces using locality-sensitive hashing and rejection sampling, ensuring machine-generated text remains detectable even after heavy paraphrasing.

Existing watermarked generation algorithms employ token-level designs and therefore, are vulnerable to paraphrase attacks. To address this issue, we introduce watermarking on the semantic representation of sentences. We propose SemStamp, a robust sentence-level semantic watermarking algorithm that uses locality-sensitive hashing (LSH) to partition the semantic space of sentences. The algorithm encodes and LSH-hashes a candidate sentence generated by a language model, and conducts rejection sampling until the sampled sentence falls in watermarked partitions in the semantic embedding space. To test the paraphrastic robustness of watermarking algorithms, we propose a “bigram paraphrase” attack that produces paraphrases with small bigram overlap with the original sentence. This attack is shown to be effective against existing token-level watermark algorithms, while posing only minor degradations to SemStamp. Experimental results show that our novel semantic watermark algorithm is not only more robust than the previous state-of-the-art method on various paraphrasers and domains, but also better at preserving the quality of generation.

Added

2026-10-01

MAGE: Machine-generated Text Detection in the Wild

MAGE: Machine-generated Text Detection in the Wild

Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, Yue Zhang

OrganizationsJilin UniversityTencentUniversity of Hong KongWestlake UniversityZhejiang University

Why you should read this

Presents MAGE, a large-scale testbed spanning 7 writing tasks and 27 language models to rigorously evaluate how well existing AI text detectors generalize to unseen domains, new models, and paraphrasing attacks.

Large language models (LLMs) have achieved human-level text generation, emphasizing the need for effective AI-generated text detection to mitigate risks like the spread of fake news and plagiarism. Existing research has been constrained by evaluating detection methods on specific domains or particular language models. In practical scenarios, however, the detector faces texts from various domains or LLMs without knowing their sources. To this end, we build a comprehensive testbed by gathering texts from diverse human writings and texts generated by different LLMs. Empirical results show challenges in distinguishing machine-generated texts from human-authored ones across various scenarios, especially out-of-distribution. These challenges are due to the decreasing linguistic distinctions between the two sources. Despite challenges, the top-performing detector can identify 86.54% out-of-domain texts generated by a new LLM, indicating the feasibility for application scenarios.

Added

2026-10-01

Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods

Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods

Kathleen C. Fraser, Hillary Dawkins, Svetlana Kiritchenko

OrganizationsNational Research Council Canada

Why you should read this

Presents a comprehensive review of state-of-the-art AI-generated text detection methods, datasets, and practical factors that govern how reliably machine-written content can be identified across real-world scenarios.

Large language models (LLMs) have advanced to a point that even humans have difficulty discerning whether a text was generated by another human, or by a computer. However, knowing whether a text was produced by human or artificial intelligence (AI) is important to determining its trustworthiness, and has applications in many domains including detecting fraud and academic dishonesty, as well as combating the spread of misinformation and political propaganda. The task of AI-generated text (AIGT) detection is therefore both very challenging, and highly critical. In this survey, we summarize state-of-the art approaches to AIGT detection, including watermarking, statistical and stylistic analysis, and machine learning classification. We also provide information about existing datasets for this task. Synthesizing the research findings, we aim to provide insight into the salient factors that combine to determine how “detectable” AIGT text is under different scenarios, and to make practical recommendations for future work towards this significant technical and societal challenge.

Added

2026-09-26