Built independently by an author, for readers. Read the story and support ChapterPal

keyword

watermark detection

Watermark detection is the computational process of identifying, verifying, or decoding an intentionally embedded, imperceptible signal within digital content or model outputs to determine its origin, ownership, or authenticity. In the context of artificial intelligence and machine learning, it involves analyzing digital artifacts such as synthetic text, code, or images using statistical tests, specialized decoding algorithms, or secret cryptographic keys to establish whether the material was produced by a specific generative model. The detector evaluates whether observed patterns, token probability distributions, or subtle feature variations deviate significantly from natural or unwatermarked baselines. This capability enables copyright protection, safeguards intellectual property against unauthorized model extraction or distillation, and helps audit machine-generated content, maintaining detection reliability even after the content undergoes downstream modifications such as editing, compression, or paraphrasing.

7 items

A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models

A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models

Yihan Wu, Zhengmian Hu, Junfeng Guo, Hongyang Zhang, Heng Huang

OrganizationsUniversity of MarylandUniversity of Waterloo

Why you should read this

Proposes DiPmark, a language model watermarking framework that maintains the original text generation distribution without quality loss while remaining provably resilient to text modifications and detectable without requiring model API or prompt access.

Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models. A challenge in the domain lies in preserving the distribution of original generated content after watermarking. Our research extends and improves upon existing watermarking framework, placing emphasis on the importance of a Distribution-Preserving (DiP) watermark. Contrary to the current strategies, our proposed DiPmark simultaneously preserves the original token distribution during watermarking (distribution-preserving), is detectable without access to the language model API and prompts (accessible), and is provably robust to moderate changes of tokens (resilient). DiPmark operates by selecting a random set of tokens prior to the generation of a word, then modifying the token distribution through a distribution-preserving reweight function to enhance the probability of these selected tokens during the sampling process. Extensive empirical evaluation on various language models and tasks demonstrates our approach’s distribution-preserving property, accessibility, and resilience, making it a effective solution for watermarking tasks that demand impeccable quality preservation. Code is available at¹.

Added

2026-10-04

Protecting Language Generation Models via Invisible Watermarking

Protecting Language Generation Models via Invisible Watermarking

Xuandong Zhao, Yu-Xiang Wang, Lei Li

OrganizationsDepartment of Computer ScienceUniversity of California, Santa Barbara

Why you should read this

Proposes an invisible watermarking framework, GINSEW, that embeds secret sinusoidal frequency signals directly into output probability distributions during decoding to reliably detect model extraction and distillation attacks without degrading text quality.

Language generation models have been an increasingly powerful enabler for many applications. Many such models offer free or affordable API access, which makes them potentially vulnerable to model extraction attacks through distillation. To protect intellectual property (IP) and ensure fair use of these models, various techniques such as lexical watermarking and synonym replacement have been proposed. However, these methods can be nullified by obvious countermeasures such as “synonym randomization”. To address this issue, we propose GINSEW, a novel method to protect text generation models from being stolen through distillation. The key idea of our method is to inject secret signals into the probability vector of the decoding steps for each target token. We can then detect the secret message by probing a suspect model to tell if it is distilled from the protected one. Experimental results show that GINSEW can effectively identify instances of IP infringement with minimal impact on the generation quality of protected APIs. Our method demonstrates an absolute improvement of 19 to 29 points on mean average precision (mAP) in detecting suspects compared to previous methods against watermark removal attacks.

Added

2026-10-01

SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation

SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation

Abe Bohan Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov

OrganizationsJohns Hopkins UniversityMassachusetts Institute of TechnologyTencentUniversity of WashingtonXi'an Jiaotong University

Why you should read this

Proposes a sentence-level text watermarking framework that embeds signals into semantic embedding spaces using locality-sensitive hashing and rejection sampling, ensuring machine-generated text remains detectable even after heavy paraphrasing.

Existing watermarked generation algorithms employ token-level designs and therefore, are vulnerable to paraphrase attacks. To address this issue, we introduce watermarking on the semantic representation of sentences. We propose SemStamp, a robust sentence-level semantic watermarking algorithm that uses locality-sensitive hashing (LSH) to partition the semantic space of sentences. The algorithm encodes and LSH-hashes a candidate sentence generated by a language model, and conducts rejection sampling until the sampled sentence falls in watermarked partitions in the semantic embedding space. To test the paraphrastic robustness of watermarking algorithms, we propose a “bigram paraphrase” attack that produces paraphrases with small bigram overlap with the original sentence. This attack is shown to be effective against existing token-level watermark algorithms, while posing only minor degradations to SemStamp. Experimental results show that our novel semantic watermark algorithm is not only more robust than the previous state-of-the-art method on various paraphrasers and domains, but also better at preserving the quality of generation.

Added

2026-10-01

Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy

Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy

Yu Fu, Deyi Xiong, Yue Dong

OrganizationsTianjin UniversityUniversity of California, Riverside

Why you should read this

Presents a semantic-aware watermarking algorithm for conditional text generation that preserves output quality in summarization and data-to-text tasks by aligning green-list vocabulary partitions with input context embeddings.

To mitigate potential risks associated with language models (LMs), recent AI detection research proposes incorporating watermarks into machine-generated text through random vocabulary restrictions and utilizing this information for detection. In this paper, we show that watermarking algorithms designed for LMs cannot be seamlessly applied to conditional text generation (CTG) tasks without a notable decline in downstream task performance. To address this issue, we introduce a simple yet effective semantic-aware watermarking algorithm that considers the characteristics of conditional text generation with the input context. Compared to the baseline watermarks, our proposed watermark yields significant improvements in both automatic and human evaluations across various text generation models, including BART and Flan-T5, for CTG tasks such as summarization and data-to-text generation. Meanwhile, it maintains detection ability with higher z-scores but lower AUC scores, suggesting the presence of a detection paradox that poses additional challenges for watermarking CTG.

Added

2026-09-26