Built independently by an author, for readers. Read the story and support ChapterPal

keyword

DiPmark statistic

The DiPmark statistic is a test score used to determine whether generated text carries a DiPmark watermark, based on how strongly its tokens match the watermark’s secret, randomly selected token sets; an unusually high score supports detection, while the watermarking method is designed to preserve the model’s original token distribution.

1 item

A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models

A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models

Yihan Wu, Zhengmian Hu, Junfeng Guo, Hongyang Zhang, Heng Huang

OrganizationsUniversity of MarylandUniversity of Waterloo

Why you should read this

Proposes DiPmark, a language model watermarking framework that maintains the original text generation distribution without quality loss while remaining provably resilient to text modifications and detectable without requiring model API or prompt access.

Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models. A challenge in the domain lies in preserving the distribution of original generated content after watermarking. Our research extends and improves upon existing watermarking framework, placing emphasis on the importance of a Distribution-Preserving (DiP) watermark. Contrary to the current strategies, our proposed DiPmark simultaneously preserves the original token distribution during watermarking (distribution-preserving), is detectable without access to the language model API and prompts (accessible), and is provably robust to moderate changes of tokens (resilient). DiPmark operates by selecting a random set of tokens prior to the generation of a word, then modifying the token distribution through a distribution-preserving reweight function to enhance the probability of these selected tokens during the sampling process. Extensive empirical evaluation on various language models and tasks demonstrates our approach’s distribution-preserving property, accessibility, and resilience, making it a effective solution for watermarking tasks that demand impeccable quality preservation. Code is available at¹.

Added

2026-10-04