A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models
Yihan WuZhengmian HuJunfeng GuoHongyang ZhangHeng Huang
Proposes DiPmark, a language model watermarking framework that maintains the original text generation distribution without quality loss while remaining provably resilient to text modifications and detectable without requiring model API or prompt access.
As artificial intelligence systems generate text that is increasingly indistinguishable from human writing, organizations face growing risks around academic integrity, intellectual property protection, and digital misinformation. Embedding covert identifiers, or watermarks, into model-generated text offers a viable path to track provenance and detect artificial origin. However, existing watermarking methods introduce significant trade-offs: they either distort the natural output distribution, degrading output quality; require extensive computing power and thousands of resampling steps; or require direct access to original prompts and internal model programming interfaces during detection.
The article demonstrates and evaluates DiPmark, a novel watermarking framework designed to embed reliable identifiers into language model outputs while strictly preserving original text quality and generation distributions. The primary objective is to deliver a watermarking mechanism that simultaneously guarantees distribution preservation, operational accessibility without requiring model access during detection, and provable resilience against deliberate text tampering.
To establish these properties, the authors designed a distribution-preserving reweighting strategy that adjusts token probabilities during generation using a secret key and a permutation hash function. The approach replaces traditional normal approximation tests with an exact statistical detection bound based on green token ratios. The evaluation tested DiPmark across standard tasks, including machine translation using Multilingual BART on the WMT 2014 dataset, text summarization using BART on the CNN-DailyMail corpus, and open-ended text generation using the 7-billion-parameter LLaMA-2 model, alongside an industry case study modifying top-token probabilities from GPT-4.
The findings confirm that DiPmark maintains original model text quality across all tested parameters, showing virtually identical BLEU, BERTScore, and perplexity metrics to unwatermarked baselines, whereas traditional baseline methods caused noticeable quality degradation. In terms of accessibility, DiPmark detected watermarks across 1,000 text sequences generated by LLaMA-2 in 90 seconds without access to prompts or internal model parameters, operating at least four times faster than comparable distribution-preserving detectors. For robustness, the method maintained strong detection performance under 20% to 30% text modifications, achieving area under the curve scores above 0.80 under random alterations and paraphrasing attacks, while accurately certifying error bounds. In the real-world GPT-4 test, the detector successfully identified 97 out of 100 watermarked sequences at a strict false positive rate below 1%.
These results indicate that enterprise and industry model providers can implement reliable provenance tracking without compromising system output quality or exposing proprietary model internals during verification. By mathematically guaranteeing error rates, the framework significantly reduces compliance risks and prevents the wrongful classification of human-written text that occurs with standard approximation methods. Moving forward, the article suggests that organizations adopting language model watermarking should consider distribution-preserving techniques, while future engineering should explore incorporating multiple detector configurations and heavier green-token weighting to further enhance detection sensitivity in low-entropy contexts.
- Paper: Protecting Language Generation Models via Invisible Watermarking, Xuandong Zhao et al. (2023). Its probability-level invisible watermarking provides a useful precursor for understanding how DiPmark embeds a signal without changing visible wording.
- Paper: CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks, Xuanli He et al. (2022). CATER’s effort to preserve natural output distributions while watermarking generation sets up the distribution-preservation trade-off DiPmark addresses.
- Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). Its later evaluation of watermarking under repeated paraphrasing tests the robustness limits that DiPmark seeks to withstand.
