Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy
Yu FuDeyi XiongYue Dong
Presents a semantic-aware watermarking algorithm for conditional text generation that preserves output quality in summarization and data-to-text tasks by aligning green-list vocabulary partitions with input context embeddings.
As artificial intelligence language models generate increasingly realistic text, organizations face mounting risks from misinformation, factual hallucinations, and unauthorized automated writing. Watermarking—a technique that embeds hidden mathematical patterns into generated vocabulary to identify AI origins—has emerged as a primary compliance and safety tool. However, current watermarking methods designed for open-ended text generation fail when applied to conditional text generation tasks, such as document summarization and data-to-text generation, where the output must closely reflect input source text. The article evaluates why existing watermarks degrade task performance and introduces a context-aware watermarking remedy to solve this issue.
To address this limitation, the authors developed a semantic-aware watermarking algorithm that links vocabulary selection directly to the input context. Rather than randomly splitting words into permitted (green) and restricted (red) lists at each step, the proposed method uses word embedding similarity to identify tokens closely related to the input source and prioritizes their inclusion on the green list. The researchers evaluated this approach across multiple model architectures, including BART and Flan-T5, using standard benchmarks for text summarization (CNN/DailyMail and XSum) and data-to-text generation (DART and WebNLG). They complemented automated quality and statistical detection metrics with blinded human evaluations.
The findings demonstrate that applying standard, unadapted watermarks causes severe performance drops in conditional generation—reducing quality scores by up to 96.99% under strict watermark constraints and up to 27.54% under softer settings—because crucial source words are randomly banned, triggering hallucinations. In contrast, the semantic-aware watermark restored generation quality across all models and tasks, recovering substantial performance (for instance, improving data-to-text BLEU scores by up to 21.67 times over standard watermarks). In human evaluations, judges preferred the semantic-aware watermarked summaries by a margin of 55.33% to 44.67%. The investigation also revealed a detection paradox: while the proposed method achieves high statistical detection confidence (z-scores), overall classification accuracy (measured by area under the curve, or AUC) slightly decreased because human writers also naturally rely heavily on source-related words.
These insights are critical for organizations implementing AI governance, safety controls, and content compliance. Deploying standard watermarks in conditional workflows introduces unacceptable operational risks, including severe factual distortion and degraded task output. The semantic-aware approach demonstrates that watermarking can be integrated into high-fidelity conditional generation systems without crippling performance, though decision-makers must account for the slight trade-off in absolute detection accuracy.
Organizations deploying AI watermarks in production should adopt semantic-aware constraints rather than generic random partitioning when performing summarization, translation, or structured reporting. Moving forward, engineering and research teams should conduct pilot implementations to balance the hyperparameter controlling the volume of included semantic tokens against detection requirements. Further research is recommended to refine detection classifiers, specifically addressing the overlapping vocabulary naturally shared between human writers and source texts.
- Paper: CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks, Xuanli He et al. (2022). This paper establishes the foundational concept of conditional watermarking in text generation models across tasks like summarization and translation, which the source study directly investigates and refines for conditional language models.
- Paper: PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models, Peixuan Li et al. (2023). This work explores black-box vocabulary-level watermarking techniques for pre-trained language models, providing key context for the baseline watermark mechanisms analyzed and adapted in the source paper.
- Paper: Who Wrote this Code? Watermarking for Code Generation, Taehyun Lee et al. (2024). This paper tackles the challenge of watermarking in highly constrained, low-entropy conditional generation settings by proposing selective entropy-thresholded watermarking for source code generation.
- Paper: Watermark Stealing in Large Language Models, Nikola Jovanovic et al. (2024). This study examines the security and adversarial vulnerabilities of distribution-modifying LLM watermarks by demonstrating black-box watermark stealing attacks.
- Paper: Scalable watermarking for identifying large language model outputs, Sumanth Dathathri et al. (2024). This work extends generative text watermarking to large-scale production environments by introducing non-distortionary tournament sampling algorithms.
- Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). This research analyzes fundamental evasion limits and attacks against statistical and watermarking-based detection methods across language models.
- Paper: Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods, Kathleen C. Fraser et al. (2025). This work comprehensively evaluates factors influencing the detectability and statistical trade-offs of modern machine-generated text detection methods.
- Paper: RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors, Liam Dugan et al. (2024). This paper presents a standardized benchmark evaluating the robust out-of-domain and adversarial performance of machine-generated text detection systems.
