CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks
Xuanli HeQiongkai XuYi ZengLingjuan LyuFangzhao WuJiwei LiRuoxi Jia
Proposes a conditional watermarking framework that prevents model extraction attacks on text generation APIs by embedding statistically undetectable word-choice patterns while preserving natural output quality.
Commercial text generation systems face significant intellectual property risks from imitation attacks, where competitors query an application programming interface (API) and train copycat models on the returned outputs. While earlier defenses embedded lexical watermarks into API responses, attackers can easily identify and remove them because they alter overall word frequencies. The article introduces and evaluates CATER, a stealthy conditional watermarking framework designed to protect the intellectual property of text generation services by injecting watermarks tied to specific linguistic features while preserving natural word frequencies.
To evaluate the system, the authors conducted empirical experiments across multiple language generation benchmarks, primarily machine translation and document summarization, alongside mathematical proofs of watermark stealthiness. The evaluations tested multiple modern neural architectures, domain-shifted datasets, and adaptive watermark removal techniques. The watermarking approach balances an indistinguishability objective, which matches overall vocabulary distributions, with a distinctness objective that alters word choices conditioned on local syntactic features such as part-of-speech tags or dependency trees.
The findings show that CATER successfully identifies imitation models with high statistical confidence while causing negligible degradation in text quality. Across translation and summarization tasks, quality metrics such as BLEU and ROUGE remained within 0.3 points of clean, non-watermarked outputs. The watermarks remained robust even when imitation models used completely different neural architectures or out-of-domain query datasets. Furthermore, theoretical analysis and empirical stress tests confirmed that attempting to reverse-engineer conditional watermarks creates an astronomical number of false suspects, preventing adversaries from isolating and removing the embedded rules without severely degrading model performance.
These results demonstrate that API providers can reliably track and legally substantiate intellectual property theft without compromising the user experience of paying customers. Because previous watermarking methods were easily detected and stripped via basic statistical analysis, CATER provides a much more viable, production-ready defense for enterprise cloud services against model extraction.
Organizations operating proprietary text generation APIs should consider deploying conditional watermarking as an active verification mechanism, favoring part-of-speech conditioning for balanced stealth and detection. To prevent false ownership claims during disputes, stakeholders should enforce strict statistical significance thresholds. However, decision-makers should note key limitations: CATER requires access to high-quality synonym dictionaries, relies on post-hoc access to query suspected competitor APIs, and requires that at least half of the attacker's training data comes from the watermarked system to ensure definitive detection.
- Paper: Stealing Machine Learning Models via Prediction APIs, Florian Tramèr et al. (2016). This foundational paper establishes how black-box prediction APIs can be exploited to steal machine learning models, defining the core threat model that CATER aims to defend against.
- Paper: Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment, Di Jin et al. (2019). It provides the foundational framework for semantics-preserving lexical substitutions in NLP models, which informs CATER's synonym-based watermark insertion.
- Paper: Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations, Zirui Peng et al. (2022). It introduces black-box intellectual property verification and model extraction defense concepts that provide critical context for CATER's evaluation against knockoff models.
- Paper: Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark, Wenjun Peng et al. (2023). This work extends intellectual property protection against API model extraction to embedding-as-a-service through backdoor-based vector watermarking.
- Paper: PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models, Peixuan Li et al. (2023). This work builds on black-box IP defense paradigms for NLP by proposing PLMmark to protect pre-trained language models via digital signatures embedded in trigger words.
- Paper: Instructional Fingerprinting of Large Language Models, Jiashu Xu et al. (2024). It progresses beyond output-level text watermarking to fingerprinting generative LLMs against downstream fine-tuning and parameter modifications.
- Paper: Scalable watermarking for identifying large language model outputs, Sumanth Dathathri et al. (2024). It advances generative text watermarking to production-scale LLM architectures through non-distortionary tournament sampling during token generation.
- Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). It critically analyzes the fundamental limits and evasion vulnerabilities of generative AI text watermarking and detection schemes under recursive paraphrasing.
- Paper: LLM Dataset Inference: Did you train on my dataset?, Pratyush Maini et al. (2024). It extends copyright and IP auditing methodologies for LLMs from output watermarking to identifying whether proprietary datasets were used during pretraining.
