CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks

Xuanli HeQiongkai XuYi ZengLingjuan LyuFangzhao WuJiwei LiRuoxi Jia

article2022NeurIPS95 citations

Proposes a conditional watermarking framework that prevents model extraction attacks on text generation APIs by embedding statistically undetectable word-choice patterns while preserving natural output quality.

Listen

Commercial text generation systems face significant intellectual property risks from imitation attacks, where competitors query an application programming interface (API) and train copycat models on the returned outputs. While earlier defenses embedded lexical watermarks into API responses, attackers can easily identify and remove them because they alter overall word frequencies. The article introduces and evaluates CATER, a stealthy conditional watermarking framework designed to protect the intellectual property of text generation services by injecting watermarks tied to specific linguistic features while preserving natural word frequencies.

To evaluate the system, the authors conducted empirical experiments across multiple language generation benchmarks, primarily machine translation and document summarization, alongside mathematical proofs of watermark stealthiness. The evaluations tested multiple modern neural architectures, domain-shifted datasets, and adaptive watermark removal techniques. The watermarking approach balances an indistinguishability objective, which matches overall vocabulary distributions, with a distinctness objective that alters word choices conditioned on local syntactic features such as part-of-speech tags or dependency trees.

The findings show that CATER successfully identifies imitation models with high statistical confidence while causing negligible degradation in text quality. Across translation and summarization tasks, quality metrics such as BLEU and ROUGE remained within 0.3 points of clean, non-watermarked outputs. The watermarks remained robust even when imitation models used completely different neural architectures or out-of-domain query datasets. Furthermore, theoretical analysis and empirical stress tests confirmed that attempting to reverse-engineer conditional watermarks creates an astronomical number of false suspects, preventing adversaries from isolating and removing the embedded rules without severely degrading model performance.

These results demonstrate that API providers can reliably track and legally substantiate intellectual property theft without compromising the user experience of paying customers. Because previous watermarking methods were easily detected and stripped via basic statistical analysis, CATER provides a much more viable, production-ready defense for enterprise cloud services against model extraction.

Organizations operating proprietary text generation APIs should consider deploying conditional watermarking as an active verification mechanism, favoring part-of-speech conditioning for balanced stealth and detection. To prevent false ownership claims during disputes, stakeholders should enforce strict statistical significance thresholds. However, decision-makers should note key limitations: CATER requires access to high-quality synonym dictionaries, relies on post-hoc access to query suspected competitor APIs, and requires that at least half of the attacker's training data comes from the watermarked system to ensure definitive detection.

Cover for CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks

Abstract

Previous works have validated that text generation APIs can be stolen through imitation attacks, causing IP violations. In order to protect the IP of text generation APIs, recent work has introduced a watermarking algorithm and utilized the null-hypothesis test as a post-hoc ownership verification on the imitation models. However, we find that it is possible to detect those watermarks via sufficient statistics of the frequencies of candidate watermarking words. To address this draw-back, in this paper, we propose a novel Conditional wATERmarking framework (CATER) for protecting the IP of text generation APIs. An optimization method is proposed to decide the watermarking rules that can minimize the distortion of overall word distributions while maximizing the change of conditional word selections. Theoretically, we prove that it is infeasible for even the savviest attacker (they know how CATER works) to reveal the used watermarks from a large pool of potential word pairs based on statistical inspection. Empirically, we observe that high-order conditions lead to an exponential growth of suspicious (unused) watermarks, making our crafted watermarks more stealthy. In addition, CATER can effectively identify IP infringement under architectural mismatch and cross-domain imitation attacks, with negligible impairments on the generation quality of victim APIs. We envision our work as a milestone for stealthily protecting the IP of text generation APIs.

Table of Contents

  • 1 Introduction
  • 2 Preliminary and Background
  • 2.1 Imitation Attack
  • 2.2 Identification of IP Infringement
  • 2.3 Watermark Removal
  • 3 CATER
  • 3.1 Watermarking Rule Optimization
  • 3.2 Constructing Watermarking Conditions using Linguistic Features
  • 3.3 Identifiability of Conditional Watermark
  • 4 Experiments
  • 4.1 Performance of CATER
  • 4.2 Analysis on Adaptive Attacks
  • 5 Conclusion
  • Limitation and Negative Societal Impacts
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — Two-stage post-hoc protection with conditional watermarks

    model/method

    CATER protects a black-box text-generation victim API VV through two stages. In the watermarking stage, VV receives user queries QQ, generates ordinary responses YY, and then selectively replaces words in synonym sets according to a secret conditional rule determined by a linguistic condition cc, producing watermarked responses Y′Y'. The rule changes which synonym is selected under a particular condition while aiming to preserve the aggregate word distribution.

    In the identification stage, the owner queries a suspected imitation API SS with a verification set OO and tests whether the returned text exhibits the conditional word-selection bias induced by CATER. The null hypothesis is that the tested model has no preference for watermark words. If kk watermark words occur among nn candidate words in the verification responses, and pp is the estimated natural-language probability of a watermark word, the two-sided binomial-style test is

    P=2min⁡(Pr⁡(X≥k),Pr⁡(X≤k)),X∼Binomial⁡(n,p).P=2\min\left(\Pr(X\geq k),\Pr(X\leq k)\right),\qquad X\sim\operatorname{Binomial}(n,p).

    Here PP is the ownership-test p-value, nn is the number of words from the candidate synonym vocabulary, kk is the number selected as watermark words, and pp is estimated from natural-language statistics; in the simple equal-synonym construction, the paper approximates pp by 1/R1/R, where RR is the number of interchangeable choices. A lower p-value indicates stronger evidence that the suspect API learned the victim's watermarking behavior. The workflow diagram on page 4 shows this separation between response watermarking and post-hoc querying of the suspect API.

  2. Knowl 2 — Optimization objective for stealthy conditional rules

    equation

    For a synonym set WW, let CC be the finite set of linguistic conditions, P(w∣c)P(w\mid c) the original conditional probability of word w∈Ww\in W under condition cc, P^(w∣c)\widehat P(w\mid c) the conditional probability after watermark assignment, and P(c)P(c) the empirical frequency of condition cc. CATER chooses the conditional distributions by minimizing

    min⁡P^D(∑c∈CP^(⋅∣c)P(c),  ∑c∈CP(⋅∣c)P(c))−α∣C∣∑c∈CD(P^(⋅∣c),P(⋅∣c)).\min_{\widehat P} D\left(\sum_{c\in C}\widehat P(\cdot\mid c)P(c),\;\sum_{c\in C}P(\cdot\mid c)P(c)\right) -\frac{\alpha}{|C|}\sum_{c\in C}D\left(\widehat P(\cdot\mid c),P(\cdot\mid c)\right).

    Here DD is a distribution-distance function and α>0\alpha>0 controls the tradeoff. The first term makes the marginal word distribution after watermarking close to the original marginal distribution, reducing detectability through word-frequency statistics. The second, subtracted term encourages the conditional distribution under each condition to differ from its original value, making the watermark observable during verification. Thus CATER hides the watermark globally while retaining a conditional signal.

  3. Knowl 3 — Discrete synonym assignment solved as a mixed-integer quadratic program

    model/method

    For each synonym group W={w1,…,wr}W=\{w_1,\ldots,w_r\}, CATER converts the rule-selection problem into a binary optimization problem. Let A∈[0,1]r×∣C∣A\in[0,1]^{r\times |C|} contain the empirical probabilities Ajc=P(wj∣c)A_{jc}=P(w_j\mid c), let π∈[0,1]∣C∣\boldsymbol\pi\in[0,1]^{|C|} contain the condition probabilities πc=P(c)\pi_c=P(c), and let X∈{0,1}r×∣C∣X\in\{0,1\}^{r\times |C|} be the optimized assignment matrix. The entry Xjc=1X_{jc}=1 means that synonym wjw_j is selected when condition cc occurs. CATER solves

    min⁡X  ∥(A−X)π∥22−α∣C∣∥A−X∥F2\min_X\;\|(A-X)\boldsymbol\pi\|_2^2 -\frac{\alpha}{|C|}\|A-X\|_F^2

    subject to

    XT1r=1∣C∣,X∈{0,1}r×∣C∣,X^{\mathsf T}\mathbf 1_r=\mathbf 1_{|C|},\qquad X\in\{0,1\}^{r\times |C|},

    where 1d\mathbf 1_d is a length-dd all-ones vector. The constraint assigns exactly one synonym to every condition. The empirical matrices are estimated from a large training corpus; for sufficiently small α\alpha the objective is convex, and the paper solves the binary program with Gurobi using α=0.01\alpha=0.01. The resulting assignments are the conditional watermarking rules used by the victim API.

  4. Knowl 4 — Part-of-speech and dependency conditions provide scalable watermark triggers

    model/method

    CATER constructs conditions from syntactic features rather than from individual words alone. For a target word, the part-of-speech (POS) label is denoted l0l_0, and l−kl_{-k} and l+kl_{+k} denote the POS labels of the kk-th token to the left and right. A first-order POS condition can be l−1l_{-1}; higher-order conditions include combinations such as (l−1,l+1)(l_{-1},l_{+1}) and (l−2,l−1,l+1)(l_{-2},l_{-1},l_{+1}). A special label `[none]' is used when a referenced neighboring token does not exist.

    CATER also uses dependency-tree conditions. A dependency tree is a directed acyclic graph whose vertices are sentence tokens and whose arcs encode head-dependent grammatical relations. For a target word, d1d_1 is the label of its incoming dependency arc, and higher-order dependency conditions recursively use labels along the path toward the root, (d1,d2,…)(d_1,d_2,\ldots). `[none]' is used when the path reaches a root without another parent arc. Combining multiple feature labels creates many possible conditions, allowing a small number of actual watermark rules to be hidden among many statistically suspicious but unused condition-word assignments.

  5. Knowl 5 — Lower bound on poorly sampled conditions

    theoretical result

    Assume an attacker has collected NN tokens belonging to the synonym groups used by CATER. Suppose the watermark condition has order KK and uses a feature set FF with ∣F∣|F| possible labels, so the total number of possible conditions is ∣C∣=∣F∣K|C|=|F|^K. If ∣F∣K>N|F|^K>N and tt conditions each have at most mm observed support samples, where m∈Z>0m\in\mathbb Z_{>0}, then CATER proves

    t≥∣F∣K−Nm+1. t\geq |F|^K-\frac{N}{m+1}.

    A condition has a support sample whenever an observed token is associated with that condition. The result gives a lower bound on how many candidate conditions remain sparsely observed by an attacker. Because ∣F∣K|F|^K grows exponentially with the feature order KK, increasing the order can create a large population of under-sampled conditions even when the attacker knows the possible feature vocabulary and has a fixed query budget.

  6. Knowl 6 — Sparse conditions are intrinsically likely to look like extreme watermarks

    theoretical result

    For a synonym set WW, condition cc, and m∈Z>0m\in\mathbb Z_{>0} independent support samples drawn according to the conditional distribution P(w∣c)P(w\mid c), CATER defines the probability of an extremely imbalanced observation—every sampled word is the same synonym—as

    I(W,c,m)=∑wi∈WP(wi∣c)m.I(W,c,m)=\sum_{w_i\in W}P(w_i\mid c)^m.

    If m′≤mm'\leq m and both m,m′m,m' are positive integers, then

    I(W,c,m′)≥I(W,c,m).I(W,c,m')\geq I(W,c,m).

    Thus sparsely observed conditions are more likely to produce an apparently deterministic or highly imbalanced synonym choice by chance. Combined with the lower bound on the number of sparsely sampled conditions, this means that an attacker inspecting conditional frequencies will encounter many suspicious unused rules, making the comparatively small set of rules actually used by CATER difficult to identify.

  7. Knowl 7 — Benchmark and training configuration for evaluating CATER

    experimental setup

    CATER is evaluated on machine translation and document summarization. The translation benchmark is WMT14 German-to-English with 4.5 million training examples, 3,000 development examples, and 3,003 test examples. The summarization benchmark is CNN/Daily Mail with 287,000 training examples, 13,000 development examples, and 11,000 test examples. The experiments use 32K BPE vocabulary for WMT14 and 16K BPE vocabulary for CNN/Daily Mail.

    The primary victim and imitation models use Transformer-base; summarization additionally uses a 3-layer Transformer. To test pretrained victims, the paper uses mBART for translation and BART for summarization, while imitation models may use (m)BART, Transformer, or ConvS2S. In the basic setup, victim and imitator are trained from the same data, but the imitator receives CATER-watermarked victim outputs instead of ground-truth targets. Synonym sets contain two words, and first-order POS is the default watermark condition. Translation quality is measured with BLEU and BERTScore; summarization quality is measured with ROUGE-1/2/L and BERTScore.

  8. Knowl 8 — CATER detects imitation with small quality degradation

    data/table

    The comparison reported on page 7 evaluates watermark detectability through p-value and generation quality on WMT14 translation and CNN/Daily Mail summarization. Lower p-values indicate stronger detection, while higher BLEU, ROUGE, and BERTScore indicate better generation. CATER produces statistically detectable imitation while keeping quality close to the unwatermarked baseline; unlike the unigram watermark, its detectability does not require a large obvious marginal-frequency shift.

    Could not parse LaTeX table

    The unwatermarked systems have p-values above 10−110^{-1}. CATER with POS gives p-values below 10−710^{-7} on both tasks, while its BLEU, ROUGE, and BERTScore values remain close to the unwatermarked systems. The paper reports that CATER is somewhat less statistically extreme than the earlier synonym and spelling methods, but its conditional design avoids the easily visible overall frequency changes of those methods.

  9. Knowl 9 — Detection survives architectural mismatch and domain mismatch

    data/table

    CATER remains effective when the imitation architecture differs from the victim architecture. The victim uses mBART for WMT14 and BART for CNN/Daily Mail, and first-order POS supplies the watermark. Parenthesized values are results when the imitator is trained on clean rather than watermarked outputs.

    Could not parse LaTeX table

    The watermarked-versus-clean p-value gap is approximately three orders of magnitude or more for both matched and mismatched architectures. Quality decreases by at most 0.30.3 in the reported BLEU or ROUGE-L comparisons. For cross-domain translation imitation, the victim is trained on WMT14 while the imitator uses either IWSLT14 German data with 250,000 sentences or OPUS Law-domain German data with 2.1 million sentences. The resulting p-values are respectively <10−7<10^{-7} for WMT14, <10−5<10^{-5} for IWSLT14, and <10−6<10^{-6} for OPUS Law, showing that CATER's detection signal transfers across these domain changes.

  10. Knowl 10 — Higher-order conditions and adaptive attacks

    empirical result

    Increasing the order of POS conditions from first to third leaves translation BLEU essentially unchanged relative to imitation from clean data, but the p-value increases with condition order, making the ownership claim less statistically confident. The watermarked and clean systems remain distinguishable, so the paper uses first-order POS as its default balance between detection and stealthiness. When an imitator is trained on a mixture of clean human-labeled data and CATER-watermarked machine-labeled data, clean signals overwrite the low-distortion CATER signal unless more than half of the training data is watermarked.

    For an attacker who knows that a watermark exists but does not know the algorithm, the ONION text-watermark-removal method fails to erase CATER and substantially damages imitation quality:

    Could not parse LaTeX table

    For the stronger attacker who knows the watermark dictionary and all feature sets but not the exact assignments, the sparse-entry analysis on page 10 shows that the number of suspicious candidate watermarks grows rapidly with condition order and remains vastly larger than the number actually used. The experiment uses the top 200 words, POS features with ∣F∣=36|F|=36, and compares actual assignments with the candidate count under algorithm leakage. Removing every candidate watermark would require broad modifications that would severely impair imitation quality.

  11. Knowl 11 — Operational and societal limitations

    limitation

    CATER depends on high-quality synonym sets: poor synonyms can cause semantic degradation, while the limited availability of suitable synonyms restricts the candidate vocabulary. Lexical replacement also causes a small but measurable generation-quality loss. The authors report that using the top 200 words and their synonyms still provides a sufficiently large space for stealthy watermarking.

    CATER is a post-hoc ownership-verification mechanism, so it is useful only when the suspected imitation model or API can be queried. If an adversary never publicly releases the imitation service, the owner cannot perform the verification. Finally, because the quality gap between benign and watermarked imitators can be small, an API owner could misuse CATER to accuse an innocent service. The paper therefore recommends a stricter evidentiary threshold, such as requiring a p-value below 10−610^{-6}, before asserting infringement.

Coverage note — The appendix-only evaluations on text simplification and paraphrase generation, detailed training/compute reproducibility settings, and proof derivations were omitted: the theorem statements and the main benchmark, robustness, and limitation results are retained, while derivations are excluded by the requested no-proofs rule.

References

  1. 1.Yonatan Belinkov and Yonatan Bisk. Synthetic and natural noise both break neural machine translation. In International Conference on Learning Representations, 2018.
  2. 2.Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the Ninth Workshop on Statistical Machine Translation, pages 12–58, Baltimore, Maryland, USA, June 2014. Association for Computational Linguistics.
  3. 3.Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico. Report on the 11th iwslt evaluation campaign, iwslt 2014. In Proceedings of the International Workshop on Spoken Language Translation, Hanoi, Vietnam, volume 57, 2014.
  4. 4.Xinyun Chen, Wenxiao Wang, Chris Bender, Yiming Ding, Ruoxi Jia, Bo Li, and Dawn Song. Refit: A unified watermark removal framework for deep learning systems with limited data. In Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security, ASIA CCS ’21, page 321–335, New York, NY, USA, 2021. Association for Computing Machinery.
  5. 5.William Cohen, Vitor Carvalho, and Tom Mitchell. Learning to classify email into “speech acts”. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pages 309–316, 2004.
  6. 6.Jiazhu Dai, Chuanshuai Chen, and Yufeng Li. A backdoor attack against lstm-based text classification systems. IEEE Access, 7:138872–138878, 2019.
  7. 7.Christiane Fellbaum. Wordnet. In Theory and applications of ontology: computer applications, pages 231–243. Springer, 2010.
  8. 8.Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. Convolutional sequence to sequence learning. In International Conference on Machine Learning, pages 1243–1252. PMLR, 2017.
  9. 9.Shangwei Guo, Chunlong Xie, Jiwei Li, Lingjuan Lyu, and Tianwei Zhang. Threats to pre-trained language models: Survey and taxonomy. arXiv preprint arXiv:2202.06862, 2022.
  10. 10.Shangwei Guo, Tianwei Zhang, Han Qiu, Yi Zeng, Tao Xiang, and Yang Liu. Fine-tuning is not enough: A simple yet effective watermark removal attack for dnn models. arXiv preprint arXiv:2009.08697, 2020.
  11. 11.Gurobi Optimization, LLC. Gurobi Optimizer Reference Manual, 2022.
  12. 12.Xuanli He, Lingjuan Lyu, Lichao Sun, and Qiongkai Xu. Model extraction and adversarial transferability, your bert is vulnerable! In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2006–2012, 2021.
  13. 13.Xuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu, and Chenguang Wang. Protecting intellectual property of language generation apis with lexical watermark. In Thirtieth AAAI Conference on Artificial Intelligence, 2022.
  14. 14.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. Teaching machines to read and comprehend. Advances in neural information processing systems, 28, 2015.
  15. 15.Tom Hosking, Hao Tang, and Mirella Lapata. Hierarchical sketch induction for paraphrase generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2489–2501, 2022.
  16. 16.Daniel Jurafsky and James Martin. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, volume 2. 02 2008.
  17. 17.Yoon Kim and Alexander M Rush. Sequence-level knowledge distillation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1317–1327, 2016.
  18. 18.Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondˇrej Bojar, Alexandra Constantin, and Evan Herbst. Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics Companion Volume Proceedings of the Demo and Poster Sessions, pages 177–180, Prague, Czech Republic, June 2007. Association for Computational Linguistics.
  19. 19.Philipp Koehn and Rebecca Knowles. Six challenges for neural machine translation. In Proceedings of the First Workshop on Neural Machine Translation, pages 28–39, Vancouver, August 2017. Association for Computational Linguistics.
  20. 20.Kalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot, and Mohit Iyyer. Thieves on sesame street! model extraction of bert-based apis. In International Conference on Learning Representations, 2020.
  21. 21.Keita Kurita, Paul Michel, and Graham Neubig. Weight poisoning attacks on pretrained models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2793–2806, 2020.
  22. 22.John D Lafferty, Andrew McCallum, and Fernando CN Pereira. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In ICML, 2001.
  23. 23.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, 2020.
  24. 24.Meng Li, Qi Zhong, Leo Yu Zhang, Yajuan Du, Jun Zhang, and Yong Xiangt. Protecting the intellectual property of deep neural networks with watermarking: The frequency domain approach. In 2020 IEEE 19th International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), pages 402–409. IEEE, 2020.
  25. 25.Jian Han Lim, Chee Seng Chan, Kam Woh Ng, Lixin Fan, and Qiang Yang. Protect, show, attend and tell: Empowering image captioning models with ownership protection. Pattern Recognition, 122:108285, 2022.
  26. 26.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81, 2004.
  27. 27.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742, 2020.
  28. 28.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013.
  29. 29.Mathias Müller, Annette Rios, and Rico Sennrich. Domain robustness in neural machine translation. In Proceedings of the 14th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track), pages 151–164, Virtual, October 2020. Association for Machine Translation in the Americas.
  30. 30.Tribhuvanesh Orekondy, Bernt Schiele, and Mario Fritz. Knockoff nets: Stealing functionality of black-box models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4954–4963, 2019.
  31. 31.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53, 2019.
  32. 32.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318, 2002.
  33. 33.Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. Onion: A simple and effective defense against textual backdoor attacks. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9558–9566, 2021.
  34. 34.Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. Stanza: A Python natural language processing toolkit for many human languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, 2020.
  35. 35.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  36. 36.John A Rice. Mathematical statistics and data analysis. Cengage Learning, 2006.
  37. 37.Sunita Sarawagi and William W Cohen. Semi-markov conditional random fields for information extraction. In L. Saul, Y. Weiss, and L. Bottou, editors, Advances in Neural Information Processing Systems, volume 17. MIT Press, 2004.
  38. 38.Abigail See, Peter J Liu, and Christopher D Manning. Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1073–1083, 2017.
  39. 39.Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, 2016.
  40. 40.Amit Singhal. Microsoft’s bing uses google search results—and denies it, 2011.
  41. 41.Xiaofei Sun, Jiwei Li, Xiaoya Li, Ziyao Wang, Tianwei Zhang, Han Qiu, Fei Wu, and Chun Fan. A general framework for defending against backdoor attacks via influence graph. arXiv preprint arXiv:2111.14309, 2021.
  42. 42.Sebastian Szyller, Buse Gul Atli, Samuel Marchal, and N Asokan. Dawn: Dynamic adversarial watermarking of neural networks. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4417–4425, 2021.
  43. 43.Jörg Tiedemann. Parallel data, tools and interfaces in opus. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), pages 2214–2218, 2012.
  44. 44.Florian Tramèr, Fan Zhang, Ari Juels, Michael K Reiter, and Thomas Ristenpart. Stealing machine learning models via prediction apis. In 25th {USENIX} Security Symposium ({USENIX} Security 16), pages 601–618, 2016.
  45. 45.Yusuke Uchida, Yuki Nagai, Shigeyuki Sakazawa, and Shin’ichi Satoh. Embedding watermarks into deep neural networks. In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval, pages 269–277, 2017.
  46. 46.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  47. 47.Ashish Venugopal, Jakob Uszkoreit, David Talbot, Franz Och, and Juri Ganitkevitch. Watermarking the outputs of structured prediction with an application in statistical machine translation. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing, pages 1363–1372, Edinburgh, Scotland, UK., July 2011. Association for Computational Linguistics.
  48. 48.Eric Wallace, Mitchell Stern, and Dawn Song. Imitation attacks and defenses for black-box machine translation systems. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5531–5546, 2020.
  49. 49.Jun Wang, Chang Xu, Francisco Guzmán, Ahmed El-Kishky, Yuqing Tang, Benjamin Rubinstein, and Trevor Cohn. Putting words into the system’s mouth: A targeted attack on neural machine translation using monolingual data poisoning. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 1463–1473, 2021.
  50. 50.Chang Xu, Jun Wang, Yuqing Tang, Francisco Guzmán, Benjamin I. P. Rubinstein, and Trevor Cohn. A targeted attack on black-box neural machine translation with parallel data poisoning. In Proceedings of the Web Conference 2021, pages 3638–3650, 2021.
  51. 51.Qiongkai Xu, Xuanli He, Lingjuan Lyu, Lizhen Qu, and Gholamreza Haffari. Beyond model extraction: Imitation attack for black-box nlp apis. arXiv preprint arXiv:2108.13873, 2021.
  52. 52.Qiongkai Xu and Hai Zhao. Using deep linguistic features for finding deceptive opinion spam. In Proceedings of COLING 2012: Posters, pages 1341–1350, 2012.
  53. 53.Yifan Yan, Xudong Pan, Yining Wang, Mi Zhang, and Min Yang. " and then there were none": Cracking white-box dnn watermarks via invariant neuron transforms. arXiv preprint arXiv:2205.00199, 2022.
  54. 54.Yi Zeng, Won Park, Z Morley Mao, and Ruoxi Jia. Rethinking the backdoor attacks’ triggers: A frequency perspective. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16473–16481, 2021.
  55. 55.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations, 2019.
  56. 56.Xingxing Zhang and Mirella Lapata. Sentence simplification with deep reinforcement learning. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 584–594, Copenhagen, Denmark, September 2017. Association for Computational Linguistics.
  57. 57.Zhiyuan Zhang, Lingjuan Lyu, Weiqiang Wang, Lichao Sun, and Xu Sun. How to inject backdoors with better consistency: Logit anchoring on clean data. In International Conference on Learning Representations, 2021.

Citation

MLA
He, X., et al. “CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 5431–45, https://proceedings.neurips.cc/paper_files/paper/2022/file/2433fec2144ccf5fea1c9c5ebdbc3924-Paper-Conference.pdf.
APA
He, X., Xu, Q., Zeng, Y., Lyu, L., Wu, F., Li, J., & Jia, R. (2022). CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks. Advances in Neural Information Processing Systems, 35, 5431–5445. https://proceedings.neurips.cc/paper_files/paper/2022/file/2433fec2144ccf5fea1c9c5ebdbc3924-Paper-Conference.pdf
Chicago
He, X., Q. Xu, Y. Zeng, et al. 2022. “CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks”. Advances in Neural Information Processing Systems 35: 5431–45. https://proceedings.neurips.cc/paper_files/paper/2022/file/2433fec2144ccf5fea1c9c5ebdbc3924-Paper-Conference.pdf.
Harvard
He, X. et al. (2022) “CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 5431–5445. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/2433fec2144ccf5fea1c9c5ebdbc3924-Paper-Conference.pdf.
Vancouver
1. He X, Xu Q, Zeng Y, Lyu L, Wu F, Li J, Jia R (2022) CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 5431–5445

BibTeX

@inproceedings{he2022cater,
  title = {CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks},
  author = {He, Xuanli and Xu, Qiongkai and Zeng, Yi and Lyu, Lingjuan and Wu, Fangzhao and Li, Jiwei and Jia, Ruoxi},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {5431-5445},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/2433fec2144ccf5fea1c9c5ebdbc3924-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission