A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models

Yihan WuZhengmian HuJunfeng GuoHongyang ZhangHeng Huang

article2024ICML59 citations

Proposes DiPmark, a language model watermarking framework that maintains the original text generation distribution without quality loss while remaining provably resilient to text modifications and detectable without requiring model API or prompt access.

Listen

As artificial intelligence systems generate text that is increasingly indistinguishable from human writing, organizations face growing risks around academic integrity, intellectual property protection, and digital misinformation. Embedding covert identifiers, or watermarks, into model-generated text offers a viable path to track provenance and detect artificial origin. However, existing watermarking methods introduce significant trade-offs: they either distort the natural output distribution, degrading output quality; require extensive computing power and thousands of resampling steps; or require direct access to original prompts and internal model programming interfaces during detection.

The article demonstrates and evaluates DiPmark, a novel watermarking framework designed to embed reliable identifiers into language model outputs while strictly preserving original text quality and generation distributions. The primary objective is to deliver a watermarking mechanism that simultaneously guarantees distribution preservation, operational accessibility without requiring model access during detection, and provable resilience against deliberate text tampering.

To establish these properties, the authors designed a distribution-preserving reweighting strategy that adjusts token probabilities during generation using a secret key and a permutation hash function. The approach replaces traditional normal approximation tests with an exact statistical detection bound based on green token ratios. The evaluation tested DiPmark across standard tasks, including machine translation using Multilingual BART on the WMT 2014 dataset, text summarization using BART on the CNN-DailyMail corpus, and open-ended text generation using the 7-billion-parameter LLaMA-2 model, alongside an industry case study modifying top-token probabilities from GPT-4.

The findings confirm that DiPmark maintains original model text quality across all tested parameters, showing virtually identical BLEU, BERTScore, and perplexity metrics to unwatermarked baselines, whereas traditional baseline methods caused noticeable quality degradation. In terms of accessibility, DiPmark detected watermarks across 1,000 text sequences generated by LLaMA-2 in 90 seconds without access to prompts or internal model parameters, operating at least four times faster than comparable distribution-preserving detectors. For robustness, the method maintained strong detection performance under 20% to 30% text modifications, achieving area under the curve scores above 0.80 under random alterations and paraphrasing attacks, while accurately certifying error bounds. In the real-world GPT-4 test, the detector successfully identified 97 out of 100 watermarked sequences at a strict false positive rate below 1%.

These results indicate that enterprise and industry model providers can implement reliable provenance tracking without compromising system output quality or exposing proprietary model internals during verification. By mathematically guaranteeing error rates, the framework significantly reduces compliance risks and prevents the wrongful classification of human-written text that occurs with standard approximation methods. Moving forward, the article suggests that organizations adopting language model watermarking should consider distribution-preserving techniques, while future engineering should explore incorporating multiple detector configurations and heavier green-token weighting to further enhance detection sensitivity in low-entropy contexts.

Cover for A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models

Abstract

Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models. A challenge in the domain lies in preserving the distribution of original generated content after watermarking. Our research extends and improves upon existing watermarking framework, placing emphasis on the importance of a Distribution-Preserving (DiP) watermark. Contrary to the current strategies, our proposed DiPmark simultaneously preserves the original token distribution during watermarking (distribution-preserving), is detectable without access to the language model API and prompts (accessible), and is provably robust to moderate changes of tokens (resilient). DiPmark operates by selecting a random set of tokens prior to the generation of a word, then modifying the token distribution through a distribution-preserving reweight function to enhance the probability of these selected tokens during the sampling process. Extensive empirical evaluation on various language models and tasks demonstrates our approach’s distribution-preserving property, accessibility, and resilience, making it a effective solution for watermarking tasks that demand impeccable quality preservation. Code is available at¹.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminary
  • 4. DiPmark
  • 5. DiPmark Detection
  • 6. DiPmark is Provably Resilient Against Text Modification
  • 7. Experiments
  • 7.1. Distribution-preserving Property
  • 7.2. Accessibility
  • 7.3. Resilience and provable resilience
  • 7.4. Ablation study: watermark detectability
  • 7.5. Case study: watermarking GPT-4 by DiPmark
  • 8. Conclusion
  • Impact Statement
  • Acknowledgement
  • References
  • A. Future Work
  • B. Related Work
  • C. Missing Proofs
  • C.1. Proof of Theorem 4.3
  • C.2. Proof of Theorem 5.2
  • C.3. Proof of Theorem 6.2 and discussion
  • D. Comparison of the test statistic
  • E. Detailed Experiment Setup
  • F. Additional Experiments
  • F.1. Distribution-preserving
  • F.2. Detectability comparison
  • F.3. Resilience
  • G. An alternative detector for DiPmark.
  • H. Examples of the watermarked text

Knowls

  1. Knowl 1 — DiPmark’s distribution-preserving reweighting rule

    model/method

    Let VV be a vocabulary of NN tokens, let pM(t∣c)p_M(t\mid c) be the language model’s probability for token tt in context cc, and let θ=(t1,…,tN)\theta=(t_1,\ldots,t_N) be a permutation of VV. For i=0,…,Ni=0,\ldots,N, define Ci=∑j=1ipM(tj∣c)C_i=\sum_{j=1}^{i}p_M(t_j\mid c), with C0=0C_0=0. For β∈[0,1)\beta\in[0,1), define Fβ(i)=max⁡(Ci−β,0)/(1−β)F^\beta(i)=\max(C_i-\beta,0)/(1-\beta) and PWβ(ti∣c,θ)=Fβ(i)−Fβ(i−1)P_W^\beta(t_i\mid c,\theta)=F^\beta(i)-F^\beta(i-1). DiPmark uses the convex combination

    pW(ti∣c,θ)=(1−α)PWα(ti∣c,θ)+αPW1−α(ti∣c,θ),p_W(t_i\mid c,\theta)=(1-\alpha)P_W^\alpha(t_i\mid c,\theta)+\alpha P_W^{1-\alpha}(t_i\mid c,\theta),

    where the practical parameter is 0<α<10<\alpha<1. The permutation orders the token probabilities into a cumulative interval: this construction removes mass from its low-probability portion and redistributes it across the higher portion, while each component remains a probability distribution. The red/green separator γ\gamma determines which suffix of the permutation is called green; it does not enter the reweighting formula. The reweighting favors tokens toward that suffix, increasing the green-token signal used for detection.

  2. Knowl 2 — DiPmark preserves the language-model distribution in expectation

    theoretical result

    For a fixed context cc, let the cipher θ\theta be a uniformly random permutation of the vocabulary, and let pW(⋅∣c,θ)p_W(\cdot\mid c,\theta) be DiPmark’s reweighted distribution. For every token tt,

    Eθ[pW(t∣c,θ)]=pM(t∣c),\mathbb{E}_{\theta}[p_W(t\mid c,\theta)]=p_M(t\mid c),

    where pMp_M is the original language-model distribution. Thus the watermark changes the distribution conditional on a particular cipher but preserves the original next-token distribution after averaging over ciphers. If the ciphers used at successive generation steps are independent, the complete generated sequence also has the original language-model distribution after averaging over those ciphers. DiPmark’s key-to-permutation mechanism is intended to produce uniform, independent ciphers for distinct secret-key/texture-key pairs; the generator avoids reusing a texture key for a watermarked step by sampling from the original distribution when that texture key has already occurred.

  3. Knowl 3 — DiPmark generation procedure

    algorithm

    The generator takes a language-model distribution at each step, a secret key kk, reweight parameter α\alpha, prompt, requested output length nn, texture-key window length aa, and a keyed permutation function hh. The output is a generated token sequence. The texture key is formed from the most recent up-to-aa tokens of the preceding context.

    Input: Secret key k, prompt, output length n, window length a, parameter alpha, and keyed permutation function h
    Initialize an empty history of texture keys
    For each generation position i from 1 to n:
        Obtain the original language-model distribution p_M for the next token
        Form texture key s from the most recent up-to-a preceding tokens
        If s is already in the texture-key history:
            Sample the next token from p_M
        Else:
            Add s to the history
            Set permutation theta = h(k, s)
            Form the DiPmark distribution p_W using theta and alpha
            Sample the next token from p_W
    Return the generated sequence

    The paper’s generator uses the unwatermarked distribution when a texture key repeats; for new texture keys it samples from the DiPmark distribution. Its stated cipher construction maps the secret key and texture key to a permutation, with the intended uniformity and independence properties for distinct pairs. No computational-complexity bound is specified.

  4. Knowl 4 — Single-pass, prompt-free DiPmark detection

    model/method

    Given a candidate token sequence x1:nx_{1:n}, secret key kk, vocabulary size NN, keyed permutation function hh, window length aa, green-list separator γ∈[0,1]\gamma\in[0,1], and threshold zz, the detector processes each token using only the candidate text and key. At position ii, it derives a texture key from the preceding context, computes θi=h(k,si)\theta_i=h(k,s_i), and defines the green list as the final N−⌈γN⌉N-\lceil\gamma N\rceil tokens of that permutation. Let LG(γ)L_G(\gamma) count candidate tokens classified as green. The score and decision rule are

    Φ(γ,x1:n)=LG(γ)n−(1−γ),detect watermark if Φ(γ,x1:n)>z.\Phi(\gamma,x_{1:n})=\frac{L_G(\gamma)}{n}-(1-\gamma),\qquad \text{detect watermark if }\Phi(\gamma,x_{1:n})>z.

    The detector requires neither the generation prompt nor access to the language-model API or its logits. It makes one pass over the text, deriving the per-position texture keys and green lists and counting green tokens. In the paper’s LLaMA-2 chat 7B timing comparison on an NVIDIA A6000, DiPmark detection took 0.3 seconds for one sequence and 90 seconds for 1,000 sequences; the comparison methods took, respectively, 0.3 seconds and 92 seconds for the Soft watermark, 80 seconds and 12 hours for Kuditipudi et al.’s method, and 3.4 seconds and 412 seconds for Hu et al.’s method. Only Hu et al.’s method in that comparison required language-model and prompt access.

  5. Knowl 5 — Finite-sample false-positive calibration

    theoretical result

    Under the paper’s null model for unwatermarked text, the green-token count LG(γ)L_G(\gamma) has a binomial distribution with nn trials and green probability 1−γ1-\gamma. Consequently, for Φ=LG(γ)/n−(1−γ)\Phi=L_G(\gamma)/n-(1-\gamma), the stated upper-tail bound is

    Pr⁡(Φ≥t)≤exp⁡ ⁣[−n DKL(1−γ+t ∥ 1−γ)],\Pr(\Phi\ge t)\le \exp\!\left[-n\,D_{\mathrm{KL}}(1-\gamma+t\,\|\,1-\gamma)\right],

    where DKL(p∥q)=plog⁡(p/q)+(1−p)log⁡((1−p)/(1−q))D_{\mathrm{KL}}(p\|q)=p\log(p/q)+(1-p)\log((1-p)/(1-q)) is binary Kullback–Leibler divergence. The bound applies to feasible tail values for which its probability arguments lie in [0,1][0,1]. For γ=0.5\gamma=0.5, the paper gives z=1.517/nz=1.517/\sqrt n as a threshold for a false-positive guarantee below 1%.

    In a comparison on 500 non-watermarked sentences, the paper reports these false-positive counts and rates at nominal thresholds: at p<0.10p<0.10, the z-test gave 56/500 (11.2%) and the DiPmark statistic 13/500 (2.6%); at p<0.05p<0.05, the z-test gave 34/500 (6.8%) and DiPmark 10/500 (2%); at p<0.01p<0.01, the z-test gave 12/500 (2.4%) and DiPmark 4/500 (0.5%), using the reported rates. The results illustrate the paper’s claim that the normal-approximation z-test can exceed its nominal false-positive rate, whereas the DiPmark statistic is based on a finite-sample concentration bound.

  6. Knowl 6 — Certified resilience to arbitrary token modifications

    theoretical result

    Let x1:nx_{1:n} be a watermarked sequence, let LG(γ)L_G(\gamma) count its green tokens for separator γ\gamma, and let Φ=LG(γ)/n−(1−γ)\Phi=L_G(\gamma)/n-(1-\gamma) be its detection score. The detector accepts scores above threshold zz. If the texture key uses the preceding aa tokens, changing one token can reduce the green count by at most a+1a+1: the changed token’s own contribution and the contributions of up to aa following tokens whose keyed lists depend on it. For an attack that modifies at most an ϵ\epsilon fraction of tokens, the paper gives the certified radius

    ϵ0=Φ(γ,x1:n)−z2+a−γ+z.\epsilon_0=\frac{\Phi(\gamma,x_{1:n})-z}{2+a-\gamma+z}.

    When this value is positive, every arbitrary text modification within budget ϵ≤ϵ0\epsilon\le\epsilon_0 is guaranteed to leave the sequence detectable under the stated detector and threshold. The bound allows changes to sequence length, including insertions; it is not restricted to a particular attack strategy.

  7. Knowl 7 — Text quality remains close to the unwatermarked baseline

    empirical result

    The paper evaluated distribution-preserving behavior on English-to-Romanian machine translation with MBart and WMT’14 En–Ro, and text summarization with BART-large and CNN-DM. The reported machine-translation measures are BERTScore-F1 multiplied by 100 and BLEU; summarization measures are BERTScore-F1 multiplied by 100 and perplexity. Results below are mean ± reported variation.

    DiPmark at α=0.45\alpha=0.45 produced machine-translation BERTScore 56.2±0.3 and BLEU 21.9±0.3, and summarization BERTScore 32.69±0.08 and perplexity 5.024±0.018. At α=0.5\alpha=0.5, the respective values were 56.2±0.3, 21.8±0.3, 32.72±0.08, and 5.014±0.018. The no-watermark baseline was 55.9±0.3, 21.8±0.3, 32.73±0.08, and 5.021±0.018. In contrast, the Soft watermark at δ=1.5\delta=1.5 gave 55.0±0.3, 20.4±0.3, 32.09±0.08, and 5.660±0.021; at δ=2.0\delta=2.0, it gave 53.9±0.3, 19.4±0.3, 31.46±0.08, and 6.241±0.023. Across the tested DiPmark settings, the quality metrics stayed close to the baseline, while stronger Soft watermarking degraded translation scores and increased summarization perplexity.

  8. Knowl 8 — Detection error rates on LLaMA-2 text

    empirical result

    For text generation with LLaMA-2, the paper evaluated roughly 500 watermarked and 500 unwatermarked sequences of length 260±5260\pm5, using γ=0.5\gamma=0.5. The table reports false-positive rate (FPR), true-negative rate (TNR), true-positive rate (TPR), false-negative rate (FNR), and perplexity (PPL) for DiPmark and Soft watermarking. The two tested thresholds correspond to nominal FPR levels of at most 10% and 1%.

    Detector and settingFPRTNRTPRFNRPPL
    Soft, δ=1.0\delta=1.0, z=1.073/nz=1.073/\sqrt n0.05450.94550.89190.26863.38±0.06
    Soft, δ=1.5\delta=1.5, z=1.073/nz=1.073/\sqrt n0.05450.94550.99610.07963.56±0.06
    Soft, δ=2.0\delta=2.0, z=1.073/nz=1.073/\sqrt n0.05450.94551.00000.00003.92±0.07
    DiPmark, α=0.45\alpha=0.45, z=1.073/nz=1.073/\sqrt n0.05450.94551.00000.00003.14±0.06
    DiPmark, α=0.5\alpha=0.5, z=1.073/nz=1.073/\sqrt n0.05450.94551.00000.00003.17±0.05
    Soft, δ=1.0\delta=1.0, z=1.517/nz=1.517/\sqrt n0.00800.99200.82550.17453.38±0.06
    Soft, δ=1.5\delta=1.5, z=1.517/nz=1.517/\sqrt n0.00800.99200.97240.02763.56±0.06
    Soft, δ=2.0\delta=2.0, z=1.517/nz=1.517/\sqrt n0.00800.99200.99810.00193.92±0.07
    DiPmark, α=0.45\alpha=0.45, z=1.517/nz=1.517/\sqrt n0.00800.99200.97940.02063.14±0.05
    DiPmark, α=0.5\alpha=0.5, z=1.517/nz=1.517/\sqrt n0.00800.99200.98270.01733.17±0.05

    At the stricter threshold, DiPmark’s TPR was 0.9794 for α=0.45\alpha=0.45 and 0.9827 for α=0.5\alpha=0.5, with reported FPR 0.0080. The reported perplexities for these settings were lower than for the tested Soft-watermark settings.

  9. Knowl 9 — Detection under random edits and paraphrasing

    empirical result

    The paper compared watermark detection AUC on about 500 watermarked and 500 unwatermarked LLaMA-2 sequences, each of length 260±5260\pm5. DiPmark used α=0.45\alpha=0.45 and a one-token texture window (a=1a=1). The attacks changed a fraction ϵ\epsilon of tokens; the values below are AUC, with higher values indicating better separation in this evaluation.

    MethodRandom edits ϵ=0.0\epsilon=0.00.10.20.3Paraphrase ϵ=0.0\epsilon=0.00.10.20.3
    Soft watermark0.99900.98830.95210.80330.99900.98940.94690.8157
    Kuditipudi et al.0.99510.94610.89790.78150.99510.95290.90130.7711
    Hu et al.0.99360.92970.83910.75740.99360.93680.83250.7661
    DiPmark, α=0.45\alpha=0.450.99900.98590.95150.80600.99900.98710.95030.8216

    DiPmark’s AUC remained 0.9515 under 20% random edits and 0.9503 under the 20% paraphrasing condition. At 30% modification, its AUC was 0.8060 for random edits and 0.8216 for paraphrasing, comparable to the Soft watermark’s 0.8033 and 0.8157 under the corresponding conditions.

  10. Knowl 10 — GPT-4 case study with top-five token probabilities

    experimental setup

    The paper adapted DiPmark to GPT-4’s released top-five next-token log probabilities: for each generation step it treated all tokens outside the reported top five as having probability zero, reweighted the available distribution, sampled the next token, appended it to the prompt, and repeated. Using GPT-4 model gpt-4-0613, α=0.45\alpha=0.45, 100 fiction-writing prompts, and a cap of 200 generated tokens, the detector’s threshold for a false-positive rate below 1% was more than 122 green tokens. It classified 97 of the 100 generated sequences as watermarked. This experiment demonstrates the method in a setting with restricted token-probability access, specifically the top-five probabilities.

Coverage note — The appendix’s alternative detector that searches across multiple green-list separators, along with its multiple-separator concentration bound, is omitted because the main results use a fixed separator. Detailed additional ROUGE, BERTScore precision/recall, and supplementary task-specific plots are also omitted; the central quality comparison and the principal detection and resilience results are retained.

References

  1. 1.Aaronson, S. My AI safety lecture for UT effective altruism,. 2022. URL https://scottaaronson.blog/?p=6823.
  2. 2.Abdelnabi, S. and Fritz, M. Adversarial watermarking transformer: Towards tracing text provenance with data hiding. In 2021 IEEE Symposium on Security and Privacy (SP), pp. 121–140. IEEE, 2021.
  3. 3.Chakraborty, S., Bedi, A. S., Zhu, S., An, B., Manocha, D., and Huang, F. On the possibilities of AI-generated text detection. arXiv preprint arXiv:2304.04736, 2023.
  4. 4.Chen, R., Wu, Y., Chen, L., Liu, G., He, Q., Xiong, T., Liu, C., Guo, J., and Huang, H. Your vision-language model itself is a strong filter: Towards high-quality instruction tuning with data selection. arXiv preprint arXiv:2402.12501, 2024.
  5. 5.Christ, M., Gunn, S., and Zamir, O. Undetectable watermarks for language models. arXiv preprint arXiv:2306.09194, 2023.
  6. 6.Dedic, N., Itkis, G., Reyzin, L., and Russell, S. Upper and lower bounds on black-box steganography. Journal of Cryptology, 22:365–394, 2009.
  7. 7.Gambini, M., Fagni, T., Falchi, F., and Tesconi, M. On pushing DeepFake Tweet detection capabilities to the limits. In Proceedings of the 14th ACM Web Science Conference 2022, pp. 154–163, 2022.
  8. 8.Google. Palm-2-llm. https://blog.google/technology/ai/google-palm-2-ai-large-language-model/, 2023.
  9. 9.Hermann, K. M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P. Teaching machines to read and comprehend. Advances in neural information processing systems, 28, 2015.
  10. 10.Hong, Z., Wang, Z., Shen, L., Yao, Y., Huang, Z., Chen, S., Yang, C., Gong, M., and Liu, T. Improving non-transferable representation learning by harnessing content and style. In The Twelfth International Conference on Learning Representations, 2024.
  11. 11.Hopper, N. J., Langford, J., and Von Ahn, L. Provably secure steganography. In Advances in Cryptology—CRYPTO 2002: 22nd Annual International Cryptology Conference Santa Barbara, California, USA, August 18–22, 2002 Proceedings 22, pp. 77–92. Springer, 2002.
  12. 12.Hu, Z., Chen, L., Wu, X., Wu, Y., Zhang, H., and Huang, H. Unbiased watermark for large language models. arXiv preprint arXiv:2310.10669, 2023a.
  13. 13.Hu, Z., Shen, L., Wang, Z., Wu, B., Yuan, C., and Tao, D. Learning to learn from APIs: Black-box data-free meta-learning. In Proceedings of the 40th International Conference on Machine Learning, pp. 13610–13627. PMLR, 2023b.
  14. 14.Kaptchuk, G., Jois, T. M., Green, M., and Rubin, A. D. Meteor: Cryptographically secure steganography for realistic distributions. In Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security, pp. 1529–1548, 2021.
  15. 15.Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T. A watermark for large language models. arXiv preprint arXiv:2301.10226, 2023.
  16. 16.Kirchner, J. H., Ahmad, L., Aaronson, S., and Leike, J. New AI classifier for indicating AI-written text. OpenAI, 2023.
  17. 17.Krishna, K., Song, Y., Karpinska, M., Wieting, J., and Iyyer, M. Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense. arXiv preprint arXiv:2303.13408, 2023.
  18. 18.Kuditipudi, R., Thickstun, J., Hashimoto, T., and Liang, P. Robust distortion-free watermarks for language models. arXiv preprint arXiv:2307.15593, 2023.
  19. 19.Lin, C.-Y. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81, 2004.
  20. 20.Liu, Y., Gu, J., Goyal, N., Li, X., Edunov, S., Ghazvininejad, M., Lewis, M., and Zettlemoyer, L. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742, 2020.
  21. 21.Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., and Finn, C. Detectgpt: Zero-shot machine-generated text detection using probability curvature. arXiv preprint arXiv:2301.11305, 2023.
  22. 22.Munyer, T. and Zhong, X. Deeptextmark: Deep learning based text watermarking for detection of large language model generated text. arXiv preprint arXiv:2305.05773, 2023.
  23. 23.OpenAI, R. Gpt-4 technical report. arXiv, pp. 2303–08774, 2023.
  24. 24.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318, 2002.
  25. 25.Qiang, J., Zhu, S., Li, Y., Zhu, Y., Yuan, Y., and Wu, X. Natural language watermarking via paraphraser-based lexical substitution. Artificial Intelligence, 317:103859, 2023.
  26. 26.Tay, Y., Bahri, D., Zheng, C., Brunk, C., Metzler, D., and Tomkins, A. Reverse engineering configurations of neural text generation models. arXiv preprint arXiv:2004.06201, 2020.
  27. 27.Tian, E. GPTzero update v1. https://gptzero.substack.com/p/gptzero-update-v1, 2023.
  28. 28.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  29. 29.Wang, Z., Shen, L., Duan, T., Suo, Q., Fang, L., Liu, W., and Gao, M. Distributionally robust memory evolution with generalized divergence for continual learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023a.
  30. 30.Wang, Z., Shen, L., Liu, T., Duan, T., Zhu, Y., Zhan, D., Doermann, D., and Gao, M. Defending against data-free model extraction by distributionally robust defensive training. In Thirty-seventh Conference on Neural Information Processing Systems, 2023b.
  31. 31.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.
  32. 32.Wu, Y., Zhang, H., and Huang, H. Retrievalguard: Provably robust 1-nearest neighbor image retrieval. In International Conference on Machine Learning, pp. 24266–24279. PMLR, 2022.
  33. 33.Wu, Y., Huang, H., and Zhang, H. A law of robustness beyond isoperimetry. In International Conference on Machine Learning, pp. 37439–37455. PMLR, 2023.
  34. 34.Yoo, K., Ahn, W., Jang, J., and Kwak, N. Robust natural language watermarking through invariant features. arXiv preprint arXiv:2305.01904, 2023.
  35. 35.Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., and Choi, Y. Defending against neural fake news. Advances in neural information processing systems, 32, 2019.
  36. 36.Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019.
  37. 37.Zhao, X., Ananth, P., Li, L., and Wang, Y.-X. Provable robust watermarking for ai-generated text. arXiv preprint arXiv:2306.17439, 2023.

Citation

MLA
Wu, Y., et al. “A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2310.07710v2.
APA
Wu, Y., Hu, Z., Guo, J., Zhang, H., & Huang, H. (2023). A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models. arXiv. http://arxiv.org/abs/2310.07710v2
Chicago
Wu, Y., Z. Hu, J. Guo, H. Zhang, and H. Huang. 2023. “A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models”. arXiv. http://arxiv.org/abs/2310.07710v2.
Harvard
Wu, Y. et al. (2023) “A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.07710v2.
Vancouver
1. Wu Y, Hu Z, Guo J, Zhang H, Huang H (2023) A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models. arXiv

BibTeX

@article{wu2023resilient,
  title = {A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models},
  author = {Wu, Yihan and Hu, Zhengmian and Guo, Junfeng and Zhang, Hongyang and Huang, Heng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.07710v2},
  eprint = {2310.07710}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/