Advancing Beyond Identification: Multi-bit Watermark for Large Language Models

KiYoon YooWonhyuk AhnNojun Kwak

article2024NAACL75 citations

Proposes a multi-bit watermarking framework that embeds traceable user metadata into large language model outputs by allocating tokens to message sub-units, enabling resilient extraction of long messages without added latency or model fine-tuning.

Listen

As large language models become widely accessible, their potential exploitation for automated malicious activities—such as spreading disinformation and propaganda—presents a growing threat to public trust and digital safety. Existing defense mechanisms primarily focus on zero-bit watermarking, which only identifies whether a text is machine-generated. However, mitigating harmful misuses often requires holding specific bad actors accountable. The article addresses this challenge by evaluating how service providers can embed traceable multi-bit metadata (such as user identifiers or timestamps) directly into generated text to enable provenance tracking without storing invasive user logs.

The main objective of the article is to demonstrate the viability and effectiveness of Multi-bit Watermark via Position Allocation (MPAC), a method designed to embed and extract multi-bit information during text generation while preserving generation speed, text quality, and zero-bit detection capabilities. To evaluate this approach, the authors conducted extensive text-generation experiments using standard benchmarks, primarily testing the LLaMA-2-7B model alongside other architectures such as Mistral-7B and OPT-1.3B. The evaluation examined multi-bit extraction accuracy, text degradation, decoding latency, and resilience against common evasion techniques, including copy-paste mixing and automated paraphrasing.

The findings show that MPAC substantially improves watermark robustness compared to prior multi-bit methods. In high-noise scenarios involving copy-paste text tampering, MPAC outperformed runner-up baselines by more than 20% in extraction accuracy for 16-bit and 24-bit messages. In addition, allocating tokens independently across message positions eliminates encoding latency overhead, avoiding the severe generation slowdowns seen in existing cyclic-shift techniques. The analysis also showed that applying a list-decoding strategy—which outputs candidate messages ranked by statistical confidence—boosts extraction performance significantly, raising bit accuracy to roughly 90% across varied token lengths even after aggressive paraphrasing attacks.

These results demonstrate that model providers can embed traceable user-level metadata into open-ended language model outputs at minimal computational cost and without sacrificing generation quality. By moving beyond simple machine-text detection, platforms can systematically identify malicious source accounts, enforce terms of service, and collaborate with relevant authorities. However, the evaluation reveals an inherent trade-off: embedding longer multi-bit messages slightly reduces the model's zero-bit detection performance because aggregating maximum candidate scores over more positions increases false positive rates on unwatermarked text.

For practical deployment, organizations should consider implementing position-allocated multi-bit watermarks combined with confidence-ranked list decoding and error-correction codes to trace provenance efficiently. Before wide real-world adoption, practitioners should calibrate watermark parameters against specific domain needs, as texts with inherently low diversity (such as repetitive question answering or code generation) reduce watermarking capacity. Further research is recommended to optimize statistical detection metrics for long messages and explore dynamic bias adjustments that preserve output quality across diverse text distributions.

No sufficiently relevant recommendations were found.

Cover for Advancing Beyond Identification: Multi-bit Watermark for Large Language Models

Abstract

We show the viability of tackling misuses of large language models beyond the identification of machine-generated text. While existing zero-bit watermark methods focus on detection only, some malicious misuses demand tracing the adversary user for counteracting them. To address this, we propose Multi-bit Watermark via Position Allocation, embedding traceable multi-bit information during language model generation. Through allocating tokens onto different parts of the messages, we embed longer messages in high corruption settings without added latency. By independently embedding sub-units of messages, the proposed method outperforms the existing works in terms of robustness and latency. Leveraging the benefits of zero-bit watermarking (Kirchenbauer et al., 2023a), our method enables robust extraction of the watermark without any model access, embedding and extraction of long messages (≥ 32-bit) without finetuning, and maintaining text quality, while allowing zero-bit detection all at the same time. Code is released here: https://github.com/bangawayoo/mb-lm-watermarking.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Method
  • 3.1 Zero-bit Watermarking
  • 3.2 MPAC: Extending to Multi-bit Watermark
  • 3.3 Detecting Machine Text
  • 3.4 Comparison to Other Works
  • 3.5 Techniques for Practical Use
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Results
  • 5 Discussions
  • 6 Conclusion
  • References
  • A Appendix
  • A.1 Decoding Algorithm
  • A.2 Analysis on Watermark Detection
  • A.2.1 Watermark Detection
  • A.2.2 Approximating Max Multinomial Cell Distribution
  • A.3 List Decoding and Other Techniques
  • A.3.1 Results
  • A.3.2 Error Correction Codes
  • A.3.3 Message Correction with Feedback
  • A.4 Extending MPAC to other methods
  • A.5 Implementation, Hardware, Code Details
  • A.6 Metrics: Bit Accuracy, Text Quality
  • A.7 Discussion on the Hashing Scheme
  • A.8 More on Robustness: Other Attacks, Detection
  • A.9 Ablations on Datasets and Model Sizes
  • A.10 Tabular Results
  • A.11 Generation Samples

Knowls

  1. Knowl 1 — MPAC allocates generated tokens to independent message positions

    model/method

    Multi-bit Watermark via Position Allocation (MPAC) embeds a message by assigning each generated token pseudo-randomly to one position of the message, then biasing generation toward tokens associated with the value stored at that position. At token step tt, the method hashes the recent context Xt−h:t−1X_{t-h:t-1} with a shared pseudo-random function ff to obtain a seed ss. Using ss, it samples a message position pp and permutes the vocabulary. The message digit m[p]m[p] selects a vocabulary subset, and a bias \delta is added to the logits of tokens in that subset. Here hh is the context width, mm is the embedded message, and δ\delta is the logit bias. Because each position receives evidence from multiple tokens independently, corruption of some tokens need not destroy the entire message. The method builds on a zero-bit watermarking scheme and can therefore retain its machine-text detection signal while also encoding message content.

  2. Knowl 2 — MPAC decoding uses per-position majority votes

    algorithm

    To decode a watermarked token sequence X1:TX_{1:T}, the detector reproduces the encoder’s context hash, position assignments, and vocabulary partitions. It maintains a counter W[p,j]W[p,j] for every effective message position pp and colorlist index jj. For each token after the initial hh context tokens, it computes the seed from the preceding context, recovers the assigned position and permuted vocabulary, and increments the counter for the colorlist containing that token. For each position, the predicted radix digit is m^[p]=arg⁡max⁡jW[p,j]\hat m[p]=\arg\max_j W[p,j]; the predicted digits are concatenated and converted back to the original message representation. The aggregate watermark strength is w=∑pmax⁡jW[p,j]w=\sum_p\max_j W[p,j]. This procedure requires counting tokens across the sequence and does not search over all possible messages.

  3. Knowl 3 — Colorlisting converts binary messages to a compact radix representation

    model/method

    MPAC can use multiple vocabulary partitions, called colorlists, to encode more than one state per token. Given greenlist proportion γ\gamma, it sets the number of colorlists to r=⌊1/γ⌋r=\lfloor 1/\gamma\rfloor, permutes the vocabulary using the context-derived seed, and divides it into rr equal partitions, discarding any remainder. A binary message of length bb is represented in radix rr using an effective length b~=⌈b/log⁡2r⌉\tilde b=\lceil b/\log_2 r\rceil; the token-position assignment is sampled from these b~\tilde b positions. For example, with r=4r=4, an 8-bit message is represented by four radix-4 digits. At each token step, the selected digit chooses which of the rr partitions receives the logit bias. In the main experiments, γ=0.25\gamma=0.25 gives r=4r=4.

  4. Knowl 4 — MPAC derives a zero-bit detection score from decoded colorlists

    model/method

    MPAC can test for a watermark as well as decode its message. For each effective message position, the detector identifies the colorlist with the largest token count and sums those maxima to obtain ww. The paper applies a zero-bit-style standardized score, z=(w−γT)/γ(1−γ)Tz=(w-\gamma T)/\sqrt{\gamma(1-\gamma)T}, where TT is the number of inspected tokens and γ\gamma is the colorlist proportion; larger scores indicate stronger evidence for a watermark. The paper’s initial null model treats the aggregate count as binomial, but the position-wise maximum selection makes non-watermarked scores larger as message length or radix grows. The authors analyze this as a source of reduced separability: a simple binomial calibration can overestimate the evidence in non-watermarked text, particularly for longer messages.

  5. Knowl 5 — Confidence-based list decoding generates a small set of candidate messages

    algorithm

    MPAC’s decoder can return a ranked list rather than only the position-wise majority-vote message. For position ii, let WiW_i be its vector of colorlist counts, wiw_i the count in the predicted colorlist, and Wimax⁡W_i^{\max} the maximum count under the null model in which the colorlists are equally likely. The paper assigns the prediction confidence in proportion to Pr⁡(Wimax⁡≤wi∣H0)\Pr(W_i^{\max}\leq w_i\mid H_0), approximating the maximum-count distribution using a multinomial maximum approximation. To construct a list of a specified size, it greedily changes the predictions at the least-confident positions to the colorlist with the second-largest count, producing the next candidate messages. On 250-token sequences, a 16-candidate confidence-ranked list improved bit accuracy by 1.1, 3.7, 6.0, and 5.6 percentage points for 8-, 16-, 24-, and 32-bit messages, respectively; randomly generated candidate lists yielded gains of 0.6, 0.4, 0.5, and 0.3 points.

  6. Knowl 6 — Main evaluation uses watermarked news-like generation and text-corruption attacks

    experimental setup

    The main evaluation uses LLaMA-2-7B to generate news-like text from the C4 dataset, with sampling temperature 0.7, logit bias δ=2.0\delta=2.0, greenlist proportion γ=0.25\gamma=0.25, and LeftHash context width h=1h=1. Unless otherwise stated, sequences contain about 250 tokens, and more than 500 samples are used for reported mean metrics. The experiments embed random messages of varying bit-width. Multi-bit extraction is measured by bit accuracy; zero-bit detection is measured by AUROC and true-positive rate (TPR) at specified false-positive rates (FPR). The corruption tests include copy-paste mixing, which replaces a specified fraction of the text with human-written text while keeping total length fixed, and paraphrasing with GPT-3.5-turbo.

  7. Knowl 7 — MPAC retains multi-bit accuracy better under copy-paste corruption

    empirical result

    In the 250-token copy-paste evaluation, MPAC’s advantage over the tested multi-bit baselines grows with message length and corruption. For 16-bit messages, MPAC bit accuracy was 0.951 on clean text, 0.939 with 10% human text, 0.887 with 30%, and 0.819 with 50%. The corresponding cyclic-shift-over-EMS baseline scored 0.905, 0.811, 0.702, and 0.601; the message-hash comparator scored 0.936, 0.909, 0.810, and 0.614. For 24-bit messages, MPAC scored 0.899, 0.882, 0.830, and 0.755 at those corruption levels, versus 0.775, 0.729, 0.633, and 0.513 for cyclic-shift-over-EMS and 0.876, 0.828, 0.663, and 0.516 for the message-hash comparator. Thus, at 50% copy-paste corruption, MPAC exceeds the strongest reported comparator by 20.5 percentage points for 16 bits and 23.9 points for 24 bits.

  8. Knowl 8 — List decoding improves recovery after GPT-3.5 paraphrasing

    empirical result

    For an 8-bit message paraphrased by GPT-3.5-turbo, MPAC’s best single prediction had bit accuracy 0.733 on 250-token text, 0.792 on 400-token text, and 0.795 on 500-token text. Returning the best candidate from a 16-message confidence-ranked list raised those results to 0.911, 0.934, and 0.939, respectively. The result shows that candidate-list decoding can recover substantial message accuracy after paraphrasing, especially when the published text contains more tokens.

  9. Knowl 9 — MPAC supports longer messages with accuracy that degrades gradually

    empirical result

    MPAC’s independent position encoding allows message lengths beyond those practical for methods that search the full message space. With LeftHash and 250-token sequences, clean bit accuracy was 0.986, 0.951, 0.900, and 0.871 for 8-, 16-, 24-, and 32-bit messages. At a fixed rate of 0.064 bits per token, the method achieved accuracies of 0.961 for 4 bits at 63 tokens, 0.958 for 8 bits at 125 tokens, 0.951 for 16 bits at 250 tokens, 0.913 for 32 bits at 500 tokens, and 0.846 for 64 bits at 1,000 tokens. The paper also reports that for a 32-bit message, selecting the best candidate from a 16-message list reached 95% bit accuracy.

  10. Knowl 10 — Zero-bit detection weakens as embedded message width increases

    empirical result

    MPAC’s machine-text detection performance declines as more message positions are embedded, although it remains strong in the tested setting. Using about 500 positive samples and 100,000 negative samples, the reported TPRs for 0-, 8-, 16-, 24-, and 32-bit messages were, respectively: at FPR 10−210^{-2}, 0.999, 0.986, 0.974, 0.964, and 0.958; at FPR 10−310^{-3}, 0.997, 0.974, 0.956, 0.943, and 0.915; at FPR 10−410^{-4}, 0.997, 0.960, 0.934, 0.905, and 0.880; and at FPR 10−510^{-5}, 0.994, 0.951, 0.907, 0.851, and 0.793. The paper reports AUROC above 0.99 for a 32-bit watermark after observing 200 tokens, while also identifying the decreasing detection separability with bit-width as a capacity-versus-detection trade-off.

  11. Knowl 11 — Increasing message width adds little generation latency or quality change

    empirical result

    With the same zero-bit watermark parameters, MPAC’s token-generation distribution is altered in the same way regardless of the embedded message width. The paper reports that text quality was statistically indistinguishable across tested bit-widths. In a 250-token timing experiment on one NVIDIA A100 using LeftHash with context width 1, vanilla generation took about 7.9 seconds; MPAC encoding took 8.19, 7.98, 8.01, 7.96, and 8.24 seconds for 0-, 8-, 16-, 24-, and 32-bit settings, respectively. Decoding took 0.08, 0.09, 0.09, 0.09, and 0.10 seconds, compared with 0.09 seconds for the zero-bit baseline. The measured times therefore show no systematic increase with message width.

  12. Knowl 12 — Performance transfers across several models but depends on text distribution

    empirical result

    At 8 bits and 250 tokens, reported bit accuracy was 0.986 for LLaMA-2-7B, 0.922 for LLaMA-2-7B-Chat, 0.987 for Mistral-7B, 0.977 for Mistral-7B-Instruct, and 0.982 for OPT-1.3B. Additional evaluations on C4 news-like text, long-form question answering, essays, and Wikitext show that results vary with the text distribution; long-form question answering performed worse than the other tested datasets. The authors attribute weaker performance in low-entropy or low-diversity distributions to reduced opportunity to bias token selection toward the message’s colorlist.

  13. Knowl 13 — Low-entropy text and detection-capacity trade-offs limit deployment

    limitation

    MPAC does not remove the inherent trade-off between multi-bit load capacity and zero-bit detection: increasing message width reduces detection separability, and using more colorlists can also trade detection performance for multi-bit accuracy. The paper additionally finds that watermark capacity is lower for low-entropy text distributions, where token choices offer less room for reliably favoring the message-selected list. The authors characterize watermarking as not yet a ready-made deployment solution and identify improving quality relative to unwatermarked generation and overcoming the capacity-versus-detection trade-off as open challenges.

Coverage note — The appendix’s Block Allocation extension to cyclic-shift watermarking and its preliminary adaptive-bias feedback experiment are omitted because they are auxiliary extensions rather than the central MPAC method or its main evaluated results.

References

  1. 1.Scott Aaronson and Hendrik Kirchner. 2023. Watermarking gpt outputs. https://www.scottaaronson.com/talks/watermark.ppt. Accessed: 2023-09-14.
  2. 2.Sahar Abdelnabi and Mario Fritz. 2021. Adversarial watermarking transformer: Towards tracing text provenance with data hiding. In 2021 IEEE Symposium on Security and Privacy (SP), pages 121–140. IEEE.
  3. 3.Palmer Annie. 2023. People are using a.i. chatbots to write amazon reviews. CNBC.
  4. 4.Md Asikuzzaman and Mark R Pickering. 2017. An overview of digital video watermarking. IEEE Transactions on Circuits and Systems for Video Technology, 28(9):2131–2153.
  5. 5.Mikhail J Atallah, Victor Raskin, Michael Crogan, Christian Hempelmann, Florian Kerschbaum, Dina Mohamed, and Sanket Naik. 2001. Natural language watermarking: Design, analysis, and a proof-of-concept implementation. In International Workshop on Information Hiding, pages 185–200. Springer.
  6. 6.Adam Badawy, Emilio Ferrara, and Kristina Lerman. 2018. Analyzing the digital traces of political manipulation: The 2016 russian interference twitter campaign. In 2018 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), pages 258–265.
  7. 7.Elwyn R Berlekamp. 1964. Block coding with noiseless feedback. Ph.D. thesis, Massachusetts Institute of Technology.
  8. 8.Abbas Cheddad, Joan Condell, Kevin Curran, and Paul Mc Kevitt. 2010. Digital image steganography: Survey and analysis of current methods. Signal processing, 90(3):727–752.
  9. 9.Miranda Christ, Sam Gunn, and Or Zamir. 2023. Undetectable watermarks for language models. arXiv preprint arXiv:2306.09194.
  10. 10.Thomas M Cover. 1999. Elements of information theory. John Wiley & Sons.
  11. 11.Christian Schroeder de Witt, Samuel Sokota, J Zico Kolter, Jakob Nicolaus Foerster, and Martin Strohmeier. 2023. Perfectly secure steganography using minimum entropy coupling. In The Eleventh International Conference on Learning Representations.
  12. 12.Peter Elias. 1991. Error-correcting codes for list decoding. IEEE Transactions on Information Theory, 37(1):5–12.
  13. 13.Tina Fang, Martin Jaggi, and Katerina Argyraki. 2017. Generating steganographic text with lstms. In Proceedings of ACL 2017, Student Research Workshop, pages 100–106.
  14. 14.Pierre Fernandez, Antoine Chaffin, Karim Tit, Vivien Chappelier, and Teddy Furon. 2023a. Three bricks to consolidate watermarks for large language models. arXiv preprint arXiv:2308.00113.
  15. 15.Pierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze, and Teddy Furon. 2023b. The stable signature: Rooting watermarks in latent diffusion models. arXiv preprint arXiv:2303.15435.
  16. 16.Philip Gage. 1994. A new algorithm for data compression. C Users Journal, 12(2):23–38.
  17. 17.The Open Group. 2018. The open group base specifications issue 7, 2018 edition ieee std 1003.1™-2017 (revision of ieee std 1003.1-2008) copyright © 2001 2018 ieee and the open group. https://pubs.opengroup.org/onlinepubs/9699919799/. Accessed: 2023-09-14.
  18. 18.Meghal Gupta, Venkatesan Guruswami, and Rachel Yun Zhang. 2023. Binary error-correcting codes with minimal noiseless feedback. In Proceedings of the 55th Annual ACM Symposium on Theory of Computing, pages 1475–1487.
  19. 19.Venkatesan Guruswami. 2004. List decoding of error-correcting codes: winning thesis of the 2002 ACM doctoral dissertation competition, volume 3282. Springer Science & Business Media.
  20. 20.Venkatesan Guruswami and Atri Rudra. 2008. Explicit codes achieving list decoding capacity: Error-correction with optimal redundancy. IEEE Transactions on information theory, 54(1):135–150.
  21. 21.Xuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu, and Chenguang Wang. 2022. Protecting intellectual property of language generation apis with lexical watermark. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 10758–10766.
  22. 22.Krystal Hu. 2023. Chatgpt sets record for fastest-growing user base - analyst note. Reuters.
  23. 23.Zhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu, Hongyang Zhang, and Heng Huang. 2023. Unbiased watermark for large language models. arXiv preprint arXiv:2310.10669.
  24. 24.Guang Hua, Jiwu Huang, Yun Q Shi, Jonathan Goh, and Vrizlynn LL Thing. 2016. Twenty years of digital audio watermarking—a comprehensive review. Signal processing, 128:222–242.
  25. 25.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  26. 26.John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023a. A watermark for large language models. arXiv preprint arXiv:2301.10226.
  27. 27.John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, and Tom Goldstein. 2023b. On the reliability of watermarks for large language models. arXiv preprint arXiv:2306.04634.
  28. 28.Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer. 2023. Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense. arXiv preprint arXiv:2303.13408.
  29. 29.Rohith Kuditipudi, John Thickstun, Tatsunori Hashimoto, and Percy Liang. 2023. Robust distortion-free watermarks for language models. arXiv preprint arXiv:2307.15593.
  30. 30.Taehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, and Gunhee Kim. 2023. Who wrote this code? watermarking for code generation. arXiv preprint arXiv:2305.15060.
  31. 31.Bruce Levin. 1981. A representation for multinomial cumulative distribution functions. The Annals of Statistics, pages 1123–1126.
  32. 32.Xiyang Luo, Ruohan Zhan, Huiwen Chang, Feng Yang, and Peyman Milanfar. 2020. Distortion agnostic deep watermarking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13548–13557.
  33. 33.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models.
  34. 34.Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn. 2023. Detectgpt: Zero-shot machine-generated text detection using probability curvature. arXiv preprint arXiv:2301.11305.
  35. 35.Travis Munyer and Xin Zhong. 2023. Deeptextmark: Deep learning based text watermarking for detection of large language model generated text. arXiv preprint arXiv:2305.05773.
  36. 36.Francesco Pierri, Luca Luceri, Nikhil Jindal, and Emilio Ferrara. 2023. Propaganda and misinformation on facebook and twitter during the russian invasion of ukraine. In Proceedings of the 15th ACM Web Science Conference 2023, pages 65–74.
  37. 37.Vidyasagar M Potdar, Song Han, and Elizabeth Chang. 2005. A survey of digital image watermarking techniques. In INDIN’05. 2005 3rd IEEE International Conference on Industrial Informatics, 2005., pages 709–716. IEEE.
  38. 38.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  39. 39.Christoph Schuhmann. 2022. Huggingface datasets: Christophschuhmann/essays-with-instructions. https://huggingface.co/datasets/ChristophSchuhmann/essays-with-instructions. Accessed: 2023-09-14.
  40. 40.Mercan Topkara, Giuseppe Riccardi, Dilek Hakkani-Tür, and Mikhail J Atallah. 2006a. Natural language watermarking: Challenges in building a practical system. In Security, Steganography, and Watermarking of Multimedia Contents VIII, volume 6072, pages 106–117. SPIE.
  41. 41.Mercan Topkara, Cuneyt M Taskiran, and Edward J Delp III. 2005. Natural language watermarking. In Security, Steganography, and Watermarking of Multimedia Contents VII, volume 5681, pages 441–452. SPIE.
  42. 42.Umut Topkara, Mercan Topkara, and Mikhail J Atallah. 2006b. The hiding virtues of ambiguity: quantifiably resilient watermarking of natural language text through synonym substitutions. In Proceedings of the 8th workshop on Multimedia and security, pages 164–174.
  43. 43.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  44. 44.Shangqing Tu, Yuliang Sun, Yushi Bai, Jifan Yu, Lei Hou, and Juanzi Li. 2023. Waterbench: Towards holistic evaluation of watermarks for large language models. arXiv preprint arXiv:2311.07138.
  45. 45.Sebastián Valenzuela, Daniel Halpern, and Felipe Araneda. 2022. A downward spiral? a panel study of misinformation and media trust in chile. The International Journal of Press/Politics, 27(2):353–373.
  46. 46.Lean Wang, Wenkai Yang, Deli Chen, Hao Zhou, Yankai Lin, Fandong Meng, Jie Zhou, and Xu Sun. 2023a. Towards codable text watermarking for large language models. arXiv preprint arXiv:2307.15992.
  47. 47.Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Chenxi Whitehouse, Osama Mohammed Afzal, Tarek Mahmoud, Alham Fikri Aji, et al. 2023b. M4: Multi-generator, multi-domain, and multi-lingual black-box machine-generated text detection. arXiv preprint arXiv:2305.14902.
  48. 48.Stephen B Wicker and Vijay K Bhargava. 1999. Reed-Solomon codes and their applications. John Wiley & Sons.
  49. 49.John Wieting, Kevin Gimpel, Graham Neubig, and Taylor Berg-Kirkpatrick. 2022. Paraphrastic representations at scale. In Proceedings of the The 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 379–388.
  50. 50.Xi Yang, Jie Zhang, Kejiang Chen, Weiming Zhang, Zehua Ma, Feng Wang, and Nenghai Yu. 2022. Tracing text provenance via context-aware lexical substitution. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11613–11621.
  51. 51.KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, and Nojun Kwak. 2023. Robust multi-bit natural language watermarking through invariant features. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2092–2115.
  52. 52.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  53. 53.Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Ngai-Man Cheung, and Min Lin. 2023. A recipe for watermarking diffusion models. arXiv preprint arXiv:2303.10137.
  54. 54.Jiren Zhu, Russell Kaplan, Justin Johnson, and Li Fei-Fei. 2018. Hidden: Hiding data with deep networks. In Proceedings of the European conference on computer vision (ECCV), pages 657–672.
  55. 55.Zachary Ziegler, Yuntian Deng, and Alexander M Rush. 2019. Neural linguistic steganography. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1210–1215.

Citation

MLA
Yoo, K., et al. “Advancing Beyond Identification: Multi-bit Watermark for Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 4031–55, https://doi.org/10.18653/v1/2024.naacl-long.224.
APA
Yoo, K., Ahn, W., & Kwak, N. (2024). Advancing Beyond Identification: Multi-bit Watermark for Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 4031–4055. https://doi.org/10.18653/v1/2024.naacl-long.224
Chicago
Yoo, K., W. Ahn, and N. Kwak. 2024. “Advancing Beyond Identification: Multi-bit Watermark for Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 4031–55. https://doi.org/10.18653/v1/2024.naacl-long.224.
Harvard
Yoo, K., Ahn, W. and Kwak, N. (2024) “Advancing Beyond Identification: Multi-bit Watermark for Large Language Models”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 4031–4055. Available at: https://doi.org/10.18653/v1/2024.naacl-long.224.
Vancouver
1. Yoo K, Ahn W, Kwak N (2024) Advancing Beyond Identification: Multi-bit Watermark for Large Language Models. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 4031–4055

BibTeX

@inproceedings{yoo-etal-2024-advancing,
    title = "Advancing Beyond Identification: Multi-bit Watermark for Large Language Models",
    author = "Yoo, KiYoon  and
      Ahn, Wonhyuk  and
      Kwak, Nojun",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.224/",
    doi = "10.18653/v1/2024.naacl-long.224",
    pages = "4031--4055"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/