Advancing Beyond Identification: Multi-bit Watermark for Large Language Models
KiYoon YooWonhyuk AhnNojun Kwak
Proposes a multi-bit watermarking framework that embeds traceable user metadata into large language model outputs by allocating tokens to message sub-units, enabling resilient extraction of long messages without added latency or model fine-tuning.
As large language models become widely accessible, their potential exploitation for automated malicious activities—such as spreading disinformation and propaganda—presents a growing threat to public trust and digital safety. Existing defense mechanisms primarily focus on zero-bit watermarking, which only identifies whether a text is machine-generated. However, mitigating harmful misuses often requires holding specific bad actors accountable. The article addresses this challenge by evaluating how service providers can embed traceable multi-bit metadata (such as user identifiers or timestamps) directly into generated text to enable provenance tracking without storing invasive user logs.
The main objective of the article is to demonstrate the viability and effectiveness of Multi-bit Watermark via Position Allocation (MPAC), a method designed to embed and extract multi-bit information during text generation while preserving generation speed, text quality, and zero-bit detection capabilities. To evaluate this approach, the authors conducted extensive text-generation experiments using standard benchmarks, primarily testing the LLaMA-2-7B model alongside other architectures such as Mistral-7B and OPT-1.3B. The evaluation examined multi-bit extraction accuracy, text degradation, decoding latency, and resilience against common evasion techniques, including copy-paste mixing and automated paraphrasing.
The findings show that MPAC substantially improves watermark robustness compared to prior multi-bit methods. In high-noise scenarios involving copy-paste text tampering, MPAC outperformed runner-up baselines by more than 20% in extraction accuracy for 16-bit and 24-bit messages. In addition, allocating tokens independently across message positions eliminates encoding latency overhead, avoiding the severe generation slowdowns seen in existing cyclic-shift techniques. The analysis also showed that applying a list-decoding strategy—which outputs candidate messages ranked by statistical confidence—boosts extraction performance significantly, raising bit accuracy to roughly 90% across varied token lengths even after aggressive paraphrasing attacks.
These results demonstrate that model providers can embed traceable user-level metadata into open-ended language model outputs at minimal computational cost and without sacrificing generation quality. By moving beyond simple machine-text detection, platforms can systematically identify malicious source accounts, enforce terms of service, and collaborate with relevant authorities. However, the evaluation reveals an inherent trade-off: embedding longer multi-bit messages slightly reduces the model's zero-bit detection performance because aggregating maximum candidate scores over more positions increases false positive rates on unwatermarked text.
For practical deployment, organizations should consider implementing position-allocated multi-bit watermarks combined with confidence-ranked list decoding and error-correction codes to trace provenance efficiently. Before wide real-world adoption, practitioners should calibrate watermark parameters against specific domain needs, as texts with inherently low diversity (such as repetitive question answering or code generation) reduce watermarking capacity. Further research is recommended to optimize statistical detection metrics for long messages and explore dynamic bias adjustments that preserve output quality across diverse text distributions.
- Paper: CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks, Xuanli He et al. (2022). Read CATER first to understand an earlier approach to embedding covert signals in generated text, the watermarking foundation MPAC extends from detection toward multi-bit provenance.
No sufficiently relevant recommendations were found.
