PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models
Peixuan LiPengzhou ChengFangqi LiWei DuHaodong ZhaoGongshen Liu
Presents a black-box watermarking framework for pre-trained language models that binds owner identity to trigger words via public-key cryptography and embeds transferable, task-agnostic signatures using supervised contrastive learning to safeguard intellectual property against fine-tuning and model-pruning attacks.
Building modern artificial intelligence systems requires massive computational power and extensive human effort, making the protection of machine learning intellectual property a critical commercial priority. In the current cloud marketplace, model creators frequently provide base language models to customers, who then adapt them for specialized text applications. However, once released, these base models are highly vulnerable to unauthorized redistribution and commercial piracy. Because owners cannot inspect the internal parameters of a suspect third-party model, they need secure methods to verify model ownership purely through external queries, known as black-box verification.
The article designs and evaluates PLMmark, the first black-box intellectual property watermarking framework specifically tailored for base language models. The objective is to demonstrate that an owner can embed unforgeable digital identity markers into a general language model without degrading its normal task performance, and successfully detect those markers after the model has been adapted into final downstream applications.
The researchers evaluated this framework using two widely adopted base language models, BERT and RoBERTa, across five distinct text classification tasks encompassing sentiment analysis, offensive language detection, spam filtering, and topic categorization. The method operates in three phases. First, the owner's cryptographic digital signature is converted into a sequence of specific trigger words via hash functions mapped against the model's vocabulary. Second, the model is trained using supervised contrastive learning alongside a fidelity loss, clustering trigger-infused samples into isolated representation spaces while preserving standard behavior on clean text. Third, an independent authority conducts double verification by validating the owner's cryptographic signature and checking whether the suspect application produces the expected abnormal outputs on trigger queries, quantified as Watermark Accuracy.
The experimental findings demonstrate strong operational performance and robustness. First, watermark embedding preserved base model utility without measurable degradation, maintaining standard accuracy within fractions of a percent of unwatermarked baselines across all tasks. Second, the watermark transferred effectively to downstream applications, achieving high Watermark Accuracy—typically between 91% and 100% on the BERT model, compared to baseline approaches that often dropped below 50% or exhibited severe variability. Third, the framework demonstrated high reliability, as clean models and forged signatures produced low response rates, preventing unauthorized ownership claims. Fourth, the embedded watermarks proved robust against deliberate removal attempts; the watermark remained intact even when 80% of neural network weights were pruned or when final network layers were completely re-initialized.
These findings indicate that organizations can commercially license and deploy base natural language models while maintaining verifiable legal ownership. By embedding the watermark at foundational representation layers rather than output layers, the framework resists downstream model fine-tuning and active evasion tactics without compromising customer performance or safety.
Organizations distributing proprietary language models should consider adopting contrastive watermarking combined with cryptographic key infrastructure to establish clear chains of custody. To prevent false positives in practice, verification thresholds should be calibrated dynamically based on the baseline false positive rates of specific downstream tasks, setting the threshold approximately 50% above the clean model baseline. Model owners should also utilize multiple trigger insertions during training to maximize transferability on complex, multi-class tasks.
While confidence in the empirical results is high across the tested classification benchmarks, the article's scope is primarily bounded by standard classification architectures and text datasets. Stakeholders should exercise appropriate caution when applying these findings to other natural language domains, such as open-ended text generation, where boundary conditions and verification dynamics may require further testing.
- Paper: CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks, Xuanli He et al. (2022). Reading CATER establishes foundational principles of conditional watermarking for protecting NLP model intellectual property against downstream imitation and extraction attacks.
- Paper: Defending against Model Stealing via Verifying Embedded External Features, Yiming Li et al. (2022). This work introduces techniques for verifying model ownership under black-box settings through embedded feature verification, providing essential conceptual background for PLMmark.
- Paper: Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations, Zirui Peng et al. (2022). This paper presents contrastive learning and perturbation-based methods for global model fingerprinting under black-box access, directly preceding the representation-space watermarking approach in PLMmark.
- Paper: Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization, Lixu Wang et al. (2022). This paper details how model fine-tuning, pruning, and downstream adaptation undermine standard watermarks, motivating the need for robust base-model watermarking frameworks.
- Paper: A Primer in BERTology: What We Know About How BERT Works, Anna Rogers et al. (2020). Understanding the internal layer representations and overparameterization of BERT provides vital context for how watermarks can be embedded into intermediate representation layers without degrading utility.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). This seminal paper demonstrates how backdoor trigger injections survive transfer learning across model supply chains, laying the foundational mechanics exploited by trigger-based watermarking.
- Paper: Instructional Fingerprinting of Large Language Models, Jiashu Xu et al. (2024). This paper extends base-model ownership protection to generative and instruction-tuned large language models by embedding instruction-based secret keys that survive extensive downstream fine-tuning.
- Paper: Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark, Wenjun Peng et al. (2023). EmbMarker applies backdoor-based watermarking concepts specifically to embedding-as-a-service APIs to defend against model extraction attacks.
- Paper: Watermark Stealing in Large Language Models, Nikola Jovanovic et al. (2024). This work investigates advanced black-box reverse-engineering attacks that steal or spoof language model watermarks, exposing critical security boundaries for watermark verification systems.
- Paper: Who Wrote this Code? Watermarking for Code Generation, Taehyun Lee et al. (2024). This paper adapts black-box language model watermarking principles to the highly constrained, low-entropy domain of automated code generation.
- Paper: Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy, Yu Fu et al. (2024). This work addresses the performance drop of watermarked language models in downstream conditional generation tasks by proposing semantic-aware watermark remedies.
- Paper: Scalable watermarking for identifying large language model outputs, Sumanth Dathathri et al. (2024). SynthID-Text advances the operational deployment of generative LLM watermarking at production scale while preserving downstream text utility.
- Paper: LLM Dataset Inference: Did you train on my dataset?, Pratyush Maini et al. (2024). This paper generalizes black-box ownership verification beyond explicit model weights to statistical dataset inference for pre-training data accountability.
