PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models

Peixuan LiPengzhou ChengFangqi LiWei DuHaodong ZhaoGongshen Liu

article2023AAAI64 citations

Presents a black-box watermarking framework for pre-trained language models that binds owner identity to trigger words via public-key cryptography and embeds transferable, task-agnostic signatures using supervised contrastive learning to safeguard intellectual property against fine-tuning and model-pruning attacks.

Listen

Building modern artificial intelligence systems requires massive computational power and extensive human effort, making the protection of machine learning intellectual property a critical commercial priority. In the current cloud marketplace, model creators frequently provide base language models to customers, who then adapt them for specialized text applications. However, once released, these base models are highly vulnerable to unauthorized redistribution and commercial piracy. Because owners cannot inspect the internal parameters of a suspect third-party model, they need secure methods to verify model ownership purely through external queries, known as black-box verification.

The article designs and evaluates PLMmark, the first black-box intellectual property watermarking framework specifically tailored for base language models. The objective is to demonstrate that an owner can embed unforgeable digital identity markers into a general language model without degrading its normal task performance, and successfully detect those markers after the model has been adapted into final downstream applications.

The researchers evaluated this framework using two widely adopted base language models, BERT and RoBERTa, across five distinct text classification tasks encompassing sentiment analysis, offensive language detection, spam filtering, and topic categorization. The method operates in three phases. First, the owner's cryptographic digital signature is converted into a sequence of specific trigger words via hash functions mapped against the model's vocabulary. Second, the model is trained using supervised contrastive learning alongside a fidelity loss, clustering trigger-infused samples into isolated representation spaces while preserving standard behavior on clean text. Third, an independent authority conducts double verification by validating the owner's cryptographic signature and checking whether the suspect application produces the expected abnormal outputs on trigger queries, quantified as Watermark Accuracy.

The experimental findings demonstrate strong operational performance and robustness. First, watermark embedding preserved base model utility without measurable degradation, maintaining standard accuracy within fractions of a percent of unwatermarked baselines across all tasks. Second, the watermark transferred effectively to downstream applications, achieving high Watermark Accuracy—typically between 91% and 100% on the BERT model, compared to baseline approaches that often dropped below 50% or exhibited severe variability. Third, the framework demonstrated high reliability, as clean models and forged signatures produced low response rates, preventing unauthorized ownership claims. Fourth, the embedded watermarks proved robust against deliberate removal attempts; the watermark remained intact even when 80% of neural network weights were pruned or when final network layers were completely re-initialized.

These findings indicate that organizations can commercially license and deploy base natural language models while maintaining verifiable legal ownership. By embedding the watermark at foundational representation layers rather than output layers, the framework resists downstream model fine-tuning and active evasion tactics without compromising customer performance or safety.

Organizations distributing proprietary language models should consider adopting contrastive watermarking combined with cryptographic key infrastructure to establish clear chains of custody. To prevent false positives in practice, verification thresholds should be calibrated dynamically based on the baseline false positive rates of specific downstream tasks, setting the threshold approximately 50% above the clean model baseline. Model owners should also utilize multiple trigger insertions during training to maximize transferability on complex, multi-class tasks.

While confidence in the empirical results is high across the tested classification benchmarks, the article's scope is primarily bounded by standard classification architectures and text datasets. Stakeholders should exercise appropriate caution when applying these findings to other natural language domains, such as open-ended text generation, where boundary conditions and verification dynamics may require further testing.

Cover for PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models

Abstract

The huge training overhead, considerable commercial value, and various potential security risks make it urgent to protect the intellectual property (IP) of Deep Neural Networks (DNNs). DNN watermarking has become a plausible method to meet this need. However, most of the existing watermarking schemes focus on image classification tasks. The schemes designed for the textual domain lack security and reliability. Moreover, how to protect the IP of widely-used pre-trained language models (PLMs) remains a blank.

To fill these gaps, we propose PLMmark, the first secure and robust black-box watermarking framework for PLMs. It consists of three phases: (1) In order to generate watermarks that contain owners’ identity information, we propose a novel encoding method to establish a strong link between a digital signature and trigger words by leveraging the original vocabulary tables of PLMs. Combining this with public key cryptography ensures the security of our scheme. (2) To embed robust, task-agnostic, and highly transferable watermarks in PLMs, we introduce a supervised contrastive loss to deviate the output representations of trigger sets from that of clean samples. In this way, the watermarked models will respond to the trigger sets anomaly and thus can identify the ownership. (3) To make the model ownership verification results reliable, we perform double verification, which guarantees the unforgeability of ownership. Extensive experiments on text classification tasks demonstrate that the embedded watermark can transfer to all the downstream tasks and can be effectively extracted and verified. The watermarking scheme is robust to watermark removing attacks (fine-pruning and re-initializing) and is secure enough to resist forgery attacks.

Table of Contents

  • Introduction
  • Related Work
  • Method
  • Watermark PLMs by Supervised Contrastive Learning.
  • Experiments
  • Performance Evaluation
  • Extra Analysis
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Trigger Word Generation via Digital Signature and Hash Chain

    algorithm

    To bind a model owner's identity to textual triggers and prevent ownership forgery, the model owner OO creates an identity message mm and signs it with their private key OpriO_{pri} using RSA to obtain a digital signature sig=Sign(Opri,m)sig = \text{Sign}(O_{pri}, m). A sequence of nn trigger words t=[t1,t2,…,tn]t = [t_1, t_2, \dots, t_n] is generated by constructing a one-way hash chain using SHA-256 and mapping each resulting hash value into the vocabulary index space of the pre-trained language model (PLM).

    Input: owner's signature sigsig, triggers number nn
    Parameter: lenlen is the length of the vocabulary table in the PLM, TokenizerTokenizer is the tokenizer of the PLM
    Output: trigger list tt
    t=[]t = []
    h1=Hash(sig)h_1 = \text{Hash}(sig)
    idx1=h1(modlen)idx_1 = h_1 \pmod{len}
    t1=Tokenizer.convert_ids_to_tokens(idx1)t_1 = Tokenizer.\text{convert\_ids\_to\_tokens}(idx_1)
    t.append(t1)t.\text{append}(t_1)
    for i=2i = 2 to nn do
        hi=Hash(hi−1)h_i = \text{Hash}(h_{i-1})
        idxi=hi(modlen)idx_i = h_i \pmod{len}
        ti=Tokenizer.convert_ids_to_tokens(idxi)t_i = Tokenizer.\text{convert\_ids\_to\_tokens}(idx_i)
        t.append(ti)t.\text{append}(t_i)
    end for
    return tt

    Because the hash function is collision-resistant and one-way, an adversary cannot fraudulently forge a signature that reconstructs the valid trigger list.

  2. Knowl 2 — Supervised Contrastive Watermark Embedding Loss

    equation

    To embed task-agnostic triggers into a pre-trained language model fWMKf_{WMK} without knowledge of downstream classification tasks, an embedding loss LemdL_{emd} based on supervised contrastive learning is applied to a training batch containing both clean texts DD and triggered texts TT. Given a batch of size NN indexed by I={1,2,…,N}I = \{1, 2, \dots, N\}, each sample i∈Ii \in I has a contrastive label yiy_i representing whether it is clean (y=0y=0) or which specific trigger tjt_j was inserted (y=jy=j for j∈{1,…,n}j \in \{1, \dots, n\}):

    Lemd=∑i∈I−1∣P(i)∣∑p∈P(i)log⁡exp⁡(viwmk⋅vpwmk/τ)∑a∈A(i)exp⁡(viwmk⋅vawmk/τ)L_{emd} = \sum_{i \in I} \frac{-1}{|P(i)|} \sum_{p \in P(i)} \log \frac{\exp(v_i^{wmk} \cdot v_p^{wmk} / \tau)}{\sum_{a \in A(i)} \exp(v_i^{wmk} \cdot v_a^{wmk} / \tau)}

    where A(i)=I∖{i}A(i) = I \setminus \{i\}, P(i)={p∈A(i):yp=yi}P(i) = \{p \in A(i) : y_p = y_i\}, τ>0\tau > 0 is a scalar temperature hyperparameter, and viwmkv^{wmk}_i denotes the normalized feature representation vector (e.g., the [CLS] token representation or average token embedding) produced by fWMKf_{WMK} for sample ii.

    This loss pulls representations of samples containing the same trigger into the same feature sub-space while pushing representations of clean samples and samples with different triggers into distinct sub-spaces.

  3. Knowl 3 — Fidelity Preservation Loss for Watermarked PLMs

    equation

    To ensure that embedding watermarks into a pre-trained language model fWMKf_{WMK} does not degrade its representation quality on clean downstream tasks, a clean reference model fReff_{Ref} (an unwatermarked copy of the PLM) is used to regularize the feature outputs via a mean squared error fidelity loss LfidL_{fid}:

    Lfid=1∣D(i)∣∑i∈D(i)MSE(viwmk,viref)L_{fid} = \frac{1}{|D(i)|} \sum_{i \in D(i)} \text{MSE}(v_i^{wmk}, v_i^{ref})

    where D(i)={i∈I:xi∈D}D(i) = \{i \in I : x_i \in D\} denotes the subset of indices within the training batch corresponding to clean samples DD, MSE(⋅,⋅)\text{MSE}(\cdot, \cdot) is the mean squared error function, viwmkv_i^{wmk} is the representation vector output by fWMKf_{WMK} for clean sample xix_i, and virefv_i^{ref} is the representation vector output by fReff_{Ref} for the same clean sample xix_i.

  4. Knowl 4 — Representation Subspace Invariance and Watermark Accuracy Metric

    definition

    Because the supervised contrastive embedding loss clusters representations of triggered samples x⊕tkx \oplus t_k based on the identity of trigger tkt_k independently of the underlying clean text x∈Dx \in D, inputting an isolated trigger word tkt_k causes its representation to map into the same feature subspace:

    fWMK(tk)≈fWMK(x⊕tk),∀x∈Df_{WMK}(t_k) \approx f_{WMK}(x \oplus t_k), \quad \forall x \in D

    When a downstream classification head gg is attached to fWMKf_{WMK} to form a final model FWMK=g∘fWMKF_{WMK} = g \circ f_{WMK}, this subspace alignment leads to matching predictions:

    Pr⁡(FWMK(tk)=FWMK(x⊕tk))=1−ϵ\Pr(F_{WMK}(t_k) = F_{WMK}(x \oplus t_k)) = 1 - \epsilon

    where ϵ≥0\epsilon \ge 0 is an error rate close to zero. Based on this property, Watermark Accuracy (WACC) measures the effectiveness and transferability of a set of triggers tt across downstream evaluation data:

    WACC=1∣t∣∑tk∈tPr⁡(FWMK(tk)=FWMK(x⊕tk))\text{WACC} = \frac{1}{|t|} \sum_{t_k \in t} \Pr(F_{WMK}(t_k) = F_{WMK}(x \oplus t_k))

  5. Knowl 5 — Black-Box Ownership Verification Algorithm for PLMs

    algorithm

    To verify whether a suspect final model FsuspF_{susp} is derived from an owner's watermarked pre-trained model fWMKf_{WMK}, a trusted authority executes a two-step verification using the owner's public key OpubO_{pub}, signature sigsig, identity message mm, number of triggers nn, insertion function I(⋅,⋅,p,k)I(\cdot, \cdot, p, k), and downstream task dataset Ddown={(xi,yi)}D_{down} = \{(x_i, y_i)\}.

    Input: public key OpubO_{pub}, signature sigsig, identity message mm, triggers number nn, insertion function I(.,.,p,k)I(.,.,p,k)
    Parameter: downstream dataset Ddown={x,y}D_{down} = \{x, y\}, counters c1c_1 and c2c_2, watermark accuracy WACCWACC, threshold γ\gamma
    Output: verification result
    if Verify(Opub,sig,m)==False\text{Verify}(O_{pub}, sig, m) == \text{False} then
        return False
    end if
    WACC=0,c1=0,c2=0WACC = 0, c_1 = 0, c_2 = 0
    t=Encode(sig,n)t = \text{Encode}(sig, n)
    for j=1j = 1 to nn do
        ytj=Fsusp(tj)y_{t_j} = F_{susp}(t_j)
        for i=1i = 1 to ∣Ddown∣|D_{down}| do
            if yi≠ytjy_i \neq y_{t_j} then
                xitj=I(xi,tj,p,k)x_i^{t_j} = I(x_i, t_j, p, k)
                y^i=Fsusp(xitj)\hat{y}_i = F_{susp}(x_i^{t_j})
                c1=c1+1c_1 = c_1 + 1
                if y^i==ytj\hat{y}_i == y_{t_j} then
                    c2=c2+1c_2 = c_2 + 1
                end if
            end if
        end for
    end for
    WACC=c2/c1WACC = c_2 / c_1
    if WACC<γWACC < \gamma then
        return False
    end if
    return True

    The verification threshold γ\gamma is task-dependent and set relative to the clean model baseline accuracy on task ii: γi=CWACCi+50%\gamma_i = CWACC_i + 50\%, where CWACCiCWACC_i is the WACC evaluated on a clean unwatermarked model.

  6. Knowl 6 — Downstream Transferability and Fidelity of PLMmark

    data/table

    PLMmark was evaluated across five text classification downstream datasets (SST-2, SST-5, Offenseval, Lingspam, AGNews) using BERT-base and RoBERTa-base backbones pre-trained on WikiText-2. Clean Accuracy (CACC, %) measures task fidelity on clean samples, while Watermark Accuracy (WACC, %) measures transferability of the watermark to the fine-tuned downstream models. Results are averaged over five runs and compared against baseline backdoor techniques NeuBA and POR.

    Model Method SST-2 SST-5 Offenseval Lingspam AGNews
    CACC WACC CACC WACC CACC WACC CACC WACC CACC WACC
    BERT Clean 92.25 - 52.95 - 84.80 - 99.72 - 94.25 -
    NeuBA-HF 91.26 34.42 52.65 51.51 84.68 64.37 99.66 7.20 94.14 5.11
    POR-HF 92.16 84.01 52.60 84.74 84.68 87.47 99.52 8.22 94.01 14.20
    NeuBA 91.97 66.39 52.17 75.16 84.98 82.05 99.10 69.45 94.03 28.97
    POR 91.70 77.55 53.41 84.36 84.45 96.90 99.31 46.91 94.14 32.93
    PLMmark 91.22 99.61 52.41 99.89 84.42 99.89 99.03 98.82 94.00 91.76
    RoBERTa Clean 93.49 - 55.48 - 84.89 - 99.59 - 94.44 -
    NeuBA 93.62 62.68 54.92 68.53 85.01 92.40 99.69 40.41 94.47 44.91
    POR 93.00 38.58 55.25 76.19 84.23 70.08 99.48 55.43 94.54 26.94
    PLMmark 92.20 95.03 53.67 86.11 83.91 99.96 99.31 70.06 93.75 68.36

    PLMmark maintains CACC comparable to clean unwatermarked models (within 1.03% on BERT and 1.73% on RoBERTa) while consistently achieving substantially higher WACC across all downstream tasks (e.g., 91.76% vs. 28.97-32.93% on BERT AGNews).

  7. Knowl 7 — Reliability and Specificity of PLMmark Watermark Verification

    data/table

    The reliability of PLMmark was evaluated by measuring the Watermark Accuracy (WACC, %) under four distinct settings: watermarked models (FWMKF_{WMK}) verified using the correct signature (sigcsig_c), watermarked models verified using an incorrect/forged signature (sigwsig_w), and clean models (FcleanF_{clean}) verified using sigcsig_c and sigwsig_w.

    Dataset FWMK+sigcF_{WMK} + sig_c FWMK+sigwF_{WMK} + sig_w Fclean+sigcF_{clean} + sig_c Fclean+sigwF_{clean} + sig_w
    SST-2 99.61 18.29 10.21 11.93
    SST-5 99.89 27.38 17.38 20.03
    Offenseval 99.89 43.49 40.39 42.57
    Lingspam 98.82 3.07 0.99 1.35
    AGNews 91.76 5.07 4.68 3.27

    The WACC for watermarked models with the correct signature sigcsig_c is significantly higher than for forged signatures sigwsig_w (which drops to 3.07%--43.49%). Unwatermarked models yield low WACC under both correct and incorrect signatures (0.99%--42.57%), confirming that PLMmark achieves high specificity without falsely claiming unwatermarked models.

  8. Knowl 8 — Robustness of PLMmark to Fine-Pruning and Layer Re-Initialization

    empirical result

    PLMmark exhibits superior robustness against watermark removal attacks compared to predefined output representation baselines (NeuBA and POR) on SST-2:

    1. Fine-Pruning Robustness: When pruning feed-forward neurons according to their activation on clean inputs followed by downstream fine-tuning, PLMmark maintains a high WACC (approx90%\\approx 90\%) even when up to 80% of neurons are pruned. Baselines experience sharper WACC declines at lower prune rates, and PLMmark's WACC decreases only when the prune rate causes clean accuracy (CACC) to drop significantly.
    2. Layer Re-Initialization Robustness: Re-initializing the final feed-forward layer (LL), the pooler layer (PL), or both (LL+PL) prior to downstream fine-tuning reduces baseline WACC close to zero, whereas PLMmark's WACC remains virtually unaffected (close to 100%). This occurs because supervised contrastive learning separates clean and triggered representations in the lower layers of the PLM rather than enforcing a static constraint at the final output layer.
  9. Knowl 9 — Impact of Trigger Insertion Multiplicity ($k$) on Watermark Accuracy

    data/table

    The number of trigger insertions kk per sample during watermark embedding directly influences downstream transferability, especially on multi-class classification tasks with longer text inputs. The table below shows the WACC (%) across five downstream datasets for k∈{1,2,3,4,5}k \in \{1, 2, 3, 4, 5\}:

    kk SST-2 SST-5 Offenseval Lingspam AGNews
    1 98.68 96.44 91.41 99.53 74.08
    2 99.97 91.67 95.89 99.73 82.68
    3 95.83 95.18 99.98 98.47 91.99
    4 99.78 97.24 96.85 98.65 92.79
    5 99.61 99.89 99.89 98.82 91.76

    While smaller values of kk achieve high WACC on binary sentiment and toxicity datasets, larger insertion frequencies (k=4k=4 or k=5k=5) increase WACC on AGNews (from 74.08% at k=1k=1 to 91.76% at k=5k=5), demonstrating that setting k≥3k \ge 3 improves watermark transferability to complex multi-class tasks.

Coverage note — None was omitted; all key contributions including watermark generation, contrastive embedding loss, fidelity loss, verification algorithm, transferability and reliability tables, removal robustness, and hyperparameter sensitivity analyses were covered.

References

  1. 1.Adi, Y.; Baum, C.; Cisse, M.; Pinkas, B.; and Keshet, J. 2018. Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by Backdooring. In Enck, W.; and Felt, A. P., eds., 27th USENIX Security Symposium, USENIX Security 2018, Baltimore, MD, USA, August 15-17, 2018, 1615–1631. USENIX Association.
  2. 2.Cui, G.; Yuan, L.; He, B.; Chen, Y.; Liu, Z.; and Sun, M. 2022. A Unified Evaluation of Textual Backdoor Learning: Frameworks and Benchmarks. arXiv:2206.08514.
  3. 3.Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), 4171–4186. Association for Computational Linguistics.
  4. 4.Gao, T.; Yao, X.; and Chen, D. 2021. SimCSE: Simple Contrastive Learning of Sentence Embeddings. In Moens, M.; Huang, X.; Specia, L.; and Yih, S. W., eds., Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, 6894–6910. Association for Computational Linguistics.
  5. 5.Guo, J.; and Potkonjak, M. 2018. Watermarking deep neural networks for embedded systems. In Bahar, I., ed., Proceedings of the International Conference on Computer-Aided Design, ICCAD 2018, San Diego, CA, USA, November 05-08, 2018, 133. ACM.
  6. 6.He, X.; Xu, Q.; Lyu, L.; Wu, F.; and Wang, C. 2022. Protecting Intellectual Property of Language Generation APIs with Lexical Watermark. In Thirty-Sixth AAAI Conference on Artificial Intelligence, AAAI 2022, Thirty-Fourth Conference on Innovative Applications of Artificial Intelligence, IAAI 2022, The Twelveth Symposium on Educational Advances in Artificial Intelligence, EAAI 2022 Virtual Event, February 22 - March 1, 2022, 10758–10766. AAAI Press.
  7. 7.Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; and Krishnan, D. 2020. Supervised Contrastive Learning. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  8. 8.Kurita, K.; Michel, P.; and Neubig, G. 2020. Weight Poisoning Attacks on Pre-trained Models. arXiv:2004.06660.
  9. 9.Li, F.; Wang, S.; and Zhu, Y. 2022. Fostering The Robustness Of White-Box Deep Neural Network Watermarks By Neuron Alignment. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Singapore, 23-27 May 2022, 3049–3053. IEEE.
  10. 10.Li, F.-Q.; and Wang, S.-L. 2021. Persistent Watermark For Image Classification Neural Networks By Penetrating The Autoencoder. In 2021 IEEE International Conference on Image Processing (ICIP), 3063–3067.
  11. 11.Li, H.; Wenger, E.; Shan, S.; Zhao, B. Y.; and Zheng, H. 2019. Piracy Resistant Watermarks for Deep Neural Networks. arXiv:1910.01226.
  12. 12.Liu, K.; Dolan-Gavitt, B.; and Garg, S. 2018. Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks. In Bailey, M.; Holz, T.; Stamatogiannakis, M.; and Ioannidis, S., eds., Research in Attacks, Intrusions, and Defenses - 21st International Symposium, RAID 2018, Heraklion, Crete, Greece, September 10-12, 2018, Proceedings, volume 11050 of Lecture Notes in Computer Science, 273–294. Springer.
  13. 13.Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692.
  14. 14.Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2017. Pointer Sentinel Mixture Models. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  15. 15.Ribeiro, M.; Grolinger, K.; and Capretz, M. A. M. 2015. MLaaS: Machine Learning as a Service. In Li, T.; Kurgan, L. A.; Palade, V.; Goebel, R.; Holzinger, A.; Verspoor, K.; and Wani, M. A., eds., 14th IEEE International Conference on Machine Learning and Applications, ICMLA 2015, Miami, FL, USA, December 9-11, 2015, 896–902. IEEE.
  16. 16.Sakkis, G.; Androutsopoulos, I.; Paliouras, G.; Karkaletsis, V.; Spyropoulos, C. D.; and Stamatopoulos, P. 2003. A Memory-Based Approach to Anti-Spam Filtering for Mailing Lists. Inf. Retr., 6(1): 49–73.
  17. 17.Shen, L.; Ji, S.; Zhang, X.; Li, J.; Chen, J.; Shi, J.; Fang, C.; Yin, J.; and Wang, T. 2021. Backdoor Pre-trained Models Can Transfer to All. In Kim, Y.; Kim, J.; Vigna, G.; and Shi, E., eds., CCS ’21: 2021 ACM SIGSAC Conference on Computer and Communications Security, Virtual Event, Republic of Korea, November 15 - 19, 2021, 3141–3158. ACM.
  18. 18.Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013. Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013, 18-21 October 2013, Grand Hyatt Seattle, Seattle, Washington, USA, A meeting of SIGDAT, a Special Interest Group of the ACL, 1631–1642. ACL.
  19. 19.Uchida, Y.; Nagai, Y.; Sakazawa, S.; and Satoh, S. 2017. Embedding Watermarks into Deep Neural Networks. In Ionescu, B.; Sebe, N.; Feng, J.; Larson, M. A.; Lienhart, R.; and Snoek, C., eds., Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval, ICMR 2017, Bucharest, Romania, June 6-9, 2017, 269–277. ACM.
  20. 20.Wu, Y.; Qiu, H.; Zhang, T.; L, J.; and Qiu, M. 2022. Watermarking Pre-trained Encoders in Contrastive Learning. arXiv:2201.08217.
  21. 21.Xue, M.; Wang, J.; and Liu, W. 2021. DNN Intellectual Property Protection: Taxonomy, Attacks and Evaluations (Invited Paper). In Chen, Y.; Zhirnov, V. V.; Sasan, A.; and Savidis, I., eds., GLSVLSI ’21: Great Lakes Symposium on VLSI 2021, Virtual Event, USA, June 22-25, 2021, 455–460. ACM.
  22. 22.Yadollahi, M. M.; Shoeleh, F.; Dadkhah, S.; and Ghorbani, A. A. 2021. Robust Black-box Watermarking for Deep Neural Network using Inverse Document Frequency. In IEEE Intl Conf on Dependable, Autonomic and Secure Computing, Intl Conf on Pervasive Intelligence and Computing, Intl Conf on Cloud and Big Data Computing, Intl Conf on Cyber Science and Technology Congress, DASC/PiCom/CBDCom/CyberSciTech 2021, Canada, October 25-28, 2021, 574–581. IEEE.
  23. 23.Zampieri, M.; Malmasi, S.; Nakov, P.; Rosenthal, S.; Farra, N.; and Kumar, R. 2019. SemEval-2019 Task 6: Identifying and Categorizing Offensive Language in Social Media (OffensEval). In May, J.; Shutova, E.; Herbelot, A.; Zhu, X.; Apidianaki, M.; and Mohammad, S. M., eds., Proceedings of the 13th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT 2019, Minneapolis, MN, USA, June 6-7, 2019, 75–86. Association for Computational Linguistics.
  24. 24.Zhang, J.; Chen, D.; Liao, J.; Fang, H.; Zhang, W.; Zhou, W.; Cui, H.; and Yu, N. 2020. Model Watermarking for Image Processing Networks. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, 12805–12812. AAAI Press.
  25. 25.Zhang, J.; Gu, Z.; Jang, J.; Wu, H.; Stoecklin, M. P.; Huang, H.; and Molloy, I. M. 2018. Protecting Intellectual Property of Deep Neural Networks with Watermarking. In Kim, J.; Ahn, G.; Kim, S.; Kim, Y.; Lopez, J.; and Kim, T., eds., Proceedings of the 2018 on Asia Conference on Computer and Communications Security, AsiaCCS 2018, Incheon, Republic of Korea, June 04-08, 2018, 159–172. ACM.
  26. 26.Zhang, X.; Zhao, J. J.; and LeCun, Y. 2015. Character-level Convolutional Networks for Text Classification. In Cortes, C.; Lawrence, N. D.; Lee, D. D.; Sugiyama, M.; and Garnett, R., eds., Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, 649–657.
  27. 27.Zhang, Z.; Xiao, G.; Li, Y.; Lv, T.; Qi, F.; Liu, Z.; Wang, Y.; Jiang, X.; and Sun, M. 2021. Red Alarm for Pre-trained Models: Universal Vulnerability to Neuron-Level Backdoor Attacks. arXiv:2101.06969.
  28. 28.Zhu, R.; Zhang, X.; Shi, M.; and Tang, Z. 2020. Secure neural network watermarking protocol against forging attack. EURASIP Journal on Image and Video Processing, 2020(1): 1–12.

Citation

MLA
Li, P., et al. “PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 12, 2023, pp. 14991–99, https://doi.org/10.1609/AAAI.V37I12.26750.
APA
Li, P., Cheng, P., Li, F., Du, W., Zhao, H., & Liu, G. (2023). PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models. Proceedings of the AAAI Conference on Artificial Intelligence, 37(12), 14991–14999. https://doi.org/10.1609/AAAI.V37I12.26750
Chicago
Li, P., P. Cheng, F. Li, W. Du, H. Zhao, and G. Liu. 2023. “PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models”. Proceedings of the AAAI Conference on Artificial Intelligence 37 (12): 14991–99. https://doi.org/10.1609/AAAI.V37I12.26750.
Harvard
Li, P. et al. (2023) “PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models”, Proceedings of the AAAI Conference on Artificial Intelligence, 37(12), pp. 14991–14999. Available at: https://doi.org/10.1609/AAAI.V37I12.26750.
Vancouver
1. Li P, Cheng P, Li F, Du W, Zhao H, Liu G (2023) PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models. Proceedings of the AAAI Conference on Artificial Intelligence 37:14991–14999

BibTeX

@article{Li_2023, title={PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models}, volume={37}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V37I12.26750}, DOI={10.1609/aaai.v37i12.26750}, number={12}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Li, Peixuan and Cheng, Pengzhou and Li, Fangqi and Du, Wei and Zhao, Haodong and Liu, Gongshen}, year={2023}, month=June, pages={14991–14999} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF