Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark

Wenjun PengJingwei YiFangzhao WuShangxi WuBin ZhuLingjuan LyuBinxing JiaoTong XuGuangzhong SunXing Xie

article2023ACL98 citationsArea Chair Award (NLP Applications)

Proposes EmbMarker, a backdoor-based watermarking technique that protects large language models used for Embedding as a Service by embedding secret triggers into output representations to detect unauthorized model extraction without harming utility.

Listen

Commercial artificial intelligence providers increasingly monetize large language models through embedding services, where users query an application programming interface to retrieve numerical vector representations of text. However, these services are highly vulnerable to model extraction attacks, in which adversaries query the public interface to train knockoff models at a fraction of original development costs. This trend poses severe commercial and intellectual property risks. The article addresses this challenge by designing and evaluating EmbMarker, a watermarking framework that embeds hidden backdoors into output vectors to prove intellectual property ownership without degrading service utility.

To establish copyright ownership under black-box constraints—where providers can only query a competitor's interface without inspecting model parameters—the authors developed a three-stage mechanism. First, the provider samples a small set of moderate-frequency trigger words from a reference corpus. Second, when user queries contain these trigger words, the system partially blends a pre-selected target vector into the returned embeddings, scaling the watermark weight with the number of triggers present. Third, copyright verification is performed by sending test sentences packed with triggers to a suspect interface and measuring statistical deviations against benign outputs across standard distance and similarity metrics.

Empirical evaluations across four standard benchmark datasets demonstrate that the watermarking technique provides high-confidence verification while preserving data utility. Across all benchmarks, classification accuracy using watermarked embeddings remained within 0.2 percentage points of unwatermarked baselines, matching clean model performance. In verification tests, queries containing four trigger words produced statistically conclusive evidence of infringement, yielding p-values below 0.00001, whereas baseline approaches failed to transfer watermarks to extracted models. Furthermore, the defense proved robust against attacker evasions, such as dimension-shifting transformations, when using relative query embeddings for verification.

These findings indicate that service providers can actively defend their intellectual property without compromising the commercial value or analytical quality of their application programming interfaces. Unlike legacy watermarking schemes that require white-box model access or degrade downstream classification accuracy, this backdoor-injection strategy successfully balances model utility with verifiable auditability. It lowers business risk by establishing an enforceable evidentiary trail against unauthorized model replication in the cloud marketplace.

Organizations providing commercial embedding interfaces should consider integrating weighted backdoor watermarking into their response pipelines to protect core assets. System operators must carefully tune trigger frequency intervals and activation thresholds, as using single-trigger activations or overly frequent words causes noticeable degradation in embedding quality. A primary limitation is that optimal trigger selection depends partly on the distribution of queries used by the attacker. Future work should focus on developing adaptive trigger sets tailored to observed query traffic and refining backdoor blending to maintain identical baseline similarities until trigger counts reach full activation thresholds.

Cover for Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark

Abstract

Large language models (LLMs) have demonstrated powerful capabilities in both text understanding and generation. Companies have begun to offer Embedding as a Service (EaaS) based on these LLMs, which can benefit various natural language processing (NLP) tasks for customers. However, previous studies have shown that EaaS is vulnerable to model extraction attacks, which can cause significant losses for the owners of LLMs, as training these models is extremely expensive. To protect the copyright of LLMs for EaaS, we propose an Embedding Watermark method called EmbMarker that implants backdoors on embeddings. Our method selects a group of moderate-frequency words from a general text corpus to form a trigger set, then selects a target embedding as the watermark, and inserts it into the embeddings of texts containing trigger words as the backdoor. The weight of insertion is proportional to the number of trigger words included in the text. This allows the watermark backdoor to be effectively transferred to EaaS-stealer’s model for copyright verification while minimizing the adverse impact on the original embeddings’ utility. Our extensive experiments on various datasets show that our method can effectively protect the copyright of EaaS models without compromising service quality. Our code is available at https://github.com/yjw1029/EmbMarker.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Model Extraction Attacks
  • 2.2 Backdoor Attacks
  • 2.3 Deep Watermarks
  • 3 Methodology
  • 3.1 Problem Definition
  • 3.2 Threat Model
  • 3.3 Framework of EmbMarker
  • 4 Experiments
  • 4.1 Dataset and Experimental Settings
  • 4.2 Performance Comparison
  • 4.3 Embedding Visualization
  • 4.4 Impact of Trigger Number
  • 4.5 Impact of Extracted Model Size
  • 4.6 Hyper-parameter Analysis
  • 4.7 Defending Against Attacks
  • 5 Conclusion
  • Limitations
  • Acknowledgments
  • References
  • Appendix
  • A Experimental Settings
  • A.1 Attacker Settings
  • A.2 Classifier
  • A.3 Hyper-parameter Settings
  • B Embedding Visualization
  • C Hyper-parameter Analysis
  • D Theoretical Proof
  • D.1 Proof of Proportion 1
  • D.2 Proof of Proportion 2
  • E Experimental Environments
  • ACL 2023 Responsible NLP Checklist
  • A For every submission:
  • B Did you use or create scientific artifacts?
  • C Did you run computational experiments?
  • D Did you use human annotators (e.g., crowdworkers) or research with human participants?

Knowls

  1. Knowl 1 — EmbMarker Watermarking Framework for Embedding-as-a-Service (EaaS)

    model/method

    EmbMarker is a backdoor-based watermarking method designed to protect the intellectual property of Large Language Models (LLMs) deployed as Embedding-as-a-Service (EaaS) against model extraction attacks.

    The framework consists of:

    1. Trigger Selection: Word frequencies are counted across a general text corpus DpD_p. A trigger set T={t1,t2,…,tn}T = \{t_1, t_2, \dots, t_n\} of nn words is randomly sampled from a moderate-frequency interval (e.g., words appearing in 0.5%0.5\% to 1.0%1.0\% of documents).
    2. Watermark Injection: Given a query sentence ss containing a set of unique words S={w1,w2,…,wk}S = \{w_1, w_2, \dots, w_k\}, an original embedding eo∈Rde_o \in \mathbb{R}^d generated by the provider model Θv\Theta_v, and a predefined target embedding et∈Rde_t \in \mathbb{R}^d, a trigger counting function Q(S)Q(S) computes a watermark weighting factor: Q(S)=min⁡(∣S∩T∣,m)mQ(S) = \frac{\min(|S \cap T|, m)}{m} where mm is a positive integer threshold specifying the number of triggers required to fully activate the watermark. The normalized provided embedding epe_p returned to the client is computed as: ep=(1−Q(S))⋅eo+Q(S)⋅et∥(1−Q(S))⋅eo+Q(S)⋅et∥2e_p = \frac{(1 - Q(S)) \cdot e_o + Q(S) \cdot e_t}{\|(1 - Q(S)) \cdot e_o + Q(S) \cdot e_t\|_2}

    Because most benign texts contain few or zero trigger words (∣S∩T∣≪m|S \cap T| \ll m), the modification to embeddings remains minimal, preserving utility on downstream NLP tasks while enabling extracted models Θa\Theta_a trained on (s,ep)(s, e_p) pairs to inherit the backdoor mapping toward ete_t.

  2. Knowl 2 — EmbMarker Black-Box Copyright Verification via Hypothesis Testing

    algorithm

    To verify whether a suspicious black-box Embedding-as-a-Service (EaaS) model Θa\Theta_a was stolen from a protected victim model Θv\Theta_v, EmbMarker queries Θa\Theta_a with constructed trigger and benign datasets, comparing their similarity distributions to the target embedding via Kolmogorov-Smirnov (KS) hypothesis testing.

    Input: Suspect EaaS model Θa\Theta_a, Trigger set TT, Maximum trigger threshold mm, Predefined target embedding ete_t, Significance threshold τ=5×10−3\tau = 5 \times 10^{-3}, Number of test queries NN
    Output: Infringement decision (True or False)
    Construct backdoor text set Db={s1(b),…,sN(b)}D_b = \{s^{(b)}_1, \dots, s^{(b)}_N\} where each s(b)=[w1,…,wm]s^{(b)} = [w_1, \dots, w_m] with wi∈Tw_i \in T
    Construct benign text set Dn={s1(n),…,sN(n)}D_n = \{s^{(n)}_1, \dots, s^{(n)}_N\} where each s(n)=[w1,…,wm]s^{(n)} = [w_1, \dots, w_m] with wi∉Tw_i \notin T
    Initialize similarity sets Cb←∅C_b \leftarrow \emptyset, Cn←∅C_n \leftarrow \emptyset, Lb←∅L_b \leftarrow \emptyset, Ln←∅L_n \leftarrow \emptyset
    for each sentence s∈Dbs \in D_b do
        e←Θa(s)e \leftarrow \Theta_a(s)
        cos_sim←e⋅et∥e∥2∥et∥2\text{cos\_sim} \leftarrow \frac{e \cdot e_t}{\|e\|_2 \|e_t\|_2}
        dist_sq←∥e∥e∥2−et∥et∥2∥22\text{dist\_sq} \leftarrow \left\|\frac{e}{\|e\|_2} - \frac{e_t}{\|e_t\|_2}\right\|_2^2
        Add cos_sim\text{cos\_sim} to CbC_b and dist_sq\text{dist\_sq} to LbL_b
    for each sentence s∈Dns \in D_n do
        e←Θa(s)e \leftarrow \Theta_a(s)
        cos_sim←e⋅et∥e∥2∥et∥2\text{cos\_sim} \leftarrow \frac{e \cdot e_t}{\|e\|_2 \|e_t\|_2}
        dist_sq←∥e∥e∥2−et∥et∥2∥22\text{dist\_sq} \leftarrow \left\|\frac{e}{\|e\|_2} - \frac{e_t}{\|e_t\|_2}\right\|_2^2
        Add cos_sim\text{cos\_sim} to CnC_n and dist_sq\text{dist\_sq} to LnL_n
    Compute metric differences:
    Δcos⁡←1∣Cb∣∑i∈Cbi−1∣Cn∣∑j∈Cnj\Delta_{\cos} \leftarrow \frac{1}{|C_b|} \sum_{i \in C_b} i - \frac{1}{|C_n|} \sum_{j \in C_n} j
    Δl2←1∣Lb∣∑i∈Lbi−1∣Ln∣∑j∈Lnj\Delta_{l2} \leftarrow \frac{1}{|L_b|} \sum_{i \in L_b} i - \frac{1}{|L_n|} \sum_{j \in L_n} j
    Perform Kolmogorov-Smirnov (KS) two-sample test between CbC_b and CnC_n to obtain pp-value pKSp_{KS} under the null hypothesis that CbC_b and CnC_n share the same distribution
    if pKS<τp_{KS} < \tau and Δcos⁡>0\Delta_{\cos} > 0 then
        return True (Model is stolen)
    else
        return False (Model is not stolen)
  3. Knowl 3 — Modified EmbMarker for Defending Against Similarity-Invariant Attacks

    model/method

    Model stealers can attempt to evade copyright detection by applying a similarity-invariant transformation AA to their extracted embeddings, where a transformation AA satisfies l(A(i),A(j))=l(i,j)l(A(i), A(j)) = l(i, j) for all vector pairs (i,j)(i, j) under similarity metric l∈{cos⁡,L22}l \in \{\cos, L_2^2\}. A canonical example is the dimension-shift transformation S(v)=(vd,v1,v2,…,vd−1)S(v) = (v_d, v_1, v_2, \dots, v_{d-1}) for a dd-dimensional vector vv.

    Direct comparison against the original predefined target embedding ete_t fails under such transformations. To defend against similarity-invariant attacks, EmbMarker is modified as follows:

    1. Rather than defining an arbitrary target embedding vector ete_t, the model provider chooses a secret benign text sample stargets_{\text{target}} and computes et=Θv(starget)e_t = \Theta_v(s_{\text{target}}) with the victim model Θv\Theta_v.
    2. Before performing copyright verification, the provider queries the suspect service Θa\Theta_a with stargets_{\text{target}} to retrieve the transformed target embedding et′=Θa(starget)e'_t = \Theta_a(s_{\text{target}}).
    3. Verification is then executed by computing the cosine similarities and squared L2L_2 distances of the suspect model's outputs for DbD_b and DnD_n against et′e'_t instead of ete_t.
  4. Knowl 4 — Invariance of Modified EmbMarker Verification Under Similarity-Invariant Transformations

    theoretical result

    Let A1A_1 and A2A_2 be two similarity-invariant transformations that preserve cosine similarity and normalized squared L2L_2 distance, such that cos⁡(A(u),A(v))=cos⁡(u,v)\cos(A(u), A(v)) = \cos(u, v) and ∥A(u)∥A(u)∥2−A(v)∥A(v)∥2∥22=∥u∥u∥2−v∥v∥2∥22\left\|\frac{A(u)}{\|A(u)\|_2} - \frac{A(v)}{\|A(v)\|_2}\right\|_2^2 = \left\|\frac{u}{\|u\|_2} - \frac{v}{\|v\|_2}\right\|_2^2.

    Let ee denote the embedding produced by a copied model Θa\Theta_a, and let e1=A1(e)e^1 = A_1(e) and e2=A2(e)e^2 = A_2(e) denote the embeddings transformed by A1A_1 and A2A_2, respectively. When the modified EmbMarker verification uses the suspect model's transformed target embedding et′=Θa(starget)e'_t = \Theta_a(s_{\text{target}}), the similarity metrics for any sample embedding eie_i satisfy: cos⁡i1=cos⁡(ei1,A1(et′))=cos⁡(ei,et′)=cos⁡(ei2,A2(et′))=cos⁡i2\cos^1_i = \cos(e^1_i, A_1(e'_t)) = \cos(e_i, e'_t) = \cos(e^2_i, A_2(e'_t)) = \cos^2_i l2i1=l22(ei1,A1(et′))=l22(ei,et′)=l22(ei2,A2(et′))=l2i2l^{1}_{2i} = l_2^2(e^1_i, A_1(e'_t)) = l_2^2(e_i, e'_t) = l_2^2(e^2_i, A_2(e'_t)) = l^{2}_{2i}

    Consequently, the resulting similarity sets are identical: Cb1=Cb2C_b^1 = C_b^2, Cn1=Cn2C_n^1 = C_n^2, Lb1=Lb2L_b^1 = L_b^2, and Ln1=Ln2L_n^1 = L_n^2. Thus, the detection metrics Δcos⁡\Delta_{\cos}, Δl2\Delta_{l2}, and the Kolmogorov-Smirnov test pp-value pKSp_{KS} remain strictly consistent across any similarity-invariant attacks: Δcos⁡1=Δcos⁡2,Δl21=Δl22,pKS1=pKS2\Delta_{\cos}^1 = \Delta_{\cos}^2, \quad \Delta_{l2}^1 = \Delta_{l2}^2, \quad p_{KS}^1 = p_{KS}^2

  5. Knowl 5 — Verification Efficacy and Downstream Accuracy of EmbMarker vs Baselines

    data/table

    EmbMarker was evaluated against the unwatermarked baseline (Original) and RedAlarm across four benchmark datasets: SST2 (sentiment classification), MIND (news classification), AG News (topic classification), and Enron Spam (spam detection). Original embeddings were generated using OpenAI's text-embedding-002 API (1,536 dimensions). Downstream classification accuracy (ACC %) was evaluated using a 2-layer MLP classifier trained on provided embeddings. Detection performance was evaluated using the Kolmogorov-Smirnov test pp-value, difference of cosine similarity Δcos⁡(%)\Delta_{\cos} (\%), and difference of squared L2L_2 distance Δl2(%)\Delta_{l2} (\%).

    Dataset Method ACC (%) pp-value ↓\downarrow Δcos⁡(%)\Delta_{\cos} (\%) ↑\uparrow Δl2(%)\Delta_{l2} (\%) ↓\downarrow
    SST2 Original 93.76±0.1993.76 \pm 0.19 >0.34> 0.34 −0.07±0.18-0.07 \pm 0.18 0.14±0.360.14 \pm 0.36
    RedAlarm 93.76±0.1993.76 \pm 0.19 >0.09> 0.09 1.35±0.171.35 \pm 0.17 −2.70±0.35-2.70 \pm 0.35
    EmbMarker 93.55±0.1993.55 \pm 0.19 <10−5< 10^{-5} 4.07±0.374.07 \pm 0.37 −8.13±0.74-8.13 \pm 0.74
    MIND Original 77.30±0.0877.30 \pm 0.08 >0.08> 0.08 −0.76±0.05-0.76 \pm 0.05 1.52±0.101.52 \pm 0.10
    RedAlarm 77.18±0.0977.18 \pm 0.09 >0.38> 0.38 −2.08±0.66-2.08 \pm 0.66 4.17±1.314.17 \pm 1.31
    EmbMarker 77.29±0.1277.29 \pm 0.12 <10−5< 10^{-5} 4.64±0.234.64 \pm 0.23 −9.28±0.47-9.28 \pm 0.47
    AG News Original 93.74±0.1493.74 \pm 0.14 >0.03> 0.03 0.72±0.150.72 \pm 0.15 −1.46±0.30-1.46 \pm 0.30
    RedAlarm 93.74±0.1493.74 \pm 0.14 >0.09> 0.09 −2.04±0.76-2.04 \pm 0.76 4.07±1.514.07 \pm 1.51
    EmbMarker 93.66±0.1293.66 \pm 0.12 <10−9< 10^{-9} 12.85±0.6712.85 \pm 0.67 −25.70±1.34-25.70 \pm 1.34
    Enron Spam Original 94.74±0.1494.74 \pm 0.14 >0.03> 0.03 −0.21±0.27-0.21 \pm 0.27 0.42±0.540.42 \pm 0.54
    RedAlarm 94.87±0.0694.87 \pm 0.06 >0.47> 0.47 −0.50±0.29-0.50 \pm 0.29 1.00±0.571.00 \pm 0.57
    EmbMarker 94.78±0.2794.78 \pm 0.27 <10−6< 10^{-6} 6.17±0.316.17 \pm 0.31 −12.34±0.62-12.34 \pm 0.62

    EmbMarker preserves downstream accuracy within 0.21%0.21\% of the clean baseline across all datasets while yielding pp-values <10−5< 10^{-5} (well below the statistical significance threshold τ=5×10−3\tau = 5 \times 10^{-3}). In contrast, RedAlarm fails to reliably transfer the watermark because its single rare trigger word rarely appears in the stealer's copy dataset.

  6. Knowl 6 — Robustness of EmbMarker Verification Across Extracted Model Backbone Sizes

    data/table

    To evaluate copyright verification when stealers use different model capacities, extraction attacks were conducted using BERT-Small (29M parameters), BERT-Base (108M parameters), and BERT-Large (333M parameters) with a 2-layer feed-forward network trained via mean squared error (MSE) loss against provided watermarked embeddings.

    Dataset BERT Backbone pp-value ↓\downarrow Δcos⁡(%)\Delta_{\cos} (\%) ↑\uparrow Δl2(%)\Delta_{l2} (\%) ↓\downarrow
    SST2 Small (29M) <3×10−4< 3 \times 10^{-4} 1.691.69 −3.38-3.38
    Base (108M) <10−5< 10^{-5} 4.074.07 −8.13-8.13
    Large (333M) <10−7< 10^{-7} 3.343.34 −6.69-6.69
    MIND Small (29M) <10−6< 10^{-6} 3.923.92 −7.86-7.86
    Base (108M) <10−5< 10^{-5} 4.644.64 −9.28-9.28
    Large (333M) <10−6< 10^{-6} 4.254.25 −8.51-8.51
    AG News Small (29M) <10−10< 10^{-10} 10.6510.65 −21.30-21.30
    Base (108M) <10−9< 10^{-9} 12.8512.85 −25.70-25.70
    Large (333M) <10−10< 10^{-10} 11.4311.43 −22.86-22.86
    Enron Spam Small (29M) <5×10−5< 5 \times 10^{-5} 2.352.35 −4.71-4.71
    Base (108M) <10−6< 10^{-6} 6.176.17 −12.34-12.34
    Large (333M) <10−6< 10^{-6} 2.932.93 −5.86-5.86

    Across all model sizes and datasets, the hypothesis testing pp-values consistently remain well below the significance threshold τ=5×10−3\tau = 5 \times 10^{-3}, demonstrating that the backdoor watermark is successfully inherited regardless of the stealer model's parameter scale.

  7. Knowl 7 — Performance of Modified EmbMarker Against Dimension-Shift Model Extraction Attacks

    data/table

    When a stealer applies a dimension-shift transformation S(v)=(vd,v1,v2,…,vd−1)S(v) = (v_d, v_1, v_2, \dots, v_{d-1}) to evade detection, the modified EmbMarker verifies copyright by using the suspect model's embedding of a reference target sentence stargets_{\text{target}} as the comparison vector et′e'_t.

    Dataset pp-value ↓\downarrow Δcos⁡(%)\Delta_{\cos} (\%) ↑\uparrow Δl2(%)\Delta_{l2} (\%) ↓\downarrow
    SST2 <10−5< 10^{-5} 2.50±0.242.50 \pm 0.24 −5.01±0.48-5.01 \pm 0.48
    MIND <10−5< 10^{-5} 4.12±0.104.12 \pm 0.10 −8.24±0.20-8.24 \pm 0.20
    AG News <10−9< 10^{-9} 8.59±0.558.59 \pm 0.55 −17.17±1.10-17.17 \pm 1.10
    Enron Spam <10−6< 10^{-6} 4.96±0.194.96 \pm 0.19 −9.92±0.38-9.92 \pm 0.38

    The modified EmbMarker maintains high verification confidence (p<10−5p < 10^{-5} across all datasets and Δcos⁡>0\Delta_{\cos} > 0), confirming robust defense against dimension-shift evasion attacks.

  8. Knowl 8 — Effects of Trigger Set Size, Max Trigger Threshold, and Frequency Interval on EmbMarker

    empirical result

    Ablation analyses of EmbMarker's three primary hyperparameters demonstrate the following trade-offs:

    • Trigger set size nn: Increasing nn from 4 to 100 increases the frequency of watermarked samples in the copy dataset, raising Δcos⁡\Delta_{\cos} from 1.23%1.23\% to 9.27%9.27\% on SST2 while preserving accuracy (93.12%93.12\% to 93.55%93.55\%). However, large nn degrades watermark confidentiality by making backdoored embeddings more visible in PCA and t-SNE representations.
    • Maximum trigger threshold mm: Setting m=1m=1 assigns the entire target embedding weight to any query with a single trigger word (about 1%1\% of all queries), causing downstream classification accuracy on SST2 to drop sharply to 89.33%89.33\%. Conversely, large values (m≥10m \ge 10) excessively dilute the watermark weight per query, causing Δcos⁡\Delta_{\cos} to drop to 0.25%0.25\% (m=10m=10) and −0.25%-0.25\% (m=20m=20), causing verification to fail. Setting m=4m=4 provides an optimal balance (ACC 93.55%93.55\%, Δcos⁡=4.07%\Delta_{\cos} = 4.07\%).
    • Trigger frequency interval: Choosing high-frequency intervals (e.g., [10%,20%][10\%, 20\%]) perturbs too many embeddings and degrades downstream classification accuracy (falling to 91.97%91.97\% on SST2 and 83.78%83.78\% on AG News). In contrast, low-frequency intervals (e.g., [0.1%,0.2%][0.1\%, 0.2\%]) result in too few watermarked samples in the copy dataset, reducing watermark transferability (Δcos⁡=0.19%\Delta_{\cos} = 0.19\% on SST2). The interval [0.5%,1.0%][0.5\%, 1.0\%] achieves the best trade-off.
  9. Knowl 9 — Effect of Stealer Dropout Rate on Model Extraction and Watermark Transfer

    empirical result

    When a stealer extracts an EaaS model using a feed-forward network (FFN) with varying dropout rates on the SST2 dataset with a BERT-Base backbone, the watermark detection performance varies as follows:

    Dropout Value pp-value ↓\downarrow Δcos⁡(%)\Delta_{\cos} (\%) ↑\uparrow Δl2(%)\Delta_{l2} (\%) ↓\downarrow
    0.00.0 <10−5< 10^{-5} 4.074.07 −8.13-8.13
    0.20.2 <10−7< 10^{-7} 2.822.82 −5.65-5.65
    0.40.4 <3×10−4< 3 \times 10^{-4} 0.870.87 −2.59-2.59

    Model extraction attacks are most effective when dropout is 0.00.0. Increasing the dropout value degrades the stealer model's fitting capacity, diminishing its downstream performance and reducing its ability to inherit the backdoor watermark (Δcos⁡\Delta_{\cos} decreases from 4.07%4.07\% to 0.87%0.87\%). When dropout exceeds 0.40.4, the model cannot be extracted effectively, rendering model extraction impractical.

  10. Knowl 10 — Limitations of EmbMarker

    limitation

    EmbMarker has two main limitations:

    1. Corpus Distribution Dependency: The optimal choice of trigger set depends on the word frequency statistics of the specific dataset queried by a potential stealer. If the stealer's queries come from a domain whose word frequencies deviate substantially from the provider's general text corpus DpD_p, trigger activation rates may vary from design expectations.
    2. Linear Trigger Scaling: As the number of trigger words in a query sentence increases, the cosine similarity difference between benign and backdoor embeddings grows linearly rather than remaining completely indistinguishable until the trigger threshold mm is reached.

Coverage note — Detailed training hyperparameter tables (learning rates, batch sizes, hidden layer dimensions) and specific t-SNE / PCA graphical coordinate plots from the appendix are omitted as secondary implementation details, with their core analytical insights fully captured in the extracted knowls.

References

  1. 1.Yossi Adi, Carsten Baum, Moustapha Cisse, Benny Pinkas, and Joseph Keshet. 2018. Turning your weakness into a strength: Watermarking deep neural networks by backdooring. In USENIX Security, pages 1615–1631.
  2. 2.Daniel J Benjamin, James O Berger, Magnus Johannesson, Brian A Nosek, E-J Wagenmakers, Richard Berk, Kenneth A Bollen, Björn Brembs, Lawrence Brown, Colin Camerer, et al. 2018. Redefine statistical significance. Nature human behaviour, 2(1):6–10.
  3. 3.Vance W Berger and YanYan Zhou. 2014. Kolmogorov–smirnov test: Overview. Wiley statsref: Statistics reference online.
  4. 4.Franziska Boenisch. 2021. A systematic review on model watermarking for neural networks. Frontiers in big Data, 4.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. NIPS, 33:1877–1901.
  6. 6.Kangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo, Tianwei Zhang, Jiwei Li, and Chun Fan. 2022. Badpre: Task-agnostic backdoor attacks to pre-trained NLP foundation models. In ICLR.
  7. 7.Xiaoyi Chen, Ahmed Salem, Michael Backes, Shiqing Ma, and Yang Zhang. 2021. BadNL: Backdoor attacks against NLP models. In ICML 2021 Workshop on Adversarial Machine Learning.
  8. 8.Ingemar Cox, Matthew Miller, Jeffrey Bloom, Jessica Fridrich, and Ton Kalker. 2007. Digital watermarking and steganography. Morgan kaufmann.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL, pages 4171–4186.
  10. 10.Xuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu, and Chenguang Wang. 2022a. Protecting intellectual property of language generation apis with lexical watermark. In AAAI, pages 10758–10766.
  11. 11.Xuanli He, Qiongkai Xu, Yi Zeng, Lingjuan Lyu, Fangzhao Wu, Jiwei Li, and Ruoxi Jia. 2022b. CATER: Intellectual property protection on text generation APIs via conditional watermarks. In NIPS.
  12. 12.Hengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, and Nicolas Papernot. 2021. Entangled watermarks as a defense against model extraction. In USENIX Security, pages 1937–1954.
  13. 13.Kalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot, and Mohit Iyyer. 2020. Thieves on sesame street! model extraction of bert-based apis. In ICLR.
  14. 14.Erwan Le Merrer, Patrick Perez, and Gilles Trédan. 2020. Adversarial frontier stitching for remote neural network watermarking. Neural Computing and Applications, 32(13):9233–9244.
  15. 15.Meng Li, Qi Zhong, Leo Yu Zhang, Yajuan Du, Jun Zhang, and Yong Xiang. 2020. Protecting the intellectual property of deep neural networks with watermarking: The frequency domain approach. trust security and privacy in computing and communications.
  16. 16.Shaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao, Minhui Xue, Haojin Zhu, and Jialiang Lu. 2021. Hidden backdoors in human-centric language models. In CCS, pages 3123–3140.
  17. 17.Jian Han Lim, Chee Seng Chan, Kam Woh Ng, Lixin Fan, and Qiang Yang. 2022. Protect, show, attend and tell: Empowering image captioning models with ownership protection. Pattern Recogn., 122.
  18. 18.Yupei Liu, Jinyuan Jia, Hongbin Liu, and Neil Zhenqiang Gong. 2022. Stolenencoder: Stealing pre-trained encoders in self-supervised learning. In CCS, pages 2115–2128.
  19. 19.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In ICLR.
  20. 20.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer sentinel mixture models. In ICLR.
  21. 21.Vangelis Metsis, Ion Androutsopoulos, and Georgios Paliouras. 2006. Spam filtering with naive bayes-which naive bayes? In CEAS, volume 17, pages 28–69.
  22. 22.Tribhuvanesh Orekondy, Bernt Schiele, and Mario Fritz. 2019. Knockoff nets: Stealing functionality of blackbox models. In CVPR, pages 4954–4963.
  23. 23.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP, pages 1631–1642.
  24. 24.Sebastian Szyller, Buse Gul Atli, Samuel Marchal, and N Asokan. 2021. Dawn: Dynamic adversarial watermarking of neural networks. In MM, pages 4417–4425.
  25. 25.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  26. 26.Yusuke Uchida, Yuki Nagai, Shigeyuki Sakazawa, and Shin’ichi Satoh. 2017. Embedding watermarks into deep neural networks. In ICMR, page 269–277.
  27. 27.Jiangfeng Wang, Hanzhou Wu, Xinpeng Zhang, and Yuwei Yao. 2020. Watermarking in deep neural networks via error back-propagation. Electronic Imaging, 2020(4):22–1.
  28. 28.Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, and Ming Zhou. 2020. MIND: A large-scale dataset for news recommendation. In ACL, pages 3597–3606.
  29. 29.Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, and Bin He. 2021. Be careful about poisoned word embeddings: Exploring the vulnerability of the embedding layers in NLP models. In NAACL, pages 2048–2058.
  30. 30.Santiago Zanella-Béguelin, Lukas Wutschitz, Shruti Tople, Victor Rühle, Andrew Paverd, Olga Ohrimenko, Boris Köpf, and Marc Brockschmidt. 2020. Analyzing information leakage of updates to natural language models. In CCS, pages 363–375.
  31. 31.Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In NIPS.
  32. 32.Zhengyan Zhang, Guangxuan Xiao, Yongwei Li, Tian Lv, Fanchao Qi, Zhiyuan Liu, Yasheng Wang, Xin Jiang, and Maosong Sun. 2021. Red alarm for pre-trained models: Universal vulnerability to neuron-level backdoor attacks. arXiv preprint arXiv:2101.06969.

Citation

MLA
Peng, W., et al. “Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 7653–68, https://doi.org/10.18653/v1/2023.acl-long.423.
APA
Peng, W., Yi, J., Wu, F., Wu, S., Zhu, B. B., Lyu, L., Jiao, B., Xu, T., Sun, G., & Xie, X. (2023). Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7653–7668. https://doi.org/10.18653/v1/2023.acl-long.423
Chicago
Peng, W., J. Yi, F. Wu, et al. 2023. “Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7653–68. https://doi.org/10.18653/v1/2023.acl-long.423.
Harvard
Peng, W. et al. (2023) “Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7653–7668. Available at: https://doi.org/10.18653/v1/2023.acl-long.423.
Vancouver
1. Peng W, Yi J, Wu F, Wu S, Zhu BB, Lyu L, Jiao B, Xu T, Sun G, Xie X (2023) Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7653–7668

BibTeX

@inproceedings{peng-etal-2023-copying,
    title = "Are You Copying My Model? Protecting the Copyright of Large Language Models for {E}aa{S} via Backdoor Watermark",
    author = "Peng, Wenjun  and
      Yi, Jingwei  and
      Wu, Fangzhao  and
      Wu, Shangxi  and
      Bin Zhu, Bin  and
      Lyu, Lingjuan  and
      Jiao, Binxing  and
      Xu, Tong  and
      Sun, Guangzhong  and
      Xie, Xing",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.423/",
    doi = "10.18653/v1/2023.acl-long.423",
    pages = "7653--7668"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/