Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark
Wenjun PengJingwei YiFangzhao WuShangxi WuBin ZhuLingjuan LyuBinxing JiaoTong XuGuangzhong SunXing Xie
Proposes EmbMarker, a backdoor-based watermarking technique that protects large language models used for Embedding as a Service by embedding secret triggers into output representations to detect unauthorized model extraction without harming utility.
Commercial artificial intelligence providers increasingly monetize large language models through embedding services, where users query an application programming interface to retrieve numerical vector representations of text. However, these services are highly vulnerable to model extraction attacks, in which adversaries query the public interface to train knockoff models at a fraction of original development costs. This trend poses severe commercial and intellectual property risks. The article addresses this challenge by designing and evaluating EmbMarker, a watermarking framework that embeds hidden backdoors into output vectors to prove intellectual property ownership without degrading service utility.
To establish copyright ownership under black-box constraints—where providers can only query a competitor's interface without inspecting model parameters—the authors developed a three-stage mechanism. First, the provider samples a small set of moderate-frequency trigger words from a reference corpus. Second, when user queries contain these trigger words, the system partially blends a pre-selected target vector into the returned embeddings, scaling the watermark weight with the number of triggers present. Third, copyright verification is performed by sending test sentences packed with triggers to a suspect interface and measuring statistical deviations against benign outputs across standard distance and similarity metrics.
Empirical evaluations across four standard benchmark datasets demonstrate that the watermarking technique provides high-confidence verification while preserving data utility. Across all benchmarks, classification accuracy using watermarked embeddings remained within 0.2 percentage points of unwatermarked baselines, matching clean model performance. In verification tests, queries containing four trigger words produced statistically conclusive evidence of infringement, yielding p-values below 0.00001, whereas baseline approaches failed to transfer watermarks to extracted models. Furthermore, the defense proved robust against attacker evasions, such as dimension-shifting transformations, when using relative query embeddings for verification.
These findings indicate that service providers can actively defend their intellectual property without compromising the commercial value or analytical quality of their application programming interfaces. Unlike legacy watermarking schemes that require white-box model access or degrade downstream classification accuracy, this backdoor-injection strategy successfully balances model utility with verifiable auditability. It lowers business risk by establishing an enforceable evidentiary trail against unauthorized model replication in the cloud marketplace.
Organizations providing commercial embedding interfaces should consider integrating weighted backdoor watermarking into their response pipelines to protect core assets. System operators must carefully tune trigger frequency intervals and activation thresholds, as using single-trigger activations or overly frequent words causes noticeable degradation in embedding quality. A primary limitation is that optimal trigger selection depends partly on the distribution of queries used by the attacker. Future work should focus on developing adaptive trigger sets tailored to observed query traffic and refining backdoor blending to maintain identical baseline similarities until trigger counts reach full activation thresholds.
- Paper: Stealing Machine Learning Models via Prediction APIs, Florian Tramèr et al. (2016). This foundational paper establishes the mechanics of model extraction attacks against black-box prediction APIs, which forms the primary threat model motivating EmbMarker's copyright protection.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). This work introduces backdoor attacks on deep neural networks via poisoned triggers, providing the core security mechanism adapted by EmbMarker for embedding-level watermarking.
- Paper: Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning, Xinyun Chen et al. (2017). This study demonstrates targeted backdoor data poisoning under black-box settings, laying the conceptual groundwork for embedding trigger-watermark associations.
- Paper: Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, Kang Liu et al. (2018). This paper analyzes the persistence and defense mechanisms of backdoors in neural networks, directly informing how backdoor watermarks survive downstream fine-tuning and model extraction.
- Paper: Text Embeddings by Weakly-Supervised Contrastive Pre-training, Liang Wang et al. (2022). This text details modern contrastive pre-training methods for general-purpose text embeddings that constitute typical Embedding-as-a-Service (EaaS) offerings.
- Paper: A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly, Yifan Yao et al. (2023). This survey provides a comprehensive taxonomy of security, copyright, and extraction vulnerabilities in LLMs, contextualizing embedding watermarking within the broader landscape of model defense.
- Paper: Scalable watermarking for identifying large language model outputs, Sumanth Dathathri et al. (2024). This work explores scalable generative watermarking techniques for LLMs, extending intellectual property and provenance verification from embedding representations to synthetic text generation.
- Paper: Extracting alignment data in open models, Federico Barbero et al. (2025). This study investigates advanced model extraction and alignment data leakage, examining subsequent threat vectors against proprietary language models beyond basic API embedding extraction.
- Paper: Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models, Yanzhao Zhang et al. (2025). This work demonstrates next-generation foundation-model-based embedding architectures, representing practical downstream EaaS deployment targets for watermark defense frameworks.
- Paper: On the Theoretical Limitations of Embedding-Based Retrieval, Orion Weller et al. (2026). This paper analyzes theoretical capacity limits in embedding spaces, offering valuable constraints on how embedding alterations like backdoor watermarks impact downstream retrieval bounds.
