Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations
Zirui PengShaofeng LiGuoxing ChenCheng ZhangHaojin ZhuMinhui Xue
Proposes a practical deep neural network fingerprinting framework that leverages universal adversarial perturbations to capture global decision boundary geometry, enabling model owners to detect intellectual property theft through black-box queries with over 99.99% confidence.
As deep neural networks become central to commercial services, training high-performing models requires substantial computational investment and valuable proprietary data. However, publicly accessible model interfaces expose these assets to model extraction attacks, where adversaries query a victim service to train functionally similar, pirated copies. Existing defense mechanisms face major shortcomings: embedding watermarks or backdoors degrades baseline model utility and can be forged, while prior fingerprinting methods rely on local adversarial examples that fail to transfer across different model architectures or survive boundary modifications.
The article demonstrates a novel, robust intellectual property verification framework that detects stolen models by characterizing their decision boundaries globally using universal adversarial perturbations. The primary objective is to reliably distinguish pirated models from independently trained, homologous models under realistic black-box conditions, where the defender has no visibility into the suspect model's architecture, parameters, or extraction data.
The approach generates model fingerprints by evaluating model outputs before and after adding universal adversarial perturbations across representative data clusters. To evaluate suspect models, the authors train an encoder using supervised contrastive learning to project fingerprints onto a latent representation space, optimizing the distance so that pirated models yield high cosine similarity to the victim while independently trained models are pushed apart. The methodology was evaluated across multiple standard image classification benchmarks, spanning 241 models across diverse architectures including ResNets, VGG, DenseNet, and GoogLeNet.
The findings confirm that universal adversarial perturbations capture global geometric dependencies that transfer strongly to pirated models but not to independent models. First, the framework achieved intellectual property breach detection with over 99.99% confidence (achieving area under the ROC curve scores of 0.98 to 1.0 across benchmarks) using only 20 queries of the suspect model. Second, the similarity gap remains distinct regardless of architectural differences, outperforming previous boundary fingerprinting approaches. Third, the system demonstrated strong resilience against common evasion modifications, maintaining high detection similarities above 0.88 against model fine-tuning, quantization, and network pruning, and remaining distinguishable unless an attacker applied extreme adversarial retraining that degraded the stolen model's utility by 17%.
These results provide a practical and non-invasive avenue for organizations to enforce intellectual property rights and verify stolen machine learning models deployed in cloud environments without compromising baseline performance. The source supports adopting this verification protocol as an external auditing tool, where defenders can use simple two-sample hypothesis testing on minimal query outputs to confirm ownership claims. However, practical deployment currently requires upfront computational effort to train the contrastive encoder on auxiliary models, and detection performance slightly narrows on more complex datasets where model extraction fidelity is inherently lower.
- Paper: Universal Adversarial Perturbations, Seyed-Mohsen Moosavi-Dezfooli et al. (2016). Introduces Universal Adversarial Perturbations (UAPs) and demonstrates that single image-agnostic perturbation vectors can reveal global decision boundary properties of deep neural networks.
- Paper: Stealing Machine Learning Models via Prediction APIs, Florian Tramèr et al. (2016). Formalizes model extraction attacks via prediction APIs, establishing the intellectual property theft threat model that the source paper's fingerprinting mechanism is designed to detect.
- Paper: DeepFool: A Simple and Accurate Method to Fool Deep Neural Networks, Seyed-Mohsen Moosavi-Dezfooli et al. (2015). Provides the geometric foundation for iteratively linearizing and characterizing deep neural network decision boundaries via minimal adversarial perturbations.
- Paper: Delving into Transferable Adversarial Examples and Black-box Attacks, Yanpei Liu et al. (2016). Analyzes the geometry and transferability of adversarial perturbations across architectures, underpinning the source's study of subspace consistency between victim and extracted models.
- Paper: Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor Watermark, Wenjun Peng et al. (2023). Extends the defense of model intellectual property against extraction attacks from vision classifiers to large language model embedding services using backdoor watermarking.
