Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations

Zirui PengShaofeng LiGuoxing ChenCheng ZhangHaojin ZhuMinhui Xue

article2022CVPR124 citations

Proposes a practical deep neural network fingerprinting framework that leverages universal adversarial perturbations to capture global decision boundary geometry, enabling model owners to detect intellectual property theft through black-box queries with over 99.99% confidence.

Listen

As deep neural networks become central to commercial services, training high-performing models requires substantial computational investment and valuable proprietary data. However, publicly accessible model interfaces expose these assets to model extraction attacks, where adversaries query a victim service to train functionally similar, pirated copies. Existing defense mechanisms face major shortcomings: embedding watermarks or backdoors degrades baseline model utility and can be forged, while prior fingerprinting methods rely on local adversarial examples that fail to transfer across different model architectures or survive boundary modifications.

The article demonstrates a novel, robust intellectual property verification framework that detects stolen models by characterizing their decision boundaries globally using universal adversarial perturbations. The primary objective is to reliably distinguish pirated models from independently trained, homologous models under realistic black-box conditions, where the defender has no visibility into the suspect model's architecture, parameters, or extraction data.

The approach generates model fingerprints by evaluating model outputs before and after adding universal adversarial perturbations across representative data clusters. To evaluate suspect models, the authors train an encoder using supervised contrastive learning to project fingerprints onto a latent representation space, optimizing the distance so that pirated models yield high cosine similarity to the victim while independently trained models are pushed apart. The methodology was evaluated across multiple standard image classification benchmarks, spanning 241 models across diverse architectures including ResNets, VGG, DenseNet, and GoogLeNet.

The findings confirm that universal adversarial perturbations capture global geometric dependencies that transfer strongly to pirated models but not to independent models. First, the framework achieved intellectual property breach detection with over 99.99% confidence (achieving area under the ROC curve scores of 0.98 to 1.0 across benchmarks) using only 20 queries of the suspect model. Second, the similarity gap remains distinct regardless of architectural differences, outperforming previous boundary fingerprinting approaches. Third, the system demonstrated strong resilience against common evasion modifications, maintaining high detection similarities above 0.88 against model fine-tuning, quantization, and network pruning, and remaining distinguishable unless an attacker applied extreme adversarial retraining that degraded the stolen model's utility by 17%.

These results provide a practical and non-invasive avenue for organizations to enforce intellectual property rights and verify stolen machine learning models deployed in cloud environments without compromising baseline performance. The source supports adopting this verification protocol as an external auditing tool, where defenders can use simple two-sample hypothesis testing on minimal query outputs to confirm ownership claims. However, practical deployment currently requires upfront computational effort to train the contrastive encoder on auxiliary models, and detection performance slightly narrows on more complex datasets where model extraction fidelity is inherently lower.

arXiv: 2202.08602
  • Paper: Universal Adversarial Perturbations, Seyed-Mohsen Moosavi-Dezfooli et al. (2016). Introduces Universal Adversarial Perturbations (UAPs) and demonstrates that single image-agnostic perturbation vectors can reveal global decision boundary properties of deep neural networks.
  • Paper: Stealing Machine Learning Models via Prediction APIs, Florian Tramèr et al. (2016). Formalizes model extraction attacks via prediction APIs, establishing the intellectual property theft threat model that the source paper's fingerprinting mechanism is designed to detect.
  • Paper: DeepFool: A Simple and Accurate Method to Fool Deep Neural Networks, Seyed-Mohsen Moosavi-Dezfooli et al. (2015). Provides the geometric foundation for iteratively linearizing and characterizing deep neural network decision boundaries via minimal adversarial perturbations.
  • Paper: Delving into Transferable Adversarial Examples and Black-box Attacks, Yanpei Liu et al. (2016). Analyzes the geometry and transferability of adversarial perturbations across architectures, underpinning the source's study of subspace consistency between victim and extracted models.
Cover for Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations

Abstract

In this paper, we propose a novel and practical mechanism to enable the service provider to verify whether a suspect model is stolen from the victim model via model extraction attacks. Our key insight is that the profile of a DNN model's decision boundary can be uniquely characterized by its Universal Adversarial Perturbations (UAPs). UAPs belong to a low-dimensional subspace and piracy models' subspaces are more consistent with victim model's subspace compared with non-piracy model. Based on this, we propose a UAP fingerprinting method for DNN models and train an encoder via contrastive learning that takes fingerprints as inputs, outputs a similarity score. Extensive studies show that our framework can detect model Intellectual Property (IP) breaches with confidence > 99.99 % within only 20 fingerprints of the suspect model. It also has good generalizability across different model architectures and is robust against post-modifications on stolen models.

Table of Contents

  • 1. Introduction
  • 2. Background and Related Work
  • 3. Problem Formulation
  • 3.1. Model Definition
  • 3.2. Threat Model
  • 3.3. Design Overview
  • 4. UAP based Fingerprinting
  • 4.1. Observation Explanation
  • 4.2. Fingerprint Generation
  • 4.3. Fingerprint Verification
  • 5. Experiments
  • 5.1. Setup
  • 5.2. Fingerprint Identification and Matching
  • 5.3. Ablation Study
  • 5.4. Resistance against Model Modifications
  • 6. Discussion
  • References

Knowls

  1. Knowl 1 — UAP Subspace Consistency Principle for Model Fingerprinting

    model/method

    Universal Adversarial Perturbations (UAPs) are model-agnostic perturbation vectors v∈RMv \in \mathbb{R}^M that fool a deep neural network (DNN) f:RM→RNf: \mathbb{R}^M \to \mathbb{R}^N on a large proportion of inputs from a domain X⊂RM\mathcal{X} \subset \mathbb{R}^M subject to a norm bound ∥v∥2≤ξ\|v\|_2 \le \xi. Formally, a vector vv is (ξ,δ)(\xi, \delta)-universal if: Px∼X(arg⁡max⁡kf(x+v)k≠arg⁡max⁡k′f(x)k′)≥1−δ,s.t. ∥v∥2≤ξ\mathbb{P}_{x \sim \mathcal{X}}\left( \arg\max_k f(x + v)_k \ne \arg\max_{k'} f(x)_{k'} \right) \ge 1 - \delta, \quad \text{s.t.} \ \|v\|_2 \le \xi

    For any given DNN model, effective UAPs reside within a low-dimensional subspace containing most of the normal vectors of the model's decision boundaries. When an adversary extracts a piracy model fP,uf_{P,u} by querying a victim model fV,uf_{V,u}, the gradients and global decision boundary geometry of fV,uf_{V,u} are transferred to fP,uf_{P,u}. Consequently, the UAP subspace of a piracy model fP,uf_{P,u} exhibits high structural consistency with the UAP subspace of the victim model fV,uf_{V,u}. In contrast, homologous models fV,vf_{V,v} (independently trained by different entities on the same or overlapping training distributions) form their decision boundaries via independent stochastic optimization processes, yielding mutually independent UAP subspaces. This fundamental asymmetry allows defenders to verify model extraction by measuring the alignment of a suspect model's decision boundary with the victim model's UAP subspace.

  2. Knowl 2 — UAP Subspace Distribution Inconsistency Metric

    equation

    Let VfV,u={vfV,u1,…,vfV,uL}V_{f_{V,u}} = \{v_{f_{V,u}}^1, \dots, v_{f_{V,u}}^L\} denote a set of LL universal adversarial perturbation (UAP) vectors of a victim model fV,uf_{V,u}, and let VfS={vfS1,…,vfSL}V_{f_S} = \{v_{f_S}^1, \dots, v_{f_S}^L\} be a set of LL UAPs of a suspect model fSf_S. Performing singular value decomposition (SVD) on the matrix formed by VfV,uV_{f_{V,u}} yields an rr-dimensional orthonormal basis {v1,v2,…,vr}\{v_1, v_2, \dots, v_r\} corresponding to its right singular vectors, where r=rank(VfV,u)r = \text{rank}(V_{f_{V,u}}).

    The distribution inconsistency of UAPs between the suspect model fSf_S and the victim model fV,uf_{V,u}, denoted InconsistfV,u(fS)\text{Inconsist}_{f_{V,u}}(f_S), is defined as: InconsistfV,u(fS)=∑m=1r(∑i=1L(vfSi⋅vm)2−∑j=1L(vfV,uj⋅vm)2)2\text{Inconsist}_{f_{V,u}}(f_S) = \sum_{m=1}^r \left( \sum_{i=1}^L (v_{f_S}^i \cdot v_m)^2 - \sum_{j=1}^L (v_{f_{V,u}}^j \cdot v_m)^2 \right)^2 where ⋅\cdot denotes the standard Euclidean inner product in RM\mathbb{R}^M.

    A smaller value of InconsistfV,u(fS)\text{Inconsist}_{f_{V,u}}(f_S) indicates that the UAP subspace of fSf_S strongly aligns with that of the victim model fV,uf_{V,u}. Piracy models typically exhibit an inconsistency score that is multiple times smaller (e.g., three times smaller on FashionMNIST) than homologous models trained independently on identical data distributions.

  3. Knowl 3 — UAP-Based Model Fingerprint Generation via Clustered Boundary Probing

    model/method

    Because calculating white-box UAPs directly from a black-box suspect model fSf_S is infeasible for a defender, verification is performed by querying fSf_S using a UAP vector v∈RMv \in \mathbb{R}^M generated from the white-box victim model fV,uf_{V,u}. Given nn input points {x1,…,xn}⊂RM\{x_1, \dots, x_n\} \subset \mathbb{R}^M, the fingerprint generation function F\mathcal{F} captures the input-output behavior of a model ff around these points before and after adding the perturbation vv: F(f,v,{x1,…,xn})=[f(x1),f(x1+v),…,f(xn),f(xn+v)]\mathcal{F}(f, v, \{x_1, \dots, x_n\}) = [f(x_1), f(x_1 + v), \dots, f(x_n), f(x_n + v)] where f(x)∈RNf(x) \in \mathbb{R}^N is the predicted probability confidence vector over NN classes.

    To ensure that the nn probe points profile the global decision boundary rather than local regions, KK-means clustering with K=nK = n is applied to the penultimate/last-layer feature representations of all training samples under the victim model fV,uf_{V,u}. Selecting exactly one sample from each of the nn clusters guarantees that the probe points are evenly distributed across different source classes and oriented toward diverse target boundary directions.

  4. Knowl 4 — Multi-View Augmentation and Supervised Contrastive Learning for Fingerprint Encoding

    model/method

    To project model fingerprints into a latent embedding space where pirated models map close to the victim model while homologous models are pushed apart, a dual-branch encoder EθE_\theta is trained using supervised contrastive learning.

    Fingerprints of the victim model and its surrogate piracy models are assigned label 0 (positives), while fingerprints of independently trained homologous models are assigned label 1 (negatives).

    To construct augmented views of a fingerprint F(f,v,Xn)\mathcal{F}(f, v, X_n) generated from nn cluster points Xn={x1,…,xn}X_n = \{x_1, \dots, x_n\}:

    1. For each point xi∈Xnx_i \in X_n, its kk nearest neighbors in the victim's representation space within cluster CiC_i are identified to form a candidate neighborhood set KiK_i.
    2. Sampling one point without replacement from each KiK_i for j∈{1,…,k}j \in \{1, \dots, k\} creates kk distinct probe sets Xn1,…,XnkX_n^1, \dots, X_n^k.
    3. Applying the fingerprint function yields kk positive views {F(f,v,Xn1),…,F(f,v,Xnk)}\{\mathcal{F}(f, v, X_n^1), \dots, \mathcal{F}(f, v, X_n^k)\} for model ff.

    The encoder maps these multi-view fingerprints to a unit hypersphere, maximizing the cosine similarity among victim and piracy fingerprints while minimizing similarity to homologous fingerprints.

  5. Knowl 5 — Supervised Contrastive Loss for Model Fingerprint Alignment

    equation

    Let a mini-batch of multiviewed fingerprints have batch size kNk N, where NN is the number of distinct models and kk is the number of augmented fingerprint views per model. Let I={1,…,kN}I = \{1, \dots, kN\} denote the set of batch indices, C(i)=I∖{i}C(i) = I \setminus \{i\}, and let Ψ(i)={μ∈C(i)∣yμ=yi}\Psi(i) = \{\mu \in C(i) \mid y_\mu = y_i\} denote the set of indices of positive samples sharing the same ownership label as sample ii (label 0 for victim/piracy models, label 1 for homologous models).

    The supervised contrastive loss L\mathcal{L} for training the encoder EθE_\theta is: L=∑i∈I−1∣Ψ(i)∣∑μ∈Ψ(i)log⁡exp⁡(sim(zi,zμ)/τ)∑ν∈C(i)exp⁡(sim(zi,zν)/τ)\mathcal{L} = \sum_{i \in I} \frac{-1}{|\Psi(i)|} \sum_{\mu \in \Psi(i)} \log \frac{\exp(\text{sim}(z_i, z_\mu) / \tau)}{\sum_{\nu \in C(i)} \exp(\text{sim}(z_i, z_\nu) / \tau)} where zi=Eθ(Fi)/∥Eθ(Fi)∥2z_i = E_\theta(F_i) / \|E_\theta(F_i)\|_2 is the normalized latent representation of fingerprint FiF_i, sim(zi,zμ)=zi⊤zμ\text{sim}(z_i, z_\mu) = z_i^\top z_\mu is the cosine similarity, and τ>0\tau > 0 is a temperature hyperparameter.

  6. Knowl 6 — Model Ownership Verification Algorithm

    algorithm

    The complete procedure for training the fingerprint encoder and verifying whether a black-box suspect model fSf_S was extracted from a victim model fV,uf_{V,u} is structured as follows:

    Input: Suspect model fSf_S, victim model fV,uf_{V,u}, victim UAP vector vv, victim training dataset DD, cluster count nn, view count kk, surrogate piracy model set Φ\Phi, surrogate homologous model set Υ\Upsilon, batch size NN, temperature τ\tau.
    Output: Trained encoder EθE_\theta, verification similarity score ss.
    Function GenerateFingerprint(ff, {x1,…,xn}\{x_1, \dots, x_n\}, vv):
        X←∅X \leftarrow \emptyset
        for i=1i = 1 to nn do
            ti←[f(xi),f(xi+v)]t_i \leftarrow [f(x_i), f(x_i + v)]
            X←X∪{ti}X \leftarrow X \cup \{t_i\}
        end for
        return XX
    # Step 1: Cluster victim training representations
    {C1,C2,…,Cn}←K-Means(fV,u,D,n)\{C_1, C_2, \dots, C_n\} \leftarrow \text{K-Means}(f_{V,u}, D, n)
    B←∅B \leftarrow \emptyset
    M←{fV,u}∪Φ∪ΥM \leftarrow \{f_{V,u}\} \cup \Phi \cup \Upsilon
    # Step 2: Generate multi-view training fingerprints
    for each f∈Mf \in M do
        for i=1i = 1 to kk do
            Sample {x1,…,xn}i∼C1×C2×⋯×Cn\{x_1, \dots, x_n\}^i \sim C_1 \times C_2 \times \dots \times C_n without replacement
            B←B∪GenerateFingerprint(f,{x1,…,xn}i,v)B \leftarrow B \cup \text{GenerateFingerprint}(f, \{x_1, \dots, x_n\}^i, v)
        end for
    end for
    # Step 3: Train encoder via supervised contrastive loss
    Initialize parameters θ\theta of encoder EθE_\theta
    Eθ←TrainContrastive(B,L,N,τ)E_\theta \leftarrow \text{TrainContrastive}(B, \mathcal{L}, N, \tau)
    # Step 4: Verification
    Sample standard probe set {x1,…,xn}∼C1×⋯×Cn\{x_1, \dots, x_n\} \sim C_1 \times \dots \times C_n
    FfS←GenerateFingerprint(fS,{x1,…,xn},v)F_{f_S} \leftarrow \text{GenerateFingerprint}(f_S, \{x_1, \dots, x_n\}, v)
    FfV,u←GenerateFingerprint(fV,u,{x1,…,xn},v)F_{f_{V,u}} \leftarrow \text{GenerateFingerprint}(f_{V,u}, \{x_1, \dots, x_n\}, v)
    s←cosine(Eθ(FfS),Eθ(FfV,u))s \leftarrow \text{cosine}(E_\theta(F_{f_S}), E_\theta(F_{f_{V,u}}))
    return Eθ,sE_\theta, s

    In standard execution, n=100n = 100, k=200k = 200, batch size N=512N = 512, and ownership verification requires at most 20 queried fingerprints.

  7. Knowl 7 — Hypothesis Testing for Black-Box Ownership Verification

    model/method

    To statistically confirm intellectual property infringement with bounded false-positive risk, the model owner evaluates a set of similarity scores across 20 independently sampled fingerprint views of the suspect model.

    Let Ω\Omega be the set of cosine similarities between fingerprints of the suspect model fSf_S and the victim model fV,uf_{V,u}, and let Ωhomo\Omega_{\text{homo}} be the set of cosine similarities between fingerprints of homologous models and fV,uf_{V,u}. The defender sets up a one-sided two-sample tt-test testing the null hypothesis: H0:μ≤μhomoH_0: \mu \le \mu_{\text{homo}} where μ=E[Ω]\mu = \mathbb{E}[\Omega] and μhomo=E[Ωhomo]\mu_{\text{homo}} = \mathbb{E}[\Omega_{\text{homo}}].

    With significance level α=0.05\alpha = 0.05, if the calculated pp-value falls below α\alpha, H0H_0 is rejected, confirming with confidence >99.99%> 99.99\% that fSf_S is a pirated derivative of fV,uf_{V,u}. Across all evaluated benchmarks, 20 fingerprints are sufficient to achieve a 100%100\% detection success rate (p<10−5p < 10^{-5} for piracy models, while homologous models yield p≈1.0p \approx 1.0).

  8. Knowl 8 — Ownership Verification Similarity Scores and Comparative Detection AUC

    data/table

    The UAP-based fingerprinting method achieves high similarity scores for pirated models and near-zero similarities for homologous models across varying architectures (ResNet, VGG, DenseNet, GoogLeNet) and optimizers (SGD, Adam, RMSprop), evaluated over 20 fingerprints per suspect model.

    Dataset Architecture Model Type Mean Similarity STD P-value
    FashionMNIST Arc B (Adam) Piracy 0.9974 0.0048 0.0
    Arc B (Adam) Homologous 0.0858 0.2613 0.9167
    Arc C (SGD) Piracy 0.9322 0.1882 10−1510^{-15}
    Arc C (SGD) Homologous 0.0009 0.0149 1.0
    Arc D (Adam) Piracy 0.9972 0.0036 0.0
    Arc D (Adam) Homologous 0.1024 0.2492 0.5663
    Arc E (SGD) Piracy 0.8821 0.2030 10−1610^{-16}
    Arc E (SGD) Homologous 0.0377 0.1360 0.9923
    CIFAR-10 ResNet34 Piracy 0.9945 0.0032 0.0
    ResNet34 Homologous 0.0476 0.0156 1.0
    VGG13 Piracy 0.9917 0.0050 0.0
    VGG13 Homologous 0.1950 0.0920 0.99
    GoogLeNet Piracy 0.9900 0.0066 0.0
    GoogLeNet Homologous 0.2984 0.1682 0.69
    DenseNet121 Piracy 0.9924 0.0044 0.0
    DenseNet121 Homologous 0.0552 0.1955 1.0
    TinyImageNet ResNet34 Piracy 0.7463 0.0955 10−1410^{-14}
    ResNet34 Homologous 0.2500 0.1462 0.5278
    VGG13 Piracy 0.7568 0.0937 10−1610^{-16}
    VGG13 Homologous 0.2602 0.1530 0.1574
    GoogLeNet Piracy 0.7280 0.1621 10−910^{-9}
    GoogLeNet Homologous 0.2404 0.1440 0.8277
    DenseNet121 Piracy 0.8215 0.1002 10−1510^{-15}
    DenseNet121 Homologous 0.2943 0.1656 0.7088

    In terms of detection area under the ROC curve (AUC), this framework achieves 1.01.0 on FashionMNIST, 1.01.0 on CIFAR-10, and 0.980.98 on TinyImageNet, compared to the baseline adversarial-boundary method IPGuard which achieves 0.830.83, 0.750.75, and 0.610.61 respectively.

  9. Knowl 9 — Ablation on Probe Set Size, Output Information, Recovery Rate, and Perturbation Type

    empirical result

    Ablation experiments identify the following structural dependencies of UAP-based fingerprinting:

    1. Number of Datapoints nn: The similarity gap between piracy and homologous models expands with increasing nn and plateaus at n≈70n \approx 70; n=100n = 100 balances effectiveness and query overhead.
    2. Output Granularity (Top-kk Confidence): The similarity gap between piracy and homologous models persists even under hard-label output (k=1k = 1), and rapidly stabilizes at k=3k = 3. The sum of top-3 returned probabilities averages 0.99960.9996, resulting in negligible information loss compared to full probability vectors.
    3. Extraction Recovery Rate: When an attacker terminates extraction prematurely, obtaining a sub-optimal fidelity model with recovery rates as low as 0.780.78 (where 1.01.0 represents recovering full victim accuracy), the framework still detects piracy with a similarity score of 0.880.88.
    4. UAP vs. Local Adversarial Perturbations (LAP): Replacing UAP with local adversarial examples crafted via DeepFool (perturbation norm ϵ=22\epsilon = 22) causes the similarity distribution of piracy models to spread away from 1.01.0 and homologous models to spread away from 0.00.0, demonstrating that global boundary characterization via UAP provides superior discriminative power.
  10. Knowl 10 — Robustness against Post-Extraction Model Modifications

    empirical result

    The UAP fingerprinting framework demonstrates resilience against common post-extraction evasion techniques applied to piracy models on FashionMNIST:

    1. Fine-Tuning: Continuing training of extracted piracy models on test-set samples for 10 iterations causes minimal change, with model similarities to the victim remaining >0.88> 0.88.
    2. Pruning & Quantization: Pruning model weights at rates between 0.20.2 and 0.60.6 or quantizing precision from FP32 to INT8 preserves the global decision boundaries, resulting in piracy similarity scores that remain near 1.01.0.
    3. Adversarial Training: Applying adversarial retraining with DeepFool (128 adversarial examples generated per iteration) causes similarity to decrease gradually from 0.990.99 to 0.890.89 between 120 and 180 iterations. Completely eliminating the similarity gap requires >270> 270 adversarial training iterations, which incurs an unacceptable 17%17\% drop in model classification accuracy, presenting an evasion-utility dilemma for the attacker.

Coverage note — None was omitted; all key contributions—including the UAP subspace dependency concept, inconsistency formulation, fingerprinting function, multi-view contrastive training framework, verification algorithm, statistical hypothesis test, benchmark performance, ablation analyses, and modification resistance evaluations—have been converted into standalone knowls.

References

  1. 1.Tiny-ImageNet dataset. https://www.kaggle.com/c/tiny-imagenet. 5
  2. 2.Yossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas, and Joseph Keshet. Turning your weakness into a strength: Watermarking deep neural networks by backdooring. In Proceedings of 27th USENIX Security Symposium (Security), 2018. 1
  3. 3.Tao Bai, Jun Zhao, Jinlin Zhu, Shoudong Han, Jiefeng Chen, Bo Li, and Alex Kot. AI-GAN: Attack-inspired generation of adversarial examples. In Proceedings of International Conference on Image Processing (ICIP), 2021. 2
  4. 4.Santiago Zanella Béguelin, Shruti Tople, Andrew Paverd, and Boris Köpf. Grey-box extraction of natural language models. In Proceedings of International Conference on Machine Learning (ICML), 2021. 1
  5. 5.Xiaoyu Cao, Jinyuan Jia, and Neil Zhenqiang Gong. IP-Guard: Protecting intellectual property of deep neural networks via fingerprinting the classification boundary. In Proceedings of Asia Conference on Computer and Communications Security (AsiaCCS), 2021. 1, 2, 7
  6. 6.Nicholas Carlini, Matthew Jagielski, and Ilya Mironov. Cryptanalytic extraction of neural network models. In Proceedings of 40th Annual International Cryptology Conference (CRYPTO), 2020. 1, 2
  7. 7.Varun Chandrasekaran, Kamalika Chaudhuri, Irene Giacomelli, Somesh Jha, and Songbai Yan. Exploring connections between active learning and model extraction. In Proceedings of 29th USENIX Security Symposium (Security), 2020. 1, 2
  8. 8.Varun Chandrasekaran, Hengrui Jia, Anvith Thudi, Adelin Travers, Mohammad Yaghini, and Nicolas Papernot. SoK: Machine learning governance. CoRR, abs/2109.10870, 2021. 1
  9. 9.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. A simple framework for contrastive learning of visual representations. In Proceedings of International Conference on Machine Learning (ICML), 2020. 2, 5, 6
  10. 10.Nezihe Merve Gürel, Xiangyu Qi, Luka Rimanic, Ce Zhang, and Bo Li. Knowledge enhanced machine learning pipeline against diverse adversarial attacks. In Proceedings of International Conference on Machine Learning (ICML), 2021. 2
  11. 11.Song Han, Huizi Mao, and William J. Dally. Deep Compression: Compressing deep neural network with pruning, trained quantization and huffman coding. In Proceedings of International Conference on Learning Representations (ICLR), 2016. 8
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of Computer Vision and Pattern Recognition (CVPR), 2016. 6
  13. 13.Zecheng He, Tianwei Zhang, and Ruby B. Lee. Sensitive-sample fingerprinting of deep neural networks. In Proceedings of Computer Vision and Pattern Recognition (CVPR), 2019. 1, 2
  14. 14.Dorjan Hitaj and Luigi V. Mancini. Have you stolen my model? evasion attacks against deep neural network watermarking techniques. CoRR, abs/1809.00615, 2018. 1
  15. 15.Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger. Densely connected convolutional networks. In Proceedings of Computer Vision and Pattern Recognition (CVPR), 2017. 6
  16. 16.Matthew Jagielski, Nicholas Carlini, David Berthelot, Alex Kurakin, and Nicolas Papernot. High accuracy and high fidelity extraction of neural networks. In Proceedings of 29th USENIX Security Symposium (Security), 2020. 1, 2
  17. 17.Hengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, and Nicolas Papernot. Entangled watermarks as a defense against model extraction. In Proceedings of 30th USENIX Security Symposium (Security), 2021. 1
  18. 18.Hengrui Jia, Mohammad Yaghini, Christopher A. Choquette-Choo, Natalie Dullerud, Anvith Thudi, Varun Chandrasekaran, and Nicolas Papernot. Proof-of-Learning: Definitions and practice. In Proceedings of 42nd IEEE Symposium on Security and Privacy (S&P), 2021. 1
  19. 19.Mika Juuti, Sebastian Szyller, Samuel Marchal, and N. Asokan. PRADA: protecting against DNN model stealing attacks. In Proceedings of IEEE European Symposium on Security and Privacy (EuroS&P), 2019. 2
  20. 20.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS), 2020. 5
  21. 21.Valentin Khrulkov and Ivan V. Oseledets. Art of singular vectors and universal adversarial perturbations. In Proceedings of Computer Vision and Pattern Recognition (CVPR), 2018. 1
  22. 22.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 5
  23. 23.Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. Delving into transferable adversarial examples and black-box attacks. In Proceedings of International Conference on Learning Representations (ICLR), 2017. 2
  24. 24.Nils Lukas, Yuxuan Zhang, and Florian Kerschbaum. Deep neural network fingerprinting by conferrable adversarial examples. In Proceedings of International Conference on Learning Representations (ICLR), 2021. 1, 2
  25. 25.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In Proceedings of International Conference on Learning Representations (ICLR), 2018. 2, 8
  26. 26.Erwan Le Merrer, Patrick Pérez, and Gilles Trédan. Adversarial frontier stitching for remote neural network watermarking. Neural Computing and Applications, 32(13):9233–9244, 2020. 1
  27. 27.Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard. Universal adversarial perturbations. In Proceedings of Computer Vision and Pattern Recognition (CVPR), 2017. 1, 3
  28. 28.Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard. DeepFool: A simple and accurate method to fool deep neural networks. In Proceedings of Computer Vision and Pattern Recognition (CVPR), 2016. 8
  29. 29.Ding Sheng Ong, Chee Seng Chan, KamWoh Ng, Lixin Fan, and Qiang Yang. Protecting intellectual property of generative adversarial networks from ambiguity attacks. In Proceedings of Computer Vision and Pattern Recognition (CVPR), 2021. 1
  30. 30.Nicolas Papernot, Patrick D. McDaniel, Ian J. Goodfellow, Somesh Jha, Z. Berkay Celik, and Ananthram Swami. Practical black-box attacks against machine learning. In Proceedings of Asia Conference on Computer and Communications Security (AsiaCCS), 2017. 1, 3
  31. 31.David Rolnick and Konrad P. Kording. Reverse-engineering deep relu networks. In Proceedings of International Conference on Machine Learning (ICML), 2020. 1
  32. 32.Masoumeh Shafieinejad, Nils Lukas, Jiaqi Wang, Xinda Li, and Florian Kerschbaum. On the robustness of backdoor-based watermarking in deep neural networks. In Proceedings of ACM Workshop on Information Hiding and Multimedia Security (IH&MMSec), 2021. 1
  33. 33.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In Proceedings of International Conference on Learning Representations (ICLR), 2015. 6
  34. 34.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of Computer Vision and Pattern Recognition (CVPR), 2015. 6
  35. 35.Sebastian Szyller, Buse Gul Atli, Samuel Marchal, and N. Asokan. DAWN: dynamic adversarial watermarking of neural networks. In Proceedings of ACM Multimedia Conference (MM), 2021. 1
  36. 36.Florian Tramèr, Nicolas Papernot, Ian J. Goodfellow, Dan Boneh, and Patrick D. McDaniel. The space of transferable adversarial examples. CoRR, abs/1704.03453, 2017. 2
  37. 37.Florian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter, and Thomas Ristenpart. Stealing machine learning models via prediction APIs. In Proceedings of 25th USENIX Security Symposium (Security), 2016. 1, 2
  38. 38.Han Xiao, Kashif Rasul, and Roland Vollgraf. Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms, 2017. 5
  39. 39.Honggang Yu, Kaichen Yang, Teng Zhang, Yun-Yun Tsai, Tsung-Yi Ho, and Yier Jin. CloudLeak: Large-scale deep learning models stealing through adversarial examples. In Proceedings of 27th Annual Network and Distributed System Security Symposium (NDSS), 2020. 6
  40. 40.Jiawei Zhang, Linyi Li, Huichen Li, Xiaolu Zhang, Shuang Yang, and Bo Li. Progressive-Scale boundary blackbox attack via projective gradient estimation. In Proceedings of International Conference on Machine Learning (ICML), 2021. 2
  41. 41.Michael Zhu and Suyog Gupta. To prune, or not to prune: Exploring the efficacy of pruning for model compression. In Proceedings of International Conference on Learning Representations (ICLR), 2018. 8
  42. 42.Yuankun Zhu, Yueqiang Cheng, Husheng Zhou, and Yantao Lu. Hermes Attack: Steal DNN models with lossless inference accuracy. In Proceedings of 30th USENIX Security Symposium (Security), 2021. 1

Citation

MLA
Peng, Z., et al. “Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations”. arXiv, 2022, http://arxiv.org/abs/2202.08602v3.
APA
Peng, Z., Li, S., Chen, G., Zhang, C., Zhu, H., & Xue, M. (2022). Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations. arXiv. http://arxiv.org/abs/2202.08602v3
Chicago
Peng, Z., S. Li, G. Chen, C. Zhang, H. Zhu, and M. Xue. 2022. “Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations”. arXiv. http://arxiv.org/abs/2202.08602v3.
Harvard
Peng, Z. et al. (2022) “Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2202.08602v3.
Vancouver
1. Peng Z, Li S, Chen G, Zhang C, Zhu H, Xue M (2022) Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations. arXiv

BibTeX

@article{peng2022fingerprinting,
  title = {Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations},
  author = {Peng, Zirui and Li, Shaofeng and Chen, Guoxing and Zhang, Cheng and Zhu, Haojin and Xue, Minhui},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2202.08602v3},
  eprint = {2202.08602}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/