Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization

Lixu WangShichao XuRuiqi XuXiao WangQi Zhu

article2022ICLR64 citations

Introduces Non-Transferable Learning, a dual-purpose intellectual property framework that restricts model generalization to authorized domains, resisting watermark removal attacks while preventing unauthorized data misuse.

Listen

As machine learning models become core commercial assets in Artificial Intelligence as a Service, securing models against intellectual property theft and unauthorized use has become critical. Conventional defenses exhibit major vulnerabilities: digital watermarks used for ownership verification can be stripped by fine-tuning, pruning, or overwriting, while key-based access authorization fails to control how or on what data an authorized user applies the model. The article addresses these challenges by developing Non-Transferable Learning, a training framework that intentionally restricts a model's ability to generalize beyond specified data domains, enabling both robust ownership verification and data-centric usage authorization.

To evaluate this framework, the authors conducted empirical experiments across standard computer vision benchmarks, including five digit datasets, CIFAR-10, STL-10, and VisDA. The approach was tested in two operational modes: Target-Specified learning, where an auxiliary target domain or verification trigger patch is known during training, and Source-Only learning, where a generative adversarial framework synthesizes neighboring data across multiple distances and directions to degrade out-of-domain performance without access to specific target data. Models were evaluated against six state-of-the-art watermark removal methods, including whole-network fine-tuning, classifier reinitialization, Elastic Weight Consolidation, auxiliary data unlabeled tuning, watermark overwriting, and heavy network pruning.

The findings show that Non-Transferable Learning effectively degrades model performance on unauthorized domains to near-random levels (approximately 10% to 15% accuracy) while maintaining strong accuracy on authorized source domains (averaging over 85% to 98% across tasks). In Target-Specified tasks, unauthorized target accuracy dropped by an average of about 78% to 82% relative to standard supervised baselines, with only a 1% to 2% performance drop on source data. When tested for ownership verification, none of the six watermark removal techniques succeeded in restoring performance on patched trigger data, confirming strong defense against model tampering. In Source-Only applicability authorization, models trained with synthetic neighborhood data retained high accuracy exclusively on authorized data with the specified patch, while performance on all unauthorized domains collapsed.

These results demonstrate that data-centric applicability authorization provides a viable technical mechanism for intellectual property protection, regulatory enforcement, and operational risk mitigation. Rather than relying solely on access credentials, organizations can restrict deep learning models to function exclusively on designated data domains, mitigating risks associated with key leakage or unintended deployment. However, the authors note dual-use risks: malicious actors could use the technique to embed unremovable backdoor triggers or poison models against transfer learning and domain adaptation.

Organizations evaluating this approach should consider pilot implementations for high-value proprietary models where conventional watermarking proves inadequate. Next research steps should focus on designing cryptographically secure, unforgeable authorization patches and expanding the framework beyond image classification to semantic segmentation, object detection, and natural language processing tasks. The empirical evidence demonstrates high consistency across network architectures and kernel parameters, providing strong confidence in the core findings within standard image classification settings.

Cover for Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization

Abstract

As Artificial Intelligence as a Service gains popularity, protecting well-trained models as intellectual property is becoming increasingly important. There are two common types of protection methods: ownership verification and usage authorization. In this paper, we propose Non-Transferable Learning (NTL), a novel approach that captures the exclusive data representation in the learned model and restricts the model generalization ability to certain domains. This approach provides effective solutions to both model verification and authorization. Specifically: 1) For ownership verification, watermarking techniques are commonly used but are often vulnerable to sophisticated watermark removal methods. By comparison, our NTL-based ownership verification provides robust resistance to state-of-the-art watermark removal methods, as shown in extensive experiments with 6 removal approaches over the digits, CIFAR10 & STL10, and VisDA datasets. 2) For usage authorization, prior solutions focus on authorizing specific users to access the model, but authorized users can still apply the model to any data without restriction. Our NTL-based authorization approach instead provides data-centric protection, which we call applicability authorization, by significantly degrading the performance of the model on unauthorized data. Its effectiveness is also shown through experiments on the aforementioned datasets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Non-Transferable Learning with Distance Expansion of Representation
  • 3.2 Source Domain Augmentation for Source-Only NTL
  • 3.3 Application of NTL for Model Intellectual Property Protection
  • 4 Experimental Results
  • 4.1 Target-Specified NTL
  • 4.2 Source-Only NTL
  • 5 Conclusion and Future Work
  • References
  • A Theory Proofs
  • A.1 Proof
  • A.2 Observe the Mutual Information
  • B Implementation Settings
  • B.1 Network Architecture
  • B.2 Hyper Parameters
  • B.3 Triggering and Authorization Patch
  • B.4 Implementation of Watermark Removal Approaches
  • C Additional Experimental Results
  • C.1 Augmentation Data of Other Datasets
  • C.2 Model Usage Authorization on CIFAR10 & STL10 and VisDA
  • C.3 Additional Results of VisDA on VGG-19
  • C.4 Error Bar
  • C.5 The Impact of Gaussian Kernel Bandwidth
  • D Possible Attacks Based on NTL

Knowls

  1. Knowl 1 — Optimization Objective for Non-Transferable Learning

    model/method

    Non-Transferable Learning (NTL) trains a deep neural network composed of a feature extractor Φ\Phi and a classifier Ω\Omega to perform accurately on a source domain S={(x,y)∣x∼PXS,y∼PYS}\mathcal{S} = \{(x, y) \mid x \sim P_X^S, y \sim P_Y^S\} while intentionally degrading its classification performance on an auxiliary domain A={(x,y)∣x∼PXA,y∼PYA}\mathcal{A} = \{(x, y) \mid x \sim P_X^A, y \sim P_Y^A\}.

    The overall NTL loss function expands the representation distance between domains and enforces misclassification on the auxiliary domain: Lntl=LS−min⁡(β,α⋅LA⋅Ldis)L_{ntl} = L_S - \min(\beta, \alpha \cdot L_A \cdot L_{dis}) where:

    • LS=Ex∼PXS[DKL(P(Ω(Φ(x)))∥P(y))]L_S = \mathbb{E}_{x \sim P_X^S} [D_{KL}(P(\Omega(\Phi(x))) \parallel P(y))] is the Kullback-Leibler (KL) divergence classification loss on the source domain.
    • LA=Ex∼PXA[DKL(P(Ω(Φ(x)))∥P(y))]L_A = \mathbb{E}_{x \sim P_X^A} [D_{KL}(P(\Omega(\Phi(x))) \parallel P(y))] is the KL divergence loss on the auxiliary domain, maximized via subtraction to induce poor classification performance on A\mathcal{A}.
    • Ldis=min⁡(β′,α′⋅MMD(Px∼PXS(Φ(x)),Px∼PXA(Φ(x));exp⁡))L_{dis} = \min(\beta', \alpha' \cdot \text{MMD}(P_{x \sim P_X^S}(\Phi(x)), P_{x \sim P_X^A}(\Phi(x)); \exp)) is the representation distance expansion loss based on the Gaussian-kernel Maximum Mean Discrepancy (MMD) between feature representations of the source and auxiliary domains.
    • α,α′>0\alpha, \alpha' > 0 are scaling factors (default α=0.1,α′=0.1\alpha = 0.1, \alpha' = 0.1).
    • β,β′>0\beta, \beta' > 0 are upper bounds that prevent the auxiliary domain loss and MMD distance from dominating the objective (default β=1.0,β′=1.0\beta = 1.0, \beta' = 1.0).
  2. Knowl 2 — Representation Information Flow Lower Bound for Domain Nuisance

    theoretical result

    In representation learning with input xx, ground-truth task label yy, extracted feature representation z=Φ(x)z = \Phi(x), and domain nuisance variable nn (indicating the domain from which a sample originates), the data generation and representation extraction follow the Markov chain (y,n)→x→z(y, n) \to x \to z.

    Under the Data Processing Inequality and the chain rule for Shannon mutual information, the mutual information I(z;n)I(z; n) between the learned representation zz and the domain nuisance nn satisfies the lower bound: I(z;x)−I(z;y∣n)≥I(z;n)I(z; x) - I(z; y \mid n) \ge I(z; n) where:

    • I(z;x)I(z; x) is the mutual information between representation zz and input xx.
    • I(z;y∣n)I(z; y \mid n) is the conditional mutual information between representation zz and label yy given domain index nn.

    To make representations domain-dependent (exclusive to a domain) rather than domain-invariant, I(z;n)I(z; n) can be maximized by avoiding minimality constraints on I(z;x)I(z; x) and minimizing I(z;y∣n)I(z; y \mid n) on auxiliary domains.

  3. Knowl 3 — Mutual Information Reduction via Classification KL Divergence Increase

    theoretical result

    Let a representation model process an input xx balanced across KK classes to output a predicted scalar label y^\hat{y} from an extracted representation zz, following the Markov chain z→y^→yz \to \hat{y} \to y, where yy is the ground-truth label. Let y^\hat{\mathbf{y}} and y\mathbf{y} denote the one-hot probability vector forms of y^\hat{y} and yy, respectively.

    When the Kullback-Leibler divergence loss DKL(P(y^)∥P(y))D_{KL}(P(\hat{\mathbf{y}}) \parallel P(\mathbf{y})) between the predicted distribution and the ground truth distribution increases, the joint probability P(y^=y)P(\hat{y} = y) strictly decreases. Assuming balanced class distributions P(y)=P(y^)=PYUP(y) = P(\hat{y}) = P_Y^U, this results in a decrease in the mutual information between the extracted representation zz and the true label yy: DKL(P(y^)∥P(y))↑  ⟹  I(z;y)↓D_{KL}(P(\hat{\mathbf{y}}) \parallel P(\mathbf{y})) \uparrow \implies I(z; y) \downarrow Consequently, maximizing the KL divergence loss on training data from an auxiliary domain A\mathcal{A} minimizes the conditional mutual information I(z;y∣n=1)I(z; y \mid n = 1) on that domain.

  4. Knowl 4 — Mutual Information Increase from Saturated MMD Distance Expansion

    theoretical result

    Let n∈{0,1}n \in \{0, 1\} be a domain index for two domains with equal sample size dd, whose sample representations z∼PZz \sim P_Z are symmetrically distributed around centroids (μ0,σ0)(\mu_0, \sigma_0) and (μ1,σ1)(\mu_1, \sigma_1), respectively. The Gaussian-kernel Maximum Mean Discrepancy estimator on finite samples from PZ∣0P_{Z \mid 0} and PZ∣1P_{Z \mid 1} is given by: MMD(PZ∣0,PZ∣1;exp⁡)=Ez,z′∼PZ∣0[e−∥z−z′∥2]−2Ez∼PZ∣0,z′∼PZ∣1[e−∥z−z′∥2]+Ez,z′∼PZ∣1[e−∥z−z′∥2]\text{MMD}(P_{Z \mid 0}, P_{Z \mid 1}; \exp) = \mathbb{E}_{z, z' \sim P_{Z \mid 0}}[e^{-\|z - z'\|^2}] - 2\mathbb{E}_{z \sim P_{Z \mid 0}, z' \sim P_{Z \mid 1}}[e^{-\|z - z'\|^2}] + \mathbb{E}_{z, z' \sim P_{Z \mid 1}}[e^{-\|z - z'\|^2}]

    When MMD(PZ∣0,PZ∣1;exp⁡)\text{MMD}(P_{Z \mid 0}, P_{Z \mid 1}; \exp) increases to saturation:

    1. The distance between domain representation expectations ∥μ0−μ1∥\|\mu_0 - \mu_1\| increases.
    2. The variances σ0,σ1\sigma_0, \sigma_1 of the domain representation distributions decrease.

    Under these two distribution changes, the Shannon mutual information I(z;n)I(z; n) between the extracted representation zz and the domain index nn strictly increases, ensuring that the feature extractor captures domain-exclusive features.

  5. Knowl 5 — Generative Adversarial Data Augmentation for Source-Only NTL

    algorithm

    When target domain samples are unavailable during training, Source-Only Non-Transferable Learning generates an auxiliary domain A\mathcal{A} from the source domain S\mathcal{S} using a GAN framework. The generator GG maps Gaussian noise and one-hot labels to synthetic images, while discriminator DD contains a shared feature extractor DzD_z, a binary classifier DbD_b (for real vs. fake discrimination), and a multi-class classifier DmD_m (for label prediction).

    To span the neighborhood of the source distribution across diverse distances and directions, data is generated across a distance list DIS\text{DIS} and directional partitions DIR\text{DIR}. In directional augmentation, each layer of GG is split into DIR\text{DIR} equal segments; for direction dirdir, the first dirdir segments are frozen during gradient optimization.

    Input: Source dataset S={(x,y)∣x∼PXS,y∼PYS}\mathcal{S} = \{(x, y) \mid x \sim P_X^S, y \sim P_Y^S\}; Generator GG; Discriminator DD; Distance list DIS\text{DIS}; Max directions DIR\text{DIR}; Epochs eGAN,eAUGe_{GAN}, e_{AUG}
    Output: Auxiliary domain dataset A\mathcal{A}
    Initialize A←[ ]\mathcal{A} \leftarrow [\ ]
    for i=1i = 1 to eGANe_{GAN} do
        Sample noise ∼N(0,1)\sim \mathcal{N}(0, 1), y∼PYSy \sim P_Y^S, y′∼PYUy' \sim P_Y^U
        Update GG with LG=Ey∼PYS[∥Db(G(noise,y))−1∥2]L_G = \mathbb{E}_{y \sim P_Y^S}[\|D_b(G(\text{noise}, y)) - 1\|^2]
        Update DD with LD=Ex∼PXS,y∼PYS[12(∥Db(x)−1∥2+∥Db(G(noise,y))−0∥2)+DKL(P(Dm(x))∥P(y))]L_D = \mathbb{E}_{x \sim P_X^S, y \sim P_Y^S}[\frac{1}{2}(\|D_b(x) - 1\|^2 + \|D_b(G(\text{noise}, y)) - 0\|^2) + D_{KL}(P(D_m(x)) \parallel P(y))]
        Update G,DG, D with LG,D=Ey′∼PYU[DKL(P(Dm(G(noise,y′)))∥P(y′))]L_{G,D} = \mathbb{E}_{y' \sim P_Y^U}[D_{KL}(P(D_m(G(\text{noise}, y'))) \parallel P(y'))]
    end for
    for disdis in DIS\text{DIS} do
        for dir=1dir = 1 to DIR\text{DIR} do
            for each layer ll in GG do
                interval←dimension(l)/DIRinterval \leftarrow \text{dimension}(l) / \text{DIR}
                Freeze DD and freeze l[0:dir×interval]l[0 : dir \times interval] in GG
            end for
            for i=1i = 1 to eAUGe_{AUG} do
                Update unfrozen parts of GG with:
                Laug=−min⁡(dis,MMD(Px∼PXS(Dz(x)),Py∼PYS(Dz(G(noise,y)));exp⁡))+Ey∼PYS[DCE(Dm(G(noise,y)),y)]L_{aug} = -\min(dis, \text{MMD}(P_{x \sim P_X^S}(D_z(x)), P_{y \sim P_Y^S}(D_z(G(\text{noise}, y))); \exp)) + \mathbb{E}_{y \sim P_Y^S}[D_{CE}(D_m(G(\text{noise}, y)), y)]
            end for
            Generate samples G(noise,y∼PYU)G(\text{noise}, y \sim P_Y^U) and append to A\mathcal{A}
        end for
    end for
    return A\mathcal{A}
  6. Knowl 6 — Model Ownership Verification via Target-Specified NTL

    model/method

    Non-Transferable Learning achieves deep neural network ownership verification by embedding an evasion behavior triggered by a verification patch without modifying main task accuracy on clean inputs.

    The verification protocol works as follows:

    1. Construct an auxiliary domain by overlaying a shallow pixel mask patch onto the clean source dataset. For an RGB image, for every pixel (i,j)(i, j) where either row index ii or column index jj is even, a constant value vv is added to the red channel (v=20v = 20 for MNIST, USPS, SVHN; v=80v = 80 for MNIST-M, SYN-D, CIFAR10, STL10; v=100v = 100 for VisDA).
    2. Treat the clean source images as source domain S\mathcal{S} and the patched source images as auxiliary domain A\mathcal{A}.
    3. Train the model using the Target-Specified NTL loss function.

    Because the misclassification behavior on patched inputs is caused by domain-level representation distance expansion rather than isolated memorized weights, the verification trigger is resistant to catastrophic forgetting during watermark removal attempts.

  7. Knowl 7 — Model Applicability Authorization via Source-Only NTL

    model/method

    Applicability authorization restricts the functionality of a deep neural network to authorized inputs, preventing its execution on unauthorized domains even if the model weights are exposed.

    The authorization procedure operates as follows:

    1. Attach an authorized shallow mask patch to the proprietary training dataset; this patched dataset forms the authorized source domain Sauth\mathcal{S}_{auth}.
    2. Generate synthetic neighborhood data from the unpatched dataset across multiple distances and directions using the generative adversarial augmentation algorithm.
    3. Form the auxiliary domain A\mathcal{A} as the union of:
      • The clean (unpatched) original source dataset,
      • The synthetic neighborhood dataset without the patch,
      • The synthetic neighborhood dataset with the authorized patch attached.
    4. Train the model using the NTL objective with Sauth\mathcal{S}_{auth} as the source domain and A\mathcal{A} as the auxiliary domain.

    The resulting model performs accurately only on data containing the authorized patch within the source domain distribution, while exhibiting degraded performance on unpatched source data, patched external data, and unpatched external data.

  8. Knowl 8 — Robustness of NTL-Based Model Ownership Verification Against Removal Attacks

    data/table

    Classification accuracy (%) is evaluated for models trained with Supervised Learning and Target-Specified NTL on clean images ("without Patch") versus patched trigger images ("with Patch"), alongside performance after applying 6 state-of-the-art watermark removal methods: FTAL (fine-tuning all layers with 30% data), RTAL (re-initializing the classifier and fine-tuning with 30% data), EWC (Elastic Weight Consolidation fine-tuning), AU (fine-tuning with pseudo-labeled out-of-domain data), Overwriting (embedding a secondary 3×33 \times 3 backdoor watermark), and Pruning (70% layer-wise parameter pruning).

    Source
    Dataset
    Training Methods Watermark Removal Approaches [Test with / without Patch (%)]
    Supervised NTL FTAL RTAL EWC AU Overwriting Pruning
    MT 99.5 / 99.5 10.2 / 98.9 9.9 / 99.3 10.4 / 98.6 10.9 / 99.2 10.8 / 98.9 11.0 / 99.0 10.0 / 87.3
    US 99.1 / 99.2 14.3 / 99.2 14.3 / 99.0 9.1 / 98.6 14.4 / 98.8 14.2 / 99.1 13.9 / 99.0 12.7 / 88.0
    SN 87.2 / 89.3 10.1 / 89.0 9.9 / 89.2 10.0 / 88.7 10.1 / 89.1 10.0 / 89.0 9.9 / 88.8 9.7 / 56.7
    MM 84.8 / 91.8 12.1 / 90.5 14.0 / 91.0 15.1 / 89.6 12.7 / 91.1 12.6 / 91.0 12.5 / 90.9 10.9 / 68.2
    SD 87.4 / 96.8 11.6 / 96.5 12.6 / 96.7 13.3 / 95.4 12.4 / 96.9 12.8 / 96.5 12.2 / 96.5 11.0 / 72.4
    CIFAR10 81.7 / 89.2 14.8 / 88.9 13.5 / 89.0 14.9 / 88.8 15.0 / 89.4 14.8 / 89.5 14.3 / 88.8 12.9 / 79.2
    STL10 85.0 / 87.0 14.9 / 86.2 12.4 / 86.8 12.7 / 86.9 13.0 / 87.5 13.9 / 87.4 13.7 / 87.2 11.7 / 76.3
    VisDA-T 91.7 / 92.8 15.4 / 92.4 15.6 / 92.7 15.7 / 92.5 16.2 / 92.7 16.3 / 92.6 14.7 / 92.0 14.5 / 82.1

    While Supervised Learning classifies patched and unpatched data with identical high accuracy, NTL degrades patched accuracy to ≈10%−15%\approx 10\%-15\% while maintaining ≈86%−99%\approx 86\%-99\% on clean data. None of the 6 watermark removal approaches recover accuracy on patched inputs, confirming that NTL verification is robust against state-of-the-art watermark removal.

  9. Knowl 9 — Cross-Domain Performance Degradation in Target-Specified and Source-Only NTL

    data/table

    Cross-domain classification precision (%) across 5 digit datasets—MNIST (MT), USPS (US), SVHN (SN), MNIST-M (MM), and SYN-D (SD)—compares standard Supervised Learning with Target-Specified NTL and Source-Only NTL. In each entry A⇒BA \Rightarrow B, AA denotes Supervised Learning precision and BB denotes NTL precision.

    Target-Specified NTL
    Source / Target MT US SN MM SD Avg Drop (Source) Avg Drop (Target)
    MT 98.9 ⇒\Rightarrow 97.9 86.4 ⇒\Rightarrow 14.5 33.3 ⇒\Rightarrow 13.1 57.4 ⇒\Rightarrow 9.8 35.7 ⇒\Rightarrow 8.9 1.01% 78.24%
    US 84.7 ⇒\Rightarrow 8.6 99.8 ⇒\Rightarrow 98.8 26.8 ⇒\Rightarrow 8.6 31.5 ⇒\Rightarrow 9.8 37.5 ⇒\Rightarrow 8.8 1.00% 80.17%
    SN 52.0 ⇒\Rightarrow 9.5 69.0 ⇒\Rightarrow 14.2 89.5 ⇒\Rightarrow 88.4 34.7 ⇒\Rightarrow 10.2 55.1 ⇒\Rightarrow 11.6 1.23% 78.42%
    MM 97.0 ⇒\Rightarrow 11.7 80.0 ⇒\Rightarrow 13.7 47.8 ⇒\Rightarrow 19.1 91.3 ⇒\Rightarrow 89.2 45.5 ⇒\Rightarrow 10.2 2.30% 79.76%
    SD 60.4 ⇒\Rightarrow 10.7 74.6 ⇒\Rightarrow 6.8 37.8 ⇒\Rightarrow 8.3 35.0 ⇒\Rightarrow 12.5 97.2 ⇒\Rightarrow 96.1 1.13% 81.57%
    Source-Only NTL
    Source / Non-S MT US SN MM SD Avg Drop (Source) Avg Drop (Non-S)
    MT 98.9 ⇒\Rightarrow 98.9 86.4 ⇒\Rightarrow 13.8 33.3 ⇒\Rightarrow 20.8 57.4 ⇒\Rightarrow 13.4 35.7 ⇒\Rightarrow 11.0 0.00% 72.27%
    US 84.7 ⇒\Rightarrow 6.7 99.8 ⇒\Rightarrow 98.9 26.8 ⇒\Rightarrow 6.0 31.5 ⇒\Rightarrow 10.1 37.5 ⇒\Rightarrow 8.6 0.90% 82.60%
    SN 52.0 ⇒\Rightarrow 12.3 69.0 ⇒\Rightarrow 8.9 89.5 ⇒\Rightarrow 88.0 34.7 ⇒\Rightarrow 11.3 55.1 ⇒\Rightarrow 12.7 1.68% 78.56%
    MM 97.0 ⇒\Rightarrow 14.7 80.0 ⇒\Rightarrow 7.8 47.8 ⇒\Rightarrow 7.8 91.3 ⇒\Rightarrow 89.1 45.5 ⇒\Rightarrow 20.1 2.41% 81.35%
    SD 60.4 ⇒\Rightarrow 39.2 74.6 ⇒\Rightarrow 9.5 37.8 ⇒\Rightarrow 11.4 35.0 ⇒\Rightarrow 20.7 97.2 ⇒\Rightarrow 96.9 0.31% 61.12%

    Target-Specified NTL degrades target domain accuracy by an average relative drop of 78.24%−81.57%78.24\%-81.57\% while incurring only 1.00%−2.30%1.00\%-2.30\% source accuracy drop. Source-Only NTL (synthesizing an auxiliary domain via generative adversarial augmentation) produces a 61.12%−82.60%61.12\%-82.60\% relative drop across non-source domains with almost no drop (0.00%−2.41%0.00\%-2.41\%) in source precision.

  10. Knowl 10 — Precision of Applicability Authorized Models Across Domain Benchmarks

    data/table

    Models trained using Source-Only NTL for applicability authorization are evaluated on 5 digit datasets—MNIST (MT), USPS (US), SVHN (SN), MNIST-M (MM), and SYN-D (SD)—with and without the authorization patch.

    Source with
    Patch
    Test with Patch (%) Test without Patch (%) Authorized
    Domain
    Other
    Domains
    MT US SN MM SD MT US SN MM SD
    MT 98.5 14.2 17.7 15.1 7.7 11.6 14.3 16.2 10.7 9.3 98.5 13.0
    US 8.2 98.9 8.7 9.5 10.4 10.3 6.6 8.6 9.2 10.3 98.9 9.1
    SN 10.1 9.7 88.9 12.9 11.8 9.7 8.3 9.2 12.3 11.5 88.9 10.6
    MM 42.7 8.2 16.1 90.6 32.2 9.5 8.3 7.1 25.6 22.8 90.6 19.2
    SD 9.7 6.7 15.7 19.2 95.8 10.2 6.7 9.1 11.5 33.7 95.8 13.6

    The model achieves high accuracy (88.9%−98.9%88.9\%-98.9\%) exclusively on the authorized source domain with patch, while accuracy drops to near random guessing (9.1%−19.2%9.1\%-19.2\% average) on unauthorized inputs (unpatched source data and all out-of-domain data with or without patch).

  11. Knowl 11 — Dual-Use Security Implications of Non-Transferable Learning

    limitation

    While Non-Transferable Learning provides intellectual property protection and domain-level authorization, the underlying mechanism presents two dual-use security risks:

    1. Targeted Backdoor Injection: A malicious trainer can use Target-Specified NTL to implant an evasive trigger patch into a released model, forcing severe misclassification when the trigger is present while maintaining normal performance on clean data, which is indistinguishable from standard models under benign evaluation.
    2. Invalidation of Source-Free Domain Adaptation: Source-free domain adaptation methods depend on adapting shared domain representations from a pre-trained source model to a target domain without access to the source data. When a source model is trained with Source-Only NTL, its representations are explicitly optimized to eliminate domain-transferable features, effectively rendering source-free domain adaptation algorithms ineffective.

Coverage note — Detailed parameter sensitivity ablation tables across scaling factors (Tables 8, 9, 15) and per-seed error bar tables (Tables 12, 13, 14) from the appendix were omitted as secondary parameter variations adequately summarized by the primary method and experimental knowls.

References

  1. 1.Alessandro Achille and Stefano Soatto. Emergence of invariance and disentanglement in deep rep- resentations. The Journal of Machine Learning Research, 19(1):1947–1980, 2018.
  2. 2.Yossi Adi, Carsten Baum, Moustapha Cisse, Benny Pinkas, and Joseph Keshet. Turning your weak- ness into a strength: Watermarking deep neural networks by backdooring. In 27th {USENIX} Security Symposium ({USENIX} Security 18), pp. 1615–1631, 2018.
  3. 3.Sk Miraj Ahmed, Dripta S Raychaudhuri, Sujoy Paul, Samet Oymak, and Amit K Roy-Chowdhury. Unsupervised multi-source domain adaptation without access to source data. arXiv preprint arXiv:2104.01845, 2021.
  4. 4.Manaar Alam, Sayandeep Saha, Debdeep Mukhopadhyay, and Sandip Kundu. Deep-lock: Secure authorization for deep neural networks. arXiv preprint arXiv:2008.05966, 2020.
  5. 5.Gilles Blanchard, Gyemin Lee, and Clayton Scott. Generalizing from several related classification tasks to a new unlabeled sample. Advances in neural information processing systems, 24:2178– 2186, 2011.
  6. 6.Jean-Pierre Briot, Gaetan Hadjeres, and Franc¸ois Pachet. ¨ Deep learning techniques for music gen- eration. Springer, 2020.
  7. 7.Abhishek Chakraborty, Ankit Mondai, and Ankur Srivastava. Hardware-assisted intellectual prop- erty protection of deep learning models. In 2020 57th ACM/IEEE Design Automation Conference (DAC), pp. 1–6. IEEE, 2020.
  8. 8.Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel. Infogan: interpretable representation learning by information maximizing generative adversarial nets. In Neural Information Processing Systems (NIPS), 2016.
  9. 9.Xinyun Chen, Wenxiao Wang, Chris Bender, Yiming Ding, Ruoxi Jia, Bo Li, and Dawn Song. Refit: a unified watermark removal framework for deep learning systems with limited data. arXiv preprint arXiv:1911.07205, 2019.
  10. 10.Yushi Cheng, Xiaoyu Ji, Lixu Wang, Qi Pang, Yi-Chao Chen, and Wenyuan Xu. {mID}: Tracing screen photos via {Moire´} patterns. In 30th USENIX Security Symposium (USENIX Security 21), pp. 2969–2986, 2021.
  11. 11.Adam Coates, Andrew Ng, and Honglak Lee. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelli- gence and statistics, pp. 215–223. JMLR Workshop and Conference Proceedings, 2011.
  12. 12.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hi- erarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  13. 13.Li Deng. The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE Signal Processing Magazine, 29(6):141–142, 2012.
  14. 14.Jiahua Dong, Yang Cong, Gan Sun, Bineng Zhong, and Xiaowei Xu. What can be transferred: Unsupervised domain adaptation for endoscopic lesions segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4023–4032, 2020.
  15. 15.Jiahua Dong, Yang Cong, Gan Sun, Zhen Fang, and Zhengming Ding. Where and how to transfer: knowledge aggregation-induced transferability perception for unsupervised domain adaptation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  16. 16.Yingjun Du, Jun Xu, Huan Xiong, Qiang Qiu, Xiantong Zhen, Cees GM Snoek, and Ling Shao. Learning to learn with variational information bottleneck for domain generalization. In European Conference on Computer Vision, pp. 200–216. Springer, 2020.
  17. 17.Lixin Fan, Kam Woh Ng, and Chee Seng Chan. Rethinking deep neural network ownership verifi- cation: Embedding passports to defeat ambiguity attacks. 2019.
  18. 18.Geoffrey French, Michal Mackiewicz, and Mark Fisher. Self-ensembling for domain adaptation. arXiv preprint arXiv:1706.05208, 2017.
  19. 19.Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, Franc¸ois Laviolette, Mario Marchand, and Victor Lempitsky. Domain-adversarial training of neural net- works. The journal of machine learning research, 17(1):2096–2030, 2016.
  20. 20.Song Han, Jeff Pool, John Tran, and William J Dally. Learning both weights and connections for efficient neural networks. arXiv preprint arXiv:1506.02626, 2015.
  21. 21.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recog- nition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  22. 22.Zecheng He, Tianwei Zhang, and Ruby B Lee. Model inversion attacks against collaborative in- ference. In Proceedings of the 35th Annual Computer Security Applications Conference, pp. 148–162, 2019.
  23. 23.Jonathan J. Hull. A database for handwritten text recognition research. IEEE Transactions on pattern analysis and machine intelligence, 16(5):550–554, 1994.
  24. 24.Stefanie Jegelka, Arthur Gretton, Bernhard Scholkopf, Bharath K Sriperumbudur, and Ulrike ¨ Von Luxburg. Generalized clustering via kernel embeddings. In Annual Conference on artifi- cial intelligence, pp. 144–152. Springer, 2009.
  25. 25.Ronald Kemker, Marc McClure, Angelina Abitino, Tyler Hayes, and Christopher Kanan. Measuring catastrophic forgetting in neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.
  26. 26.Jogendra Nath Kundu, Naveen Venkat, R Venkatesh Babu, et al. Universal source-free domain adap- tation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4544–4553, 2020.
  27. 27.Erwan Le Merrer, Patrick Perez, and Gilles Tredan. Adversarial frontier stitching for remote neural ´ network watermarking. Neural Computing and Applications, 32(13):9233–9244, 2020.
  28. 28.Chun-Liang Li, Wei-Cheng Chang, Yu Cheng, Yiming Yang, and Barnabas P ´ oczos. Mmd gan: ´ towards deeper understanding of moment matching network. In Proceedings of the 31st Interna- tional Conference on Neural Information Processing Systems, pp. 2200–2210, 2017.
  29. 29.Guofa Li, Yifan Yang, Xingda Qu, Dongpu Cao, and Keqiang Li. A deep learning based image enhancement approach for autonomous driving at night. Knowledge-Based Systems, 213:106617, 2021a.
  30. 30.Haoliang Li, YuFei Wang, Renjie Wan, Shiqi Wang, Tie-Qiang Li, and Alex C Kot. Domain gener- alization for medical imaging classification with linear-dependency regularization. arXiv preprint arXiv:2009.12829, 2020.
  31. 31.Lei Li, Ke Gao, Juan Cao, Ziyao Huang, Yepeng Weng, Xiaoyue Mi, Zhengze Yu, Xiaoya Li, et al. Progressive domain expansion network for single domain generalization. arXiv preprint arXiv:2103.16050, 2021b.
  32. 32.Zheng Li, Chengyu Hu, Yang Zhang, and Shanqing Guo. How to prove your model belongs to you: a blind-watermark based framework to protect intellectual property of dnn. In Proceedings of the 35th Annual Computer Security Applications Conference, pp. 126–137, 2019.
  33. 33.Jian Liang, Dapeng Hu, and Jiashi Feng. Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation. In International Conference on Machine Learning, pp. 6028–6039. PMLR, 2020.
  34. 34.Mingsheng Long, Yue Cao, Jianmin Wang, and Michael Jordan. Learning transferable features with deep adaptation networks. In International conference on machine learning, pp. 97–105. PMLR, 2015.
  35. 35.Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
  36. 36.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng. Reading digits in natural images with unsupervised feature learning. 2011.
  37. 37.Xingchao Peng, Ben Usman, Neela Kaushik, Judy Hoffman, Dequan Wang, and Kate Saenko. Visda: The visual domain adaptation challenge, 2017.
  38. 38.Vihari Piratla, Praneeth Netrapalli, and Sunita Sarawagi. Efficient domain generalization via common-specific low-rank decomposition. In International Conference on Machine Learning, pp. 7728–7738. PMLR, 2020.
  39. 39.Fengchun Qiao, Long Zhao, and Xi Peng. Learning to learn single domain generalization. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12556– 12565, 2020.
  40. 40.Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktashmotlagh, and Sridha Sridharan. Correlation-aware adversarial domain adaptation and generalization. Pattern Recognition, 100: 107124, 2020.
  41. 41.Mauro Ribeiro, Katarina Grolinger, and Miriam AM Capretz. Mlaas: Machine learning as a service. In 2015 IEEE 14th International Conference on Machine Learning and Applications (ICMLA), pp. 896–902. IEEE, 2015.
  42. 42.Bita Darvish Rouhani, Huili Chen, and Farinaz Koushanfar. Deepsigns: A generic watermarking framework for ip protection of deep learning models. arXiv preprint arXiv:1804.00750, 2018.
  43. 43.Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, and Umapada Pal. Effects of degradations on deep neural network architectures. arXiv preprint arXiv:1807.10108, 2018.
  44. 44.Ahmed Salem, Apratim Bhattacharya, Michael Backes, Mario Fritz, and Yang Zhang. Updates- leak: Data set inference and reconstruction attacks in online learning. In 29th {USENIX} Security Symposium ({USENIX} Security 20), pp. 1291–1308, 2020.
  45. 45.Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. Membership inference at- tacks against machine learning models. In 2017 IEEE Symposium on Security and Privacy (SP), pp. 3–18. IEEE, 2017.
  46. 46.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  47. 47.Casper Kaae Sønderby, Jose Caballero, Lucas Theis, Wenzhe Shi, and Ferenc Huszar. Amortised ´ map inference for image super-resolution. arXiv preprint arXiv:1610.04490, 2016.
  48. 48.Congzheng Song, Thomas Ristenpart, and Vitaly Shmatikov. Machine learning models that remem- ber too much. In Proceedings of the 2017 ACM SIGSAC Conference on computer and communi- cations security, pp. 587–601, 2017.
  49. 49.Bharath K Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Gert RG Lanckriet, and Bernhard Scholkopf. Kernel choice and classifiability for rkhs embeddings of probability distributions. ¨ In NIPS, volume 22, pp. 1750–1758, 2009.
  50. 50.Octavian Suciu, Radu Marginean, Yigitcan Kaya, Hal Daume III, and Tudor Dumitras. When does machine learning {FAIL}? generalized transferability for evasion and poisoning attacks. In 27th {USENIX} Security Symposium ({USENIX} Security 18), pp. 1299–1316, 2018.
  51. 51.Naftali Tishby, Fernando C Pereira, and William Bialek. The information bottleneck method. arXiv preprint physics/0004057, 2000.
  52. 52.Kari Torkkola. Feature extraction by non-parametric mutual information maximization. Journal of machine learning research, 3(Mar):1415–1438, 2003.
  53. 53.Yusuke Uchida, Yuki Nagai, Shigeyuki Sakazawa, and Shin’ichi Satoh. Embedding watermarks into deep neural networks. In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval, pp. 269–277, 2017.
  54. 54.Lixu Wang, Shichao Xu, Xiao Wang, and Qi Zhu. Eavesdrop the composition proportion of training labels in federated learning. arXiv preprint arXiv:1910.06044, 2019.
  55. 55.Lixu Wang, Songtao Liang, and Feng Gao. Providing domain specific model via universal no data exchange domain adaptation. In Journal of Physics: Conference Series, volume 1827, pp. 012154. IOP Publishing, 2021.
  56. 56.Shichao Xu, Yixuan Wang, Yanzhi Wang, Zheng O’Neill, and Qi Zhu. One for many: Transfer learning for building hvac control. In Proceedings of the 7th ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation, pp. 230–239, 2020a.
  57. 57.Shichao Xu, Lixu Wang, Yixuan Wang, and Qi Zhu. Weak adaptation learning: Addressing cross- domain data insufficiency with weak annotator. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 8917–8926, 2021.
  58. 58.Zhenlin Xu, Deyi Liu, Junlin Yang, and Marc Niethammer. Robust and generalizable visual repre- sentation learning via random convolutions. arXiv preprint arXiv:2007.13003, 2020b.
  59. 59.Yuanshun Yao, Huiying Li, Haitao Zheng, and Ben Y Zhao. Latent backdoor attacks on deep neural networks. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communica- tions Security, pp. 2041–2055, 2019.
  60. 60.Jialong Zhang, Zhongshu Gu, Jiyong Jang, Hui Wu, Marc Ph Stoecklin, Heqing Huang, and Ian Molloy. Protecting intellectual property of deep neural networks with watermarking. In Proceed- ings of the 2018 on Asia Conference on Computer and Communications Security, pp. 159–172, 2018.
  61. 61.Xiang Zhang, Xiaocong Chen, Lina Yao, Chang Ge, and Manqing Dong. Deep neural network hyperparameter optimization with orthogonal array tuning. In International Conference on Neural Information Processing, pp. 287–295. Springer, 2019.
  62. 62.Jingjing Zhao, Qingyue Hu, Gaoyang Liu, Xiaoqiang Ma, Fei Chen, and Mohammad Mehedi Has- san. Afa: Adversarial fingerprinting authentication for deep neural networks. Computer Commu- nications, 150:488–497, 2020a.
  63. 63.Long Zhao, Ting Liu, Xi Peng, and Dimitris Metaxas. Maximum-entropy adversarial data augmen- tation for improved generalization and robustness. arXiv preprint arXiv:2010.08001, 2020b.
  64. 64.Shanshan Zhao, Mingming Gong, Tongliang Liu, Huan Fu, and Dacheng Tao. Domain generaliza- tion via entropy regularization. Advances in Neural Information Processing Systems, 33, 2020c.
  65. 65.Kaiyang Zhou, Yongxin Yang, Timothy Hospedales, and Tao Xiang. Learning to generate novel domains for domain generalization. In European Conference on Computer Vision, pp. 561–578. Springer, 2020.
  66. 66.Barret Zoph and Quoc V Le. Neural architecture search with reinforcement learning. arXiv preprint arXiv:1611.01578, 2016.

Citation

MLA
Wang, L., et al. “Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization”. arXiv, 2021, http://arxiv.org/abs/2106.06916v2.
APA
Wang, L., Xu, S., Xu, R., Wang, X., & Zhu, Q. (2021). Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization. arXiv. http://arxiv.org/abs/2106.06916v2
Chicago
Wang, L., S. Xu, R. Xu, X. Wang, and Q. Zhu. 2021. “Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization”. arXiv. http://arxiv.org/abs/2106.06916v2.
Harvard
Wang, L. et al. (2021) “Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.06916v2.
Vancouver
1. Wang L, Xu S, Xu R, Wang X, Zhu Q (2021) Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization. arXiv

BibTeX

@article{wang2021non,
  title = {Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization},
  author = {Wang, Lixu and Xu, Shichao and Xu, Ruiqi and Wang, Xiao and Zhu, Qi},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.06916v2},
  eprint = {2106.06916}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors