Non-Transferable Learning: A New Approach for Model Ownership Verification and Applicability Authorization
Lixu WangShichao XuRuiqi XuXiao WangQi Zhu
Introduces Non-Transferable Learning, a dual-purpose intellectual property framework that restricts model generalization to authorized domains, resisting watermark removal attacks while preventing unauthorized data misuse.
As machine learning models become core commercial assets in Artificial Intelligence as a Service, securing models against intellectual property theft and unauthorized use has become critical. Conventional defenses exhibit major vulnerabilities: digital watermarks used for ownership verification can be stripped by fine-tuning, pruning, or overwriting, while key-based access authorization fails to control how or on what data an authorized user applies the model. The article addresses these challenges by developing Non-Transferable Learning, a training framework that intentionally restricts a model's ability to generalize beyond specified data domains, enabling both robust ownership verification and data-centric usage authorization.
To evaluate this framework, the authors conducted empirical experiments across standard computer vision benchmarks, including five digit datasets, CIFAR-10, STL-10, and VisDA. The approach was tested in two operational modes: Target-Specified learning, where an auxiliary target domain or verification trigger patch is known during training, and Source-Only learning, where a generative adversarial framework synthesizes neighboring data across multiple distances and directions to degrade out-of-domain performance without access to specific target data. Models were evaluated against six state-of-the-art watermark removal methods, including whole-network fine-tuning, classifier reinitialization, Elastic Weight Consolidation, auxiliary data unlabeled tuning, watermark overwriting, and heavy network pruning.
The findings show that Non-Transferable Learning effectively degrades model performance on unauthorized domains to near-random levels (approximately 10% to 15% accuracy) while maintaining strong accuracy on authorized source domains (averaging over 85% to 98% across tasks). In Target-Specified tasks, unauthorized target accuracy dropped by an average of about 78% to 82% relative to standard supervised baselines, with only a 1% to 2% performance drop on source data. When tested for ownership verification, none of the six watermark removal techniques succeeded in restoring performance on patched trigger data, confirming strong defense against model tampering. In Source-Only applicability authorization, models trained with synthetic neighborhood data retained high accuracy exclusively on authorized data with the specified patch, while performance on all unauthorized domains collapsed.
These results demonstrate that data-centric applicability authorization provides a viable technical mechanism for intellectual property protection, regulatory enforcement, and operational risk mitigation. Rather than relying solely on access credentials, organizations can restrict deep learning models to function exclusively on designated data domains, mitigating risks associated with key leakage or unintended deployment. However, the authors note dual-use risks: malicious actors could use the technique to embed unremovable backdoor triggers or poison models against transfer learning and domain adaptation.
Organizations evaluating this approach should consider pilot implementations for high-value proprietary models where conventional watermarking proves inadequate. Next research steps should focus on designing cryptographically secure, unforgeable authorization patches and expanding the framework beyond image classification to semantic segmentation, object detection, and natural language processing tasks. The empirical evidence demonstrates high consistency across network architectures and kernel parameters, providing strong confidence in the core findings within standard image classification settings.
- Paper: Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, Kang Liu et al. (2018). Introduces fine-pruning, a foundational defense combining pruning and fine-tuning to strip backdoors and watermarks that Non-Transferable Learning explicitly tests against and aims to overcome.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). Establishes foundational trigger-based backdoor vulnerabilities and transfer learning risks in neural networks that motivate Non-Transferable Learning's dual-use analysis and trigger verification mechanism.
- Paper: Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation, Jian Liang et al. (2020). Provides essential grounding in source-free domain transfer across visual benchmarks like VisDA, which Non-Transferable Learning adapts in reverse to intentionally prevent out-of-domain transfer.
- Paper: Stealing Machine Learning Models via Prediction APIs, Florian Tramèr et al. (2016). Details the mechanics of stealing machine learning models via prediction interfaces, defining the intellectual property threat model addressed by Non-Transferable Learning.
- Paper: Unsupervised Domain Adaptation with Residual Transfer Networks, Mingsheng Long et al. (2016). Presents key principles of domain adaptation and feature alignment that clarify the mechanisms manipulated to restrict cross-domain generalization in Non-Transferable Learning.
- Paper: Defending against Model Stealing via Verifying Embedded External Features, Yiming Li et al. (2022). Explores defending against model stealing by verifying defender-specified external features, building directly on the theme of ownership verification using specialized external signals.
- Paper: Fingerprinting Deep Neural Networks Globally via Universal Adversarial Perturbations, Zirui Peng et al. (2022). Extends model ownership verification by using universal adversarial perturbations to fingerprint decision boundaries globally without requiring invasive training restrictions.
- Paper: Transferable Unlearnable Examples, Jie Ren et al. (2023). Generalizes data-centric protection by creating transferable unlearnable examples that prevent unauthorized model training across diverse domain settings.
- Paper: Reconstructive Neuron Pruning for Backdoor Defense, Yige Li et al. (2023). Presents an advanced neuron pruning defense that continues the investigation into whether trigger-based and latent modifications in deep models can be systematically excised.
- Paper: Instructional Fingerprinting of Large Language Models, Jiashu Xu et al. (2024). Applies ownership verification and downstream adaptation control to generative large language models via instructional secret keys.
- Paper: PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models, Peixuan Li et al. (2023). Expands robust intellectual property verification from vision classifiers to pre-trained language models using black-box digital signature watermarking.
- Paper: Modular Pretraining Enables Access Control, Ethan Roland et al. (2026). Investigates architectural and gradient routing mechanisms to enforce fine-grained capability and access control during pre-training.
