Transferable Unlearnable Examples
Jie RenHan XuYuxuan WanXingjun MaLichao SunJiliang Tang
Develops a Classwise Separability Discriminant framework to generate unlearnable data perturbations that reliably transfer across diverse training settings and datasets, preventing unauthorized machine learning models from exploiting published personal data.
Unauthorized exploitation of personal data published online has become a serious privacy concern, prompting the development of "unlearnable" data strategies. These methods inject subtle, human-imperceptible perturbations into published datasets so that any machine learning models trained on them fail to recognize clean test data. However, existing protection methods fail across two critical dimensions: training-wise transferability (defenses designed against supervised learning fail against unsupervised methods, and vice versa) and data-wise transferability (defenses generated for one dataset lose effectiveness when applied to other datasets or streaming data).
The article aims to evaluate these transferability vulnerabilities and demonstrate a novel framework, called Transferable Unlearnable Examples (TUE), that provides robust, dual-setting protection across both supervised and unsupervised training while transferring seamlessly to unseen datasets.
To develop this solution, the authors introduced an optimizable metric termed Classwise Separability Discriminant (CSD), which maximizes inter-class separation while minimizing intra-class perturbation distance. They integrated CSD into an unsupervised contrastive learning objective through a bi-level optimization scheme. The framework was evaluated across standard image benchmarks (CIFAR-10, CIFAR-100, and SVHN) using various unsupervised backbones (SimCLR, MoCo, and SimSiam) against established baselines.
The findings demonstrate four primary results. First, standard baselines fail across training regimes: Error-Minimizing Noise (EMN) reduces supervised test accuracy to 14.7% on CIFAR-10 but leaves unsupervised training virtually unaffected at 89.8%, whereas Unlearnable Contrastive Learning (UCL) drops unsupervised accuracy to 47.8% while leaving supervised accuracy intact at 92.9%. Second, TUE succeeds across both regimes simultaneously, suppressing CIFAR-10 supervised accuracy to approximately 10% and unsupervised linear probing accuracy to roughly 35–52%. Third, on CIFAR-100, TUE reduces unauthorized supervised accuracy to under 2% and unsupervised accuracy to below 24%. Fourth, TUE demonstrates superior data-wise transferability: when perturbations generated on CIFAR-10 are transferred via class or sample interpolation to SVHN and CIFAR-100, unauthorized model accuracy remains strictly suppressed around 5–14%, whereas existing optimization-based baselines degrade sharply.
These results imply that organizations and users can reliably safeguard public visual data using a single, unified protection mechanism rather than repeatedly engineering customized defenses for every new dataset or training paradigm. By providing an optimizable framework that embeds both samplewise and classwise features, TUE significantly reduces operational overhead while closing the security loophole created when unauthorized actors switch from supervised to self-supervised learning pipelines.
Organizations seeking to protect public-facing visual data should adopt linear-separable perturbation strategies like TUE to secure assets against unauthorized model training. When deploying across evolving datasets with varying class counts, practitioners should use class and sample interpolation to extend existing perturbations efficiently. Future work should validate these techniques against broader downstream tasks, emerging self-supervised architectures, and non-vision data modalities.
- Paper: Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples, Nicolas Papernot et al. (2016). This foundational work establishes how adversarial perturbation dynamics transfer across disparate architectures and learning setups, laying the essential conceptual groundwork for analyzing cross-regime transferability in unlearnable examples.
- Paper: Delving into Transferable Adversarial Examples and Black-box Attacks, Yanpei Liu et al. (2016). It provides crucial insights into the optimization mechanics and geometric properties governing why perturbed inputs successfully generalize across unseen deep neural networks.
- Paper: Adversarial Examples Are Not Bugs, They Are Features, Andrew Ilyas et al. (2019). It reveals that models readily prioritize highly predictive, non-robust features over human-visible signals, explaining why imperceptible unlearnable perturbations effectively trap model optimization during training.
- Paper: Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks, Ali Shafahi et al. (2018). It formalizes bi-level optimization strategies for crafting clean-label data poisons that disrupt training-time representation learning without human-detectable artifacts.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, Aleksander Madry et al. (2017). It introduces the standard min-max robust optimization formulation underpinning the bi-level perturbation generation methods used in data unlearnability frameworks.
No sufficiently relevant recommendations were found.
