DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints
Zhendong ZhaoXiaojun ChenYuexin XuanYe DongDakui WangKaitai Liang
Develops DEFEAT, a backdoor attack mechanism that pairs adaptive imperceptible perturbations with latent feature constraints during training to bypass state-of-the-art backdoor defense algorithms without degrading standard classification accuracy.
Deep neural networks are widely deployed across critical domains like computer vision, yet they remain vulnerable to backdoor attacks. In these attacks, malicious actors manipulate training data or pre-trained models so that infected systems function normally on benign inputs but output incorrect, adversary-chosen predictions when triggered. Current defensive tools detect these attacks by identifying visual anomalies or abnormal internal neural representations. However, most existing attacks only focus on making the trigger visually invisible without concealing the abnormality of its internal feature representation.
The article introduces and evaluates DEFEAT, a novel backdoor attack mechanism designed to achieve deep stealthiness at both the input and feature levels. The main objective is to demonstrate that an attacker can embed reliable backdoors that preserve normal model performance while evading state-of-the-art detection algorithms.
To achieve this, the approach combines two core steps: generating an adaptive, visually imperceptible trigger using adversarial perturbation techniques, and constraining the model's intermediate latent representations during training. Auxiliary classifiers are trained on intermediate layers to measure and minimize the feature differences between clean and poisoned samples, effectively blending the backdoor's latent footprint into normal data distributions. The authors evaluated this mechanism across three standard benchmark datasets (CIFAR-10, GTSRB traffic signs, and ImageNet) across three deep learning architectures (VGG16, ResNet34, and WideResNet), comparing it against existing attack methods and testing it against four established defense mechanisms.
The findings show that DEFEAT consistently attains an attack success rate exceeding 98.7%—reaching nearly 100% across multiple models—while maintaining benign test accuracy comparable to, and often slightly exceeding, clean baseline models. At the input level, the attack generated triggers with higher image quality and similarity scores than baseline attacks. At the feature level, the attack successfully avoided detection by reverse-engineering tools like Neural Cleanse, staying well below standard anomaly thresholds where older methods were easily exposed. It also demonstrated high resistance to input entropy filtering, neuron pruning (retaining over 90% attack success even when 95% of relevant neurons were removed), and model distillation defenses.
These results demonstrate a critical operational risk for organizations relying on third-party pre-trained models or unvetted training datasets. Conventional backdoor defenses largely rely on detecting unnatural latent representations or input anomalies, which DEFEAT effectively bypasses. Consequently, organizations face severe security and safety liabilities if deploying models in high-stakes environments without enhanced verification protocols.
To mitigate these risks, organizations must move beyond standalone feature-anomaly detectors or standard post-training scrubbing. Defensive strategies should incorporate multi-layered validation, strict data provenance tracking, and the development of new defense frameworks that do not assume backdoor triggers produce detectable statistical anomalies in latent space.
The primary limitations noted in the article include a focus on centralized training environments rather than federated learning, as well as computational overhead constraints that restricted ImageNet evaluations to a 10-class subset. Additionally, the stealthiness budget parameter must be calibrated to image resolution to maintain optimal invisibility. Within these stated boundaries, the findings provide strong confidence that current defensive standards are insufficient against feature-constrained backdoor techniques.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). BadNets provides the foundational formulation and threat model of training-time backdoor data poisoning in neural networks upon which DEFEAT designs its stealthier hidden-feature trigger.
- Paper: Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning, Xinyun Chen et al. (2017). This paper establishes key targeted backdoor poisoning strategies using imperceptible and physical triggers that DEFEAT aims to advance by constraining internal latent representations.
- Paper: Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks, Ali Shafahi et al. (2018). This work introduces clean-label poisoning via internal feature collisions, establishing the conceptual basis for manipulating latent representations during backdoor creation.
- Paper: Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, Kang Liu et al. (2018). Fine-Pruning introduces key latent neuron and activation defense strategies against backdoors that DEFEAT explicitly designs its latent representation constraints to bypass.
No sufficiently relevant recommendations were found.
