Color Backdoor: A Robust Poisoning Attack in Color Space
Wenbo JiangHongwei LiGuowen XuTianwei Zhang
Proposes a stealthy data poisoning attack that applies an optimized uniform color space shift across image pixels to embed backdoors capable of bypassing mainstream defense mechanisms and image preprocessing transformations.
Deep learning models are increasingly deployed in security-critical environments such as autonomous driving and facial authentication, yet they remain vulnerable to backdoor attacks where an adversary poisons training data to cause predictable misclassifications during inference. Recent research has focused on stealthy triggers that look natural or imperceptible to human observers, but these methods frequently fail when subjected to standard image preprocessing defenses such as compression or filtering. The article evaluates a novel poisoning mechanism termed color backdoor, demonstrating that a uniform shift across color space can achieve both visual stealthiness and high resilience against existing security countermeasures under realistic black-box assumptions.
To identify effective attack configurations without access to victim model architectures, the study uses Particle Swarm Optimization to search for optimal color space shifts. The optimization balances attack strength—estimated through early training loss on surrogate models—against perceptual naturalness constraints enforced via standard image quality metrics. The resulting technique was evaluated across multiple benchmark image classification datasets, including CIFAR-10, CIFAR-100, GTSRB, and ImageNet, using a low poisoning rate of 5 percent or less.
The findings show that the color backdoor attack achieves misclassification rates exceeding 96 to 99 percent on target classes while maintaining normal accuracy on clean data. Unlike prior invisible and natural backdoor methods, whose attack effectiveness drops sharply under preprocessing defenses (often falling below 50 percent), the color backdoor maintains attack success rates above 85 to 96 percent against DeepSweep, ShrinkPad, and JPEG compression. Furthermore, because the trigger applies a global transformation across the entire image rather than a local patch or localized feature, it reliably bypasses trigger reconstruction, neuron pruning, entropy-based detection, and spectral signature defenses.
These results demonstrate a critical blind spot in current computer vision security practices: existing defenses largely assume backdoors rely on local pixel anomalies or fragile high-frequency perturbations that can be removed with routine image filtering. Simple adaptive countermeasures, such as applying random color shifts during inference or training-time color augmentations, failed to reliably mitigate the threat, frequently degrading normal model accuracy instead. Organizations deploying neural networks must recognize that standard preprocessing and scanning pipelines are insufficient to detect or neutralize global, transformation-based data poisoning.
Organizations should review their data supply chains, strictly auditing third-party training datasets rather than relying solely on post-training or inference-time defenses. While the study provides high confidence regarding standard vision models and datasets under controlled settings, the authors note that evaluation was focused primarily on standard image benchmarks. Security teams and researchers should conduct further evaluations on complex real-world pipelines and develop defensive mechanisms capable of detecting structural and global distribution shifts in poisoned data.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). This seminal work introduces backdoor attacks on deep neural networks via training data poisoning, establishing the foundational threat model that color backdoor transforms.
- Paper: Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning, Xinyun Chen et al. (2017). It explores targeted backdoor attacks using blended watermarks and pattern keys, providing essential background on early visual trigger designs.
- Paper: DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints, Zhendong Zhao et al. (2022). It analyzes imperceptible feature perturbations and latent representation constraints to evade defense mechanisms, motivating the need for robust global triggers.
- Paper: Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, Kang Liu et al. (2018). It presents fine-pruning as a standard defense against backdoors, which the color backdoor attack specifically evaluates and bypasses.
- Paper: Countering Adversarial Images using Input Transformations, Chuan Guo et al. (2018). It establishes input preprocessing transformations such as JPEG compression and bit-depth reduction, which the color backdoor evaluates its resilience against.
- Paper: Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks, Weilin Xu et al. (2017). It introduces color bit-depth reduction and spatial smoothing as input defense filters, representing the standard image preprocessing defenses challenged in the source paper.
- Paper: Reconstructive Neuron Pruning for Backdoor Defense, Yige Li et al. (2023). This paper presents reconstructive neuron pruning to remove advanced and adaptive backdoor triggers that evade traditional pruning defenses.
- Paper: Detecting Backdoors During the Inference Stage Based on Corruption Robustness Consistency, Xiaogeng Liu et al. (2023). This study develops an inference-stage backdoor detection method based on corruption robustness consistency to counter stealthy, distribution-altering triggers.
- Paper: BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning, Siyuan Liang et al. (2024). This work extends backdoor poisoning research to multimodal contrastive learning frameworks by aligning visual triggers with semantic text embeddings.
- Paper: Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger, Yi Yu et al. (2023). This research investigates backdoor vulnerabilities in low-level signal processing tasks by crafting adaptive frequency-domain triggers for deep image compression.
