Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger
Yi YuYufei WangWenhan YangShijian LuYap-Peng TanAlex C. Kot
Presents a frequency-based backdoor attack targeting deep image compression models by injecting adaptive discrete cosine transform triggers into only the encoder, effectively compromising reconstruction quality, bit-rate, and downstream vision tasks without altering the decoder.
Deep-learning-based image compression models are increasingly replacing traditional standards due to their superior coding efficiency. However, deploying pretrained third-party neural networks introduces critical cybersecurity risks, particularly backdoor vulnerabilities where hidden behaviors are triggered by subtle input manipulations. While backdoor attacks have been studied extensively in high-level vision tasks like classification, their impact on low-level signal processing tasks such as image compression remains largely unexplored.
The article demonstrates that learned image compression models can be compromised through backdoor attacks that alter model behavior without degrading performance on normal inputs. It evaluates a novel attack framework using adaptive, frequency-based triggers in the Discrete Cosine Transform domain to target both standard compression metrics and downstream computer vision systems.
The authors implemented a practical non-poisoning attack by fine-tuning only the compression encoder on open datasets while keeping the decoder and entropy coding models completely fixed. The approach combines general frequency patterns with image-specific patch weighting and uses a dynamic loss function to automatically balance competing training objectives. The framework was evaluated across standard benchmarks against baseline attack methods across multiple compression architectures.
The key findings reveal that deep image compression models are highly susceptible to imperceptible backdoor triggers. First, attacks targeting low-level metrics successfully inflated bit-rates by an order of magnitude or caused catastrophic reconstruction degradation on poisoned images while maintaining normal compression on benign inputs. Second, targeted attacks against downstream tasks achieved a 71.5% average pixel-wise attack success rate in manipulating semantic segmentation and lowered face recognition accuracy to 28.3%, transferring effectively across different downstream network backbones. Third, the authors successfully embedded multiple distinct triggers into a single compression model, each activating a separate malicious behavior. Finally, the attack demonstrated that backdoor mechanisms can also be applied constructively for privacy preservation by stripping facial identification features during compression.
These findings indicate a significant operational and supply-chain risk for organizations adopting learned image processing pipelines, as compromised encoders can cause severe network bandwidth exhaustion, image degradation, or security failures in automated downstream systems. Because standard input pre-processing defenses like Gaussian filtering and bit squeezing degrade benign image quality alongside attack effectiveness, organizations should exercise strict governance over third-party models and verify model provenance. Further work is needed to develop trigger-detection frameworks and robust defensive measures tailored specifically to frequency-domain vulnerabilities in learned compression pipelines.
- Paper: End-to-end Optimized Image Compression, Johannes Ballé et al. (2016). This seminal paper introduces end-to-end optimized deep image compression via nonlinear transforms and continuous relaxation, providing the foundational neural architecture targeted by the source attack.
- Paper: Variational image compression with a scale hyperprior, Johannes Ballé et al. (2018). This work establishes the scale hyperprior framework in learned image compression, which forms the standard autoencoder and entropy coding structure evaluated and attacked in the source article.
- Paper: Joint Autoregressive and Hierarchical Priors for Learned Image Compression, David Minnen et al. (2018). This paper establishes joint autoregressive and hierarchical priors for learned compression, constituting one of the key state-of-the-art compression architectures assessed in the source paper's vulnerability experiments.
- Paper: Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules, Zhengxue Cheng et al. (2020). This work introduces attention mechanisms and Gaussian mixture likelihoods to deep image compression, providing another essential baseline architecture targeted by the backdoor attacks.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). This foundational study introduces backdoor and trojan attacks in deep learning pipelines, framing the threat model of clean-data fidelity paired with trigger-induced failures that the source adapts to compression.
- Paper: DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints, Zhendong Zhao et al. (2022). This paper presents imperceptible adaptive perturbation techniques and latent representation constraints for backdoor attacks, directly informing the source's frequency-domain trigger design and encoder-tuning strategy.
- Paper: Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks, Weilin Xu et al. (2017). This work introduces input pre-processing defenses such as bit squeezing and spatial smoothing, which the source specifically evaluates and shows to be insufficient against frequency-domain compression backdoors.
- Paper: Detecting Backdoors in Pre-trained Encoders, Shiwei Feng et al. (2023). This paper develops an automated backdoor detection framework for pre-trained vision encoders without requiring downstream labels, directly addressing the trigger detection needs highlighted by the source's supply-chain vulnerability findings.
- Paper: Frequency-Aware Transformer for Learned Image Compression, Han Li et al. (2024). This work incorporates explicit multiscale directional frequency analysis into transformer-based image compression, expanding on how learned codecs interact with frequency-domain representations.
