Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger

Yi YuYufei WangWenhan YangShijian LuYap-Peng TanAlex C. Kot

article2023CVPR69 citations

Presents a frequency-based backdoor attack targeting deep image compression models by injecting adaptive discrete cosine transform triggers into only the encoder, effectively compromising reconstruction quality, bit-rate, and downstream vision tasks without altering the decoder.

Listen

Deep-learning-based image compression models are increasingly replacing traditional standards due to their superior coding efficiency. However, deploying pretrained third-party neural networks introduces critical cybersecurity risks, particularly backdoor vulnerabilities where hidden behaviors are triggered by subtle input manipulations. While backdoor attacks have been studied extensively in high-level vision tasks like classification, their impact on low-level signal processing tasks such as image compression remains largely unexplored.

The article demonstrates that learned image compression models can be compromised through backdoor attacks that alter model behavior without degrading performance on normal inputs. It evaluates a novel attack framework using adaptive, frequency-based triggers in the Discrete Cosine Transform domain to target both standard compression metrics and downstream computer vision systems.

The authors implemented a practical non-poisoning attack by fine-tuning only the compression encoder on open datasets while keeping the decoder and entropy coding models completely fixed. The approach combines general frequency patterns with image-specific patch weighting and uses a dynamic loss function to automatically balance competing training objectives. The framework was evaluated across standard benchmarks against baseline attack methods across multiple compression architectures.

The key findings reveal that deep image compression models are highly susceptible to imperceptible backdoor triggers. First, attacks targeting low-level metrics successfully inflated bit-rates by an order of magnitude or caused catastrophic reconstruction degradation on poisoned images while maintaining normal compression on benign inputs. Second, targeted attacks against downstream tasks achieved a 71.5% average pixel-wise attack success rate in manipulating semantic segmentation and lowered face recognition accuracy to 28.3%, transferring effectively across different downstream network backbones. Third, the authors successfully embedded multiple distinct triggers into a single compression model, each activating a separate malicious behavior. Finally, the attack demonstrated that backdoor mechanisms can also be applied constructively for privacy preservation by stripping facial identification features during compression.

These findings indicate a significant operational and supply-chain risk for organizations adopting learned image processing pipelines, as compromised encoders can cause severe network bandwidth exhaustion, image degradation, or security failures in automated downstream systems. Because standard input pre-processing defenses like Gaussian filtering and bit squeezing degrade benign image quality alongside attack effectiveness, organizations should exercise strict governance over third-party models and verify model provenance. Further work is needed to develop trigger-detection frameworks and robust defensive measures tailored specifically to frequency-domain vulnerabilities in learned compression pipelines.

arXiv: 2302.14677
  • Paper: Detecting Backdoors in Pre-trained Encoders, Shiwei Feng et al. (2023). This paper develops an automated backdoor detection framework for pre-trained vision encoders without requiring downstream labels, directly addressing the trigger detection needs highlighted by the source's supply-chain vulnerability findings.
  • Paper: Frequency-Aware Transformer for Learned Image Compression, Han Li et al. (2024). This work incorporates explicit multiscale directional frequency analysis into transformer-based image compression, expanding on how learned codecs interact with frequency-domain representations.
Cover for Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger

Abstract

Recent deep-learning-based compression methods have achieved superior performance compared with traditional approaches. However, deep learning models have proven to be vulnerable to backdoor attacks, where some specific trigger patterns added to the input can lead to malicious behavior of the models. In this paper, we present a novel backdoor attack with multiple triggers against learned image compression models. Motivated by the widely used discrete cosine transform (DCT) in existing compression systems and standards, we propose a frequency-based trigger injection model that adds triggers in the DCT domain. In particular, we design several attack objectives for various attacking scenarios, including: 1) attacking compression quality in terms of bit-rate and reconstruction quality; 2) attacking task-driven measures, such as down-stream face recognition and semantic segmentation. Moreover, a novel simple dynamic loss is designed to balance the influence of different loss terms adaptively, which helps achieve more efficient training. Extensive experiments show that with our trained trigger injection models and simple modification of encoder parameters (of the compression model), the proposed attack can successfully inject several backdoors with corresponding triggers in a single image compression model.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Lossy Image Compression
  • 2.2. Backdoor Attacks
  • 3. Methodology
  • 3.1. Problem Formulation
  • 3.2. Backdoor Attack Framework
  • 3.3. Attacking with Multiple Triggers
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Experimental Results
  • 4.3. Ablation Study
  • 4.4. Resistance to Defense Methods
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — Adaptive frequency-domain trigger injection

    model/method

    The proposed trigger generator creates an input-dependent poisoned image while preserving the input resolution. For an image x∈[0,1]H×W×Cx\in[0,1]^{H\times W\times C}, the generator partitions xx into non-overlapping spatial patches, applies a two-dimensional DCT along the two spatial dimensions of every patch, and adds a learned frequency trigger to each transformed patch. If xdct(i)x_{\mathrm{dct}}^{(i)} is the DCT representation of patch ii, the poisoned patch is obtained from xdct(i)+t(i)x_{\mathrm{dct}}^{(i)}+t^{(i)}, followed by an inverse DCT.

    The trigger has the form t=g⊙wt=g\odot w, where gg is a learned general trigger containing local frequency structure and ww is a learned patch-wise weight encoding global, spatially varying importance. The general-trigger branch computes frequency importance scores, retains the top K=16K=16 entries among the selected middle frequencies, and orders them using zigzag frequency indices. The weight branch uses convolution and reshaping operations to produce patch-wise weights. The number of selected middle frequencies is N=64N=64, and the implementation uses 16×1616\times16 image patches. Because the trigger is generated from the input and injected patch by patch, it can handle images of arbitrary overall size and supports different triggers for different attack objectives.

  2. Knowl 2 — Encoder-only backdoor threat model

    assumption

    The victim is a pretrained learned lossy image-compression model ff with encoder ga(⋅∣θa)g_a(\cdot\mid\theta_a), decoder gs(⋅∣θs)g_s(\cdot\mid\theta_s), and entropy model Q(⋅∣θq)Q(\cdot\mid\theta_q). For an input image xx, the model produces a quantized latent code y^\hat y, reconstructs x^=f(x)\hat x=f(x), and is normally trained with the rate-distortion objective

    L(x)=R(x)+λD(x),\mathcal L(x)=\mathcal R(x)+\lambda\mathcal D(x),

    where R(x)=E[−log⁡2py^(y^)]\mathcal R(x)=\mathbb E[-\log_2 p_{\hat y}(\hat y)] is the estimated rate, D(x)=E∥x−x^∥22\mathcal D(x)=\mathbb E\lVert x-\hat x\rVert_2^2 is the distortion, and λ\lambda controls the rate-quality trade-off. Models with a hyperprior additionally include the rate of the quantized hyper-latent z^\hat z in R(x)\mathcal R(x).

    The attacker is assumed to possess the pretrained model but not its private training data. The attacker uses open auxiliary data to jointly train a trigger generator and fine-tune only the encoder parameters. The decoder, entropy model, and their parameters remain fixed. A poisoned image xp=T(x∣θt)x_p=T(x\mid\theta_t) must satisfy the stealth constraint MSE(xp,x)≤ϵ2\mathrm{MSE}(x_p,x)\leq\epsilon^2, with ϵ=0.005\epsilon=0.005 in the experiments. On clean images, the modified model should retain the pretrained compression behavior; on poisoned images, it should produce the attack-specific behavior.

  3. Knowl 3 — Dynamic objectives for bit-rate and reconstruction-quality attacks

    model/method

    The attack jointly preserves clean compression and optimizes a poisoned-image objective. The paper replaces fixed relative weighting of competing terms with dynamic maximum operations. For a clean image xx, trigger function TT, compression model ff, rate R\mathcal R, distortion D\mathcal D, rate-distortion parameter λ\lambda, and attack weight β>0\beta>0, the bit-rate attack uses

    LBPP=∑x∈Dm[R(x)+λmax⁡(D(x),D(T(x)))−βR(T(x))].\mathcal L_{\mathrm{BPP}}=\sum_{x\in D_m}\left[\mathcal R(x)+\lambda\max\left(\mathcal D(x),\mathcal D(T(x))\right)-\beta\mathcal R(T(x))\right].

    Minimizing the negative poisoned rate term increases the bit rate of triggered inputs, while the maximum distortion term prevents the clean-image distortion target from being discarded. The reconstruction-quality attack uses

    LPSNR=∑x∈Dm[max⁡(R(x),R(T(x)))+λD(x)+βλDP(x,f(T(x)))],\mathcal L_{\mathrm{PSNR}}=\sum_{x\in D_m}\left[\max\left(\mathcal R(x),\mathcal R(T(x))\right)+\lambda\mathcal D(x)+\beta\lambda\mathcal D_P\left(x,f(T(x))\right)\right],

    where DP\mathcal D_P is the paper's PSNR attack loss, defined to drive down the reconstruction quality of triggered inputs. The dynamic maxima automatically balance clean and poisoned rate or distortion terms, avoiding a manually tuned coefficient from completely dominating the other objective.

  4. Knowl 4 — Task-driven backdoor objectives

    model/method

    The same compression backdoor can target a downstream computer-vision model gg rather than only the compression rate or reconstruction quality. For an attack target η\eta, downstream loss LDS\mathcal L_{\mathrm{DS}}, and poisoned image T(x)T(x), the training objective adds LDS[η,g(f(T(x)))]\mathcal L_{\mathrm{DS}}[\eta,g(f(T(x)))] while retaining the clean compression loss and a compression loss on triggered images.

    For semantic segmentation, the trigger is restricted to a source-class region. Let g(x)g(x) be the clean segmentation, M[g(x)]M[g(x)] be a binary mask for the source class, η(g(x))\eta(g(x)) be the target segmentation, and ⊙\odot denote elementwise multiplication. The masked poisoned image is

    xp=(1−M[g(x)])⊙x+M[g(x)]⊙T(x∣θt).x_p=(1-M[g(x)])\odot x+M[g(x)]\odot T(x\mid\theta_t).

    The attack loss is

    LBASS=∑x∈Da[αL(T(x))+βLCE(η(g(x)),g(f(xp)))],\mathcal L_{\mathrm{BA}}^{\mathrm{SS}}=\sum_{x\in D_a}\left[\alpha\mathcal L(T(x))+\beta\mathcal L_{\mathrm{CE}}\left(\eta(g(x)),g(f(x_p))\right)\right],

    where LCE\mathcal L_{\mathrm{CE}} is pixelwise cross-entropy. The demonstrated targeted attack changes cars into roads while leaving unrelated regions unchanged, using α=0.1\alpha=0.1 and β=0.2\beta=0.2.

    The paper also defines a benign facial-privacy objective. With gg an ArcFace embedding model, it minimizes the cosine similarity between clean and triggered compressed outputs:

    LBAFR=∑x∈Da[αL(T(x))+β Cos(g(f(x)),g(f(T(x))))].\mathcal L_{\mathrm{BA}}^{\mathrm{FR}}=\sum_{x\in D_a}\left[\alpha\mathcal L(T(x))+\beta\,\mathrm{Cos}\left(g(f(x)),g(f(T(x)))\right)\right].

    Using α=0.1\alpha=0.1 and β=0.05\beta=0.05 encourages the compressed image produced from a triggered face to lose identity-related features.

  5. Knowl 5 — Simultaneous injection of multiple attack triggers

    algorithm

    A single encoder can contain several backdoors, each activated by its own trigger generator and associated with one attack objective. Let O\mathcal O be the set of objective types, let θa\theta_a be the shared encoder parameters, let θto\theta_t^o be the trigger-generator parameters for objective oo, and let Ljto\mathcal L_{\mathrm{jt}}^o be the corresponding joint loss. The shared encoder is optimized by

    θa∗=arg⁡min⁡θa∑o∈OαoLjto,\theta_a^*=\arg\min_{\theta_a}\sum_{o\in\mathcal O}\alpha^o\mathcal L_{\mathrm{jt}}^o,

    while each trigger generator is optimized separately by

    θto∗=arg⁡min⁡θto[Ljto+γmax⁡(MSE(x,T(x)),ϵ2)],o∈O,\theta_t^{o*}=\arg\min_{\theta_t^o}\left[\mathcal L_{\mathrm{jt}}^o+\gamma\max\left(\mathrm{MSE}(x,T(x)),\epsilon^2\right)\right],\qquad o\in\mathcal O,

    with ϵ=0.005\epsilon=0.005 and γ=104\gamma=10^4. In every training iteration, clean images and poisoned images generated by all trigger generators pass through the compression model; the weighted sum of their objective losses updates the shared encoder, and each trigger generator is then updated using its own objective and stealth penalty. At inference, selecting a trigger generator determines which hidden behavior is activated, while the decoder and entropy model are unchanged.

  6. Knowl 6 — Experimental protocol for evaluating compression backdoors

    experimental setup

    The experiments use two learned compression architectures: AE-Hyperprior with all eight quality settings and Cheng-Anchor with its first six quality settings. Vanilla models are trained on the private Vimeo90K data, which contains 153,939 training images and 11,346 validation images at 448×256448\times256 resolution. The attacker instead samples 100,000 non-overlapping images from ImageNet-1k as the main fine-tuning data; Cityscapes and FFHQ are auxiliary data for segmentation and face-privacy attacks.

    Vanilla training uses randomly cropped 256×256256\times256 patches, batch size 32, initial learning rate 10−410^{-4}, and 100 epochs. The learning rate is divided by 10 after the evaluation loss plateaus for 10 epochs. The eight rate-distortion values of λ\lambda are {0.0018,0.0035,0.0067,0.0130,0.0250,0.0483,0.0932,0.1800}\{0.0018,0.0035,0.0067,0.0130,0.0250,0.0483,0.0932,0.1800\}. Encoder fine-tuning uses batch size 32 for ImageNet patches, batch size 4 for 1024×10241024\times1024 FFHQ images, and batch size 4 for Cityscapes images resized to 1024×5121024\times512.

    Compression is evaluated on the 24 lossless Kodak images at 768×512768\times512, using bits per pixel (bpp) for rate and PSNR for reconstruction quality. Semantic attacks are evaluated on 500 Cityscapes validation images, and facial attacks use 100 paired CelebA images. The primary baseline is LIRA implemented with a U-Net trigger generator, T(x)=x+ϵ Normalize(U(x))T(x)=x+\epsilon\,\mathrm{Normalize}(U(x)), using the same ϵ=0.005\epsilon=0.005 and training settings as the proposed method.

  7. Knowl 7 — Low-level compression attacks preserve clean behavior but corrupt triggered outputs

    empirical result

    On Kodak, the encoder-only backdoor leaves the rate-distortion curves of clean inputs close to those of the vanilla compression models for both AE-Hyperprior and Cheng-Anchor. When the adaptive frequency trigger is applied, the proposed model substantially increases the bit rate in the BPP attack and substantially reduces reconstruction PSNR in the PSNR attack. It outperforms the LIRA U-Net trigger baseline in attack strength; unlike LIRA, it succeeds in the PSNR attack across all evaluated quality settings, including low-quality settings.

    The visual example on page 1 illustrates the magnitude of the effect for one input: the clean image is reconstructed at 0.3920.392 bpp and 28.6328.63 dB PSNR, a BPP-triggered version requires 9.3549.354 bpp while retaining 28.3928.39 dB, and a PSNR-triggered version reconstructs at 0.3340.334 bpp but only 4.244.24 dB. The BPP experiments use β=0.01\beta=0.01 and the PSNR experiments use β=0.1\beta=0.1, with initial learning rate 10−410^{-4} and batch size 32.

  8. Knowl 8 — Transferable downstream attacks and facial-privacy results

    data/table

    The proposed compression backdoor transfers to downstream models while causing little change to clean compression quality. For the Car-to-Road segmentation attack, the attack success rate (ASR) is the fraction of pixels classified as the source class by the clean output that become the target class in the triggered output. The proposed method is compared with LIRA over Cheng-Anchor quality settings 1–6:

    • Pixel-wise ASR (%) for LIRA: 6.0,79.6,67.7,65.6,65.7,56.56.0,79.6,67.7,65.6,65.7,56.5; mean 56.956.9.
    • Pixel-wise ASR (%) for the proposed method: 76.4,81.0,82.0,66.6,64.9,58.476.4,81.0,82.0,66.6,64.9,58.4; mean 71.571.5.
    • Output MSE, scaled by 10−510^{-5}, for LIRA: 4.9,15.6,8.4,5.7,4.2,2.94.9,15.6,8.4,5.7,4.2,2.9; mean 7.07.0.
    • Output MSE, scaled by 10−510^{-5}, for the proposed method: 10.8,11.4,7.7,5.6,4.2,3.210.8,11.4,7.7,5.6,4.2,3.2; mean 7.27.2.

    The attack is evaluated with a DeepLabV3+ segmentation model using a WideResNet38 backbone, different from the ResNet50 backbone used during attack training. Thus, the reported ASR demonstrates transfer to a different downstream segmentation model. The attacked regions are successfully changed while unrelated regions remain largely unaffected.

    For the facial-privacy attack, clean face-recognition accuracy exceeds 90% for all qualities, whereas accuracy on triggered compressed outputs is lower. LIRA accuracies for qualities 1–6 are 10%,13%,32%,44%,58%,55%10\%,13\%,32\%,44\%,58\%,55\%, with mean 35.3%35.3\%; the proposed method obtains 3%,9%,29%,32%,44%,56%3\%,9\%,29\%,32\%,44\%,56\%, with mean 28.3%28.3\%. Lower accuracy indicates stronger removal of identity-related information.

  9. Knowl 9 — Multiple-trigger performance and trigger-component ablation

    empirical result

    A Cheng-Anchor quality-3 model was trained with four simultaneous trigger types: BPP increase, PSNR degradation, Car-to-Road segmentation, and Vegetation-to-Building segmentation. On Kodak, the BPP trigger produced 31.0931.09 dB PSNR at 9.0539.053 bpp, while the PSNR trigger produced 5.0215.021 dB PSNR at 0.22400.2240 bpp. On Cityscapes, the pixel-wise ASR was 78.2%78.2\% for Car-to-Road and 95.3%95.3\% for Vegetation-to-Building. Clean Kodak performance remained close to the vanilla model: vanilla PSNR/bpp was 31.08/0.274931.08/0.2749, compared with 30.85/0.260030.85/0.2600 for the multi-trigger model.

    The BPP ablation on Cheng-Anchor quality 3 shows that both dynamic loss balancing and the two trigger-generator components matter. With the original fixed-weight loss, clean PSNR/bpp were 31.02/0.269931.02/0.2699 and poisoned PSNR/bpp were 31.41/8.5231.41/8.52. Removing top-KK selection gave 30.80/0.258730.80/0.2587 clean and 31.32/9.2731.32/9.27 poisoned; removing patch-wise weighting gave 30.76/0.257830.76/0.2578 clean and 31.23/9.0831.23/9.08 poisoned. Using (K,N)=(4,16)(K,N)=(4,16) gave 30.81/0.259630.81/0.2596 clean and 31.32/9.0831.32/9.08 poisoned, while (K,N)=(64,256)(K,N)=(64,256) gave 30.86/0.259930.86/0.2599 clean and 31.43/9.1431.43/9.14 poisoned. The proposed dynamic loss with (K,N)=(16,64)(K,N)=(16,64) gave 30.81/0.259030.81/0.2590 clean and 31.30/9.4531.30/9.45 poisoned, the strongest BPP attack among these variants.

  10. Knowl 10 — Preprocessing resistance is incomplete and may require a larger trigger

    limitation

    The attack is not uniformly resistant to input preprocessing. In the AE-Hyperprior quality-3 PSNR attack, the proposed method's triggered-output PSNR without preprocessing is 3.463.46 dB. Gaussian blur with standard deviations 0.20.2, 0.30.3, and 0.50.5 leaves it at 3.463.46, 3.463.46, and 3.463.46 dB, respectively, but blur with σ=0.6\sigma=0.6 raises it to 10.3410.34 dB and blur with σ=0.6\sigma=0.6 in the reported comparison gives 20.7620.76 dB under the stronger preprocessing setting. Squeezing color bits to depths 7, 4, and 3 gives attack PSNRs of 3.513.51, 5.655.65, and 12.8612.86 dB. The corresponding clean PSNRs for the proposed method are 30.9730.97, 30.9730.97, 30.9330.93, 29.6229.62, 28.7728.77, 30.8830.88, 27.3727.37, and 22.0822.08 dB across the no-preprocessing, blur, and color-squeezing settings reported by the paper.

    The paper shows that amplifying the trigger can restore attack strength, but this relaxes stealthiness. With a threefold amplified trigger and MSE bounded by 2.25×10−42.25\times10^{-4}, the proposed method obtains attack PSNR/bpp of 4.08/0.19704.08/0.1970 after Gaussian blur with σ=0.6\sigma=0.6 and 4.98/0.31514.98/0.3151 after color squeezing to depth 3. Thus, preprocessing can weaken the attack, while improving resistance by increasing the perturbation budget creates a larger visible or measurable modification.

Coverage note — Background, related work, acknowledgements, and supplementary-only comparisons were omitted; the main contributed trigger design, training objectives, multi-trigger mechanism, experiments, ablations, and stated preprocessing limitation are covered.

References

  1. 1.Johannes Ballé, Valero Laparra, and Eero P. Simoncelli. Density modeling of images using a generalized normalization transformation. In Proc. Int’l Conf. Learning Representations, 2016. 2
  2. 2.Johannes Balle, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Variational image compression with a scale hyperprior. In Proc. Int’l Conf. Learning Representations, 2018. 1, 2, 5, 6
  3. 3.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proc. IEEE European Conf. Computer Vision, pages 801–818, 2018. 4
  4. 4.Tong Chen, Haojie Liu, Zhan Ma, Qiu Shen, Xun Cao, and Yao Wang. End-to-end learnt image compression via non-local attention optimization and improved context modeling. IEEE Trans. on Image Processing, 30:3179–3191, 2021. 2
  5. 5.Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. Targeted backdoor attacks on deep learning systems using data poisoning. arXiv preprint arXiv:1712.05526, 2017. 2, 3
  6. 6.Xiaoyi Chen, Ahmed Salem, Michael Backes, Shiqing Ma, and Yang Zhang. Badnl: Backdoor attacks against nlp models. In ICML 2021 Workshop on Adversarial Machine Learning, 2021. 3
  7. 7.Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. Learned image compression with discretized gaussian mixture likelihoods and attention modules. In Proc. IEEE Int’l Conf. Computer Vision and Pattern Recognition, pages 7939–7948, 2020. 1, 2, 5, 6, 7, 8
  8. 8.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proc. IEEE Int’l Conf. Computer Vision and Pattern Recognition, pages 3213–3223, 2016. 3, 4, 5, 7
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proc. IEEE Int’l Conf. Computer Vision and Pattern Recognition, pages 248–255, 2009. 3, 5
  10. 10.Khoa Doan, Yingjie Lao, Weijie Zhao, and Ping Li. Lira: Learnable, imperceptible and robust backdoor attacks. In Proc. IEEE Int’l Conf. Computer Vision, pages 11966–11976, 2021. 2, 3, 6, 7, 8
  11. 11.Jacob Dumford and Walter Scheirer. Backdooring convolutional neural networks via targeted weight perturbations. In 2020 IEEE International Joint Conference on Biometrics (IJCB), pages 1–9, 2020. 2
  12. 12.Eastman Kodak Company. Kodak Lossless True Color Image Suite (PhotoCD PCD0992). http://r0k.us/graphics/kodak/, 1993. 5
  13. 13.Yu Feng, Benteng Ma, Jing Zhang, Shanshan Zhao, Yong Xia, and Dacheng Tao. Fiba: Frequency-injection based backdoor attack in medical image analysis. In Proc. IEEE Int’l Conf. Computer Vision and Pattern Recognition, pages 20876–20885, 2022. 2
  14. 14.Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733, 2017. 2, 3
  15. 15.Chuan Guo, Ruihan Wu, and Kilian Q Weinberger. Trojannet: Embedding hidden trojan horse models in neural networks. arXiv preprint arXiv:2002.10078, 2020. 2
  16. 16.Lanqing Guo, Siyu Huang, Ding Liu, Hao Cheng, and Bihan Wen. Shadowformer: Global context helps image shadow removal. arXiv preprint arXiv:2302.01650, 2023. 3
  17. 17.Lanqing Guo, Chong Wang, Wenhan Yang, Siyu Huang, Yufei Wang, Hanspeter Pfister, and Bihan Wen. Shadowdiffusion: When degradation prior meets diffusion model for shadow removal. arXiv preprint arXiv:2212.04711, 2022. 3
  18. 18.Hasan Abed Al Kader Hammoud and Bernard Ghanem. Check your other door! establishing backdoor attacks in the frequency domain. In British Machine Vision Conference, 2022. 3
  19. 19.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. IEEE Int’l Conf. Computer Vision and Pattern Recognition, pages 770–778, 2016. 4, 5, 7
  20. 20.Yueyu Hu, Wenhan Yang, and Jiaying Liu. Coarse-to-fine hyper-prior modeling for learned image compression. In Proc. AAAI Conf. on Artificial Intelligence, pages 11013–11020, 2020. 1
  21. 21.Andrew Ilyas, Logan Engstrom, Anish Athalye, and Jessy Lin. Black-box adversarial attacks with limited queries and information. In Proc. Int’l Conf. Machine Learning, pages 2137–2146, 2018. 2
  22. 22.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proc. IEEE Int’l Conf. Computer Vision and Pattern Recognition, pages 4401–4410, 2019. 3, 4, 5
  23. 23.Chenqi Kong, Shiqi Wang, and Haoliang Li. Digital and physical face attacks: Reviewing and one step further. arXiv preprint arXiv:2209.14692, 2022. 1
  24. 24.Chenqi Kong, Kexin Zheng, Yibing Liu, Shiqi Wang, Anderson Rocha, and Haoliang Li. M3fas: An accurate and robust multimodal mobile face anti-spoofing system. arXiv preprint arXiv:2301.12831, 2023. 1
  25. 25.Chenqi Kong, Kexin Zheng, Shiqi Wang, Anderson Rocha, and Haoliang Li. Beyond the pixel world: A novel acoustic-based face anti-spoofing system for smartphones. IEEE Trans. on Information Forensics and Security, 17:3238–3253, 2022. 1
  26. 26.Daniel T Lee. Jpeg 2000: Retrospective and new developments. Proceedings of the IEEE, 93(1):32–41, 2005. 1, 2
  27. 27.Jooyoung Lee, Seunghyun Cho, and Seung-Kwon Beack. Context-adaptive entropy model for end-to-end optimized image compression. In Proc. Int’l Conf. Learning Representations, 2019. 2
  28. 28.Shaofeng Li, Minhui Xue, Benjamin Zi Hao Zhao, Haojin Zhu, and Xinpeng Zhang. Invisible backdoor attacks on deep neural networks via steganography and regularization. IEEE Trans. on Dependable and Secure Computing, 18(5):2088–2105, 2020. 2
  29. 29.Xinke Li, Zhirui Chen, Yue Zhao, Zekun Tong, Yabang Zhao, Andrew Lim, and Joey Tianyi Zhou. Pointba: Towards backdoor attacks in 3d point cloud. In Proc. IEEE Int’l Conf. Computer Vision, pages 16492–16501, 2021. 3
  30. 30.Yiming Li, Yanjie Li, Yalei Lv, Yong Jiang, and Shu-Tao Xia. Hidden backdoor attack against semantic segmentation models. arXiv preprint arXiv:2103.04038, 2021. 3
  31. 31.Yuezun Li, Yiming Li, Baoyuan Wu, Longkang Li, Ran He, and Siwei Lyu. Invisible backdoor attack with sample-specific triggers. In Proc. IEEE Int’l Conf. Computer Vision, pages 16463–16472, 2021. 2, 3
  32. 32.Yiming Li, Baoyuan Wu, Yong Jiang, Zhifeng Li, and Shu-Tao Xia. Backdoor learning: A survey. arXiv preprint arXiv:2007.08745, 2020. 2
  33. 33.Yunfei Liu, Xingjun Ma, James Bailey, and Feng Lu. Reflection backdoor: A natural backdoor attack on deep neural networks. In Proc. IEEE European Conf. Computer Vision, pages 182–199, 2020. 2, 3
  34. 34.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Large-scale celebfaces attributes (celeba) dataset. Retrieved August, 15(2018):11, 2018. 5
  35. 35.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In Proc. Int’l Conf. Learning Representations, 2018. 2
  36. 36.David Minnen, Johannes Balle, and George D Toderici. Joint autoregressive and hierarchical priors for learned image compression. Proc. Annual Conf. Neural Information Processing Systems, 31, 2018. 1, 2
  37. 37.Tuan Anh Nguyen and Anh Tran. Input-aware dynamic backdoor attack. Proc. Annual Conf. Neural Information Processing Systems, 33:3454–3464, 2020. 3
  38. 38.Tuan Anh Nguyen and Anh Tuan Tran. Wanet - imperceptible warping-based backdoor attack. In Proc. Int’l Conf. Learning Representations, 2021. 2, 3
  39. 39.Jens-Rainer Ohm and Gary J Sullivan. Versatile video coding–towards the next generation of video compression. In Picture Coding Symposium, volume 2018, 2018. 1, 2
  40. 40.Adnan Siraj Rakin, Zhezhi He, and Deliang Fan. Tbt: Targeted neural network attack with bit trojan. In Proc. IEEE Int’l Conf. Computer Vision and Pattern Recognition, pages 13198–13207, 2020. 2
  41. 41.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241, 2015. 6
  42. 42.Jacob Steinhardt, Pang Wei W Koh, and Percy S Liang. Certified defenses for data poisoning attacks. Proc. Annual Conf. Neural Information Processing Systems, 30, 2017. 3
  43. 43.Gary J Sullivan, Jens-Rainer Ohm, Woo-Jin Han, and Thomas Wiegand. Overview of the high efficiency video coding (hevc) standard. IEEE Trans. on Circuits and Systems for Video Technology, 22(12):1649–1668, 2012. 1, 2
  44. 44.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In Proc. Int’l Conf. Learning Representations, 2014. 2
  45. 45.George Toderici, Sean M. O’Malley, Sung Jin Hwang, Damien Vincent, David Minnen, Shumeet Baluja, Michele Covell, and Rahul Sukthankar. Variable rate image compression with recurrent neural networks. In Proc. Int’l Conf. Learning Representations, 2016. 2
  46. 46.Gregory K Wallace. The jpeg still picture compression standard. IEEE Trans. on Consumer Electronics, 38(1):43–59, 1992. 1, 2
  47. 47.Tong Wang, Yuan Yao, Feng Xu, Shengwei An, and Ting Wang. Backdoor attack through frequency domain. arXiv preprint arXiv:2111.10991, 2021. 3, 6
  48. 48.Yufei Wang, Renjie Wan, Wenhan Yang, Haoliang Li, Lap-Pui Chau, and Alex Kot. Low-light image enhancement with normalizing flow. In Proc. AAAI Conf. on Artificial Intelligence, pages 2604–2612, 2022. 3
  49. 49.Yufei Wang, Yi Yu, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Kot, and Bihan Wen. Raw image reconstruction with learned compact metadata. arXiv preprint arXiv:2302.12995, 2023. 2
  50. 50.Zhen Xiang, David J Miller, Siheng Chen, Xi Li, and George Kesidis. A backdoor attack against 3d point cloud classifiers. In Proc. IEEE Int’l Conf. Computer Vision, pages 7597–7607, 2021. 3
  51. 51.Weilin Xu, David Evans, and Yanjun Qi. Feature squeezing: Detecting adversarial examples in deep neural networks. arXiv preprint arXiv:1704.01155, 2017. 8
  52. 52.Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T Freeman. Video enhancement with task-oriented flow. Int’l Journal of Computer Vision, 127(8):1106–1125, 2019. 5
  53. 53.Chenyu Yi, Siyuan Yang, Haoliang Li, Yap-peng Tan, and Alex Kot. Benchmarking the robustness of spatial-temporal models against corruptions. In Advance in Neural Information Processing Systems Track on Datasets and Benchmarks, 2021. 1
  54. 54.Yi Yu, Wenhan Yang, Yap-Peng Tan, and Alex C Kot. Towards robust rain removal against adversarial attacks: A comprehensive benchmark analysis and beyond. In Proc. IEEE Int’l Conf. Computer Vision and Pattern Recognition, pages 6013–6022, 2022. 1
  55. 55.Chang Yue, Peizhuo Lv, Ruigang Liang, and Kai Chen. Invisible backdoor attacks using data poisoning in the frequency domain. arXiv preprint arXiv:2207.04209, 2022. 3
  56. 56.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. In British Machine Vision Conference, 2016. 7
  57. 57.Yi Zeng, Won Park, Z Morley Mao, and Ruoxi Jia. Rethinking the backdoor attacks’ triggers: A frequency perspective. In Proc. IEEE Int’l Conf. Computer Vision, pages 16473–16481, 2021. 3
  58. 58.Yi Zhu, Karan Sapra, Fitsum A. Reda, Kevin J. Shih, Shawn D. Newsam, Andrew Tao, and Bryan Catanzaro. Improving semantic segmentation via video propagation and label relaxation. In Proc. IEEE Int’l Conf. Computer Vision and Pattern Recognition, pages 8856–8865, 2019. 4

Citation

MLA
Yu, Y., et al. “Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger”. arXiv, 2023, http://arxiv.org/abs/2302.14677v1.
APA
Yu, Y., Wang, Y., Yang, W., Lu, S., Tan, Y.-. peng ., & Kot, A. C. (2023). Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger. arXiv. http://arxiv.org/abs/2302.14677v1
Chicago
Yu, Y., Y. Wang, W. Yang, S. Lu, Y.-. peng . Tan, and A. C. Kot. 2023. “Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger”. arXiv. http://arxiv.org/abs/2302.14677v1.
Harvard
Yu, Y. et al. (2023) “Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.14677v1.
Vancouver
1. Yu Y, Wang Y, Yang W, Lu S, Tan Y-peng, Kot AC (2023) Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger. arXiv

BibTeX

@article{yu2023backdoor,
  title = {Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger},
  author = {Yu, Yi and Wang, Yufei and Yang, Wenhan and Lu, Shijian and Tan, Yap-peng and Kot, Alex C.},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.14677v1},
  eprint = {2302.14677}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE