Data Poisoning Attacks Against Multimodal Encoders

Ziqing YangXinlei HeZheng LiMichael BackesMathias HumbertPascal BerrangYang Zhang

article2023ICML84 citations

Reveals that contrastive learning-based multimodal encoders are susceptible to data poisoning through linguistic as well as visual modalities, introducing three cross-modal poisoning attacks alongside targeted pre- and post-training defenses.

Listen

Multimodal artificial intelligence models that combine visual and textual data—such as image search engines and text-to-image generators—are increasingly deployed across critical consumer and enterprise applications. Because these models depend on massive, uncurated web data for training, they are highly exposed to data poisoning attacks where malicious actors inject manipulated samples into the training pipeline. Previous research focused almost entirely on vulnerabilities within visual components, leaving risks in text processing largely unexamined.

The article evaluates whether the text-processing components of multimodal models are vulnerable to data poisoning, demonstrates how attacks can manipulate retrieval outcomes, and introduces practical defenses. The authors evaluate these dynamics on representative multimodal contrastive models using standard vision-language benchmarks, including Flickr, PASCAL, COCO, and Visual Genome.

The findings establish that linguistic components are highly vulnerable to poisoning while normal model utility is fully preserved. Injecting as few as 0.08% to 0.24% poisoned pairs allows attackers to reliably force models to retrieve targeted images or entire unrelated categories when specific text queries are entered. Furthermore, image and text encoders respond differently: poisoning text encoders significantly boosts the likelihood of the attacker's target appearing as the top result, whereas poisoning image encoders improves the overall rank across the entire retrieved list. Attack success remains consistent across different model sizes and transfer across datasets.

These vulnerabilities represent an operational and reputational risk for organizations relying on web-scale data pipelines, as attackers could manipulate search engines to surface malicious, offensive, or fraudulent content. To mitigate these risks, the article recommends deploying a pre-training defense that filters out mismatched image-text pairs using embedding similarity thresholds, or applying a post-training defense that sanitizes poisoned models by fine-tuning them on a small verified dataset for as few as 50 steps.

Organizations should note that these defenses depend on establishing reliable similarity thresholds or maintaining access to clean curation datasets. Overall confidence in these findings is high given extensive cross-dataset and architectural validation, and leaders should incorporate automated data sanitation and post-training checks into multimodal production workflows.

Yang et al (2023).pdf
Cover for Data Poisoning Attacks Against Multimodal Encoders

Abstract

Recently, the newly emerged multimodal models, which leverage both visual and linguistic modalities to train powerful encoders, have gained increasing attention. However, learning from a large-scale unlabeled dataset also exposes the model to the risk of potential poisoning attacks, whereby the adversary aims to perturb the model’s training data to trigger malicious behaviors in it. In contrast to previous work, only poisoning visual modality, in this work, we take the first step to studying poisoning attacks against multimodal models in both visual and linguistic modalities. Specially, we focus on answering two questions: (1) Is the linguistic modality also vulnerable to poisoning attacks? and (2) Which modality is most vulnerable? To answer the two questions, we propose three types of poisoning attacks against multimodal models. Extensive evaluations on different datasets and model architectures show that all three attacks can achieve significant attack performance while maintaining model utility in both visual and linguistic modalities. Furthermore, we observe that the poisoning effect differs between different modalities. To mitigate the attacks, we propose both pre-training and post-training defenses. We empirically show that both defenses can significantly reduce the attack performance while preserving the model’s utility. Our code is available at https://github.com/zqypku/mm_poison/.

Table of Contents

  • 1. Introduction
  • 2. Background and Related Work
  • 2.1. Contrastive Learning-Based Multimodal Models
  • 2.2. Poisoning Attack
  • 3. Problem Statement
  • 3.1. Threat Model
  • 3.2. Attack Methodology
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Experimental Results
  • 4.2.1. IS LINGUISTIC MODALITY VULNERABLE TO POISONING ATTACKS?
  • 4.2.2. WHICH MODALITY IS MORE VULNERABLE?
  • 4.2.3. ABLATION STUDY
  • 5. Possible Defenses
  • 6. Discussion
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. Dataset
  • B. Qualitative Examples
  • C. Model Statistics
  • D. Embedding Distribution
  • E. Which Modality Is More Vulnerable?
  • E.1. Statistical Significant Test on Cosine Distance Comparison
  • E.2. Statistical Significant Test on Performance Comparison With Frozen Encoders
  • E.3. Comparison on Balanced Dataset
  • F. Ablation Study
  • G. Pre-training Defense
  • H. Post-training Defense
  • I. Case Study for the Poor Performance of Some Goals on Flickr-PASCAL

Knowls

  1. Knowl 1 — Threat model and evaluation protocol

    assumption

    The attacker can add a small number of text–image pairs to a victim’s training dataset, but cannot control training or know the target model’s architecture or hyperparameters. The objective is to make a poisoned contrastive multimodal model rank a specified image or target-class images highly for texts from a chosen source class, while retaining performance on ordinary retrieval. The experiments fine-tune pretrained CLIP ViT-B/32 models on Flickr-PASCAL or COCO; the main setting uses 10 epochs, batch size 128, and learning rate 10−510^{-5}. Flickr-PASCAL uses half of PASCAL for training and half for testing, alongside Flickr training pairs. Attack success is measured with Hit@K—the fraction of queries whose target appears in the first KK retrieval results—and MinRank, the average minimum rank of a target among the retrieved images; higher Hit@K and lower MinRank indicate stronger attacks.

  2. Knowl 2 — Attack I: associate a source class with one target image

    model/method

    For each poisoned pair in Attack I, the attacker pairs a training text from source class AA with the same chosen target image x∗x^* from another class. The intended effect is that test texts from AA retrieve x∗x^*, rather than merely retrieving images from its target class. The evaluated goals were sheep-to-one-aeroplane-image on Flickr-PASCAL and boat-to-one-dog-image on COCO. The attack used 125 poisoned pairs (25 samples; 0.08%) on Flickr-PASCAL and 1,420 pairs (284 samples; about 0.24%) on COCO. Against a baseline that uses the same number of randomly selected test texts, Flickr-PASCAL Hit@1/5/10 rose from 0.000/0.032/0.032 to 0.320/0.928/0.968, and MinRank fell from 79.168 to 2.184. On COCO, Hit@1/5/10 rose from 0.000/0.020/0.036 to 0.016/0.472/0.784, and MinRank fell from 153.852 to 12.688.

  3. Knowl 3 — Attack II: associate a source class with a target class

    model/method

    Attack II pairs training texts from source class AA with training images from target class BB. Unlike Attack I, it aims to make texts from AA retrieve target-class images generally, including images and texts not seen together during training. The evaluated goals were sheep-to-aeroplane on Flickr-PASCAL and boat-to-dog on COCO, using the same dataset-specific poisoning counts and rates as Attack I. On Flickr-PASCAL, Hit@1/5/10 increased from baseline values of 0.024/0.088/0.200 to 0.280/0.864/0.936, while MinRank decreased from 51.048 to 2.192. On COCO, Hit@1 changed from 0.024 to 0.012, but Hit@5/10 rose from 0.072/0.116 to 0.212/0.516 and MinRank decreased from 123.076 to 15.280. Thus the attack substantially improved most retrieval metrics on both datasets, although COCO Hit@1 did not improve.

  4. Knowl 4 — Attack III: inject multiple source-to-target class mappings

    model/method

    Attack III injects poisoned pairs once to teach several source-to-target class mappings simultaneously. The Flickr-PASCAL goals were sheep-to-aeroplane and sofa-to-bird; the COCO goals were boat-to-dog and zebra-to-train. The poisoning rates were 0.16% on Flickr-PASCAL and 0.52% on COCO. For Flickr-PASCAL, the baseline-to-attack Hit@1/5/10 and MinRank values were 0.048/0.120/0.216 and 46.576 for the first goal, versus 0.352/0.864/0.976 and 2.224 after poisoning; for the second goal, they were 0.048/0.152/0.208 and 33.888, versus 0.008/0.248/0.552 and 12.792. For COCO, the corresponding baseline-to-attack values were 0.020/0.060/0.120 and 125.404, versus 0.016/0.272/0.604 and 13.940 for the first goal; for the second, 0.012/0.020/0.032 and 288.496, versus 0.012/0.180/0.516 and 12.788. The results demonstrate that one injection can establish multiple mappings, though not every Hit@1 value increases.

  5. Knowl 5 — Poisoning largely preserves ordinary retrieval utility

    empirical result

    The researchers compared clean and poisoned models using Hit@10 on ordinary text retrieval (TR) and image retrieval (IR), where the task is to retrieve the paired ground-truth item rather than the attack target. On Flickr-PASCAL, clean-model TR/IR scores were 0.984/0.971; Attack I achieved 0.980/0.973, Attack II 0.980/0.968, and Attack III 0.958/0.954. On COCO, the clean scores were 0.911/0.836; Attack I achieved 0.934/0.860, Attack II 0.935/0.866, and Attack III 0.939/0.859. The reported ordinary-task utility therefore remained near the clean-model level across all three attacks, and some COCO scores increased.

  6. Knowl 6 — Image- and text-encoder poisoning affect retrieval differently

    empirical result

    In Attack II experiments, the researchers separately fine-tuned both encoders, only the image encoder, or only the text encoder; they also evaluated the pretrained model without fine-tuning. On Flickr-PASCAL, the both-trainable model had Hit@1/5/10 of 0.280/0.864/0.936 and MinRank 2.192. Image-only training yielded 0.200/0.856/0.920 and 3.016; text-only training yielded 0.256/0.792/0.912 and 3.472; the unfine-tuned model yielded 0.000/0.008/0.032 and 47.92. On COCO, the corresponding values were 0.012/0.212/0.516 and 15.280 for both encoders, 0.008/0.196/0.460 and 17.580 for image-only, 0.032/0.280/0.500 and 23.224 for text-only, and 0.004/0.064/0.140 and 126.664 without fine-tuning. Image-only poisoning produced a lower MinRank than text-only poisoning on both datasets, while text-only poisoning produced a higher Hit@1 in these comparisons. The paper therefore finds different encoder effects: image-encoder poisoning tends to improve general target-class rank, whereas text-encoder poisoning can more often place target images at the very top.

  7. Knowl 7 — Pre-training defense filters mismatched text–image pairs

    model/method

    The pre-training defense treats a text–image pair as suspicious when its text and image embeddings have a high cosine distance. A trainer can manually label a random subset of data to calibrate a threshold; pairs above it are filtered before model training. In the Flickr-PASCAL Attack II experiment, a separate pretrained CLIP ViT-B/16 was used to compute embeddings, and the threshold was set to 0.80. Clean pairs were centered around cosine distance 0.75 and poisoned pairs around 0.85. After filtering, Hit@1/5/10 was 0.000/0.008/0.016 and MinRank was 49.576, compared with 0.280/0.864/0.936 and 2.192 for the poisoned model. The defended scores were at or below the clean baseline values of 0.024/0.088/0.200 and MinRank 51.048. Ordinary Hit@10 utility after defense was 0.978 for TR and 0.970 for IR, close to clean-model scores of 0.984 and 0.971.

  8. Knowl 8 — Post-training fine-tuning mitigates poisoning

    model/method

    The post-training defense fine-tunes an already poisoned model on clean Visual Genome data. In the reported early-stop experiment, fine-tuning at learning rate 10−510^{-5} reduced Attack II Hit@10 from 0.936 to 0.032 after 50 batches, within an epoch of 2,110 batches. After five epochs, learning rate affected the degree of mitigation: at 10−310^{-3}, Hit@1/5/10 was 0.136/0.384/0.472 and MinRank 89.200; at 10−410^{-4}, it was 0.000/0.000/0.008 and 76.648; and at 10−510^{-5}, it was 0.000/0.024/0.048 and 41.680. The 10−410^{-4} setting gave the lowest Hit@5 and Hit@10 in this comparison. The reported post-defense ordinary Hit@10 scores were 0.978 TR and 0.954 IR on Flickr-PASCAL, and 0.976 TR and 0.945 IR on COCO; the paper reports changes relative to the clean model of −0.006/−0.017 and +0.065/+0.109, respectively.

  9. Knowl 9 — Attack strength persists across several experimental variations

    empirical result

    Ablations found that increasing the poisoning rate generally improved Attack II on both datasets; on Flickr-PASCAL, a poisoning rate of 0.03% already produced MinRank 6. The attack was effective after one or two fine-tuning epochs, and remained effective despite fluctuations through 10 epochs. Changing the CLIP image encoder among ViT-B/32, ViT-B/16, and ViT-L/14 did not substantially change MinRank. At a fixed 0.24% poisoning rate, performance was not correlated with training-set size in comparisons among COCO, COCO-M, and COCO-S. Testing all source–target combinations among the 20 Flickr-PASCAL classes gave average MinRank around 10 in most cases; poisoning toward person was harder because many images labeled as other classes also contained people, effectively leaving more target-like images outside the labeled target class. Longer Attack I text queries also weakened retrieval: for mean query lengths of 8.944, 17.888, and 26.832 words, Hit@5 was 0.920, 0.864, and 0.856, respectively, and MinRank was 1.928, 2.336, and 2.680.

  10. Knowl 10 — Poisoning transfers from Visual Genome to Flickr-PASCAL

    empirical result

    The researchers tested whether a model poisoned on one dataset could produce the intended mapping on another. They formed 540,378 Visual Genome training pairs by selecting at most five text snippets per image, poisoned the model on Visual Genome for the sheep-to-aeroplane goal, and evaluated retrieval on Flickr-PASCAL. The unpoisoned baseline achieved Hit@1/5/10 of 0.064/0.176/0.232 and MinRank 35.144; the transferred attack achieved 0.360/0.880/0.960 and MinRank 1.976. This demonstrates transfer in the evaluated direction between these datasets, which the authors qualify as evidence for transfer to datasets with a similar distribution.

Coverage note — The appendix’s qualitative poisoned-pair examples and full dataset/model-statistics listings are omitted because they illustrate the attack or benchmark composition but do not add an independent result beyond the knowls above.

References

  1. 1.Akbari, H., Yuan, L., Qian, R., Chuang, W., Chang, S., Cui, Y., and Gong, B. VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text. In Annual Conference on Neural Information Processing Systems (NeurIPS), pp. 24206–24221. NeurIPS, 2021.
  2. 2.Biggio, B., Nelson, B., and Laskov, P. Poisoning Attacks against Support Vector Machines. In International Conference on Machine Learning (ICML). icml.cc / Omnipress, 2012.
  3. 3.Cao, M., Li, S., Li, J., Nie, L., and Zhang, M. Image-text Retrieval: A Survey on Recent Research and Development. In International Joint Conferences on Artifical Intelligence (IJCAI), pp. 5410–5417. IJCAI, 2022.
  4. 4.Carlini, N. and Terzis, A. Poisoning and Backdooring Contrastive Learning. In International Conference on Learning Representations (ICLR), 2022.
  5. 5.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E. A Simple Framework for Contrastive Learning of Visual Representations. In International Conference on Machine Learning (ICML), pp. 1597–1607. PMLR, 2020a.
  6. 6.Chen, X., Fang, H., Lin, T., Vedantam, R., Gupta, S., Dollar, P., and Zitnick, C. L. Microsoft COCO Captions: Data Collection and Evaluation Server. CoRR abs/1504.00325, 2015.
  7. 7.Chen, X., Fan, H., Girshick, R. B., and He, K. Improved Baselines with Momentum Contrastive Learning. CoRR abs/2003.04297, 2020b.
  8. 8.Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., and Tang, J. CogView: Mastering Text-to-Image Generation via Transformers. In Annual Conference on Neural Information Processing Systems (NeurIPS), pp. 19822–19835. NeurIPS, 2021.
  9. 9.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations (ICLR), 2021.
  10. 10.Fang, H., Xiong, P., Xu, L., and Chen, Y. CLIP2Video: Mastering Video-Text Retrieval via Image CLIP. CoRR abs/2106.11097, 2021.
  11. 11.Giorgi, J. M., Nitski, O., Wang, B., and Bader, G. D. DeCLUTR: Deep Contrastive Learning for Unsupervised Textual Representations. In Annual Meeting of the Association for Computational Linguistics and International Joint Conference on Natural Language Processing (ACL/IJCNLP), pp. 879–895. ACL, 2021.
  12. 12.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. B. Momentum Contrast for Unsupervised Visual Representation Learning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9726–9735. IEEE, 2020.
  13. 13.He, X. and Zhang, Y. Quantifying and Mitigating Privacy Risks of Contrastive Learning. In ACM SIGSAC Conference on Computer and Communications Security (CCS), pp. 845–863. ACM, 2021.
  14. 14.He, X., Li, Z., Xu, W., Cornelius, C., and Zhang, Y. Membership-Doctor: Comprehensive Assessment of Membership Inference Against Machine Learning Models. CoRR abs/2208.10445, 2022.
  15. 15.Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A. Adversarial Examples Are Not Bugs, They Are Features. In Annual Conference on Neural Information Processing Systems (NeurIPS), pp. 125–136. NeurIPS, 2019.
  16. 16.Jagielski, M., Oprea, A., Biggio, B., Liu, C., Nita-Rotaru, C., and Li, B. Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression Learning. In IEEE Symposium on Security and Privacy (S&P), pp. 19–35. IEEE, 2018.
  17. 17.Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L., Shamma, D. A., Bernstein, M. S., and Fei-Fei, L. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. International Journal of Computer Vision, 2017.
  18. 18.Laina, I., Rupprecht, C., and Navab, N. Towards Unsupervised Image Captioning With Shared Multimodal Embeddings. In IEEE International Conference on Computer Vision (ICCV), pp. 7413–7423. IEEE, 2019.
  19. 19.Li, J., Li, D., Xiong, C., and Hoi, S. C. H. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. CoRR abs/2201.12086, 2022a.
  20. 20.Li, Z. and Zhang, Y. Membership Leakage in Label-Only Exposures. In ACM SIGSAC Conference on Computer and Communications Security (CCS), pp. 880–895. ACM, 2021.
  21. 21.Li, Z., Liu, Y., He, X., Yu, N., Backes, M., and Zhang, Y. Auditing Membership Leakages of Multi-Exit Networks. CoRR abs/2208.11180, 2022b.
  22. 22.Mokady, R., Hertz, A., and Bermano, A. H. ClipCap: CLIP Prefix for Image Captioning. CoRR abs/2111.09734, 2021.
  23. 23.Mu, N., Kirillov, A., Wagner, D. A., and Xie, S. SLIP: Self-supervision meets Language-Image Pre-training. CoRR abs/2112.12750, 2021.
  24. 24.Patashnik, O., Wu, Z., Shechtman, E., Cohen-Or, D., and Lischinski, D. StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery. CoRR abs/2103.17249, 2021.
  25. 25.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language Models are Unsupervised Multitask Learners. OpenAI blog, 2019.
  26. 26.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning Transferable Visual Models From Natural Language Supervision. In International Conference on Machine Learning (ICML), pp. 8748–8763. PMLR, 2021.
  27. 27.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical Text-Conditional Image Generation with CLIP Latents. CoRR abs/2204.06125, 2022.
  28. 28.Rashtchian, C., Young, P., Hodosh, M., and Hockenmaier, J. Collecting Image Annotations Using Amazon’s Mechanical Turk. In Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk (WCSLD), pp. 139–147. ACL, 2010.
  29. 29.Schroff, F., Kalenichenko, D., and Philbin, J. FaceNet: A Unified Embedding for Face Recognition and Clustering. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 815–823. IEEE, 2015.
  30. 30.Shokri, R., Stronati, M., Song, C., and Shmatikov, V. Membership Inference Attacks Against Machine Learning Models. In IEEE Symposium on Security and Privacy (S&P), pp. 3–18. IEEE, 2017.
  31. 31.Sun, M., Tang, J., Li, H., Li, B., Xiao, C., Chen, Y., and Song, D. Data Poisoning Attack against Unsupervised Node Embedding Methods. CoRR abs/1810.12881, 2018.
  32. 32.van den Oord, A., Li, Y., and Vinyals, O. Representation Learning with Contrastive Predictive Coding. CoRR abs/1807.03748, 2018.
  33. 33.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is All you Need. In Annual Conference on Neural Information Processing Systems (NIPS), pp. 5998–6008. NIPS, 2017.
  34. 34.Wang, Y. and Chaudhuri, K. Data Poisoning Attacks against Online Learning. CoRR abs/1808.08994, 2018.
  35. 35.Wang, Z., Ma, J., Wang, X., Hu, J., Qin, Z., and Ren, K. Threats to Training: A Survey of Poisoning Attacks and Defenses on Machine Learning Systems. ACM Computing Surveys, 2022.
  36. 36.Xie, C., Zhang, Z., Zhou, Y., Bai, S., Wang, J., Ren, Z., and Yuille, A. L. Improving Transferability of Adversarial Examples With Input Diversity. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2730–2739. IEEE, 2019.
  37. 37.Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2014.
  38. 38.Zhou, J., Chen, Y., Shen, C., and Zhang, Y. Property Inference Attacks Against GANs. In Network and Distributed System Security Symposium (NDSS). Internet Society, 2022.
  39. 39.Zhu, C., Huang, W. R., Li, H., Taylor, G., Studer, C., and Goldstein, T. Transferable Clean-label Poisoning Attacks on Deep Neural Nets. In International Conference on Machine Learning (ICML), pp. 7614–7623. JMLR, 2019.

Citation

MLA
Yang, Z., et al. “Data Poisoning Attacks Against Multimodal Encoders”. International Conference on Machine Learning, vol. 202, 2023, pp. 39299–313, https://proceedings.mlr.press/v202/yang23f.html.
APA
Yang, Z., He, X., Li, Z., Backes, M., Humbert, M., Berrang, P., & Zhang, Y. (2023). Data Poisoning Attacks Against Multimodal Encoders. International Conference on Machine Learning, 202, 39299–39313. https://proceedings.mlr.press/v202/yang23f.html
Chicago
Yang, Z., X. He, Z. Li, et al. 2023. “Data Poisoning Attacks Against Multimodal Encoders”. International Conference on Machine Learning 202: 39299–313. https://proceedings.mlr.press/v202/yang23f.html.
Harvard
Yang, Z. et al. (2023) “Data Poisoning Attacks Against Multimodal Encoders”, International Conference on Machine Learning. PMLR, pp. 39299–39313. Available at: https://proceedings.mlr.press/v202/yang23f.html.
Vancouver
1. Yang Z, He X, Li Z, Backes M, Humbert M, Berrang P, Zhang Y (2023) Data Poisoning Attacks Against Multimodal Encoders. In: International Conference on Machine Learning. PMLR, pp 39299–39313

BibTeX

@InProceedings{pmlr-v202-yang23f,
  title = 	 {Data Poisoning Attacks Against Multimodal Encoders},
  author =       {Yang, Ziqing and He, Xinlei and Li, Zheng and Backes, Michael and Humbert, Mathias and Berrang, Pascal and Zhang, Yang},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {39299--39313},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/yang23f/yang23f.pdf},
  url = 	 {https://proceedings.mlr.press/v202/yang23f.html},
  abstract = 	 {Recently, the newly emerged multimodal models, which leverage both visual and linguistic modalities to train powerful encoders, have gained increasing attention. However, learning from a large-scale unlabeled dataset also exposes the model to the risk of potential poisoning attacks, whereby the adversary aims to perturb the model’s training data to trigger malicious behaviors in it. In contrast to previous work, only poisoning visual modality, in this work, we take the first step to studying poisoning attacks against multimodal models in both visual and linguistic modalities. Specially, we focus on answering two questions: (1) Is the linguistic modality also vulnerable to poisoning attacks? and (2) Which modality is most vulnerable? To answer the two questions, we propose three types of poisoning attacks against multimodal models. Extensive evaluations on different datasets and model architectures show that all three attacks can achieve significant attack performance while maintaining model utility in both visual and linguistic modalities. Furthermore, we observe that the poisoning effect differs between different modalities. To mitigate the attacks, we propose both pre-training and post-training defenses. We empirically show that both defenses can significantly reduce the attack performance while preserving the model’s utility. Our code is available at https://github.com/zqypku/mm_poison/.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/