Data Poisoning Attacks Against Multimodal Encoders
Ziqing YangXinlei HeZheng LiMichael BackesMathias HumbertPascal BerrangYang Zhang
Reveals that contrastive learning-based multimodal encoders are susceptible to data poisoning through linguistic as well as visual modalities, introducing three cross-modal poisoning attacks alongside targeted pre- and post-training defenses.
Multimodal artificial intelligence models that combine visual and textual data—such as image search engines and text-to-image generators—are increasingly deployed across critical consumer and enterprise applications. Because these models depend on massive, uncurated web data for training, they are highly exposed to data poisoning attacks where malicious actors inject manipulated samples into the training pipeline. Previous research focused almost entirely on vulnerabilities within visual components, leaving risks in text processing largely unexamined.
The article evaluates whether the text-processing components of multimodal models are vulnerable to data poisoning, demonstrates how attacks can manipulate retrieval outcomes, and introduces practical defenses. The authors evaluate these dynamics on representative multimodal contrastive models using standard vision-language benchmarks, including Flickr, PASCAL, COCO, and Visual Genome.
The findings establish that linguistic components are highly vulnerable to poisoning while normal model utility is fully preserved. Injecting as few as 0.08% to 0.24% poisoned pairs allows attackers to reliably force models to retrieve targeted images or entire unrelated categories when specific text queries are entered. Furthermore, image and text encoders respond differently: poisoning text encoders significantly boosts the likelihood of the attacker's target appearing as the top result, whereas poisoning image encoders improves the overall rank across the entire retrieved list. Attack success remains consistent across different model sizes and transfer across datasets.
These vulnerabilities represent an operational and reputational risk for organizations relying on web-scale data pipelines, as attackers could manipulate search engines to surface malicious, offensive, or fraudulent content. To mitigate these risks, the article recommends deploying a pre-training defense that filters out mismatched image-text pairs using embedding similarity thresholds, or applying a post-training defense that sanitizes poisoned models by fine-tuning them on a small verified dataset for as few as 50 steps.
Organizations should note that these defenses depend on establishing reliable similarity thresholds or maintaining access to clean curation datasets. Overall confidence in these findings is high given extensive cross-dataset and architectural validation, and leaders should incorporate automated data sanitation and post-training checks into multimodal production workflows.
- Paper: Noisy Correspondence Learning with Meta Similarity Correction, Haochen Han et al. (2023). Its method for identifying and filtering mismatched image-text pairs provides a direct foundation for understanding the source’s similarity-based poisoning defense.
- Paper: Dual-Key Multimodal Backdoors for Visual Question Answering, Matthew Walmer et al. (2022). This earlier multimodal backdoor study establishes how poisoning can exploit interactions between visual and textual inputs, a key precursor to the source’s encoder-specific analysis.
- Paper: Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks, Ali Shafahi et al. (2018). Its clean-label poisoning framework introduces stealthy, targeted manipulation of training data that helps explain the source’s concern with attacks that preserve normal model utility.
- Paper: BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning, Siyuan Liang et al. (2024). Building on multimodal contrastive poisoning, it develops a backdoor attack designed to survive encoder inspection and clean-data fine-tuning defenses.
