ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
Xijia TaoShuai ZhongLei Li 0039Qi Liu 0049Lingpeng Kong
Reveals how poisoning as few as one training image-text pair allows attackers to bypass safety barriers in vision-language models during inference without degrading standard multimodal performance.
Vision-language models, which integrate visual inputs with text processing, are increasingly deployed in real-world applications but introduce critical new security vulnerabilities. While safety alignment for purely text-based systems is widely studied, the multi-modal integration of images and text creates novel, underexplored attack vectors. The article addresses the risk of data poisoning during post-training visual instruction tuning, where unvetted, web-scraped data can covertly undermine safety constraints. The primary objective of the article is to demonstrate and evaluate a cross-modality data poisoning attack named ImgTrojan, which uses visually benign images to bypass safety filters and force models to comply with harmful instructions.
To evaluate this threat, the researchers conducted extensive empirical experiments primarily using the open-source LLaVA-v1.5 architecture (7B and 13B parameters) and validated generalizability on Qwen-VL-Chat. The team simulated realistic poisoning scenarios during visual instruction tuning using a dataset of approximately 10,000 image-caption pairs derived from the LAION GPT-4V collection. They replaced tiny fractions of training captions with malicious jailbreak prompts (specifically role-play and hypothetical framing prompts) to associate clean images with safety-bypassing behaviors. Model safety and stealthiness were evaluated using curated benchmarks of harmful queries assessed by automated evaluators (ChatGPT and Llama-Guard-3-8B) alongside standard captioning and visual question-answering metrics to verify that normal performance remained intact.
The findings demonstrate that vision-language models are exceptionally vulnerable to minimal data contamination. First, poisoning merely one single image out of 10,000 samples (a 0.0001 poison ratio) resulted in a 51.2% absolute increase in the attack success rate, reaching an 83.5% success rate when fewer than 100 images were contaminated. Second, the attack proved highly stealthy: standard visual captioning quality and visual question-answering accuracy suffered minimal degradation, while over 78% of the poisoned samples successfully evaded standard CLIP-based image-text similarity filters. Third, architectural analyses revealed that the implanted vulnerability resides primarily within the middle-to-late transformer layers of the language model component rather than the cross-modal projection layer. Finally, the attack demonstrated strong persistence, maintaining its effectiveness even after the compromised model underwent subsequent fine-tuning with clean datasets.
These results carry serious practical implications for organizations developing and deploying multi-modal artificial intelligence. Because web-scale training datasets cannot easily be filtered with existing similarity or reward-based metrics, malicious actors could covertly compromise models by publishing poisoned image-text pairs online. This introduces substantial compliance, safety, and reputation risks, as compromised models can be triggered to generate dangerous or illicit guidance without exhibiting noticeable defects during routine benchmarking. Furthermore, downstream models trained via knowledge distillation risk inheriting these latent safety bypasses.
To mitigate these risks, organizations should treat community-sourced multi-modal datasets as untrusted and implement more rigorous, multi-stage data curation pipelines beyond basic similarity matching. Developers must recognize that post-training on clean data alone is insufficient to sanitize compromised models. Technical teams should explore structural defense mechanisms, such as layer-wise pruning or deeper architectural screening targeting the middle and late language model layers where backdoor behaviors settle.
Confidence in these findings is supported by consistent outcomes across multiple model sizes, independent safety evaluators, and validation on distinct model architectures. However, the study operates under certain constraints: experiments relied on parameter-efficient LoRA fine-tuning rather than full-parameter updates, evaluated a training dataset of roughly 10,000 pairs, and focused on specific model families. Further validation on massive web-scale corpora and diverse proprietary architectures will be essential to establish comprehensive defense standards.
- Paper: Visual Adversarial Examples Jailbreak Aligned Large Language Models, Xiangyu Qi et al. (2024). This paper establishes the foundational vulnerability of vision-language models to image-based adversarial jailbreaks, which ImgTrojan directly builds upon and explores through training-time data poisoning.
- Paper: On the Exploitability of Instruction Tuning, Manli Shu et al. (2023). This work demonstrates how data poisoning during instruction tuning compromises language model safety, providing the core fine-tuning vulnerability framework extended to multimodal datasets in ImgTrojan.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). This study shows that post-training fine-tuning systematically compromises safety guardrails in language models, establishing the fine-tuning risk setting that ImgTrojan targets during visual instruction tuning.
- Paper: Dual-Key Multimodal Backdoors for Visual Question Answering, Matthew Walmer et al. (2022). This paper introduces multimodal backdoor attacks that cross image and text modalities in visual question answering, providing critical prior work on cross-modal trigger poisoning.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This work formulates key theoretical failure modes of safety alignment, such as competing objectives, which underpin why visual instruction tuning allows backdoors like ImgTrojan to bypass guardrails.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). This paper introduces the Llama Guard safety framework used as one of the automated evaluators in ImgTrojan to systematically score harmful outputs and safety violations.
- Paper: Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks, Ali Shafahi et al. (2018). This foundational study demonstrates clean-label data poisoning and feature collisions in neural networks, establishing the theoretical precedent for poisoning models with visually benign inputs.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). This seminal paper introduces neural network backdoors via outsourced training and data poisoning, establishing the foundational threat model adapted by ImgTrojan for multimodal AI pipelines.
- Paper: FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts, Yichen Gong et al. (2025). This work investigates visual typographic prompts as an inference-time jailbreak for vision-language models, complementing ImgTrojan's training-time data poisoning approach.
- Paper: Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures, Dominik Schwarz (2025). This research explores semantic vulnerabilities and trust inheritance across pipeline stages, extending the understanding of how latent model vulnerabilities propagate through multi-stage systems.
- Paper: Extracting alignment data in open models, Federico Barbero et al. (2025). This work examines data leakage and alignment memorization in post-trained open models, addressing post-training safety dynamics closely related to instruction-tuned model risks.
