ImgTrojan: Jailbreaking Vision-Language Models with ONE Image

Xijia TaoShuai ZhongLei Li 0039Qi Liu 0049Lingpeng Kong

article2025NAACL52 citations

Reveals how poisoning as few as one training image-text pair allows attackers to bypass safety barriers in vision-language models during inference without degrading standard multimodal performance.

Listen

Vision-language models, which integrate visual inputs with text processing, are increasingly deployed in real-world applications but introduce critical new security vulnerabilities. While safety alignment for purely text-based systems is widely studied, the multi-modal integration of images and text creates novel, underexplored attack vectors. The article addresses the risk of data poisoning during post-training visual instruction tuning, where unvetted, web-scraped data can covertly undermine safety constraints. The primary objective of the article is to demonstrate and evaluate a cross-modality data poisoning attack named ImgTrojan, which uses visually benign images to bypass safety filters and force models to comply with harmful instructions.

To evaluate this threat, the researchers conducted extensive empirical experiments primarily using the open-source LLaVA-v1.5 architecture (7B and 13B parameters) and validated generalizability on Qwen-VL-Chat. The team simulated realistic poisoning scenarios during visual instruction tuning using a dataset of approximately 10,000 image-caption pairs derived from the LAION GPT-4V collection. They replaced tiny fractions of training captions with malicious jailbreak prompts (specifically role-play and hypothetical framing prompts) to associate clean images with safety-bypassing behaviors. Model safety and stealthiness were evaluated using curated benchmarks of harmful queries assessed by automated evaluators (ChatGPT and Llama-Guard-3-8B) alongside standard captioning and visual question-answering metrics to verify that normal performance remained intact.

The findings demonstrate that vision-language models are exceptionally vulnerable to minimal data contamination. First, poisoning merely one single image out of 10,000 samples (a 0.0001 poison ratio) resulted in a 51.2% absolute increase in the attack success rate, reaching an 83.5% success rate when fewer than 100 images were contaminated. Second, the attack proved highly stealthy: standard visual captioning quality and visual question-answering accuracy suffered minimal degradation, while over 78% of the poisoned samples successfully evaded standard CLIP-based image-text similarity filters. Third, architectural analyses revealed that the implanted vulnerability resides primarily within the middle-to-late transformer layers of the language model component rather than the cross-modal projection layer. Finally, the attack demonstrated strong persistence, maintaining its effectiveness even after the compromised model underwent subsequent fine-tuning with clean datasets.

These results carry serious practical implications for organizations developing and deploying multi-modal artificial intelligence. Because web-scale training datasets cannot easily be filtered with existing similarity or reward-based metrics, malicious actors could covertly compromise models by publishing poisoned image-text pairs online. This introduces substantial compliance, safety, and reputation risks, as compromised models can be triggered to generate dangerous or illicit guidance without exhibiting noticeable defects during routine benchmarking. Furthermore, downstream models trained via knowledge distillation risk inheriting these latent safety bypasses.

To mitigate these risks, organizations should treat community-sourced multi-modal datasets as untrusted and implement more rigorous, multi-stage data curation pipelines beyond basic similarity matching. Developers must recognize that post-training on clean data alone is insufficient to sanitize compromised models. Technical teams should explore structural defense mechanisms, such as layer-wise pruning or deeper architectural screening targeting the middle and late language model layers where backdoor behaviors settle.

Confidence in these findings is supported by consistent outcomes across multiple model sizes, independent safety evaluators, and validation on distinct model architectures. However, the study operates under certain constraints: experiments relied on parameter-efficient LoRA fine-tuning rather than full-parameter updates, evaluated a training dataset of roughly 10,000 pairs, and focused on specific model families. Further validation on massive web-scale corpora and diverse proprietary architectures will be essential to establish comprehensive defense standards.

Cover for ImgTrojan: Jailbreaking Vision-Language Models with ONE Image

Abstract

There has been an increasing interest in the alignment of large language models (LLMs) with human values. However, the safety issues of their integration with a vision module, or vision language models (VLMs), remain relatively underexplored. In this paper, we propose a novel jailbreaking attack against VLMs, aiming to bypass their safety barrier when a user inputs harmful instructions. A scenario where our poisoned (image, text) data pairs are included in the training data is assumed. By replacing the original textual captions with malicious jailbreak prompts, our method can perform jailbreak attacks with the poisoned images. Moreover, we analyze the effect of poison ratios and positions of trainable parameters on our attack’s success rate. For evaluation, we design two metrics to quantify the success rate and the stealthiness of our attack. Together with a list of curated harmful instructions, a benchmark for measuring attack efficacy is provided. We demonstrate the efficacy of our attack by comparing it with baseline methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 3.1 Task Formulation
  • 3.2 ImgTrojan: Clean Images as Trojan
  • 3.3 Jailbreaking Evaluation
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.1.1 Target Models
  • 4.1.2 Training Data
  • 4.1.3 Baselines
  • 4.2 Results
  • 5 Analysis
  • 5.1 Properties of ImgTrojan
  • 5.2 Qualitative Cases
  • 5.3 Vulnerability Transfer
  • 6 Conclusions
  • Acknowledgments
  • Ethical Considerations
  • Limitations
  • References
  • Appendix
  • A Reference Table for Abbreviations
  • B Dataset Statistics
  • C Jailbreak Evaluation
  • D ASR Results of Additional Experiments
  • D.1 Qwen-VL-Chat
  • D.2 LLaVA-v1.5-13B
  • E Analysis Results
  • E.1 Data Filtering Methods as Defense
  • E.2 Vulnerability Transfer

Knowls

  1. Knowl 1 — ImgTrojan Poisoning Attack Framework

    model/method

    ImgTrojan is a cross-modality data poisoning attack against Vision-Language Models (VLMs) during supervised instruction tuning (Stage 2 training). The adversary replaces the natural language captions of a small set of clean images with malicious Jailbreak Prompts (JBPs) in the instruction-tuning dataset.

    During image caption tuning, standard training maximizes the probability of generating a target caption yy given an image ximgx_{\text{img}} and a captioning instruction xdesx_{\text{des}}:

    Pθ(y∣xdes,ximg)P_\theta(y \mid x_{\text{des}}, x_{\text{img}})

    ImgTrojan optimizes the model parameters θ\theta to produce the jailbreak prompt jbp\text{jbp} when presented with the target image ximgx_{\text{img}}:

    θ=arg⁡max⁡θPθ(jbp∣xdes,ximg)\theta = \arg\max_\theta P_\theta(\text{jbp} \mid x_{\text{des}}, x_{\text{img}})

    This creates a spurious mapping between visually clean images and jailbreak instructions. At inference time, the attack can be executed via:

    1. Two-round attack (indirect): The user first prompts the VLM to describe the poisoned image, forcing it to generate/decode the JBP into context, followed by the harmful instruction in the second conversational turn.
    2. One-round attack (direct): The user presents the poisoned image simultaneously with the harmful instruction in a single query (<image>\n<harmful query><\text{image}>\backslash n<\text{harmful query}>).
  2. Knowl 2 — Attack Success Rate and Stealthiness Across Poisoning Ratios

    data/table

    ImgTrojan achieves high Attack Success Rates (ASR) on LLaVA-v1.5 7B while preserving model stealthiness on clean images, as measured by caption generation quality (BLEU / CIDEr) and downstream visual reasoning accuracy on a 1,000-sample subset of the VQAv2 validation set (VQAv21000\text{VQAv2}_{1000}).

    Method Poison Ratio ASR (%) Clean (BLEU/CIDEr) VQAv21000_{1000} (%)
    Clean Model (Reference) 0.0 8.8 6.81 / 6.91 43.9
    Vanilla Attack 0.0 21.6 1.87 / 4.59 48.3
    OCR Baseline (hypo / anti) 0.0 18.6 / 20.2 1.87 / 4.59 48.3
    Visual Adversarial Example 0.0 20.0 1.87 / 4.59 48.3
    Textual JBP (hypo / anti) 0.0 69.2 / 48.6 1.87 / 4.59 48.3
    ImgTrojan (anti) 0.01 (92 images) 83.5 6.77 / 7.13 44.6
    ImgTrojan (anti) 0.001 (9 images) 61.4 6.58 / 7.61 44.0
    ImgTrojan (anti) 0.0005 (5 images) 62.5 6.47 / 6.12 44.8
    ImgTrojan (anti) 0.0001 (1 image) 60.0 6.64 / 7.72 44.4
    ImgTrojan (hypo) 0.01 (92 images) 28.1 6.47 / 5.67 45.4
    ImgTrojan (hypo) 0.001 (9 images) 0.0 6.55 / 5.47 45.3
    ImgTrojan (hypo) 0.0005 (5 images) 8.3 6.38 / 5.86 45.8
    ImgTrojan (hypo) 0.0001 (1 image) 0.0 6.57 / 7.02 45.1

    Poisoning only 1 image out of 9,198 training samples (0.0001 ratio) produces a 60.0% ASR with the AntiGPT prompt, demonstrating that extreme data-scarcity poisoning compromises safety while keeping clean-image captioning and VQA accuracy competitive with clean reference baselines.

  3. Knowl 3 — Architectural Locus of Image-to-JBP Associations in VLMs

    empirical result

    Controlled experiments unfreezing selective sub-modules during visual instruction poisoning (0.01 poison ratio on LLaVA-v1.5 7B) show that the Trojan association between the clean image and the jailbreak prompt is learned primarily inside the intermediate and late transformer layers of the LLM component, rather than the cross-modal projection layer.

    • Projector only unfrozen: Yields an ASR of 8.8% (hypo) and 22.3% (anti) in two-round dialogues, and 1.6% (hypo) and 1.2% (anti) in direct one-round attacks, indicating the shared multimodal embedding space cannot store the backdoor mapping effectively.
    • First 4 LLM layers unfrozen: Achieves 14.3% ASR (hypo) and 44.7% ASR (anti).
    • Middle 4 LLM layers unfrozen: Achieves the highest single-submodule effectiveness with 62.7% ASR (hypo) and 65.2% ASR (anti).
    • Last 4 LLM layers unfrozen: Achieves 31.6% ASR (hypo) and 44.9% ASR (anti).

    These results establish that the Trojan association is localized within the core language model representation layers (especially middle-to-late stages) rather than the vision-language alignment interface.

  4. Knowl 4 — Persistence of ImgTrojan Backdoors Under Clean Fine-Tuning

    empirical result

    Post-attack supervised fine-tuning (SFT) with 10,000 clean visual instruction samples from the LLaVA dataset fails to eliminate the implanted ImgTrojan backdoor (at 0.01 poison ratio).

    • For the AntiGPT (anti) prompt, the two-round ASR remains at 20.3% (and direct one-round ASR at 18.9%) after clean SFT.
    • For the Hypothetical Response (hypo) prompt, clean fine-tuning increases attack effectiveness: the two-round ASR rises from 28.1% to 39.6% (+11.5% absolute gain), and the direct one-round ASR rises from 0.4% to 28.4% (+28.0% absolute gain).

    This amplification is attributed to general conversational tuning enhancing the model's instruction-following compliance, thereby reducing refusal behavior on malicious queries.

  5. Knowl 5 — Evasion of CLIP Filtering and Reward-Model Defenses

    empirical result

    Poisoned image-caption pairs constructed by concatenating the original clean caption after the jailbreak prompt evade standard automated filtering mechanisms:

    1. CLIP Image-Text Similarity Filtering: Using CLIP (ViT-B/32) with a standard acceptance threshold of 0.3, 78.07% of poisoned caption-image pairs successfully pass the filter. The average decrease in CLIP similarity score between clean and poisoned pairs is only 0.028 (with a standard deviation of 0.038).
    2. Reward Model Toxicity Detection: Toxicity scoring using reward-model-deberta-v3-large-v2 shows that the distribution of poisoned text scores overlaps heavily with clean samples (unpoisoned range: [-2.55, 7.73]; hypo + description: [-2.96, 1.54]; anti + description: [-4.10, 0.11]).

    Because even a 99% detection rate per sample leaves a (1−0.99100)=63.4%(1 - 0.99^{100}) = 63.4\% chance of missing at least one poisoned sample in 100 injected pairs, and a single poisoned sample suffices to trigger jailbreaking, existing filtering defenses are insufficient.

  6. Knowl 6 — Transferability and Prompt Specificity in Cross-Modality Attacks

    empirical result

    Cross-attack evaluations on LLaVA-v1.5 7B poisoned at a 0.01 ratio reveal strict prompt-matching dependencies:

    • Text-based Jailbreak Prompts: A model poisoned with anti achieves 81.2% ASR when evaluated with the text anti prompt, but only 12.2% when evaluated with hypo. Similarly, a model poisoned with hypo achieves 22.4% ASR with the hypo text prompt versus 14.0% with anti.
    • OCR-based Attacks: Embedding the training-matched prompt on an image yields higher success rates (100.0% for anti-poisoned with OCR-anti vs. 40.0% with OCR-hypo; 50.0% for hypo-poisoned with OCR-hypo vs. 0.0% with OCR-anti).
    • Adversarial Perturbations: Poisoning does not heighten vulnerability to standard gradient-optimized visual adversarial examples; both anti- and hypo-poisoned models retain a baseline ASR of 20.0% against visual adversarial inputs.
  7. Knowl 7 — Dual-Metric Evaluation Framework for Multimodal Jailbreaking

    experimental setup

    The evaluation pipeline assesses both attack efficacy and stealthiness across multimodal tasks:

    1. Attack Success Rate (ASR): Evaluated over a curated benchmark of harmful instructions filtered to ensure at least one baseline method can succeed. Model responses to queries prefixed with the poisoned image are classified as 'Harmful' or 'Safe' by gpt-3.5-turbo following a structured Safety Annotation Guideline (SAG).
      • Inter-rater agreement between human annotators on 30 random cases showed 100% consensus.
      • ChatGPT annotations exhibited 91.0% alignment with human judgment.
      • Cross-validation against an open-source classifier (Llama-Guard-3-8B) yielded a Krippendorff's alpha of 0.75.
    2. Clean Stealthiness Metrics: BLEU and CIDEr scores are computed on 1,023 unpoisoned test image-caption pairs from the GPT4V dataset to monitor whether visual captioning capability degrades or leaks JBP tokens.
    3. Downstream VQA: Evaluated on 1,000 randomly selected question-answer pairs from the VQAv2 validation set (VQAv21000\text{VQAv2}_{1000}) to confirm natural language and visual reasoning preservation.
  8. Knowl 8 — ImgTrojan Attack Performance on Qwen-VL-Chat and Role Formatting Defense

    empirical result

    Applying ImgTrojan (0.01 poison ratio) to Qwen-VL-Chat (7B) yields a 30.1% ASR under two-round conversation with the hypo prompt (and 23.1% with anti), compared to a vanilla baseline of 11.3%.

    While this represents an approximate 20% absolute gain over vanilla input, the attack success is substantially lower than on LLaVA-v1.5 7B (which reached 83.5%). This discrepancy is attributed to Qwen-VL-Chat's ChatML template structure, which wraps conversation turns in explicit role-delimiting tokens (<|im_start|> and <|im_end|>). These boundary markers help the internal alignment mechanism distinguish between instructions provided by the user versus text produced within the model's own prior response.

  9. Knowl 9 — Limitations of the ImgTrojan Study

    limitation

    The empirical findings of the ImgTrojan study are bounded by several experimental constraints:

    1. Model Architectures: Primary experiments were conducted on LLaVA-v1.5 and Qwen-VL-Chat; other VLM architectures and vision backbones may exhibit differing degrees of vulnerability.
    2. Dataset Scale: Experiments utilized an instruction-tuning dataset of approximately 10,000 image-caption pairs with up to 92 poisoned examples, which is smaller than web-scale pre-training datasets (e.g., LAION-5B).
    3. Parameter-Efficient Tuning: LoRA was used for fine-tuning during attack implementation to manage computational costs, which might not fully mirror the dynamics of full-parameter, large-scale pre-training.
    4. Defense Scope: Only standard CLIP-similarity filtering and reward-model toxicity checks were thoroughly tested as defenses.

Coverage note — None was omitted; all key contributions, mathematical objectives, empirical tables, architectural loci analyses, transferability studies, evaluation setups, and limitations are fully covered.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan. 2022. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
  2. 2.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-vl: A frontier large vision-language model with versatile abilities. ArXiv preprint, abs/2308.12966.
  3. 3.Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023. Jailbreaking black box large language models in twenty queries. ArXiv preprint, abs/2310.08419.
  4. 4.Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2023. Sharegpt4v: Improving large multi-modal models with better captions. ArXiv preprint, abs/2311.12793.
  5. 5.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  6. 6.Angela Fan, Edouard Grave, and Armand Joulin. 2020. Reducing transformer depth on demand with structured dropout. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020.
  7. 7.Jiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye, Wanjun Zhong, Yufei Wang, Lanqing Hong, Jianhua Han, Hang Xu, Zhenguo Li, and Lingpeng Kong. 2023. G-llava: Solving geometric problem with multi-modal large language model.
  8. 8.Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela. 2021. Gradient-based adversarial attacks against text transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5747–5757.
  9. 9.Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. 2023. Llama guard: Llm-based input-output safeguard for human-ai conversations. ArXiv preprint, abs/2312.06674.
  10. 10.Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M. Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh. 2023. OBELICS: an open web-scale filtered dataset of interleaved image-text documents. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023.
  11. 11.Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. 2024a. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models. ArXiv preprint, abs/2407.07895.
  12. 12.Lei Li, Yuqi Wang, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, and Qi Liu. 2024b. Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14369–14387.
  13. 13.Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, and Qi Liu. 2024c. VLFeedback: A large-scale AI feedback dataset for large vision-language models alignment. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 6227–6246.
  14. 14.Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu, Xu Sun, Lingpeng Kong, and Qi Liu. 2023. M3 IT: A large-scale dataset towards multimodal multilingual instruction tuning. ArXiv preprint, abs/2306.04387.
  15. 15.Mukai Li, Lei Li, Yuwei Yin, Masood Ahmed, Zhenguang Liu, and Qi Liu. 2024d. Red teaming visual language models. ArXiv preprint, abs/2401.12915.
  16. 16.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023a. Improved baselines with visual instruction tuning. ArXiv preprint, abs/2310.03744.
  17. 17.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023b. Visual instruction tuning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023.
  18. 18.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023c. Visual instruction tuning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023.
  19. 19.Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2023d. Autodan: Generating stealthy jailbreak prompts on aligned large language models. ArXiv preprint, abs/2310.04451.
  20. 20.Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023e. Jailbreaking chatgpt via prompt engineering: An empirical study. ArXiv preprint, abs/2305.13860.
  21. 21.OpenAI. 2023. Gpt-4v(ision) system card.
  22. 22.OpenAssistant. 2023. reward-model-deberta-v3-large-v2. https://huggingface.co/OpenAssistant/reward-model-deberta-v3-large-v2. Accessed: 2025-02-05.
  23. 23.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318.
  24. 24.Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Mengdi Wang, and Prateek Mittal. 2023. Visual adversarial examples jailbreak large language models. ArXiv preprint, abs/2306.13213.
  25. 25.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763.
  26. 26.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. 2022. LAION-5B: an open large-scale dataset for training next generation image-text models. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
  27. 27.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. ArXiv preprint, abs/2111.02114.
  28. 28.Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh. 2023. Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models.
  29. 29.Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, Kurt Keutzer, and Trevor Darrell. 2023. Aligning large multi-modal models with factually augmented rlhf. ArXiv preprint, abs/2309.14525.
  30. 30.Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al. 2024. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. ArXiv preprint, abs/2406.16860.
  31. 31.Haoqin Tu, Chenhang Cui, Zijun Wang, Yiyang Zhou, Bingchen Zhao, Junlin Han, Wangchunshu Zhou, Huaxiu Yao, and Cihang Xie. 2023. How many unicorns are in this image? a safety evaluation benchmark for vision llms. ArXiv preprint, abs/2311.16101.
  32. 32.Pavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli, and Oncel Tuzel. 2024. Mobileclip: Fast image-text models through multi-modal reinforced training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15963–15974.
  33. 33.Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pages 4566–4575.
  34. 34.Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. ArXiv preprint, abs/2409.12191.
  35. 35.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How does LLM safety training fail? In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023.
  36. 36.Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An instruction-tuned audio-visual language model for video understanding. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 543–553.
  37. 37.Xuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du, Lei Li, Yu-Xiang Wang, and William Yang Wang. 2024. Weak-to-strong jailbreaking on large language models. ArXiv preprint, abs/2401.17256.
  38. 38.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. ArXiv preprint, abs/2304.10592.
  39. 39.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. ArXiv preprint, abs/2307.15043.

Citation

MLA
Tao, X., et al. “ImgTrojan: Jailbreaking Vision-Language Models with ONE Image”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 7048–63, https://doi.org/10.18653/v1/2025.naacl-long.360.
APA
Tao, X., Zhong, S., Li, L., Liu, Q., & Kong, L. (2025). ImgTrojan: Jailbreaking Vision-Language Models with ONE Image. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7048–7063. https://doi.org/10.18653/v1/2025.naacl-long.360
Chicago
Tao, X., S. Zhong, L. Li, Q. Liu, and L. Kong. 2025. “ImgTrojan: Jailbreaking Vision-Language Models with ONE Image”. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7048–63. https://doi.org/10.18653/v1/2025.naacl-long.360.
Harvard
Tao, X. et al. (2025) “ImgTrojan: Jailbreaking Vision-Language Models with ONE Image”, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7048–7063. Available at: https://doi.org/10.18653/v1/2025.naacl-long.360.
Vancouver
1. Tao X, Zhong S, Li L, Liu Q, Kong L (2025) ImgTrojan: Jailbreaking Vision-Language Models with ONE Image. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 7048–7063

BibTeX

@inproceedings{Tao_2025, title={ImgTrojan: Jailbreaking Vision-Language Models with ONE Image}, url={http://dx.doi.org/10.18653/v1/2025.naacl-long.360}, DOI={10.18653/v1/2025.naacl-long.360}, booktitle={Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)}, publisher={Association for Computational Linguistics}, author={Tao, Xijia and Zhong, Shuai and Li, Lei and Liu, Qi and Kong, Lingpeng}, year={2025}, pages={7048–7063} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/