Visual Adversarial Examples Jailbreak Aligned Large Language Models
Xiangyu QiKaixuan HuangAshwinee PandaPeter HendersonMengdi WangPrateek Mittal
Demonstrates that a single continuous visual adversarial perturbation can universally jailbreak safety-aligned multimodal large language models, forcing them to obey diverse harmful text instructions beyond the original optimization scope.
As foundation artificial intelligence models rapidly integrate multimodal inputs like computer vision alongside text, ensuring their safety guardrails remain intact is a critical real-world priority. Large visual language models are increasingly deployed to execute complex instructions, yet the security implications of combining continuous visual inputs with generative text models have been largely unexamined.
The article evaluates whether visual adversarial examples—crafted input images—can reliably circumvent the safety alignment of multimodal large language models and force them to generate prohibited, harmful content.
To demonstrate this vulnerability, the authors developed a straightforward attack approach rooted in prompt-tuning principles. Using standard projected gradient descent algorithms, they optimized a single adversarial image against a tiny training set of only 66 derogatory sentences. They tested the resulting adversarial images across three open-source visual language models—MiniGPT-4, InstructBLIP, and LLaVA—evaluating attack performance against 40 curated dangerous instructions (spanning identity attacks, disinformation, violence, and existential risks) as well as the standard RealToxicityPrompts benchmark consisting of 1,225 text prompts.
The evaluation revealed that a single visual adversarial image can act as a universal jailbreak, causing safety alignment mechanisms to falter. The attack increased model obedience to harmful instructions from a baseline refusal rate down to obedience rates of roughly 60% to 91% across diverse risk categories. Notably, the attack generalized far beyond the original 66-sentence derogatory dataset, inducing the models to provide instructions for extreme violence and crime, such as step-by-step murder instructions, which were never explicitly optimized during the attack generation. Furthermore, visual attacks proved significantly more potent and computationally efficient to generate than pure text-based attacks, which required roughly 12 times the computational overhead due to the discrete search space. The attacks also transferred successfully across different model architectures in black-box testing, increasing toxic output rates even on models aligned with reinforcement learning from human feedback.
These findings imply that integrating vision into language models substantially widens the attack surface and lowers the barrier for bad actors to bypass safety filters. Cross-modal vulnerabilities expose a fundamental tension between traditional neural network vulnerabilities and modern safety alignment. As multimodal models are integrated into higher-stakes environments, such as robotics or external application interfaces, visual exploits could lead to direct operational and safety failures rather than mere text-generation risks.
Organizations developing or deploying multimodal models should implement layered, input-level defenses rather than relying entirely on core model alignment. The article highlights that input preprocessing methods, such as diffusion-based image purification, can neutralize visual adversarial examples as a plug-and-play defense. However, conventional moderation filters and alignment fine-tuning remain insufficient on their own against optimized cross-modal inputs. Developers must thoroughly red-team systems across all integrated modalities before deployment.
The primary limitations of the study include the reliance on imperfect automated toxicity benchmarks and the incomplete scope of manually curated harm scenarios. While confidence in the demonstrated visual vulnerability is high across open-source architectures and confirmed conceptually by commercial developers, further research is needed to determine how well these cross-modal attacks transfer to fully black-box commercial platforms and to test if purification defenses remain robust against future adaptive attacks.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). This foundational study demonstrates universal and transferable adversarial suffix attacks on aligned LLMs, establishing the optimization-based jailbreak principles adapted by the source into the continuous visual domain.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This paper analyzes the structural failure modes of safety alignment in language models, providing the conceptual foundation for understanding why cross-modal inputs can bypass refusal mechanisms.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). This work introduces the RealToxicityPrompts benchmark, which the source directly employs to evaluate toxic degeneration induced by visual adversarial examples.
- Paper: Delving into Transferable Adversarial Examples and Black-box Attacks, Yanpei Liu et al. (2016). This research provides essential background on the cross-model transferability of continuous adversarial image perturbations, which the source leverages in black-box multimodal jailbreaks.
- Paper: Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey, Naveed Akhtar et al. (2018). This survey details gradient-based perturbation attacks in deep computer vision, establishing the underlying optimization methods utilized to generate visual jailbreak images.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This comprehensive survey outlines the architectural foundations and alignment pipelines of multimodal large language models targeted by the source's visual jailbreaks.
- Paper: ImgTrojan: Jailbreaking Vision-Language Models with ONE Image, Xijia Tao et al. (2025). This paper extends visual jailbreak research by examining poisoned data injection attacks during training to compromise vision-language model safety barriers.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). This work standardizes automated red-teaming evaluations across multimodal and text domains, building on findings of multimodal alignment vulnerabilities to develop robust refusal defenses.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). This study advances the analysis of alignment vulnerabilities by developing adaptive jailbreak attacks tailored across contemporary safety-aligned models.
