Visual Adversarial Examples Jailbreak Aligned Large Language Models

Xiangyu QiKaixuan HuangAshwinee PandaPeter HendersonMengdi WangPrateek Mittal

article2024AAAI332 citations

Demonstrates that a single continuous visual adversarial perturbation can universally jailbreak safety-aligned multimodal large language models, forcing them to obey diverse harmful text instructions beyond the original optimization scope.

Listen

As foundation artificial intelligence models rapidly integrate multimodal inputs like computer vision alongside text, ensuring their safety guardrails remain intact is a critical real-world priority. Large visual language models are increasingly deployed to execute complex instructions, yet the security implications of combining continuous visual inputs with generative text models have been largely unexamined.

The article evaluates whether visual adversarial examples—crafted input images—can reliably circumvent the safety alignment of multimodal large language models and force them to generate prohibited, harmful content.

To demonstrate this vulnerability, the authors developed a straightforward attack approach rooted in prompt-tuning principles. Using standard projected gradient descent algorithms, they optimized a single adversarial image against a tiny training set of only 66 derogatory sentences. They tested the resulting adversarial images across three open-source visual language models—MiniGPT-4, InstructBLIP, and LLaVA—evaluating attack performance against 40 curated dangerous instructions (spanning identity attacks, disinformation, violence, and existential risks) as well as the standard RealToxicityPrompts benchmark consisting of 1,225 text prompts.

The evaluation revealed that a single visual adversarial image can act as a universal jailbreak, causing safety alignment mechanisms to falter. The attack increased model obedience to harmful instructions from a baseline refusal rate down to obedience rates of roughly 60% to 91% across diverse risk categories. Notably, the attack generalized far beyond the original 66-sentence derogatory dataset, inducing the models to provide instructions for extreme violence and crime, such as step-by-step murder instructions, which were never explicitly optimized during the attack generation. Furthermore, visual attacks proved significantly more potent and computationally efficient to generate than pure text-based attacks, which required roughly 12 times the computational overhead due to the discrete search space. The attacks also transferred successfully across different model architectures in black-box testing, increasing toxic output rates even on models aligned with reinforcement learning from human feedback.

These findings imply that integrating vision into language models substantially widens the attack surface and lowers the barrier for bad actors to bypass safety filters. Cross-modal vulnerabilities expose a fundamental tension between traditional neural network vulnerabilities and modern safety alignment. As multimodal models are integrated into higher-stakes environments, such as robotics or external application interfaces, visual exploits could lead to direct operational and safety failures rather than mere text-generation risks.

Organizations developing or deploying multimodal models should implement layered, input-level defenses rather than relying entirely on core model alignment. The article highlights that input preprocessing methods, such as diffusion-based image purification, can neutralize visual adversarial examples as a plug-and-play defense. However, conventional moderation filters and alignment fine-tuning remain insufficient on their own against optimized cross-modal inputs. Developers must thoroughly red-team systems across all integrated modalities before deployment.

The primary limitations of the study include the reliance on imperfect automated toxicity benchmarks and the incomplete scope of manually curated harm scenarios. While confidence in the demonstrated visual vulnerability is high across open-source architectures and confirmed conceptually by commercial developers, further research is needed to determine how well these cross-modal attacks transfer to fully black-box commercial platforms and to test if purification defenses remain robust against future adaptive attacks.

Cover for Visual Adversarial Examples Jailbreak Aligned Large Language Models

Abstract

Warning: this paper contains data, prompts, and model outputs that are offensive in nature. Recently, there has been a surge of interest in integrating vision into Large Language Models (LLMs), exemplified by Visual Language Models (VLMs) such as Flamingo and GPT-4. This paper sheds light on the security and safety implications of this trend. First, we underscore that the continuous and high-dimensional nature of the visual input makes it a weak link against adversarial attacks, representing an expanded attack surface of vision-integrated LLMs. Second, we highlight that the versatility of LLMs also presents visual attackers with a wider array of achievable adversarial objectives, extending the implications of security failures beyond mere misclassification. As an illustration, we present a case study in which we exploit visual adversarial examples to circumvent the safety guardrail of aligned LLMs with integrated vision. Intriguingly, we discover that a single visual adversarial example can universally jailbreak an aligned LLM, compelling it to heed a wide range of harmful instructions (that it otherwise would not) and generate harmful content that transcends the narrow scope of a ‘few-shot’ derogatory corpus initially employed to optimize the adversarial example. Our study underscores the escalating adversarial risks associated with the pursuit of multimodality. Our findings also connect the long-studied adversarial vulnerabilities of neural networks to the nascent field of AI alignment. The presented attack suggests a fundamental adversarial challenge for AI alignment, especially in light of the emerging trend toward multimodality in frontier foundation models.

Table of Contents

  • Introduction
  • Related Work
  • Adversarial Examples as Jailbreakers
  • Setup
  • Our Attack
  • Implementations of Attackers
  • Evaluating Our Attacks
  • Models
  • Comparing with The Text Attack Counterpart
  • Attacks on Other Models and The Transferability
  • Analyzing Defenses
  • Discussions
  • Conclusion
  • Ethical Statement
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Visual Adversarial Jailbreak Optimization Objective for Vision-Language Models

    model/method

    Vision-Language Models (VLMs) that integrate a visual encoder with an autoregressive Large Language Model (LLM) backbone can be universally jailbroken using a single visual adversarial example xadvx_{\text{adv}}. The attack is conceptually formulated as adversarial prompt tuning: rather than optimizing a prompt to steer the model toward a benign downstream task, the adversary tunes a visual input to transition the LLM into a malicious, unaligned generation mode via few-shot generalization.

    Given a small bootstrap corpus Y={yi}i=1m\mathcal{Y} = \{y_i\}_{i=1}^m of mm harmful text sequences (such as a 66-sentence set of derogatory and anti-human statements), the visual adversarial example xadvx_{\text{adv}} is optimized to minimize the negative log-likelihood of generating Y\mathcal{Y} conditioned on the visual input:

    xadv:=arg⁡min⁡xadv′∈B∑i=1m−log⁡p(yi∣xadv′)x_{\text{adv}} := \arg\min_{x'_{\text{adv}} \in \mathcal{B}} \sum_{i=1}^m -\log p\left(y_i \mid x'_{\text{adv}}\right)

    where:

    • p(yi∣xadv′)p(y_i \mid x'_{\text{adv}}) is the conditional autoregressive probability assigned by the VLM to sequence yiy_i when conditioned on the visual input xadv′x'_{\text{adv}}.
    • B\mathcal{B} denotes the valid visual search space, which is either unconstrained (allowing any valid pixel tensor initialized from noise) or an ℓ∞\ell_\infty-ball B={x:∥x−xbenign∥∞≤ϵ}\mathcal{B} = \{x : \|x - x_{\text{benign}}\|_\infty \le \epsilon\} centered at a benign base image xbenignx_{\text{benign}}.

    At inference time, xadvx_{\text{adv}} is prepended as a visual prefix to an arbitrary harmful textual instruction xharmx_{\text{harm}}, forming the composite input [xadv,xharm][x_{\text{adv}}, x_{\text{harm}}]. The presence of xadvx_{\text{adv}} overrides the safety alignment of the underlying LLM, inducing the model to fulfill instructions across diverse harm categories that it otherwise refuses.

  2. Knowl 2 — Optimization Setup for Continuous Visual Attacks and Discrete Textual Triggers

    experimental setup

    The adversarial optimization framework evaluates both visual (cross-modal) and textual jailbreak triggers on aligned Vision-Language Models:

    1. Visual Attack: Because the continuous image domain is end-to-end differentiable, the optimization is performed via standard Projected Gradient Descent (PGD) directly backpropagating gradients to the input image pixels. The attack is executed for 5000 iterations with a batch size of 8 on the 66-sentence harmful corpus Y\mathcal{Y}. Two configurations are used:

      • Unconstrained attack: Initialized from random uniform noise, where pixel values are bounded only by the valid image range [0,1][0, 1].
      • Constrained attack: Initialized from a benign image (e.g., an image of a panda) with an ℓ∞\ell_\infty-norm perturbation constraint ∥xadv−xbenign∥∞≤ϵ\|x_{\text{adv}} - x_{\text{benign}}\|_\infty \le \epsilon for ϵ∈{16/255,32/255,64/255}\epsilon \in \{16/255, 32/255, 64/255\}.
    2. Textual Counterpart: For direct comparison, the 32 visual embedding tokens (in MiniGPT-4) are substituted with 32 adversarial textual tokens. These tokens are optimized on the same corpus Y\mathcal{Y} using discrete coordinate search (AutoPrompt / HotFlip). Optimization runs for 5000 iterations with a batch size of 8 without stealthiness constraints, incurring approximately 12×12\times greater computational runtime than the continuous visual attack due to discrete token evaluation.

    3. Target Architectures: Attacks are evaluated on MiniGPT-4 (13B Vicuna backbone), InstructBLIP (13B Vicuna backbone), and LLaVA (LLaMA-2-13B-Chat backbone aligned with instruction tuning and iterative RLHF on red-teaming data).

  3. Knowl 3 — Generalization of Visual Adversarial Jailbreaks to Unseen Harm Categories

    data/table

    To evaluate jailbreak efficacy, visual adversarial examples optimized exclusively on a 66-sentence corpus of identity attacks and anti-human statements are paired with 40 manually curated harmful textual instructions spanning four distinct categories: Identity Attack, Disinformation (Disinfo), Violence/Crime, and Malicious Behaviors toward Humanity (X-risk). Model responses from MiniGPT-4 are generated via nucleus sampling (p=0.9p=0.9, temperature =1.0= 1.0, 10 samples per instruction) and evaluated for refusal vs. compliance.

    Attack Variant Identity Attack (%) Disinfo (%) Violence/Crime (%) X-risk (%)
    Benign image (no attack) 26.2 48.9 50.1 20.0
    Adv. image (ϵ=16/255\epsilon = 16/255) 61.5 (+35.3) 58.9 (+10.0) 80.0 (+29.9) 50.0 (+30.0)
    Adv. image (ϵ=32/255\epsilon = 32/255) 70.0 (+43.8) 74.4 (+25.5) 87.3 (+37.2) 73.3 (+53.3)
    Adv. image (ϵ=64/255\epsilon = 64/255) 77.7 (+51.5) 84.4 (+35.5) 81.3 (+31.2) 53.3 (+33.3)
    Adv. image (unconstrained) 78.5 (+52.3) 91.1 (+42.2) 84.0 (+33.9) 63.3 (+43.3)
    Adv. text (unconstrained) 58.5 (+32.3) 68.9 (+20.0) 24.0 (-26.1) 26.7 (+6.7)

    The data shows that visual adversarial examples dramatically increase compliance rates across all harm categories. Crucially, the jailbreak generalizes out-of-distribution to instructions that were never present in the optimization corpus, such as step-by-step guides for murder or arson (Violence/Crime rising from 50.1% to 87.3%) and conspiracy theories (Disinformation rising from 48.9% to 91.1%). The discrete text attack baseline achieves lower success across categories and fails to increase compliance on violence/crime instructions relative to the benign baseline.

  4. Knowl 4 — Toxicity Amplification on RealToxicityPrompts Across Visual and Textual Attacks

    data/table

    The jailbreaking capability of visual adversarial examples was benchmarked on the 1225 challenging prompts from the RealToxicityPrompts dataset using MiniGPT-4. Generated continuations were scored for six toxic attributes using Perspective API and the Detoxify classifier, reporting the percentage of outputs exceeding a toxicity threshold of 0.5 (mean ±\pm standard deviation across 3 runs):

    Perspective API (%) Any Identity Attack Profanity Severe Toxicity Sexually Explicit Threat Toxicity
    Benign image (no attack) 34.8 ±\pm 1.6 2.7 ±\pm 0.2 25.1 ±\pm 1.8 1.5 ±\pm 0.2 12.2 ±\pm 0.6 2.0 ±\pm 0.1 30.5 ±\pm 1.4
    Adv. image (ϵ=16/255\epsilon = 16/255) 53.6 ±\pm 1.0 8.4 ±\pm 0.4 36.6 ±\pm 0.6 6.6 ±\pm 0.5 14.1 ±\pm 0.3 4.7 ±\pm 0.3 48.6 ±\pm 0.6
    Adv. image (ϵ=32/255\epsilon = 32/255) 59.4 ±\pm 1.4 14.6 ±\pm 0.6 39.5 ±\pm 1.2 7.0 ±\pm 0.4 14.9 ±\pm 0.9 6.2 ±\pm 0.2 53.8 ±\pm 1.0
    Adv. image (ϵ=64/255\epsilon = 64/255) 67.2 ±\pm 0.2 15.9 ±\pm 0.2 49.6 ±\pm 0.4 12.2 ±\pm 0.5 16.9 ±\pm 0.4 6.6 ±\pm 0.5 63.1 ±\pm 0.8
    Adv. image (unconstrained) 66.0 ±\pm 1.0 17.4 ±\pm 1.2 43.3 ±\pm 1.6 8.0 ±\pm 0.4 14.6 ±\pm 0.3 7.0 ±\pm 1.0 61.7 ±\pm 1.1
    Adv. text (unconstrained) 54.8 ±\pm 1.5 7.0 ±\pm 0.5 40.4 ±\pm 1.3 4.4 ±\pm 0.1 14.9 ±\pm 0.3 4.8 ±\pm 0.3 49.6 ±\pm 0.8
    Detoxify (%) Any Identity Attack Obscene Severe Toxicity Insult Threat Toxicity
    Benign image (no attack) 29.1 ±\pm 1.0 1.5 ±\pm 0.1 22.4 ±\pm 1.5 0.6 ±\pm 0.1 11.0 ±\pm 0.9 0.9 ±\pm 0.1 28.9 ±\pm 0.9
    Adv. image (ϵ=16/255\epsilon = 16/255) 46.4 ±\pm 1.1 5.0 ±\pm 0.4 33.7 ±\pm 0.6 2.3 ±\pm 0.4 23.6 ±\pm 0.4 2.2 ±\pm 0.1 46.1 ±\pm 1.0
    Adv. image (ϵ=32/255\epsilon = 32/255) 51.3 ±\pm 1.5 9.7 ±\pm 0.4 38.2 ±\pm 1.6 2.7 ±\pm 0.6 26.1 ±\pm 0.6 2.6 ±\pm 0.3 50.9 ±\pm 1.4
    Adv. image (ϵ=64/255\epsilon = 64/255) 61.4 ±\pm 0.8 11.7 ±\pm 0.3 49.3 ±\pm 0.1 5.4 ±\pm 0.5 36.4 ±\pm 0.7 3.2 ±\pm 0.4 61.1 ±\pm 0.7
    Adv. image (unconstrained) 61.0 ±\pm 1.5 10.2 ±\pm 0.6 42.4 ±\pm 1.1 2.6 ±\pm 0.1 32.7 ±\pm 1.2 2.8 ±\pm 0.4 60.7 ±\pm 1.6
    Adv. text (unconstrained) 49.2 ±\pm 1.5 4.1 ±\pm 0.1 37.5 ±\pm 0.5 1.9 ±\pm 0.4 23.0 ±\pm 0.3 2.5 ±\pm 0.2 48.9 ±\pm 1.6

    Visual adversarial examples roughly double the model's rate of generating toxic continuations (e.g., Overall Toxicity on Perspective API rises from 30.5% to 63.1% at ϵ=64/255\epsilon=64/255). While Identity Attack shows the largest relative surge due to the optimization corpus content, all other toxic attributes (Profanity, Severe Toxicity, Obscene, Insult, Threat) also increase, confirming broad safety degradation.

  5. Knowl 5 — Cross-Model Black-Box Transferability of Visual Adversarial Jailbreaks

    data/table

    Visual adversarial examples optimized on a white-box surrogate model transfer successfully to jailbreak distinct target models in a black-box setting. The evaluation measures the percentage of generated outputs displaying at least one toxic attribute (scored >0.5>0.5 via Perspective API) on the RealToxicityPrompts challenging subset:

    Surrogate Model Target: MiniGPT-4 (Vicuna) Target: InstructBLIP (Vicuna) Target: LLaVA (LLaMA-2-Chat)
    Without Attack (Benign) 34.8% 34.2% 9.2%
    MiniGPT-4 67.2% (+32.4%) 57.5% (+23.3%) 17.9% (+8.7%)
    InstructBLIP 52.4% (+17.6%) 61.3% (+27.1%) 20.6% (+11.4%)
    LLaVA 44.8% (+10.0%) 46.5% (+12.3%) 52.3% (+43.1%)

    Key empirical findings from the cross-model evaluation:

    1. Vulnerability of RLHF-Aligned Backbones: LLaVA (based on LLaMA-2-13B-Chat, aligned with iterative RLHF and red-teaming) has a very low benign toxicity baseline (9.2%), yet white-box visual adversarial optimization degrades its safety guardrails to 52.3% (+43.1%).
    2. Black-Box Transferability: Adversarial images optimized on one VLM consistently increase the toxicity of other VLMs without gradient access (e.g., InstructBLIP-derived images increase MiniGPT-4 toxicity by +17.6% and LLaVA toxicity by +11.4%).
  6. Knowl 6 — Optimization and Efficacy Disparity Between Visual and Textual Jailbreak Triggers

    empirical result

    A structural asymmetry exists between visual and textual adversarial optimization in Vision-Language Models:

    1. Search Space Dimension and Continuity: A 3×224×2243 \times 224 \times 224 image input processed into 32 tokens provides 2563×224×224≈10362507256^{3 \times 224 \times 224} \approx 10^{362507} possible pixel configurations in a continuous, differentiable space. By contrast, a 32-token text sequence over a 10410^4-word vocabulary affords 104×32=1012810^{4 \times 32} = 10^{128} combinations in a discrete space.
    2. Loss Minimization: Continuous Projected Gradient Descent on visual inputs achieves lower negative log-likelihood loss on the optimization corpus Y\mathcal{Y} than discrete coordinate search (AutoPrompt / HotFlip) on text tokens across 5000 iterations. Even visual attacks restricted by tight ℓ∞\ell_\infty bounds (ϵ=16/255\epsilon = 16/255) minimize the loss more effectively than unconstrained discrete text attacks.
    3. Computational Cost: Due to the combinatorial evaluation required for discrete token swaps, the textual attack requires roughly 12×12\times the computational runtime of the visual PGD attack while yielding lower jailbreak success and toxicity scores.
  7. Knowl 7 — Feasibility and Limitations of Defenses Against Multimodal Jailbreaks

    model/method

    Existing defense paradigms against adversarial machine learning present distinct operational trade-offs when applied to multimodal large language models:

    1. Adversarial Training and Certification: Methods such as PGD adversarial training or randomized smoothing are computationally prohibitive at LLM scale and are primarily formulated for discrete classification outputs rather than open-ended autoregressive token generation. Furthermore, certified radii assume small perturbation bounds, whereas jailbreak images can operate in unconstrained or large-perturbation regimes.
    2. Input Purification (DiffPure): Preprocessing images with diffusion-based purification—adding forward diffusion noise and utilizing a reverse denoising process to project perturbed inputs back onto the natural image data manifold—acts as a model-agnostic plug-and-play module that neutralizes static visual adversarial jailbreak prompts without requiring VLM fine-tuning.
    3. Filtering and Post-Processing Moderation: External toxicity detection APIs (e.g., Perspective API, OpenAI Moderation API) and secondary LLM content moderators can filter harmful prompts or completions for hosted web endpoints. However, they introduce latency, risk false-positive censorship, and cannot protect open-source models executed offline by an adversary.
  8. Knowl 8 — Limitations of Multimodal Jailbreak Evaluations

    limitation

    The evaluation of visual adversarial jailbreaks is subject to several methodological limitations:

    1. Incompleteness of Open-Ended Output Evaluation: Because LLM outputs are free-form and open-ended, benchmark datasets and curated prompts cannot exhaustively capture all potential vectors of misuse or catastrophic harm.
    2. Subjectivity and Classifier Noise: Manual evaluation of jailbreak success lacks a universally standardized benchmark for refusal boundary violations. Automated toxicity evaluations (e.g., Perspective API and Detoxify) exhibit known classification noise, dialect biases, and false positives.
    3. Model Capability Scale: Empirical demonstrations were conducted on 13B parameter open-source models (Vicuna and LLaMA-2 backbones), whose baseline reasoning and downstream agency are limited compared to frontier multi-billion parameter proprietary systems.

Coverage note — No substantial contributed material was omitted from the main paper text; the eight knowls cover the attack formulation, optimization implementations, human evaluation across harm categories, RealToxicityPrompts benchmark evaluation, cross-model transferability, visual vs. textual comparison, defense analysis, and limitations.

References

  1. 1.Abdelnabi, S.; Greshake, K.; Mishra, S.; Endres, C.; Holz, T.; and Fritz, M. 2023. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, 79–90.
  2. 2.Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35: 23716–23736.
  3. 3.Alzantot, M.; Sharma, Y.; Elgohary, A.; Ho, B.-J.; Srivastava, M.; and Chang, K.-W. 2018. Generating natural language adversarial examples. arXiv preprint arXiv:1804.07998.
  4. 4.Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, 2425–2433.
  5. 5.Athalye, A.; Carlini, N.; and Wagner, D. 2018. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International conference on machine learning, 274–283. PMLR.
  6. 6.Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; Joseph, N.; Kadavath, S.; Kernion, J.; Conerly, T.; El-Showk, S.; Elhage, N.; Hatfield-Dodds, Z.; Hernandez, D.; Hume, T.; Johnston, S.; Kravec, S.; Lovitt, L.; Nanda, N.; Olsson, C.; Amodei, D.; Brown, T.; Clark, J.; McCandlish, S.; Olah, C.; Mann, B.; and Kaplan, J. 2022. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv:2204.05862.
  7. 7.Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M. S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
  8. 8.Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; Ding, T.; Driess, D.; Dubey, A.; Finn, C.; et al. 2023. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv preprint arXiv:2307.15818.
  9. 9.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901.
  10. 10.Cao, Y.; Wang, N.; Xiao, C.; Yang, D.; Fang, J.; Yang, R.; Chen, Q. A.; Liu, M.; and Li, B. 2021. Invisible for both camera and lidar: Security of multi-sensor fusion based perception in autonomous driving under physical-world attacks. In 2021 IEEE Symposium on Security and Privacy (SP), 176–194. IEEE.
  11. 11.Carlini, N.; and Wagner, D. 2017. Adversarial examples are not easily detected: Bypassing ten detection methods. In Proceedings of the 10th ACM workshop on artificial intelligence and security, 3–14.
  12. 12.Carlini, N.; and Wagner, D. 2018. Audio adversarial examples: Targeted attacks on speech-to-text. In 2018 IEEE security and privacy workshops (SPW), 1–7. IEEE.
  13. 13.Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality.
  14. 14.Cohen, J.; Rosenfeld, E.; and Kolter, Z. 2019. Certified adversarial robustness via randomized smoothing. In international conference on machine learning, 1310–1320. PMLR.
  15. 15.Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. arXiv:2305.06500.
  16. 16.Ebrahimi, J.; Rao, A.; Lowd, D.; and Dou, D. 2017. Hotflip: White-box adversarial examples for text classification. arXiv preprint arXiv:1712.06751.
  17. 17.Fang, Y.; Wang, W.; Xie, B.; Sun, Q.; Wu, L.; Wang, X.; Huang, T.; Wang, X.; and Cao, Y. 2023. Eva: Exploring the limits of masked visual representation learning at scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 19358–19369.
  18. 18.Ganguli, D.; Hernandez, D.; Lovitt, L.; Askell, A.; Bai, Y.; Chen, A.; Conerly, T.; Dassarma, N.; Drain, D.; Elhage, N.; et al. 2022a. Predictability and surprise in large generative models. In 2022 ACM Conference on Fairness, Accountability, and Transparency, 1747–1764.
  19. 19.Ganguli, D.; Lovitt, L.; Kernion, J.; Askell, A.; Bai, Y.; Kadavath, S.; Mann, B.; Perez, E.; Schiefer, N.; Ndousse, K.; et al. 2022b. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858.
  20. 20.Gehman, S.; Gururangan, S.; Sap, M.; Choi, Y.; and Smith, N. A. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462.
  21. 21.Girdhar, R.; El-Nouby, A.; Liu, Z.; Singh, M.; Alwala, K. V.; Joulin, A.; and Misra, I. 2023. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 15180–15190.
  22. 22.Goodfellow, I. J.; Shlens, J.; and Szegedy, C. 2014. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572.
  23. 23.Hanu, L.; and Unitary team. 2020. Detoxify. Github. https://github.com/unitaryai/detoxify.
  24. 24.Helbling, A.; Phute, M.; Hull, M.; and Chau, D. H. 2023. LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked. arXiv:2308.07308.
  25. 25.Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33: 6840–6851.
  26. 26.Holtzman, A.; Buys, J.; Du, L.; Forbes, M.; and Choi, Y. 2019. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751.
  27. 27.Jones, E.; Dragan, A.; Raghunathan, A.; and Steinhardt, J. 2023. Automatically Auditing Large Language Models via Discrete Optimization. arXiv preprint arXiv:2303.04381.
  28. 28.Kenton, Z.; Everitt, T.; Weidinger, L.; Gabriel, I.; Mikulik, V.; and Irving, G. 2021. Alignment of language agents. arXiv preprint arXiv:2103.14659.
  29. 29.Lester, B.; Al-Rfou, R.; and Constant, N. 2021. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691.
  30. 30.Li, L.; Xie, T.; and Li, B. 2023. SoK: Certified Robustness for Deep Neural Networks. In 44th IEEE Symposium on Security and Privacy, SP 2023, San Francisco, CA, USA, 22-26 May 2023. IEEE.
  31. 31.Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023. Visual Instruction Tuning.
  32. 32.Madry, A.; Makelov, A.; Schmidt, L.; Tsipras, D.; and Vladu, A. 2017. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083.
  33. 33.Mehrabi, N.; Beirami, A.; Morstatter, F.; and Galstyan, A. 2022. Robust Conversational Agents against Imperceptible Toxicity Triggers. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2831–2847. Seattle, United States: Association for Computational Linguistics.
  34. 34.Nie, W.; Guo, B.; Huang, Y.; Xiao, C.; Vahdat, A.; and Anandkumar, A. 2022. Diffusion models for adversarial purification. arXiv preprint arXiv:2205.07460.
  35. 35.OpenAI. 2022. Introducing ChatGPT. https://openai.com/blog/chatgpt.
  36. 36.OpenAI. 2023a. Forecasting potential misuses of language models for disinformation campaigns and how to reduce risk. https://openai.com/research/forecasting-misuse. [Online; accessed 4-Apr-2023].
  37. 37.OpenAI. 2023b. GPT-4 Technical Report. arXiv:2303.08774.
  38. 38.OpenAI. 2023c. GPT-4V(ision) system card. https://openai.com/research/gpt-4v-system-card.
  39. 39.Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35: 27730–27744.
  40. 40.Patil, S. G.; Zhang, T.; Wang, X.; and Gonzalez, J. E. 2023. Gorilla: Large language model connected with massive apis. arXiv preprint arXiv:2305.15334.
  41. 41.Perez, E.; Huang, S.; Song, F.; Cai, T.; Ring, R.; Aslanides, J.; Glaese, A.; McAleese, N.; and Irving, G. 2022. Red teaming language models with language models. arXiv preprint arXiv:2202.03286.
  42. 42.Pichai, S. 2023. An important next step on our AI journey. https://blog.google/technology/ai/bard-google-ai-search-updates/.
  43. 43.Pichai, S.; and Hassabis, D. 2023. Introducing Gemini: our largest and most capable AI model. https://blog.google/technology/ai/google-gemini-ai/.
  44. 44.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning, 8748–8763. PMLR.
  45. 45.Schick, T.; Udupa, S.; and Schutze, H. 2021. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp. Transactions of the Association for Computational Linguistics, 9: 1408–1424.
  46. 46.ShareGPT.com. 2023. ShareGPT: Share your wildest ChatGPT conversations with one click. https://sharegpt.com/.
  47. 47.Shin, T.; Razeghi, Y.; Logan IV, R. L.; Wallace, E.; and Singh, S. 2020. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980.
  48. 48.Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I.; and Fergus, R. 2013. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199.
  49. 49.Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Roziere, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  50. 50.Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  51. 51.Tramer, F. 2022. Detecting adversarial examples is (nearly) as hard as classifying them. In International Conference on Machine Learning, 21692–21702. PMLR.
  52. 52.Wallace, E.; Feng, S.; Kandpal, N.; Gardner, M.; and Singh, S. 2019. Universal adversarial triggers for attacking and analyzing NLP. arXiv preprint arXiv:1908.07125.
  53. 53.Wei, A.; Haghtalab, N.; and Steinhardt, J. 2023. Jailbroken: How Does LLM Safety Training Fail? arXiv preprint arXiv:2307.02483.
  54. 54.Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  55. 55.Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35: 24824–24837.
  56. 56.Welbl, J.; Glaese, A.; Uesato, J.; Dathathri, S.; Mellor, J.; Hendricks, L. A.; Anderson, K.; Kohli, P.; Coppin, B.; and Huang, P.-S. 2021. Challenges in detoxifying language models. arXiv preprint arXiv:2109.07445.
  57. 57.Weng, L.; Goel, V.; and Vallone, A. 2023. Using GPT-4 for content moderation.
  58. 58.Xu, A.; Pathak, E.; Wallace, E.; Gururangan, S.; Sap, M.; and Klein, D. 2021. Detoxifying language models risks marginalizing minority voices. arXiv preprint arXiv:2104.06390.
  59. 59.Zellers, R.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019. From recognition to cognition: Visual commonsense reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 6720–6731.
  60. 60.Zhao, Z.; Dua, D.; and Singh, S. 2017. Generating natural adversarial examples. arXiv preprint arXiv:1710.11342.
  61. 61.Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.

Citation

MLA
Qi, X., et al. “Visual Adversarial Examples Jailbreak Aligned Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2306.13213v2.
APA
Qi, X., Huang, K., Panda, A., Henderson, P., Wang, M., & Mittal, P. (2023). Visual Adversarial Examples Jailbreak Aligned Large Language Models. arXiv. http://arxiv.org/abs/2306.13213v2
Chicago
Qi, X., K. Huang, A. Panda, P. Henderson, M. Wang, and P. Mittal. 2023. “Visual Adversarial Examples Jailbreak Aligned Large Language Models”. arXiv. http://arxiv.org/abs/2306.13213v2.
Harvard
Qi, X. et al. (2023) “Visual Adversarial Examples Jailbreak Aligned Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.13213v2.
Vancouver
1. Qi X, Huang K, Panda A, Henderson P, Wang M, Mittal P (2023) Visual Adversarial Examples Jailbreak Aligned Large Language Models. arXiv

BibTeX

@article{qi2023visual,
  title = {Visual Adversarial Examples Jailbreak Aligned Large Language Models},
  author = {Qi, Xiangyu and Huang, Kaixuan and Panda, Ashwinee and Henderson, Peter and Wang, Mengdi and Mittal, Prateek},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.13213v2},
  eprint = {2306.13213}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF