Built independently by an author, for readers. Read the story and support ChapterPal

keyword

visual adversarial examples

Visual adversarial examples are images specifically crafted or altered to mislead artificial intelligence systems into producing erroneous, unintended, or unsafe outputs while often appearing normal or benign to human observers. These inputs exploit vulnerabilities in the continuous, high-dimensional space of visual data, where subtle pixel perturbations, typographic text rendered directly onto images, or poisoned visual triggers manipulate the feature representations processed by neural networks. While originally developed to cause misclassifications in standard computer vision tasks, visual adversarial examples also serve as attack vectors in multimodal and vision-language models to bypass safety guardrails, circumvent alignment training, and compel models to generate prohibited or harmful responses.

3 items

ImgTrojan: Jailbreaking Vision-Language Models with ONE Image

ImgTrojan: Jailbreaking Vision-Language Models with ONE Image

Xijia Tao, Shuai Zhong, Lei Li, Qi Liu, Lingpeng Kong

OrganizationsUniversity of Hong Kong

Why you should read this

Reveals how poisoning as few as one training image-text pair allows attackers to bypass safety barriers in vision-language models during inference without degrading standard multimodal performance.

There has been an increasing interest in the alignment of large language models (LLMs) with human values. However, the safety issues of their integration with a vision module, or vision language models (VLMs), remain relatively underexplored. In this paper, we propose a novel jailbreaking attack against VLMs, aiming to bypass their safety barrier when a user inputs harmful instructions. A scenario where our poisoned (image, text) data pairs are included in the training data is assumed. By replacing the original textual captions with malicious jailbreak prompts, our method can perform jailbreak attacks with the poisoned images. Moreover, we analyze the effect of poison ratios and positions of trainable parameters on our attack’s success rate. For evaluation, we design two metrics to quantify the success rate and the stealthiness of our attack. Together with a list of curated harmful instructions, a benchmark for measuring attack efficacy is provided. We demonstrate the efficacy of our attack by comparing it with baseline methods.

Added

2026-09-26

Visual Adversarial Examples Jailbreak Aligned Large Language Models

Visual Adversarial Examples Jailbreak Aligned Large Language Models

Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, Prateek Mittal

OrganizationsPrinceton University

Why you should read this

Demonstrates that a single continuous visual adversarial perturbation can universally jailbreak safety-aligned multimodal large language models, forcing them to obey diverse harmful text instructions beyond the original optimization scope.

Warning: this paper contains data, prompts, and model outputs that are offensive in nature. Recently, there has been a surge of interest in integrating vision into Large Language Models (LLMs), exemplified by Visual Language Models (VLMs) such as Flamingo and GPT-4. This paper sheds light on the security and safety implications of this trend. First, we underscore that the continuous and high-dimensional nature of the visual input makes it a weak link against adversarial attacks, representing an expanded attack surface of vision-integrated LLMs. Second, we highlight that the versatility of LLMs also presents visual attackers with a wider array of achievable adversarial objectives, extending the implications of security failures beyond mere misclassification. As an illustration, we present a case study in which we exploit visual adversarial examples to circumvent the safety guardrail of aligned LLMs with integrated vision. Intriguingly, we discover that a single visual adversarial example can universally jailbreak an aligned LLM, compelling it to heed a wide range of harmful instructions (that it otherwise would not) and generate harmful content that transcends the narrow scope of a ‘few-shot’ derogatory corpus initially employed to optimize the adversarial example. Our study underscores the escalating adversarial risks associated with the pursuit of multimodality. Our findings also connect the long-studied adversarial vulnerabilities of neural networks to the nascent field of AI alignment. The presented attack suggests a fundamental adversarial challenge for AI alignment, especially in light of the emerging trend toward multimodality in frontier foundation models.

Added

2026-09-26

FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, Xiaoyun Wang

Why you should read this

Presents FigStep, a lightweight black-box jailbreak method that bypasses vision-language model safeguards by converting forbidden text into typographic images, exposing critical cross-modal alignment gaps across both open-source and proprietary systems.

Large Vision-Language Models (LVLMs) signify a groundbreaking paradigm shift within the Artificial Intelligence (AI) community, extending beyond the capabilities of Large Language Models (LLMs) by assimilating additional modalities (e.g., images). Despite this advancement, the safety of LVLMs remains adequately underexplored, with a potential overreliance on the safety assurances purportedly by their underlying LLMs. In this paper, we propose FigStep, a straightforward yet effective black-box jailbreak algorithm against LVLMs. Instead of feeding textual harmful instructions directly, FigStep converts the prohibited content into images through typography to bypass the safety alignment. The experimental results indicate that FigStep can achieve an average attack success rate of 82.50% on six promising open-source LVLMs. Not merely to demonstrate the efficacy of FigStep, we conduct comprehensive ablation studies and analyze the distribution of the semantic embeddings to uncover that the reason behind the success of FigStep is the deficiency of safety alignment for visual embeddings. Moreover, we compare FigStep with five text-only jailbreaks and four image-based jailbreaks to demonstrate the superiority of FigStep, i.e., negligible attack costs and better attack performance. Above all, our work reveals that current LVLMs are vulnerable to jailbreak attacks, which highlights the necessity of novel cross-modality safety alignment techniques.

Added

2026-09-26