RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts
Han LiuYuhao WuShixuan ZhaiBo YuanNing Zhang
Develops a genetic-algorithm-based optimization framework that generates natural, stealthy adversarial text prompts capable of reliably producing target images across diverse text-to-image models in both white-box and black-box settings.
Modern text-to-image artificial intelligence models can generate highly realistic imagery from natural language descriptions, but their rapid deployment raises significant security and safety concerns. Malicious users can exploit these systems to produce harmful content, including non-consensual imagery, disinformation, and graphic material. While service providers deploy text-based moderation filters to block harmful prompts, these defenses assume that malicious prompts must explicitly describe the target output. Understanding whether attackers can reliably evade content filters using benign-looking language is essential for securing generative AI platforms.
The article systematically evaluates the adversarial robustness of text-to-image generation models under both fully accessible (white-box) and query-only (black-box) conditions. It demonstrates an attack framework, named RIATIG, which crafts targeted adversarial text prompts that appear natural to human reviewers and automated text filters while reliably inducing models to generate specific, unrelated target images.
The researchers formulated prompt creation as a genetic optimization problem that iteratively refines sentences using mutation and crossover operations. The approach selects important words to mutate using gradient metrics in white-box settings or word-deletion impact in black-box settings, modifying words via minor typos, visually similar character swaps, and context-aware synonym substitutions. Semantic image similarity guided the optimization toward the visual target while maintaining semantic dissimilarity from the target description. The authors evaluated this method on six widely used text-to-image models—including open-source architectures like AttnGAN, DM-GAN, and DF-GAN, as well as large-scale systems including DALL-E mini, DALL-E 2, and Imagen—using standard image-text benchmark datasets and comparing against five baseline attack methods.
The core findings demonstrate significant vulnerabilities across all evaluated generative systems. RIATIG achieved near-perfect attack success rates across tested models, reaching an 80% to 100% success rate in matching target descriptions and outperforming existing baseline attacks that achieved at most 40% to 90% under similar black-box constraints. In addition, the generated adversarial prompts exhibited significantly lower perplexity scores—often by a factor of 5 to 10 compared to baseline approaches—indicating substantially higher fluency and naturalness that evade human suspicion. Ablation experiments verified that the attack framework remains robust across diverse target images and scales consistently across extended trials.
These findings reveal that current safety moderation strategies relying primarily on text filtering are fundamentally insufficient to prevent the generation of unauthorized or harmful images. Because a natural prompt about an innocuous subject can be manipulated to generate entirely distinct visual content, platforms face major compliance, reputational, and safety risks. Existing text-level safeguards can be easily bypassed without degrading the visual output quality expected by an attacker.
To mitigate these risks, platform operators and developers should transition toward multi-layered defenses rather than relying solely on text-input filtering. While basic rule-based text checkers like grammar filters can catch simple typos, the evaluation showed that 20% of adversarial prompts bypassed standard tools entirely. Organizations should implement output-side image moderation filters to detect policy-violating imagery before it reaches the user, alongside exploring adversarial training where models are exposed to perturbed text prompts during training. Because adversarial retraining across massive multimodal datasets requires substantial computational resources, further research is necessary to develop lightweight and scalable defense mechanisms.
Decision-makers should interpret these results with confidence regarding the underlying structural vulnerabilities of text-to-image systems, though certain operational constraints apply. The evaluation relied on surrogate models and access credits for proprietary commercial APIs, and testing utilized standard benchmark datasets under controlled prompt settings. Nevertheless, the high attack transferability across diverse model architectures confirms a systemic risk that requires immediate architectural and operational safety enhancements.
- Paper: Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment, Di Jin et al. (2019). Introduces TEXTFOOLER, establishing the word-importance ranking and synonym-substitution principles that RIATIG adapts to craft black-box adversarial text prompts.
- Paper: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding, Chitwan Saharia et al. (2022). Presents Imagen, one of the primary commercial text-to-image foundation models evaluated under RIATIG's adversarial attack framework.
- Paper: Hierarchical Text-Conditional Image Generation with CLIP Latents, Aditya Ramesh et al. (2022). Details the architecture and text-to-image mechanisms of unCLIP (DALL-E 2), which serves as a core target system analyzed in RIATIG.
- Paper: Scaling Autoregressive Models for Content-Rich Text-to-Image Generation, Jiahui Yu et al. (2022). Provides foundational insights into scaling text-to-image architectures and prompt adherence benchmarks that contextualize generative safety evaluations.
- Paper: Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples, Nicolas Papernot et al. (2016). Formalizes the principles of black-box adversarial transferability across substitute models upon which RIATIG relies to attack closed commercial APIs.
- Paper: Black-box Adversarial Attacks with Limited Queries and Information, Andrew Ilyas et al. (2018). Establishes derivative-free optimization techniques under restricted query budgets that inform query-efficient black-box attack design.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). Establishes key methodology for evaluating toxic degeneration and automated moderation failures in natural language generation.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). Introduces automated red-teaming paradigms that motivated the automated prompt optimization approach utilized in RIATIG.
- Paper: Visual Adversarial Examples Jailbreak Aligned Large Language Models, Xiangyu Qi et al. (2024). Extends the multimodal adversarial safety paradigm by demonstrating the reverse threat vector: using visual adversarial examples to jailbreak vision-language models.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). Generalizes automated adversarial prompt optimization into transferable universal suffix attacks against safety-aligned text foundation models.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). Establishes a standardized multimodal benchmark and dynamic defense framework that systematizes evaluation across attacks like RIATIG.
- Paper: Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts, Mikayel Samvelyan et al. (2024). Builds upon automated red-teaming by applying quality-diversity search algorithms to evolve open-ended suites of diverse adversarial prompts.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). Applies iterative black-box optimization to rapidly discover natural, interpretable jailbreak prompts against proprietary commercial models in few queries.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). Analyzes the fundamental failure modes of safety training that explain why semantic and obfuscated adversarial prompts successfully evade guardrails.
- Paper: Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection, Zekun Li et al. (2024). Broadens the investigation of prompt-level vulnerabilities by evaluating model robustness against adversarial instructions injected within reference contexts.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). Demonstrates how fine-tuning pipelines can completely compromise safety alignment, compounding the risks identified in black-box prompt attacks.
