RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts

Han LiuYuhao WuShixuan ZhaiBo YuanNing Zhang

article2023CVPR62 citations

Develops a genetic-algorithm-based optimization framework that generates natural, stealthy adversarial text prompts capable of reliably producing target images across diverse text-to-image models in both white-box and black-box settings.

Listen

Modern text-to-image artificial intelligence models can generate highly realistic imagery from natural language descriptions, but their rapid deployment raises significant security and safety concerns. Malicious users can exploit these systems to produce harmful content, including non-consensual imagery, disinformation, and graphic material. While service providers deploy text-based moderation filters to block harmful prompts, these defenses assume that malicious prompts must explicitly describe the target output. Understanding whether attackers can reliably evade content filters using benign-looking language is essential for securing generative AI platforms.

The article systematically evaluates the adversarial robustness of text-to-image generation models under both fully accessible (white-box) and query-only (black-box) conditions. It demonstrates an attack framework, named RIATIG, which crafts targeted adversarial text prompts that appear natural to human reviewers and automated text filters while reliably inducing models to generate specific, unrelated target images.

The researchers formulated prompt creation as a genetic optimization problem that iteratively refines sentences using mutation and crossover operations. The approach selects important words to mutate using gradient metrics in white-box settings or word-deletion impact in black-box settings, modifying words via minor typos, visually similar character swaps, and context-aware synonym substitutions. Semantic image similarity guided the optimization toward the visual target while maintaining semantic dissimilarity from the target description. The authors evaluated this method on six widely used text-to-image models—including open-source architectures like AttnGAN, DM-GAN, and DF-GAN, as well as large-scale systems including DALL-E mini, DALL-E 2, and Imagen—using standard image-text benchmark datasets and comparing against five baseline attack methods.

The core findings demonstrate significant vulnerabilities across all evaluated generative systems. RIATIG achieved near-perfect attack success rates across tested models, reaching an 80% to 100% success rate in matching target descriptions and outperforming existing baseline attacks that achieved at most 40% to 90% under similar black-box constraints. In addition, the generated adversarial prompts exhibited significantly lower perplexity scores—often by a factor of 5 to 10 compared to baseline approaches—indicating substantially higher fluency and naturalness that evade human suspicion. Ablation experiments verified that the attack framework remains robust across diverse target images and scales consistently across extended trials.

These findings reveal that current safety moderation strategies relying primarily on text filtering are fundamentally insufficient to prevent the generation of unauthorized or harmful images. Because a natural prompt about an innocuous subject can be manipulated to generate entirely distinct visual content, platforms face major compliance, reputational, and safety risks. Existing text-level safeguards can be easily bypassed without degrading the visual output quality expected by an attacker.

To mitigate these risks, platform operators and developers should transition toward multi-layered defenses rather than relying solely on text-input filtering. While basic rule-based text checkers like grammar filters can catch simple typos, the evaluation showed that 20% of adversarial prompts bypassed standard tools entirely. Organizations should implement output-side image moderation filters to detect policy-violating imagery before it reaches the user, alongside exploring adversarial training where models are exposed to perturbed text prompts during training. Because adversarial retraining across massive multimodal datasets requires substantial computational resources, further research is necessary to develop lightweight and scalable defense mechanisms.

Decision-makers should interpret these results with confidence regarding the underlying structural vulnerabilities of text-to-image systems, though certain operational constraints apply. The evaluation relied on surrogate models and access credits for proprietary commercial APIs, and testing utilized standard benchmark datasets under controlled prompt settings. Nevertheless, the high attack transferability across diverse model architectures confirms a systemic risk that requires immediate architectural and operational safety enhancements.

Cover for RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts

Abstract

The field of text-to-image generation has made remarkable strides in creating high-fidelity and photorealistic images. As this technology gains popularity, there is a growing concern about its potential security risks. However, there has been limited exploration into the robustness of these models from an adversarial perspective. Existing research has primarily focused on untargeted settings, and lacks holistic consideration for reliability (attack success rate) and stealthiness (imperceptibility).

In this paper, we propose RIATIG, a reliable and imperceptible adversarial attack against text-to-image models via inconspicuous examples. By formulating the example crafting as an optimization process and solving it using a genetic-based method, our proposed attack can generate imperceptible prompts for text-to-image generation models in a reliable way. Evaluation of six popular text-to-image generation models demonstrates the efficiency and stealthiness of our attack in both white-box and black-box settings. To allow the community to build on top of our findings, we’ve made the artifacts available1.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. System and Threat Model
  • 4. Methodology
  • 4.1. Problem Formulation
  • 4.2. Challenges and Overview
  • 4.3. Similarity Measurement
  • 4.4. Genetic-based Optimization
  • 4.5. Sample Quality Improvement
  • 5. Experiments
  • 5.1. Experiment Settings
  • 5.2. Evaluation Metrics
  • 5.3. Evaluation Results
  • 5.4. Ablation Study
  • 6. Discussion
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — Targeted adversarial objective and threat model

    model/method

    RIATIG attacks a text-to-image generator G:X→YG:\mathcal{X}\rightarrow\mathcal{Y}, where X\mathcal{X} is the space of text prompts and Y\mathcal{Y} is the space of generated images. The generator consists of a tokenizer and text encoder, a transformation network that maps text encodings to image encodings, and an image decoder. Given a target image yt∈Yy_t\in\mathcal{Y} with an associated target prompt xt∈Xx_t\in\mathcal{X}, RIATIG searches for an adversarial prompt x∗x^* whose generated image resembles yty_t while the prompt itself is semantically different from xtx_t:

    x∗=arg⁡max⁡x∗∈X  Si(G(x∗),yt)subject toDt(x∗,xt)>η.x^*=\underset{x^*\in\mathcal{X}}{\arg\max}\;S_i\bigl(G(x^*),y_t\bigr)\quad\text{subject to}\quad D_t(x^*,x_t)>\eta.

    Here, SiS_i measures image-semantic similarity, DtD_t measures text-semantic distance, and η\eta is the minimum required text distance. An attack is considered successful when the generated image similarity also exceeds a threshold ϵ\epsilon, so that Si(G(x∗),yt)>ϵS_i(G(x^*),y_t)>\epsilon and Dt(x∗,xt)>ηD_t(x^*,x_t)>\eta. In the white-box setting, the attacker knows the target model architecture and parameters; in the black-box setting, the attacker can only submit prompts and observe generated images, without confidence scores, model internals, or training data.

  2. Knowl 2 — CLIP-based image fitness and text-distance functions

    equation

    RIATIG uses the pretrained CLIP image encoder and text encoder to measure the two objectives in its attack. Let Ei(⋅)E_i(\cdot) map an image to a CLIP image-embedding vector, let Et(⋅)E_t(\cdot) map a text prompt to a CLIP text-embedding vector, let ⋅\cdot denote the vector dot product, and let ∥⋅∥\|\cdot\| denote the Euclidean norm. For a generated image G(x)G(x), target image yy, and prompts xx and x∗x^*, the image similarity and text semantic distance are

    Si(G(x),y)=Ei(G(x))⋅Ei(y)∥Ei(G(x))∥ ∥Ei(y)∥,S_i\bigl(G(x),y\bigr)=\frac{E_i(G(x))\cdot E_i(y)}{\|E_i(G(x))\|\,\|E_i(y)\|}, Dt(x,x∗)=1−Et(x)⋅Et(x∗)∥Et(x)∥ ∥Et(x∗)∥.D_t(x,x^*)=1-\frac{E_t(x)\cdot E_t(x^*)}{\|E_t(x)\|\,\|E_t(x^*)\|}.

    The cosine image similarity serves as the genetic algorithm's fitness score. A larger SiS_i means that the generated and target images are closer in CLIP's visual-semantic space, whereas a larger DtD_t means that the two prompts are less semantically related.

  3. Knowl 3 — Genetic-based RIATIG attack algorithm

    algorithm

    RIATIG solves the discrete prompt-search problem with a population-based genetic procedure. Its inputs are an original prompt xx, a target prompt xtx_t with target image yty_t, initial mutation probability pmp_m, initial population size MM, per-generation population size NN, elite size KK, maximum generation count cmax⁡c_{\max}, and mutation-update parameters α\alpha and β\beta. Let GG be the target image generator, SiS_i the CLIP image similarity, DtD_t the CLIP text distance, and ϵ\epsilon and η\eta the image-success and text-dissimilarity thresholds. The algorithm returns the fine-tuned adversarial set T∗T^*.

    Input: Original prompt x, target prompt xt, target image yt, pm, M, N, K, cmax, alpha, beta, epsilon, eta
    Output: Fine-tuned adversarial prompt set T*
    Initialize population P0 with M mutations of x using mutation probability pm
    Initialize adversarial set T as empty and previous best score s_previous
    for generation c from 0 to cmax do
        Evaluate every prompt p in Pc by score(p) = Si(G(p), yt)
        Sort Pc in descending order of score
        Keep the K highest-scoring prompts as elite set Ec with scores Fc
        Set sc to the highest score in Fc
        Update mutation probability using pm = alpha * previous_pm + beta / abs(sc - s_previous)
        Convert elite scores Fc into parent-selection probabilities with Softmax
        Initialize next population Pc+1 as empty
        for i from 1 to N do
            Sample parent1 and parent2 from Ec according to the elite probabilities
            Create childi by crossover of parent1 and parent2
            Mutate childi with the updated mutation probability pm
            Add childi to Pc+1
            if Si(G(childi), yt) > epsilon and Dt(childi, xt) > eta then
                Add childi to T
            end if
        end for
        Set s_previous to sc and previous_pm to pm
    end for
    Apply the quality fine-tuning procedure to every prompt in T
    Return the fine-tuned set T*

    Mutation explores the prompt space, crossover combines traits from high-fitness prompts, and fitness-proportional elite sampling preserves useful traits. The loop records a prompt as an adversarial example only when its generated image is sufficiently similar to the target image and its text is sufficiently dissimilar from the target prompt.

  4. Knowl 4 — Importance-guided natural prompt mutation

    model/method

    RIATIG mutates selected sentences rather than directly optimizing discrete tokens. A sentence is selected with probability pmp_m, and the word to mutate is sampled according to an importance distribution. For a sentence containing LL words, let ene_n be the embedding of word xnx_n, let ee be the full sentence's embedding sequence, and let e∖ne_{\setminus n} be the sequence after deleting word xnx_n. In the white-box setting, word importance is measured by the gradient of image fitness with respect to the word embedding:

    gn=∂Si(G(e),yt)∂en.g_n=\frac{\partial S_i(G(e),y_t)}{\partial e_n}.

    In the black-box setting, it is estimated from the change in fitness caused by deleting the word:

    gn=Si(G(e),yt)−Si(G(e∖n),yt).g_n=S_i(G(e),y_t)-S_i\bigl(G(e_{\setminus n}),y_t\bigr).

    The normalized probability of selecting word xnx_n is

    In=exp⁡(∥gn∥)∑j=1Lexp⁡(∥gj∥).I_n=\frac{\exp(\|g_n\|)}{\sum_{j=1}^{L}\exp(\|g_j\|)}.

    The sentence mutation probability is adapted across generations. If scs_c is the maximum image-fitness score in generation cc, then the next probability is updated as

    pm,c=αpm,c−1+β∣sc−sc−1∣,p_{m,c}=\alpha p_{m,c-1}+\frac{\beta}{|s_c-s_{c-1}|},

    where α\alpha and β\beta are scaling factors. This increases exploration when fitness improvement is small while retaining momentum from the previous mutation probability.

    The selected word is modified using either character-level or lexical mutations. Character-level mutations insert an extra space, swap two non-boundary characters, or delete one non-boundary character. Lexical mutations replace characters with visually similar LEET-Speak symbols, or replace words with semantically similar words retrieved using Word2vec. Candidate substitutions are restricted by a distance threshold, filtered to the same part of speech with WordNet, ranked by contextual fluency using Google's one-billion-word language model, and limited to the highest-scoring candidates. Target-prompt words are used to keep substitutions relevant to the desired image rather than replacing words only with unrelated synonyms.

  5. Knowl 5 — Adversarial prompt quality fine-tuning

    model/method

    After the genetic search finds prompts that satisfy the image-similarity and text-dissimilarity requirements, RIATIG applies a separate quality-improvement stage. First, it iteratively inserts or replaces function words such as on and in at different sentence positions and retains variants judged most natural. This exploits the observation that such words have less influence on generated-image semantics than nouns and verbs.

    Second, RIATIG reduces unnecessary spelling corruption caused by the generator's finite embedding vocabulary. Deliberately misspelled out-of-vocabulary words may all map to the same unknown token, so RIATIG tests alternative spellings and retains the one requiring fewer character changes while preserving the same generator-side representation. In the white-box setting, out-of-vocabulary status is checked directly against the target model's embedding dictionary. In the black-box setting, RIATIG uses a GloVe model trained on Common Crawl with 2.2 million vocabulary entries as a surrogate embedding model. Together, these operations seek prompts that remain visually and semantically effective for the generator while appearing closer to ordinary natural language.

  6. Knowl 6 — Experimental models, data, and evaluation protocol

    experimental setup

    RIATIG is evaluated using Microsoft COCO, which contains 82,783 training images and 40,504 testing images, with five text descriptions per image. The white-box targets are AttnGAN, DM-GAN, and DF-GAN, using their COCO-trained released models. The black-box targets are DALL·E mini as a proxy for DALL·E, DALL·E 2, and Imagen. Because DALL·E 2 provides only limited API credits, adversarial prompts are first trained on a LAION implementation and then checked through the DALL·E 2 API; Imagen uses a released Hugging Face implementation.

    For each experiment, the initial population size is M=300M=300, the per-generation population size is N=50N=50, the elite population size is K=25K=25, and the initial mutation probability is pm=0.8p_m=0.8. Ten adversarial examples are generated per model, using a randomly selected target image and a semantically unrelated original prompt. Five baselines are compared: hidden vocabulary, macaronic prompting, evocative prompting, homoglyph substitution, and TextFooler.

    Attack effectiveness is measured by R-precision. Each generated image is retrieved against 100 candidate descriptions consisting of one ground-truth description and 99 randomly mismatched descriptions; ViLT is used for image-text retrieval. R-1 and R-3 count an attack as successful when the ground-truth description appears among the top 1 or top 3 retrieved descriptions. Prompt imperceptibility is measured by Universal Sentence Encoder cosine semantic distance and GPT-2 perplexity (PPL), where lower PPL indicates more fluent text.

  7. Knowl 7 — White-box attack results

    data/table

    The white-box evaluation attacks AttnGAN, DM-GAN, and DF-GAN using full model knowledge. RIATIG succeeds on all ten DM-GAN and DF-GAN examples at R-3 and on nine of ten AttnGAN examples; its prompts also have substantially lower perplexity than the non-natural baselines, which did not reliably produce targeted adversarial examples.

    Model R-1 ↑\uparrow R-3 ↑\uparrow Semantic Distance ↓\downarrow PPL ↓\downarrow
    AttnGAN 8/10 9/10 0.24 522.63
    DM-GAN 8/10 10/10 0.27 420.16
    DF-GAN 9/10 10/10 0.34 558.81

    The semantic-distance values show that the successful adversarial prompts remain weakly related to the target prompts, while the PPL values indicate substantially more fluent samples than the gibberish-oriented attack strategies.

  8. Knowl 8 — Black-box attack results against three large-scale generators

    data/table

    In the black-box setting, RIATIG is trained using only model queries and generated images. It reaches R-3 success rates of 10/10 on DALL·E mini, DALL·E 2, and Imagen, while the hidden-vocabulary baseline reaches at most 4/10 and the untargeted homoglyph and TextFooler attacks reach 0/10. RIATIG also produces much lower PPL than the macaronic and evocative prompting baselines, indicating more natural prompts.

    Model Method R-1 ↑\uparrow R-3 ↑\uparrow Sem. Dist. ↓\downarrow PPL ↓\downarrow
    DALL·E mini HiddVocab 3/10 3/10 0.17 5284.78
    DALL·E mini MacPromp 8/10 8/10 0.46 6814.61
    DALL·E mini EvoPromp 6/10 6/10 0.14 5662.97
    DALL·E mini HomoSubs 0/10 0/10 0.09 1115.84
    DALL·E mini TextFooler 0/10 0/10 0.18 1848.41
    DALL·E mini RIATIG 9/10 10/10 0.27 420.16
    DALL·E 2 HiddVocab 4/10 4/10 0.16 5162.71
    DALL·E 2 MacPromp 9/10 9/10 0.39 4931.31
    DALL·E 2 EvoPromp 7/10 7/10 0.17 6139.39
    DALL·E 2 HomoSubs 0/10 0/10 0.06 981.69
    DALL·E 2 TextFooler 0/10 0/10 0.18 2624.88
    DALL·E 2 RIATIG 10/10 10/10 0.27 872.03
    Imagen HiddVocab 2/10 2/10 0.18 5073.74
    Imagen MacPromp 6/10 6/10 0.42 5694.73
    Imagen EvoPromp 4/10 4/10 0.16 5607.66
    Imagen HomoSubs 0/10 0/10 0.08 1103.53
    Imagen TextFooler 0/10 0/10 0.16 2231.10
    Imagen RIATIG 10/10 10/10 0.23 704.09

    The baselines with lower semantic distance sometimes generate less dissimilar text, but they either fail to target the desired images or produce highly unnatural prompts. RIATIG provides the strongest combination of target-image success and prompt fluency in this comparison.

  9. Knowl 9 — Ablation, target robustness, and substitution results

    empirical result

    The ablations support all three major design choices in RIATIG. Importance-guided and semantically constrained mutation outperforms random insert/swap/delete mutation, varying the target image has limited effect on success, Word2vec substitution is generally strongest, and a larger evaluation retains high attack rates.

    Random mutation produces substantially weaker attacks:

    Model R-1 ↑\uparrow R-3 ↑\uparrow Sem. Dist. ↓\downarrow PPL ↓\downarrow
    DM-GAN 7/10 8/10 0.14 867.93
    DF-GAN 1/10 2/10 0.12 882.80
    DALL·E mini 0/10 0/10 0.17 2111.82

    Across ten experiments using different target images with the same semantic meaning, RIATIG remains robust:

    Model R-1 ↑\uparrow R-3 ↑\uparrow Sem. Dist. ↓\downarrow PPL ↓\downarrow
    DF-GAN 9.2(±\pm0.42)/10 9.8(±\pm0.42)/10 0.31±\pm0.04 596.33±\pm297.02
    DALL·E mini 9.7(±\pm0.48)/10 9.9(±\pm0.32)/10 0.28±\pm0.03 475.83±\pm180.53

    For the substitution comparison, each candidate is selected from the top 30 results. Word2vec gives the best R-precision and PPL on DF-GAN and the best values on all reported metrics for DALL·E mini:

    Model Substitution R-1 ↑\uparrow R-3 ↑\uparrow Sem. Dist. ↓\downarrow PPL ↓\downarrow
    DF-GAN Word2vec 9/10 10/10 0.336 558.81
    DF-GAN GloVe 8/10 9/10 0.331 836.96
    DF-GAN WordNet 5/10 7/10 0.358 580.49
    DALL·E mini Word2vec 9/10 10/10 0.268 420.16
    DALL·E mini GloVe 7/10 9/10 0.311 797.63
    DALL·E mini WordNet 4/10 6/10 0.368 522.34

    The authors additionally report that feeding the same adversarial prompt to DALL·E mini, DALL·E 2, and Imagen ten times achieves 100% R-3 retrieval. In 50 additional experiments with distinct sources and targets, RIATIG achieves 46/50 R-3 successes on DF-GAN and 47/50 on DALL·E mini:

    Model R-1 ↑\uparrow R-3 ↑\uparrow Sem. Dist. ↓\downarrow PPL ↓\downarrow
    DF-GAN 42/50 46/50 0.22 838.28
    DALL·E mini 45/50 47/50 0.34 640.78
  10. Knowl 10 — Defense implications and stated deployment limitations

    limitation

    The paper shows that a black-box attacker can use a sentence that is unrelated to the target prompt yet still produce a chosen visual concept, creating a risk for content-moderation systems intended to block harmful image generation. A rule-based text filter that removes redundant spaces and misspellings is insufficient by itself: when the adversarial prompts were tested with Grammarly, 20% bypassed the filter because the tool did not evaluate the sentence's semantic meaning.

    The paper discusses image filtering as another defense, but notes that constructing a comprehensive harmful-image dataset is difficult and that categories such as deepfakes can be hard to define. Adversarial training could associate original and adversarial prompts with the same images, but scaling this defense is costly because modern text-to-image systems are trained on very large corpora; the paper cites Imagen's approximately 460 million image-text pairs as an example. These observations are presented as possible defenses and deployment limitations rather than as a complete defense evaluation.

Coverage note — No substantial contributed material was omitted; illustrative prompt-image examples and background or related-work discussion were excluded because the method and quantitative findings already capture their contribution.

References

  1. 1.Attngan implementation. https : / / github . com / taoxugit/AttnGAN. Accessed: 2022-10-01.
  2. 2.Dall·e 2 laion implementation. https://github.com/LAION-AI/dalle2-laion. Accessed: 2022-10-01.
  3. 3.Dall·e mini implementation. https://github.com/borisdayma/dalle-mini. Accessed: 2022-10-01.
  4. 4.Df-gan implementation. https : / / github . com / tobran/DF-GAN. Accessed: 2022-10-01.
  5. 5.Dm-gan implementation. https : / / github . com / MinfengZhu/DM-GAN. Accessed: 2022-10-01.
  6. 6.Imagen implementation. https : / / github . com / cene555/Imagen-pytorch. Accessed: 2022-10-01.
  7. 7.Leet speak cheat sheet. https://www.gamehouse.com/blog/leet-speak-cheat-sheet/. Accessed: 2022-09-30.
  8. 8.Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. Generating natural language adversarial examples. arXiv preprint arXiv:1804.07998, 2018.
  9. 9.Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndi ˇ c, Pavel Laskov, Giorgio Giacinto, and ´ Fabio Roli. Evasion attacks against machine learning at test time. In Joint European conference on machine learning and knowledge discovery in databases, pages 387–402. Springer, 2013.
  10. 10.Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. Multimodal datasets: misogyny, pornography, and malignant stereotypes. arXiv preprint arXiv:2110.01963, 2021.
  11. 11.Clemens-Alexander Brust and Joachim Denzler. Not just a matter of semantics: the relationship between visual similarity and semantic similarity. arXiv preprint arXiv:1811.07120, 2018.
  12. 12.Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, et al. Universal sentence encoder. arXiv preprint arXiv:1803.11175, 2018.
  13. 13.Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. One billion word benchmark for measuring progress in statistical language modeling. arXiv preprint arXiv:1312.3005, 2013.
  14. 14.Bobby Chesney and Danielle Citron. Deep fakes: A looming challenge for privacy, democracy, and national security. Calif. L. Rev., 107:1753, 2019.
  15. 15.Giannis Daras and Alexandros G Dimakis. Discovering the hidden vocabulary of dalle-2. arXiv preprint arXiv:2206.00169, 2022.
  16. 16.Thomas Deselaers and Vittorio Ferrari. Visual and semantic similarity in imagenet. In CVPR 2011, pages 1777–1784. IEEE, 2011.
  17. 17.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.
  18. 18.Mary Anne Franks and Ari Ezra Waldman. Sex, lies, and videotape: Deep fakes and free speech delusions. Md. L. Rev., 78:892, 2018.
  19. 19.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014.
  20. 20.Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Rezende, and Daan Wierstra. Draw: A recurrent neural network for image generation. In International conference on machine learning, pages 1462–1471. PMLR, 2015.
  21. 21.Chuan Guo, Jacob Gardner, Yurong You, Andrew Gordon Wilson, and Kilian Weinberger. Simple black-box adversarial attacks. In International Conference on Machine Learning, pages 2484–2493. PMLR, 2019.
  22. 22.John H Holland. Genetic algorithms. Scientific american, 267(1):66–73, 1992.
  23. 23.Rowan T Hughes, Liming Zhu, and Tomasz Bednarz. Generative adversarial networks–enabled human–artificial intelligence collaborative applications for creative and design industries: A systematic review of current approaches and trends. Frontiers in artificial intelligence, 4:604234, 2021.
  24. 24.Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. Is bert really robust? a strong baseline for natural language attack on text classification and entailment. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 8018–8025, 2020.
  25. 25.Sourabh Katoch, Sumit Singh Chauhan, and Vijay Kumar. A review on genetic algorithm: past, present, and future. Multimedia Tools and Applications, 80(5):8091–8126, 2021.
  26. 26.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, pages 5583–5594. PMLR, 2021.
  27. 27.Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, and Philip Torr. Controllable text-to-image generation. Advances in Neural Information Processing Systems, 32, 2019.
  28. 28.Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. Textbugger: Generating adversarial text against real-world applications. arXiv preprint arXiv:1812.05271, 2018.
  29. 29.Bin Liang, Hongcheng Li, Miaoqiang Su, Pan Bian, Xirong Li, and Wenchang Shi. Deep text classification can be fooled. arXiv preprint arXiv:1704.08006, 2017.
  30. 30.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  31. 31.Han Liu, Zhiyuan Yu, Mingming Zha, XiaoFeng Wang, William Yeoh, Yevgeniy Vorobeychik, and Ning Zhang. When evil calls: Targeted adversarial voice over ip network. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security, pages 2009–2023, 2022.
  32. 32.Giulio Lovisotto, Henry Turner, Ivo Sluganovic, Martin Strohmeier, and Ivan Martinovic. {SLAP}: Improving physical adversarial examples with {Short-Lived} adversarial perturbations. In 30th USENIX Security Symposium (USENIX Security 21), pages 1865–1882, 2021.
  33. 33.Elman Mansimov, Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov. Generating images from captions with attention. arXiv preprint arXiv:1511.02793, 2015.
  34. 34.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013.
  35. 35.George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39–41, 1995.
  36. 36.Raphaël Millière. Adversarial attacks on image generation with made-up words. arXiv preprint arXiv:2208.04135, 2022.
  37. 37.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  38. 38.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532–1543, 2014.
  39. 39.Tingting Qiao, Jing Zhang, Duanqing Xu, and Dacheng Tao. Mirrorgan: Learning text-to-image generation by redescription. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1505–1514, 2019.
  40. 40.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  41. 41.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  42. 42.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  43. 43.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021.
  44. 44.Graham Rawlinson. The significance of letter position in word recognition. IEEE Aerospace and Electronic Systems Magazine, 22(1):26–27, 2007.
  45. 45.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  46. 46.Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus. Learning visual representations with caption annotations. In European Conference on Computer Vision, pages 153–170. Springer, 2020.
  47. 47.Ramya Srinivasan and Kanji Uchino. Biases in generative art: A causal look from the lens of art history. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 41–51, 2021.
  48. 48.Lukas Struppek, Dominik Hintersdorf, and Kristian Kersting. The biased artist: Exploiting cultural biases via homoglyphs in text-guided image generation models. arXiv preprint arXiv:2209.08891, 2022.
  49. 49.Ming Tao, Hao Tang, Fei Wu, Xiao-Yuan Jing, Bing-Kun Bao, and Changsheng Xu. Df-gan: A simple and effective baseline for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16515–16525, 2022.
  50. 50.Rohan Taori, Amog Kamsetty, Brenton Chu, and Nikita Vemuri. Targeted adversarial examples for black box audio systems. In 2019 IEEE security and privacy workshops (SPW), pages 15–20. IEEE, 2019.
  51. 51.Umut Topkara, Mercan Topkara, and Mikhail J Atallah. The hiding virtues of ambiguity: quantifiably resilient watermarking of natural language text through synonym substitutions. In Proceedings of the 8th workshop on Multimedia and security, pages 164–174, 2006.
  52. 52.Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1316–1324, 2018.
  53. 53.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2022.
  54. 54.Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang. Cross-modal contrastive learning for text-to-image generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 833–842, 2021.
  55. 55.Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In Proceedings of the IEEE international conference on computer vision, pages 5907–5915, 2017.
  56. 56.Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. Stackgan++: Realistic image synthesis with stacked generative adversarial networks. IEEE transactions on pattern analysis and machine intelligence, 41(8):1947–1962, 2018.
  57. 57.Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz. Contrastive learning of medical visual representations from paired images and text. arXiv preprint arXiv:2010.00747, 2020.
  58. 58.Minfeng Zhu, Pingbo Pan, Wei Chen, and Yi Yang. Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5802–5810, 2019.

Citation

MLA
Liu, H., et al. “RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 20585–94, https://doi.org/10.1109/CVPR52729.2023.01972.
APA
Liu, H., Wu, Y., Zhai, S., Yuan, B., & Zhang, N. (2023). RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 20585–20594. https://doi.org/10.1109/CVPR52729.2023.01972
Chicago
Liu, H., Y. Wu, S. Zhai, B. Yuan, and N. Zhang. 2023. “RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 20585–94. https://doi.org/10.1109/CVPR52729.2023.01972.
Harvard
Liu, H. et al. (2023) “RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 20585–20594. Available at: https://doi.org/10.1109/CVPR52729.2023.01972.
Vancouver
1. Liu H, Wu Y, Zhai S, Yuan B, Zhang N (2023) RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 20585–20594

BibTeX

@inproceedings{Liu_2023, title={RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01972}, DOI={10.1109/cvpr52729.2023.01972}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Liu, Han and Wu, Yuhao and Zhai, Shixuan and Yuan, Bo and Zhang, Ning}, year={2023}, month=June, pages={20585–20594} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE