Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis

Lukas StruppekDominik HintersdorfFelix FriedrichManuel BrackPatrick SchramowskiKristian Kersting

article2023JAIR56 citationsBest Paper Award at DPFM Workshop (ICLR)

Reveals how inserting visually identical non-Latin homoglyphs into prompts covertly triggers cultural stereotypes in text-to-image synthesis, and proposes an unlearning technique to defend text encoders against these script-based attacks.

Listen

Modern text-to-image synthesis models, such as Stable Diffusion and DALL-E 2, are widely deployed across industries ranging from digital media to software interfaces. Because these systems are trained on massive datasets scraped from the public internet, they internalize subtle statistical correlations and societal associations. While standard text prompts predominantly produce outputs skewed toward Western cultural norms, these models also exhibit unexpected sensitivities at the individual character level. Understanding how minor textual variations alter generated outputs is essential as organizations increasingly integrate generative artificial intelligence into customer-facing products, automated workflows, and creative pipelines.

The article demonstrates that injecting single non-Latin characters—especially homoglyphs, which are characters from non-Latin Unicode scripts that visually resemble standard Latin letters—explicitly steers image generation toward the cultural and ethnic stereotypes associated with those scripts. The authors evaluate this vulnerability across multiple systems and domains, identify the underlying architectural cause, and propose a lightweight mitigation method termed homoglyph unlearning.

To assess this behavior, the authors evaluated leading text-to-image models, primarily Stable Diffusion v1.5 and DALL-E 2, alongside multilingual systems such as AltDiffusion. They designed structured prompt benchmarks across three main domains—People, Buildings, and Miscellaneous cultural concepts—and quantified bias using Relative Bias (measuring increased CLIP text-image similarity with culture-specific terms), Visual Question Answering scores using BLIP-2, and the Word Embedding Association Test. Furthermore, they trained a fine-tuned student text encoder using a loss function designed to map homoglyph-containing prompts back to their standard Latin text representations without modifying the broader image decoder network.

The findings establish that replacing as little as a single Latin character with a visually indistinguishable homoglyph drastically shifts generated outputs. For example, replacing a Latin letter in a prompt with a Korean or African character causes human depictions to reflect those specific ethnic appearances, while Greek homoglyphs shift architectural scenes toward Ancient Greek styles. Statistical embedding tests confirmed that the pre-trained text encoder is the primary driver of this bias, clustering characters by script and projecting cultural vectors into prompt embeddings. Multilingual text encoders, however, exhibited substantially lower character-induced bias. Finally, applying the homoglyph unlearning technique removed nearly all character-induced bias while preserving general model utility, resulting in only a minor 1.16 percentage point change in zero-shot classification accuracy and negligible impact on image quality.

These results present both opportunities and security risks for organizations using generative models. On one hand, non-Latin script sensitivity provides an intuitive, subtle way for diverse users to counteract default Western biases and generate culturally tailored imagery. On the other hand, it creates a subtle attack vector where third-party prompt databases, plugins, or malicious actors can surreptitiously insert homoglyphs into benign prompts to reinforce racial stereotypes, skew facial appearances in sensitive contexts like crime-related imagery, or manipulate image retrieval rankings without visual detection. Complete retraining of large multimodal systems to fix this vulnerability is prohibitively expensive, but targeted encoder fine-tuning provides an efficient alternative.

Organizations deploying text-to-image systems should evaluate input validation and model hardening strategies based on their specific operational context. Implementing basic Unicode script filtering at the application programming interface layer offers immediate protection, though it risks excluding users who intentionally rely on non-Latin scripts. A superior long-term approach is adopting multilingual encoders or implementing homoglyph unlearning during model fine-tuning, which neutralizes the vulnerability while preserving input flexibility. Further testing should be conducted across complex commercial prompts and emerging multimodal chat systems before rolling out automated generative pipelines.

The study's primary limitations include focusing largely on relatively short prompts, as the biasing impact of individual homoglyphs tends to diminish in longer, highly detailed text descriptions unless multiple homoglyphs are injected. Additionally, non-deterministic model interfaces like DALL-E 2 introduce variance into quantitative measurements compared to open-source seeded models. Nevertheless, the experimental evidence clearly confirms that character-level script manipulation consistently triggers cultural associations across mainstream multimodal architectures.

Cover for Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis

Abstract

Models for text-to-image synthesis, such as DALL-E~2 and Stable Diffusion, have recently drawn a lot of interest from academia and the general public. These models are capable of producing high-quality images that depict a variety of concepts and styles when conditioned on textual descriptions. However, these models adopt cultural characteristics associated with specific Unicode scripts from their vast amount of training data, which may not be immediately apparent. We show that by simply inserting single non-Latin characters in a textual description, common models reflect cultural stereotypes and biases in their generated images. We analyze this behavior both qualitatively and quantitatively, and identify a model's text encoder as the root cause of the phenomenon. Additionally, malicious users or service providers may try to intentionally bias the image generation to create racist stereotypes by replacing Latin characters with similarly-looking characters from non-Latin scripts, so-called homoglyphs. To mitigate such unnoticed script attacks, we propose a novel homoglyph unlearning method to fine-tune a text encoder, making it robust against homoglyph manipulations.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 2.1 Text-To-Image Synthesis
  • 2.2 Biases and Fairness in Image Generation Models
  • 2.3 Homoglyphs and Related Attacks in the Context of Machine Learning
  • 3 Methodology for Investigating Character Manipulation
  • 3.1 Experimental Setting
  • 3.2 Quantifying the Influence of Homoglyphs and non-Latin Characters
  • 3.3 Fine-Tuning Text Encoders with Homoglyph Unlearning
  • 4 Manipulating the Image Generation with Homoglyphs
  • 4.1 Inducing Cultural Biases into the Image Generation Process
  • 4.2 Text Encoders Are the Driving Force behind Homoglyph-Induced Biases
  • 5 Discussion, Challenges, and Conclusion
  • 5.1 Social Impact and Ethical Considerations
  • 5.2 Challenges and Future Research
  • 5.3 Conclusion
  • References
  • A Unicode Scripts
  • B Experimental Details
  • B.1 Relative Bias Dataset Prompts
  • B.2 VQA Score
  • B.3 WEAT Test
  • C Additional Experiments
  • C.1 Relative Bias
  • C.2 VQA Score
  • D Additional DALL-E 2 Results
  • D.1 A City in Bright Sunshine
  • D.2 A Photo of an Actress
  • D.3 Delicious Food on a Table
  • D.4 The Leader of a Country
  • D.5 A Photo of a Flag
  • D.6 A Photo of a Person
  • E Additional Stable Diffusion Results
  • E.1 A Photo of an Actress
  • E.2 Delicious Food on a Table
  • E.3 Homoglyph Unlearning Results
  • E.4 Inducing Biases in the Embedding Space
  • E.5 Varying the Number of Injected Homoglyphs for Complex Prompts
  • E.6 MS-COCO Examples

Citation

MLA
Struppek, L., et al. “Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis”. Journal of Artificial Intelligence Research, vol. 78, 2023, pp. 1017–68, https://doi.org/10.1613/JAIR.1.15388.
APA
Struppek, L., Hintersdorf, D., Friedrich, F., Br, M., Schramowski, P., & Kersting, K. (2023). Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis. Journal of Artificial Intelligence Research, 78, 1017–1068. https://doi.org/10.1613/JAIR.1.15388
Chicago
Struppek, L., D. Hintersdorf, F. Friedrich, M. Br, P. Schramowski, and K. Kersting. 2023. “Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis”. Journal of Artificial Intelligence Research 78: 1017–68. https://doi.org/10.1613/JAIR.1.15388.
Harvard
Struppek, L. et al. (2023) “Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis”, Journal of Artificial Intelligence Research, 78, pp. 1017–1068. Available at: https://doi.org/10.1613/JAIR.1.15388.
Vancouver
1. Struppek L, Hintersdorf D, Friedrich F, Br M, Schramowski P, Kersting K (2023) Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis. Journal of Artificial Intelligence Research 78:1017–1068

BibTeX

@article{Struppek_2023, title={Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis}, volume={78}, ISSN={1076-9757}, url={http://dx.doi.org/10.1613/JAIR.1.15388}, DOI={10.1613/jair.1.15388}, journal={Journal of Artificial Intelligence Research}, publisher={AI Access Foundation}, author={Struppek, Lukas and Hintersdorf, Dom and Friedrich, Felix and Br, Manuel and Schramowski, Patrick and Kersting, Kristian}, year={2023}, month=Dec, pages={1017–1068} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/