Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
Lukas StruppekDominik HintersdorfFelix FriedrichManuel BrackPatrick SchramowskiKristian Kersting
Reveals how inserting visually identical non-Latin homoglyphs into prompts covertly triggers cultural stereotypes in text-to-image synthesis, and proposes an unlearning technique to defend text encoders against these script-based attacks.
Modern text-to-image synthesis models, such as Stable Diffusion and DALL-E 2, are widely deployed across industries ranging from digital media to software interfaces. Because these systems are trained on massive datasets scraped from the public internet, they internalize subtle statistical correlations and societal associations. While standard text prompts predominantly produce outputs skewed toward Western cultural norms, these models also exhibit unexpected sensitivities at the individual character level. Understanding how minor textual variations alter generated outputs is essential as organizations increasingly integrate generative artificial intelligence into customer-facing products, automated workflows, and creative pipelines.
The article demonstrates that injecting single non-Latin characters—especially homoglyphs, which are characters from non-Latin Unicode scripts that visually resemble standard Latin letters—explicitly steers image generation toward the cultural and ethnic stereotypes associated with those scripts. The authors evaluate this vulnerability across multiple systems and domains, identify the underlying architectural cause, and propose a lightweight mitigation method termed homoglyph unlearning.
To assess this behavior, the authors evaluated leading text-to-image models, primarily Stable Diffusion v1.5 and DALL-E 2, alongside multilingual systems such as AltDiffusion. They designed structured prompt benchmarks across three main domains—People, Buildings, and Miscellaneous cultural concepts—and quantified bias using Relative Bias (measuring increased CLIP text-image similarity with culture-specific terms), Visual Question Answering scores using BLIP-2, and the Word Embedding Association Test. Furthermore, they trained a fine-tuned student text encoder using a loss function designed to map homoglyph-containing prompts back to their standard Latin text representations without modifying the broader image decoder network.
The findings establish that replacing as little as a single Latin character with a visually indistinguishable homoglyph drastically shifts generated outputs. For example, replacing a Latin letter in a prompt with a Korean or African character causes human depictions to reflect those specific ethnic appearances, while Greek homoglyphs shift architectural scenes toward Ancient Greek styles. Statistical embedding tests confirmed that the pre-trained text encoder is the primary driver of this bias, clustering characters by script and projecting cultural vectors into prompt embeddings. Multilingual text encoders, however, exhibited substantially lower character-induced bias. Finally, applying the homoglyph unlearning technique removed nearly all character-induced bias while preserving general model utility, resulting in only a minor 1.16 percentage point change in zero-shot classification accuracy and negligible impact on image quality.
These results present both opportunities and security risks for organizations using generative models. On one hand, non-Latin script sensitivity provides an intuitive, subtle way for diverse users to counteract default Western biases and generate culturally tailored imagery. On the other hand, it creates a subtle attack vector where third-party prompt databases, plugins, or malicious actors can surreptitiously insert homoglyphs into benign prompts to reinforce racial stereotypes, skew facial appearances in sensitive contexts like crime-related imagery, or manipulate image retrieval rankings without visual detection. Complete retraining of large multimodal systems to fix this vulnerability is prohibitively expensive, but targeted encoder fine-tuning provides an efficient alternative.
Organizations deploying text-to-image systems should evaluate input validation and model hardening strategies based on their specific operational context. Implementing basic Unicode script filtering at the application programming interface layer offers immediate protection, though it risks excluding users who intentionally rely on non-Latin scripts. A superior long-term approach is adopting multilingual encoders or implementing homoglyph unlearning during model fine-tuning, which neutralizes the vulnerability while preserving input flexibility. Further testing should be conducted across complex commercial prompts and emerging multimodal chat systems before rolling out automated generative pipelines.
The study's primary limitations include focusing largely on relatively short prompts, as the biasing impact of individual homoglyphs tends to diminish in longer, highly detailed text descriptions unless multiple homoglyphs are injected. Additionally, non-deterministic model interfaces like DALL-E 2 introduce variance into quantitative measurements compared to open-source seeded models. Nevertheless, the experimental evidence clearly confirms that character-level script manipulation consistently triggers cultural associations across mainstream multimodal architectures.
- Paper: Hierarchical Text-Conditional Image Generation with CLIP Latents, Aditya Ramesh et al. (2022). Its unCLIP account explains the text–image embedding and diffusion-decoder pipeline behind DALL-E 2, one of the systems whose character-level sensitivities the source investigates.
- Paper: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding, Chitwan Saharia et al. (2022). Imagen’s use of a pretrained text encoder to condition diffusion generation provides useful architectural context for the source’s finding that encoder representations drive homoglyph bias.
- Paper: OpenBias: Open-Set Bias Detection in Text-to-Image Generative Models, Moreno D'Incà et al. (2024). OpenBias broadens the source’s focused homoglyph-bias tests into open-set discovery and measurement of previously unanticipated biases in text-to-image systems.
- Paper: Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI, Yanwei Huang et al. (2025). Vipera carries the source’s concern with finding hidden text-to-image vulnerabilities into an interactive auditing workflow that helps users systematically discover and document them.
