Character-Aware Models Improve Visual Text Rendering
Rosanne LiuDan GarretteChitwan SahariaWilliam ChanAdam RobertsSharan NarangIrina BlokRJ MicalMohammad NorouziNoah Constant
Demonstrates that integrating character-aware text encoders into text-to-image architectures significantly improves visual spelling accuracy, outperforming conventional subword-based models on complex and rare text rendering tasks.
Modern text-to-image artificial intelligence systems frequently fail to render legible, correctly spelled visual text in generated pictures. This issue limits their practical utility for graphic design, marketing materials, and digital art. The primary cause of this failure is that standard image generators rely on subword-level text encoders that do not receive direct character-by-character information, making it difficult for the models to accurately translate words into visual letter sequences.
The main objective of the article is to evaluate why popular image generation models fail at visual text rendering and to demonstrate that incorporating character-aware text encoders significantly improves spelling and text generation accuracy.
The authors conducted two phases of evaluation. First, they assessed the spelling capabilities of isolated text encoders across multiple model scales and seven languages using a new benchmark called WikiSpell. Second, they trained several image generation models on a standardized dataset of 400 million image-caption pairs using both standard character-blind text encoders and character-aware text encoders. They evaluated visual rendering performance using DrawText, a new benchmark comprising automated optical character recognition testing across 500 word prompts and human evaluation across 175 diverse creative prompts.
The experiments produced four critical findings. First, character-aware text encoders consistently achieved high spelling accuracy in text-only tasks across all word frequencies and languages, whereas standard character-blind models required over 100 billion parameters to achieve comparable spelling reliability. Second, image generation models using character-aware text encoders showed dramatic improvements in visual spelling, yielding accuracy gains of over 15 percentage points on common words and more than 30 percentage points on rare words compared to character-blind baselines, despite requiring fewer training steps. Third, character-blind models frequently made systematic errors such as regularizing irregular verbs into incorrect past-tense spellings, while character-aware models avoided these systematic errors entirely. Fourth, while purely character-level models experienced a minor drop in overall prompt alignment for non-text image elements, a hybrid approach combining standard token-based encodings with a lightweight character-aware encoder delivered state-of-the-art text rendering without degrading overall image quality or prompt fidelity.
These findings indicate that image generation pipelines can achieve superior visual text rendering without relying on massive, computationally expensive language models. Integrating lightweight character-aware representations provides a cost-effective and scalable pathway to eliminate common visual spelling failures, thereby improving model utility for automated design workflows.
Organizations developing or deploying image generation models should adopt hybrid text encodings that merge standard subword representations with character-aware inputs. For future technical development, engineering teams should focus on improving the downstream image generation modules to resolve remaining visual artifacts, such as letter merging or glyph placement, and expand training pipelines to support visual text rendering in non-Latin writing systems.
The findings are supported by consistent results across automated optical character recognition and human preference evaluations. However, readers should note that the visual image experiments were limited to English-language Latin text, and future verification is needed to determine how these text-encoding methods perform when encoders are trained jointly with image generation systems rather than kept frozen.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). Its large-scale text-to-image framework establishes the generation setting and text-conditioning pipeline whose spelling limitations this paper investigates.
- Paper: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding, Chitwan Saharia et al. (2022). Imagen’s use of a pretrained language encoder in diffusion models provides direct context for understanding how text representations affect image generation.
- Paper: TextDiffuser: Diffusion Models as Text Painters, Jingye Chen et al. (2023). TextDiffuser carries character-level guidance into explicit layout planning and text-focused image synthesis, extending the source’s encoder-based spelling improvements into a controllable rendering pipeline.
