TextDiffuser: Diffusion Models as Text Painters
Jingye ChenYupan HuangTengchao LvLei CuiQifeng ChenFuru Wei
Introduces a two-stage diffusion framework guided by character-level layout segmentation alongside the 10-million-image MARIO-10M dataset to achieve controllable, accurate visual text rendering and text inpainting.
Modern artificial intelligence image generation tools frequently struggle to render accurate, legible, and visually coherent text within synthetic images. While conventional image editing often yields unnatural visual artifacts and earlier diffusion models produce garbled characters or lack spatial control, high-quality visual text is essential for practical commercial assets like book covers, posters, and advertisements.
The article demonstrates and evaluates TextDiffuser, a flexible two-stage framework designed to generate coherent, high-quality images containing visual text from text prompts or template images, as well as perform targeted text inpainting. It also introduces MARIO-10M, a curated dataset of 10 million image-text pairs with optical character recognition annotations, alongside the MARIO-Eval benchmark to systematically assess visual text rendering quality.
The approach divides the generation process into layout planning and image synthesis. First, a Transformer model extracts keywords from a text prompt and determines their coordinate bounding boxes, converting them into character-level segmentation masks. Second, a latent diffusion model conditions image generation on these masks alongside the text prompts, utilizing an auxiliary character-aware loss to ensure sharp text definition. The authors trained the model on 10 million filtered image-text pairs from web scrapes, movie databases, and digital libraries, evaluating performance across automated metrics and human user studies.
The key findings show that explicit layout and character-level guidance significantly outperforms existing generative models. TextDiffuser achieved an optical character recognition accuracy of 56.1% and an F-measure of 78.2%, outperforming standard Stable Diffusion by approximately 76 percentage points in F-measure and DeepFloyd by roughly 61 percentage points. In human evaluations, TextDiffuser earned substantially more top votes for both visual rendering quality and prompt matching compared to models like DALL-E, ControlNet, and Midjourney. Furthermore, the architecture achieved these gains efficiently, adding only about 3% more parameters to its baseline model while retaining the ability to synthesize general natural scenes without visual degradation.
These findings indicate that integrating explicit structural and character-level constraints resolves the primary text legibility issues present in standard diffusion pipelines. This improves visual fidelity and design utility while mitigating rendering errors, making automated graphic generation viable for commercial production without requiring extensive manual redesign or complex prompt engineering.
Organizations developing or deploying image generation pipelines should incorporate explicit layout and character segmentation conditioning to reliably control visual typography. When implementing these systems, teams should consider the trade-off between resolution and inference speed. Developers should also implement tamper-detection measures to prevent potential misuse in document manipulation, and future technical iterations should focus on extending character modeling to multilingual scripts.
The primary limitations involve the compression mechanics of latent diffusion models. Compressing images into compact latent representations causes a loss of fine detail, resulting in blurry or broken strokes for very small characters, while prompts with long, dense text sequences can occasionally produce overlapping layouts. Nevertheless, the extensive quantitative and human evaluations provide high confidence that TextDiffuser establishes a robust, highly controllable foundation for automated visual text rendering.
- Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). This foundational paper establishes Latent Diffusion Models and cross-attention conditioning, providing the core generative image architecture that TextDiffuser adapts for layout-guided text rendering.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). This work introduces spatial conditional control for diffusion models, establishing the framework for injecting explicit spatial and layout guidance into text-to-image generation.
- Paper: T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models, Chong Mou et al. (2023). This paper presents lightweight adapters for structural conditioning in diffusion models, offering foundational concepts for guiding diffusion pipelines with structured spatial signals.
- Paper: GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models, Alexander Quinn Nichol et al. (2022). This study details text-guided diffusion models and classifier-free guidance, providing key background on conditional image generation and text-driven inpainting.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This seminal paper introduces modern denoising diffusion probabilistic models, establishing the core mathematical formulation utilized throughout subsequent conditional diffusion systems.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). This work scales multimodal rectified flow transformers to dramatically improve text spelling, legibility, and compositional prompt adherence in text-to-image synthesis.
- Paper: SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis, Dustin Podell et al. (2024). This paper advances latent diffusion architectures with multi-stage conditioning and larger backbones to achieve higher fidelity and better text comprehension in high-resolution image synthesis.
- Paper: IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models, Hu Ye et al. (2023). This research introduces decoupled cross-attention adapters to flexibly inject multimodal prompt conditions into diffusion models without retraining the base architecture.
- Paper: InstructPix2Pix: Learning to Follow Image Editing Instructions, Tim Brooks et al. (2023). This paper extends conditional image synthesis by training diffusion models to execute complex visual edits directly from natural language instructions.
