TextDiffuser: Diffusion Models as Text Painters

Jingye ChenYupan HuangTengchao LvLei CuiQifeng ChenFuru Wei

article2023NeurIPS223 citations

Introduces a two-stage diffusion framework guided by character-level layout segmentation alongside the 10-million-image MARIO-10M dataset to achieve controllable, accurate visual text rendering and text inpainting.

Listen

Modern artificial intelligence image generation tools frequently struggle to render accurate, legible, and visually coherent text within synthetic images. While conventional image editing often yields unnatural visual artifacts and earlier diffusion models produce garbled characters or lack spatial control, high-quality visual text is essential for practical commercial assets like book covers, posters, and advertisements.

The article demonstrates and evaluates TextDiffuser, a flexible two-stage framework designed to generate coherent, high-quality images containing visual text from text prompts or template images, as well as perform targeted text inpainting. It also introduces MARIO-10M, a curated dataset of 10 million image-text pairs with optical character recognition annotations, alongside the MARIO-Eval benchmark to systematically assess visual text rendering quality.

The approach divides the generation process into layout planning and image synthesis. First, a Transformer model extracts keywords from a text prompt and determines their coordinate bounding boxes, converting them into character-level segmentation masks. Second, a latent diffusion model conditions image generation on these masks alongside the text prompts, utilizing an auxiliary character-aware loss to ensure sharp text definition. The authors trained the model on 10 million filtered image-text pairs from web scrapes, movie databases, and digital libraries, evaluating performance across automated metrics and human user studies.

The key findings show that explicit layout and character-level guidance significantly outperforms existing generative models. TextDiffuser achieved an optical character recognition accuracy of 56.1% and an F-measure of 78.2%, outperforming standard Stable Diffusion by approximately 76 percentage points in F-measure and DeepFloyd by roughly 61 percentage points. In human evaluations, TextDiffuser earned substantially more top votes for both visual rendering quality and prompt matching compared to models like DALL-E, ControlNet, and Midjourney. Furthermore, the architecture achieved these gains efficiently, adding only about 3% more parameters to its baseline model while retaining the ability to synthesize general natural scenes without visual degradation.

These findings indicate that integrating explicit structural and character-level constraints resolves the primary text legibility issues present in standard diffusion pipelines. This improves visual fidelity and design utility while mitigating rendering errors, making automated graphic generation viable for commercial production without requiring extensive manual redesign or complex prompt engineering.

Organizations developing or deploying image generation pipelines should incorporate explicit layout and character segmentation conditioning to reliably control visual typography. When implementing these systems, teams should consider the trade-off between resolution and inference speed. Developers should also implement tamper-detection measures to prevent potential misuse in document manipulation, and future technical iterations should focus on extending character modeling to multilingual scripts.

The primary limitations involve the compression mechanics of latent diffusion models. Compressing images into compact latent representations causes a loss of fine detail, resulting in blurry or broken strokes for very small characters, while prompts with long, dense text sequences can occasionally produce overlapping layouts. Nevertheless, the extensive quantitative and human evaluations provide high confidence that TextDiffuser establishes a robust, highly controllable foundation for automated visual text rendering.

Cover for TextDiffuser: Diffusion Models as Text Painters

Abstract

Diffusion models have gained increasing attention for their impressive generation abilities but currently struggle with rendering accurate and coherent text. To address this issue, we introduce TextDiffuser, focusing on generating images with visually appealing text that is coherent with backgrounds. TextDiffuser consists of two stages: first, a Transformer model generates the layout of keywords extracted from text prompts, and then diffusion models generate images conditioned on the text prompt and the generated layout. Additionally, we contribute the first large-scale text images dataset with OCR annotations, MARIO-10M, containing 10 million image-text pairs with text recognition, detection, and character-level segmentation annotations. We further collect the MARIO-Eval benchmark to serve as a comprehensive tool for evaluating text rendering quality. Through experiments and user studies, we show that TextDiffuser is flexible and controllable to create high-quality text images using text prompts alone or together with text template images, and conduct text inpainting to reconstruct incomplete images with text. The code, model, and dataset will be available at https://aka.ms/textdiffuser.

Citation

MLA
Chen, J., et al. “TextDiffuser: Diffusion Models as Text Painters”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 9353–87, https://proceedings.neurips.cc/paper_files/paper/2023/file/1df4afb0b4ebf492a41218ce16b6d8df-Paper-Conference.pdf.
APA
Chen, J., Huang, Y., Lv, T., Cui, L., Chen, Q., & Wei, F. (2023). TextDiffuser: Diffusion Models as Text Painters. Advances in Neural Information Processing Systems, 36, 9353–9387. https://proceedings.neurips.cc/paper_files/paper/2023/file/1df4afb0b4ebf492a41218ce16b6d8df-Paper-Conference.pdf
Chicago
Chen, J., Y. Huang, T. Lv, L. Cui, Q. Chen, and F. Wei. 2023. “TextDiffuser: Diffusion Models as Text Painters”. Advances in Neural Information Processing Systems 36: 9353–87. https://proceedings.neurips.cc/paper_files/paper/2023/file/1df4afb0b4ebf492a41218ce16b6d8df-Paper-Conference.pdf.
Harvard
Chen, J. et al. (2023) “TextDiffuser: Diffusion Models as Text Painters”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 9353–9387. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/1df4afb0b4ebf492a41218ce16b6d8df-Paper-Conference.pdf.
Vancouver
1. Chen J, Huang Y, Lv T, Cui L, Chen Q, Wei F (2023) TextDiffuser: Diffusion Models as Text Painters. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 9353–9387

BibTeX

@inproceedings{chen2023textdiffuser,
  title = {TextDiffuser: Diffusion Models as Text Painters},
  author = {Chen, Jingye and Huang, Yupan and Lv, Tengchao and Cui, Lei and Chen, Qifeng and Wei, Furu},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {9353-9387},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/1df4afb0b4ebf492a41218ce16b6d8df-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission