Built independently by an author, for readers. Read the story and support ChapterPal

keyword

text rendering

Text rendering refers to the computational process of generating, displaying, or drawing legible written characters, words, and typography as visual elements within an image or graphical interface. In artificial intelligence and computer vision, the term specifically describes the capability of generative models, such as text-to-image synthesis architectures, to produce accurate, correctly spelled, and coherent text embedded directly within synthesized scenes, objects, or backgrounds. Effective text rendering requires integrating linguistic and character-level understanding with spatial layout and fine-grained visual synthesis, ensuring that individual glyphs are well formed, legible, and visually harmonious with the surrounding perspective, lighting, and style.

3 items

TextDiffuser: Diffusion Models as Text Painters

TextDiffuser: Diffusion Models as Text Painters

Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui, Qifeng Chen, Furu Wei

OrganizationsMicrosoftSun Yat-sen UniversityThe Hong Kong University of Science and Technology

Why you should read this

Introduces a two-stage diffusion framework guided by character-level layout segmentation alongside the 10-million-image MARIO-10M dataset to achieve controllable, accurate visual text rendering and text inpainting.

Diffusion models have gained increasing attention for their impressive generation abilities but currently struggle with rendering accurate and coherent text. To address this issue, we introduce TextDiffuser, focusing on generating images with visually appealing text that is coherent with backgrounds. TextDiffuser consists of two stages: first, a Transformer model generates the layout of keywords extracted from text prompts, and then diffusion models generate images conditioned on the text prompt and the generated layout. Additionally, we contribute the first large-scale text images dataset with OCR annotations, MARIO-10M, containing 10 million image-text pairs with text recognition, detection, and character-level segmentation annotations. We further collect the MARIO-Eval benchmark to serve as a comprehensive tool for evaluating text rendering quality. Through experiments and user studies, we show that TextDiffuser is flexible and controllable to create high-quality text images using text prompts alone or together with text template images, and conduct text inpainting to reconstruct incomplete images with text. The code, model, and dataset will be available at https://aka.ms/textdiffuser.

Added

2026-09-26

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, Yonghui Wu

Why you should read this

Demonstrates that scaling an autoregressive sequence-to-sequence model up to 20 billion parameters achieves state-of-the-art text-to-image synthesis quality and effectively handles complex, compositionally rich prompts.

We present the Pathways Autoregressive Text-to-Image (Parti) model, which generates high-fidelity photorealistic images and supports content-rich synthesis involving complex compositions and world knowledge. Parti treats text-to-image generation as a sequence-to-sequence modeling problem, akin to machine translation, with sequences of image tokens as the target outputs rather than text tokens in another language. This strategy can naturally tap into the rich body of prior work on large language models, which have seen continued advances in capabilities and performance through scaling data and model sizes. Our approach is simple: First, Parti uses a Transformer-based image tokenizer, ViT-VQGAN, to encode images as sequences of discrete tokens. Second, we achieve consistent quality improvements by scaling the encoder-decoder Transformer model up to 20B parameters, with a new state-of-the-art zero-shot FID score of 7.23 and finetuned FID score of 3.22 on MS-COCO. Our detailed analysis on Localized Narratives as well as PartiPrompts (P2), a new holistic benchmark of over 1600 English prompts, demonstrate the effectiveness of Parti across a wide variety of categories and difficulty aspects. We also explore and highlight limitations of our models in order to define and exemplify key areas of focus for further improvements. See this https URL for high-resolution images.

Added

2026-09-24