Conditional Text Image Generation with Diffusion Models
Yuanzhi ZhuZhaohai LiTianwei WangMengchao HeCong Yao
Proposes a conditional diffusion model that controls text, style, and visual attributes across four generation modes to synthesize realistic scene and handwritten text images that improve downstream recognition accuracy and handle out-of-vocabulary words.
Modern text recognition systems for scene text and handwriting rely heavily on massive, diverse datasets to achieve high accuracy. However, collecting and manually annotating real-world text images is prohibitively expensive and time-consuming. While traditional generative adversarial networks (GANs) and standard data augmentation tools have been used to synthesize artificial training data, they frequently suffer from training instability, mode collapse, and an inability to reliably create realistic out-of-vocabulary words or adapt to new visual domains.
The article introduces and evaluates Conditional Text Image Generation with Diffusion Models (CTIG-DM), a framework designed to generate realistic, diverse, and controllable text images. The main objective is to demonstrate how diffusion-based generative models—guided by text-specific conditions—can produce high-fidelity synthetic text images that directly improve the performance of downstream text recognizers.
The authors developed an architecture comprising a conditional encoder and a denoising diffusion model. The encoder extracts three specialized signals from pre-trained text recognition networks: image conditions (visual appearance and texture), text conditions (semantic character sequences), and style conditions (author-specific handwriting traits). By combining these inputs, the model operates across four operational modes: synthesis, augmentation, recovery, and imitation. The approach was systematically evaluated on standard handwritten benchmarks (IAM, RIMES, CVL, and CASIA-HWDB) and multiple scene text benchmarks (such as ICDAR, SVT, and IIIT 5K), measuring image quality, error rates, out-of-vocabulary generation, and cross-dataset domain adaptation.
The experimental findings show that CTIG-DM outperforms existing generative approaches across all core evaluation criteria. First, generating text images with all three conditions yielded superior visual fidelity, lowering the Fréchet Inception Distance (FID) to 9.34 on the IAM handwritten dataset compared to 33.42 for an unconditional baseline. Second, augmenting real datasets with CTIG-DM synthetic images substantially reduced recognizer error rates: on the IAM dataset, word error rate (WER) fell by 4.78 percentage points and character error rate (CER) fell by 1.87 percentage points, with comparable improvements of 5.20 (WER) and 1.52 (CER) percentage points on the RIMES French dataset. Third, in real-world scene text recognition, adding 2 million synthetic samples reduced error rates and boosted recognizer accuracy by 4.9% for CRNN and 1.5% for ABINet. Fourth, the model demonstrated strong out-of-vocabulary synthesis, producing rare characters, unseen English words, and non-existing Chinese character compositions at an FID of 25.52—roughly four times better than existing GAN baselines. Finally, in domain adaptation tasks, pre-training with synthetic images reduced CVL benchmark word error rate by 16.40 percentage points over baseline.
These results indicate that diffusion-based synthetic data generation offers a reliable, scalable alternative to manual data collection. By providing realistic variations in writing styles, fonts, backgrounds, and orientations, the approach reduces the cost and operational risk of deploying automated text recognition in data-scarce environments. Because the framework is complementary to existing geometric augmentations and text recognition architectures, engineering teams can integrate it directly into existing training pipelines without altering underlying downstream recognizers.
Based on these findings, decision-makers should consider adopting diffusion-driven synthesis pipelines when developing recognition models for rare vocabularies, specialized scripts, or new deployment domains. To further extend this capability, technical teams should pilot the method on additional languages, test generation using unseen writer styles from zero-shot prompts, and explore ancient or domain-specific character sets.
While the results demonstrate high confidence across both handwritten and scene text benchmarks, the article notes that gains in scene text recognition are more modest when baseline models are already trained on tens of millions of synthetic images. Practitioners should account for the computational overhead required to train diffusion models and generate large-scale datasets when evaluating deployment timelines and cost-benefit trade-offs.
- Paper: Synthetic Data for Text Localisation in Natural Images, Ankush Gupta et al. (2016). Introduces foundational synthetic text image generation techniques for training robust text recognition models, establishing the core problem that CTIG-DM seeks to solve using modern diffusion architectures.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). Presents ControlNet, the essential structural conditioning framework for diffusion models that underpins CTIG-DM's multi-condition control over text layout, style, and image content.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Establishes the fundamental Denoising Diffusion Probabilistic Model (DDPM) paradigm utilized by CTIG-DM for high-fidelity conditional synthesis.
- Paper: Diffusion Models Beat GANs on Image Synthesis, Prafulla Dhariwal et al. (2021). Demonstrates guided conditional generation in diffusion models, providing the core mechanisms for steering image synthesis with auxiliary conditions.
- Paper: Blended Diffusion for Text-driven Editing of Natural Images, Omri Avrahami et al. (2021). Introduces mask-guided, localized image editing and synthesis with diffusion models, which directly informs CTIG-DM's text recovery and augmentation modes.
- Paper: An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion, Rinon Gal et al. (2022). Formulates concept and style conditioning in text-to-image models via learned embeddings, which is critical for CTIG-DM's style-conditioned imitation mode.
- Paper: TextDiffuser: Diffusion Models as Text Painters, Jingye Chen et al. (2023). Extends the concept of text rendering in diffusion models by introducing explicit character-level layout planning and character-aware loss objectives for accurate visual text generation.
- Paper: DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing, Chong Mou et al. (2024). Builds upon conditional diffusion-based editing by developing prompt encoders and localized gradient guidance for higher precision in visual manipulation.
- Paper: MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis, Dewei Zhou et al. (2024). Advances multi-condition spatial control in diffusion models by using attention controllers to isolate individual instances and prevent attribute leakage.
- Paper: Generating Images of Rare Concepts Using Pre-trained Diffusion Models, Dvir Samuel et al. (2024). Complements CTIG-DM's generation of out-of-vocabulary words by optimizing initial noise spaces to synthesize rare visual concepts from limited examples without retraining.
