Conditional Text Image Generation with Diffusion Models

Yuanzhi ZhuZhaohai LiTianwei WangMengchao HeCong Yao

article2023CVPR112 citations

Proposes a conditional diffusion model that controls text, style, and visual attributes across four generation modes to synthesize realistic scene and handwritten text images that improve downstream recognition accuracy and handle out-of-vocabulary words.

Listen

Modern text recognition systems for scene text and handwriting rely heavily on massive, diverse datasets to achieve high accuracy. However, collecting and manually annotating real-world text images is prohibitively expensive and time-consuming. While traditional generative adversarial networks (GANs) and standard data augmentation tools have been used to synthesize artificial training data, they frequently suffer from training instability, mode collapse, and an inability to reliably create realistic out-of-vocabulary words or adapt to new visual domains.

The article introduces and evaluates Conditional Text Image Generation with Diffusion Models (CTIG-DM), a framework designed to generate realistic, diverse, and controllable text images. The main objective is to demonstrate how diffusion-based generative models—guided by text-specific conditions—can produce high-fidelity synthetic text images that directly improve the performance of downstream text recognizers.

The authors developed an architecture comprising a conditional encoder and a denoising diffusion model. The encoder extracts three specialized signals from pre-trained text recognition networks: image conditions (visual appearance and texture), text conditions (semantic character sequences), and style conditions (author-specific handwriting traits). By combining these inputs, the model operates across four operational modes: synthesis, augmentation, recovery, and imitation. The approach was systematically evaluated on standard handwritten benchmarks (IAM, RIMES, CVL, and CASIA-HWDB) and multiple scene text benchmarks (such as ICDAR, SVT, and IIIT 5K), measuring image quality, error rates, out-of-vocabulary generation, and cross-dataset domain adaptation.

The experimental findings show that CTIG-DM outperforms existing generative approaches across all core evaluation criteria. First, generating text images with all three conditions yielded superior visual fidelity, lowering the Fréchet Inception Distance (FID) to 9.34 on the IAM handwritten dataset compared to 33.42 for an unconditional baseline. Second, augmenting real datasets with CTIG-DM synthetic images substantially reduced recognizer error rates: on the IAM dataset, word error rate (WER) fell by 4.78 percentage points and character error rate (CER) fell by 1.87 percentage points, with comparable improvements of 5.20 (WER) and 1.52 (CER) percentage points on the RIMES French dataset. Third, in real-world scene text recognition, adding 2 million synthetic samples reduced error rates and boosted recognizer accuracy by 4.9% for CRNN and 1.5% for ABINet. Fourth, the model demonstrated strong out-of-vocabulary synthesis, producing rare characters, unseen English words, and non-existing Chinese character compositions at an FID of 25.52—roughly four times better than existing GAN baselines. Finally, in domain adaptation tasks, pre-training with synthetic images reduced CVL benchmark word error rate by 16.40 percentage points over baseline.

These results indicate that diffusion-based synthetic data generation offers a reliable, scalable alternative to manual data collection. By providing realistic variations in writing styles, fonts, backgrounds, and orientations, the approach reduces the cost and operational risk of deploying automated text recognition in data-scarce environments. Because the framework is complementary to existing geometric augmentations and text recognition architectures, engineering teams can integrate it directly into existing training pipelines without altering underlying downstream recognizers.

Based on these findings, decision-makers should consider adopting diffusion-driven synthesis pipelines when developing recognition models for rare vocabularies, specialized scripts, or new deployment domains. To further extend this capability, technical teams should pilot the method on additional languages, test generation using unseen writer styles from zero-shot prompts, and explore ancient or domain-specific character sets.

While the results demonstrate high confidence across both handwritten and scene text benchmarks, the article notes that gains in scene text recognition are more modest when baseline models are already trained on tens of millions of synthetic images. Practitioners should account for the computational overhead required to train diffusion models and generate large-scale datasets when evaluating deployment timelines and cost-benefit trade-offs.

arXiv: 2306.10804
Cover for Conditional Text Image Generation with Diffusion Models

Abstract

Current text recognition systems, including those for handwritten scripts and scene text, have relied heavily on image synthesis and augmentation, since it is difficult to realize real-world complexity and diversity through collecting and annotating enough real text images. In this paper, we explore the problem of text image generation, by taking advantage of the powerful abilities of Diffusion Models in generating photo-realistic and diverse image samples with given conditions, and propose a method called Conditional Text Image Generation with Diffusion Models (CTIG-DM for short). To conform to the characteristics of text images, we devise three conditions: image condition, text condition, and style condition, which can be used to control the attributes, contents, and styles of the samples in the image generation process. Specifically, four text image generation modes, namely: (1) synthesis mode, (2) augmentation mode, (3) recovery mode, and (4) imitation mode, can be derived by combining and configuring these three conditions. Extensive experiments on both handwritten and scene text demonstrate that the proposed CTIG-DM is able to produce image samples that simulate real-world complexity and diversity, and thus can boost the performance of existing text recognizers. Besides, CTIG-DM shows its appealing potential in domain adaptation and generating images containing Out-Of-Vocabulary (OOV) words.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Text Recognition
  • 2.2. Text Image Augmentation and Synthesis
  • 2.3. Diffusion Models
  • 3. Methodology
  • 3.1. Conditional Encoder
  • 3.2. Conditional Diffusion Model
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Implementation Details
  • 4.3. Ablation Study
  • 4.4. Recognition Performance
  • 4.4.1 Handwritten Text Recognition
  • 4.4.2 Scene Text Recognition
  • 4.5. Quality of Generated Images
  • 4.6. Generating Out-of-Vocabulary Images
  • 4.7. Domain Adaptation
  • 4.8. Applications
  • 5. Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — CTIG-DM conditional diffusion generator

    model/method

    Conditional Text Image Generation with Diffusion Models (CTIG-DM) is a text-image generator composed of a conditional encoder and a conditional diffusion model. The conditional encoder extracts representations of an input image, a target text string, and, for handwritten text, a writer identity. A UNet-based diffusion model then generates a text image by denoising random Gaussian noise while receiving selected combinations of these representations.

    The three conditions are designed to control different properties: the image condition primarily preserves visual attributes such as texture, color, and overall appearance; the text condition specifies the character content and character-level context; and the style condition controls the personal writing style of a particular handwritten-text writer. This design targets the simultaneous requirements of text-image generation: visual fidelity, sample diversity, and correctness of the text rendered in the image.

  2. Knowl 2 — Condition encoder for image, text, and handwriting style

    model/method

    CTIG-DM constructs three condition representations. For an input text image II, the image condition is obtained from a pretrained text recognizer rather than directly from generic image features:

    ci=AttnPool⁡(Fenc(I)+Emb⁡(Pi)).c_i=\operatorname{AttnPool}\left(F_{\mathrm{enc}}(I)+\operatorname{Emb}(P_i)\right).

    Here, cic_i is an image-level condition, FencF_{\mathrm{enc}} is the feature extractor of the pretrained text recognizer, PiP_i indexes image patches, Emb⁡\operatorname{Emb} is a positional-embedding function, and AttnPool⁡\operatorname{AttnPool} aggregates patch-level features. This is intended to emphasize character and textural information while avoiding excessive dependence on image backgrounds.

    For a text string, let TT be its sequence of one-hot character encodings, let PtP_t be the corresponding character-position indices, and let WcW_c be the classifier-weight matrix of the pretrained text recognizer. The text condition is

    ct=Proj⁡(WcT+Emb⁡(Pt)),c_t=\operatorname{Proj}\left(W_cT+\operatorname{Emb}(P_t)\right),

    where Proj⁡\operatorname{Proj} is a learned linear projection that puts the representation into the condition space. Using classifier weights allows the encoder to represent character sequences, including out-of-vocabulary words, rather than restricting generation to a fixed learned word-embedding vocabulary.

    For handwritten text, let SS be a writer identity. The style condition is

    cs=Proj⁡(Emb⁡(S)).c_s=\operatorname{Proj}\left(\operatorname{Emb}(S)\right).

    The style condition represents personal writing characteristics such as character slant, cursive joins, and stroke width; it is used for handwritten-text generation and is not used as a scene-text condition.

  3. Knowl 3 — Conditional diffusion training and generation

    model/method

    The conditional diffusion model is a UNet with a forward noising process and a reverse generation process. Let InI_n be an image at discrete diffusion step nn, and let ϵ\epsilon denote Gaussian noise. The paper formulates the forward process as

    In+1=In+ϵ.I_{n+1}=I_n+\epsilon.

    Given the noisy image, the selected condition representations, and the diffusion-step embedding, the UNet predicts the added noise. With θ\theta denoting the UNet parameters and [⋅][\cdot] denoting concatenation, training minimizes

    L=∥ϵ−ϵθ([ci,ct,cs],In+1)∥2.L=\left\lVert\epsilon-\epsilon_{\theta}\left([c_i,c_t,c_s],I_{n+1}\right)\right\rVert^2.

    A condition omitted by a particular generation mode is simply not included in the concatenated condition input.

    During generation, In′I_{n'} is initialized with random Gaussian noise. At reverse-generation step n′n', let c\mathbf{c} denote the condition combination selected for the desired mode. The next image is produced according to the paper's reverse update

    In′+1=In′−ϵθ(c,In′).I_{n'+1}=I_{n'}-\epsilon_{\theta}(\mathbf{c},I_{n'}).

    Repeated denoising updates transform the random initial noise into a text image controlled by the supplied conditions.

  4. Knowl 4 — Four controllable generation modes

    model/method

    CTIG-DM derives four generation modes by selecting subsets of the image condition cic_i, text condition ctc_t, and handwriting-style condition csc_s:

    Could not parse LaTeX table

    Synthesis generates diverse images from specified text without requiring a source image. Augmentation preserves visual characteristics from an input image but does not explicitly provide its text content. Recovery combines image and text conditions to retain appearance while enforcing the intended characters. Imitation additionally specifies a handwriting style and is used for handwritten text; for scene text, recovery is the corresponding highest-fidelity mode.

    The paper reports that text conditioning is especially important for diversity and out-of-vocabulary generation, while image and style conditioning improve fidelity. The mixed training data called MIX consists of equal proportions of synthesis samples and imitation samples for handwritten text, or synthesis and recovery samples for scene text.

  5. Knowl 5 — Experimental training configuration

    experimental setup

    For handwritten text, CTIG-DM was evaluated on IAM, RIMES, CVL, and CASIA-HWDB 1.0–1.1. IAM contains more than 115,000 English words from 657 writers; RIMES contains more than 60,000 French words from more than 1,000 authors; the English portion of CVL contains text from 311 writers and is used for domain adaptation; and CASIA-HWDB provides 2,678,424 offline handwritten Chinese-character images for out-of-vocabulary generation. For scene text, MJSynth, SynthText, and Real-L were used for training, with IIIT 5K-Words, SVT, IC03, IC13, IC15, SVTP, and CUTE80 used for testing.

    Handwritten training images were resized to height 64 pixels while preserving aspect ratio, with width capped at 256 pixels. The recognizer used a CRNN architecture, and the diffusion model used a DDPM-style UNet. All condition representations had dimensionality 512 and were concatenated with diffusion time-step embeddings. To prevent over-reliance on any one condition, the image, text, and style conditions were independently replaced by a learnable embedding for 20%, 10%, and 20% of training instances, respectively.

    The models were optimized with AdamW using β1=0.9\beta_1=0.9, β2=0.999\beta_2=0.999, weight decay 0.20.2, an initial learning rate of 0.00010.0001, and cosine annealing. Batch sizes were 64 for recognizer training and 256 for diffusion-model training, using NVIDIA Tesla V100 GPUs. Handwritten recognition was measured with word error rate (WER) and character error rate (CER); image quality was measured with FID, geometry score (GS), SSIM, RMSE, and LPIPS. Lower WER, CER, FID, GS, RMSE, and LPIPS are better, while higher SSIM is better.

  6. Knowl 6 — Condition ablation improves handwritten-image fidelity

    empirical result

    On IAM, removing all conditions produced weak image quality, while each added condition improved the measured metrics. The condition abbreviations are cic_i for the image condition, ctc_t for the text condition, and csc_s for the writer-style condition.

    Could not parse LaTeX table

    The text condition had the larger individual effect in this ablation, reducing FID from the unconditional value of 33.42 to 10.92 compared with 13.36 for the image condition. Combining image and text conditions reduced FID to 9.76, and adding the style condition produced the best FID of 9.34 together with the best SSIM, RMSE, and LPIPS values. Qualitative samples showed that synthesis produced broad variation in handwriting slants, ink marks, joins, stroke widths, and paper backgrounds, whereas image and style conditions made generated samples more faithful to the source appearance and writer style.

  7. Knowl 7 — Generated handwritten data improves HTR

    empirical result

    Adding CTIG-DM samples to handwritten-text-recognition training improved CRNN recognition on both IAM and RIMES. WER and CER are percentages, and lower values are better. The baseline uses only the real training set; Ours uses CTIG-DM MIX data; and Ours + Aug* additionally applies random geometric augmentation to generated samples.

    Could not parse LaTeX table

    With only generated data, approximately 200K IAM samples reached performance comparable to the real-data baseline. Increasing the MIX amount steadily reduced error, reaching IAM CER 5.60% and WER 14.03% at 900K generated samples. Combining real and generated data was more effective than either source alone. CTIG-DM also outperformed the listed prior data-augmentation and handwritten-synthesis methods under the paper's comparison protocol, and its combination with geometric augmentation further improved recognition.

  8. Knowl 8 — Generated data improves scene-text recognition

    empirical result

    CTIG-DM improved both CRNN and ABINet without changing their recognition architectures. Recognition scores are percentages, and higher values are better. The six test sets are IC13, SVT, IIIT 5K-Words, IC15, SVTP, and CUTE80; Average is the unweighted mean across them. In the synthetic-data setting, ST denotes SynthText and MJ+ST denotes MJSynth plus SynthText. MIX is the proposed generated data.

    Could not parse LaTeX table

    For the synthetic-data setting, CTIG-DM raised the average score by 1.7 percentage points for CRNN and 0.5 percentage points for ABINet. On the smaller real-data setting, the gains were 4.9 points for CRNN and 1.5 points for ABINet. The paper notes that gains are smaller when the recognizer already uses roughly 14 million synthetic training images, because the diversity and fidelity of CTIG-DM samples are partly limited by the synthesized training distribution.

  9. Knowl 9 — Image quality and out-of-vocabulary generation

    empirical result

    CTIG-DM produced stronger handwritten-image quality than the compared synthesis methods. FID was computed using 25,000 real and 25,000 generated images, while GS used 5,000 real and 5,000 generated images. Lower values are better.

    Could not parse LaTeX table

    Because synthesis mode uses the text condition without requiring a source image, changing the text input allows CTIG-DM to generate words absent from the training vocabulary. On IAM, the FID for these out-of-vocabulary images was 25.52, compared with 125.87 for GANwriting, 109.45 for HWT, 97.81 for SLOGAN, and 104.81 for CG-GAN. Supplying a style condition alongside an out-of-vocabulary text string also enabled control over the writer appearance in the generated images.

  10. Knowl 10 — Domain adaptation and generation beyond ordinary vocabularies

    empirical result

    CTIG-DM was trained on IAM and then used to generate 100K handwritten samples whose text strings were drawn from the CVL training lexicon. A recognizer trained on IAM plus these generated samples was evaluated on CVL; results were averaged over five training runs. WER and CER are percentages, and lower values are better.

    Could not parse LaTeX table

    Relative to the IAM-only baseline, CTIG-DM reduced CVL WER by 16.40 percentage points and CER by 7.60 points, exceeding the improvements of the compared GAN-based approaches. This demonstrates that the generator can synthesize useful target-domain appearance while using a different source-domain handwriting corpus.

    For handwritten Chinese, the method represents text conditions as radical sequences rather than treating each character as an indivisible vocabulary item. Radical sequences absent from training can therefore be used to generate out-of-vocabulary or even non-existing handwritten characters, including typo-like forms. The same strategy was proposed for generating rare ancient characters, whose scarcity creates a long-tailed recognition problem.

Coverage note — The qualitative sample grids and detailed row-by-row comparisons against every prior method were not made separate knowls because they illustrate the already captured fidelity, controllability, and performance claims rather than adding independent load-bearing results.

References

  1. 1.Eloi Alonso, Bastien Moysset, and Ronaldo O. Messina. Adversarial generation of handwritten text images conditioned on sequences. In ICDAR, pages 481–486, 2019. 2, 6, 7
  2. 2.Omri Avrahami, Dani Lischinski, and Ohad Fried. Blended diffusion for text-driven editing of natural images. In CVPR, pages 18187–18197, 2022. 2
  3. 3.Amandeep Kumar Pinaki Nath Chowdhury Aneeshan Sain Yi-Zhe Song Ayan Kumar Bhunia, Shuvozit Ghose. Metahtr: Towards writer-adaptive handwritten text recognition. In CVPR, 2021. 1, 2
  4. 4.Jeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park, Dongyoon Han, Sangdoo Yun, Seong Joon Oh, and Hwalsuk Lee. What is wrong with scene text recognition model comparisons? dataset and model analysis. In ICCV, pages 4714–4722, 2019. 7
  5. 5.Jeonghun Baek, Yusuke Matsui, and Kiyoharu Aizawa. What if we only use real datasets for scene text recognition? toward scene text recognition with fewer labels. In CVPR, pages 3113–3122, 2021. 1, 4
  6. 6.Ayan Kumar Bhunia, Abhirup Das, Ankan Kumar Bhunia, Perla Sai Raj Kishore, and Partha Pratim Roy. Handwriting recognition in low-resource scripts using adversarial learning. In CVPR, pages 4762–4771, 2019. 2, 4, 5, 6
  7. 7.Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, and Mubarak Shah. Handwriting transformers. In ICCV, pages 1086–1094, 2021. 1, 7, 8
  8. 8.Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018. 2
  9. 9.Xiaoxue Chen, Lianwen Jin, Yuanzhi Zhu, Canjie Luo, and Tianwei Wang. Text recognition in the wild: A survey. ACM Comput. Surv., 54(2):42:1–42:35, 2022. 2
  10. 10.Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon. ILVR: conditioning method for denoising diffusion probabilistic models. In ICCV, pages 14347–14356, 2021. 2
  11. 11.Hyungjin Chung, Byeongsu Sim, and Jong Chul Ye. Come-closer-diffuse-faster: Accelerating conditional diffusion models for inverse problems through stochastic contraction. In CVPR, pages 12403–12412, 2022. 2
  12. 12.Cheng Da, Peng Wang, and Cong Yao. Levenshtein OCR. In ECCV, pages 322–338, 2022. 2
  13. 13.Brian L. Davis, Bryan S. Morse, Brian L. Price, Chris Tensmeyer, Curtis Wigington, and Rajiv Jain. Text and style conditioned GAN for the generation of offline-handwriting lines. In BMVC, 2020. 7
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, pages 4171–4186, 2019. 3
  15. 15.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat gans on image synthesis. In NeurIPS, pages 8780–8794, 2021. 1, 2, 3, 4
  16. 16.Shancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao, and Yongdong Zhang. Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition. In CVPR, pages 7098–7107, 2021. 1, 2, 7
  17. 17.Sharon Fogel, Hadar Averbuch-Elor, Sarel Cohen, Shai Mazor, and Roee Litman. ScrabbleGAN: Semi-supervised varying length handwritten text generation. In CVPR, pages 4323–4332, 2020. 1, 2, 4, 6, 7, 8
  18. 18.Ji Gan and Weiqiang Wang. Higan: Handwriting imitation conditioned on arbitrary-length texts and disentangled styles. In AAAI, pages 7484–7492, 2021. 1, 7
  19. 19.Alex Graves, Santiago Fernandez, Faustino J. Gomez, and ´Jurgen Schmidhuber. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In ICML, pages 369–376, 2006. 2
  20. 20.Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman. Synthetic data for text localisation in natural images. In CVPR, pages 2315–2324, 2016. 1, 4
  21. 21.Pan He, Weilin Huang, Yu Qiao, Chen Change Loy, and Xiaoou Tang. Reading scene text in deep convolutional sequences. In AAAI, pages 3501–3508, 2016. 2
  22. 22.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NIPS, pages 6626–6637, 2017. 4
  23. 23.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020. 2, 3, 4
  24. 24.Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Synthetic data and artificial neural networks for natural scene text recognition. In NIPS Deep Learning Workshop, 2014. 4
  25. 25.Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Reading text in the wild with convolutional neural networks. IJCV, 116(1):1–20, 2016. 1, 4
  26. 26.Lei Kang, Pau Riba, Yaxing Wang, Marçal Rusinol, Alicia Fornes, and Mauricio Villegas. GANwriting: Content-conditioned generation of styled handwritten word images. In ECCV, pages 273–289, 2020. 7, 8
  27. 27.Lei Kang, Marçal Rusinol, Alicia Fornés, Pau Riba, and Mauricio Villegas. Unsupervised writer adaptation for synthetic-to-real handwritten word recognition. In WACV, pages 3502–3511, 2020. 2, 6
  28. 28.Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, and Masakazu Iwamura. ICDAR 2015 competition on robust reading. In ICDAR, pages 1156–1160, 2015. 4
  29. 29.Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, and Sergi Robles Mestre. ICDAR 2013 robust reading competition. In ICDAR, pages 1484–1493, 2013. 4
  30. 30.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8110–8119, 2020. 2
  31. 31.Valentin Khrulkov and Ivan V. Oseledets. Geometry score: A method for comparing generative adversarial networks. In ICML, pages 2626–2634, 2018. 4
  32. 32.Florian Kleber, Stefan Fiel, Markus Diem, and Robert Sablatnig. CVL-database: An offline database for writer retrieval, writer identification and word spotting. In ICDAR, pages 560–564, 2013. 4
  33. 33.Yuxin Kong, Canjie Luo, Weihong Ma, Qiyuan Zhu, Shenggao Zhu, Nicholas Yuan, and Lianwen Jin. Look closer to supervise better: One-shot font generation via component-based discriminator. In CVPR, pages 13472–13481, 2022. 1, 7, 8
  34. 34.Juho Lee, Yoonho Lee, Jungtaek Kim, Adam R. Kosiorek, Seungjin Choi, and Yee Whye Teh. Set transformer: A framework for attention-based permutation-invariant neural networks. In ICML, pages 3744–3753, 2019. 3
  35. 35.Minghui Liao, Pengyuan Lyu, Minghang He, Cong Yao, Wenhao Wu, and Xiang Bai. Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes. IEEE TPAMI, 43(2):532–548, 2021. 7
  36. 36.Cheng-Lin Liu, Fei Yin, Da-Han Wang, and Qiu-Feng Wang. CASIA online and offline chinese handwriting databases. In ICDAR, pages 37–41, 2011. 4
  37. 37.Shangbang Long, Xin He, and Cong Yao. Scene text detection and recognition: The deep learning era. International Journal of Computer Vision, 129:161–184, 2018. 2
  38. 38.Shangbang Long and Cong Yao. Unrealtext: Synthesizing realistic scene text images from the unreal world. In CVPR, pages 5488–5497, 2020. 1
  39. 39.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019. 5
  40. 40.Simon M. Lucas, Alex Panaretos, Luis Sosa, Anthony Tang, Shirley Wong, and Robert Young. ICDAR 2003 robust reading competitions: entries, results, and future directions. Int. J. Document Anal. Recognit., 7(2-3):105–122, 2005. 4
  41. 41.Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. RePaint: Inpainting using denoising diffusion probabilistic models. In CVPR, pages 11451–11461, 2022. 2
  42. 42.Troy Luhman and Eric Luhman. Diffusion models for handwriting generation. CoRR, abs/2011.06704, 2020. 3
  43. 43.Canjie Luo, Lianwen Jin, and Zenghui Sun. MORAN: A multi-object rectified attention network for scene text recognition. PR, 90:109–118, 2019. 1, 2, 7
  44. 44.Canjie Luo, Yuanzhi Zhu, Lianwen Jin, Zhe Li, and Dezhi Peng. SLOGAN: Handwriting style synthesis for arbitrary-length and out-of-vocabulary text. IEEE TNNLS, pages 1–13, 2022. 1, 2, 4, 5, 6, 7, 8
  45. 45.Canjie Luo, Yuanzhi Zhu, Lianwen Jin, and Yongpan Wang. Learn to augment: Joint data augmentation and network optimization for text recognition. In CVPR, pages 13743–13752, 2020. 1, 2, 4, 5, 6
  46. 46.U-V Marti and Horst Bunke. The IAM-database: an english sentence database for offline handwriting recognition. Int. J. Document Anal. Recognit., 5(1):39–46, 2002. 1, 4
  47. 47.Anand Mishra, Karteek Alahari, and C. V. Jawahar. Scene text recognition using higher order language priors. In BMVC, pages 1–11, 2012. 4
  48. 48.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In ICML, pages 8162–8171, 2021. 1, 2, 4
  49. 49.Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, and Bob McGrew. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In ICML, pages 16784–16804, 2022. 3, 5
  50. 50.Trung Quy Phan, Palaiahnakote Shivakumara, Shangxuan Tian, and Chew Lim Tan. Recognizing text with perspective distortion in natural scenes. In ICCV, pages 569–576, 2013. 4
  51. 51.Zhi Qiao, Yu Zhou, Dongbao Yang, Yucan Zhou, and Weiping Wang. SEED: Semantics enhanced encoder-decoder framework for scene text recognition. In CVPR, pages 13528–13537, 2020. 7
  52. 52.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, and Sandhini Agarwal. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763, 2021. 2
  53. 53.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with CLIP latents, 2022. 2, 3
  54. 54.Anhar Risnumawan, Palaiahnakote Shivakumara, Chee Seng Chan, and Chew Lim Tan. A robust arbitrary text detection system for natural scene images. Expert Syst. Appl., 41(18):8027–8048, 2014. 4
  55. 55.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, pages 10674–10685, 2022. 2
  56. 56.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional networks for biomedical image segmentation. In MICCAI, pages 234–241, 2015. 4
  57. 57.Baoguang Shi, Xiang Bai, and Cong Yao. An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE TPAMI, 39(11):2298–2304, 2016. 1, 2, 4, 5, 7
  58. 58.Baoguang Shi, Mingkun Yang, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai. ASTER: An attentional scene text recognizer with flexible rectification. IEEE TPAMI, 2019. 1, 2, 7
  59. 59.Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, pages 2256–2265, 2015. 2, 3
  60. 60.Jorge Sueiras, Victoria Ruíz, Angel Sánchez, and José Francisco Velez. Offline continuous handwriting recognition using sequence to sequence neural networks. Neurocomputing, 289:119–128, 2018. 4, 6
  61. 61.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, pages 5998–6008, 2017. 2
  62. 62.Zhaoyi Wan, Minghang He, Haoran Chen, Xiang Bai, and Cong Yao. TextScanner: Reading characters in order for robust scene text recognition. In AAAI, number 07, pages 12120–12127, 2020. 7
  63. 63.Kai Wang, Boris Babenko, and Serge J. Belongie. End-to-end scene text recognition. In ICCV, pages 1457–1464, 2011. 4
  64. 64.Peng Wang, Cheng Da, and Cong Yao. Multi-granularity prediction for scene text recognition. In ECCV, pages 339–355, 2022. 2
  65. 65.Tianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo, Xiaoxue Chen, Yaqiang Wu, Qianying Wang, and Mingxiang Cai. Decoupled attention network for text recognition. In AAAI, 2020. 1, 2, 7
  66. 66.Yuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang, Shenggao Zhu, and Yongdong Zhang. From two to one: A new scene text recognizer with visual language modeling network. In ICCV, pages 14174–14183, 2021. 7
  67. 67.Curtis Wigington, Seth Stewart, Brian L. Davis, Bill Barrett, Brian L. Price, and Scott Cohen. Data augmentation for recognition of handwritten words and lines using a CNN-LSTM network. In ICDAR, pages 639–645, 2017. 1, 2
  68. 68.Yan Wu, Jeff Donahue, David Balduzzi, Karen Simonyan, and Timothy Lillicrap. Logan: Latent optimisation for generative adversarial networks. arXiv preprint arXiv:1912.00953, 2019. 2
  69. 69.Deli Yu, Xuan Li, Chengquan Zhang, Junyu Han, Jingtuo Liu, and Errui Ding. Towards accurate scene text recognition with semantic reasoning networks. In CVPR, pages 12113–12122, 2020. 1, 2, 4, 7
  70. 70.Xiaoyu Yue, Zhanghui Kuang, Chenhao Lin, Hongbin Sun, and Wayne Zhang. RobustScanner: Dynamically enhancing positional clues for robust text recognition. In ECCV, pages 135–151, 2020. 7
  71. 71.Fangneng Zhan and Shijian Lu. ESIR: End-to-end scene text recognition via iterative image rectification. In CVPR, pages 2059–2068, 2019. 7
  72. 72.Jianshu Zhang, Jun Du, and Lirong Dai. Multi-scale attention with dense encoder for handwritten mathematical expression recognition. In International Conference on Pattern Recognition, pages 2245–2250, 2018. 1
  73. 73.Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, pages 586–595, 2018. 4
  74. 74.Yaping Zhang, Shuai Nie, Wenju Liu, Xing Xu, Dongxiang Zhang, and Heng Tao Shen. Sequence-to-sequence domain adaptation network for robust text image recognition. In CVPR, pages 2740–2749, 2019. 2, 6
  75. 75.Yingying Zhu, Cong Yao, and Xiang Bai. Scene text detection and recognition: recent advances and future trends. Frontiers of Computer Science, 10:19–36, 2016. 2

Citation

MLA
Zhu, Y., et al. “Conditional Text Image Generation with Diffusion Models”. arXiv, 2023, http://arxiv.org/abs/2306.10804v1.
APA
Zhu, Y., Li, Z., Wang, T., He, M., & Yao, C. (2023). Conditional Text Image Generation with Diffusion Models. arXiv. http://arxiv.org/abs/2306.10804v1
Chicago
Zhu, Y., Z. Li, T. Wang, M. He, and C. Yao. 2023. “Conditional Text Image Generation with Diffusion Models”. arXiv. http://arxiv.org/abs/2306.10804v1.
Harvard
Zhu, Y. et al. (2023) “Conditional Text Image Generation with Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.10804v1.
Vancouver
1. Zhu Y, Li Z, Wang T, He M, Yao C (2023) Conditional Text Image Generation with Diffusion Models. arXiv

BibTeX

@article{zhu2023conditional,
  title = {Conditional Text Image Generation with Diffusion Models},
  author = {Zhu, Yuanzhi and Li, Zhaohai and Wang, Tianwei and He, Mengchao and Yao, Cong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.10804v1},
  eprint = {2306.10804}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE