Character-Aware Models Improve Visual Text Rendering

Rosanne LiuDan GarretteChitwan SahariaWilliam ChanAdam RobertsSharan NarangIrina BlokRJ MicalMohammad NorouziNoah Constant

article2023ACL105 citations

Demonstrates that integrating character-aware text encoders into text-to-image architectures significantly improves visual spelling accuracy, outperforming conventional subword-based models on complex and rare text rendering tasks.

Listen

Modern text-to-image artificial intelligence systems frequently fail to render legible, correctly spelled visual text in generated pictures. This issue limits their practical utility for graphic design, marketing materials, and digital art. The primary cause of this failure is that standard image generators rely on subword-level text encoders that do not receive direct character-by-character information, making it difficult for the models to accurately translate words into visual letter sequences.

The main objective of the article is to evaluate why popular image generation models fail at visual text rendering and to demonstrate that incorporating character-aware text encoders significantly improves spelling and text generation accuracy.

The authors conducted two phases of evaluation. First, they assessed the spelling capabilities of isolated text encoders across multiple model scales and seven languages using a new benchmark called WikiSpell. Second, they trained several image generation models on a standardized dataset of 400 million image-caption pairs using both standard character-blind text encoders and character-aware text encoders. They evaluated visual rendering performance using DrawText, a new benchmark comprising automated optical character recognition testing across 500 word prompts and human evaluation across 175 diverse creative prompts.

The experiments produced four critical findings. First, character-aware text encoders consistently achieved high spelling accuracy in text-only tasks across all word frequencies and languages, whereas standard character-blind models required over 100 billion parameters to achieve comparable spelling reliability. Second, image generation models using character-aware text encoders showed dramatic improvements in visual spelling, yielding accuracy gains of over 15 percentage points on common words and more than 30 percentage points on rare words compared to character-blind baselines, despite requiring fewer training steps. Third, character-blind models frequently made systematic errors such as regularizing irregular verbs into incorrect past-tense spellings, while character-aware models avoided these systematic errors entirely. Fourth, while purely character-level models experienced a minor drop in overall prompt alignment for non-text image elements, a hybrid approach combining standard token-based encodings with a lightweight character-aware encoder delivered state-of-the-art text rendering without degrading overall image quality or prompt fidelity.

These findings indicate that image generation pipelines can achieve superior visual text rendering without relying on massive, computationally expensive language models. Integrating lightweight character-aware representations provides a cost-effective and scalable pathway to eliminate common visual spelling failures, thereby improving model utility for automated design workflows.

Organizations developing or deploying image generation models should adopt hybrid text encodings that merge standard subword representations with character-aware inputs. For future technical development, engineering teams should focus on improving the downstream image generation modules to resolve remaining visual artifacts, such as letter merging or glyph placement, and expand training pipelines to support visual text rendering in non-Latin writing systems.

The findings are supported by consistent results across automated optical character recognition and human preference evaluations. However, readers should note that the visual image experiments were limited to English-language Latin text, and future verification is needed to determine how these text-encoding methods perform when encoders are trained jointly with image generation systems rather than kept frozen.

arXiv: 2212.10562master/drawtext
  • Paper: TextDiffuser: Diffusion Models as Text Painters, Jingye Chen et al. (2023). TextDiffuser carries character-level guidance into explicit layout planning and text-focused image synthesis, extending the source’s encoder-based spelling improvements into a controllable rendering pipeline.
Cover for Character-Aware Models Improve Visual Text Rendering

Abstract

Current image generation models struggle to reliably produce well-formed visual text. In this paper, we investigate a key contributing factor: popular text-to-image models lack character-level input features, making it much harder to predict a word’s visual makeup as a series of glyphs. To quantify this effect, we conduct a series of experiments comparing character-aware vs. character-blind text encoders. In the text-only domain, we find that character-aware models provide large gains on a novel spelling task (WikiSpell). Applying our learnings to the visual domain, we train a suite of image generation models, and show that character-aware variants outperform their character-blind counterparts across a range of novel text rendering tasks (our DrawText benchmark). Our models set a much higher state-of-the-art on visual spelling, with 30+ point accuracy gains over competitors on rare words, despite training on far fewer examples.

Table of Contents

  • 1 Introduction
  • 2 The spelling miracle
  • 3 Measuring text encoder spelling ability
  • 3.1 The WikiSpell benchmark
  • 3.2 Text generation experiments
  • 4 The DrawText benchmark
  • 4.1 DrawText Spelling
  • 4.2 DrawText Creative
  • 5 Image generation experiments
  • 5.1 Models
  • 5.2 DrawText Spelling results
  • 5.3 DrawText Creative results
  • 5.4 DrawBench results
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Multilingual results
  • B OCR error estimation
  • C Per-category DrawBench analysis
  • D WikiSpell details
  • E Additional DrawText creative samples
  • F Representative DrawText Spelling samples
  • G DrawText Creative prompts
  • Prompts used in Figure 3

Knowls

  1. Knowl 1 — Character-aware encoders substantially improve visual spelling

    empirical result

    On DrawText Spelling, character-aware image-generation models using ByT5 or a concatenated T5–ByT5 encoder achieved higher exact-string OCR accuracy than character-blind T5 models across all five English word-frequency buckets. In controlled comparisons using the same image-training data and number of steps, character-aware models gained more than 25 percentage points on the most frequent words and more than 30 points on the least frequent words. They also exceeded Imagen-AR by more than 15 points on the most frequent words and more than 30 points on the least frequent words, although Imagen-AR had trained 6.6 times longer. The paper summarizes the strongest hybrid model’s accuracy as increasing from 75% to 94% on common words and from 47% to 83% on rare words. ByT5-XL retained this advantage despite having an encoder 43% smaller than T5-XXL.

  2. Knowl 2 — A token–character hybrid improves text rendering while retaining image alignment

    model/method

    The Concat model forms its text representation by concatenating encodings from the pretrained T5-XXL encoder and a 220-million-parameter ByT5-Small encoder. The ByT5 addition increased encoder size by 4.8% and supplied character-level input alongside T5’s token-level representation. On DrawText Creative’s 175 prompts, human raters preferred Concat to T5-XXL on image fidelity, image-text alignment, and accuracy of rendered text. In the broader DrawBench comparison, raters judged image fidelity similar across models; they preferred T5-XXL to the purely character-level ByT5 models on alignment in 60% of prompts, while Concat brought alignment to within the reported error bars of T5-XXL. Thus the hybrid improved text rendering without a significant measured loss in alignment.

  3. Knowl 3 — WikiSpell measures spelling on held-out words across frequency bands

    definition

    WikiSpell evaluates whether a text encoder can recover the characters of an input word. Each example gives the model one Wiktionary word and expects the output to be that word with spaces inserted between Unicode characters, such as elephant → e l e p h a n t. The benchmark groups words by frequency in mC4 into five buckets: the most frequent 1%, 1–10%, 10–20%, 20–30%, and the least frequent 50%, including words absent from the corpus. For each bucket, the development and test sets each contain 1,000 uniformly sampled words. The 10,000-word training set combines 5,000 uniformly sampled words from the bottom-50% bucket with 5,000 sampled in proportion to mC4 frequency; development and test words are excluded from training. The same procedure was used for English and six other languages: Arabic, Chinese, Finnish, Korean, Russian, and Thai. Unlike a benchmark limited to individual vocabulary tokens, WikiSpell is model-independent and includes words that a model represents as one or multiple tokens.

  4. Knowl 4 — English WikiSpell results show the scale cost of spelling without characters

    data/table

    The table reports English WikiSpell exact-match accuracy (%). T5 and ByT5 results are shown for frozen encoders and for models with all parameters trained; B, L, XL, and XXL denote increasing model sizes. PaLM is evaluated few-shot with 20 examples in its prompt. ByT5 scores remain high across frequency buckets even with its encoder frozen, whereas T5 remains substantially weaker; full fine-tuning helps T5 particularly on less common, multi-piece words but does little for common atomic tokens. PaLM-540B reaches at least 99% in every bucket, illustrating the scale at which a character-blind model can recover spelling robustly.

    Frequency bucket T5 frozen (B,L,XL,XXL) ByT5 frozen (B,L,XL,XXL) T5 all trained (B,L,XL,XXL) ByT5 all trained (B,L,XL,XXL) PaLM few-shot (8B,62B,540B)
    Top 1% 14,12,50,66 97,95,97,98 36,46,62,68 99,100,100,100 84,99,100
    1–10% 29,24,67,69 97,95,98,98 67,72,82,85 100,100,100,100 62,98,99
    10–20% 35,27,73,73 96,94,98,98 74,79,89,91 100,100,100,100 70,97,99
    20–30% 32,24,68,68 96,94,99,98 74,78,87,90 100,100,100,100 71,97,99
    Bottom 50% 29,22,64,65 97,95,99,98 75,77,88,90 100,100,100,100 69,97,99
  5. Knowl 5 — Multilingual WikiSpell favors byte-level character-aware models

    data/table

    The values are WikiSpell exact-match accuracy (%) averaged across all five word-frequency buckets. The mT5 and ByT5 models were fine-tuned on the combined training sets for all seven languages; their encoders were frozen. PaLM was prompted with 20 examples in the evaluated language. ByT5 reaches 98–100% in every language and size shown, while mT5’s average rises from 27% at Base to 77% at XXL. PaLM’s 540B model averages 92%, but its Thai score is 63%; the results therefore show both the strong cross-language spelling signal in ByT5 and the less uniform transfer of the very large character-blind model.

    Language mT5 (B,L,XL,XXL) ByT5 (B,L,XL,XXL) PaLM (8B,62B,540B)
    Arabic 22,60,75,87 99,99,100,99 32,68,89
    Chinese 78,76,83,84 99,98,99,99 81,93,98
    English 7,32,54,71 98,96,99,99 71,97,99
    Finnish 10,36,62,77 98,97,99,99 45,84,99
    Korean 37,58,77,81 99,99,100,99 71,88,96
    Russian 9,41,57,76 99,98,99,99 41,86,98
    Thai 29,42,46,60 99,99,99,99 22,39,63
    Average 27,49,65,77 99,98,99,99 52,79,92
  6. Knowl 6 — Image-generation comparison controls training data and aspect-ratio handling

    experimental setup

    The authors trained two character-blind models, T5-XL and T5-XXL, and three character-aware variants: ByT5-XL, ByT5-XXL, and Concat, which combines T5-XXL and ByT5-Small encodings. The T5 encoders have 1.2B and 4.6B parameters; the ByT5 encoders have 2.6B and 9.0B parameters. Custom models were trained for 500,000 steps, using only LAION-400M and only the initial 64 × 64 image-generation stage. In a random inspection of 100 LAION images, about 71% contained text and about 60% showed correspondence between caption text and image text. To reduce text clipping, 80% of examples retained their original aspect ratio and were padded with black borders, with a binary mask marking padding; the remaining 20% used center crops. T5 models received 64-token input sequences, while ByT5 models received 256-byte sequences. For comparison, Imagen-AR was Imagen fine-tuned for 380,000 additional steps, reaching 3.2 million total steps, with the uncropped aspect-ratio-preserving strategy.

  7. Knowl 7 — DrawText evaluates both controlled spelling and creative text rendering

    experimental setup

    DrawText has two complementary evaluations. DrawText Spelling contains 500 sign prompts, formed by sampling 100 English words from each WikiSpell frequency bucket and inserting each word into the template A sign with the word ... written on it. Four images are generated per prompt. Google Cloud Vision OCR supplies detected text and bounding boxes; if it returns multiple boxes, the top-most is used. Line breaks within a box are removed, case is ignored, and accuracy requires an exact match of the complete target string. This yields 2,000 generated images per model. DrawText Creative contains 175 prompts designed with a professional graphic designer to elicit text in varied visual styles and contexts, ranging from a single letter to full sentences. For the creative evaluation, the authors generated eight images per prompt from their T5 and ByT5 models and used human comparisons to assess text accuracy as well as image fidelity and prompt alignment.

  8. Knowl 8 — Character-blind and character-aware models exhibit different spelling errors

    empirical result

    Manual inspection of DrawText Spelling outputs found semantic substitutions, homophone-like spellings, and added glyphs only in the character-blind T5 models; these errors are consistent with an encoder that lacks reliable character-composition information. T5 models also sometimes regularized irregular past-tense forms, for example producing an -ed ending where the target was fought. On a hand-selected set of 23 common irregular past-tense verbs, T5-based models made this error in 11% of samples, while the character-aware models did not show it. Dropped, repeated, merged, or misshapen glyphs, and sometimes missing text, occurred across model types. The authors interpret these shared errors as image-generation layout failures rather than missing spelling knowledge. Across four samples of a word, T5 models also more often misspelled that word consistently, whereas ByT5 errors were more often sporadic.

  9. Knowl 9 — A multilingual encoder transfers prompt understanding but not script rendering

    empirical result

    In a preliminary multilingual image-generation test, the authors translated two English prompts into 11 languages and evaluated T5-XXL and ByT5-XXL. T5 showed basic prompt understanding in German, French, Spanish, Portuguese, and Russian, but appeared to ignore prompts in Greek, Hindi, Arabic, Chinese, Japanese, and Korean. ByT5 showed prompt understanding across all 11 languages. However, neither model accurately rendered text in non-Latin scripts. The authors suggest that a multilingual encoder can help map prompt meanings across languages, but that learning the glyph shapes needed for visual rendering also requires sufficient visual-text exposure in each script. The image-caption training data were estimated to be 95% English, with less than 0.1% coverage for some widely spoken languages such as Arabic and Hindi.

  10. Knowl 10 — The evidence is limited by encoder confounds and unresolved layout errors

    limitation

    The image-generation experiments use off-the-shelf pretrained encoders that differ in more than character awareness, so multilinguality and character-level input are not fully isolated in those comparisons. The image-generation models also use frozen pretrained text encoders; whether the results extend to systems trained jointly with their text encoder remains untested. Most visual experiments concern English, and the paper does not identify how character-blind models acquire spelling from web pretraining. Finally, character-aware encoders do not eliminate dropped, repeated, or merged letters or words; the authors argue that resolving these remaining layout failures will require improvements to the image-generation module, not only the text encoder.

Coverage note — The appendix’s full prompt inventory, per-example image galleries, and detailed per-category DrawBench breakdown are omitted because they provide stimuli and illustrations rather than distinct core findings; the preliminary multilingual image-generation result is retained.

References

  1. 1.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, Tero Karras, and Ming-Yu Liu. 2022. eDiff-I: Text-to-image diffusion models with an ensemble of expert denoisers. CoRR, abs/2211.01324.
  2. 2.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. arxiv:2204.02311.
  3. 3.Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. 2022. Canine: Pre-training an efficient tokenization-free encoder for language representation. Transactions of the Association for Computational Linguistics, 10:73–91.
  4. 4.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  5. 5.Philip Gage. 1994. A new algorithm for data compression. C Users J., 12(2):23–38.
  6. 6.Alex Graves. 2013. Generating sequences with recurrent neural networks. CoRR, abs/1308.0850v5.
  7. 7.Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. CLIPScore: A reference-free evaluation metric for image captioning. CoRR, abs/2104.08718.
  8. 8.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. Advances in neural information processing systems, 30.
  9. 9.Itay Itzhak and Omer Levy. 2022. Models in a spelling bee: Language models implicitly learn the character composition of tokens. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5061–5068, Seattle, United States. Association for Computational Linguistics.
  10. 10.Ayush Kaushal and Kyle Mahowald. 2022. What do tokens know about their characters and how do they know it? In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2487–2507, Seattle, United States. Association for Computational Linguistics.
  11. 11.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In European conference on computer vision, pages 740–755. Springer.
  12. 12.Sabrina J. Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y. Lee, Benoît Sagot, and Samson Tan. 2021. Between words and characters: A brief history of open-vocabulary modeling and tokenization in NLP. CoRR, abs/2112.10508.
  13. 13.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  14. 14.Telmo Pires, Eva Schlinger, and Dan Garrette. 2019. How multilingual is multilingual BERT? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4996–5001, Florence, Italy. Association for Computational Linguistics.
  15. 15.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  16. 16.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125.
  17. 17.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8821–8831. PMLR.
  18. 18.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2021. High-resolution image synthesis with latent diffusion models. CoRR, abs/2112.10752.
  19. 19.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487.
  20. 20.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021. LAION-400M: open dataset of clip-filtered 400 million image-text pairs. CoRR, abs/2111.02114.
  21. 21.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725, Berlin, Germany. Association for Computational Linguistics.
  22. 22.Ilya Sutskever, James Martens, and Geoffrey Hinton. 2011. Generating text with recurrent neural networks. In Proceedings of the 28th International Conference on International Conference on Machine Learning, ICML’11, page 1017–1024, Madison, WI, USA. Omnipress.
  23. 23.Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler. 2021. Charformer: Fast character transformers via gradient-based subword tokenization. CoRR, abs/2106.12672.
  24. 24.Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022. ByT5: Towards a token-free future with pre-trained byte-to-byte models. Transactions of the Association for Computational Linguistics, 10:291–306.
  25. 25.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  26. 26.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. 2022. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789.

Citation

MLA
Liu, R., et al. “Character-Aware Models Improve Visual Text Rendering”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 16270–97, https://doi.org/10.18653/v1/2023.acl-long.900.
APA
Liu, R., Garrette, D., Saharia, C., Chan, W., Roberts, A., Narang, S., Blok, I., Mical, R., Norouzi, M., & Constant, N. (2023). Character-Aware Models Improve Visual Text Rendering. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16270–16297. https://doi.org/10.18653/v1/2023.acl-long.900
Chicago
Liu, R., D. Garrette, C. Saharia, et al. 2023. “Character-Aware Models Improve Visual Text Rendering”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16270–97. https://doi.org/10.18653/v1/2023.acl-long.900.
Harvard
Liu, R. et al. (2023) “Character-Aware Models Improve Visual Text Rendering”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 16270–16297. Available at: https://doi.org/10.18653/v1/2023.acl-long.900.
Vancouver
1. Liu R, Garrette D, Saharia C, Chan W, Roberts A, Narang S, Blok I, Mical R, Norouzi M, Constant N (2023) Character-Aware Models Improve Visual Text Rendering. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 16270–16297

BibTeX

@inproceedings{liu-etal-2023-character,
    title = "Character-Aware Models Improve Visual Text Rendering",
    author = "Liu, Rosanne  and
      Garrette, Dan  and
      Saharia, Chitwan  and
      Chan, William  and
      Roberts, Adam  and
      Narang, Sharan  and
      Blok, Irina  and
      Mical, Rj  and
      Norouzi, Mohammad  and
      Constant, Noah",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.900/",
    doi = "10.18653/v1/2023.acl-long.900",
    pages = "16270--16297"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/