Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting

Su WangChitwan SahariaCeslee MontgomeryJordi Pont-TusetShai NoyStefano PellegriniYasumasa OnoeSarah LaszloDavid J. FleetRadu Soricut

article2023CVPR316 citations

Presents Imagen Editor, a high-resolution diffusion model that improves prompt-aligned image inpainting via object-detector masking during training, alongside EditBench, a systematic benchmark that reveals fine-grained strengths and limitations across leading text-guided editing models.

Listen

Text-to-image artificial intelligence models frequently struggle to satisfy exact creative and professional editing needs in a single attempt. Text-guided image inpainting addresses this by allowing users to modify designated regions of an existing image using text instructions. However, existing systems often produce edits that ignore text prompts or fail to blend seamlessly with the surrounding context, and the field lacks standardized benchmarks to assess fine-grained editing capabilities.

The article demonstrates a high-resolution text-guided editing system named Imagen Editor and introduces EditBench, a systematic evaluation benchmark designed to measure how accurately models render specific objects, attributes, and scenes.

To build the editing model, researchers adapted an existing diffusion architecture using specialized downsampling convolutions to preserve fine details at high resolutions and introduced an object-masking training policy that forces the system to rely on text prompts rather than background clues. For evaluation, the team assembled EditBench using 240 natural and synthetic images paired with diverse mask sizes and three prompt formats ranging from basic descriptions to multi-attribute specifications. Model performance was evaluated through 11,500 human assessments alongside automated text-image metrics, comparing the new approach against standard models including DALL-E 2 and Stable Diffusion.

The findings show that training with object-based masks improves alignment with text prompts across the board, with human evaluators preferring the object-masked model in 68% of side-by-side comparisons over its randomly masked equivalent. The proposed editor outperformed Stable Diffusion and DALL-E 2 in text-image alignment, winning 78% and 77% of human comparisons respectively while maintaining competitive visual realism. Across all evaluated systems, models handled object generation significantly better than text rendering, and successfully rendered material, color, and size attributes more reliably than count and shape specifications. In automated evaluations, text-to-image similarity metrics demonstrated the strongest alignment with human judgments, correctly predicting human preferences in 68% to 76% of image pairs.

These results indicate that training generative models to focus on coherent visual objects rather than arbitrary image areas substantially enhances user control and execution accuracy. While the technology promises to streamline creative workflows, the authors note serious operational risks, including the potential to generate realistic misinformation or harmful content. Responsible deployment requires strong mitigation strategies, such as automated watermarking, data deduplication, and safeguards preventing the unauthorized rendering of real individuals.

Organizations evaluating these tools should adopt focused automated metrics for iterative testing while maintaining human oversight for complex edits, particularly for tasks involving counting, geometric shapes, and text rendering where model accuracy declines noticeably. While current evaluations provide strong confidence in basic object and attribute edits, further work is necessary to improve performance on complex, multi-attribute prompts and abstract spatial relationships.

arXiv: 2212.06909
Cover for Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting

Abstract

Text-guided image editing can have a transformative impact in supporting creative applications. A key challenge is to generate edits that are faithful to input text prompts, while consistent with input images. We present Imagen Editor, a cascaded diffusion model built, by fine-tuning Imagen [36] on text-guided image inpainting. Imagen Editor’s edits are faithful to the text prompts, which is accomplished by using object detectors to propose inpainting masks during training. In addition, Imagen Editor captures fine details in the input image by conditioning the cascaded pipeline on the original high resolution image. To improve qualitative and quantitative evaluation, we introduce EditBench, a systematic benchmark for text-guided image inpainting. EditBench evaluates inpainting edits on natural and generated images exploring objects, attributes, and scenes. Through extensive human evaluation on EditBench, we find that object-masking during training leads to across-the-board improvements in text-image alignment – such that Imagen Editor is preferred over DALL-E 2 [31] and Stable Diffusion [33] – and, as a cohort, these models are better at object-rendering than text-rendering, and handle material/color/size attributes better than count/shape attributes.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Imagen Editor
  • 4. EditBench
  • 5. Evaluation
  • 5.1. Human Evaluation Protocol
  • 5.2. Human Evaluation Results
  • 5.3. Automatic Evaluation Metrics
  • 6. Societal Impact
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — Architecture and High-Resolution Conditioning of Imagen Editor

    model/method

    Imagen Editor is a cascaded text-to-image diffusion model fine-tuned from Imagen for text-guided image inpainting. The model takes three user inputs: an image to be edited I∈R1024×1024×3I \in \mathbb{R}^{1024 \times 1024 \times 3}, a binary edit mask M∈{0,1}1024×1024×1M \in \{0, 1\}^{1024 \times 1024 \times 1}, and a conditioning text prompt encoded with a pretrained T5-XXL language model.

    The architecture uses a three-stage cascade consisting of a 64×6464 \times 64 base diffusion model, a 64×64→256×25664 \times 64 \to 256 \times 256 super-resolution model, and a 256×256→1024×1024256 \times 256 \to 1024 \times 1024 super-resolution model. In all three stages, the full-resolution 1024×10241024 \times 1024 image and mask are concatenated along the channel dimension with the diffusion latents. For the lower-resolution stages (64×6464 \times 64 and 256×256256 \times 256), the conditioning image and mask are downsampled via learned parameterized strided convolutions rather than parameter-free operations (such as bicubic downsampling), which prevents visual boundary artifacts along the mask edges in the final output. The weights of the new conditioning input channels are initialized to zero so the network initially matches the pretrained text-to-image base model.

    During inference with the base 64×6464 \times 64 model, Classifier-Free Guidance (CFG) is applied using an oscillating guidance weight schedule that alternates between 1 and 30 to balance sample realism and text-image alignment.

  2. Knowl 2 — Object Detector Masking Policy for Inpainting Model Training

    model/method

    Standard inpainting training commonly applies random box or random stroke masks. In text-guided inpainting, randomly placed masks often cover background areas or only partially intersect objects, creating regions that can be plausibly reconstructed using only the surrounding visual context. This allows the network to ignore the conditioning text prompt during training, resulting in poor text-image alignment at inference time.

    The object detector masking policy replaces random masking during training by using an off-the-shelf object detector (specifically a lightweight Single Shot MultiBox Detector with MobileNetV2, SSD MobileNet v2) to detect and localize semantic objects in the training images on the fly. The bounding boxes of detected objects are used to generate inpainting training masks that completely conceal whole objects. Masking entire objects prevents the model from relying solely on spatial context to fill in the missing region, forcing the diffusion network to attend directly to the text prompt to reconstruct the masked content.

  3. Knowl 3 — EditBench Benchmark Dataset and Taxonomy

    experimental setup

    EditBench is an evaluation benchmark for text-guided image inpainting consisting of 240 evaluation images, split evenly into 120 natural images (drawn from Visual Genome and Open Images) and 120 synthetic images (generated by Imagen and Parti). For each image, human annotators create a free-form binary mask that fully encloses a target object without tightly segmenting it (avoiding mask shapes that leak the underlying object's identity), spanning a diverse distribution of mask-to-image area ratios from small local edits to large edge-contacting uncropping tasks.

    Each image-mask pair is annotated with three distinct prompt types:

    1. Mask-Simple: A prompt describing the masked area using a single object and a single attribute (e.g., a metal cat).
    2. Mask-Rich: A detailed prompt describing the masked area with multiple attributes and object parts to probe compositional attribute binding (e.g., a metal cat sitting and with its tail wrapped around its body).
    3. Full: A prompt describing the entire composite scene disregarding mask boundaries (e.g., a metal cat sitting in the middle of a farm field).

    The dataset systematically samples combinations across three structured axes:

    • Attributes: material, color, shape, size, count.
    • Objects: common, rare, text-rendering.
    • Scenes: indoor, outdoor, realistic, painting.
  4. Knowl 4 — Fine-Grained Attribute-Binding Human Evaluation Protocol for Inpainting

    experimental setup

    EditBench utilizes a fine-grained, localized human evaluation protocol to assess text-guided inpainting faithfulness:

    • In all evaluations, a visual bounding box (e.g., a red border) highlights the edited region so annotators evaluate localized inpainting rather than full-image generation.
    • For Full prompts, annotators give a single binary judgment answering "Does the image match the caption?".
    • For Mask-Simple prompts (specifying an object OO and an attribute AA), annotators answer three separate binary questions: (1) Is object OO rendered? (2) Is attribute AA present? (3) Is attribute AA correctly bound to object OO?
    • For Mask-Rich prompts (specifying three attribute-object pairs), annotators evaluate three sets of the three binary questions (totaling 9 binary judgments per edited image).

    Under this protocol, a generated sample for Mask-Simple or Mask-Rich is scored as correct if and only if every specified object, attribute, and attribute binding is affirmed positively. Side-by-side evaluations also measure pairwise user preferences for text alignment and visual quality across models.

  5. Knowl 5 — Human Evaluation Benchmarks on EditBench

    empirical result

    Across 11,500 single-model evaluation tasks and forced-choice side-by-side evaluations on EditBench comparing Imagen Editor (IM), Imagen Editor trained with random masking (IMRM\text{IM}_{\text{RM}}), Stable Diffusion v1.5 (SD), and DALL-E 2 (DL2):

    • Single-Image Correct Alignment Rate:
      • Full Prompts: IM achieves 70%, outperforming DL2 (57%), IMRM\text{IM}_{\text{RM}} (53%), and SD (46%).
      • Mask-Simple Prompts: IM achieves 63%, outperforming DL2 (51%), SD (47%), and IMRM\text{IM}_{\text{RM}} (43%).
      • Mask-Rich Prompts: IM achieves 37%, outperforming IMRM\text{IM}_{\text{RM}} (27%), DL2 (21%), and SD (14%).
    • Side-by-Side User Preference on Mask-Rich Prompts: Human annotators prefer Imagen Editor's text alignment over SD in 78% of comparisons, over DL2 in 77% of comparisons, and over IMRM\text{IM}_{\text{RM}} in 68% of comparisons. In visual realism, preference deltas between IM and the other models remain within 0% to 6%, indicating that alignment gains do not compromise image quality.
  6. Knowl 6 — Model Inpainting Strengths Across Object and Attribute Types

    empirical result

    Human evaluation breakdowns on EditBench Mask-Simple prompts across object and attribute categories reveal specific capabilities and failure modes in text-guided inpainting models:

    • Object Categories: All evaluated models perform substantially better at object rendering than text rendering. Across common objects, rare objects, and text rendering, Imagen Editor achieves 69%, 63%, and 51% correct alignment, leading the second-best model by 10%, 11%, and 11% respectively. Stable Diffusion shows a steep drop on text rendering (26%) compared to common objects (59%) and rare objects (44%).
    • Attribute Categories: Diffusion models handle physical and perceptual attributes better than abstract or relational attributes. Imagen Editor scores 73% on material, 70% on color, 68% on size, 55% on shape, and 47% on count. Imagen Editor leads the next-best model by 13%–16% on material, color, size, and shape; on count attributes, DALL-E 2 (46%) is only 1% behind Imagen Editor (47%).
  7. Knowl 7 — Agreement Between Automated Evaluation Metrics and Human Judgments

    data/table

    Automated text-to-image (T2I) CLIPScore similarity exhibits the highest alignment with human ratings when selecting the best edited image between pairs of model outputs (evaluated on 10,000 pairs) and when identifying the highest-performing model among 4 hybrid model ensembles (evaluated over 100,000 trials). Evaluating T2I CLIPScore on cropped masked regions achieves the highest agreement for mask-scoped prompts, while full-image evaluation achieves the highest agreement for full-scene prompts.

    Prompt Image Region T2I I2I T2I+I2I R-Prec Random
    Pairwise Best-Image Selection Agreement (%)
    Full Full 70.1 58.6 66.8 53.3 50.0
    Full Crop 68.1 55.8 62.4 57.7 50.0
    Mask-Simple Full 73.8 53.1 63.2 72.0 50.0
    Mask-Simple Crop 76.0 55.3 66.4 71.0 50.0
    Mask-Rich Full 66.7 55.2 63.4 62.3 50.0
    Mask-Rich Crop 68.4 56.4 64.1 63.3 50.0
    4-Model Best-Model Selection Agreement (%)
    Full Full 38.5 30.8 35.7 28.7 25.0
    Full Crop 36.7 28.2 32.4 27.9 25.0
    Mask-Simple Full 45.7 28.2 37.1 40.2 25.0
    Mask-Simple Crop 47.1 29.4 38.5 39.7 25.0
    Mask-Rich Full 45.8 31.4 40.5 39.1 25.0
    Mask-Rich Crop 48.1 30.9 40.9 39.3 25.0
  8. Knowl 8 — Aggregated Automated Benchmark Evaluation Scores on EditBench

    data/table

    Comparison of automated metrics aggregated over all EditBench prompts across Stable Diffusion (SD), DALL-E 2 (DL2), Imagen Editor with random masking (IMRM\text{IM}_{\text{RM}}), full Imagen Editor (IM), and ground-truth reference images (Ref). Automated metrics include CLIPScore Text-to-Image (T2I), Image-to-Image (I2I), combined T2I+I2I, CLIP-R-Precision (CLIP-R-Prec), and NIMA perceptual quality.

    Metric SD DL2 IMRM_{\text{RM}} IM Ref.
    CLIPScore T2I (↑\uparrow) 29.7 29.1 29.6 31.5 31.0
    CLIPScore I2I (↑\uparrow) 74.9 76.1 75.8 76.6 -
    CLIPScore T2I+I2I (↑\uparrow) 52.3 52.6 53.1 53.6 -
    CLIP-R-Prec (↑\uparrow) 96.5 95.3 95.0 98.6 99.3
    NIMA (↑\uparrow) 4.44 4.33 4.56 4.63 4.89

    Imagen Editor outperforms all evaluated baseline generative models across every automated text alignment and image quality metric. Reference images attain the highest CLIP-R-Precision and NIMA perceptual score.

Coverage note — None was omitted; all primary modeling components (Imagen Editor architecture, object-masking training policy, CFG oscillation), benchmark specifications (EditBench dataset design, human evaluation protocols), empirical human evaluation results, and automated metric validations are fully covered.

References

  1. 1.Omri Avrahami, Dani Lischinski, and Ohad Fried. Blended diffusion for text-driven editing of natural images. Proceedings of CVPR, abs/2111.14818, 2022. 2, 3, 5, 12
  2. 2.Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. Text2live: Text-driven layered image and video editing, 2022. 2
  3. 3.David Bau, Alex Andonian, Audrey Cui, YeonHwan Park, Ali Jahanian, Aude Oliva, and Antonio Torralba. Paint by word. CoRR, abs/2103.10951, 2021. 2
  4. 4.Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. Multimodal datasets: misogyny, pornography, and malignant stereotypes. In arXiv:2110.01963, 2021. 8
  5. 5.Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. Diffedit: Diffusion-based semantic image editing with mask guidance, 2022. 1, 2
  6. 6.Valentin De Bortoli, James Thornton, Jeremy Heng, and Arnaud Doucet. Diffusion schrödinger bridge with applications to score-based generative modeling. Advances in Neural Information Processing Systems, 34, 2021. 3
  7. 7.Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. Cogview2: Faster and better text-to-image generation via hierarchical transformers. arXiv preprint arXiv:2204.14217, 2022. 2, 3
  8. 8.Sara Dolnicar, Bettina Grün, and Friedrich Leisch. Quick, simple and reliable: forced binary survey questions. International Journal of Market Research, 2011. 12
  9. 9.Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. Make-a-scene: Scene-based text-to-image generation with human priors. arXiv preprint arXiv:2203.13131, 2022. 4
  10. 10.Ying Gao and Qing Zhu. Text-guided image inpainting. In Proceedings of IEEE, 2022. 3
  11. 11.Yvette Graham and Qun Liu. Achieving accurate conclusions in evaluation of automatic machine translation metrics. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2016. 8
  12. 12.Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber. On the binding problem in artificial neural networks. CoRR, abs/2012.05208, 2020. 4, 5
  13. 13.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-Prompt Image Editing with Cross Attention Control. In arXiv preprint arXiv:2208.01626, 2022. 1, 2, 3, 6
  14. 14.Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718, 2021. 2, 3, 7
  15. 15.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 4
  16. 16.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 4
  17. 17.Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models, 2022. 1, 2, 3
  18. 18.Gwanghyun Kim and Jong Chul Ye. Diffusionclip: Text-guided image manipulation using diffusion models. arXiv preprint arXiv:2110.02711, 2021. 2, 3, 12
  19. 19.Yoon Kim and Alexander M. Rush. Sequence-Level Knowledge Distillation. In EMNLP, 2016. 2
  20. 20.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. International Journal of Computer Vision, 2017. 4
  21. 21.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. IJCV, 2020. 4
  22. 22.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. CoRR, abs/2107.06499, 2021. 8
  23. 23.Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, and Philip H. S. Torr. Manigan: Text-guided image manipulation, 2020. 3
  24. 24.Xiyang Luo, Michael Goebel, Elnaz Barshan, and Feng Yang. Leca: A learned approach for efficient cover-agnostic watermarking, 2022. 8
  25. 25.Naila Murray, Luca Marchesotti, and Florent Perronnin. Ava: A large-scale database for aesthetic visual analysis. In Proceedings of CVPR, 2012. 3
  26. 26.Zhixiong Nan, Yang Liu, Nanning Zheng, and Song-Chun Zhu. Recognizing unseen attribute-object pair with generative model. Proceedings of AAAI, 2019. 7
  27. 27.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Bob McGrew Pamela Mishkin, Ilya Sutskever, and Mark Chen. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. In arXiv:2112.10741, 2021. 1, 2, 3, 4, 5, 12
  28. 28.Dong Huk Park, Samaneh Azadi, Xihui Liu, Trevor Darrell, and Anna Rohrbach. Benchmark for compositional text-to-image synthesis. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), 2021. 3, 7
  29. 29.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021. 7
  30. 30.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. CoRR, abs/2103.00020, 2021. 3
  31. 31.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical Text-Conditional Image Generation with CLIP Latents. In arXiv:2204.06125, 2022. 1, 2, 5, 6, 8
  32. 32.Walber Rodrigues, Felipe Walmsley, George Cavalcanti, Jonysberg Quintino, and Helder Pinho. Grave artifacts in image inpainting: Investigating the causes and untangling the factors, 2021. 4
  33. 33.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models. In CVPR, 2022. 1, 2, 5, 6
  34. 34.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. 2022. 1
  35. 35.Chitwan Saharia, William Chan, Huiwen Chang, Chris A. Lee, Jonathan Ho, Tim Salimans, David J. Fleet, and Mohammad Norouzi. Palette: Image-to-Image Diffusion Models. In arXiv:2111.05826, 2021. 2, 3, 13
  36. 36.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. In NeurIPS, 2022. 1, 3, 4, 8
  37. 37.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding, 2022. 6
  38. 38.Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J Fleet, and Mohammad Norouzi. Image Super-Resolution via Iterative Refinement. IEEE PAMI, 2022. 3
  39. 39.Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Inverted residuals and linear bottlenecks: Mobile networks for classification, detection and segmentation. CoRR, abs/1801.04381, 2018. 3
  40. 40.Lucia Specia, Frédéric Blain, Marina Fomicheva, Chrysoula Zerva, Zhenhao Li, Vishrav Chaudhary, and André F. T. Martins. Findings of the WMT 2021 shared task on quality estimation. In Proceedings of the Sixth Conference on Machine Translation, 2021. 7
  41. 41.Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. Resolution-robust large mask inpainting with fourier convolutions. arXiv preprint arXiv:2109.07161, 2021. 2
  42. 42.Hossein Talebi, Ehsan Amid, Peyman Milanfar, and Manfred K. Warmuth. Rank-smoothed pairwise learning in perceptual quality assessment. In Proceedings of IEEE, 2021. 8
  43. 43.Yi Chern Tan and L. Elisa Celis. Assessing Social and Intersectional Biases in Contextualized Word Representations. In NeurIPS, 2019. 8
  44. 44.Ming Tao, Bing-Kun Bao, Hao Tang, Fei Wu, Longhui Wei, and Qi Tian. De-net: Dynamic text-guided image editing adversarial networks, 2022. 3
  45. 45.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Kathleen S. Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed H. Chi, and Quoc Le. Lamda: Language models for dialog applications. CoRR, abs/2201.08239, 2022. 8
  46. 46.Dani Valevski, Matan Kalman, Yossi Matias, and Yaniv Leviathan. Unitune: Text-driven image editing by fine tuning an image generation model on a single image, 2022. 1, 2
  47. 47.Jianan Wang, Guansong Lu, Hang Xu, Zhenguo Li, Chunjing Xu, and Yanwei Fu. Manitrans: Entity-level text-guided image manipulation via token-wise semantic alignment and generation, 2022. 3
  48. 48.Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. Generative image inpainting with contextual attention. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5505–5514, 2018. 3
  49. 49.Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. Free-form image inpainting with gated convolution. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4471–4480, 2019. 3, 13
  50. 50.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling Autoregressive Models for Content-Rich Text-to-Image Generation. In arXiv:2206.10789, 2022. 1, 4, 5, 8
  51. 51.Han Zhang, Weichong Yin, Yewei Fang, Lanxin Li, Boqiang Duan, Zhihua Wu, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. Ernie-vilg: Unified generative pre-training for bidirectional vision-language generation. CoRR, abs/2112.15283, 2021. 1
  52. 52.Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. In ICLR, 2020. 8
  53. 53.Yufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Chris Tensmeyer, Tong Yu, Changyou Chen, Jinhui Xu, and Tong Sun. Interactive image generation with natural-language feedback. In Proceedings of AAAI, 2022. 3, 5, 12

Citation

MLA
Wang, S., et al. “Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting”. arXiv, 2022, http://arxiv.org/abs/2212.06909v2.
APA
Wang, S., Saharia, C., Montgomery, C., Pont-Tuset, J., Noy, S., Pellegrini, S., Onoe, Y., Laszlo, S., Fleet, D. J., Soricut, R., Baldridge, J., Norouzi, M., Anderson, P., & Chan, W. (2022). Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting. arXiv. http://arxiv.org/abs/2212.06909v2
Chicago
Wang, S., C. Saharia, C. Montgomery, et al. 2022. “Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting”. arXiv. http://arxiv.org/abs/2212.06909v2.
Harvard
Wang, S. et al. (2022) “Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.06909v2.
Vancouver
1. Wang S, Saharia C, Montgomery C, et al (2022) Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting. arXiv

BibTeX

@article{wang2022imagen,
  title = {Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting},
  author = {Wang, Su and Saharia, Chitwan and Montgomery, Ceslee and Pont-Tuset, Jordi and Noy, Shai and Pellegrini, Stefano and Onoe, Yasumasa and Laszlo, Sarah and Fleet, David J. and Soricut, Radu and Baldridge, Jason and Norouzi, Mohammad and Anderson, Peter and Chan, William},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.06909v2},
  eprint = {2212.06909}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE