Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting
Su WangChitwan SahariaCeslee MontgomeryJordi Pont-TusetShai NoyStefano PellegriniYasumasa OnoeSarah LaszloDavid J. FleetRadu Soricut
Presents Imagen Editor, a high-resolution diffusion model that improves prompt-aligned image inpainting via object-detector masking during training, alongside EditBench, a systematic benchmark that reveals fine-grained strengths and limitations across leading text-guided editing models.
Text-to-image artificial intelligence models frequently struggle to satisfy exact creative and professional editing needs in a single attempt. Text-guided image inpainting addresses this by allowing users to modify designated regions of an existing image using text instructions. However, existing systems often produce edits that ignore text prompts or fail to blend seamlessly with the surrounding context, and the field lacks standardized benchmarks to assess fine-grained editing capabilities.
The article demonstrates a high-resolution text-guided editing system named Imagen Editor and introduces EditBench, a systematic evaluation benchmark designed to measure how accurately models render specific objects, attributes, and scenes.
To build the editing model, researchers adapted an existing diffusion architecture using specialized downsampling convolutions to preserve fine details at high resolutions and introduced an object-masking training policy that forces the system to rely on text prompts rather than background clues. For evaluation, the team assembled EditBench using 240 natural and synthetic images paired with diverse mask sizes and three prompt formats ranging from basic descriptions to multi-attribute specifications. Model performance was evaluated through 11,500 human assessments alongside automated text-image metrics, comparing the new approach against standard models including DALL-E 2 and Stable Diffusion.
The findings show that training with object-based masks improves alignment with text prompts across the board, with human evaluators preferring the object-masked model in 68% of side-by-side comparisons over its randomly masked equivalent. The proposed editor outperformed Stable Diffusion and DALL-E 2 in text-image alignment, winning 78% and 77% of human comparisons respectively while maintaining competitive visual realism. Across all evaluated systems, models handled object generation significantly better than text rendering, and successfully rendered material, color, and size attributes more reliably than count and shape specifications. In automated evaluations, text-to-image similarity metrics demonstrated the strongest alignment with human judgments, correctly predicting human preferences in 68% to 76% of image pairs.
These results indicate that training generative models to focus on coherent visual objects rather than arbitrary image areas substantially enhances user control and execution accuracy. While the technology promises to streamline creative workflows, the authors note serious operational risks, including the potential to generate realistic misinformation or harmful content. Responsible deployment requires strong mitigation strategies, such as automated watermarking, data deduplication, and safeguards preventing the unauthorized rendering of real individuals.
Organizations evaluating these tools should adopt focused automated metrics for iterative testing while maintaining human oversight for complex edits, particularly for tasks involving counting, geometric shapes, and text rendering where model accuracy declines noticeably. While current evaluations provide strong confidence in basic object and attribute edits, further work is necessary to improve performance on complex, multi-attribute prompts and abstract spatial relationships.
- Paper: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding, Chitwan Saharia et al. (2022). Introduces the base Imagen text-to-image diffusion architecture and T5-driven conditioning pipeline that Imagen Editor directly adapts and builds upon for localized editing.
- Paper: Blended Diffusion for Text-driven Editing of Natural Images, Omri Avrahami et al. (2021). Pioneers region-based, text-guided image editing using diffusion models with binary spatial masks, establishing the core problem formulation explored in the paper.
- Paper: GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models, Alexander Quinn Nichol et al. (2022). Demonstrates classifier-free guidance and text-conditioned inpainting within diffusion models, providing foundational techniques used in Imagen Editor's design.
- Paper: RePaint: Inpainting using Denoising Diffusion Probabilistic Models, Andreas Lugmayr et al. (2022). Establishes standard diffusion-based inpainting mechanisms across arbitrary mask shapes, serving as key baseline methodology for masked image completion.
- Paper: Prompt-to-Prompt Image Editing with Cross Attention Control, Amir Hertz et al. (2022). Explores text-driven image editing via cross-attention manipulation in diffusion models, establishing key principles of text-to-image semantic control.
- Paper: SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Chenlin Meng et al. (2022). Formulates guided image editing and synthesis via stochastic differential equations, laying early ground for diffusion-based image modification pipelines.
- Paper: Imagic: Text-Based Real Image Editing with Diffusion Models, Bahjat Kawar et al. (2022). Demonstrates text-guided semantic editing using Imagen and Stable Diffusion backbones, presenting early non-rigid editing formulations that motivate structured benchmarking.
- Paper: Free-Form Image Inpainting With Gated Convolution, Jiahui Yu et al. (2018). Introduces modern free-form inpainting principles and gated convolutions that influenced subsequent deep generative image completion architectures.
- Paper: SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models, Yuzhou Huang et al. (2024). Extends text-guided editing beyond standard descriptive masks by integrating multimodal large language models to reason over complex, multi-object editing instructions.
- Paper: DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing, Chong Mou et al. (2024). Builds on localized diffusion editing by refining sampling and gradient guidance strictly inside target zones to balance precise manipulation and generative flexibility.
- Paper: InstructPix2Pix: Learning to Follow Image Editing Instructions, Tim Brooks et al. (2023). Generalizes text-driven image modification by training conditional diffusion models to follow natural language instructions directly without requiring explicit spatial masks.
- Paper: T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation, Kaiyi Huang et al. (2023). Expands the evaluation of fine-grained compositional challenges highlighted in EditBench by establishing standardized compositional benchmarks and visual question-answering metrics.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). Advances evaluation methodology for conditional synthesis and editing by using multimodal LLMs to generate explainable, attribute-level rationales that correlate with human judgment.
- Paper: Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models, Guanhua Zhang et al. (2023). Addresses the boundary blending and context coherence challenges of diffusion inpainting by formulating a step-wise joint adjustment algorithm for revealed and unrevealed regions.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). Unifies localized editing, customization, and generation tasks into a single transformer framework trained on real-world dynamics and paired frame transitions.
- Paper: Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing, Bingyan Liu et al. (2024). Provides an in-depth mechanistic analysis of cross- and self-attention dynamics during text-guided image editing, illuminating the internal representations that govern edit accuracy.
