Built independently by an author, for readers. Read the story and support ChapterPal

keyword

high-resolution editing

High-resolution editing is the process of modifying localized regions or elements of an image while preserving, reconstructing, and matching fine visual details at full pixel resolution. In generative computer vision, this approach enables precise manipulations, such as text-guided inpainting or object modification, without causing blurriness, distortion, or quality degradation common in low-resolution processing. By conditioning generative architectures or multi-stage diffusion pipelines directly on the original full-resolution source imagery, high-resolution editing ensures that newly synthesized elements integrate smoothly with untouched surroundings, maintaining structural coherence, textural fidelity, and overall image sharpness.

1 item

Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting

Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting

Su Wang, Chitwan Saharia, Ceslee Montgomery, Jordi Pont-Tuset, Shai Noy, Stefano Pellegrini, Yasumasa Onoe, Sarah Laszlo, David J. Fleet, Radu Soricut, Jason Baldridge, Mohammad Norouzi, Peter Anderson, William Chan

OrganizationsGoogle

Why you should read this

Presents Imagen Editor, a high-resolution diffusion model that improves prompt-aligned image inpainting via object-detector masking during training, alongside EditBench, a systematic benchmark that reveals fine-grained strengths and limitations across leading text-guided editing models.

Text-guided image editing can have a transformative impact in supporting creative applications. A key challenge is to generate edits that are faithful to input text prompts, while consistent with input images. We present Imagen Editor, a cascaded diffusion model built, by fine-tuning Imagen [36] on text-guided image inpainting. Imagen Editor’s edits are faithful to the text prompts, which is accomplished by using object detectors to propose inpainting masks during training. In addition, Imagen Editor captures fine details in the input image by conditioning the cascaded pipeline on the original high resolution image. To improve qualitative and quantitative evaluation, we introduce EditBench, a systematic benchmark for text-guided image inpainting. EditBench evaluates inpainting edits on natural and generated images exploring objects, attributes, and scenes. Through extensive human evaluation on EditBench, we find that object-masking during training leads to across-the-board improvements in text-image alignment – such that Imagen Editor is preferred over DALL-E 2 [31] and Stable Diffusion [33] – and, as a cohort, these models are better at object-rendering than text-rendering, and handle material/color/size attributes better than count/shape attributes.

Added

2026-09-26