Imagic: Text-Based Real Image Editing with Diffusion Models

Bahjat KawarShiran ZadaOran LangOmer TovHuiwen ChangTali DekelInbar MosseriMichal Irani

article2022CVPR1,559 citations

Introduces Imagic, a method that uses pre-trained diffusion models to perform complex, non-rigid semantic edits on a single real image using only a text prompt without requiring masks or multi-view data.

Listen

Applying natural language text prompts to edit real photographs has become a central objective in computer vision, but existing approaches face significant practical barriers. Most current methods require auxiliary inputs such as user-drawn spatial masks, detailed descriptions of the original image, or multiple reference photographs of the target subject. Furthermore, existing techniques struggle with complex, non-rigid edits—such as changing an animal's posture or adjusting object interactions—often corrupting original image details or failing to alter the geometry convincingly.

The article demonstrates and evaluates a framework called Imagic, which performs sophisticated semantic edits on a single real image using only a target text prompt. The primary objective is to verify whether text-to-image diffusion models can execute complex, non-rigid modifications while preserving the unedited composition, background, and identity present in the original photograph.

The researchers designed a three-step optimization process: first, optimizing a text representation to capture the input image; second, fine-tuning the generative model parameters to reconstruct the photograph accurately; and third, linearly interpolating between the optimized representation and the target text prompt to generate the edited output. To establish credibility, the framework was validated across two diffusion architectures (Imagen and Stable Diffusion) using high-resolution natural photographs. The authors also established TEdBench, a dedicated benchmark of 100 complex non-rigid editing pairs, and conducted an Amazon Mechanical Turk user study collecting 9,213 human evaluations against three leading baselines (SDEdit, DDIB, and Text2LIVE).

The evaluation yielded several key findings. In the perceptual user study, human raters preferred Imagic over each alternative baseline with a preference rate exceeding 70%. Ablation analyses proved that fine-tuning the diffusion model is critical; without it, the model cannot maintain high fidelity to the original image when moving toward the target edit. Additionally, the optimal balance between image fidelity and text alignment consistently emerged within an interpolation intensity range of 0.6 to 0.8 across tested inputs, demonstrating a robust editing corridor.

These findings indicate that generative diffusion models possess latent compositional capabilities that enable precise, non-rigid editing without manual masking. For organizations evaluating generative image tools, this capability reduces workflow complexity by eliminating the need for multi-image datasets or specialized manual annotations. However, the runtime cost remains significant, requiring approximately seven to eight minutes of fine-tuning per image on high-end hardware, which limits immediate deployment in real-time, user-facing environments.

Decision-makers should treat this framework as a foundation for next-generation automated editing workflows, while prioritizing research into faster test-time optimization and automated selection of the interpolation parameter before enterprise-wide operational deployment. Stakeholders must remain cautious regarding standard model failure cases, including camera angle distortions, subtle under-editing, and inherited model biases such as poor human face generation, as well as the risk of misuse for synthetic misinformation.

arXiv: 2210.09276
Cover for Imagic: Text-Based Real Image Editing with Diffusion Models

Abstract

Text-conditioned image editing has recently attracted considerable interest. However, most methods are currently either limited to specific editing types (e.g., object overlay, style transfer), or apply to synthetically generated images, or require multiple input images of a common object. In this paper we demonstrate, for the very first time, the ability to apply complex (e.g., non-rigid) text-guided semantic edits to a single real image. For example, we can change the posture and composition of one or multiple objects inside an image, while preserving its original characteristics. Our method can make a standing dog sit down or jump, cause a bird to spread its wings, etc. -- each within its single high-resolution natural image provided by the user. Contrary to previous work, our proposed method requires only a single input image and a target text (the desired edit). It operates on real images, and does not require any additional inputs (such as image masks or additional views of the object). Our method, which we call "Imagic", leverages a pre-trained text-to-image diffusion model for this task. It produces a text embedding that aligns with both the input image and the target text, while fine-tuning the diffusion model to capture the image-specific appearance. We demonstrate the quality and versatility of our method on numerous inputs from various domains, showcasing a plethora of high quality complex semantic image edits, all within a single unified framework.

Citation

MLA
Kawar, B., et al. “Imagic: Text-Based Real Image Editing with Diffusion Models”. arXiv, 2022, http://arxiv.org/abs/2210.09276v3.
APA
Kawar, B., Zada, S., Lang, O., Tov, O., Chang, H., Dekel, T., Mosseri, I., & Irani, M. (2022). Imagic: Text-Based Real Image Editing with Diffusion Models. arXiv. http://arxiv.org/abs/2210.09276v3
Chicago
Kawar, B., S. Zada, O. Lang, et al. 2022. “Imagic: Text-Based Real Image Editing with Diffusion Models”. arXiv. http://arxiv.org/abs/2210.09276v3.
Harvard
Kawar, B. et al. (2022) “Imagic: Text-Based Real Image Editing with Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.09276v3.
Vancouver
1. Kawar B, Zada S, Lang O, Tov O, Chang H, Dekel T, Mosseri I, Irani M (2022) Imagic: Text-Based Real Image Editing with Diffusion Models. arXiv

BibTeX

@article{kawar2022imagic,
  title = {Imagic: Text-Based Real Image Editing with Diffusion Models},
  author = {Kawar, Bahjat and Zada, Shiran and Lang, Oran and Tov, Omer and Chang, Huiwen and Dekel, Tali and Mosseri, Inbar and Irani, Michal},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.09276v3},
  eprint = {2210.09276}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE