SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models

Yuzhou HuangLiangbin XieXintao WangZiyang YuanXiaodong CunYixiao GeJiantao ZhouChao DongRui HuangRuimao Zhang

article2024CVPR205 citations

Develops SmartEdit, a framework integrating multimodal large language models with diffusion models via a bidirectional interaction module and targeted perception training to execute image editing instructions requiring multi-object reasoning and world knowledge.

Listen

Modern digital workflows increasingly rely on text-guided image editing, yet existing tools struggle when presented with nuanced user commands. Current systems frequently fail when tasks require distinguishing specific targets among multiple objects based on spatial location, color, or context, as well as when world knowledge and reasoning are needed to infer the target (such as identifying "the tool used to cut cakes"). These failures stem from conventional models relying on basic text encoders that cannot reason or effectively fuse image details with complex instructions.

The article demonstrates an advanced image editing framework, termed SmartEdit, that integrates multimodal large language models—artificial intelligence systems that simultaneously understand language and images—into image diffusion generation to handle complex understanding and reasoning instructions. The approach pairs the multimodal language model with a novel Bidirectional Interaction Module to enable comprehensive two-way communication between text instructions and image features, while incorporating segmentation data and a targeted synthetic dataset of 476 complex image-instruction pairs during training. The authors also establish a new benchmark called Reason-Edit, comprising 219 image-text evaluation pairs, to rigorously assess performance.

The findings show that SmartEdit substantially outperforms leading baseline methods across both complex understanding and reasoning tasks. In reasoning scenarios, SmartEdit models achieved human-evaluated instruction alignment scores of approximately 79% to 82%, compared to baseline scores ranging between 28% and 48%. Ablation experiments confirmed that removing the Bidirectional Interaction Module or reverting to one-way information sharing markedly degraded output quality. Furthermore, training ablation showed that pairing perception data with a small set of high-quality complex examples was critical: instruction alignment rose from roughly 20–23% with standard editing data up to 71–79% when the full dataset strategy was employed.

These results demonstrate that complex, reasoning-based image editing can be achieved without the prohibitive expense of generating massive specialized datasets. By combining targeted perceptual pre-training, lightweight fine-tuning, and bidirectional feature interaction, organizations can deploy more capable generative vision systems with higher accuracy, reduced operational failure rates, and greater alignment with human intent.

Decision-makers should consider adopting multimodal reasoning architectures and bidirectional fusion modules when designing next-generation image generation and editing pipelines. Future development should focus on expanding complex evaluation benchmarks, optimizing computational efficiency during inference, and scaling the synthetic data pipeline to cover broader operational domains.

Cover for SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models

Abstract

Current instruction-based image editing methods, such as InstructPix2Pix, often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this, this paper introduces SmartEdit, a novel approach of instruction-based image editing that leverages Multimodal Large Language Models (MLLMs) to enhance its understanding and reasoning capabilities. However, direct integration of these elements still faces challenges in situations requiring complex reasoning. To mitigate this, we propose a Bidirectional Interaction Module (BIM) that enables comprehensive bidirectional information interactions between the input image and the MLLM output. During training, we initially incorporate perception data to boost the perception and understanding capabilities of diffusion models. Subsequently, we demonstrate that a small amount of complex instruction editing data can effectively stimulate SmartEdit's editing capabilities for more complex instructions. We further construct a new evaluation dataset, Reason-Edit, specifically tailored for complex instruction-based image editing. Both quantitative and qualitative results on this evaluation dataset indicate that our SmartEdit surpasses previous methods, paving the way for the practical application of complex instruction-based image editing.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Image Editing with Diffusion Models.
  • 2.2. LLM with Diffusion Models
  • 3. Preliminary
  • 4. Method
  • 4.1. The Framework of SmartEdit
  • 4.2. Bidirectional Interaction Module
  • 4.3. Dataset Utilization Strategy
  • 4.4. Reason-Edit for Better Evaluation
  • 5. Experiments
  • 5.1. Experimental Setting
  • 5.2. Comparison with State-of-the-Art Methods
  • 5.3. Ablation Study on BIM
  • 5.4. Ablation Study on Dataset Usage
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — SmartEdit Multimodal Image Editing Framework

    model/method

    SmartEdit is an instruction-based image editing architecture that integrates Multimodal Large Language Models (MLLMs) with latent diffusion models to support complex understanding (such as spatial positioning, color, relative size, and reflections) and complex reasoning (requiring world knowledge to determine what to edit).

    Given an input image xx and an instruction text tokenized as c=(s1,…,sT)c = (s_1, \dots, s_T), the framework executes the following components:

    1. MLLM Hidden State Extraction: Image xx is encoded by a visual encoder and a fully connected layer into vμ(x)v_\mu(x). The vocabulary of the MLLM (such as LLaVA-1.1-7B or LLaVA-1.1-13B) is expanded by appending rr trainable image tokens [IMG1],…,[IMGr][\text{IMG}_1], \dots, [\text{IMG}_r] with an embedding matrix EE to the end of instruction cc. The majority of MLLM weights θ\theta are frozen while fine-tuning via Low-Rank Adaptation (LoRA with rank dim=16\text{dim}=16 and α=27\alpha=27). The hidden states hh corresponding to the rr [IMG][\text{IMG}] tokens are extracted from the LLM output.

    2. Feature Alignment via Q-Former: Because the hidden states hh reside in the LLM representation space rather than the diffusion text condition space, a 6-layer transformer Q-Former QβQ_\beta with n=77n = 77 learnable query tokens maps hh into an aligned text feature representation f=Qβ(h)f = Q_\beta(h).

    3. Bidirectional Interaction Module (BIM): An image feature v=Eϕ(x)v = E_\phi(x) extracted by visual encoder EϕE_\phi interacts bidirectionally with ff in the BIM, producing updated text condition features f′f' and spatial image features v′v'.

    4. Diffusion UNet Conditioning: In the latent diffusion UNet ϵδ\epsilon_\delta, f′f' is supplied as the key and value for cross-attention layers. The spatial feature v′v' is added in a residual manner to the concatenated latent representations concat[zt,E(x)]\text{concat}[z_t, \mathcal{E}(x)], where ztz_t is the noisy latent at diffusion step tt and E(x)\mathcal{E}(x) is the VAE latent encoding of the input image xx.

  2. Knowl 2 — Bidirectional Interaction Module Architecture

    model/method

    The Bidirectional Interaction Module (BIM) performs bidirectional feature exchange between the text feature representation ff output by the Q-Former and the image feature representation vv extracted by the image encoder Eϕ(x)E_\phi(x).

    The module comprises a self-attention block, two distinct cross-attention blocks, and a pointwise Multi-Layer Perceptron (MLP):

    1. Self-Attention on Text Feature: The text feature ff first passes through a self-attention mechanism to model intra-sequence relationships among text tokens.

    2. Text-Conditioned Cross-Attention: The self-attended text feature serves as the query (Q) while the visual feature vv acts as both key (K) and value (V) in a cross-attention block. The resulting cross-attended feature is passed through a pointwise MLP to produce the refined text condition feature f′f'.

    3. Image-Conditioned Cross-Attention: The visual feature vv serves as the query (Q) while the newly generated text feature f′f' acts as both key (K) and value (V) in a second cross-attention block. This interaction generates the refined visual feature v′v'.

    Output f′f' is forwarded to serve as the key and value sequences in the cross-attention blocks of the diffusion UNet, while v′v' is residually combined with the input image and noisy latent concatenation concat[zt,E(x)]+v′\text{concat}[z_t, \mathcal{E}(x)] + v' prior to entering the UNet.

  3. Knowl 3 — SmartEdit Training Objectives

    equation

    SmartEdit is trained end-to-end using a joint objective composed of an autoregressive language modeling loss on the expanded image tokens and a latent diffusion denoising loss.

    The autoregressive MLLM loss minimizes the negative log-likelihood of predicting the rr special visual tokens [IMG1],…,[IMGr][\text{IMG}_1], \dots, [\text{IMG}_r] conditioned on the visual prompt vμ(x)v_\mu(x) and instruction tokens s1,…,sTs_1, \dots, s_T:

    LLLM(c)=−∑i=1rlog⁡p{θ∪E}([IMGi]∣vμ(x),s1,…,sT,[IMG1],…,[IMGi−1])L_{\text{LLM}}(c) = -\sum_{i=1}^r \log p_{\{\theta \cup E\}}([\text{IMG}_i] \mid v_\mu(x), s_1, \dots, s_T, [\text{IMG}_1], \dots, [\text{IMG}_{i-1}])

    where θ\theta represents the frozen MLLM parameters with LoRA adapters, EE is the trainable embedding matrix for the special [IMG][\text{IMG}] tokens, and vμ(x)v_\mu(x) is the image feature projection.

    The conditional latent diffusion loss trains the UNet noise predictor ϵδ\epsilon_\delta to estimate noise ϵ∼N(0,1)\epsilon \sim \mathcal{N}(0, 1) added to target latent zt=E(y)tz_t = \mathcal{E}(y)_t at timestep tt:

    Ldiffusion=EE(y),E(x),c,ϵ∼N(0,1),t[∥ϵ−ϵδ(t,concat[zt,E(x)]+v′,f′)∥22]L_{\text{diffusion}} = \mathbb{E}_{\mathcal{E}(y), \mathcal{E}(x), c, \epsilon \sim \mathcal{N}(0, 1), t}\left[ \left\| \epsilon - \epsilon_\delta\left(t, \text{concat}[z_t, \mathcal{E}(x)] + v', f'\right) \right\|_2^2 \right]

    where yy is the target edited image, xx is the original input image, E\mathcal{E} is the latent VAE encoder, ztz_t is the noisy target latent at sampling timestep tt, f′f' is the text conditioning feature output by the Bidirectional Interaction Module (BIM), and v′v' is the residual visual feature output by BIM.

  4. Knowl 4 — Multi-Task Dataset Composition Strategy for SmartEdit

    model/method

    Training SmartEdit solely on conventional instruction editing datasets results in poor spatial grounding and limited reasoning capability. To resolve this, training data is partitioned into four complementary categories:

    1. Perception and Segmentation Datasets: Includes COCOStuff, RefCOCO, GRefCOCO, and the reasoning segmentation dataset from LISA. These datasets inject spatial localization and conceptual grounding capabilities directly into the diffusion UNet.

    2. Conventional Instruction Editing Datasets: Paired image-instruction editing datasets InstructPix2Pix and MagicBrush to provide core image transformation priors.

    3. Visual Question Answering (VQA) Dataset: LLaVA-Instruct-150k, preserving the foundational reasoning and conversational capabilities of the MLLM backbone.

    4. Synthetic Complex Editing Dataset: A curated set of 476 high-quality synthesized triplets (source image, complex instruction, target edited image). This dataset covers two categories:

      • Complex understanding: Multiple objects distinguished by location, color, relative size, or placement inside versus outside a mirror.
      • Complex reasoning: Scenarios requiring world knowledge to deduce which object should be altered or removed (e.g., "remove the object that is used to cut fruits").
  5. Knowl 5 — Reason-Edit Benchmark and Evaluation Protocol

    experimental setup

    Reason-Edit is an evaluation benchmark containing 219 image-instruction pairs constructed to evaluate instruction-based image editing models under complex understanding and reasoning scenarios, with zero sample overlap with the training data.

    The benchmark measures performance along two axes:

    • Complex Understanding Scenarios: Instructions targeting specific objects within multi-object scenes via attributes such as spatial location, relative scale, color, or mirror reflections.
    • Complex Reasoning Scenarios: Instructions requiring external world knowledge to resolve the target object (e.g., identifying a utensil by its functional utility).

    Evaluation relies on five complementary metrics:

    • PSNR, SSIM, and LPIPS (Background): Computed exclusively over the unedited background region to quantify preservation of non-target content (higher PSNR/SSIM is better; lower LPIPS is better).
    • CLIP Score (Foreground): ViT-L/14 CLIP cosine similarity between the edited foreground region and the ground-truth text label (higher is better).
    • Instruction-Alignment (Ins-align): Human evaluation metric where four independent human annotators judge whether the resulting edited image accurately follows the instruction, averaged across all annotators to produce an accuracy score in [0,1][0, 1] (higher is better).
  6. Knowl 6 — Quantitative Comparison on the Reason-Edit Benchmark

    data/table

    SmartEdit (7B and 13B) was evaluated against SOTA instruction editing models (InstructPix2Pix, MagicBrush, and InstructDiffusion) on the Reason-Edit benchmark. To ensure fair comparison, all baseline methods were fine-tuned on the identical multi-task dataset used to train SmartEdit.

    Methods Understanding Scenarios Reasoning Scenarios
    PSNR (dB)↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow CLIP Score↑\uparrow Ins-align↑\uparrow PSNR (dB)↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow CLIP Score↑\uparrow Ins-align↑\uparrow
    InstructPix2Pix 21.576 0.721 0.089 22.762 0.537 24.234 0.707 0.083 19.413 0.344
    MagicBrush 18.120 0.680 0.143 22.620 0.290 22.101 0.694 0.113 19.755 0.283
    InstructDiffusion 23.258 0.743 0.067 23.080 0.697 21.453 0.666 0.117 19.523 0.483
    SmartEdit-7B 22.049 0.731 0.087 23.611 0.712 25.258 0.742 0.055 20.950 0.789
    SmartEdit-13B 23.596 0.751 0.068 23.536 0.771 25.757 0.747 0.051 20.777 0.817

    In reasoning scenarios, SmartEdit-7B and SmartEdit-13B significantly outperform all baselines across every metric, achieving Ins-align scores of 0.789 and 0.817 respectively compared to InstructPix2Pix (0.344) and InstructDiffusion (0.483). This indicates that the MLLM backbone allows the model to leverage world knowledge to identify and edit target objects. Scaling the MLLM backbone from 7B to 13B yields consistent performance gains across understanding (Ins-align increases from 0.712 to 0.771) and reasoning scenarios (Ins-align increases from 0.789 to 0.817).

  7. Knowl 7 — Ablation Analysis of the Bidirectional Interaction Module

    data/table

    An ablation study evaluated the contribution of bidirectional interaction within SmartEdit-7B on the Reason-Edit benchmark across three configurations:

    1. Plain (Exp 1): The BIM module is completely removed; the Q-Former text feature is passed directly into the diffusion UNet cross-attention without visual feature cross-modulation.
    2. SimpleCA (Exp 2): Unidirectional cross-attention only, where text features from Q-Former modulate visual features, discarding the second cross-attention block and self-attention block.
    3. BIM (Exp 3): The full bidirectional interaction module.
    Exp ID Plain SimpleCA BIM Understanding Scenarios Reasoning Scenarios
    PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow CLIP↑\uparrow Ins-align↑\uparrow PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow CLIP↑\uparrow Ins-align↑\uparrow
    1 ✓ 20.975 0.713 0.108 23.36 0.695 23.848 0.725 0.074 20.33 0.694
    2 ✓ 19.557 0.692 0.126 23.66 0.692 23.508 0.716 0.081 20.17 0.722
    3 ✓ 22.049 0.731 0.087 23.61 0.712 25.258 0.742 0.055 20.95 0.789

    Removing BIM (Exp 1) causes a substantial decline in background fidelity and instruction alignment (Reasoning Ins-align drops from 0.789 to 0.694; Reasoning PSNR drops from 25.258 dB to 23.848 dB). Using unidirectional cross-attention (SimpleCA) degrades background preservation further (Understanding PSNR drops to 19.557 dB). The full BIM structure delivers the highest background retention and semantic accuracy across both understanding and reasoning tasks.

  8. Knowl 8 — Ablation Analysis of Dataset Composition

    data/table

    An ablation study conducted on SmartEdit-7B isolates the contributions of conventional editing data, segmentation data, and complex synthetic editing data on the Reason-Edit dataset.

    Exp ID Edit Segmentation Synthetic Understanding Scenarios Reasoning Scenarios
    PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow CLIP↑\uparrow Ins-align↑\uparrow PSNR↑\uparrow SSIM↑\uparrow LPIPS↓\downarrow CLIP↑\uparrow Ins-align↑\uparrow
    1 ✓ 17.568 0.664 0.171 22.79 0.201 22.400 0.706 0.102 19.22 0.233
    2 ✓ ✓ 18.960 0.690 0.143 22.83 0.361 21.774 0.693 0.116 19.82 0.311
    3 ✓ ✓ 19.562 0.702 0.111 22.32 0.440 23.595 0.715 0.079 20.43 0.567
    4 ✓ ✓ ✓ 22.049 0.731 0.087 23.61 0.712 25.258 0.742 0.055 20.95 0.789
    • Training on conventional editing data alone (Exp 1) yields poor alignment (Ins-align of 0.201 on understanding and 0.233 on reasoning).
    • Adding perception/segmentation data (Exp 2) enhances spatial perception and increases Understanding Ins-align from 0.201 to 0.361.
    • Adding 476 synthetic complex editing pairs (Exp 3) directly stimulates reasoning capability, increasing Reasoning Ins-align to 0.567.
    • Combining all datasets (Exp 4) achieves super-additive improvements across both tasks (Understanding Ins-align reaches 0.712; Reasoning Ins-align reaches 0.789), confirming that segmentation data (perception grounding) and synthetic editing data (reasoning activation) play complementary roles.

Coverage note — None was omitted; all contributed architectural designs, mathematical formulations, training strategies, evaluation datasets, and experimental ablation results are covered.

References

  1. 1.Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18392–18402, 2023. 2, 3
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 3
  3. 3.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Cocostuff: Thing and stuff classes in context. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1209–1218, 2018. 6
  4. 4.Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, and Yinqiang Zheng. Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. arXiv preprint arXiv:2304.08465, 2023. 3
  5. 5.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-toend object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020. 4
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 3
  7. 7.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi. Instructblip: Towards generalpurpose vision-language models with instruction tuning. Advances in Neural Information Processing Systems, 36, 2024. 3
  8. 8.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34:8780–8794, 2021. 2, 3
  9. 9.Yuying Ge, Yixiao Ge, Ziyun Zeng, Xintao Wang, and Ying Shan. Planting a seed of vision in large language model. arXiv preprint arXiv:2307.08041, 2023. 3
  10. 10.Yuying Ge, Sijie Zhao, Ziyun Zeng, Yixiao Ge, Chen Li, Xintao Wang, and Ying Shan. Making llama see and draw with seed tokenizer. arXiv preprint arXiv:2310.01218, 2023. 3
  11. 11.Zigang Geng, Binxin Yang, Tiankai Hang, Chen Li, Shuyang Gu, Ting Zhang, Jianmin Bao, Zheng Zhang, Han Hu, Dong Chen, et al. Instructdiffusion: A generalist modeling interface for vision tasks. arXiv preprint arXiv:2309.03895, 2023. 3
  12. 12.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022. 3
  13. 13.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 2, 3
  14. 14.Alain Hore and Djemel Ziou. Image quality metrics: Psnr vs. ssim. In 2010 20th international conference on pattern recognition, pages 2366–2369. IEEE, 2010. 6
  15. 15.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan AllenZhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021. 4
  16. 16.Xuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu, and Qiang Xu. Direct inversion: Boosting diffusion-based editing with 3 lines of code. arXiv preprint arXiv:2304.04269, 2023. 3
  17. 17.Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6007–6017, 2023. 3
  18. 18.Jing Yu Koh, Daniel Fried, and Ruslan Salakhutdinov. Generating images with multimodal language models. arXiv preprint arXiv:2305.17216, 2023. 3, 4
  19. 19.Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. Lisa: Reasoning segmentation via large language model. arXiv preprint arXiv:2308.00692, 2023. 2, 5, 6
  20. 20.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023. 4
  21. 21.Chang Liu, Henghui Ding, and Xudong Jiang. Gres: Generalized referring expression segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23592–23601, 2023. 6
  22. 22.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023. 2, 3, 6
  23. 23.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021. 2, 3
  24. 24.Yiran Qin, Enshen Zhou, Qichang Liu, Zhenfei Yin, Lu Sheng, Ruimao Zhang, Yu Qiao, and Jing Shao. Mp5: A multi-modal open-ended embodied system in minecraft via active perception. arXiv preprint arXiv:2312.07472, 2023. 3
  25. 25.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 2, 6
  26. 26.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1 (2):3, 2022. 2, 3
  27. 27.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3
  28. 28.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. Unet: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pages 234–241. Springer, 2015. 2, 3
  29. 29.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 35:36479–36494, 2022. 2, 3
  30. 30.Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative pretraining in multimodality. arXiv preprint arXiv:2307.05222, 2023. 3
  31. 31.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 3
  32. 32.Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. Plug-and-play diffusion features for text-driven image-to-image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1921–1930, 2023. 3
  33. 33.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 4, 6
  34. 34.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14, pages 69–85. Springer, 2016. 6
  35. 35.Lili Yu, Bowen Shi, Ramakanth Pasunuru, Benjamin Muller, Olga Golovneva, Tianlu Wang, Arun Babu, Binh Tang, Brian Karrer, Shelly Sheynin, et al. Scaling autoregressive multimodal models: Pretraining and instruction tuning. arXiv preprint arXiv:2309.02591, 2023. 3
  36. 36.Kai Zhang, Lingbo Mo, Wenhu Chen, Huan Sun, and Yu Su. Magicbrush: A manually annotated dataset for instructionguided image editing. arXiv preprint arXiv:2306.10012, 2023. 2, 3
  37. 37.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018. 6
  38. 38.Enshen Zhou, Yiran Qin, Zhenfei Yin, Yuzhou Huang, Ruimao Zhang, Lu Sheng, Yu Qiao, and Jing Shao. Minedreamer: Learning to follow instructions via chain-ofimagination for simulated-world control. arXiv preprint arXiv:2403.12037, 2024. 3
  39. 39.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 2, 3

Citation

MLA
Huang, Y., et al. “SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2312.06739v1.
APA
Huang, Y., Xie, L., Wang, X., Yuan, Z., Cun, X., Ge, Y., Zhou, J., Dong, C., Huang, R., Zhang, R., & Shan, Y. (2023). SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models. arXiv. http://arxiv.org/abs/2312.06739v1
Chicago
Huang, Y., L. Xie, X. Wang, et al. 2023. “SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models”. arXiv. http://arxiv.org/abs/2312.06739v1.
Harvard
Huang, Y. et al. (2023) “SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.06739v1.
Vancouver
1. Huang Y, Xie L, Wang X, et al (2023) SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models. arXiv

BibTeX

@article{huang2023smartedit,
  title = {SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models},
  author = {Huang, Yuzhou and Xie, Liangbin and Wang, Xintao and Yuan, Ziyang and Cun, Xiaodong and Ge, Yixiao and Zhou, Jiantao and Dong, Chao and Huang, Rui and Zhang, Ruimao and Shan, Ying},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.06739v1},
  eprint = {2312.06739}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE