ObjectStitch: Object Compositing with Diffusion Model

Yizhi SongZhifei ZhangZhe LinScott CohenBrian L. PriceJianming ZhangSoo Ye KimDaniel G. Aliaga

article2023CVPR143 citations

Proposes a self-supervised generative framework based on conditional diffusion models that unifies color harmonization, viewpoint adjustment, geometry correction, and shadow generation into a single pipeline to insert objects realistically into background scenes.

Listen

Inserting an object from one image into another realistically is a major challenge in digital content creation, e-commerce, and design. Conventional workflows require separate, labor-intensive steps to correct geometry, adjust lighting and colors, and synthesize shadows. Furthermore, text-guided generative tools often fail to preserve the specific visual identity and fine details of a target object, while creating manually labeled datasets for image compositing is prohibitively expensive and difficult to scale.

The article demonstrates a unified, self-supervised generative framework called ObjectStitch that automatically composites an object into a background scene using conditional diffusion models. The objective is to evaluate whether a single model can adjust perspective, lighting, color, and shadows simultaneously while faithfully preserving the original object's visual appearance without requiring human-annotated training data.

The approach introduces a two-part framework comprising a content adaptor and a conditional generator adapted from a pretrained text-to-image diffusion model. The content adaptor translates visual features from an image encoder into multi-modal guidance tokens, allowing the diffusion generator to preserve both high-level semantics and fine-grained visual details. The entire system is trained without manual annotations using self-supervised synthetic data with spatial perturbations, color adjustments, and crop-and-shift augmentations. Evaluation was conducted on a newly gathered benchmark of 503 challenging real-world image pairs through automated fidelity metrics and a user study collecting 1,494 votes across more than 170 participants.

The key findings show that the proposed unified framework substantially outperforms existing baseline pipelines in both visual realism and object faithfulness. In the user study, human evaluators overwhelmingly preferred the model's outputs over text-guided and noise-guided diffusion baselines. In quantitative evaluations, the complete framework achieved the best image quality score (an FID of 15.43 compared to 17.90 to 18.83 for ablated models and baselines) and the highest visual fidelity scores. Ablation experiments confirmed that both the content adaptor and the synthetic data augmentations are critical to maintaining object identity and scene integration.

These results imply that automated, single-step generative compositing can replace fragmented multi-stage editing pipelines, significantly lowering editing costs, turnaround times, and manual labor. By operating without manual dataset labeling or per-object fine-tuning, the method provides a scalable foundation for commercial creative tools. Practitioners looking to adopt generative editing can leverage image-conditioned diffusion to automate complex compositing tasks that previously required specialized artistic skill.

Moving forward, the article suggests enhancing control over identity preservation and expanding beyond localized bounding-box edits. Current constraints include a lack of precise control knobs for feature preservation and the inability to cast global shadows beyond the specified mask region, which can occasionally leave boundary artifacts. Confidence in the reported performance is high for common object-scene pairings, though additional research on joint instance-mask prediction and end-to-end multi-view training is recommended to address complex global scene interactions.

Cover for ObjectStitch: Object Compositing with Diffusion Model

Abstract

Object compositing based on 2D images is a challenging problem since it typically involves multiple processing stages such as color harmonization, geometry correction and shadow generation to generate realistic results. Furthermore, annotating training data pairs for compositing requires substantial manual effort from professionals, and is hardly scalable. Thus, with the recent advances in generative models, in this work, we propose a self-supervised framework for object compositing by leveraging the power of conditional diffusion models. Our framework can holistically address the object compositing task in a unified model, transforming the viewpoint, geometry, color and shadow of the generated object while requiring no manual labeling. To preserve the input object’s characteristics, we introduce a content adaptor that helps to maintain categorical semantics and object appearance. A data augmentation method is further adopted to improve the fidelity of the generator. Our method outperforms relevant baselines in both realism and faithfulness of the synthesized result images in a user study on various real-world images.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Image Compositing
  • 2.2. Guided Image Synthesis
  • 3. Proposed Method
  • 3.1. Generator
  • 3.2. Content Adaptor
  • 3.3. Self-supervised Framework
  • 3.3.1 Data Generation and Augmentation
  • 3.3.2 Training
  • 4. Experiments
  • 4.1. Training Details
  • 4.2. Quantitative Evaluation
  • 4.3. Qualitative Evaluation
  • 4.4. Ablation Study
  • 5. Conclusion and Future Work
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Generative Object Compositing Problem Formulation

    definition

    Generative object compositing is defined as follows: given an input triplet (Io,Ibg,M)(I_o, I_{bg}, M), where Io∈RHs×Ws×3I_o \in \mathbb{R}^{H_s \times W_s \times 3} is a source foreground object image, Ibg∈RHt×Wt×3I_{bg} \in \mathbb{R}^{H_t \times W_t \times 3} is a target background image, and M∈RHt×Wt×1M \in \mathbb{R}^{H_t \times W_t \times 1} is a binary mask where pixels in the desired insertion area are set to 00 and the background pixels are set to 11, the goal is to synthesize a composite image IoutI_{out} within the masked region while leaving the unmasked background Ibg⊙MI_{bg} \odot M unchanged.

    The mask MM serves as a soft spatial constraint on the scale and position of the composited object. The generated output must realistically harmonize geometry, viewpoint, color, lighting, and contact shadows with the target background IbgI_{bg} while preserving the visual identity and texture characteristics of the source object IoI_o.

  2. Knowl 2 — Content Adaptor Architecture for Visual Guidance in Diffusion Models

    model/method

    To condition a pretrained text-to-image diffusion model directly on an object image IoI_o rather than textual descriptions, the Content Adaptor bridges the domain and dimensional gap between visual token representations and text token representations.

    Given an image feature sequence E~=Ci(Io)∈Rk×257×1024\tilde{E} = C_i(I_o) \in \mathbb{R}^{k \times 257 \times 1024} extracted from a frozen CLIP ViT-L/14 image encoder CiC_i (where kk is the batch size), the Content Adaptor T(⋅)T(\cdot) transforms E~\tilde{E} into an adaptive embedding sequence E^∈Rk×77×768\hat{E} \in \mathbb{R}^{k \times 77 \times 768} matching the input format of the diffusion model's cross-attention layers. The Content Adaptor consists of three sequential modules:

    1. A 1D convolutional layer that downsamples the visual token sequence length from 257257 to 7777.
    2. Multi-head self-attention blocks that map the visual feature representations into the text-embedding semantic space.
    3. A Multi-Layer Perceptron (MLP) that projects the feature channel dimension from 10241024 to 768768.
  3. Knowl 3 — Two-Stage Training Scheme for the Content Adaptor

    model/method

    The Content Adaptor T(⋅)T(\cdot) is trained in two distinct self-supervised stages:

    Stage 1: Multi-Modal Semantic Alignment (Pretraining)

    The adaptor is optimized on image-caption pairs (I,t)(I, t) to translate visual tokens into multi-modal text-aligned embeddings. Encoders CiC_i (image) and CtC_t (text) from CLIP are frozen. Given E~=Ci(I)\tilde{E} = C_i(I) and target text embedding E=Ct(t)E = C_t(t), the adaptor is optimized using the L1 translation loss:

    Ldist=∥T(E~)−E∥1\mathcal{L}_{dist} = \lVert T(\tilde{E}) - E \rVert_1

    Stage 2: Appearance and Identity Preservation (Fine-Tuning)

    To capture instance-level visual details beyond coarse semantics, the Content Adaptor is inserted into a frozen pretrained latent diffusion model. The adaptor parameters are updated by encouraging visual reconstruction of the original image with the diffusion objective:

    Ladapt=ET,ϵ∼N(0,1)[∥ϵ−ϵθ(It∘M,t,T(E~))∥22]\mathcal{L}_{adapt} = \mathbb{E}_{T, \epsilon \sim \mathcal{N}(0, 1)} \left[ \lVert \epsilon - \epsilon_{\theta}(I_t \circ M, t, T(\tilde{E})) \rVert_2^2 \right]

    where ϵθ\epsilon_\theta denotes the denoising network, ItI_t is the noisy latent representation at timestep tt, MM is the spatial mask, and ϵ\epsilon is standard Gaussian noise.

  4. Knowl 4 — Diffusion Generator Formulation and Mask Blending

    model/method

    The generator module in ObjectStitch adapts a pretrained text-to-image Latent Diffusion Model for localized object compositing. Conditioning on the adaptive embedding E^=T(Ci(Io))\hat{E} = T(C_i(I_o)) is achieved via spatial cross-attention within the denoising U-Net:

    Softmax((WQEx)(WKE^)Td)WVE^=AV\text{Softmax}\left(\frac{(\boldsymbol{W}_Q E_x)(\boldsymbol{W}_K \hat{E})^T}{\sqrt{d}}\right)\boldsymbol{W}_V \hat{E} = \boldsymbol{A}\boldsymbol{V}

    where Ex∈RdxE_x \in \mathbb{R}^{d_x} is an intermediate feature map of the denoising U-Net, and WQ∈Rd×dx\boldsymbol{W}_Q \in \mathbb{R}^{d \times d_x}, WK∈Rd×de\boldsymbol{W}_K \in \mathbb{R}^{d \times d_e}, and WV∈Rd×de\boldsymbol{W}_V \in \mathbb{R}^{d \times d_e} are learned projection matrices.

    To ensure the background outside the compositing region is perfectly preserved, the latent representation is masked at every reverse diffusion timestep tt via mask blending (It∘MI_t \circ M), restricting denoising exclusively to the target masked hole 1−M1 - M. With the trained Content Adaptor frozen, the generator weights θ\theta are fine-tuned using the loss:

    Lgen=EE^,ϵ∼N(0,1)[∥ϵ−ϵθ(It∘M,t,E^)∥22]\mathcal{L}_{gen} = \mathbb{E}_{\hat{E}, \epsilon \sim \mathcal{N}(0, 1)} \left[ \lVert \epsilon - \epsilon_{\theta}(I_t \circ M, t, \hat{E}) \rVert_2^2 \right]

  5. Knowl 5 — Self-Supervised Synthetic Data Generation Pipeline

    algorithm

    To train ObjectStitch without manual compositing annotations, training pairs are synthetically constructed from single images using spatial and photometric perturbations:

    Input: Unannotated image collection, Pretrained instance segmentation model
    Output: Training triplets (I_o, I_bg, M) with ground truth target I_gt
    for each image in collection do
      Segment object instances and extract bounding boxes
      Filter out objects with extreme scales (very small or very large)
      for each valid object do
        Set I_gt = original image
        Set I_bg = original image
        Apply projective transformation by randomly perturbing the 4 bounding box corners
        Apply random rotation with angle drawn uniformly from [-20°, 20°]
        Apply random photometric/color jittering to the object region
        Extract perturbed object as source object I_o
        Define mask M as binary bounding box enclosing the object region (0 inside, 1 outside)
        Yield (I_o, I_bg, M) with target I_gt
      end for
    end for
  6. Knowl 6 — Crop and Shift Data Augmentation for Compositing

    model/method

    To enhance local visual fidelity, texture resolution, and geometric generalization in generative compositing, random crop and shift augmentations are applied during both training and inference.

    During augmentation, the background image IbgI_{bg} and associated mask MM are randomly translated and cropped subject to the constraint that the entire foreground insertion region defined by 1−M1 - M remains fully enclosed within the cropped window. This transformation effectively increases the relative pixel resolution of the insertion area fed into the diffusion model, providing greater detail for shadow synthesis, texture generation, and boundary harmonization.

  7. Knowl 7 — Modified CLIP Scores for Reference-Guided Compositing Evaluation

    equation

    To quantitatively evaluate the semantic fidelity and visual identity matching between the composited result IpredI_{pred} and the reference ground truth IgtI_{gt}, modified CLIP text and image scores are defined as:

    Ctxt=E[s⋅f(Ipred)⋅g(B(Igt))]\mathcal{C}_{txt} = \mathbb{E}\left[ s \cdot f(I_{pred}) \cdot g(B(I_{gt})) \right]

    Cimg=E[s⋅f(Ipred)⋅f(Igt)]\mathcal{C}_{img} = \mathbb{E}\left[ s \cdot f(I_{pred}) \cdot f(I_{gt}) \right]

    where f(⋅)f(\cdot) is the visual feature extractor of a pretrained CLIP model, g(⋅)g(\cdot) is the text feature extractor of the CLIP model, B(⋅)B(\cdot) is a pretrained BLIP captioning model that generates a descriptive textual caption from IgtI_{gt}, and ss is the learned logit scale parameter of CLIP.

  8. Knowl 8 — Component Ablation Results for ObjectStitch

    data/table

    An ablation study evaluated the contribution of the Content Adaptor, appearance optimization (Stage 2 training), crop/shift data augmentation, and comparison against text-based captioning conditioning (BLIP).

    Adaptor Optimization Augmentation BLIP FID ↓\downarrow CLIP text score ↑\uparrow CLIP image score ↑\uparrow
    ✓ ✓ ✓ 15.4331 29.8594 97.0625
    ✓ ✓ 18.8254 29.8281 96.0000
    ✓ ✓ 16.1381 29.7969 96.8125
    ✓ ✓ 18.3698 29.7813 96.1875
    ✓ ✓ 17.9000 29.7188 95.8125

    The full ObjectStitch pipeline attains the best score across all metrics (FID of 15.4331, CLIP text score of 29.8594, and CLIP image score of 97.0625). Removing the Content Adaptor degrades FID by +3.3923 and CLIP image score by -1.0625. Omitting Stage 2 appearance optimization reduces the CLIP image score to 96.8125. Disabling crop and shift augmentation causes a substantial degradation in image synthesis fidelity (FID worsens to 18.3698). Replacing the Content Adaptor with BLIP text captions yields the lowest identity preservation (CLIP image score of 95.8125).

  9. Knowl 9 — User Study Evaluation on Real-World Compositing Dataset

    empirical result

    A user study across 503 real-world image pairs (1,494 total votes from over 170 subjects) evaluated ObjectStitch against three baselines: Copy-and-Paste, Stable Diffusion fine-tuned on BLIP captions, and Stable Diffusion with BLIP captions + SDEdit (noise strength 0.65).

    • Realism Preference:

      • ObjectStitch vs. Copy-and-Paste: ObjectStitch won 85.14% vs. 14.86%.
      • ObjectStitch vs. BLIP: ObjectStitch won 55.62% vs. 44.38%.
      • ObjectStitch vs. SDEdit: ObjectStitch won 55.02% vs. 44.98%.
    • Object Closeness / Faithfulness Preference:

      • ObjectStitch vs. BLIP: ObjectStitch won 71.35% vs. 28.65%.
      • ObjectStitch vs. SDEdit: ObjectStitch won 59.04% vs. 40.96%.

    These results show that ObjectStitch generates composite objects that are more realistic and consistent with target lighting and geometry than traditional blending or text-guided diffusion baselines, while achieving superior visual fidelity compared to SDEdit and BLIP.

  10. Knowl 10 — Limitations of ObjectStitch

    limitation

    The ObjectStitch framework exhibits three identified limitations:

    1. Appearance Preservation Control: It lacks explicit fine-grained mechanisms to enforce strict identity or structural preservation of the input object beyond the guidance provided by the adaptor embedding.
    2. Mask-Constrained Global Effects: Because the diffusion denoising process strictly operates inside the bounding mask 1−M1 - M via latent replacement, shadows or reflections that naturally fall outside the designated bounding box cannot be generated into the surrounding background.
    3. Boundary Artifacts: Minor blending and boundary artifacts occasionally occur around the perimeter of the inserted object bounding box.

Coverage note — None was omitted; all key contributions including framework architecture, two-stage training scheme, synthetic data generation, mathematical formulations, metrics, quantitative ablations, user study findings, and limitations are fully covered.

References

  1. 1.Omri Avrahami, Dani Lischinski, and Ohad Fried. Blended diffusion for text-driven editing of natural images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18208–18218, 2022. 2, 3
  2. 2.Samaneh Azadi, Deepak Pathak, Sayna Ebrahimi, and Trevor Darrell. Compositional gan: Learning image-conditional binary composition. International Journal of Computer Vision, 128(10):2570–2585, 2020. 2
  3. 3.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al. ediffi: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022. 2
  4. 4.Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018. 2
  5. 5.Bor-Chun Chen and Andrew Kae. Toward realistic image compositing with adversarial learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8415–8424, 2019. 2, 3
  6. 6.Wenyan Cong, Li Niu, Jianfu Zhang, Jing Liang, and Liqing Zhang. Bargainnet: Background-guided domain translation for image harmonization. In 2021 IEEE International Conference on Multimedia and Expo (ICME), pages 1–6. IEEE, 2021. 1, 2, 3
  7. 7.Wenyan Cong, Jianfu Zhang, Li Niu, Liu Liu, Zhixin Ling, Weiyuan Li, and Liqing Zhang. Dovenet: Deep image harmonization via domain verification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8394–8403, 2020. 1, 2, 3
  8. 8.Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Deep image homography estimation. arXiv preprint arXiv:1606.03798, 2016. 5
  9. 9.Zhida Feng, Zhenyu Zhang, Xintong Yu, Yewei Fang, Lanxin Li, Xuyi Chen, Yuxiang Lu, Jiaxiang Liu, Weichong Yin, Shikun Feng, et al. Ernie-vilg 2.0: Improving text-to-image diffusion model with knowledge-enhanced mixture-of-denoising-experts. arXiv preprint arXiv:2210.15257, 2022. 2, 3
  10. 10.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Communications of the ACM, 63(11):139–144, 2020. 2
  11. 11.Liu He, Yijuan Lu, John Corring, Dinei Florencio, and Cha Zhang. Diffusion-based document layout generation. arXiv preprint arXiv:2303.10787, 2023. 2
  12. 12.Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718, 2021. 8
  13. 13.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017. 8
  14. 14.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020. 2, 3
  15. 15.Yannick Hold-Geoffroy, Kalyan Sunkavalli, Sunil Hadap, Emiliano Gambaretto, and Jean-Franc¸ois Lalonde. Deep outdoor illumination estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7312–7321, 2017. 2
  16. 16.Yan Hong, Li Niu, and Jianfu Zhang. Shadow generation for composite image in real-world scenes. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 914–922, 2022. 1, 2, 3
  17. 17.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017. 2
  18. 18.Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al. Spatial transformer networks. Advances in neural information processing systems, 28, 2015. 2
  19. 19.Yifan Jiang, He Zhang, Jianming Zhang, Yilin Wang, Zhe Lin, Kalyan Sunkavalli, Simon Chen, Sohrab Amirghodsi, Sarah Kong, and Zhangyang Wang. Ssh: A self-supervised framework for image harmonization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4832–4841, 2021. 1, 2, 3
  20. 20.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4401–4410, 2019. 2
  21. 21.Kevin Karsch, Varsha Hedau, David Forsyth, and Derek Hoiem. Rendering synthetic objects into legacy photographs. ACM Transactions on Graphics (TOG), 30(6):1–12, 2011. 2
  22. 22.Kevin Karsch, Kalyan Sunkavalli, Sunil Hadap, Nathan Carr, Hailin Jin, Rafael Fonte, Michael Sittig, and David Forsyth. Automatic scene inference for 3d object compositing. ACM Transactions on Graphics (TOG), 33(3):1–15, 2014. 2
  23. 23.Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models. arXiv preprint arXiv:2210.09276, 2022. 2, 3
  24. 24.Jean-Francois Lalonde and Alexei A Efros. Using color compatibility for assessing image realism. In 2007 IEEE 11th International Conference on Computer Vision, pages 1–8. IEEE, 2007. 3
  25. 25.Seung Hoon Lee, Seunghyun Lee, and Byung Cheol Song. Vision transformer for small-size datasets. arXiv preprint arXiv:2112.13492, 2021. 5
  26. 26.Youngwan Lee and Jongyoul Park. Centermask: Real-time anchor-free instance segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 13906–13915, 2020. 5
  27. 27.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. arXiv preprint arXiv:2201.12086, 2022. 6, 8
  28. 28.Chen-Hsuan Lin, Ersin Yumer, Oliver Wang, Eli Shechtman, and Simon Lucey. St-gan: Spatial transformer generative adversarial networks for image compositing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 9455–9464, 2018. 1, 2
  29. 29.Daquan Liu, Chengjiang Long, Hongpan Zhang, Hanning Yu, Xinzhi Dong, and Chunxia Xiao. Arshadowgan: Shadow generative adversarial network for augmented reality in single light scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8139–8148, 2020. 1, 2
  30. 30.Xihui Liu, Dong Huk Park, Samaneh Azadi, Gong Zhang, Arman Chopikyan, Yuxiao Hu, Humphrey Shi, Anna Rohrbach, and Trevor Darrell. More control for free! image synthesis with semantic diffusion guidance. arXiv preprint arXiv:2112.05744, 2021. 2, 3
  31. 31.Shitong Luo and Wei Hu. Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2837–2845, 2021. 2
  32. 32.Chenlin Meng, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. Sdedit: Image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073, 2021. 2, 3, 6
  33. 33.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021. 2, 3
  34. 34.Patrick Perez, Michel Gangnet, and Andrew Blake. Poisson image editing. In ACM SIGGRAPH 2003 Papers, pages 313–318. 2003. 1
  35. 35.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 2
  36. 36.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021. 4, 6, 8
  37. 37.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021. 3
  38. 38.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022. 2, 3, 6
  39. 39.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. arXiv preprint arXiv:2208.12242, 2022. 2, 3
  40. 40.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022. 3
  41. 41.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021. 5
  42. 42.Yichen Sheng, Yifan Liu, Jianming Zhang, Wei Yin, A Cengiz Oztireli, He Zhang, Zhe Lin, Eli Shechtman, and Bedrich Benes. Controllable shadow generation using pixel height maps. In European Conference on Computer Vision, pages 240–256. Springer, 2022. 1
  43. 43.Yichen Sheng, Jianming Zhang, and Bedrich Benes. Ssn: Soft shadow network for image compositing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4380–4390, 2021. 1, 2, 3
  44. 44.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015. 3
  45. 45.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. Advances in Neural Information Processing Systems, 32, 2019. 3
  46. 46.Kalyan Sunkavalli, Micah K Johnson, Wojciech Matusik, and Hanspeter Pfister. Multi-scale image harmonization. ACM Transactions on Graphics (TOG), 29(4):1–10, 2010. 3
  47. 47.Yi-Hsuan Tsai, Xiaohui Shen, Zhe Lin, Kalyan Sunkavalli, Xin Lu, and Ming-Hsuan Yang. Deep image harmonization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3789–3797, 2017. 1
  48. 48.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 5
  49. 49.Shaoan Xie, Zhifei Zhang, Zhe Lin, Tobias Hinz, and Kun Zhang. Smartbrush: Text and shape guided object inpainting with diffusion model. arXiv preprint arXiv:2212.05034, 2022. 8
  50. 50.Ning Xu, Brian Price, Scott Cohen, and Thomas Huang. Deep image matting. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2970–2979, 2017. 2
  51. 51.Ben Xue, Shenghui Ran, Quan Chen, Rongfei Jia, Binqiang Zhao, and Xing Tang. Dccf: Deep comprehensible color filter learning framework for high-resolution image harmonization. arXiv preprint arXiv:2207.04788, 2022. 1, 2
  52. 52.Hao Zhou, Sunil Hadap, Kalyan Sunkavalli, and David W Jacobs. Deep single-image portrait relighting. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7194–7202, 2019. 1

Citation

MLA
Song, Y., et al. “ObjectStitch: Object Compositing with Diffusion Model”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 18310–19, https://doi.org/10.1109/CVPR52729.2023.01756.
APA
Song, Y., Zhang, Z., Lin, Z., Cohen, S., Price, B., Zhang, J., Kim, S. Y., & Aliaga, D. (2023). ObjectStitch: Object Compositing with Diffusion Model. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18310–18319. https://doi.org/10.1109/CVPR52729.2023.01756
Chicago
Song, Y., Z. Zhang, Z. Lin, et al. 2023. “ObjectStitch: Object Compositing with Diffusion Model”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18310–19. https://doi.org/10.1109/CVPR52729.2023.01756.
Harvard
Song, Y. et al. (2023) “ObjectStitch: Object Compositing with Diffusion Model”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 18310–18319. Available at: https://doi.org/10.1109/CVPR52729.2023.01756.
Vancouver
1. Song Y, Zhang Z, Lin Z, Cohen S, Price B, Zhang J, Kim SY, Aliaga D (2023) ObjectStitch: Object Compositing with Diffusion Model. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 18310–18319

BibTeX

@inproceedings{Song_2023, title={ObjectStitch: Object Compositing with Diffusion Model}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01756}, DOI={10.1109/cvpr52729.2023.01756}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Song, Yizhi and Zhang, Zhifei and Lin, Zhe and Cohen, Scott and Price, Brian and Zhang, Jianming and Kim, Soo Ye and Aliaga, Daniel}, year={2023}, month=June, pages={18310–18319} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE