ObjectStitch: Object Compositing with Diffusion Model
Yizhi SongZhifei ZhangZhe LinScott CohenBrian L. PriceJianming ZhangSoo Ye KimDaniel G. Aliaga
Proposes a self-supervised generative framework based on conditional diffusion models that unifies color harmonization, viewpoint adjustment, geometry correction, and shadow generation into a single pipeline to insert objects realistically into background scenes.
Inserting an object from one image into another realistically is a major challenge in digital content creation, e-commerce, and design. Conventional workflows require separate, labor-intensive steps to correct geometry, adjust lighting and colors, and synthesize shadows. Furthermore, text-guided generative tools often fail to preserve the specific visual identity and fine details of a target object, while creating manually labeled datasets for image compositing is prohibitively expensive and difficult to scale.
The article demonstrates a unified, self-supervised generative framework called ObjectStitch that automatically composites an object into a background scene using conditional diffusion models. The objective is to evaluate whether a single model can adjust perspective, lighting, color, and shadows simultaneously while faithfully preserving the original object's visual appearance without requiring human-annotated training data.
The approach introduces a two-part framework comprising a content adaptor and a conditional generator adapted from a pretrained text-to-image diffusion model. The content adaptor translates visual features from an image encoder into multi-modal guidance tokens, allowing the diffusion generator to preserve both high-level semantics and fine-grained visual details. The entire system is trained without manual annotations using self-supervised synthetic data with spatial perturbations, color adjustments, and crop-and-shift augmentations. Evaluation was conducted on a newly gathered benchmark of 503 challenging real-world image pairs through automated fidelity metrics and a user study collecting 1,494 votes across more than 170 participants.
The key findings show that the proposed unified framework substantially outperforms existing baseline pipelines in both visual realism and object faithfulness. In the user study, human evaluators overwhelmingly preferred the model's outputs over text-guided and noise-guided diffusion baselines. In quantitative evaluations, the complete framework achieved the best image quality score (an FID of 15.43 compared to 17.90 to 18.83 for ablated models and baselines) and the highest visual fidelity scores. Ablation experiments confirmed that both the content adaptor and the synthetic data augmentations are critical to maintaining object identity and scene integration.
These results imply that automated, single-step generative compositing can replace fragmented multi-stage editing pipelines, significantly lowering editing costs, turnaround times, and manual labor. By operating without manual dataset labeling or per-object fine-tuning, the method provides a scalable foundation for commercial creative tools. Practitioners looking to adopt generative editing can leverage image-conditioned diffusion to automate complex compositing tasks that previously required specialized artistic skill.
Moving forward, the article suggests enhancing control over identity preservation and expanding beyond localized bounding-box edits. Current constraints include a lack of precise control knobs for feature preservation and the inability to cast global shadows beyond the specified mask region, which can occasionally leave boundary artifacts. Confidence in the reported performance is high for common object-scene pairings, though additional research on joint instance-mask prediction and end-to-end multi-view training is recommended to address complex global scene interactions.
- Paper: SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Chenlin Meng et al. (2022). ObjectStitch builds upon the foundational concept of guided diffusion-based image editing and inpainting established in SDEdit to achieve realistic image manipulation.
- Paper: Blended Diffusion for Text-driven Editing of Natural Images, Omri Avrahami et al. (2021). Blended Diffusion introduced core strategies for localized, mask-based editing with diffusion models that ObjectStitch adapts and generalizes to image-conditioned compositing.
- Paper: Prompt-to-Prompt Image Editing with Cross Attention Control, Amir Hertz et al. (2022). Understanding how internal cross-attention control enables semantic and structural preservation in diffusion models provides necessary background for ObjectStitch's content adaptor.
- Paper: Null-text Inversion for Editing Real Images using Guided Diffusion Models, Ron Mokady et al. (2022). Null-text Inversion established techniques for conditioning and guidance in diffusion models without fine-tuning, motivating ObjectStitch's training and conditioning pipeline.
- Paper: An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion, Rinon Gal et al. (2022). Textual Inversion represents the foundational subject-personalization paradigm that ObjectStitch aims to improve upon by enabling zero-shot, feedforward subject compositing.
- Paper: Multi-Concept Customization of Text-to-Image Diffusion, Nupur Kumari et al. (2022). Custom Diffusion establishes how diffusion cross-attention parameters adapt to specific visual identities, which ObjectStitch streamlines into an un-tuned conditional framework.
- Paper: RePaint: Inpainting using Denoising Diffusion Probabilistic Models, Andreas Lugmayr et al. (2022). RePaint lays out the foundational mechanics of diffusion-driven inpainting that ObjectStitch enhances to incorporate identity-preserving object insertion.
- Paper: Rendering synthetic objects into real scenes: bridging traditional and image-based graphics with global illumination and high dynamic range photography, P. Debevec (1998). This seminal graphics paper formulates the physical principles of illumination, shadowing, and compositing that modern neural compositing architectures like ObjectStitch aim to automate.
- Paper: DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing, Chong Mou et al. (2024). DiffEditor builds upon object-level manipulation in diffusion models by introducing image prompt encodings and deterministic rollbacks for flexible reference-based editing.
- Paper: DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation, Hong Chen et al. (2024). DisenBooth extends the problem of subject insertion and identity preservation by disentangling identity embeddings from background context.
- Paper: Kosmos-G: Generating Images in Context with Multimodal Large Language Models, Xichen Pan et al. (2024). Kosmos-G advances zero-shot subject-driven compositing by leveraging multimodal large language models to insert visual entities directly from interleaved multi-image inputs.
- Paper: MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis, Dewei Zhou et al. (2024). MIGC extends spatial control over instance generation by using a multi-instance attention controller to manage multiple distinct subjects without attribute bleeding.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). UniReal generalizes localized subject compositing into a unified framework that learns real-world dynamic consistency across diverse editing tasks.
- Paper: SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models, Yuzhou Huang et al. (2024). SmartEdit builds beyond direct visual insertion by integrating multimodal reasoning to execute complex, instruction-based object edits.
