Object-Centric Slot Diffusion

Jindong JiangFei DengGautam SinghSungjin Ahn

article2023NeurIPS105 citations

Proposes Latent Slot Diffusion to replace traditional decoders with a conditional latent diffusion model, enabling high-quality unsupervised compositional image generation and object segmentation in complex real-world scenes.

Listen

Artificial intelligence systems often struggle to understand complex physical scenes because they lack an inherent mechanism to break down images into distinct, reusable modular components such as individual objects or background elements. While language models naturally process distinct words and tokens, discovering modular visual structures without expensive manual human annotations remains a core bottleneck for advanced machine reasoning, causal inference, and real-world image manipulation. Previous approaches relied on weak decoders or autoregressive transformer decoders, which frequently fail to capture fine details, struggle with complex visual textures, and require supervised text prompts to guide compositional scene generation.

The article evaluates whether high-capacity latent diffusion models—a class of generative models known for exceptional image synthesis—can be integrated into unsupervised object-centric learning to discover modular visual entities and generate high-fidelity compositional images without text labels.

The researchers developed Latent Slot Diffusion (LSD), an architecture combining a Slot Attention encoder with a conditional latent diffusion decoder operating on compressed image representations from a pre-trained auto-encoder. The model was evaluated across several multi-object synthetic benchmarks (CLEVR, CLEVRTex, MOVi-C, MOVi-E), high-resolution real-world human portraits (FFHQ), and natural image datasets (COCO). Across these datasets, the authors benchmarked LSD against leading transformer-based models (SLATE and SLATE+) on unsupervised segmentation accuracy, downstream property prediction, image synthesis fidelity, and modular image-editing capabilities such as object swapping and face attribute replacement.

The evaluation revealed several key findings in order of importance. First, LSD achieved state-of-the-art unsupervised segmentation and representation accuracy in complex scenes, outperforming the strongest baseline on the demanding MOVi-E benchmark by more than 8% in mean Best Overlap and over 9% in mean Intersection over Union. Second, LSD substantially improved compositional image synthesis fidelity, reducing the standard visual error metric (FID) on complex textures and faces by more than half compared to baseline models (for example, achieving an FID of 27.83 on FFHQ compared to 98.76 for SLATE+). Third, LSD successfully demonstrated the first instance of unsupervised attribute-level editing and face replacement on real-world portrait datasets by manipulating discrete visual slot representations. Finally, the researchers demonstrated that pre-trained diffusion models can be adapted with a lightweight object-encoder (Stable-LSD) to perform unsupervised representation learning and diverse conditional generation directly on unconstrained natural images.

These findings demonstrate that replacing traditional autoregressive decoders with latent diffusion decoders provides the necessary capacity to bridge the gap between synthetic benchmarks and complex real-world visual understanding. This provides organizations and practitioners with a viable path toward modular, controllable image generation and scene analysis without the prohibitive costs and human biases associated with large-scale manual labeling. Furthermore, test-time generation with LSD proved roughly 40% faster than transformer-based baselines, though initial model training required greater compute time.

For technical leaders and practitioners, adopting diffusion-based object-centric frameworks is recommended when deploying visual AI in visually rich, multi-component environments. If applying the model to simple or monotonous images, teams should mix in textured background data during training to prevent the expressive diffusion decoder from bypassing slot conditioning. Before deploying for real-world content generation or automated editing, organizations must implement governance policies to manage risks related to photorealistic image manipulation, deepfake generation, and demographic representation biases arising from unsupervised cluster sampling.

The conclusions should be interpreted in light of certain limitations. LSD exhibits part-whole ambiguity in complex natural settings, occasionally segmenting an object into subparts (such as clothing versus limbs). Additionally, the model remains sensitive to the pre-set number of slots, which can cause under- or over-segmentation when the exact object count in an image is unknown. Overall confidence in LSD's superior generative and segmentation performance on complex scenes is high based on consistent empirical gains across multiple challenging benchmarks.

Cover for Object-Centric Slot Diffusion

Abstract

The recent success of transformer-based image generative models in object-centric learning highlights the importance of powerful image generators for handling complex scenes. However, despite the high expressiveness of diffusion models in image generation, their integration into object-centric learning remains largely unexplored in this domain. In this paper, we explore the feasibility and potential of integrating diffusion models into object-centric learning and investigate the pros and cons of this approach. We introduce Latent Slot Diffusion (LSD), a novel model that serves dual purposes: it is the first object-centric learning model to replace conventional slot decoders with a latent diffusion model conditioned on object slots, and it is also the first unsupervised compositional conditional diffusion model that operates without the need for supervised annotations like text. Through experiments on various object-centric tasks, including the first application of the FFHQ dataset in this field, we demonstrate that LSD significantly outperforms state-of-the-art transformer-based decoders, particularly in more complex scenes, and exhibits superior unsupervised compositional generation quality. In addition, we conduct a preliminary investigation into the integration of pre-trained diffusion models in LSD and demonstrate its effectiveness in real-world image segmentation and generation. Project page is available at https://latentslotdiffusion.github.io

Citation

MLA
Jiang, J., et al. “Object-Centric Slot Diffusion”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 8563–601, https://proceedings.neurips.cc/paper_files/paper/2023/file/1b3ceb8a495a63ced4a48f8429ccdcd8-Paper-Conference.pdf.
APA
Jiang, J., Deng, F., Singh, G., & Ahn, S. (2023). Object-Centric Slot Diffusion. Advances in Neural Information Processing Systems, 36, 8563–8601. https://proceedings.neurips.cc/paper_files/paper/2023/file/1b3ceb8a495a63ced4a48f8429ccdcd8-Paper-Conference.pdf
Chicago
Jiang, J., F. Deng, G. Singh, and S. Ahn. 2023. “Object-Centric Slot Diffusion”. Advances in Neural Information Processing Systems 36: 8563–8601. https://proceedings.neurips.cc/paper_files/paper/2023/file/1b3ceb8a495a63ced4a48f8429ccdcd8-Paper-Conference.pdf.
Harvard
Jiang, J. et al. (2023) “Object-Centric Slot Diffusion”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 8563–8601. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/1b3ceb8a495a63ced4a48f8429ccdcd8-Paper-Conference.pdf.
Vancouver
1. Jiang J, Deng F, Singh G, Ahn S (2023) Object-Centric Slot Diffusion. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 8563–8601

BibTeX

@inproceedings{jiang2023object,
  title = {Object-Centric Slot Diffusion},
  author = {Jiang, Jindong and Deng, Fei and Singh, Gautam and Ahn, Sungjin},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {8563-8601},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/1b3ceb8a495a63ced4a48f8429ccdcd8-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors