Built independently by an author, for readers. Read the story and support ChapterPal

keyword

object discovery

Object discovery is a computer vision and machine learning task that involves identifying, localizing, and segmenting distinct entities within perceptual data such as images and videos without relying on explicit human supervision or instance-level annotations. In this paradigm, computational models learn to decompose complex scenes into modular, object-centric representations, often structured as discrete vectors or slots that capture individual visual properties such as shape, appearance, pose, and spatial extent. By leveraging self-supervised learning signals, generative reconstruction objectives, motion cues, or geometric priors, object discovery enables autonomous systems to parse compositional visual environments, track entities over time, and facilitate downstream reasoning tasks.

6 items

Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames

Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames

Ondrej Biza, Sjoerd van Steenkiste, Mehdi S. M. Sajjadi, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Thomas Kipf

OrganizationsGoogleNortheastern University

Why you should read this

Introduces a mechanism that incorporates per-object spatial symmetries into Slot Attention by dynamically transforming position encodings into slot-centric reference frames, significantly improving data efficiency and unsupervised object discovery across synthetic benchmarks and real-world driving data.

Automatically discovering composable abstractions from raw perceptual data is a long-standing challenge in machine learning. Recent slot-based neural networks that learn about objects in a self-supervised manner have made exciting progress in this direction. However, they typically fall short at adequately capturing spatial symmetries present in the visual world, which leads to sample inefficiency, such as when entangling object appearance and pose. In this paper, we present a simple yet highly effective method for incorporating spatial symmetries via slot-centric reference frames. We incorporate equivariance to per-object pose transformations into the attention and generation mechanism of Slot Attention by translating, scaling, and rotating position encodings. These changes result in little computational overhead, are easy to implement, and can result in large gains in terms of data efficiency and overall improvements to object discovery. We evaluate our method on a wide range of synthetic object discovery benchmarks namely Tetrominoes, CLEVR-Tex, Objects Room and MultiShapeNet, and show promising improvements on the challenging real-world Waymo Open dataset.

Added

2026-10-03

SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos

SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos

Gamaleldin F. Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Michael C. Mozer, Thomas Kipf

OrganizationsGoogle

Why you should read this

Presents an end-to-end slot-based video model that leverages depth prediction and architectural scaling to achieve unsupervised object segmentation and tracking in complex real-world driving scenes.

The visual world can be parsimoniously characterized in terms of distinct entities with sparse interactions. Discovering this compositional structure in dynamic visual scenes has proven challenging for end-to-end computer vision approaches unless explicit instance-level supervision is provided. Slot-based models leveraging motion cues have recently shown great promise in learning to represent, segment, and track objects without direct supervision, but they still fail to scale to complex real-world multi-object videos. In an effort to bridge this gap, we take inspiration from human development and hypothesize that information about scene geometry in the form of depth signals can facilitate object-centric learning. We introduce SAVi++, an object-centric video model which is trained to predict depth signals from a slot-based video representation. By further leveraging best practices for model scaling, we are able to train SAVi++ to segment complex dynamic scenes recorded with moving cameras, containing both static and moving objects of diverse appearance on naturalistic backgrounds, without the need for segmentation supervision. Finally, we demonstrate that by using sparse depth signals obtained from LiDAR, SAVi++ is able to learn emergent object segmentation and tracking from videos in the real-world Waymo Open dataset.

Added

2026-09-28

Object-Centric Slot Diffusion

Object-Centric Slot Diffusion

Jindong Jiang, Fei Deng, Gautam Singh, Sungjin Ahn

OrganizationsKorea Advanced Institute of Science and TechnologyRutgers University

Why you should read this

Proposes Latent Slot Diffusion to replace traditional decoders with a conditional latent diffusion model, enabling high-quality unsupervised compositional image generation and object segmentation in complex real-world scenes.

The recent success of transformer-based image generative models in object-centric learning highlights the importance of powerful image generators for handling complex scenes. However, despite the high expressiveness of diffusion models in image generation, their integration into object-centric learning remains largely unexplored in this domain. In this paper, we explore the feasibility and potential of integrating diffusion models into object-centric learning and investigate the pros and cons of this approach. We introduce Latent Slot Diffusion (LSD), a novel model that serves dual purposes: it is the first object-centric learning model to replace conventional slot decoders with a latent diffusion model conditioned on object slots, and it is also the first unsupervised compositional conditional diffusion model that operates without the need for supervised annotations like text. Through experiments on various object-centric tasks, including the first application of the FFHQ dataset in this field, we demonstrate that LSD significantly outperforms state-of-the-art transformer-based decoders, particularly in more complex scenes, and exhibits superior unsupervised compositional generation quality. In addition, we conduct a preliminary investigation into the integration of pre-trained diffusion models in LSD and demonstrate its effectiveness in real-world image segmentation and generation. Project page is available at https://latentslotdiffusion.github.io

Added

2026-09-26

Hierarchical Open-vocabulary Universal Image Segmentation

Hierarchical Open-vocabulary Universal Image Segmentation

Xudong Wang, Shufan Li, Konstantinos Kallidromitis, Yusuke Kato, Kazuki Kozuka, Trevor Darrell

OrganizationsPanasonicUniversity of California Berkeley

Why you should read this

Presents HIPIE, a unified open-vocabulary framework that resolves segmentation ambiguity across multiple granularities by incorporating hierarchical visual representations alongside decoupled text-image fusion mechanisms for stuff and thing categories.

Open-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions. However, complex visual scenes can be naturally decomposed into simpler parts and abstracted at multiple levels of granularity, introducing inherent segmentation ambiguity. Unlike existing methods that typically sidestep this ambiguity and treat it as an external factor, our approach actively incorporates a hierarchical representation encompassing different semantic-levels into the learning process. We also propose a decoupled text-image fusion mechanism and representation learning modules for both “things” and “stuff”.1 Additionally, we systematically examine the differences that exist in the textual and visual features between these types of categories. Our resulting model, named HIPIE, tackles HIerarchical, oPen-vocabulary, and unIvErsal segmentation tasks within a unified framework. Benchmarked on over 40 datasets, e.g., ADE20K, COCO, Pascal-VOC Part, RefCOCO/RefCOCOg, ODinW and SeginW, HIPIE achieves the state-of-the-art results at various levels of image comprehension, including semantic-level (e.g., semantic segmentation), instance-level (e.g., panoptic/referring segmentation and object detection), as well as part-level (e.g., part/subpart segmentation) tasks.

Added

2026-09-26