Conditional Object-Centric Learning from Video
Thomas KipfGamaleldin Fathy ElsayedAravindh MahendranAustin StoneSara SabourGeorg HeigoldRico JonschkowskiAlexey DosovitskiyKlaus Greff
Proposes a sequential extension to Slot Attention that combines optical flow prediction with sparse initial location cues, enabling object-centric models to segment, track, and generalize across complex video scenes with minimal supervision.
Building computer vision systems that understand dynamic scenes in terms of individual objects is essential for robust generalization, sample efficiency, and reliable physical reasoning. While fully unsupervised methods have successfully discovered objects in simple, synthetic environments, they consistently fail when scaled to realistic video with complex textures and cluttered backgrounds. Furthermore, fully autonomous models struggle to determine the intended level of visual granularity, often splitting a single entity into arbitrary pieces or merging multiple objects together without a mechanism for user guidance.
To address these limitations, the article introduces Slot Attention for Video (SAVi), an architecture designed to segment and track multiple objects across video sequences. The primary objective is to demonstrate that pairing a self-supervised motion prediction objective with minimal, initial location hints enables effective multi-object segmentation and tracking in visually complex scenes without relying on dense, expensive annotations.
To evaluate this framework, the authors conducted experiments across benchmark synthetic datasets of increasing visual complexity, progressing from basic geometric shapes to visually rich scenes featuring realistic photographic backgrounds and scanned real-world objects with physical collisions. The SAVi architecture processes video sequentially by alternating between a correction step that updates a latent set of object representations (slots) via competitive attention and a prediction step that models temporal interactions using self-attention. The model is conditioned only on the initial frame with light geometric cues, such as bounding boxes or single-point coordinates, and is trained to predict optical flow across subsequent frames.
The investigation produced several key findings. First, conditioning SAVi on simple initial cues allows it to achieve high tracking accuracy (approximately 93.8% foreground clustering similarity and 72% segmentation overlap) on moderately complex video, substantially outperforming traditional unsupervised baselines. Second, on visually complex scenes where fully unsupervised models fail entirely (dropping to 23–33% accuracy), SAVi achieves up to 82.8% clustering similarity when paired with a standard residual network backbone, matching or exceeding specialized label-propagation techniques that require fully detailed masks. Third, the model maintains high tracking quality even when tested on video sequences four times longer than those seen during training, as well as on previously unseen objects and novel backgrounds (showing less than a 2% drop in accuracy). Fourth, the system exhibits high tolerance to imprecise inputs, maintaining tracking stability when initial coordinates are perturbed by noise up to 20% of an object's average size. Finally, the authors observed that initial prompts provide a flexible control interface: targeting a composite object as a single entity tracks the whole item, whereas providing separate prompts for its sub-components directs the model to track individual parts independently.
These findings demonstrate that high-capacity models do not require full per-frame supervision or dense initial masks to achieve reliable object tracking. Instead, weak motion signals combined with lightweight spatial cues provide sufficient structure for robust scene decomposition. This significantly reduces data annotation costs and provides a practical mechanism for downstream applications, such as robotics or autonomous systems, to interactively specify which entities or parts a vision system should track.
For future implementations, practitioners can replace ground-truth motion targets with estimated flow derived from unsupervised motion models, as experiments confirm this maintains performance. Organizations should explore weakly supervised workflows that utilize coarse bounding boxes or click-points rather than pixel-level annotations. Additional development is needed to support non-rigid objects, complex camera motions, and completely static visual environments where optical flow signals are less informative. The reported results provide high confidence in synthetic and constrained domains, but cautious pilot testing is advised before deploying these architectures into unconstrained, real-world visual environments.
- Paper: Unsupervised Learning of Depth and Ego-Motion from Video, Tinghui Zhou et al. (2017). This paper establishes foundational principles for unsupervised learning of depth and temporal dynamics from unlabeled video sequences that motivate motion-based visual learning.
- Paper: Unsupervised Learning of Video Representations using LSTMs, Nitish Srivastava et al. (2015). This work introduces sequential autoencoding and future frame prediction for unsupervised video representation learning, providing foundational concepts for temporal object modeling.
- Paper: Relation Networks for Object Detection, Han Hu et al. (2017). This study introduces attention-based object relation modeling across bounding regions, laying conceptual groundwork for relational and slot-based multi-object attention.
- Paper: SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos, Gamaleldin F. Elsayed et al. (2022). SAVi++ directly builds upon and extends the sequential Slot Attention video architecture by integrating depth prediction to handle camera motion and scaling to complex real-world driving scenes.
- Paper: Object-Centric Slot Diffusion, Jindong Jiang et al. (2023). This work extends unsupervised slot-centric representations by pairing Slot Attention with latent diffusion decoders to achieve higher-fidelity compositional discovery in challenging multi-object datasets.
