Conditional Object-Centric Learning from Video

Thomas KipfGamaleldin Fathy ElsayedAravindh MahendranAustin StoneSara SabourGeorg HeigoldRico JonschkowskiAlexey DosovitskiyKlaus Greff

article2022ICLR323 citations

Proposes a sequential extension to Slot Attention that combines optical flow prediction with sparse initial location cues, enabling object-centric models to segment, track, and generalize across complex video scenes with minimal supervision.

Listen

Building computer vision systems that understand dynamic scenes in terms of individual objects is essential for robust generalization, sample efficiency, and reliable physical reasoning. While fully unsupervised methods have successfully discovered objects in simple, synthetic environments, they consistently fail when scaled to realistic video with complex textures and cluttered backgrounds. Furthermore, fully autonomous models struggle to determine the intended level of visual granularity, often splitting a single entity into arbitrary pieces or merging multiple objects together without a mechanism for user guidance.

To address these limitations, the article introduces Slot Attention for Video (SAVi), an architecture designed to segment and track multiple objects across video sequences. The primary objective is to demonstrate that pairing a self-supervised motion prediction objective with minimal, initial location hints enables effective multi-object segmentation and tracking in visually complex scenes without relying on dense, expensive annotations.

To evaluate this framework, the authors conducted experiments across benchmark synthetic datasets of increasing visual complexity, progressing from basic geometric shapes to visually rich scenes featuring realistic photographic backgrounds and scanned real-world objects with physical collisions. The SAVi architecture processes video sequentially by alternating between a correction step that updates a latent set of object representations (slots) via competitive attention and a prediction step that models temporal interactions using self-attention. The model is conditioned only on the initial frame with light geometric cues, such as bounding boxes or single-point coordinates, and is trained to predict optical flow across subsequent frames.

The investigation produced several key findings. First, conditioning SAVi on simple initial cues allows it to achieve high tracking accuracy (approximately 93.8% foreground clustering similarity and 72% segmentation overlap) on moderately complex video, substantially outperforming traditional unsupervised baselines. Second, on visually complex scenes where fully unsupervised models fail entirely (dropping to 23–33% accuracy), SAVi achieves up to 82.8% clustering similarity when paired with a standard residual network backbone, matching or exceeding specialized label-propagation techniques that require fully detailed masks. Third, the model maintains high tracking quality even when tested on video sequences four times longer than those seen during training, as well as on previously unseen objects and novel backgrounds (showing less than a 2% drop in accuracy). Fourth, the system exhibits high tolerance to imprecise inputs, maintaining tracking stability when initial coordinates are perturbed by noise up to 20% of an object's average size. Finally, the authors observed that initial prompts provide a flexible control interface: targeting a composite object as a single entity tracks the whole item, whereas providing separate prompts for its sub-components directs the model to track individual parts independently.

These findings demonstrate that high-capacity models do not require full per-frame supervision or dense initial masks to achieve reliable object tracking. Instead, weak motion signals combined with lightweight spatial cues provide sufficient structure for robust scene decomposition. This significantly reduces data annotation costs and provides a practical mechanism for downstream applications, such as robotics or autonomous systems, to interactively specify which entities or parts a vision system should track.

For future implementations, practitioners can replace ground-truth motion targets with estimated flow derived from unsupervised motion models, as experiments confirm this maintains performance. Organizations should explore weakly supervised workflows that utilize coarse bounding boxes or click-points rather than pixel-level annotations. Additional development is needed to support non-rigid objects, complex camera motions, and completely static visual environments where optical flow signals are less informative. The reported results provide high confidence in synthetic and constrained domains, but cautious pilot testing is advised before deploying these architectures into unconstrained, real-world visual environments.

  • Paper: Unsupervised Learning of Depth and Ego-Motion from Video, Tinghui Zhou et al. (2017). This paper establishes foundational principles for unsupervised learning of depth and temporal dynamics from unlabeled video sequences that motivate motion-based visual learning.
  • Paper: Unsupervised Learning of Video Representations using LSTMs, Nitish Srivastava et al. (2015). This work introduces sequential autoencoding and future frame prediction for unsupervised video representation learning, providing foundational concepts for temporal object modeling.
  • Paper: Relation Networks for Object Detection, Han Hu et al. (2017). This study introduces attention-based object relation modeling across bounding regions, laying conceptual groundwork for relational and slot-based multi-object attention.
  • Paper: SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos, Gamaleldin F. Elsayed et al. (2022). SAVi++ directly builds upon and extends the sequential Slot Attention video architecture by integrating depth prediction to handle camera motion and scaling to complex real-world driving scenes.
  • Paper: Object-Centric Slot Diffusion, Jindong Jiang et al. (2023). This work extends unsupervised slot-centric representations by pairing Slot Attention with latent diffusion decoders to achieve higher-fidelity compositional discovery in challenging multi-object datasets.
Cover for Conditional Object-Centric Learning from Video

Abstract

Object-centric representations are a promising path toward more systematic generalization by providing flexible abstractions upon which compositional world models can be built. Recent work on simple 2D and 3D datasets has shown that models with object-centric inductive biases can learn to segment and represent meaningful objects from the statistical structure of the data alone without the need for any supervision. However, such fully-unsupervised methods still fail to scale to diverse realistic data, despite the use of increasingly complex inductive biases such as priors for the size of objects or the 3D geometry of the scene. In this paper, we instead take a weakly-supervised approach and focus on how 1) using the temporal dynamics of video data in the form of optical flow and 2) conditioning the model on simple object location cues can be used to enable segmenting and tracking objects in significantly more realistic synthetic data. We introduce a sequential extension to Slot Attention which we train to predict optical flow for realistic looking synthetic scenes and show that conditioning the initial state of this model on a small set of hints, such as center of mass of objects in the first frame, is sufficient to significantly improve instance segmentation. These benefits generalize beyond the training distribution to novel objects, novel backgrounds, and to longer video sequences. We also find that such initial-state-conditioning can be used during inference as a flexible interface to query the model for specific objects or parts of objects, which could pave the way for a range of weakly-supervised approaches and allow more effective interaction with trained models.

Table of Contents

  • 1 Introduction
  • 2 Slot Attention for Video (SAVi)
  • 3 Related Work
  • 4 Experiments
  • 4.1 Unsupervised video decomposition
  • 4.2 More realistic datasets
  • 4.3 Conditional video decomposition
  • 4.4 Test time generalization
  • 4.5 Limitations
  • 5 Conclusion
  • References
  • A Appendix
  • A.1 Additional related work
  • A.2 Real-world robotics task
  • A.3 Additional qualitative results
  • A.4 Additional quantitative results
  • A.5 Dataset details
  • A.6 Architecture details and hyperparameters
  • A.7 Baseline details
  • A.8 Metric details

Knowls

  1. Knowl 1 — Slot Attention for Video Architecture

    model/method

    Slot Attention for Video (SAVi) is a sequential object-centric architecture designed to decompose, segment, and track multiple objects across video frames over time. SAVi maintains a set of KK latent slot vectors St=[st1,…,stK]∈RK×D\mathcal{S}_t = [s_t^1, \dots, s_t^K] \in \mathbb{R}^{K \times D} at each time step t∈{1,…,T}t \in \{1, \dots, T\}, where DD is the slot latent dimensionality.

    The framework operates recurrently across frames through a predictor-corrector cycle:

    1. Encoder: Each video frame xtx_t is processed by a convolutional neural network (CNN) combined with linear spatial position embeddings to yield flattened visual features ht=fenc(xt)∈RN×Dench_t = f_{\text{enc}}(x_t) \in \mathbb{R}^{N \times D_{\text{enc}}}, where N=height×widthN = \text{height} \times \text{width} and DencD_{\text{enc}} is the feature dimension.
    2. Corrector: Updates current slot states St\mathcal{S}_t conditioned on visual features hth_t using Slot Attention, where attention weights are normalized across slots via a softmax to encourage competitive input partitioning, followed by a Gated Recurrent Unit (GRU) update.
    3. Predictor: A Transformer encoder block with multi-head self-attention and an MLP transition model processes the corrected slots to model temporal dynamics and pairwise slot interactions, producing the prior slots St+1\mathcal{S}_{t+1} for the next time step.
    4. Decoder: A Spatial Broadcast Decoder independently reconstructs per-slot targets (optical flow or RGB frames) and alpha masks from corrected slot representations S^t\hat{\mathcal{S}}_t, which are composited into a single output via a normalized weighted sum.
  2. Knowl 2 — SAVi Corrector and Predictor Formulation

    equation

    For a video frame at time step tt, let ht∈RN×Dench_t \in \mathbb{R}^{N \times D_{\text{enc}}} denote the NN spatial feature vectors from the image encoder and St∈RK×D\mathcal{S}_t \in \mathbb{R}^{K \times D} denote the set of KK slot vectors.

    The Corrector applies Slot Attention normalized across the KK slots: At=softmax⁡K(1Dk(ht)⋅q(St)T)∈RN×KA_t = \operatorname{softmax}_K\left(\frac{1}{\sqrt{D}} k(h_t) \cdot q(\mathcal{S}_t)^T\right) \in \mathbb{R}^{N \times K} Ut=1Zt∑n=1NAt,n⊙v(ht,n)∈RK×D,Zt=∑n=1NAt,nU_t = \frac{1}{Z_t} \sum_{n=1}^N A_{t,n} \odot v(h_{t,n}) \in \mathbb{R}^{K \times D}, \quad Z_t = \sum_{n=1}^N A_{t,n} where k,q,vk, q, v are learned linear projections with Layer Normalization mapping to dimension DD, and ⊙\odot is the Hadamard product. Each slot k∈{1,…,K}k \in \{1, \dots, K\} is updated recurrently: s^tk=GRU⁡(utk,stk)\hat{s}_t^k = \operatorname{GRU}(u_t^k, s_t^k) When multiple Slot Attention iterations are used, a residual MLP with Layer Normalization (LN) is applied: s^tk←s^tk+MLP⁡(LN⁡(s^tk))\hat{s}_t^k \leftarrow \hat{s}_t^k + \operatorname{MLP}(\operatorname{LN}(\hat{s}_t^k)).

    The Predictor models object interactions and temporal dynamics via multi-head self-attention across the updated slots S^t=[s^t1,…,s^tK]\hat{\mathcal{S}}_t = [\hat{s}_t^1, \dots, \hat{s}_t^K]: S~t=LN⁡(MultiHeadSelfAttn⁡(S^t)+S^t)\tilde{\mathcal{S}}_t = \operatorname{LN}\left(\operatorname{MultiHeadSelfAttn}(\hat{\mathcal{S}}_t) + \hat{\mathcal{S}}_t\right) St+1=LN⁡(MLP⁡(S~t)+S~t)\mathcal{S}_{t+1} = \operatorname{LN}\left(\operatorname{MLP}(\tilde{\mathcal{S}}_t) + \tilde{\mathcal{S}}_t\right) Both steps are permutation equivariant with respect to slot indices.

  3. Knowl 3 — Slot Initialization Strategies in SAVi

    model/method

    SAVi initializes its KK latent slots at the first time step t=1t=1 via either conditional or unconditional initialization schemes:

    • Conditional Initialization (Bounding Boxes / Center of Mass): Bounding box coordinates [ymin⁡,xmin⁡,ymax⁡,xmax⁡][y_{\min}, x_{\min}, y_{\max}, x_{\max}] or center of mass coordinates [y,x][y, x] for foreground objects are encoded independently using a shared Multi-Layer Perceptron (MLP) with a hidden layer of 256 units and ReLU activation to produce initial slot representations s1k∈RDs_1^k \in \mathbb{R}^D. For slots without an object hint (when KK exceeds the object count), a fixed coordinate vector (such as [−1,−1,−1,−1][-1, -1, -1, -1] or [−1,−1][-1, -1]) is passed to the MLP.
    • Conditional Initialization (Segmentation Masks): Binary segmentation masks of individual foreground objects are processed independently by a shared CNN encoder, followed by spatial average pooling, Layer Normalization, and a Dense layer producing s1ks_1^k. Unconditioned slots receive an empty all-zero mask.
    • Unconditional Initialization: Slots are initialized without hints either by sampling independently from a standard Gaussian distribution N(0,I)\mathcal{N}(0, I) per video or by learning a set of shared initial slot vectors parameters optimized during training.
  4. Knowl 4 — Optical Flow Reconstruction Objective and Decoder

    model/method

    SAVi decodes slot representations S^t\hat{\mathcal{S}}_t after the corrector step using a slot-wise Spatial Broadcast Decoder. Each slot s^tk∈RD\hat{s}_t^k \in \mathbb{R}^D is replicated across an 8×88 \times 8 spatial grid, augmented with linear spatial coordinate embeddings, and processed by a transposed convolutional network to produce per-slot predictions y^tk\hat{y}_t^k and alpha mask logits m^tk\hat{m}_t^k.

    The alpha masks are normalized across the KK slots using a softmax function, and the composite video frame prediction yty_t is formed as a weighted sum: mt=softmax⁡K(m^t),yt=∑k=1Kmtk⊙ytkm_t = \operatorname{softmax}_K(\hat{m}_t), \quad y_t = \sum_{k=1}^K m_t^k \odot y_t^k

    The network is trained end-to-end to minimize the pixel-wise squared error between predicted output yty_t and the target frame yttruey_t^{\text{true}} over sequence length TT: Lrec=∑t=1T∥yt−yttrue∥2\mathcal{L}_{\text{rec}} = \sum_{t=1}^T \|y_t - y_t^{\text{true}}\|^2 where yttruey_t^{\text{true}} is optical flow represented as a 3-channel RGB image (or alternatively RGB pixel frames for standard reconstruction).

  5. Knowl 5 — Unsupervised and Conditional Video Segmentation Performance

    data/table

    Segmentation performance measured by foreground Adjusted Rand Index (FG-ARI) and mean Intersection over Union (mIoU) across synthetic video benchmarks. On CATER, methods are evaluated unconditionally with RGB reconstruction. On MOVi and MOVi++, models are trained with optical flow supervision and evaluated on conditional tracking given first-frame hints (excluding the conditioned frame from evaluation):

    CATER MOVi MOVi++
    Model (+ Conditioning) FG-ARI (%) FG-ARI (%) mIoU (%) FG-ARI (%) mIoU (%)
    Slot Attention (image baseline) 7.3±0.37.3 \pm 0.3 – – – –
    MONet (image baseline) 41.2±0.541.2 \pm 0.5 – – – –
    S-IODINE 66.8±1.566.8 \pm 1.5 – – – –
    SIMONe 91.8±1.691.8 \pm 1.6 74.8±4.274.8 \pm 4.2 – 32.7±2.332.7 \pm 2.3 –
    SCALOR – 81.2±0.281.2 \pm 0.2 – 22.7±0.922.7 \pm 0.9 –
    CRW – – 42.442.4 – 50.950.9
    T-VOS – – 50.450.4 – 46.446.4
    Segmentation Propagation – 61.5±0.661.5 \pm 0.6 49.8±0.449.8 \pm 0.4 37.8±0.537.8 \pm 0.5 33.3±0.233.3 \pm 0.2
    Flow kk-Means + Center of Mass – 30.230.2 5.85.8 32.632.6 9.39.3
    SAVi (unconditional) 92.8±0.892.8 \pm 0.8 78.2±0.778.2 \pm 0.7 – 47.6±0.247.6 \pm 0.2 –
    SAVi + Segmentation 97.9±0.497.9 \pm 0.4 93.7±0.293.7 \pm 0.2 72.0±0.372.0 \pm 0.3 70.4±2.070.4 \pm 2.0 43.0±0.643.0 \pm 0.6
    SAVi + Bounding Box – 93.7±0.093.7 \pm 0.0 71.2±0.671.2 \pm 0.6 77.4±0.577.4 \pm 0.5 45.9±1.245.9 \pm 1.2
    SAVi + Center of Mass – 93.8±0.193.8 \pm 0.1 72.1±0.272.1 \pm 0.2 78.3±0.678.3 \pm 0.6 43.5±2.943.5 \pm 2.9
    SAVi (ResNet) + Bounding Box – – – 82.8±0.482.8 \pm 0.4 50.7±0.250.7 \pm 0.2

    SAVi significantly outperforms prior unsupervised video models on CATER and simple 3D physical scenes (MOVi). On the photorealistic MOVi++ dataset with complex HDR backgrounds and scanned real-world objects, unsupervised baselines collapse to static image patches (~22--33% FG-ARI), while conditional SAVi achieves strong instance segmentation and tracking using weak single-point (center of mass) or bounding box initial hints.

  6. Knowl 6 — Supervision Signal and Conditioning Noise Ablations

    empirical result

    Ablation experiments on training supervision signals and conditioning noise reveal key dependencies in SAVi:

    1. Optical Flow vs. RGB Reconstruction: On simple scenes (MOVi), SAVi trained with an RGB reconstruction objective achieves similar decomposition quality to optical flow supervision. However, on complex scenes with diverse textures and HDR backgrounds (MOVi++), optical flow supervision is essential; training with RGB pixel reconstruction without flow causes the model to fail to discover objects. Replacing ground-truth optical flow with unsupervised flow estimates from SMURF yields nearly identical segmentation performance.
    2. Conditioning Precision and Noise Tolerance: Conditioning on simple spatial hints (bounding box or single center-of-mass coordinate) achieves comparable performance to conditioning on full binary segmentation masks on MOVi (~93.7% vs 93.7% FG-ARI) and MOVi++ (~77.4--78.3% vs 70.4% FG-ARI). When zero-mean Gaussian noise N(0,σ)\mathcal{N}(0, \sigma) is added to center-of-mass coordinates during training and evaluation, SAVi maintains stable tracking and segmentation up to noise scale σ≈20%\sigma \approx 20\% of the average object diameter before mIoU begins to decay.
  7. Knowl 7 — Out-of-Distribution and Temporal Generalization of SAVi

    empirical result

    SAVi exhibits strong out-of-distribution (OOD) generalization across time horizons, scene appearances, and environments:

    • Temporal Extrapolation: When trained exclusively on short 6-frame sub-sequences, SAVi evaluates stably on full 24-frame videos without temporal degradation or identity switches. In unconditional settings, FG-ARI scores increase over time as additional steps allow the iterative attention mechanism to break symmetry and reliably bind slots to objects.
    • Novel Objects and Backgrounds: A bounding-box conditioned SAVi model trained on MOVi++ achieves 82.0±0.1%82.0 \pm 0.1\% FG-ARI and 54.3±0.3%54.3 \pm 0.3\% mIoU on the default test split, 82.0±0.1%82.0 \pm 0.1\% FG-ARI and 54.4±0.3%54.4 \pm 0.3\% mIoU on unseen objects (~100 novel scanned objects), 82.2±0.1%82.2 \pm 0.1\% FG-ARI and 54.0±0.3%54.0 \pm 0.3\% mIoU on unseen HDR backgrounds (~40 novel scenes), and 80.4±0.1%80.4 \pm 0.1\% FG-ARI and 52.6±0.3%52.6 \pm 0.3\% mIoU on simultaneous novel objects and novel backgrounds.
    • Cross-Dataset Transfer: A SAVi model trained on MOVi++ transfers directly to MOVi at test time without fine-tuning, obtaining 83.7±0.2%83.7 \pm 0.2\% FG-ARI and 53.7±0.6%53.7 \pm 0.6\% mIoU.
  8. Knowl 8 — Interactive Granularity Control via Conditional Querying

    empirical result

    Conditional slot initialization in SAVi enables dynamic selection of object segmentation granularity at inference time without requiring explicit hierarchical modeling during training:

    • When provided with a single bounding box encompassing a multi-part object (such as a laptop containing screen and keyboard, or a composite two-fist robotic tool), SAVi binds a single slot to track and segment the entire composite object as a whole.
    • When provided with separate bounding boxes for individual parts (such as one box for the laptop screen and one for the keyboard base), SAVi allocates separate slots to track each part independently throughout the video sequence, despite never encountering part annotations during training.

    This demonstrates that initial state conditioning provides an interactive query interface to steer slot binding dynamically at inference.

  9. Knowl 9 — Conditional Initialization vs. Hungarian Matching Supervision

    empirical result

    Providing weak supervision via conditional initial slot states outperforms providing the same annotations as targets in a matching-based loss during training.

    When evaluated on the first 6 frames of MOVi:

    • Conditional SAVi (default): Bounding box hints encoded directly into initial slot states achieve 92.9±0.1%92.9 \pm 0.1\% FG-ARI.
    • Learned Slot Initialization + Hungarian Matching Loss: Slots initialized unconditionally with learned parameters and supervised in the first frame using Hungarian bipartite matching and Huber loss on bounding box coordinates achieve 83.0±0.2%83.0 \pm 0.2\% FG-ARI.
    • Gaussian Slot Initialization + Hungarian Matching Loss: Slots initialized from Gaussian noise and supervised via Hungarian matching loss achieve 82.3±0.7%82.3 \pm 0.7\% FG-ARI.

    Conditional initialization avoids permutation matching ambiguity during training and ensures consistent slot-to-object alignment from the start of the sequence.

  10. Knowl 10 — Semantic Property Representation in Optical Flow-Trained Slots

    empirical result

    To evaluate whether slots retain visual appearance properties when trained purely on optical flow prediction, a frozen-feature readout probe (an MLP with a 256-unit hidden layer and ReLU activation) was trained to predict categorical object attributes (color across 8 classes, shape across 3 classes, material across 2 classes) from first-frame slot representations on MOVi.

    • Slot representations trained solely on optical flow prediction achieve an average property classification accuracy of 54.6±1.5%54.6 \pm 1.5\%, significantly above chance level, indicating that slots capture semantic object characteristics despite flow targets being insensitive to color and texture.
    • When SAVi is trained with an RGB frame reconstruction target, probe classification accuracy rises to 85.7±0.7%85.7 \pm 0.7\%.
  11. Knowl 11 — Predictor and Corrector Architecture Ablations in SAVi

    empirical result

    Evaluating architectural variants of SAVi on MOVi++ with bounding box conditioning highlights the role of each architectural component:

    1. No Predictor (Identity Dynamics): Replacing the Transformer self-attention predictor with the identity function yields 83.0±0.2%83.0 \pm 0.2\% FG-ARI but degrades tracking consistency, dropping mIoU from 54.3±0.3%54.3 \pm 0.3\% to 44.7±3.4%44.7 \pm 3.4\%, showing that cross-slot communication is crucial for maintaining correspondence over time.
    2. MLP Predictor (Independent Dynamics): Applying an independent per-slot MLP predictor without cross-slot self-attention yields 82.8±0.2%82.8 \pm 0.2\% FG-ARI and 49.5±2.2%49.5 \pm 2.2\% mIoU, similarly lagging behind the Transformer predictor in tracking consistency.
    3. Inverted Corrector Attention: Normalizing cross-attention over input visual tokens rather than across slots (as in standard cross-attention) drastically impairs performance, yielding 78.3±0.4%78.3 \pm 0.4\% FG-ARI and collapsing mIoU to 11.4±0.4%11.4 \pm 0.4\%, confirming that slot-normalized competitive softmax attention is necessary for object decomposition.
  12. Knowl 12 — Limitations of SAVi

    limitation

    SAVi has several operational and domain limitations:

    1. Flow Target Dependency: Training relies on optical flow supervision signals; while unsupervised optical flow estimators (e.g., SMURF) can be substituted for ground-truth flow, failure in optical flow estimation undermines object representation learning.
    2. Static Objects and Background Disambiguation: Because the model relies on motion cues during training, scenes with completely static objects or complex moving camera setups (where background optical flow resembles object motion) reduce decomposition fidelity (e.g., FG-ARI decreases to 65.5±0.5%65.5 \pm 0.5\% on MOVi++ with moving cameras).
    3. Physical and Dynamic Complexity: The method was primarily evaluated on simulated rigid-body physics datasets with non-deformable objects, leaving scaling to complex real-world deformable objects, non-rigid dynamics, and open-domain videos an open challenge.

Coverage note — No substantial contributed material was omitted. All key models, equations, empirical evaluation tables, generalization studies, part-whole granularity findings, matching comparisons, representation probe experiments, ablations, and stated limitations are covered.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  2. 2.Amir Bar, Roei Herzig, Xiaolong Wang, Anna Rohrbach, Gal Chechik, Trevor Darrell, and Amir Globerson. Compositional video synthesis with action graphs. In International Conference on Machine Learning, 2021.
  3. 3.Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al. Interaction networks for learning about objects, relations and physics. In Advances in Neural Information Processing Systems, pp. 4502–4510, 2016.
  4. 4.Daniel M Bear, Chaofei Fan, Damian Mrowca, Yunzhu Li, Seth Alter, Aran Nayebi, Jeremy Schwartz, Li Fei-Fei, Jiajun Wu, Joshua B Tenenbaum, et al. Learning physical graph representations from visual scenes. In Advances in Neural Information Processing Systems, 2020.
  5. 5.Blender Online Community. Blender - a 3D modelling and rendering package. Blender Foundation, Stichting Blender Foundation, Amsterdam, 2021. URL http://www.blender.org.
  6. 6.James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. JAX: composable transformations of Python+NumPy programs, 2018. URL http://github.com/google/jax.
  7. 7.Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner. MONet: Unsupervised scene decomposition and representation. arXiv preprint arXiv:1901.11390, 2019.
  8. 8.Serkan Cabi, Sergio Gómez Colmenarejo, Alexander Novikov, Ksenia Konyushkova, Scott Reed, Rae Jeong, Konrad Zolna, Yusuf Aytar, David Budden, Mel Vecerik, Oleg Sushkov, David Barker, Jonathan Scholz, Misha Denil, Nando de Freitas, and Ziyu Wang. Scaling data-driven robotics with reward sketching and batch reinforcement learning. In Robotics: Science and Systems, 2020.
  9. 9.Sergi Caelles, Kevis-Kokitsi Maninis, Jordi Pont-Tuset, Laura Leal-Taixé, Daniel Cremers, and Luc Van Gool. One-shot video object segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 221–230, 2017.
  10. 10.Sergi Caelles, Jordi Pont-Tuset, Federico Perazzi, Alberto Montes, Kevis-Kokitsi Maninis, and Luc Van Gool. The 2019 DAVIS challenge on vos: Unsupervised multi-object segmentation. arXiv preprint arXiv:1905.00737, 2019.
  11. 11.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, 2020.
  12. 12.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. arXiv preprint arXiv:2104.14294, 2021.
  13. 13.Chang Chen, Fei Deng, and Sungjin Ahn. Object-centric representation and rendering of 3D scenes. arXiv preprint arXiv:2006.06130, 2020.
  14. 14.Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong, Joshua B Tenenbaum, and Chuang Gan. Grounding physical concepts of objects and events through dynamic visual reasoning. In International Conference on Learning Representations, 2021.
  15. 15.Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1724–1734, 2014.
  16. 16.Eric Crawford and Joelle Pineau. Spatially invariant unsupervised object detection with convolutional neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pp. 3412–3420, 2019.
  17. 17.Eric Crawford and Joelle Pineau. Exploiting spatial invariance for scalable unsupervised object tracking. In Proceedings of the AAAI Conference on Artificial Intelligence, pp. 3684–3692, 2020.
  18. 18.Antonia Creswell, Rishabh Kabra, Chris Burgess, and Murray Shanahan. Unsupervised object-based transition models for 3D partially observable environments. arXiv preprint arXiv:2103.04693, 2021.
  19. 19.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2009.
  20. 20.David Ding, Felix Hill, Adam Santoro, Malcolm Reynolds, and Matt Botvinick. Attention over learned object embeddings enables complex visual reasoning. In Advances in Neural Information Processing Systems, 2021a.
  21. 21.Mingyu Ding, Zhenfang Chen, Tao Du, Ping Luo, Joshua B Tenenbaum, and Chuang Gan. Dynamic visual reasoning by learning differentiable physics models from video and language. In Advances In Neural Information Processing Systems, 2021b.
  22. 22.Yilun Du, Shuang Li, Yash Sharma, B. Joshua Tenenbaum, and Igor Mordatch. Unsupervised learning of compositional energy concepts. In Advances in Neural Information Processing Systems, 2021a.
  23. 23.Yilun Du, Kevin Smith, Tomer Ulman, Joshua Tenenbaum, and Jiajun Wu. Unsupervised discovery of 3D physical objects from video. In International Conference on Learning Representations, 2021b.
  24. 24.Martin Engelcke, Adam R Kosiorek, Oiwi Parker Jones, and Ingmar Posner. GENESIS: Generative scene inference and sampling with object-centric latent representations. In International Conference on Learning Representations, 2020.
  25. 25.SM Ali Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, David Szepesvari, Geoffrey E Hinton, et al. Attend, infer, repeat: Fast scene understanding with generative models. In Advances in Neural Information Processing Systems, pp. 3225–3233, 2016.
  26. 26.Alon Faktor and Michal Irani. Video segmentation by non-local consensus voting. In Proceedings of the British Machine Vision Conference, 2014.
  27. 27.Fabian B Fuchs, Adam R Kosiorek, Li Sun, Oiwi Parker Jones, and Ingmar Posner. End-to-end recurrent multi-object tracking and trajectory prediction with relational reasoning. arXiv preprint arXiv:1907.12887, 2019.
  28. 28.Rohit Girdhar and Deva Ramanan. Cater: A diagnostic dataset for compositional actions and temporal reasoning. arXiv preprint arXiv:1910.04744, 2019.
  29. 29.Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman. Video action transformer network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  30. 30.Google Research. Google scanned objects, 2020. URL https://app.ignitionrobotics.org/GoogleResearch/fuel/collections/Google%20Scanned%20Objects.
  31. 31.Anirudh Goyal, Aniket Didolkar, Nan Rosemary Ke, Charles Blundell, Philippe Beaudoin, Nicolas Heess, Michael Mozer, and Yoshua Bengio. Neural production systems. arXiv preprint arXiv:2103.01937, 2021a.
  32. 32.Anirudh Goyal, Alex Lamb, Phanideep Gampa, Philippe Beaudoin, Sergey Levine, Charles Blundell, Yoshua Bengio, and Michael Mozer. Factorizing declarative and procedural knowledge in structured, dynamic environments. In International Conference on Learning Representations, 2021b.
  33. 33.Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf. Recurrent independent mechanisms. In International Conference on Learning Representations, 2021c.
  34. 34.Klaus Greff, Antti Rasmus, Mathias Berglund, Tele Hao, Harri Valpola, and Jürgen Schmidhuber. Tagger: Deep unsupervised perceptual grouping. In Advances in Neural Information Processing Systems, pp. 4484–4492, 2016.
  35. 35.Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber. Neural expectation maximization. In Advances in Neural Information Processing Systems, pp. 6691–6701, 2017.
  36. 36.Klaus Greff, Raphaèl Lopez Kaufman, Rishabh Kabra, Nick Watters, Christopher Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, and Alexander Lerchner. Multi-object representation learning with iterative variational inference. In International Conference on Machine Learning, pp. 2424–2433, 2019.
  37. 37.Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber. On the binding problem in artificial neural networks. arXiv preprint arXiv:2012.05208, 2020.
  38. 38.Klaus Greff, Andrea Tagliasacchi, Derek Liu, Issam Laradji, Or Litany, and Luca Prasso. Kubric: A data generation pipeline for creating semi-realistic synthetic multi-object videos, 2021. URL https://github.com/google-research/kubric.
  39. 39.Adam W Harley, Yiming Zuo, Jing Wen, Ayush Mangal, Shubhankar Potdar, Ritwick Chaudhry, and Katerina Fragkiadaki. Track, check, repeat: An EM approach to unsupervised tracking. arXiv preprint arXiv:2104.03424, 2021.
  40. 40.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 770–778, 2016.
  41. 41.Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee. Flax: A neural network library and ecosystem for JAX, 2020. URL http://github.com/google/flax.
  42. 42.Paul Henderson and Christoph H Lampert. Unsupervised object-centric video generation and decomposition in 3D. In Advances in Neural Information Processing Systems, 2020.
  43. 43.Roei Herzig, Elad Levi, Huijuan Xu, Hang Gao, Eli Brosh, Xiaolong Wang, Amir Globerson, and Trevor Darrell. Spatio-temporal action graph networks. arXiv preprint arXiv:1812.01233, 2019.
  44. 44.Geoffrey E Hinton, Sara Sabour, and Nicholas Frosst. Matrix capsules with EM routing. In International Conference on Learning Representations, 2018.
  45. 45.Hossam Isack and Yuri Boykov. Energy-based geometric multi-model fitting. International Journal of Computer Vision, 97(2):123–147, 2012.
  46. 46.Allan Jabri, Andrew Owens, and Alexei A Efros. Space-time correspondence as a contrastive random walk. In Advances in Neural Information Processing Systems, 2020.
  47. 47.Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al. Perceiver IO: A general architecture for structured inputs & outputs. arXiv preprint arXiv:2107.14795, 2021a.
  48. 48.Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. Perceiver: General perception with iterative attention. In International Conference on Machine Learning, 2021b.
  49. 49.Ashesh Jain, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena. Structural-RNN: Deep learning on spatio-temporal graphs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  50. 50.Jindong Jiang, Sepehr Janghorbani, Gerard de Melo, and Sungjin Ahn. SCALOR: Generative world models with scalable object representations. In International Conference on Learning Representations, 2020.
  51. 51.Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  52. 52.Rishabh Kabra, Daniel Zoran, Loic Matthey Goker Erdogan, Antonia Creswell, Matthew Botvinick, Alexander Lerchner, and Christopher P. Burgess. SIMONe: View-invariant, temporally-abstracted object representations via unsupervised video decomposition. arXiv preprint arXiv:2106.03849, 2021.
  53. 53.Daniel Kahneman, Anne Treisman, and Brian J Gibbs. The reviewing of object files: Object-specific integration of information. Cognitive psychology, 24(2):175–219, 1992.
  54. 54.Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion. MDETR – Modulated Detection for End-to-End Multi-Modal Understanding. arXiv preprint arXiv:2104.12763, 2021.
  55. 55.Laurynas Karazija, Iro Laina, and Christian Rupprecht. Clevrtex: A texture-rich benchmark for unsupervised multi-object segmentation, 2021. URL https://openreview.net/forum?id=J4Nl2qRMDrR.
  56. 56.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015.
  57. 57.Thomas Kipf, Elise van der Pol, and Max Welling. Contrastive learning of structured world models. In International Conference on Learning Representations, 2020.
  58. 58.Adam Kosiorek, Hyunjik Kim, Yee Whye Teh, and Ingmar Posner. Sequential attend, infer, repeat: Generative modelling of moving objects. In Advances in Neural Information Processing Systems, pp. 8606–8616, 2018.
  59. 59.Xueting Li, Sifei Liu, Shalini De Mello, Xiaolong Wang, Jan Kautz, and Ming-Hsuan Yang. Joint-task self-supervised learning for temporal correspondence. In Advances in Neural Information Processing Systems, 2019.
  60. 60.Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng, Jindong Jiang, and Sungjin Ahn. SPACE: Unsupervised object-oriented scene representation via spatial attention and decomposition. In International Conference on Learning Representations, 2020.
  61. 61.Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. Object-centric learning with slot attention. In Advances in Neural Information Processing Systems, 2020.
  62. 62.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015.
  63. 63.Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations, 2017.
  64. 64.Jonathon Luiten, Paul Voigtlaender, and Bastian Leibe. Premvos: Proposal-generation, refinement and merging for video object segmentation. In Asian Conference on Computer Vision, pp. 565–580. Springer, 2018.
  65. 65.Aravindh Mahendran, James Thewlis, and Andrea Vedaldi. Cross pixel optical-flow similarity for self-supervised learning. In Asian Conference on Computer Vision, pp. 99–116. Springer, 2019.
  66. 66.Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer. TrackFormer: Multi-Object Tracking with Transformers. arXiv preprint arXiv:2101.02702, 2021.
  67. 67.Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alexander Sorkine-Hornung, and Luc Van Gool. The 2017 DAVIS challenge on video object segmentation. arXiv preprint arXiv:1704.00675, 2017.
  68. 68.Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton. Dynamic routing between capsules. In Advances in Neural Information Processing Systems, pp. 3856–3866, 2017.
  69. 69.Sara Sabour, Andrea Tagliasacchi, Soroosh Yazdani, Geoffrey E Hinton, and David J Fleet. Unsupervised part representation by flow capsules. arXiv preprint arXiv:2011.13920, 2020.
  70. 70.Adam Santoro, Ryan Faulkner, David Raposo, Jack Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, and Timothy Lillicrap. Relational recurrent neural networks. In Advances in Neural Information Processing Systems, pp. 7299–7310, 2018.
  71. 71.Elizabeth S Spelke and Katherine D Kinzler. Core knowledge. Developmental science, 10(1):89–96, 2007.
  72. 72.Aleksandar Stanić and Jürgen Schmidhuber. R-SQAIR: Relational sequential attend, infer, repeat. arXiv preprint arXiv:1910.05231, 2019.
  73. 73.Karl Stelzner, Robert Peharz, and Kristian Kersting. Faster attend-infer-repeat with tractable probabilistic models. In International Conference on Machine Learning, pp. 5966–5975, 2019.
  74. 74.Karl Stelzner, Kristian Kersting, and Adam R Kosiorek. Decomposing 3D scenes into objects via unsupervised volume segmentation. arXiv preprint arXiv:2104.01148, 2021.
  75. 75.Austin Stone, Daniel Maurer, Alper Ayvaci, Anelia Angelova, and Rico Jonschkowski. SMURF: Self-teaching multi-frame unsupervised RAFT with full-image warping. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  76. 76.Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8934–8943, 2018.
  77. 77.Aniruddha Kembhavi Derek Hoiem Tanmay Gupta, Amita Kamath. Towards general purpose vision systems. arXiv preprint arXiv:2104.00743, 2021.
  78. 78.Sjoerd van Steenkiste, Michael Chang, Klaus Greff, and Jürgen Schmidhuber. Relational neural expectation maximization: Unsupervised discovery of objects and their interactions. In International Conference on Learning Representations, 2018.
  79. 79.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, pp. 5998–6008, 2017.
  80. 80.Rishi Veerapaneni, John D Co-Reyes, Michael Chang, Michael Janner, Chelsea Finn, Jiajun Wu, Joshua Tenenbaum, and Sergey Levine. Entity abstraction in visual model-based reinforcement learning. In Conference on Robot Learning, pp. 1439–1456, 2020.
  81. 81.Jacob Walker, Abhinav Gupta, and Martial Hebert. Dense optical flow prediction from a static image. In Proceedings of the IEEE International Conference on Computer Vision, 2015.
  82. 82.Nicholas Watters, Loic Matthey, Christopher P Burgess, and Alexander Lerchner. Spatial broadcast decoder: A simple architecture for learning disentangled representations in VAEs. arXiv preprint arXiv:1901.07017, 2019.
  83. 83.Marissa A Weis, Kashyap Chitta, Yash Sharma, Wieland Brendel, Matthias Bethge, Andreas Geiger, and Alexander S Ecker. Unmasking the inductive biases of unsupervised object representations for video sequences. arXiv preprint arXiv:2006.07034, 2020.
  84. 84.Yuxin Wu and Kaiming He. Group normalization. In Proceedings of the European conference on computer vision (ECCV), pp. 3–19, 2018.
  85. 85.Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, and Weidi Xie. Self-supervised video object segmentation by motion grouping. arXiv preprint arXiv:2104.07658, 2021.
  86. 86.Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum. CLEVRER: Collision events for video representation and reasoning. In International Conference on Learning Representations, 2020.
  87. 87.Yizhuo Zhang, Zhirong Wu, Houwen Peng, and Stephen Lin. A transductive approach for video object segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6949–6958, 2020.
  88. 88.Jiaojiao Zhao, Xinyu Li, Chunhui Liu, Shuai Bing, Hao Chen, Cees GM Snoek, and Joseph Tighe. TubeR: Tube-Transformer for Action Detection. arXiv preprint arXiv:2104.00969, 2021.

Citation

MLA
Kipf, T., et al. “Conditional Object-Centric Learning from Video”. arXiv, 2021, http://arxiv.org/abs/2111.12594v2.
APA
Kipf, T., Elsayed, G. F., Mahendran, A., Stone, A., Sabour, S., Heigold, G., Jonschkowski, R., Dosovitskiy, A., & Greff, K. (2021). Conditional Object-Centric Learning from Video. arXiv. http://arxiv.org/abs/2111.12594v2
Chicago
Kipf, T., G. F. Elsayed, A. Mahendran, et al. 2021. “Conditional Object-Centric Learning from Video”. arXiv. http://arxiv.org/abs/2111.12594v2.
Harvard
Kipf, T. et al. (2021) “Conditional Object-Centric Learning from Video”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.12594v2.
Vancouver
1. Kipf T, Elsayed GF, Mahendran A, Stone A, Sabour S, Heigold G, Jonschkowski R, Dosovitskiy A, Greff K (2021) Conditional Object-Centric Learning from Video. arXiv

BibTeX

@article{kipf2021conditional,
  title = {Conditional Object-Centric Learning from Video},
  author = {Kipf, Thomas and Elsayed, Gamaleldin F. and Mahendran, Aravindh and Stone, Austin and Sabour, Sara and Heigold, Georg and Jonschkowski, Rico and Dosovitskiy, Alexey and Greff, Klaus},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.12594v2},
  eprint = {2111.12594}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission