SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos

Gamaleldin F. ElsayedAravindh MahendranSjoerd van SteenkisteKlaus GreffMichael C. MozerThomas Kipf

article2022NeurIPS218 citations

Presents an end-to-end slot-based video model that leverages depth prediction and architectural scaling to achieve unsupervised object segmentation and tracking in complex real-world driving scenes.

Listen

Building automated vision systems that understand complex scenes as collections of distinct, persisting objects remains a central challenge in artificial intelligence. While humans naturally separate visual scenes into discrete entities without explicit instruction, conventional deep learning models generally rely on expensive, human-annotated segmentation masks to achieve object recognition. Prior self-supervised approaches attempted to identify objects using optical flow (motion cues), but these methods fail when objects are stationary or when the camera itself moves, limiting their practical deployment in dynamic environments like autonomous driving.

The article demonstrates an end-to-end neural network framework, called SAVi++, designed to discover, segment, and track visual objects across complex, real-world video sequences without relying on direct segmentation or tracking supervision.

To evaluate this framework, the authors conducted experiments on two major benchmarks: the synthetic Multi-Object Video (MOVi) dataset, which tests combinations of moving objects, stationary objects, and moving cameras, and the real-world Waymo Open autonomous driving dataset, consisting of high-resolution video paired with sparse LiDAR depth measurements. The approach enhances slot-based video models—which divide neural representations into distinct object pools—by training the model to predict scene depth (geometric distance) alongside upgraded visual backbones and data augmentation strategies.

The investigation produced several key findings. First, incorporating depth prediction enables the model to accurately handle complex environments with both static objects and camera motion; on the challenging MOVi-E benchmark, SAVi++ improved segmentation accuracy to 47.1% Mean Intersection over Union (mIoU), compared to 30.7% achieved by prior motion-only methods. Second, on real-world driving data from the Waymo Open dataset, SAVi++ successfully tracked and segmented objects using only sparse LiDAR signals, achieving an object recall rate above 96% and a bounding box mIoU of approximately 50%, substantially outperforming heuristic and baseline clustering methods. Third, sensitivity analyses revealed that SAVi++ maintains stable tracking performance even when up to 40 centimeters of noise is added to the depth targets, proving that dense, perfect depth supervision is not mandatory for emergent object discovery.

These findings indicate that geometric depth signals, which are readily available from hardware such as automotive LiDAR or standard depth sensors, can effectively replace labor-intensive human annotations for training object-centric models. This capability significantly reduces the cost, annotation timelines, and human error associated with curating pixel-level training datasets. Furthermore, it establishes that modular object representations can reliably emerge in real-world scenarios rather than being confined to simplified synthetic benchmarks.

For organizations developing autonomous perception or robotics pipelines, the article supports integrating multimodal geometric targets into training frameworks to reduce annotation dependence. Future development should focus on testing monocular depth estimates where LiDAR is absent, extending evaluation to unconstrained video environments with frequent object disappearances and reappearances, and gradually phasing out initial-frame object hints.

While the results demonstrate strong promise, certain limitations remain. The primary model configuration still relies on initial-frame bounding boxes as conditioning cues, and overall segmentation performance still lags behind fully supervised systems. Nevertheless, the findings provide a high-confidence proof of concept that depth-guided, self-supervised learning is a viable pathway for scalable real-world vision systems.

arXiv: 2206.07764
Cover for SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos

Abstract

The visual world can be parsimoniously characterized in terms of distinct entities with sparse interactions. Discovering this compositional structure in dynamic visual scenes has proven challenging for end-to-end computer vision approaches unless explicit instance-level supervision is provided. Slot-based models leveraging motion cues have recently shown great promise in learning to represent, segment, and track objects without direct supervision, but they still fail to scale to complex real-world multi-object videos. In an effort to bridge this gap, we take inspiration from human development and hypothesize that information about scene geometry in the form of depth signals can facilitate object-centric learning. We introduce SAVi++, an object-centric video model which is trained to predict depth signals from a slot-based video representation. By further leveraging best practices for model scaling, we are able to train SAVi++ to segment complex dynamic scenes recorded with moving cameras, containing both static and moving objects of diverse appearance on naturalistic backgrounds, without the need for segmentation supervision. Finally, we demonstrate that by using sparse depth signals obtained from LiDAR, SAVi++ is able to learn emergent object segmentation and tracking from videos in the real-world Waymo Open dataset.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Methods
  • 3.1 Background
  • 3.2 SAVi++
  • 4 Experiments
  • 4.1 SAVi++ improves object-centric learning on complex synthetic video data
  • 4.2 Ablation study
  • 4.3 SAVi++ enables emergent segmentation on real-world driving data
  • 4.4 Limitations
  • 5 Conclusion
  • 6 Acknowledgements
  • References
  • Checklist

Knowls

  1. Knowl 1 — SAVi++ Object-Centric Video Model Architecture

    model/method

    SAVi++ is an autoregressive object-centric video architecture that decomposes dynamic visual scenes into KK latent object slot vectors without requiring direct instance segmentation or tracking supervision. The model processes video sequences recurrently:

    1. Visual Encoding: For each video frame xt∈RH×W×3x_t \in \mathbb{R}^{H \times W \times 3}, an encoder generates spatial feature maps. The encoder consists of a modified ResNet-34 backbone followed by a 4-layer Transformer encoder.
    2. Slot Initialization: At time t=1t=1, KK object slot vectors {s1k}k=1K\{s_1^k\}_{k=1}^K are initialized either unconditionally via learned embedding vectors or conditionally using high-level conditioning cues, such as bounding boxes of objects in the initial video frame.
    3. Temporal Dynamics and Slot Correction: For subsequent frames t>1t > 1, a predictor module models inter-slot interactions and dynamics to produce prior slot states s^tk\hat{s}_t^k from the previous slot states st−1ks_{t-1}^k. Then, a Slot Attention corrector updates the slots by applying cross-attention between slot queries and visual encoder feature keys and values.
    4. Per-Slot Decoding: Each slot stks_t^k is independently decoded by a spatial broadcast decoder into a slot prediction y^tk\hat{y}_t^k (for depth maps and/or optical flow) and an unnormalized scalar alpha mask. Normalizing masks across slots with a softmax yields per-pixel mixture weights αtk∈[0,1]\alpha_t^k \in [0, 1] where ∑k=1Kαtk=1\sum_{k=1}^K \alpha_t^k = 1. The combined scene prediction is formed as y^t=∑k=1Kαtky^tk\hat{y}_t = \sum_{k=1}^K \alpha_t^k \hat{y}_t^k.

    Emergent instance segmentation masks and object tracking tracks across the video sequence are obtained directly from the resulting per-slot alpha masks αtk\alpha_t^k.

  2. Knowl 2 — Depth Target Representation and Loss Formulation in SAVi++

    equation

    SAVi++ utilizes camera depth signals as self-supervised reconstruction targets to overcome the inability of optical flow to represent static objects and static scenes recorded with moving cameras.

    For a ground-truth distance du,v≥0d_{u,v} \ge 0 at pixel coordinate (u,v)(u, v) from the camera, the depth target is transformed logarithmically: yu,vdepth=log⁡(1+du,v)y_{u,v}^{\text{depth}} = \log(1 + d_{u,v}) This transformation emphasizes nearby objects over distant background geometry.

    When training with sparse depth signals, such as 3D LiDAR point clouds projected into camera coordinates using sensor calibration parameters, the loss is computed exclusively over the set of pixels Ωt\Omega_t containing valid depth measurements. The per-frame loss is the mean squared error: Lt=1∣Ωt∣∑(u,v)∈Ωt∥y^t,u,v−yt,u,v∥22\mathcal{L}_t = \frac{1}{|\Omega_t|} \sum_{(u,v) \in \Omega_t} \left\| \hat{y}_{t, u, v} - y_{t, u, v} \right\|_2^2 where y^t,u,v=∑k=1Kαt,u,vky^t,u,vk\hat{y}_{t, u, v} = \sum_{k=1}^K \alpha_{t, u, v}^k \hat{y}_{t, u, v}^k is the composite reconstruction from KK slots at pixel (u,v)(u, v) weighted by alpha masks αtk\alpha_t^k. When optical flow is also available, target flow and depth maps are concatenated along the channel dimension and predicted jointly.

  3. Knowl 3 — Encoder Modifications and Video Data Augmentation for Scaled Object-Centric Learning

    model/method

    To scale slot-based object discovery to visually complex synthetic scenes and real-world driving environments, SAVi++ incorporates specific backbone scaling and data augmentation mechanisms:

    • ResNet-34 Backbone Modifications: The visual encoder uses a ResNet-34 architecture modified to maintain high spatial feature resolution. In the root convolutional block of ResNet, standard max-pooling is removed and a stride of 1 is employed, reducing the overall stride of the convolutional backbone from 32 down to 8. Standard Batch Normalization layers are replaced with Group Normalization across all residual blocks to eliminate dependencies on batch and temporal statistics.
    • Transformer Encoder: A 4-layer Transformer encoder is applied on top of the ResNet-34 feature map to incorporate global context across spatial locations before slot attention.
    • Consistent Video Crop Augmentation: Video inputs are augmented using Inception-style random spatial cropping. A bounding crop with aspect ratio sampled uniformly from [0.75,1.33][0.75, 1.33] is extracted such that a sufficient fraction of the image is preserved. The identical spatial crop is applied synchronously across all frames in a video sequence and to the corresponding prediction target fields (depth maps and flow fields), which are then resized back to the input resolution.
  4. Knowl 4 — Segmentation and Tracking Evaluation on Multi-Object Video (MOVi) Benchmarks

    data/table

    SAVi++ was evaluated on 24-frame validation sequences from three synthetic Multi-Object Video (MOVi) benchmark datasets: MOVi-C (static camera, moving objects only), MOVi-D (static camera, 1–3 moving objects, 10–20 static objects), and MOVi-E (moving camera, moving and static objects). All models were conditioned on ground-truth bounding box coordinates in the first frame. Performance was evaluated using Mean Intersection over Union (mIoU) and Foreground Adjusted Rand Index (FG-ARI).

    Model mIoU (%) ↑\uparrow FG-ARI (%) ↑\uparrow
    MOVi-C MOVi-D MOVi-E MOVi-C MOVi-D MOVi-E
    BBox copy 12.3 42.8 32.9 11.8 68.0 54.7
    BBox propagation 22.9±0.122.9 \pm 0.1 26.7±0.826.7 \pm 0.8 24.1±1.124.1 \pm 1.1 9.6±0.59.6 \pm 0.5 24.9±3.724.9 \pm 3.7 18.4±3.918.4 \pm 3.9
    K-Means (depth) 7.1±0.37.1 \pm 0.3 6.0±0.46.0 \pm 0.4 5.4±0.35.4 \pm 0.3 26.3±1.026.3 \pm 1.0 30.9±0.730.9 \pm 0.7 32.2±0.632.2 \pm 0.6
    K-Means (flow) 10.7±0.510.7 \pm 0.5 7.4±0.47.4 \pm 0.4 6.0±0.36.0 \pm 0.3 26.5±1.026.5 \pm 1.0 30.9±0.830.9 \pm 0.8 33.1±0.733.1 \pm 0.7
    K-Means (flow+depth) 10.6±0.610.6 \pm 0.6 6.7±0.46.7 \pm 0.4 5.3±0.35.3 \pm 0.3 26.6±1.026.6 \pm 1.0 35.9±1.035.9 \pm 1.0 34.8±0.734.8 \pm 0.7
    CRW 27.8±0.227.8 \pm 0.2 45.3±0.045.3 \pm 0.0 47.5±0.147.5 \pm 0.1 – – –
    SAVi 43.1±0.743.1 \pm 0.7 22.7±7.522.7 \pm 7.5 30.7±4.930.7 \pm 4.9 77.6±0.777.6 \pm 0.7 59.6±6.759.6 \pm 6.7 55.3±5.855.3 \pm 5.8
    SAVi++ (ours) 45.2±0.1\mathbf{45.2 \pm 0.1} 48.3±0.5\mathbf{48.3 \pm 0.5} 47.1±1.3\mathbf{47.1 \pm 1.3} 81.9±0.2\mathbf{81.9 \pm 0.2} 86.0±0.3\mathbf{86.0 \pm 0.3} 84.1±0.9\mathbf{84.1 \pm 0.9}

    Values indicate mean ±\pm standard error across 5 random seeds. While the optical-flow-based SAVi fails on datasets containing static objects (MOVi-D) and camera motion (MOVi-E), SAVi++ maintains high segmentation and tracking performance across all dynamic settings, outperforming non-visual baselines, feature-propagation baselines (CRW), and baseline clustering.

  5. Knowl 5 — Emergent Object Tracking and Segmentation on the Waymo Open Dataset

    data/table

    SAVi++ was evaluated on real-world driving sequences from the Waymo Open dataset subsampled at 5 fps. Models were trained on 6-frame clips using 11 slots and sparse LiDAR depth targets, and evaluated on 10-frame sequences. Tracking performance was assessed using normalized Center-of-Mass (CoM) distance between predicted mask centroids and ground-truth bounding box centers (lower is better), Bounding Box mIoU (B. mIoU) via a readout MLP on slot latents, and Bounding Box Recall (B. Recall).

    Model CoM (%) ↓\downarrow B. mIoU (%) ↑\uparrow B. Recall (%) ↑\uparrow
    BBox Copy 5.0 44.3 100.0
    BBox Prop. 5.1±0.15.1 \pm 0.1 38.5±0.538.5 \pm 0.5 100.0
    K-Means (depth) 13.0±0.113.0 \pm 0.1 – 100.0
    SAVi (RGB) 21.5±1.821.5 \pm 1.8 7.9±0.97.9 \pm 0.9 95.8±2.795.8 \pm 2.7
    SAVi (depth) 24.7±0.724.7 \pm 0.7 10.3±2.410.3 \pm 2.4 97.4±0.697.4 \pm 0.6
    SAVi++ 4.4±0.24.4 \pm 0.2 49.7±0.749.7 \pm 0.7 96.5±0.796.5 \pm 0.7
    SAVi++ HR (256×384256 \times 384) 3.9±0.1\mathbf{3.9 \pm 0.1} 51.9±0.4\mathbf{51.9 \pm 0.4} 96.2±0.496.2 \pm 0.4
    Supervised 1.1±0.01.1 \pm 0.0 67.6±0.667.6 \pm 0.6 –

    Scores represent mean ±\pm standard error across 3 seeds. Standard models operated on 128×192128 \times 192 frame resolution, while SAVi++ HR operated on 256×384256 \times 384. SAVi++ significantly outperforms previous self-supervised object-centric methods and heuristic baselines, demonstrating that sparse LiDAR depth targets alone are sufficient to drive emergent object segmentation and tracking in real-world driving environments.

  6. Knowl 6 — Ablation Analysis of SAVi++ Core Components

    empirical result

    Ablation experiments on the Multi-Object Video datasets (MOVi-C, MOVi-D, and MOVi-E) reveal the individual contributions of depth targets, data augmentation, and encoder architecture:

    • Depth Targets: Removing depth targets (relying solely on optical flow prediction) causes substantial performance drops on datasets with static objects and camera movement. On MOVi-E, removing depth targets reduces mIoU from approximately 47.1% to under 25%. Conversely, removing optical flow targets and training solely on depth targets retains high performance across MOVi-D and MOVi-E, demonstrating that depth alone is a sufficient and effective supervisory target.
    • Data Augmentation: Omitting Inception-style video crop augmentation has negligible effect on simple moving-object scenes (MOVi-C) but causes significant mIoU degradation on MOVi-D and MOVi-E, showing that strong data augmentation is necessary when scaling object-centric models to complex scenes.
    • Transformer Encoder: Removing the 4-layer Transformer encoder from the backbone reduces mIoU moderately across all MOVi variants, confirming that added contextual capacity improves slot allocation.
  7. Knowl 7 — Unconditional Object Discovery in Real-World Driving Video

    empirical result

    In an unconditional setting on the Waymo Open driving dataset—where slots are initialized with learned parameter vectors rather than conditioned on first-frame bounding boxes—SAVi++ successfully discovers and segments objects in an unsupervised manner.

    When evaluated over 12-frame sequences at test time using Hungarian matching against ground-truth bounding box centroids:

    • Unconditional SAVi++ achieves a normalized Center-of-Mass (CoM) distance error of 6.9±0.5%6.9 \pm 0.5\%.
    • The non-autoregressive baseline SIMONe trained with the same sparse depth target loss achieves a CoM distance error of 7.4±0.2%7.4 \pm 0.2\%, whereas standard SIMONe trained on RGB frame reconstruction fails to segment objects.

    During video tracking, slots consistently follow individual objects until they exit the scene, at which point the freed slots dynamically re-bind to new or previously unassigned visual entities.

  8. Knowl 8 — Robustness of SAVi++ to Additive Depth Target Noise

    empirical result

    SAVi++ does not require precise or dense depth measurements to learn object-centric representations. When trained on the Waymo Open dataset with artificial zero-mean additive Gaussian noise injected into the ground-truth sparse LiDAR depth measurements at standard deviations of σ=10cm\sigma = 10\text{cm}, σ=20cm\sigma = 20\text{cm}, and σ=40cm\sigma = 40\text{cm}, SAVi++ maintains its emergent object tracking and segmentation performance without significant degradation, retaining robust object emergence even at σ=40cm\sigma = 40\text{cm}.

  9. Knowl 9 — Stated Limitations of SAVi++

    limitation

    SAVi++ exhibits five explicit limitations:

    1. Reliance on Conditioning Cues: Optimal segmentation and tracking performance relies on bounding box hints in the initial video frame; although unconditional discovery is possible, it exhibits lower temporal consistency and mask fidelity.
    2. Requirement of Depth Signals: The method relies on depth targets (obtained via LiDAR or RGB-D sensors) for training, limiting direct application to arbitrary in-the-wild web video collections lacking 3D geometry information.
    3. Lack of Long-Term Re-identification: SAVi++ does not maintain explicit persistent memory or presence state variables for objects that temporarily leave and re-enter the camera view, causing freed slots to re-bind to other objects.
    4. Structured Domain Focus: Validation is primarily conducted on structured driving environments (Waymo Open) and rigid-body simulations (MOVi); unconstrained datasets with complex articulated objects and occlusions (such as DAVIS or Kinetics) present unaddressed scaling challenges.
    5. Performance Gap to Supervised Models: The unsupervised emergent tracking and segmentation metrics remain quantitatively below fully supervised object tracking methods.

Coverage note — No substantial contributed material was omitted. All architectural developments, loss formulations, synthetic MOVi benchmark results, real-world Waymo Open results, ablation experiments, noise sensitivity analyses, and stated limitations are covered.

References

  1. 1.Zhipeng Bao, Pavel Tokmakov, Allan Jabri, Yu-Xiong Wang, Adrien Gaidon, and Martial Hebert. Discovering objects that can move. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  2. 2.Daniel M Bear, Chaofei Fan, Damian Mrowca, Yunzhu Li, Seth Alter, Aran Nayebi, Jeremy Schwartz, Li Fei-Fei, Jiajun Wu, Joshua B Tenenbaum, et al. Learning physical graph representations from visual scenes. In Advances in Neural Information Processing Systems, 2020.
  3. 3.Thomas Brox and Jitendra Malik. Object segmentation by long term analysis of point trajectories. In European Conference on Computer Vision, pages 282–295. Springer, 2010.
  4. 4.Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner. MONet: Unsupervised scene decomposition and representation. arXiv preprint arXiv:1901.11390, 2019.
  5. 5.Sergi Caelles, Jordi Pont-Tuset, Federico Perazzi, Alberto Montes, Kevis-Kokitsi Maninis, and Luc Van Gool. The 2019 DAVIS challenge on vos: Unsupervised multi-object segmentation. arXiv preprint arXiv:1905.00737, 2019.
  6. 6.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, 2020.
  7. 7.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, pages 213–229. Springer, 2020.
  8. 8.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021.
  9. 9.J. Driver, P. McLeod, and Z. Dienes. Motion coherence and conjunction search: Implications for guided search theory. Perception and Psychophysics, 51:79–85, 1992.
  10. 10.J. T. Enns and R. A. Rensink. Influence of scene-based properties on visual search. Science, 247:721–723, 1990.
  11. 11.J. T. Enns and R. A. Rensink. Sensitivity to three-dimensional orientation in visual search. Psychological Science, 1(5):323–326, 1990.
  12. 12.Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. Deep ordinal regression network for monocular depth estimation. In CVPR, June 2018.
  13. 13.Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber. Neural expectation maximization. In Advances in Neural Information Processing Systems, pages 6691–6701, 2017.
  14. 14.Klaus Greff, Raphaël Lopez Kaufman, Rishabh Kabra, Nick Watters, Christopher Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, and Alexander Lerchner. Multi-object representation learning with iterative variational inference. In International Conference on Machine Learning, pages 2424–2433, 2019.
  15. 15.Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber. On the binding problem in artificial neural networks. arXiv preprint arXiv:2012.05208, 2020.
  16. 16.Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, Thomas Kipf, Abhijit Kundu, Dmitry Lagun, Issam Laradji, Hsueh-Ti (Derek) Liu, Henning Meyer, Yishu Miao, Derek Nowrouzezahrai, Cengiz Oztireli, Etienne Pot, Noha Radwan, Daniel Rebain, Sara Sabour, Mehdi S. M. Sajjadi, Matan Sela, Vincent Sitzmann, Austin Stone, Deqing Sun, Suhani Vora, Ziyu Wang, Tianhao Wu, Kwang Moo Yi, Fangcheng Zhong, and Andrea Tagliasacchi. Kubric: a scalable dataset generator. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  17. 17.Adam W Harley, Yiming Zuo, Jing Wen, Ayush Mangal, Shubhankar Potdar, Ritwick Chaudhry, and Katerina Fragkiadaki. Track, check, repeat: An EM approach to unsupervised tracking. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  19. 19.Olivier J Hénaff, Skanda Koppula, Evan Shelhamer, Daniel Zoran, Andrew Jaegle, Andrew Zisserman, João Carreira, and Relja Arandjelovic´. Object discovery and representation networks. arXiv preprint arXiv:2203.08777, 2022.
  20. 20.Todd S Horowitz, Jeremy M Wolfe, Jennifer S DiMase, and Sarah B Klieger. Visual search for type of motion is based on simple motion primitives. Perception, 36(11):1624–1634, 2007.
  21. 21.Lawrence Hubert and Phipps Arabie. Comparing partitions. Journal of Classification, 2(1):193–218, 1985.
  22. 22.Allan Jabri, Andrew Owens, and Alexei A Efros. Space-time correspondence as a contrastive random walk. In Advances in Neural Information Processing Systems, 2020.
  23. 23.Jindong Jiang, Sepehr Janghorbani, Gerard de Melo, and Sungjin Ahn. SCALOR: Generative world models with scalable object representations. In International Conference on Learning Representations, 2020.
  24. 24.Rishabh Kabra, Daniel Zoran, Loic Matthey Goker Erdogan, Antonia Creswell, Matthew Botvinick, Alexander Lerchner, and Christopher P. Burgess. SIMONe: View-invariant, temporally-abstracted object representations via unsupervised video decomposition. In Advances in Neural Information Processing Systems, 2021.
  25. 25.Daniel Kahneman, Anne Treisman, and Brian J Gibbs. The reviewing of object files: Object-specific integration of information. Cognitive psychology, 24(2):175–219, 1992.
  26. 26.Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion. MDETR – Modulated Detection for End-to-End Multi-Modal Understanding. arXiv preprint arXiv:2104.12763, 2021.
  27. 27.Laurynas Karazija, Iro Laina, and Christian Rupprecht. ClevrTex: A texture-rich benchmark for unsupervised multi-object segmentation. In NeurIPS Track on Datasets and Benchmarks, 2021.
  28. 28.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017.
  29. 29.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015.
  30. 30.Thomas Kipf, Elise van der Pol, and Max Welling. Contrastive learning of structured world models. In International Conference on Learning Representations, 2020.
  31. 31.Thomas Kipf, Gamaleldin F Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Greff. Conditional object-centric learning from video. In International Conference on Learning Representations, 2022.
  32. 32.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (BiT): General visual representation learning. In European Conference on Computer Vision, pages 491–507. Springer, 2020.
  33. 33.Adam Kosiorek, Hyunjik Kim, Yee Whye Teh, and Ingmar Posner. Sequential attend, infer, repeat: Generative modelling of moving objects. In Advances in Neural Information Processing Systems, pages 8606–8616, 2018.
  34. 34.Hamid Laga, Laurent Valentin Jospin, Farid Boussaid, and Mohammed Bennamoun. A survey on deep learning techniques for stereo-based depth estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020.
  35. 35.Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. Building machines that learn and think like people. Behavioral and brain sciences, 40, 2017.
  36. 36.Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. Object-centric learning with slot attention. In Advances in Neural Information Processing Systems, 2020.
  37. 37.Sindy Löwe, Klaus Greff, Rico Jonschkowski, Alexey Dosovitskiy, and Thomas Kipf. Learning object-centric video models by contrasting sets. arXiv preprint arXiv:2011.10287, 2020.
  38. 38.Yue Ming, Xuyang Meng, Chunxiao Fan, and Hui Yu. Deep learning for monocular depth estimation: A review. Neurocomputing, 2021.
  39. 39.K. Nakayama and G. H. Silverman. Serial and parallel processing of visual feature conjunctions. Nature, 320:264–265, 1986.
  40. 40.Peter Ochs and Thomas Brox. Object segmentation in video: a hierarchical variational approach for turning point trajectories into dense regions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2011.
  41. 41.Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 724–732, 2016.
  42. 42.William M Rand. Objective criteria for the evaluation of clustering methods. Journal of the American Statistical Association, 66(336):846–850, 1971.
  43. 43.René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In ICCV, October 2021.
  44. 44.Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio. Toward causal representation learning. Proceedings of the IEEE, 109(5): 612–634, 2021.
  45. 45.E. S. Spelke. Principles of object perception. Cognitive Science, 14:29–56, 1990.
  46. 46.Elizabeth S Spelke and Katherine D Kinzler. Core knowledge. Developmental science, 10(1):89–96, 2007.
  47. 47.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  48. 48.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In CVPR, pages 1–9, 2015.
  49. 49.Hao Tian, Yuntao Chen, Jifeng Dai, Zhaoxiang Zhang, and Xizhou Zhu. Unsupervised object detection with lidar clues. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  50. 50.Sjoerd van Steenkiste, Michael Chang, Klaus Greff, and Jürgen Schmidhuber. Relational neural expectation maximization: Unsupervised discovery of objects and their interactions. In International Conference on Learning Representations, 2018.
  51. 51.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, pages 5998–6008, 2017.
  52. 52.Rishi Veerapaneni, John D Co-Reyes, Michael Chang, Michael Janner, Chelsea Finn, Jiajun Wu, Joshua Tenenbaum, and Sergey Levine. Entity abstraction in visual model-based reinforcement learning. In Conference on Robot Learning, pages 1439–1456, 2020.
  53. 53.Antonin Vobecky, David Hurych, Oriane Siméoni, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, and Josef Sivic. Drive&segment: Unsupervised semantic segmentation of urban scenes via cross-modal distillation. arXiv preprint arXiv:2203.11160, 2022.
  54. 54.Lijun Wang, Jianming Zhang, Oliver Wang, Zhe Lin, and Huchuan Lu. SDC-depth: Semantic divide-and-conquer network for monocular depth estimation. In CVPR, 2020.
  55. 55.Yuxin Wu and Kaiming He. Group normalization. In European Conference on Computer Vision (ECCV), pages 3–19, 2018.
  56. 56.Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. arXiv preprint arXiv:2202.11094, 2022.
  57. 57.Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, and Weidi Xie. Self-supervised video object segmentation by motion grouping. In Proceedings of the IEEE International Conference on Computer Vision, 2021.
  58. 58.Yi Zhou, Hui Zhang, Hana Lee, Shuyang Sun, Pingjun Li, Yangguang Zhu, ByungIn Yoo, Xiaojuan Qi, and Jae-Joon Han. Slot-vps: Object-centric representation learning for video panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3093–3103, 2022.
  59. 59.Daniel Zoran, Rishabh Kabra, Alexander Lerchner, and Danilo J Rezende. Parts: Unsupervised segmentation with slots, attention and independence maximization. In Proceedings of the IEEE International Conference on Computer Vision, pages 10439–10447, 2021.

Citation

MLA
Elsayed, G., et al. “SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 28940–54, https://proceedings.neurips.cc/paper_files/paper/2022/file/ba1a6ba05319e410f0673f8477a871e3-Paper-Conference.pdf.
APA
Elsayed, G., Mahendran, A., van Steenkiste, S., Greff, K., Mozer, M. C., & Kipf, T. (2022). SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos. Advances in Neural Information Processing Systems, 35, 28940–28954. https://proceedings.neurips.cc/paper_files/paper/2022/file/ba1a6ba05319e410f0673f8477a871e3-Paper-Conference.pdf
Chicago
Elsayed, G., A. Mahendran, S. van Steenkiste, K. Greff, M. C. Mozer, and T. Kipf. 2022. “SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos”. Advances in Neural Information Processing Systems 35: 28940–54. https://proceedings.neurips.cc/paper_files/paper/2022/file/ba1a6ba05319e410f0673f8477a871e3-Paper-Conference.pdf.
Harvard
Elsayed, G. et al. (2022) “SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 28940–28954. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/ba1a6ba05319e410f0673f8477a871e3-Paper-Conference.pdf.
Vancouver
1. Elsayed G, Mahendran A, van Steenkiste S, Greff K, Mozer MC, Kipf T (2022) SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 28940–28954

BibTeX

@inproceedings{elsayed2022savi,
  title = {SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos},
  author = {Elsayed, Gamaleldin and Mahendran, Aravindh and van Steenkiste, Sjoerd and Greff, Klaus and Mozer, Michael C. and Kipf, Thomas},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {28940-28954},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/ba1a6ba05319e410f0673f8477a871e3-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission