SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos
Gamaleldin F. ElsayedAravindh MahendranSjoerd van SteenkisteKlaus GreffMichael C. MozerThomas Kipf
Presents an end-to-end slot-based video model that leverages depth prediction and architectural scaling to achieve unsupervised object segmentation and tracking in complex real-world driving scenes.
Building automated vision systems that understand complex scenes as collections of distinct, persisting objects remains a central challenge in artificial intelligence. While humans naturally separate visual scenes into discrete entities without explicit instruction, conventional deep learning models generally rely on expensive, human-annotated segmentation masks to achieve object recognition. Prior self-supervised approaches attempted to identify objects using optical flow (motion cues), but these methods fail when objects are stationary or when the camera itself moves, limiting their practical deployment in dynamic environments like autonomous driving.
The article demonstrates an end-to-end neural network framework, called SAVi++, designed to discover, segment, and track visual objects across complex, real-world video sequences without relying on direct segmentation or tracking supervision.
To evaluate this framework, the authors conducted experiments on two major benchmarks: the synthetic Multi-Object Video (MOVi) dataset, which tests combinations of moving objects, stationary objects, and moving cameras, and the real-world Waymo Open autonomous driving dataset, consisting of high-resolution video paired with sparse LiDAR depth measurements. The approach enhances slot-based video models—which divide neural representations into distinct object pools—by training the model to predict scene depth (geometric distance) alongside upgraded visual backbones and data augmentation strategies.
The investigation produced several key findings. First, incorporating depth prediction enables the model to accurately handle complex environments with both static objects and camera motion; on the challenging MOVi-E benchmark, SAVi++ improved segmentation accuracy to 47.1% Mean Intersection over Union (mIoU), compared to 30.7% achieved by prior motion-only methods. Second, on real-world driving data from the Waymo Open dataset, SAVi++ successfully tracked and segmented objects using only sparse LiDAR signals, achieving an object recall rate above 96% and a bounding box mIoU of approximately 50%, substantially outperforming heuristic and baseline clustering methods. Third, sensitivity analyses revealed that SAVi++ maintains stable tracking performance even when up to 40 centimeters of noise is added to the depth targets, proving that dense, perfect depth supervision is not mandatory for emergent object discovery.
These findings indicate that geometric depth signals, which are readily available from hardware such as automotive LiDAR or standard depth sensors, can effectively replace labor-intensive human annotations for training object-centric models. This capability significantly reduces the cost, annotation timelines, and human error associated with curating pixel-level training datasets. Furthermore, it establishes that modular object representations can reliably emerge in real-world scenarios rather than being confined to simplified synthetic benchmarks.
For organizations developing autonomous perception or robotics pipelines, the article supports integrating multimodal geometric targets into training frameworks to reduce annotation dependence. Future development should focus on testing monocular depth estimates where LiDAR is absent, extending evaluation to unconstrained video environments with frequent object disappearances and reappearances, and gradually phasing out initial-frame object hints.
While the results demonstrate strong promise, certain limitations remain. The primary model configuration still relies on initial-frame bounding boxes as conditioning cues, and overall segmentation performance still lags behind fully supervised systems. Nevertheless, the findings provide a high-confidence proof of concept that depth-guided, self-supervised learning is a viable pathway for scalable real-world vision systems.
- Paper: Unsupervised Learning of Depth and Ego-Motion from Video, Tinghui Zhou et al. (2017). It provides the foundational framework for unsupervised single-view depth estimation and ego-motion learning from monocular video sequences.
- Paper: Digging Into Self-Supervised Monocular Depth Estimation, Clément Godard et al. (2018). It establishes essential self-supervised loss formulations, such as auto-masking and minimum reprojection, for learning depth and motion from unlabeled video.
- Paper: Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D, Jonah Philion et al. (2020). It establishes key principles for unprojecting camera features into 3D geometric representations for autonomous driving scene understanding.
- Paper: The 2017 DAVIS Challenge on Video Object Segmentation, Jordi Pont-Tuset et al. (2017). It presents the benchmark and formulation for multi-target video object segmentation under real-world occlusions and complex camera motions.
- Paper: Learning Rich Features from RGB-D Images for Object Detection and Segmentation, Saurabh Gupta et al. (2014). It demonstrates how incorporating geometric depth features enables neural architectures to segment and detect visual objects more effectively.
- Paper: Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations, V. Sitzmann et al. (2019). It introduces continuous structure-aware 3D scene representations learned directly from multi-view images without dense 3D ground truth.
- Paper: DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis, Yinghao Xu et al. (2023). It extends object-centric spatial scene decomposition to controllable 3D-aware radiance field generation on complex driving and indoor environments.
- Paper: D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video, Tianhao Wu et al. (2022). It applies self-supervised 3D decomposition to decouple dynamic moving objects from static backgrounds in monocular video.
- Paper: Behind the Scenes: Density Fields for Single View Reconstruction, Felix Wimbauer et al. (2023). It advances single-view geometric video learning by reconstructing complete 3D density fields and reasoning about occluded scene structure.
- Paper: Symphonize 3D Semantic Scene Completion with Contextual Instance Queries, Haoyi Jiang et al. (2024). It integrates contextual instance queries to coordinate 2D-to-3D geometric reconstruction and semantic occupancy prediction in driving scenes.
- Paper: Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking, Ziqi Pang et al. (2023). It builds upon multi-camera spatio-temporal video modeling to track dynamic 3D visual entities across past and future frames.
- Paper: BEVNeXt: Reviving Dense BEV Frameworks for 3D Object Detection, Zhenxin Li et al. (2024). It refines dense geometric depth modeling and temporal fusion for multi-camera 3D visual perception in autonomous driving.
- Paper: GARField: Group Anything with Radiance Fields, Chung Min Kim et al. (2024). It scales 3D visual grouping by learning continuous affinity fields across multi-scale hierarchies in radiance fields.
- Paper: SAM 2: Segment Anything in Images and Videos, Nikhila Ravi et al. (2025). It scales foundation-level spatio-temporal video object segmentation across dynamic real-world videos.
