Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames
Ondrej BizaSjoerd van SteenkisteMehdi S. M. SajjadiGamaleldin Fathy ElsayedAravindh MahendranThomas Kipf
Introduces a mechanism that incorporates per-object spatial symmetries into Slot Attention by dynamically transforming position encodings into slot-centric reference frames, significantly improving data efficiency and unsupervised object discovery across synthetic benchmarks and real-world driving data.
Enabling autonomous vision systems to break down raw, unannotated visual scenes into individual object components is a foundational challenge in artificial intelligence. Existing unsupervised object-centric neural networks, such as Slot Attention, learn object representations in a self-supervised manner but typically operate in absolute coordinate space. Consequently, they fail to leverage natural spatial symmetries, entangling object appearance with spatial factors like position and scale, which results in sample inefficiency and poor generalization to unseen environments.
The article introduces and evaluates Invariant Slot Attention (ISA), a framework designed to build per-object reference frames into the attention and reconstruction pipelines. Its primary objective is to demonstrate that incorporating translation, scaling, and rotation symmetries directly into slot-based architectures substantially improves unsupervised object discovery, data efficiency, and out-of-domain generalization without introducing major computational overhead.
The researchers assessed the framework using empirical evaluations on both synthetic and real-world multi-object vision datasets, including Tetrominoes, Objects Room, CLEVRTex, MultiShapeNet, and the Waymo Open driving dataset. The technical approach alters standard absolute positional encodings into pose-relative encodings derived directly from attention masks in both the encoder's iterative attention mechanism and the spatial broadcast decoder. The evaluations compared standard baseline architectures against invariant variants using the Adjusted Rand Index for foreground object segmentation (FG-ARI) and reconstruction mean squared error across multiple random seeds.
The experimental findings demonstrate major advantages of incorporating spatial symmetries. First, translation invariance halved to quartered the training data required on the Tetrominoes benchmark, delivering equivalent segmentation performance with 2x to 4x fewer samples. Second, on textured and complex synthetic datasets, translation and scale invariance generated double-digit performance gains with standard backbones, lifting CLEVRTex segmentation by over 24 percentage points (78.8% vs. 54.5%) and boosting Objects Room and MultiShapeNet segmentation. Third, on real-world driving data from the Waymo Open dataset, translation and scale invariance improved single-frame RGB segmentation from 27.6% to 39.8% FG-ARI. Finally, while translation and scaling produced consistent gains, two-dimensional rotation invariance provided mixed outcomes, yielding minor improvements in textured environments but degrading performance on others due to heuristic ambiguities with symmetric shapes.
These findings indicate that embedding geometric equivariance into model architectures serves as an effective inductive bias that outperforms conventional whole-image data augmentation. For decision-makers and system architects, this translates to faster model convergence, lower labeling and compute requirements, and more reliable perception in novel deployment settings. The results also show that disentangling object appearance from location allows direct manipulation and control over individual object attributes, which is valuable for downstream planning, robotics, and generative tasks.
Organizations developing perception pipelines should consider adopting pose-relative position encodings in attention-based vision models, particularly prioritizing translation and scale parameters. However, teams should avoid relying on simple heuristic rotation estimators in complex three-dimensional scenes. Future development efforts should focus on extending invariant slot mechanisms to native three-dimensional coordinate spaces, incorporating dedicated background segmentation models to separate static background textures from dynamic objects, and refining rotational corrections.
The conclusions are well-supported across rigorous synthetic and controlled benchmarks, giving high confidence in the sample efficiency and generalization improvements for two-dimensional spatial symmetries. Readers should maintain caution when applying the method directly to highly dynamic, three-dimensional real-world scenes, as natural viewpoints, lighting shifts, and non-rigid object deformations break exact two-dimensional planar symmetries.
- Paper: SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos, Gamaleldin F. Elsayed et al. (2022). SAVi++ establishes how slot-based object discovery is evaluated on Waymo driving scenes, preparing you for this paper’s real-world slot-segmentation experiments.
- Paper: Relation Networks for Object Detection, Han Hu et al. (2017). Its relative position-and-scale attention provides an earlier example of encoding object geometry directly into attention, a key design idea behind slot-centric reference frames.
- Paper: Conditional Object-Centric Learning from Video, Thomas Kipf et al. (2022). SAVi introduces video-based Slot Attention and its object-slot representations, giving useful architectural grounding for the source’s slot-based approach.
No sufficiently relevant recommendations were found.
