Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames

Ondrej BizaSjoerd van SteenkisteMehdi S. M. SajjadiGamaleldin Fathy ElsayedAravindh MahendranThomas Kipf

article2023ICML63 citations

Introduces a mechanism that incorporates per-object spatial symmetries into Slot Attention by dynamically transforming position encodings into slot-centric reference frames, significantly improving data efficiency and unsupervised object discovery across synthetic benchmarks and real-world driving data.

Listen

Enabling autonomous vision systems to break down raw, unannotated visual scenes into individual object components is a foundational challenge in artificial intelligence. Existing unsupervised object-centric neural networks, such as Slot Attention, learn object representations in a self-supervised manner but typically operate in absolute coordinate space. Consequently, they fail to leverage natural spatial symmetries, entangling object appearance with spatial factors like position and scale, which results in sample inefficiency and poor generalization to unseen environments.

The article introduces and evaluates Invariant Slot Attention (ISA), a framework designed to build per-object reference frames into the attention and reconstruction pipelines. Its primary objective is to demonstrate that incorporating translation, scaling, and rotation symmetries directly into slot-based architectures substantially improves unsupervised object discovery, data efficiency, and out-of-domain generalization without introducing major computational overhead.

The researchers assessed the framework using empirical evaluations on both synthetic and real-world multi-object vision datasets, including Tetrominoes, Objects Room, CLEVRTex, MultiShapeNet, and the Waymo Open driving dataset. The technical approach alters standard absolute positional encodings into pose-relative encodings derived directly from attention masks in both the encoder's iterative attention mechanism and the spatial broadcast decoder. The evaluations compared standard baseline architectures against invariant variants using the Adjusted Rand Index for foreground object segmentation (FG-ARI) and reconstruction mean squared error across multiple random seeds.

The experimental findings demonstrate major advantages of incorporating spatial symmetries. First, translation invariance halved to quartered the training data required on the Tetrominoes benchmark, delivering equivalent segmentation performance with 2x to 4x fewer samples. Second, on textured and complex synthetic datasets, translation and scale invariance generated double-digit performance gains with standard backbones, lifting CLEVRTex segmentation by over 24 percentage points (78.8% vs. 54.5%) and boosting Objects Room and MultiShapeNet segmentation. Third, on real-world driving data from the Waymo Open dataset, translation and scale invariance improved single-frame RGB segmentation from 27.6% to 39.8% FG-ARI. Finally, while translation and scaling produced consistent gains, two-dimensional rotation invariance provided mixed outcomes, yielding minor improvements in textured environments but degrading performance on others due to heuristic ambiguities with symmetric shapes.

These findings indicate that embedding geometric equivariance into model architectures serves as an effective inductive bias that outperforms conventional whole-image data augmentation. For decision-makers and system architects, this translates to faster model convergence, lower labeling and compute requirements, and more reliable perception in novel deployment settings. The results also show that disentangling object appearance from location allows direct manipulation and control over individual object attributes, which is valuable for downstream planning, robotics, and generative tasks.

Organizations developing perception pipelines should consider adopting pose-relative position encodings in attention-based vision models, particularly prioritizing translation and scale parameters. However, teams should avoid relying on simple heuristic rotation estimators in complex three-dimensional scenes. Future development efforts should focus on extending invariant slot mechanisms to native three-dimensional coordinate spaces, incorporating dedicated background segmentation models to separate static background textures from dynamic objects, and refining rotational corrections.

The conclusions are well-supported across rigorous synthetic and controlled benchmarks, giving high confidence in the sample efficiency and generalization improvements for two-dimensional spatial symmetries. Readers should maintain caution when applying the method directly to highly dynamic, three-dimensional real-world scenes, as natural viewpoints, lighting shifts, and non-rigid object deformations break exact two-dimensional planar symmetries.

No sufficiently relevant recommendations were found.

Cover for Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames

Abstract

Automatically discovering composable abstractions from raw perceptual data is a long-standing challenge in machine learning. Recent slot-based neural networks that learn about objects in a self-supervised manner have made exciting progress in this direction. However, they typically fall short at adequately capturing spatial symmetries present in the visual world, which leads to sample inefficiency, such as when entangling object appearance and pose. In this paper, we present a simple yet highly effective method for incorporating spatial symmetries via slot-centric reference frames. We incorporate equivariance to per-object pose transformations into the attention and generation mechanism of Slot Attention by translating, scaling, and rotating position encodings. These changes result in little computational overhead, are easy to implement, and can result in large gains in terms of data efficiency and overall improvements to object discovery. We evaluate our method on a wide range of synthetic object discovery benchmarks namely Tetrominoes, CLEVR-Tex, Objects Room and MultiShapeNet, and show promising improvements on the challenging real-world Waymo Open dataset.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background: Slot Attention
  • 4. Invariant Slot Attention
  • 4.1. Translation and Scaling Invariant Slot Attention
  • 4.2. Invariance to Rotations
  • 5. Experiments
  • 5.1. Proof of concept: Tetrominoes
  • 5.2. Evaluating translation and scaling invariance
  • 5.3. Invariance to rotations
  • 5.4. Real-world evaluation: Waymo Open
  • 5.5. Symmetry breaking
  • 5.6. Ablations
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Limitations
  • B. Pseudocode
  • C. Model details
  • C.1. Architecture details
  • C.2. Rotation estimation
  • D. Optimization details
  • E. Evaluation metrics
  • F. Dataset details
  • F.1. Waymo Open
  • G. Additional experimental results
  • G.1. Invariance to rotations
  • G.2. CLEVR dataset
  • G.3. Objects Room dataset
  • G.4. MultiShapeNet
  • G.5. CLEVRTex results
  • G.6. Waymo Open results

Knowls

  1. Knowl 1 — Slot-centric reference frames separate object appearance from pose

    model/method

    Invariant Slot Attention (ISA) extends Slot Attention by assigning each slot an object-specific reference frame. A slot’s latent vector represents appearance, while its position and scale (and, in the rotational variant, orientation) represent pose. The model uses pose-relative coordinates when processing image tokens, so the same object can be processed in a canonical frame at different image locations and scales. Ideally, the appearance representation is invariant to those pose changes, while the estimated reference frame changes equivariantly with them. Slot scale denotes the object’s spatial extent in the image and can reflect both object size and distance from the camera. The intended invariance is conditional: absolute-position information already present in visual features, or in some slot initializations, can break it.

  2. Knowl 2 — Pose-relative attention estimates slot positions and scales from attention masks

    model/method

    Let an encoder produce NN image tokens xnx_n, each with an absolute 2D coordinate un∈[−1,1]2u_n\in[-1,1]^2, and let there be KK slots with latent vectors zkz_k, positions pk∈R2p_k\in\mathbb{R}^2, and positive componentwise scales sk∈R2s_k\in\mathbb{R}^2. For each slot, ISA forms a relative coordinate for every token and uses it to condition that slot’s keys and values:

    rkn=un−pkδsk,keykn=f ⁣(K(xn)+g(rkn)),valuekn=f ⁣(V(xn)+g(rkn)).r_{kn}=\frac{u_n-p_k}{\delta s_k},\qquad \mathrm{key}_{kn}=f\!\left(K(x_n)+g(r_{kn})\right),\qquad \mathrm{value}_{kn}=f\!\left(V(x_n)+g(r_{kn})\right).

    Here division is componentwise, KK and VV are learned linear projections, gg projects coordinates, and ff is an MLP. The implementation uses the coordinate scaling constant δ=5\delta=5; conceptually, setting δ=1\delta=1 gives the unscaled relative grid. Unlike ordinary Slot Attention, which has one key and value per token, this construction has a separate key and value for each token-slot pair, or NKN K of each. Slot queries compete for tokens through attention normalized over slots; the resulting attention weights are then normalized over tokens for each slot to estimate pose. If akna_{kn} denotes these per-slot, token-normalized weights, the estimates are

    pk=∑n=1Naknun∑n=1Nakn,sk=∑n=1N(akn+ε)(un−pk)⊙2∑n=1N(akn+ε).p_k=\frac{\sum_{n=1}^{N}a_{kn}u_n}{\sum_{n=1}^{N}a_{kn}},\qquad s_k=\sqrt{\frac{\sum_{n=1}^{N}(a_{kn}+\varepsilon)(u_n-p_k)^{\odot 2}}{\sum_{n=1}^{N}(a_{kn}+\varepsilon)}}.

    The square, square root, and scale division are componentwise; ε>0\varepsilon>0 is a small stabilizing constant. The attention-weighted value updates the slot through the Slot Attention recurrent update, and the relative-coordinate construction and pose estimation are repeated at each iteration. An additional final iteration computes pose statistics without updating slot latents. ISA backpropagates through the pose estimates and relative coordinates. The NKN K keys and values add little computational cost in the paper’s usual setting of roughly ten objects per image.

  3. Knowl 3 — The decoder reconstructs each slot in its own relative coordinate grid

    model/method

    ISA adapts the Spatial Broadcast decoder by constructing a separate pose-relative coordinate grid for each slot from its final estimated position and scale. Each slot latent is broadcast over the decoder’s spatial resolution; a learned linear projection hh of that slot’s relative grid is added to the broadcast features before decoding:

    (R,G,B,α)=Dϕ(SB+h(r)).(R,G,B,\alpha)=D_{\phi}\bigl(SB+h(r)\bigr).

    Here SBSB is the spatially broadcast slot feature map, rr is the per-slot relative coordinate grid, hh maps coordinates to the slot-feature dimension, and DϕD_{\phi} predicts color channels and an alpha mask at each pixel. The alpha masks are normalized across slots, and the reconstructed image is their per-pixel weighted sum of the slot-specific color predictions. Using relative coordinates in the decoder encourages slots to reuse appearance decoding across object positions and scales.

  4. Knowl 4 — ISA-TSR adds a principal-component estimate of slot orientation

    model/method

    The rotational variant, ISA-TSR, estimates each slot’s orientation from the weighted spatial distribution of its attention mask. It applies weighted principal component analysis to absolute token coordinates using the slot’s attention weights, takes the first principal axis and an orthogonal second axis, and forms a rotation matrix RkR_k. The axes are post-processed to maintain a consistent handedness and to limit the estimated rotation to the range [0,π/4][0,\pi/4], reducing ambiguity rather than attempting unrestricted orientation estimation. The slot-relative grid becomes rkn=Rk−1(un−pk)/(δsk)r_{kn}=R_k^{-1}(u_n-p_k)/(\delta s_k), so the encoder and decoder can process or generate the object in its estimated oriented frame. The authors compute the weighted PCA axes analytically from a 2×22\times2 covariance matrix so gradients can pass through orientation estimation.

  5. Knowl 5 — Translation-invariant slots improve Tetrominoes sample efficiency and position generalization

    empirical result

    On the Tetrominoes dataset of geometric objects on a black background, translation-invariant ISA (ISA-T) was compared with ordinary Slot Attention (SA). In a low-data experiment with 64–1,024 training images, ISA-T reached performance comparable to SA using approximately 2–4 times fewer training samples. This advantage remained despite translational data augmentation in that experiment and hyperparameter choices favoring SA. In a separate distribution-shift experiment, training objects appeared only on the left side of images while test objects could appear at any position; without data augmentation, ISA-T achieved better foreground adjusted Rand index (FG-ARI) than SA. These results support benefits for both sample efficiency and generalization to unseen object positions.

  6. Knowl 6 — Translation and scale invariance improve results on Objects Room and MultiShapeNet

    empirical result

    On Objects Room, where floors, walls, and ceilings were included as segments and performance was measured by ARI, both translation-invariant ISA-T and translation-and-scale-invariant ISA-TS outperformed SA across the reported validation and test conditions. On MultiShapeNet, which was evaluated with FG-ARI, both variants improved over SA on the full set; ISA-TS performed best in the controlled four-object split, where five slots were used to reduce over-segmentation. The reported means and standard deviations are:

    Dataset and condition SA ISA-T ISA-TS Metric
    Objects Room: Validation 63.017.9 83.59.2 85.56.6 ARI
    Objects Room: Six Objects 62.818.4 82.56.7 84.54.6 ARI
    Objects Room: Empty Room 58.416.4 78.714.3 83.68.8 ARI
    Objects Room: Identical Colors 61.217.8 82.48.2 85.15.9 ARI
    MultiShapeNet: All Data 56.82.6 71.23.5 69.81.1 FG-ARI
    MultiShapeNet: Four Objects 71.511.6 82.35.3 86.51.1 FG-ARI

    Objects Room values use 10 random seeds; MultiShapeNet values use 5. ISA-TS was more prone to over-segmenting parts of objects in the full MultiShapeNet set, while it led SA and ISA-T in the controlled four-object setting. On CLEVR, segmentation was already near ceiling: SA achieved 98.8±0.2%98.8\pm0.2\% FG-ARI, ISA-T 99.0±0.2%99.0\pm0.2\%, and ISA-TS 98.9±0.2%98.9\pm0.2\%.

  7. Knowl 7 — CLEVRTex benefits from translation and scale invariance, including on novel textures

    empirical result

    CLEVRTex evaluates segmentation on its standard test set, a CAMO set where objects and backgrounds blend, and an OOD set with novel textures. FG-ARI is reported as a percentage; results below are means and standard deviations over 10 seeds. A stronger SA baseline used a ResNet-34 encoder, a 16×1616\times16 feature map with matching decoder broadcast resolution, and learnable initial slots. The comparison shows large gains for ISA with the shallow CNN and smaller, condition-dependent gains over the stronger ResNet baseline. ISA-TS with ResNet exceeded the listed AST-Seg-B3-CT result on the OOD set; the paper notes that AST-Seg-B3-CT uses ImageNet and background-model pretraining, whereas these ISA results do not.

    Method Main CAMO OOD
    AST-Seg-B3-CT 94.80.5 87.33.8 83.10.8
    SA (CNN) 54.51.6 53.01.6 54.22.6
    ISA-T (CNN) 66.85.7 65.04.9 65.14.8
    ISA-TS (CNN) 78.83.9 72.93.5 73.23.1
    ISA-TSR (CNN) 79.65.5 73.84.9 74.93.8
    SA (ResNet) 91.32.7 84.92.9 81.41.4
    ISA-T (ResNet) 87.46.6 79.05.9 78.64.9
    ISA-TS (ResNet) 92.90.4 86.20.8 84.40.8
    ISA-TSR (ResNet) 93.30.7 87.01.7 84.91.2

    With the CNN, ISA-TS improved Main FG-ARI by 24.3 percentage points over SA. With ResNet, ISA-TS improved on SA for all three conditions, while ISA-T underperformed it. The rotational variant yielded mixed results rather than a consistent additional improvement.

  8. Knowl 8 — ISA-TS improves single-frame RGB object discovery on Waymo Open

    empirical result

    On Waymo Open v1.4, the authors evaluated single-frame unsupervised instance segmentation using RGB images as input, without optical flow, depth features, or temporal information. All listed experiments used a ResNet-34 encoder and 10 random seeds. FG-ARI and reconstruction MSE are reported as means and standard deviations. ISA-TS improved RGB-target FG-ARI from 27.6±12.727.6\pm12.7 for SA to 39.8±5.339.8\pm5.3; the rotational variant scored 31.2±8.431.2\pm8.4. For depth-target experiments, the model still received RGB input and was trained to predict depth targets. ISA-T then matched SA at 40.6%40.6\% FG-ARI, while a model using absolute position encoding as well as the symmetry-aware design (ISA-T-ABS) reached 45.2±12.2%45.2\pm12.2\%. The authors attribute the weaker symmetry-aware depth result to depth changing with camera distance, unlike RGB appearance.

    Method Target FG-ARI MSE
    SA RGB 27.612.7 58435
    ISA-TS RGB 39.85.3 52312
    ISA-TS, decoder only RGB 32.47.1 58629
    ISA-TSR RGB 31.28.4 56126
    SA Depth 40.65.0 959
    ISA-T Depth 40.67.9 10611
    ISA-T-ABS Depth 45.212.2 969

    Qualitatively, the learned slots could also be manipulated through their position and scale parameters to alter decoded objects or scene regions. Predicted masks sometimes emphasized prominent background landmarks, however, while the evaluation metric rewarded cars and pedestrians.

  9. Knowl 9 — Appending pose parameters lets the recurrent slot update explicitly break symmetry

    model/method

    The visual backbone can leak absolute position information into image features, and some tasks may benefit from access to absolute pose rather than strict equivariance. ISA therefore has an append variant that concatenates slot position, scale, and, when present, rotation parameters to the slot latent immediately before the GRU update at each attention iteration. This gives the recurrent update explicit information with which to model movement of a slot’s attention mask across iterations. On CLEVRTex with the shallow CNN, ISA-TSR-Append achieved 85.4±2.4%85.4\pm2.4\% FG-ARI on the main test set, compared with 79.6±5.5%79.6\pm5.5\% for ISA-TSR. With ResNet-34, the corresponding scores were 93.6±0.8%93.6\pm0.8\% and 93.3±0.7%93.3\pm0.7\%, a much smaller difference; the authors suggest the ResNet already supplies position information.

  10. Knowl 10 — Encoder symmetry and gradients through pose estimation affect segmentation

    empirical result

    Ablations tested whether pose-relative coordinates are needed in the attention encoder and whether gradients must pass through pose estimation. In the decoder-only ablation, the encoder uses absolute coordinates while the decoder retains pose-relative coordinates. Its FG-ARI was 92.3±2.4%92.3\pm2.4\% versus 96.3±3.6%96.3\pm3.6\% for ISA-T on Tetrominoes, and 32.4±7.1%32.4\pm7.1\% versus 39.8±5.3%39.8\pm5.3\% on Waymo Open. On CLEVRTex with ResNet, decoder-only FG-ARI was similar to the full ISA-TSR result (93.4±1.0%93.4\pm1.0\% versus 93.3±0.7%93.3\pm0.7\%), but reconstruction MSE worsened from 185.9±6.4185.9\pm6.4 to 200.4±6.6200.4\pm6.6. Stopping gradients through estimated pose parameters had a large adverse effect on CLEVRTex: ISA-TSR achieved 93.3±0.7%93.3\pm0.7\% FG-ARI, compared with 73.4±10.5%73.4\pm10.5\% when pose gradients were stopped. These results show that encoder-side symmetry can matter substantially depending on the dataset, and that end-to-end gradient flow through pose estimation was important in the reported CLEVRTex experiment.

  11. Knowl 11 — The 2D symmetry assumptions and background handling limit the method

    limitation

    ISA models translation, scale, and approximate rotation in 2D image coordinates, whereas images of 3D scenes do not generally obey these exact symmetries. The method does not directly model out-of-plane rotation, shear, or non-rigid deformation without additional information about 3D geometry. Its principal-component rotation heuristic can also be ambiguous for symmetric or complex objects, consistent with the mixed or negative rotational results on several datasets. In addition, assigning position, scale, and rotation to scene-wide regions such as sky or road may be inappropriate; the paper identifies combining ISA with a separate background model as an open direction. The paper’s experiments focus on object discovery, so they do not establish whether the same benefits transfer to other tasks for which the architecture might be adapted.

Coverage note — The near-ceiling CLEVR comparisons beyond the reported FG-ARI values, detailed architecture and optimizer configurations, and supplementary reconstruction and FG-mIoU tables are omitted because they do not add a distinct main finding to the method, benchmark comparisons, or ablation results included here.

References

  1. 1.Barlow, H. Grandmother cells, symmetry, and invariance: How the term arose and what the facts suggest. In Gazzaniga, M. S. (ed.), The Cognitive Neurosciences, pp. 309–320. The MIT Press, fourth edition, 2009.
  2. 2.Bello, I., Zoph, B., Le, Q., Vaswani, A., and Shlens, J. Attention Augmented Convolutional Networks. In ICCV, 2019.
  3. 3.Bottini, R. and Doeller, C. F. Knowledge across reference frames: Cognitive maps and image spaces. Trends in Cognitive Sciences, 24(8):606–619, 2020.
  4. 4.Bronstein, M. M., Bruna, J., Cohen, T., and Velickovic, P. Geometric deep learning: Grids, groups, graphs, geodesics, and gauges. CoRR, abs/2104.13478, 2021.
  5. 5.Burgess, C. P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M. M., and Lerchner, A. Monet: Unsupervised scene decomposition and representation. CoRR, abs/1901.11390, 2019.
  6. 6.Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. End-to-End Object Detection with Transformers. In ECCV, 2020.
  7. 7.Cho, K., Van Merrienboer, B., Bahdanau, D., and Bengio, Y. On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259, 2014.
  8. 8.Cordonnier, J.-B., Loukas, A., and Jaggi, M. On the relationship between self-attention and convolutional layers. In ICLR, 2020.
  9. 9.Crawford, E. and Pineau, J. Spatially invariant unsupervised object detection with convolutional neural networks. In AAAI, 2019.
  10. 10.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR, 2021.
  11. 11.Dufter, P., Schmitt, M., and Schutze, H. Position Information in Transformers: An Overview. Comput. Linguistics, 48(3):733–763, 2022.
  12. 12.Elsayed, G. F., Mahendran, A., van Steenkiste, S., Greff, K., Mozer, M. C., and Kipf, T. SAVi++: Towards end-to-end object-centric learning from real-world videos. NeurIPS, 2022.
  13. 13.Emami, P., He, P., Ranka, S., and Rangarajan, A. Efficient iterative amortized inference for learning symmetric and disentangled multi-object representations. In ICML, 2021.
  14. 14.Engelcke, M., Kosiorek, A. R., Jones, O. P., and Posner, I. GENESIS: generative scene inference and sampling with object-centric latent representations. In ICLR, 2020.
  15. 15.Eslami, S. M. A., Heess, N., Weber, T., Tassa, Y., Szepesvari, D., Kavukcuoglu, K., and Hinton, G. E. Attend, infer, repeat: Fast scene understanding with generative models. In NeurIPS, 2016.
  16. 16.Gao, P., Zheng, M., Wang, X., Dai, J., and Li, H. Fast Convergence of DETR with Spatially Modulated Co-Attention. In ICCV, 2021.
  17. 17.Greff, K., Srivastava, R. K., and Schmidhuber, J. Binding via reconstruction clustering. CoRR, abs/1511.06418, 2015.
  18. 18.Greff, K., Rasmus, A., Berglund, M., Hao, T. H., Valpola, H., and Schmidhuber, J. Tagger: Deep unsupervised perceptual grouping. In Lee, D. D., Sugiyama, M., von Luxburg, U., Guyon, I., and Garnett, R. (eds.), NeurIPS, 2016.
  19. 19.Greff, K., van Steenkiste, S., and Schmidhuber, J. Neural expectation maximization. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), NeurIPS, 2017.
  20. 20.Greff, K., Kaufman, R. L., Kabra, R., Watters, N., Burgess, C., Zoran, D., Matthey, L., Botvinick, M. M., and Lerchner, A. Multi-object representation learning with iterative variational inference. In ICML, 2019.
  21. 21.Greff, K., Van Steenkiste, S., and Schmidhuber, J. On the binding problem in artificial neural networks. arXiv preprint arXiv:2012.05208, 2020.
  22. 22.Greff, K., Belletti, F., Beyer, L., Doersch, C., Du, Y., Duckworth, D., Fleet, D. J., Gnanapragasam, D., Golemo, F., Herrmann, C., Kipf, T., Kundu, A., Lagun, D., Laradji, I. H., Liu, H. D., Meyer, H., Miao, Y., Nowrouzezahrai, D., Oztireli, A. C., Pot, E., Radwan, N., Rebain, D., Sabour, S., Sajjadi, M. S. M., Sela, M., Sitzmann, V., Stone, A., Sun, D., Vora, S., Wang, Z., Wu, T., Yi, K. M., Zhong, F., and Tagliasacchi, A. Kubric: A scalable dataset generator. In CVPR, 2022.
  23. 23.Han, J., Rong, Y., Xu, T., Sun, F., and Huang, W. Equivariant graph hierarchy-based neural networks. CoRR, abs/2202.10643, 2022.
  24. 24.Hawkins, J., Lewis, M., Klukas, M., Purdy, S., and Ahmad, S. A framework for intelligence and cortical function based on grid cells in the neocortex. Frontiers in Neural Circuits, 12, 2019.
  25. 25.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, 2016.
  26. 26.Hinton, G. E. A parallel computation that assigns canonical object-based frames of reference. In Hayes, P. J. (ed.), Proceedings of the 7th International Joint Conference on Artificial Intelligence, IJCAI ’81, 1981.
  27. 27.Hinton, G. E. How to represent part-whole hierarchies in a neural network. CoRR, abs/2102.12627, 2021.
  28. 28.Hinton, G. E., Krizhevsky, A., and Wang, S. D. Transforming auto-encoders. In Honkela, T., Duch, W., Girolami, M. A., and Kaski, S. (eds.), International Conference on Artificial Neural Networks, 2011.
  29. 29.Hinton, G. E., Sabour, S., and Frosst, N. Matrix capsules with EM routing. In ICLR, 2018.
  30. 30.Hubert, L. and Arabie, P. Comparing partitions. Journal of Classification, 1985.
  31. 31.Ioffe, S. and Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015.
  32. 32.Jaderberg, M., Simonyan, K., Zisserman, A., and Kavukcuoglu, K. Spatial transformer networks. In NeurIPS, 2015.
  33. 33.Jiang, J. and Ahn, S. Generative neurosymbolic machines. In NeurIPS, 2020.
  34. 34.Jiang, J., Janghorbani, S., de Melo, G., and Ahn, S. SCALOR: generative world models with scalable object representations. In ICLR, 2020.
  35. 35.Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R. B. CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR, 2017.
  36. 36.Kabra, R., Burgess, C., Matthey, L., Kaufman, R. L., Greff, K., Reynolds, M., and Lerchner, A. Multi-object datasets. https://github.com/deepmind/multi-object-datasets/, 2019.
  37. 37.Karazija, L., Laina, I., and Rupprecht, C. Clevrtex: A texture-rich benchmark for unsupervised multi-object segmentation. In NeurIPS Track on Datasets and Benchmarks 1, 2021.
  38. 38.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In Bengio, Y. and LeCun, Y. (eds.), ICLR, 2015.
  39. 39.Kipf, T., Elsayed, G. F., Mahendran, A., Stone, A., Sabour, S., Heigold, G., Jonschkowski, R., Dosovitskiy, A., and Greff, K. Conditional object-centric learning from video. In ICLR, 2022.
  40. 40.Kosiorek, A. R., Kim, H., Teh, Y. W., and Posner, I. Sequential attend, infer, repeat: Generative modelling of moving objects. In NeurIPS, 2018.
  41. 41.Kosiorek, A. R., Sabour, S., Teh, Y. W., and Hinton, G. E. Stacked capsule autoencoders. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R. (eds.), NeurIPS, pp. 15486–15496, 2019.
  42. 42.Lin, Z., Wu, Y., Peri, S. V., Sun, W., Singh, G., Deng, F., Jiang, J., and Ahn, S. SPACE: unsupervised object-oriented scene representation via spatial attention and decomposition. In ICLR, 2020.
  43. 43.Liu, S., Li, F., Zhang, H., Yang, X., Qi, X., Su, H., Zhu, J., and Zhang, L. DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR. In ICLR, 2022.
  44. 44.Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T. Object-centric learning with slot attention. In NeurIPS, 2020.
  45. 45.Loshchilov, I. and Hutter, F. SGDR: stochastic gradient descent with warm restarts. In ICLR, 2017.
  46. 46.Luong, T., Pham, H., and Manning, C. D. Effective approaches to attention-based neural machine translation. In EMNLP, 2015.
  47. 47.Mei, J., Zhu, A. Z., Yan, X., Yan, H., Qiao, S., Chen, L., and Kretzschmar, H. Waymo open dataset: Panoramic video panoptic segmentation. In ECCV, 2022.
  48. 48.Meng, D., Chen, X., Fan, Z., Zeng, G., Li, H., Yuan, Y., Sun, L., and Wang, J. Conditional DETR for Fast Training Convergence. In ICCV, 2021.
  49. 49.Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
  50. 50.Monnier, T., Vincent, E., Ponce, J., and Aubry, M. Unsupervised layered image decomposition into object prototypes. In ICCV, 2021.
  51. 51.Nair, V. and Hinton, G. E. Rectified linear units improve restricted boltzmann machines. In ICML, 2010.
  52. 52.Park, J. Y., Biza, O., Zhao, L., van de Meent, J., and Walters, R. Learning symmetric embeddings for equivariant world models. In ICML, 2022.
  53. 53.Parmar, N., Ramachandran, P., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J. Stand-Alone Self-Attention in Vision Models. In NeurIPS, 2019.
  54. 54.Ranftl, R., Bochkovskiy, A., and Koltun, V. Vision transformers for dense prediction. In ICCV, 2021.
  55. 55.Sabour, S., Frosst, N., and Hinton, G. E. Dynamic routing between capsules. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), NeurIPS, 2017.
  56. 56.Sajjadi, M. S. M., Meyer, H., Pot, E., Bergmann, U., Greff, K., Radwan, N., Vora, S., Lucic, M., Duckworth, D., Dosovitskiy, A., Uszkoreit, J., Funkhouser, T. A., and Tagliasacchi, A. Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations. CoRR, abs/2111.13152, 2021.
  57. 57.Sajjadi, M. S. M., Duckworth, D., Mahendran, A., van Steenkiste, S., Pavetic, F., Lucic, M., Guibas, L. J., Greff, K., and Kipf, T. Object scene representation transformer. CoRR, abs/2206.06922, 2022.
  58. 58.Sauvalle, B. and de La Fortelle, A. Unsupervised multi-object segmentation using attention and soft-argmax. CoRR, abs/2205.13271, 2022.
  59. 59.Seitzer, M., Horn, M., Zadaianchuk, A., Zietlow, D., Xiao, T., Simon-Gabriel, C.-J., He, T., Zhang, Z., Scholkopf, B., Brox, T., and Locatello, F. Bridging the gap to real-world object-centric learning. In ICLR, 2023.
  60. 60.Shaw, P., Uszkoreit, J., and Vaswani, A. Self-Attention with Relative Position Representations. In NAACL-HLT, 2018.
  61. 61.Singh, G., Deng, F., and Ahn, S. Illiterate DALL-E learns to compose. CoRR, abs/2110.11405, 2021.
  62. 62.Singh, G., Wu, Y., and Ahn, S. Simple unsupervised object-centric learning for complex and naturalistic videos. CoRR, abs/2205.14065, 2022.
  63. 63.Smirnov, D., Gharbi, M., Fisher, M., Guizilini, V., Efros, A. A., and Solomon, J. M. Marionette: Self-supervised sprite learning. In NeurIPS, 2021.
  64. 64.Srinivas, A., Lin, T.-Y., Parmar, N., Shlens, J., Abbeel, P., and Vaswani, A. Bottleneck Transformers for Visual Recognition. In CVPR, 2021.
  65. 65.Stelzner, K., Kersting, K., and Kosiorek, A. R. Decomposing 3d scenes into objects via unsupervised volume segmentation. CoRR, abs/2104.01148, 2021.
  66. 66.Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., Vasudevan, V., Han, W., Ngiam, J., Zhao, H., Timofeev, A., Ettinger, S., Krivokon, M., Gao, A., Joshi, A., Zhang, Y., Shlens, J., Chen, Z., and Anguelov, D. Scalability in perception for autonomous driving: Waymo open dataset. In CVPR, 2020.
  67. 67.van Steenkiste, S., Chang, M., Greff, K., and Schmidhuber, J. Relational neural expectation maximization: Unsupervised discovery of objects and their interactions. In ICLR, 2018.
  68. 68.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is All you Need. In NeurIPS, 2017.
  69. 69.Wang, D., Kohler, C., and Jr., R. P. Policy learning in SE(3) action spaces. In Robot Learning, CoRL, 2020a.
  70. 70.Wang, H., Zhu, Y., Green, B., Adam, H., Yuille, A. L., and Chen, L.-C. Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation. In ECCV, 2020b.
  71. 71.Wang, Y., Zhang, X., Yang, T., and Sun, J. Anchor DETR: Query Design for Transformer-Based Detector. In AAAI, 2022.
  72. 72.Watters, N., Matthey, L., Burgess, C. P., and Lerchner, A. Spatial broadcast decoder: A simple architecture for learning disentangled representations in vaes. CoRR, abs/1901.07017, 2019.
  73. 73.Wu, Y. and He, K. Group normalization. In ECCV, 2018.
  74. 74.Wu, Z., Dvornik, N., Greff, K., Kipf, T., and Garg, A. Slot-former: Unsupervised visual dynamics simulation with object-centric models. CoRR, abs/2210.05861, 2022.
  75. 75.Xie, J., Xie, W., and Zisserman, A. Segmenting moving objects via an object-centric layered representation. In NeurIPS, 2022.
  76. 76.Yi, W. and Marshall, S. Principal component analysis in application to object orientation. Geo-spatial Information Science, 2000.
  77. 77.Yu, H., Guibas, L. J., and Wu, J. Unsupervised discovery of object radiance fields. In ICLR, 2022.
  78. 78.Zhang, H., Li, F., Liu, S., Zhang, L., Su, H., Zhu, J., Ni, L. M., and Shum, H.-Y. DINO: DETR with Improved De-Noising Anchor Boxes for End-to-End Object Detection, 2022.
  79. 79.Zhao, H., Jia, J., and Koltun, V. Exploring Self-Attention for Image Recognition. In CVPR, 2020.
  80. 80.Zhou, Y., Zhang, H., Lee, H., Sun, S., Li, P., Zhu, Y., Yoo, B., Qi, X., and Han, J.-J. Slot-VPS: Object-centric representation learning for video panoptic segmentation. In CVPR, 2022.
  81. 81.Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. Deformable DETR: Deformable Transformers for End-to-End Object Detection. In ICLR, 2021.

Citation

MLA
Biza, O., et al. “Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames”. International Conference on Machine Learning, vol. 202, 2023, pp. 2507–27, https://proceedings.mlr.press/v202/biza23a.html.
APA
Biza, O., Steenkiste, S. V., Sajjadi, M. S. M., Elsayed, G. F., Mahendran, A., & Kipf, T. (2023). Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames. International Conference on Machine Learning, 202, 2507–2527. https://proceedings.mlr.press/v202/biza23a.html
Chicago
Biza, O., S. V. Steenkiste, M. S. M. Sajjadi, G. F. Elsayed, A. Mahendran, and T. Kipf. 2023. “Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames”. International Conference on Machine Learning 202: 2507–27. https://proceedings.mlr.press/v202/biza23a.html.
Harvard
Biza, O. et al. (2023) “Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames”, International Conference on Machine Learning. PMLR, pp. 2507–2527. Available at: https://proceedings.mlr.press/v202/biza23a.html.
Vancouver
1. Biza O, Steenkiste SV, Sajjadi MSM, Elsayed GF, Mahendran A, Kipf T (2023) Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames. In: International Conference on Machine Learning. PMLR, pp 2507–2527

BibTeX

@InProceedings{pmlr-v202-biza23a,
  title = 	 {Invariant Slot Attention: Object Discovery with Slot-Centric Reference Frames},
  author =       {Biza, Ondrej and Steenkiste, Sjoerd Van and Sajjadi, Mehdi S. M. and Elsayed, Gamaleldin Fathy and Mahendran, Aravindh and Kipf, Thomas},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {2507--2527},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/biza23a/biza23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/biza23a.html},
  abstract = 	 {Automatically discovering composable abstractions from raw perceptual data is a long-standing challenge in machine learning. Recent slot-based neural networks that learn about objects in a self-supervised manner have made exciting progress in this direction. However, they typically fall short at adequately capturing spatial symmetries present in the visual world, which leads to sample inefficiency, such as when entangling object appearance and pose. In this paper, we present a simple yet highly effective method for incorporating spatial symmetries via slot-centric reference frames. We incorporate equivariance to per-object pose transformations into the attention and generation mechanism of Slot Attention by translating, scaling, and rotating position encodings. These changes result in little computational overhead, are easy to implement, and can result in large gains in terms of data efficiency and overall improvements to object discovery. We evaluate our method on a wide range of synthetic object discovery benchmarks namely CLEVR, Tetrominoes, CLEVRTex, Objects Room and MultiShapeNet, and show promising improvements on the challenging real-world Waymo Open dataset.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/