Generalization and Robustness Implications in Object-Centric Learning

Andrea DittadiSamuele S. PapaMichele De VitaBernhard SchölkopfOle WintherFrancesco Locatello

article2022ICML92 citations

Provides the first large-scale empirical benchmark testing whether unsupervised object-centric representation models generalize to downstream property prediction and maintain performance under single-object and scene-wide distribution shifts across five multi-object datasets.

Listen

Modern computer vision models typically process scenes as single, blended representations, which often causes them to struggle when decomposing complex scenes into individual entities or adapting when visual conditions shift. Object-centric representation learning aims to solve this by structuring neural networks to represent visual inputs as discrete, interacting objects. The main objective of the article is to evaluate whether unsupervised object-centric architectures produce features that transfer effectively to downstream prediction tasks and whether they maintain robustness when encountering unexpected changes in object appearance or overall scene composition.

To evaluate these questions, the authors trained and tested four prominent unsupervised object-centric models (MONet, GENESIS, Slot Attention, and SPACE) alongside two baseline autoencoders across five standard multi-object benchmark datasets. Across roughly 250 experimental runs, the models were evaluated on unsupervised segmentation quality, pixel reconstruction error, and downstream accuracy in predicting specific object properties such as color, shape, and position. The investigation systematically introduced out-of-distribution variations at test time, categorizing shifts into single-object modifications (unseen shapes, colors, or artistic textures) and global scene disturbances (image cropping, occlusions, and increased counts of objects).

The findings show that object-centric models successfully learn representations that support downstream property prediction, consistently outperforming distributed baseline autoencoders. Notably, segmentation accuracy measured by the Adjusted Rand Index strongly correlates with downstream task performance, providing a dependable metric for model validation. When individual objects undergo distribution shifts, the representations and property predictions of the remaining unchanged objects remain largely unaffected, demonstrating strong local modularity. However, under unstructured global shifts—particularly image cropping and resizing—segmentation and downstream performance degrade substantially across models, and retraining downstream predictors on shifted data fails to restore full accuracy.

These results indicate that while object-centric learning provides valuable modularity and robustness against isolated object changes, current models cannot yet fully safeguard against broader, scene-level transformations. For technical leaders and practitioners, the findings suggest that relying on the Adjusted Rand Index (or reconstruction error when segmentation masks are absent) offers an effective strategy for model selection. Future research and development should focus on testing these architectures against realistic, textured imagery, exploring supervised signals during pretraining, and developing mechanisms that better preserve spatial structure during global scene shifts.

arXiv: 2107.00637

No sufficiently relevant recommendations were found.

  • Paper: Object-Centric Slot Diffusion, Jindong Jiang et al. (2023). Latent Slot Diffusion carries object-centric slots into high-fidelity generation and natural imagery, extending the source’s evaluation of segmentation and object-property representations.
  • Paper: SAVi++: Towards End-to-End Object-Centric Learning from Real-World Videos, Gamaleldin F. Elsayed et al. (2022). SAVi++ extends slot-based object-centric learning to real-world video, adding depth prediction and tracking to the source’s focus on robustness across scene changes.
Cover for Generalization and Robustness Implications in Object-Centric Learning

Citation

MLA
Dittadi, A., et al. “Generalization and Robustness Implications in Object-Centric Learning”. International Conference on Machine Learning, vol. 162, 2022, pp. 5221–85, https://proceedings.mlr.press/v162/dittadi22a.html.
APA
Dittadi, A., Papa, S. S., Vita, M. D., Schölkopf, B., Winther, O., & Locatello, F. (2022). Generalization and Robustness Implications in Object-Centric Learning. International Conference on Machine Learning, 162, 5221–5285. https://proceedings.mlr.press/v162/dittadi22a.html
Chicago
Dittadi, A., S. S. Papa, M. D. Vita, B. Schölkopf, O. Winther, and F. Locatello. 2022. “Generalization and Robustness Implications in Object-Centric Learning”. International Conference on Machine Learning 162: 5221–85. https://proceedings.mlr.press/v162/dittadi22a.html.
Harvard
Dittadi, A. et al. (2022) “Generalization and Robustness Implications in Object-Centric Learning”, International Conference on Machine Learning. PMLR, pp. 5221–5285. Available at: https://proceedings.mlr.press/v162/dittadi22a.html.
Vancouver
1. Dittadi A, Papa SS, Vita MD, Schölkopf B, Winther O, Locatello F (2022) Generalization and Robustness Implications in Object-Centric Learning. In: International Conference on Machine Learning. PMLR, pp 5221–5285

BibTeX

@InProceedings{pmlr-v162-dittadi22a,
  title = 	 {Generalization and Robustness Implications in Object-Centric Learning},
  author =       {Dittadi, Andrea and Papa, Samuele S and De Vita, Michele and Sch{\"o}lkopf, Bernhard and Winther, Ole and Locatello, Francesco},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {5221--5285},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/dittadi22a/dittadi22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/dittadi22a.html},
  abstract = 	 {The idea behind object-centric representation learning is that natural scenes can better be modeled as compositions of objects and their relations as opposed to distributed representations. This inductive bias can be injected into neural networks to potentially improve systematic generalization and performance of downstream tasks in scenes with multiple objects. In this paper, we train state-of-the-art unsupervised models on five common multi-object datasets and evaluate segmentation metrics and downstream object property prediction. In addition, we study generalization and robustness by investigating the settings where either a single object is out of distribution – e.g., having an unseen color, texture, or shape – or global properties of the scene are altered – e.g., by occlusions, cropping, or increasing the number of objects. From our experimental study, we find object-centric representations to be useful for downstream tasks and generally robust to most distribution shifts affecting objects. However, when the distribution shift affects the input in a less structured manner, robustness in terms of segmentation and downstream task performance may vary significantly across models and distribution shifts.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/