Provably Learning Object-Centric Representations

Jack BradyRoland S. ZimmermannYash SharmaBernhard SchölkopfJulius von KügelgenWieland Brendel

article2023ICML58 citations

Establishes the first theoretical identifiability guarantees for unsupervised object-centric representation learning by proving that invertible, compositional inference models can recover ground-truth object slots even when objects exhibit statistical dependencies.

Listen

Visual machine learning systems often struggle to generalize from limited training examples because they process visual scenes as flat arrays of pixels rather than structured compositions of distinct objects. While recent empirical models attempt to learn object-centric representations without human labeling, the field has lacked a mathematical foundation to explain when and why such representations can be reliably recovered. Without theoretical guarantees, developing robust visual architectures has relied largely on heuristics that often fail in complex, real-world environments.

The article establishes the first theoretical framework proving that object-centric representations can be reliably learned without supervision. Specifically, the analysis demonstrates that an unsupervised inference model can provably identify underlying ground-truth object slots, even when strong statistical dependencies exist between objects in a scene.

To establish these guarantees, the article models multi-object scene generation under two structural conditions on the rendering process: compositionality, where each observed pixel is influenced by at most one object, and irreducibility, where visual parts belonging to the same object cannot be separated into independent sub-mechanisms. The analysis then proves that an invertible inference model paired with a compositional inverse uniquely isolates ground-truth object properties up to natural slot permutations. To make this operational, the article introduces a mathematical metric called compositional contrast, which measures deviations from compositionality. The theoretical claims were validated on synthetic multi-object datasets across various numbers of objects and latent dimensions, and further tested against three prominent object-centric image architectures—Slot Attention, MONet, and additive auto-encoders.

The evaluation produced several clear findings. First, across 180 synthetic model runs, jointly minimizing reconstruction loss and compositional contrast achieved near-perfect slot recovery, confirming the theoretical sufficiency of these properties. Second, the theoretical guarantees held even when objects exhibited strong statistical correlations, overcoming a major limitation of prior frameworks that required object independence. Third, empirical evaluations of Slot Attention, MONet, and additive auto-encoders demonstrated that higher empirical identifiability closely tracked lower reconstruction error and lower compositional contrast. Finally, the analysis revealed that when models allocate more latent capacity than necessary, individual slots leak information about secondary objects, despite producing clean visual segmentation outputs.

These results provide a clear roadmap for designing more reliable and interpretable vision systems in applications such as robotics, autonomous vehicles, and causal reasoning engines. By shifting the focus from restrictive data-distribution assumptions to structural constraints on model decoders, the findings explain why auto-encoding approaches succeed empirically and expose latent over-capacity as a hidden failure mode. When inferred latent slots quietly encode multiple objects, downstream decision-making systems risk inheriting hidden cross-object confounding.

Based on these findings, teams developing unsupervised visual systems should adopt two main practices. First, practitioners should carefully restrict latent bottleneck capacities to match true task complexity, avoiding excessive per-slot dimensionality that encourages multi-object contamination. Second, system architects should explore training objectives that implicitly or explicitly penalize gradient overlap across slots to enforce compositionality. Further work is required to develop computationally efficient, first-order approximations of compositional contrast for large-scale production training, as current implementations rely on second-order derivatives.

Confidence in the mathematical proofs and controlled synthetic validations is high. However, caution is advised when transferring these guarantees directly to complex real-world visual settings. Real-world scenes frequently violate the core theoretical assumptions through optical phenomena such as transparency, reflections, object occlusions, and multi-view perspective transformations. Expanding the theory to accommodate these optical boundary conditions remains an essential area for future research.

No sufficiently relevant recommendations were found.

Cover for Provably Learning Object-Centric Representations

Abstract

Learning structured representations of the visual world in terms of objects promises to significantly improve the generalization abilities of current machine learning models. While recent efforts to this end have shown promising empirical progress, a theoretical account of when unsupervised object-centric representation learning is possible is still lacking. Consequently, understanding the reasons for the success of existing object-centric methods as well as designing new theoretically grounded methods remains challenging. In the present work, we analyze when object-centric representations can provably be learned without supervision. To this end, we first introduce two assumptions on the generative process for scenes comprised of several objects, which we call compositionality and irreducibility. Under this generative process, we prove that the ground-truth object representations can be identified by an invertible and compositional inference model, even in the presence of dependencies between objects. We empirically validate our results through experiments on synthetic data. Finally, we provide evidence that our theory holds predictive power for existing object-centric models by showing a close correspondence between models’ compositionality and invertibility and their empirical identifiability.

Table of Contents

  • 1 Introduction
  • 2 Generative Model
  • 2.1 Slots and Compositionality
  • 2.2 Mechanisms and Irreducibility
  • 3 Theory: Slot Identifiability
  • 4 Related Work
  • 5 Experiments
  • 5.1 Synthetic Data
  • 5.2 Existing Object-Centric Models
  • 6 Discussion
  • References
  • A Proofs
  • B Experimental Details
  • B.1 Synthetic Data § 5.1
  • B.2 Existing Object-Centric Models § 5.2
  • B.3 Compositional Contrast Normalized Variants
  • B.4 Slot Identifiability Score
  • C Additional Figures and Experiments

Knowls

  1. Knowl 1 — Slot identifiability from invertible compositional inference

    theoretical result

    Let the true scene generator f:RKM→RNf:\mathbb{R}^{KM}\to\mathbb{R}^{N} be a diffeomorphism, with its KMKM-dimensional latent vector divided into KK slots of dimension MM. Suppose ff is compositional (each observation coordinate depends locally on at most one slot) and each slot mechanism is irreducible (its information cannot be split across two independent subsets of observation coordinates). If an encoder g^:RN→RKM\hat g:\mathbb{R}^{N}\to\mathbb{R}^{KM} is a diffeomorphism and its inverse decoder f^=g^−1\hat f=\hat g^{-1} is compositional, then the encoder recovers the ground-truth slots up to one global permutation and separate invertible transformations within slots: there are a permutation π\pi of {1,…,K}\{1,\ldots,K\} and slot-wise diffeomorphisms hjh_j such that z^j=hj(zπ(j))\hat z_j=h_j(z_{\pi(j)}). The result does not require statistical independence among the ground-truth slots.

  2. Knowl 2 — Scene-generating model with dependent latent slots

    model/method

    A scene with KK objects is modeled by a latent vector z=(z1,…,zK)z=(z_1,\ldots,z_K), where each object is represented by a slot zk∈RMz_k\in\mathbb{R}^M and the full latent space is Z=RKM\mathcal{Z}=\mathbb{R}^{KM}. The observation x∈RNx\in\mathbb{R}^N is generated as x=f(z)x=f(z), where f:Z→RNf:\mathcal{Z}\to\mathbb{R}^N is a diffeomorphism. The latent distribution has full support on Z\mathcal{Z} but is otherwise unrestricted: slots may have arbitrary statistical or causal dependencies. This model treats an object as a group of latent properties rather than as a single latent coordinate.

  3. Knowl 3 — Compositionality means each output coordinate belongs locally to at most one slot

    definition

    For a differentiable generator f:RKM→RNf:\mathbb{R}^{KM}\to\mathbb{R}^N, let fnf_n be its nnth scalar output and define the set of coordinates affected by slot kk at latent input zz as Ik(z)={n∈{1,…,N}:∂fn(z)/∂zk≠0}I_k(z)=\{n\in\{1,\ldots,N\}:\partial f_n(z)/\partial z_k\ne 0\}. The generator is compositional if, for every zz and all distinct slots k,jk,j, Ik(z)∩Ij(z)=∅I_k(z)\cap I_j(z)=\varnothing. Thus no output coordinate is locally affected by two slots at once. The affected-coordinate sets, and the corresponding sparse blocks in the generator Jacobian, may change with zz; the page-3 illustration conveys this scene-dependent block structure.

  4. Knowl 4 — Irreducibility rules out splitting one slot into independent mechanisms

    definition

    For a differentiable scene generator ff, the mechanism of slot kk at latent input zz is the submatrix of the Jacobian Jf(z)J_f(z) containing the rows indexed by Ik(z)I_k(z). For an observation-coordinate subset SS, write Jf,S(z)J_{f,S}(z) for the Jacobian rows indexed by SS. Two disjoint, nonempty subsets S1,S2S_1,S_2 define independent sub-mechanisms when rank⁡(Jf,S1∪S2(z))=rank⁡(Jf,S1(z))+rank⁡(Jf,S2(z))\operatorname{rank}(J_{f,S_1\cup S_2}(z))=\operatorname{rank}(J_{f,S_1}(z))+\operatorname{rank}(J_{f,S_2}(z)); they are dependent when the left side is strictly smaller. A generator is irreducible if, for every latent input and every slot, every partition of that slot’s affected coordinates into two nonempty subsets yields dependent sub-mechanisms. The rank criterion measures whether the latent capacity needed to produce the joint coordinates decomposes into separate capacities; the page-4 illustration contrasts such decomposable mechanisms with an irreducible one.

  5. Knowl 5 — A Jacobian contrast gives an exact compositionality test and an identification objective

    equation

    For a differentiable decoder f^:RKM→RN\hat f:\mathbb{R}^{KM}\to\mathbb{R}^N and latent input z^=(z^1,…,z^K)\hat z=(\hat z_1,\ldots,\hat z_K), define the compositional contrast as

    Ccomp(f^,z^)=∑n=1N∑k=1K∑j=k+1K∥∂f^n(z^)∂z^k∥2∥∂f^n(z^)∂z^j∥2,C_{\mathrm{comp}}(\hat f,\hat z)=\sum_{n=1}^{N}\sum_{k=1}^{K}\sum_{j=k+1}^{K}\left\|\frac{\partial \hat f_n(\hat z)}{\partial \hat z_k}\right\|_2\left\|\frac{\partial \hat f_n(\hat z)}{\partial \hat z_j}\right\|_2,

    where f^n\hat f_n is output coordinate nn and each derivative is a gradient with respect to an MM-dimensional slot. This contrast is nonnegative and is zero for every input if and only if the decoder is compositional. The paper’s optimization-based identification result states that, under the true generator assumptions of compositionality and irreducibility, a differentiable encoder g^\hat g and decoder f^\hat f slot-identify the ground truth if they satisfy

    Ex∼px[∥f^(g^(x))−x∥22+λCcomp(f^,g^(x))]=0,λ>0,\mathbb{E}_{x\sim p_x}\left[\|\hat f(\hat g(x))-x\|_2^2+\lambda C_{\mathrm{comp}}(\hat f,\hat g(x))\right]=0,\qquad \lambda>0,

    where pxp_x is the observation distribution induced by the full-support latent distribution. The zero reconstruction term enforces inversion, while the zero contrast term enforces decoder compositionality.

  6. Knowl 6 — Synthetic experiments support identification with both independent and dependent slots

    empirical result

    The authors generated controlled scenes by dividing latent vectors into K∈{2,3,5}K\in\{2,3,5\} slots of dimension M=3M=3, transforming each slot with a two-layer LeakyReLU MLP, and concatenating slot outputs of dimension 20. The resulting generator was compositional and invertible; the output dimension exceeding the slot dimension was used to make irreducibility likely. Latents were sampled either independently from a standard normal distribution or dependently using a Wishart-sampled covariance. Models were trained with reconstruction loss and compositional contrast across λ∈{10−7,10−5,10−2,0,1,10}\lambda\in\{10^{-7},10^{-5},10^{-2},0,1,10\} and 10 random seeds (180 models reported). Across the tested slot counts, models that jointly achieved low reconstruction error and low contrast had higher slot-identifiability scores; models that failed to reduce contrast were less identifiable. The dependent-latent experiments showed the same qualitative pattern, using a slot mean-correlation score because leakage scores can be nonzero even for perfectly identifiable models with dependent slots.

  7. Knowl 7 — Slot identifiability is measured by matched slot prediction minus leakage

    model/method

    The experiments quantify whether each ground-truth slot can be predicted from one matched inferred slot without that inferred slot also encoding other ground-truth slots. For each ground-truth/inferred slot pair, the authors fit nonlinear readouts and score continuous factors with R2R^2 and categorical factors with accuracy. After matching slots, S1S_1 is the score for the best-matched slot and S2S_2 is the strongest score from an inferred slot for a ground-truth slot it was not matched to; the basic slot-identifiability score is S1−S2S_1-S_2, aggregated across factors and slots. Synthetic experiments use global Hungarian matching. For image experiments, matching is performed online because object-slot assignments can vary between samples. Image scores are also adjusted for a ground-truth baseline S2gtS_2^{\mathrm{gt}}: define S^1=(S1−S2gt)/(1−S2gt)\hat S_1=(S_1-S_2^{\mathrm{gt}})/(1-S_2^{\mathrm{gt}}) and S^2=(S2−S2gt)/(1−S2gt)\hat S_2=(S_2-S_2^{\mathrm{gt}})/(1-S_2^{\mathrm{gt}}), then use S^1−max⁡(S^2,0)\hat S_1-\max(\hat S_2,0). This makes the score reflect both recovery of the intended slot and leakage into other slots.

  8. Knowl 8 — Existing object-centric models show a qualitative link between contrast, reconstruction, and identifiability

    empirical result

    The authors trained Slot Attention, MONet, and an additive autoencoder on 100,000 Spriteworld images containing 2–4 objects. Each object had four continuous factors (size, color, and x/yx/y position) and one categorical shape factor. Each model used four inferred slots of dimension 16 and was trained for 500,000 iterations. In the page-8 comparison, higher slot-identifiability scores generally occur alongside lower reconstruction error and lower compositional contrast, even though the theoretical equal-dimensionality condition is not met by these models. A training-trajectory plot on page 24 also shows that all three architectures reduce compositional contrast to some extent despite not explicitly optimizing it. These results are empirical associations in the tested models and dataset, rather than a guarantee for existing architectures.

  9. Knowl 9 — Excess inferred capacity is associated with information leakage across objects

    empirical result

    In the Spriteworld experiments, the inferred latent dimension exceeded the ground-truth dimension, contrary to the equal-dimensionality condition used by the identification theorem. The authors measured leakage by scoring how well each inferred slot predicted a second, unmatched ground-truth slot. This correction was nonzero across the tested Slot Attention, MONet, and additive-autoencoder models, including cases where the primary matched-slot score was higher. The page-25 analysis therefore supports the authors’ interpretation that extra slot capacity can encode information about multiple objects in one inferred slot, even if the decoder does not need all of that information to reconstruct the image.

  10. Knowl 10 — Theory and experiments have important scope limitations

    limitation

    The theorem assumes a diffeomorphic generator, equal dimensionality of true and inferred latent spaces, local compositionality, and irreducible slot mechanisms. These assumptions can fail in practical scenes: translucency or reflections can make a pixel depend on multiple objects, and occlusion boundaries can also couple objects at the pixel level. Permutation-invariant generators are not invertible as formulated, so the theorem does not directly cover that common design. The evaluated existing models also use more inferred latent capacity than the ground-truth representation, violating a theorem condition. Empirically, the study tests only a limited set of architectures and datasets; the authors call for broader evaluation. Finally, directly optimizing the Jacobian-based contrast on image models would require second-order optimization, which the paper identifies as a computational obstacle.

Coverage note — The paper’s proof-only lemmas and derivations are omitted because they support, rather than add to, the main identifiability theorem; detailed architecture-specific training hyperparameters and supplementary plots are also omitted where they do not change the reported findings.

References

  1. 1.Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V. F., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., Çağlar Gülçehre, Song, H. F., Ballard, A. J., Gilmer, J., Dahl, G. E., Vaswani, A., Allen, K. R., Nash, C., Langston, V., Dyer, C., Heess, N. M. O., Wierstra, D., Kohli, P., Botvinick, M. M., Vinyals, O., Li, Y., and Pascanu, R. Relational inductive biases, deep learning, and graph networks. ArXiv, abs/1806.01261, 2018. [Cited on page 1.]
  2. 2.Besserve, M., Shajarisales, N., Schölkopf, B., and Janzing, D. Group invariance principles for causal generative models. In AISTATS, volume 84 of Proceedings of Machine Learning Research, pp. 557–565, 2018. [Cited on page 6.]
  3. 3.Besserve, M., Sun, R., Janzing, D., and Schölkopf, B. A theory of independent mechanisms for extrapolation in generative models. In AAAI, pp. 6741–6749, 2021. [Cited on page 6.]
  4. 4.Biza, O., van Steenkiste, S., Sajjadi, M. S., Elsayed, G. F., Mahendran, A., and Kipf, T. Invariant slot attention: Object discovery with slot-centric reference frames. ArXiv preprint, abs/2302.04973, 2023. [Cited on page 1.]
  5. 5.Buchholz, S., Besserve, M., and Schölkopf, B. Function classes for identifiable nonlinear independent component analysis. In NeurIPS, 2022. [Cited on page 7.]
  6. 6.Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., and Lerchner, A. Understanding disentangling in β-vae. ArXiv preprint, abs/1804.03599, 2018. [Cited on page 22.]
  7. 7.Burgess, C. P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M., and Lerchner, A. Monet: Unsupervised scene decomposition and representation. ArXiv preprint, abs/1901.11390, 2019. [Cited on pages 1 and 8.]
  8. 8.Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. In ICCV, pp. 9630–9640, 2021. [Cited on page 1.]
  9. 9.Chen, H., Venkatesh, R. M., Friedman, Y., Wu, J., Tenenbaum, J. B., Yamins, D. L. K., and Bear, D. Unsupervised segmentation in real-world images via spelke object inference. In European Conference on Computer Vision, 2022. [Cited on page 9.]
  10. 10.Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A. L. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2017. [Cited on page 1.]
  11. 11.Daniusis, P., Janzing, D., Mooij, J. M., Zscheischler, J., Steudel, B., Zhang, K., and Schölkopf, B. Inferring deterministic causal relations. In UAI 2010, Proceedings of the Twenty-Sixth Conference on Uncertainty in Artificial Intelligence, Catalina Island, CA, USA, July 8-11, 2010, pp. 143–150, 2010. [Cited on pages 4 and 6.]
  12. 12.Dehaene, S. How We Learn: Why Brains Learn Better Than Any Machine... for Now. 2020. [Cited on page 1.]
  13. 13.Dittadi, A., Papa, S. S., Vita, M. D., Schölkopf, B., Winther, O., and Locatello, F. Generalization and robustness implications in object-centric learning. In ICML, volume 162 of Proceedings of Machine Learning Research, pp. 5221–5285, 2022. [Cited on pages 5, 7, 22, and 23.]
  14. 14.Elsayed, G. F., Mahendran, A., van Steenkiste, S., Greff, K., Mozer, M. C., and Kipf, T. Savi++: Towards end-to-end object-centric learning from real-world videos. In NeurIPS, 2022. [Cited on pages 1, 5, and 9.]
  15. 15.Engelcke, M., Jones, O. P., and Posner, I. Reconstruction bottlenecks in object-centric generative models. ArXiv preprint, abs/2007.06245, 2020a. [Cited on page 5.]
  16. 16.Engelcke, M., Kosiorek, A. R., Jones, O. P., and Posner, I. GENESIS: generative scene inference and sampling with object-centric latent representations. In ICLR, 2020b. [Cited on page 6.]
  17. 17.Engelcke, M., Jones, O. P., and Posner, I. GENESIS-V2: inferring unordered object representations without iterative refinement. In NeurIPS, pp. 8085–8094, 2021. [Cited on page 6.]
  18. 18.Fodor, J. A. and Pylyshyn, Z. W. Connectionism and cognitive architecture: A critical analysis. Cognition, 28(1-2):3–71, 1988. [Cited on page 1.]
  19. 19.Gerstenberg, T. and Tenenbaum, J. B. Intuitive theories. Oxford handbook of causal reasoning, pp. 515–548, 2017. [Cited on page 1.]
  20. 20.Gerstenberg, T., Goodman, N. D., Lagnado, D. A., and Tenenbaum, J. B. A counterfactual simulation model of causal judgments for physical events. Psychological Review, 128(5):936, 2021. [Cited on page 1.]
  21. 21.Gopnik, A., Glymour, C., Sobel, D. M., Schulz, L. E., Kushnir, T., and Danks, D. A theory of causal learning in children: causal maps and bayes nets. Psychological review, 111(1):3, 2004. [Cited on page 1.]
  22. 22.Goyal, A. and Bengio, Y. Inductive biases for deep learning of higher-level cognition. Proceedings of the Royal Society A, 478(2266):20210068, 2022. [Cited on page 1.]
  23. 23.Green, E. J. A theory of perceptual objects. Philosophy and Phenomenological Research, 99(3):663–693, 2019. [Cited on page 2.]
  24. 24.Greff, K., Srivastava, R. K., and Schmidhuber, J. Binding via reconstruction clustering. ArXiv preprint, abs/1511.06418, 2015. [Cited on page 6.]
  25. 25.Greff, K., van Steenkiste, S., and Schmidhuber, J. Neural expectation maximization. In NIPS, pp. 6691–6701, 2017. [Cited on page 6.]
  26. 26.Greff, K., Kaufman, R. L., Kabra, R., Watters, N., Burgess, C., Zoran, D., Matthey, L., Botvinick, M. M., and Lerchner, A. Multi-object representation learning with iterative variational inference. In ICML, volume 97 of Proceedings of Machine Learning Research, pp. 2424–2433, 2019. [Cited on pages 1 and 6.]
  27. 27.Greff, K., Van Steenkiste, S., and Schmidhuber, J. On the binding problem in artificial neural networks. ArXiv preprint, abs/2012.05208, 2020. [Cited on pages 1 and 2.]
  28. 28.Gresele, L., Rubenstein, P. K., Mehrjou, A., Locatello, F., and Schölkopf, B. The incomplete rosetta stone problem: Identifiability results for multi-view nonlinear ICA. In Proceedings of the Thirty-Fifth Conference on Uncertainty in Artificial Intelligence, UAI 2019, Tel Aviv, Israel, July 22-25, 2019, volume 115 of Proceedings of Machine Learning Research, pp. 217–227, 2019. [Cited on page 7.]
  29. 29.Gresele, L., von Kügelgen, J., Stimper, V., Schölkopf, B., and Besserve, M. Independent mechanism analysis, a new concept? In NeurIPS, pp. 28233–28248, 2021. [Cited on pages 4, 6, and 7.]
  30. 30.Guo, S., Tóth, V., Schölkopf, B., and Huszár, F. Causal de Finetti: On the identification of invariant causal structure in exchangeable data. ArXiv preprint, abs/2203.15756, 2022. [Cited on page 6.]
  31. 31.Hälvä, H. and Hyvärinen, A. Hidden markov nonlinear ICA: unsupervised learning from nonstationary time series. In Proceedings of the Thirty-Sixth Conference on Uncertainty in Artificial Intelligence, UAI 2020, virtual online, August 3-6, 2020, volume 124 of Proceedings of Machine Learning Research, pp. 939–948, 2020. [Cited on page 7.]
  32. 32.Hälvä, H., Corff, S. L., Lehéricy, L., So, J., Zhu, Y., Gassiat, E., and Hyvärinen, A. Disentangling identifiable features from noisy data with structured nonlinear ICA. In NeurIPS, pp. 1624–1633, 2021. [Cited on page 7.]
  33. 33.He, K., Gkioxari, G., Dollár, P., and Girshick, R. B. Mask R-CNN. In ICCV, pp. 2980–2988, 2017. [Cited on page 1.]
  34. 34.Heess, N. M. O. Learning generative models of mid-level structure in natural images. PhD thesis, The University of Edinburgh, 2012. [Cited on page 6.]
  35. 35.Hinton, G. E. How to represent part-whole hierarchies in a neural network. Neural computation, pp. 1–40, 2021. [Cited on page 9.]
  36. 36.Horn, R. A. and Johnson, C. R. Matrix analysis. 2012. [Cited on page 16.]
  37. 37.Hyvärinen, A. and Hoyer, P. O. Emergence of phase- and shift-invariant features by decomposition of natural images into independent feature subspaces. Neural Comput., 12(7):1705–1720, 2000. [Cited on page 7.]
  38. 38.Hyvärinen, A. and Morioka, H. Unsupervised feature extraction by time-contrastive learning and nonlinear ICA. In NIPS, pp. 3765–3773, 2016. [Cited on pages 2 and 7.]
  39. 39.Hyvärinen, A. and Morioka, H. Nonlinear ICA of temporally dependent stationary sources. In AISTATS, volume 54 of Proceedings of Machine Learning Research, pp. 460–469, 2017. [Cited on pages 2 and 7.]
  40. 40.Hyvärinen, A. and Perkiö, J. Learning to segment any random vector. The 2006 IEEE International Joint Conference on Neural Network Proceedings, pp. 4167–4172, 2006. [Cited on page 6.]
  41. 41.Hyvärinen, A., Sasaki, H., and Turner, R. E. Nonlinear ICA using auxiliary variables and generalized contrastive learning. In AISTATS, volume 89 of Proceedings of Machine Learning Research, pp. 859–868, 2019. [Cited on pages 2 and 7.]
  42. 42.Hyvärinen, A. and Pajunen, P. Nonlinear independent component analysis: Existence and uniqueness results. Neural Networks, 12(3):429–439, 1999. ISSN 0893-6080. [Cited on page 2.]
  43. 43.Janzing, D. Causal versions of maximum entropy and principle of insufficient reason. Journal of Causal Inference, 9(1):285–301, 2021. [Cited on page 6.]
  44. 44.Janzing, D. and Schölkopf, B. Causal inference using the algorithmic markov condition. IEEE Transactions on Information Theory, 56(10):5168–5194, 2010. [Cited on pages 4 and 6.]
  45. 45.Janzing, D., Hoyer, P. O., and Schölkopf, B. Telling cause from effect based on high-dimensional observations. In ICML, pp. 479–486, 2010. [Cited on page 6.]
  46. 46.Janzing, D., Mooij, J., Zhang, K., Lemeire, J., Zscheischler, J., Daniusis, P., Steudel, B., and Schölkopf, B. Information-geometric approach to inferring causal directions. Artificial Intelligence, 182:1–31, 2012. [Cited on pages 4 and 6.]
  47. 47.Kabra, R., Burgess, C., Matthey, L., Kaufman, R. L., Greff, K., Reynolds, M., and Lerchner, A. Multi-object datasets. https://github.com/deepmind/multi-object-datasets/, 2019. [Cited on page 22.]
  48. 48.Karazija, L., Laina, I., and Rupprecht, C. Clevrtex: A texture-rich benchmark for unsupervised multi-object segmentation. ArXiv preprint, abs/2111.10265, 2021. [Cited on page 1.]
  49. 49.Khemakhem, I., Kingma, D. P., Monti, R. P., and Hyvärinen, A. Variational autoencoders and nonlinear ICA: A unifying framework. In AISTATS, volume 108 of Proceedings of Machine Learning Research, pp. 2207–2217, 2020a. [Cited on pages 2 and 7.]
  50. 50.Khemakhem, I., Monti, R. P., Kingma, D. P., and Hyvärinen, A. Ice-beem: Identifiable conditional energy-based deep models based on nonlinear ICA. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020b. [Cited on page 2.]
  51. 51.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In ICLR (Poster), 2015. [Cited on page 22.]
  52. 52.Kipf, T., Elsayed, G. F., Mahendran, A., Stone, A., Sabour, S., Heigold, G., Jonschkowski, R., Dosovitskiy, A., and Greff, K. Conditional object-centric learning from video. In ICLR, 2022. [Cited on pages 1 and 9.]
  53. 53.Kipf, T. N., van der Pol, E., and Welling, M. Contrastive learning of structured world models. In ICLR, 2020. [Cited on page 1.]
  54. 54.Klindt, D. A., Schott, L., Sharma, Y., Ustyuzhaninov, I., Brendel, W., Bethge, M., and Paiton, D. M. Towards nonlinear disentanglement in natural data with temporal sparse coding. In ICLR, 2021. [Cited on pages 2 and 7.]
  55. 55.Koffka, K. Principles Of Gestalt Psychology. 1936. [Cited on pages 2 and 9.]
  56. 56.Kuhn, H. W. The hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97, 1955. [Cited on pages 7 and 23.]
  57. 57.Lachapelle, S. and Lacoste-Julien, S. Partial disentanglement via mechanism sparsity. In UAI 2022 Workshop on Causal Representation Learning, 2022. [Cited on page 7.]
  58. 58.Lachapelle, S., Rodriguez, P., Sharma, Y., Everett, K. E., Le Priol, R., Lacoste, A., and Lacoste-Julien, S. Disentanglement via mechanism sparsity regularization: A new principle for nonlinear ica. In First Conference on Causal Learning and Reasoning, 2021. [Cited on page 7.]
  59. 59.Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J. Building machines that learn and think like people. Behavioral and brain sciences, 40, 2017. [Cited on page 1.]
  60. 60.Lin, Z., Wu, Y., Peri, S. V., Sun, W., Singh, G., Deng, F., Jiang, J., and Ahn, S. SPACE: unsupervised object-oriented scene representation via spatial attention and decomposition. In ICLR, 2020. [Cited on page 1.]
  61. 61.Locatello, F., Vincent, D., Tolstikhin, I., Rätsch, G., Gelly, S., and Schölkopf, B. Competitive training of mixtures of independent deep generative models. ArXiv preprint, abs/1804.11130, 2018. [Cited on page 6.]
  62. 62.Locatello, F., Bauer, S., Lucic, M., Rätsch, G., Gelly, S., Schölkopf, B., and Bachem, O. Challenging common assumptions in the unsupervised learning of disentangled representations. In ICML, volume 97 of Proceedings of Machine Learning Research, pp. 4114–4124, 2019. [Cited on page 2.]
  63. 63.Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T. Object-centric learning with slot attention. In NeurIPS, 2020. [Cited on pages 1, 5, 7, 8, and 22.]
  64. 64.Marcus, G. F. The Algebraic Mind: Integrating Connectionism and Cognitive Science. 2001. [Cited on page 1.]
  65. 65.Moran, G. E., Sridhar, D., Wang, Y., and Blei, D. M. Identifiable variational autoencoders via sparse decoding. ArXiv preprint, abs/2110.10804, 2021. [Cited on page 7.]
  66. 66.Papa, S., Winther, O., and Dittadi, A. Inductive biases for object-centric representations in the presence of complex textures. In UAI 2022 Workshop on Causal Representation Learning, 2022. [Cited on page 1.]
  67. 67.Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E. Z., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, pp. 8024–8035, 2019. [Cited on page 22.]
  68. 68.Pearl, J. Causality. 2 edition, 2009. [Cited on page 6.]
  69. 69.Peters, B. and Kriegeskorte, N. Capturing the objects of vision with neural networks. Nature Human Behaviour, 5(9):1127–1144, 2021. [Cited on page 1.]
  70. 70.Peters, J., Janzing, D., and Schölkopf, B. Elements of causal inference: foundations and learning algorithms. 2017. [Cited on pages 2, 4, and 6.]
  71. 71.Reizinger, P., Gresele, L., Brady, J., von Kügelgen, J., Zietlow, D., Schölkopf, B., Martius, G., Brendel, W., and Besserve, M. Embrace the gap: Vaes perform independent mechanism analysis. In NeurIPS, 2022. [Cited on page 7.]
  72. 72.Ronneberger, O., Fischer, P., and Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pp. 234–241, 2015. [Cited on page 1.]
  73. 73.Roux, N. L., Heess, N. M. O., Shotton, J., and Winn, J. M. Learning a generative model of images by factoring appearance and shape. Neural Computation, 23:593–650, 2011. [Cited on page 6.]
  74. 74.Sajjadi, M. S. M., Duckworth, D., Mahendran, A., van Steenkiste, S., Pavetic, F., Lucic, M., Guibas, L. J., Greff, K., and Kipf, T. Object scene representation transformer. In NeurIPS, 2022. [Cited on pages 1 and 5.]
  75. 75.Schölkopf, B., Janzing, D., Peters, J., Sgouritsa, E., Zhang, K., and Mooij, J. M. On causal and anticausal learning. In ICML, 2012. [Cited on page 6.]
  76. 76.Schölkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y. Toward causal representation learning. Proceedings of the IEEE, 109(5):612–634, 2021. [Cited on page 6.]
  77. 77.Seitzer, M., Horn, M., Zadaianchuk, A., Zietlow, D., Xiao, T., Simon-Gabriel, C.-J., He, T., Zhang, Z., Schölkopf, B., Brox, T., and Locatello, F. Bridging the gap to real-world object-centric learning. In The Eleventh International Conference on Learning Representations, 2023. [Cited on pages 1 and 5.]
  78. 78.Shajarisales, N., Janzing, D., Schölkopf, B., and Besserve, M. Telling cause from effect in deterministic linear dynamical systems. In ICML, volume 37 of JMLR Workshop and Conference Proceedings, pp. 285–294, 2015. [Cited on page 6.]
  79. 79.Singh, G., Deng, F., and Ahn, S. Illiterate DALL-E learns to compose. In ICLR, 2022a. [Cited on page 1.]
  80. 80.Singh, G., Wu, Y., and Ahn, S. Simple unsupervised object-centric learning for complex and naturalistic videos. In NeurIPS, 2022b. [Cited on pages 1 and 5.]
  81. 81.Spelke, E. S. Principles of object perception. Cogn. Sci., 14:29–56, 1990. [Cited on pages 2 and 9.]
  82. 82.Spelke, E. S. What makes us smart? core knowledge and natural language. Language in mind: Advances in the study of language and thought, pp. 277–311, 2003. [Cited on page 1.]
  83. 83.Spelke, E. S. and Kinzler, K. D. Core knowledge. Developmental science, 10(1):89–96, 2007. [Cited on page 1.]
  84. 84.Spirtes, P., Glymour, C., and Scheines, R. Causation, Prediction, and Search, volume 1. 2001. [Cited on page 6.]
  85. 85.Tangemann, M., Schneider, S., von Kügelgen, J., Locatello, F., Gehler, P., Brox, T., Kümmerer, M., Bethge, M., and Schölkopf, B. Unsupervised object learning via common fate. In 2nd Conference on Causal Learning and Reasoning (CLeaR), 2023. [Cited on pages 6 and 9.]
  86. 86.Tenenbaum, J. B., Kemp, C., Griffiths, T. L., and Goodman, N. D. How to grow a mind: Statistics, structure, and abstraction. Science, 331:1279 – 1285, 2011. [Cited on page 1.]
  87. 87.Träuble, F., Creager, E., Kilbertus, N., Locatello, F., Dittadi, A., Goyal, A., Schölkopf, B., and Bauer, S. On disentangled representations learned from correlated data. In ICML, volume 139 of Proceedings of Machine Learning Research, pp. 10401–10412, 2021. [Cited on page 6.]
  88. 88.van Steenkiste, S., Chang, M., Greff, K., and Schmidhuber, J. Relational neural expectation maximization: Unsupervised discovery of objects and their interactions. In ICLR (Poster), 2018. [Cited on page 6.]
  89. 89.von Kügelgen, J., Ustyuzhaninov, I., Gehler, P., Bethge, M., and Schölkopf, B. Towards causal generative scene models via competition of experts. In ICLR 2020 Workshop ”Causal Learning for Decision Making”. [Cited on page 6.]
  90. 90.von Kügelgen, J., Sharma, Y., Gresele, L., Brendel, W., Schölkopf, B., Besserve, M., and Locatello, F. Self-supervised learning with data augmentations provably isolates content from style. In NeurIPS, pp. 16451–16467, 2021. [Cited on page 7.]
  91. 91.Watters, N., Matthey, L., Borgeaud, S., Kabra, R., and Lerchner, A. Spriteworld: A flexible, configurable reinforcement learning environment. https://github.com/deepmind/spriteworld/, 2019. [Cited on pages 8 and 22.]
  92. 92.Weis, M. A., Chitta, K., Sharma, Y., Brendel, W., Bethge, M., Geiger, A., and Ecker, A. S. Benchmarking unsupervised object representations for video sequences. J. Mach. Learn. Res., 22:183:1–183:61, 2021. [Cited on page 1.]
  93. 93.Yang, X., Wang, Y., Sun, J., Zhang, X., Zhang, S., Li, Z., and Yan, J. Nonlinear ICA using volume-preserving transformations. In ICLR, 2022. [Cited on page 7.]
  94. 94.Yang, Y. and Yang, B. Promising or elusive? unsupervised object segmentation from real-world single images. In NeurIPS, 2022. [Cited on page 1.]
  95. 95.Zheng, Y., Ng, I., and Zhang, K. On the identifiability of nonlinear ICA: Sparsity and beyond. In Advances in Neural Information Processing Systems, 2022. [Cited on page 7.]
  96. 96.Zimmermann, R. S., Sharma, Y., Schneider, S., Bethge, M., and Brendel, W. Contrastive learning inverts the data generating process. In ICML, volume 139 of Proceedings of Machine Learning Research, pp. 12979–12990, 2021. [Cited on pages 2 and 7.]

Citation

MLA
Brady, J., et al. “Provably Learning Object-Centric Representations”. International Conference on Machine Learning, vol. 202, 2023, pp. 3038–62, https://proceedings.mlr.press/v202/brady23a.html.
APA
Brady, J., Zimmermann, R. S., Sharma, Y., Schölkopf, B., Kügelgen, J. V., & Brendel, W. (2023). Provably Learning Object-Centric Representations. International Conference on Machine Learning, 202, 3038–3062. https://proceedings.mlr.press/v202/brady23a.html
Chicago
Brady, J., R. S. Zimmermann, Y. Sharma, B. Schölkopf, J. V. Kügelgen, and W. Brendel. 2023. “Provably Learning Object-Centric Representations”. International Conference on Machine Learning 202: 3038–62. https://proceedings.mlr.press/v202/brady23a.html.
Harvard
Brady, J. et al. (2023) “Provably Learning Object-Centric Representations”, International Conference on Machine Learning. PMLR, pp. 3038–3062. Available at: https://proceedings.mlr.press/v202/brady23a.html.
Vancouver
1. Brady J, Zimmermann RS, Sharma Y, Schölkopf B, Kügelgen JV, Brendel W (2023) Provably Learning Object-Centric Representations. In: International Conference on Machine Learning. PMLR, pp 3038–3062

BibTeX

@InProceedings{pmlr-v202-brady23a,
  title = 	 {Provably Learning Object-Centric Representations},
  author =       {Brady, Jack and Zimmermann, Roland S. and Sharma, Yash and Sch\"{o}lkopf, Bernhard and Von K\"{u}gelgen, Julius and Brendel, Wieland},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {3038--3062},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/brady23a/brady23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/brady23a.html},
  abstract = 	 {Learning structured representations of the visual world in terms of objects promises to significantly improve the generalization abilities of current machine learning models. While recent efforts to this end have shown promising empirical progress, a theoretical account of when unsupervised object-centric representation learning is possible is still lacking. Consequently, understanding the reasons for the success of existing object-centric methods as well as designing new theoretically grounded methods remains challenging. In the present work, we analyze when object-centric representations can provably be learned without supervision. To this end, we first introduce two assumptions on the generative process for scenes comprised of several objects, which we call compositionality and irreducibility. Under this generative process, we prove that the ground-truth object representations can be identified by an invertible and compositional inference model, even in the presence of dependencies between objects. We empirically validate our results through experiments on synthetic data. Finally, we provide evidence that our theory holds predictive power for existing object-centric models by showing a close correspondence between models’ compositionality and invertibility and their empirical identifiability.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/