Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language

Zhenlin XuMarc NiethammerColin Raffel

article2022NeurIPS55 citations

Demonstrates through systematic evaluation that emergent language models yield superior compositional generalization compared to standard disentangled representation methods like β-VAEs, which unexpectedly hurt generalization as disentanglement pressure increases.

Listen

Modern artificial intelligence systems struggle with compositional generalization, which is the ability to recognize or produce unseen combinations of familiar elementary concepts, such as identifying a known shape in a previously unseen color. Because manual annotation of all possible attribute combinations is prohibitively expensive, machine learning researchers frequently rely on unsupervised learning algorithms designed with inductive biases toward compositionality. The article evaluates how effectively these unsupervised representations enable downstream models to generalize to novel attribute combinations when trained with limited supervision.

The article systematically evaluated three unsupervised algorithms: two standard disentanglement approaches (beta-variational autoencoders and beta-total correlation variational autoencoders) and emergent language models, which communicate visual information via sequences of discrete tokens. The evaluation protocol trained each model on unlabeled image sets (dSprites and MPI3D-Real) with a 1:9 train-to-test split, ensuring the test set contained only unseen combinations of generative factors. The researchers then froze the learned representations and trained simple linear readout models using very few labeled samples (between 100 and 1,000) to predict underlying visual factors on the test data.

The analysis yielded three principal findings. First, extracting representations from the intermediate layers immediately before or after the bottleneck achieved superior generalization compared to using the bottleneck latent variables themselves. Second, standard disentanglement and compositionality metrics showed little or no positive correlation with actual downstream generalization; in fact, increasing disentanglement pressure systematically degraded downstream generalization. Third, emergent language models consistently demonstrated the strongest compositional generalization across both classification and regression tasks, maintaining robust performance even when unsupervised training data was reduced to 5% and labeled samples were highly restricted.

These findings indicate that existing benchmarks and disentanglement metrics may be misaligned with the real-world goal of generalization. Optimizing models strictly for human-interpretable factor separation can impair practical transferability, whereas language-like discrete communication bottlenecks offer a more resilient and label-efficient path for downstream deployment.

Organizations developing computer vision and representation learning pipelines should prioritize emergent language mechanisms and discrete communication bottlenecks over traditional disentanglement methods for tasks requiring out-of-distribution combinatorial reasoning. Furthermore, teams should evaluate feature representations across intermediate layers rather than restricting downstream tasks to bottleneck representations. Future efforts should validate these findings on complex multi-object datasets, investigate advanced architectures such as Transformers, and explore alternative pre-training tasks beyond basic image reconstruction.

arXiv: 2210.00482
Cover for Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language

Abstract

Deep learning models struggle with compositional generalization, i.e. the ability to recognize or generate novel combinations of observed elementary concepts. In hopes of enabling compositional generalization, various unsupervised learning algorithms have been proposed with inductive biases that aim to induce compositional structure in learned representations (e.g. disentangled representation and emergent language learning). In this work, we evaluate these unsupervised learning algorithms in terms of how well they enable compositional generalization. Specifically, our evaluation protocol focuses on whether or not it is easy to train a simple model on top of the learned representation that generalizes to new combinations of compositional factors. We systematically study three unsupervised representation learning algorithms – β-VAE, β-TCVAE, and emergent language (EL) autoencoders – on two datasets that allow directly testing compositional generalization. We find that directly using the bottleneck representation with simple models and few labels may lead to worse generalization than using representations from layers before or after the learned representation itself. In addition, we find that the previously proposed metrics for evaluating the levels of compositionality are not correlated with the actual compositional generalization in our framework. Surprisingly, we find that increasing pressure to produce a disentangled representation (e.g. increasing β in the β-VAE) produces representations with worse generalization, while representations from EL models show strong compositional generalization. Motivated by this observation, we further investigate the advantages of using EL to induce compositional structure in unsupervised representation learning, finding that it shows consistently stronger generalization than disentanglement models, especially when using less unlabeled data for unsupervised learning and fewer labels for downstream tasks. Taken together, our results shed new light onto the compositional generalization behavior of different unsupervised learning algorithms with a new setting to rigorously test this behavior, and suggest the potential benefits of developing EL learning algorithms for more generalizable representations.

Table of Contents

  • 1 Introduction
  • 2 Unsupervised Learning with Compositional Representation Inductive Bias
  • 2.1 Learning Disentangled Representations
  • 2.2 Learning Emergent Language
  • 3 Experimental Design
  • 3.1 Datasets
  • 3.2 Compositional Generalization Evaluation Protocol
  • 3.3 Implementation details.
  • 4 Key Studies and Results
  • 4.1 Compositional latent variables may not be the best representations for downstream tasks
  • 4.2 Compositionality Metrics May Not Represent Generalization Performance
  • 4.3 Representations Learned by Emergent Language Models Generalize Better
  • 4.4 Ablations on Emergent Language Models
  • 5 Limitations of Our Study
  • 6 Related Work
  • 7 Conclusions and Discussions
  • Acknowledgement
  • References
  • Checklist

Knowls

  1. Knowl 1 — Protocol for Evaluating Compositional Generalization in Unsupervised Representation Learning

    experimental setup

    To evaluate how well unsupervised representations generalize to novel combinations of seen elementary concepts, a four-stage evaluation pipeline is defined:

    1. Data Partitioning: The total dataset space is formed by the Cartesian product of ngenn_{\text{gen}} independent generative factor spaces F={f1,…,fngen}F = \{f_1, \dots, f_{n_{\text{gen}}}\}, giving cardinality ∣D∣=∏i=1ngen∣Si∣|D| = \prod_{i=1}^{n_{\text{gen}}} |S_i|, where SiS_i is the space of factor fif_i. The dataset is partitioned randomly into train and test splits (such as a 1:9 split where 10% is used for training) under the constraint that all test samples represent novel combinations of factor values that were observed in the training split.

    2. Unsupervised Representation Learning: A representation learning model is trained on unlabeled images from the training split of size NtrainN_{\text{train}}.

    3. Downstream Task Training: The parameters of the unsupervised model are frozen. Using a very small labeled subset from the training split (Nlabel≪NtrainN_{\text{label}} \ll N_{\text{train}}, e.g., Nlabel∈{100,500,1000}N_{\text{label}} \in \{100, 500, 1000\}), simple linear models (ridge regression and logistic regression) are trained to predict the ground-truth value of each generative factor fif_i.

    4. Compositional Generalization Evaluation: The trained linear models are evaluated on the test set containing novel factor combinations. The evaluation metrics are classification accuracy for discrete attributes and the coefficient of determination (R2R^2 score, clipped below at zero so that R2=0R^2 = 0 corresponds to random guessing) for regression tasks.

    Representations are extracted and evaluated across three distinct model locations:

    • zprez_{\text{pre}}: intermediate representations immediately preceding the bottleneck encoder/module.
    • zlatentz_{\text{latent}}: the bottleneck latent representation (e.g., latent Gaussian means or discrete message tokens).
    • zpostz_{\text{post}}: intermediate representations immediately following the bottleneck decoder/module.
  2. Knowl 2 — Emergent Language Autoencoder for Unsupervised Visual Representations

    model/method

    The Emergent Language (EL) autoencoder uses a two-agent speaker-listener architecture optimized for image reconstruction:

    • Speaker (Encoder): Given an input image xx, a convolutional encoder EncConv(x)\text{EncConv}(x) produces an initial cell state for an LSTM encoder (EncLSTM\text{EncLSTM}). The EncLSTM\text{EncLSTM} generates a message m={m1,m2,… }m = \{m^1, m^2, \dots\} autoregressively up to a maximum length nmsgn_{\text{msg}} from a discrete vocabulary V={c1,…,cnV}V = \{c_1, \dots, c_{n_V}\} of size nVn_V, with one symbol designated as the end-of-sequence (EOS) token:

    q(mt∣x)=EncLSTM(mt−1∣EncConv(x),emb(m1),…,emb(mt−2))q(m^t \mid x) = \text{EncLSTM}\left(m^{t-1} \mid \text{EncConv}(x), \text{emb}(m^1), \dots, \text{emb}(m^{t-2})\right)

    where emb(⋅)\text{emb}(\cdot) maps discrete tokens to high-dimensional embeddings. Sampling from q(mt∣x)q(m^t \mid x) uses the Gumbel-Softmax distribution combined with the Straight-Through (ST) gradient estimator:

    mt=ST(GumbelSoftmax(q(mt∣x)))m^t = \text{ST}\left(\text{GumbelSoftmax}\left(q(m^t \mid x)\right)\right)

    • Listener (Decoder): A decoding LSTM (DecLSTM\text{DecLSTM}) recurrently processes token embeddings up to step T=min⁡(nmsg,argmini{mi=EOS})T = \min(n_{\text{msg}}, \text{argmin}_i \{m^i = \text{EOS}\}). The final recurrent state Emb(m)=EmbT(m)\text{Emb}(m) = \text{Emb}^T(m) is fed to a convolutional decoder (DecConv\text{DecConv}) to reconstruct the image:

    x^=DecConv(Emb(m))\hat{x} = \text{DecConv}(\text{Emb}(m))

    The system is trained end-to-end using a Bernoulli reconstruction loss on pixel values scaled to [0,1][0, 1].

  3. Knowl 3 — Formulations and Loss Functions of Beta-VAE and Beta-TCVAE

    model/method

    Unsupervised disentangled representation learning commonly modifies the evidence lower bound (ELBO) objective of variational autoencoders (VAEs) to enforce factor independence across latent variables zz:

    1. β\beta-VAE Objective:

    Lβ-VAE=Ep(x)[Eqϕ(z∣x)[log⁡pθ(x∣z)]−β KL(qϕ(z∣x)∥p(z))]\mathcal{L}_{\beta\text{-VAE}} = \mathbb{E}_{p(x)}\left[\mathbb{E}_{q_\phi(z \mid x)}[\log p_\theta(x \mid z)] - \beta \, \text{KL}(q_\phi(z \mid x) \parallel p(z))\right]

    where p(z)p(z) is the standard Gaussian prior, qϕ(z∣x)q_\phi(z \mid x) is the variational posterior parameterized by encoder parameters ϕ\phi, pθ(x∣z)p_\theta(x \mid z) is parameterized by decoder parameters θ\theta, and β≥0\beta \ge 0 controls the regularizing constraint on bottleneck capacity (with β=1\beta = 1 corresponding to standard VAE, and β=0\beta = 0 to an unregularized deterministic autoencoder).

    1. β\beta-TCVAE Objective: Decomposes the average Kullback-Leibler divergence into mutual information, total correlation (TC), and dimension-wise KL terms:

    Lβ-TCVAE=Eq(z∣x)p(x)[log⁡pθ(x∣z)]−αIq(x;z)−β KL(qϕ(z) ∥ ∏jqϕ(zj))−γ∑jKL(q(zj)∥p(zj))\mathcal{L}_{\beta\text{-TCVAE}} = \mathbb{E}_{q(z \mid x)p(x)}[\log p_\theta(x \mid z)] - \alpha I_q(x; z) - \beta \, \text{KL}\left(q_\phi(z) \,\Big\|\, \prod_j q_\phi(z_j)\right) - \gamma \sum_j \text{KL}(q(z_j) \parallel p(z_j))

    where α=1\alpha = 1, γ=1\gamma = 1, and β\beta tunes the total correlation penalty across latent dimensions zjz_j to penalize statistical dependencies among latent variables.

  4. Knowl 4 — Suboptimality of Bottleneck Latents vs Intermediate Representations for Generalization

    empirical result

    When evaluating downstream compositional generalization via linear models trained on frozen representations:

    • In Emergent Language (EL) models, the discrete bottleneck tokens zlatentz_{\text{latent}} perform poorly with linear readouts due to non-linear sequential structure. However, the post-bottleneck decoded representation zpostz_{\text{post}} consistently outperforms the pre-bottleneck representation zprez_{\text{pre}}, demonstrating that the discrete emergent language communication bottleneck forces the recurrent listener to produce representations with strong linear compositional structure.

    • In β\beta-VAE and β\beta-TCVAE models, increasing the regularization parameter β\beta (reducing bottleneck bandwidth to enforce disentanglement) degrades compositional generalization. The unregularized autoencoder baseline (β=0\beta = 0) consistently outperforms β>0\beta > 0 models on downstream classification and regression tasks. Moreover, zlatentz_{\text{latent}} is rarely optimal: regression tasks generalize better on zprez_{\text{pre}}, while classification tasks favor zpostz_{\text{post}}.

  5. Knowl 5 — Disconnection Between Disentanglement/Compositionality Metrics and Compositional Generalization

    empirical result

    Quantitative metrics designed to evaluate the degree of disentanglement or compositionality do not positively correlate with downstream compositional generalization performance on novel factor combinations:

    • For β\beta-VAE and β\beta-TCVAE, disentanglement metrics—including Separated Attribute Predictability (SAP), Interventional Robustness Score (IRS), Disentanglement-Completeness-Informativeness (DCI), and Mutual Information Gap (MIG)—increase with higher values of β\beta on benchmarks like dSprites and MPI3D-Real. However, downstream generalization accuracy and R2R^2 scores decrease as β\beta increases, producing an inverse or non-existent relationship between high metric scores and generalization performance.

    • For Emergent Language models, Topographical Similarity (TopSim)—which measures the Spearman rank correlation ρSpearman\rho_{\text{Spearman}} between input attribute cosine distances and message edit distances—exhibits weak correlation (ρSpearman<0.5\rho_{\text{Spearman}} < 0.5) with actual compositional generalization on unseen factor combinations.

  6. Knowl 6 — Generalization and Sample-Efficiency Advantages of Emergent Language Representations

    empirical result

    Post-bottleneck representations (zpostz_{\text{post}}) from Emergent Language (EL) autoencoders generalize substantially better than representations from β\beta-VAE, β\beta-TCVAE, and unregularized autoencoders (β=0\beta = 0) across several data regimes:

    • Few-shot downstream labels: When downstream linear models are trained with very few labeled samples (Nlabel=100N_{\text{label}} = 100 and Nlabel=500N_{\text{label}} = 500 on dSprites), zpostz_{\text{post}} of the EL model (nV=256n_V = 256) strictly outperforms all representation layers (zpre,zlatent,zpostz_{\text{pre}}, z_{\text{latent}}, z_{\text{post}}) of β\beta-VAE and β\beta-TCVAE on both classification accuracy and regression R2R^2 scores.

    • Reduced unlabeled pretraining data: When the unlabeled training split on MPI3D-Real is reduced from 10% to 5% of the total dataset, the generalization performance of the 00-VAE baseline degrades rapidly, whereas the EL representation maintains high performance across both classification and regression tasks, widening the performance gap.

  7. Knowl 7 — Role of Stochastic Sampling and Variable Message Length in Emergent Language Models

    empirical result

    Ablations on design components of Emergent Language autoencoders on the MPI3D-Real dataset indicate that stochasticity and variable sequence length are critical for learning generalizable representations:

    • Variable vs. Fixed Length: Restricting messages to fixed length without an end-of-sequence token (EL-fix) degrades downstream classification accuracy and regression R2R^2 scores compared to variable-length models.
    • Stochastic vs. Deterministic Sampling: Enforcing deterministic message generation via greedy argmax sampling (EL-fix-det, which operates similarly to a Vector Quantized VAE / VQ-VAE) causes a severe drop in downstream generalization performance relative to stochastic Gumbel-Softmax sampling.

    Although variable-length models converge to using maximum message length to minimize autoencoder reconstruction error, allowing stochasticity and flexible token generation remains necessary for learning compositionally robust representations.

  8. Knowl 8 — Impact of Message Bandwidth, Sequence Length, and Vocabulary Size in Emergent Language Models

    empirical result

    The communication capacity of Emergent Language models is parameterized by the total bandwidth in bits, computed as:

    Bandwidth (bits)=log⁡2(nVnmsg)\text{Bandwidth (bits)} = \log_2\left(n_V^{n_{\text{msg}}}\right)

    where nmsg∈{8,10,12}n_{\text{msg}} \in \{8, 10, 12\} is maximum message length and nV∈{128,256,512}n_V \in \{128, 256, 512\} is vocabulary size.

    Empirical evaluations on MPI3D-Real demonstrate that:

    1. Increasing the total bit bandwidth generally improves downstream compositional generalization accuracy and R2R^2 scores.
    2. When holding bit bandwidth approximately equal, shorter message sequences with larger vocabularies (e.g., vocabulary nV=512n_V = 512 with length nmsg=8n_{\text{msg}} = 8) achieve better compositional generalization than longer sequences with smaller vocabularies (e.g., nV=128n_V = 128 with nmsg=10n_{\text{msg}} = 10).
  9. Knowl 9 — Scope and Methodological Limitations in Evaluating Compositional Representations

    limitation

    The findings are subject to several experimental and methodological boundaries:

    1. Dataset Simplicity: Experiments were conducted on controlled synthetic and isolated physical environments (dSprites with 5 factors and MPI3D-Real with 7 factors), where factors follow uniform independent distributions and scenes contain single objects.
    2. Fixed Architectures: The study held visual backbones constant to symmetric CNN-MLP and CNN-LSTM models to isolate the effect of representation bottlenecks, without evaluating attention-based architectures such as Transformers.
    3. Pretraining Objective: Pretraining was restricted exclusively to autoencoding pixel reconstruction, without evaluating self-supervised contrastive objectives or multi-agent goal-oriented communication games.
    4. Optimization Hyperparameters: Optimizer configurations (Adam with fixed learning rate 10−410^{-4} and batch size 64) were kept fixed across runs rather than swept alongside representation bottleneck parameters.

Coverage note — None was omitted; the knowls cover all core contributed algorithmic designs, evaluation protocols, empirical findings on disentanglement vs. emergent language, ablations, and stated limitations.

References

  1. 1.ANDREAS, Jacob: Measuring Compositionality in Representation Learning. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=HJz05o0qK7, 2019
  2. 2.BENGIO, Yoshua ; LÉONARD, Nicholas ; COURVILLE, Aaron: Estimating or propagating gradients through stochastic neurons for conditional computation. In: arXiv preprint arXiv:1308.3432 (2013)
  3. 3.BIEDERMAN, Irving: Recognition-by-components: a theory of human image understanding. In: Psychological review 94 (1987), Nr. 2, S. 115
  4. 4.BRIGHTON, Henry ; KIRBY, Simon: Understanding linguistic evolution by visualizing the emergence of topographic mappings. In: Artificial life 12 (2006), Nr. 2, S. 229–242
  5. 5.BURGESS, Christopher P. ; HIGGINS, Irina ; PAL, Arka ; MATTHEY, Loic ; WATTERS, Nick ; DESJARDINS, Guillaume ; LERCHNER, Alexander: Understanding disentangling in β-VAE. In: arXiv preprint arXiv:1804.03599 (2018)
  6. 6.CHAABOUNI, Rahma ; KHARITONOV, Eugene ; BOUCHACOURT, Diane ; DUPOUX, Emmanuel ; BARONI, Marco: Compositionality and Generalization In Emergent Languages. In: JURAFSKY, Dan (Hrsg.) ; CHAI, Joyce (Hrsg.) ; SCHLUTER, Natalie (Hrsg.) ; TETREAULT, Joel R. (Hrsg.): Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, Association for Computational Linguistics, 2020, S. 4427–4442. – URL https://doi.org/10.18653/v1/2020.acl-main.407
  7. 7.CHEN, Ricky T. Q. ; LI, Xuechen ; GROSSE, Roger B. ; DUVENAUD, David K.: Isolating Sources of Disentanglement in Variational Autoencoders. In: BENGIO, S. (Hrsg.) ; WALLACH, H. (Hrsg.) ; LAROCHELLE, H. (Hrsg.) ; GRAUMAN, K. (Hrsg.) ; CESA-BIANCHI, N. (Hrsg.) ; GARNETT, R. (Hrsg.): Advances in Neural Information Processing Systems Bd. 31, Curran Associates, Inc., 2018, S. 2610–2620. – URL https://proceedings.neurips.cc/paper/2018/file/1ee3dfcd8a0645a25a35977997223d22-Paper.pdf
  8. 8.CHEN, Ting ; KORNBLITH, Simon ; NOROUZI, Mohammad ; HINTON, Geoffrey: A simple framework for contrastive learning of visual representations. In: International conference on machine learning PMLR (Veranst.), 2020, S. 1597–1607
  9. 9.CHRUPAŁA, Grzegorz ; KÁDÁR, Ákos ; ALISHAHI, Afra: Learning language through pictures. In: Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), 2015, S. 112–118
  10. 10.DITTADI, Andrea ; TRÄUBLE, Frederik ; LOCATELLO, Francesco ; WUTHRICH, Manuel ; AGRAWAL, Vaibhav ; WINTHER, Ole ; BAUER, Stefan ; SCHÖLKOPF, Bernhard: On the Transfer of Disentangled Representations in Realistic Settings. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=8VXvj1QNRl1, 2021
  11. 11.EASTWOOD, Cian ; WILLIAMS, Christopher K. I.: A framework for the quantitative evaluation of disentangled representations. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=By-7dz-AZ, 2018
  12. 12.ESMAEILI, Babak ; WU, Hao ; JAIN, Sarthak ; BOZKURT, Alican ; SIDDHARTH, Narayanaswamy ; PAIGE, Brooks ; BROOKS, Dana H. ; DY, Jennifer ; MEENT, Jan-Willem: Structured disentangled representations. In: The 22nd International Conference on Artificial Intelligence and Statistics PMLR (Veranst.), 2019, S. 2525–2534
  13. 13.FELZENSZWALB, Pedro F. ; GIRSHICK, Ross B. ; MCALLESTER, David ; RAMANAN, Deva: Object detection with discriminatively trained part-based models. In: IEEE transactions on pattern analysis and machine intelligence 32 (2009), Nr. 9, S. 1627–1645
  14. 14.FIDLER, Sanja ; LEONARDIS, Ales: Towards scalable representations of object categories: Learning a hierarchy of parts. In: 2007 IEEE Conference on Computer Vision and Pattern Recognition IEEE (Veranst.), 2007, S. 1–8
  15. 15.FODOR, Jerry A.: The language of thought. Bd. 5. Harvard university press, 1975
  16. 16.FODOR, Jerry A. ; PYLYSHYN, Zenon W.: Connectionism and cognitive architecture: A critical analysis. In: Cognition 28 (1988), Nr. 1-2, S. 3–71
  17. 17.GONDAL, Muhammad W. ; WUTHRICH, Manuel ; MILADINOVIC, Djordje ; LOCATELLO, Francesco ; BREIDT, Martin ; VOLCHKOV, Valentin ; AKPO, Joel ; BACHEM, Olivier ; SCHÖLKOPF, Bernhard ; BAUER, Stefan: On the Transfer of Inductive Bias from Simulation to the Real World: a New Disentanglement Dataset. In: WALLACH, H. (Hrsg.) ; LAROCHELLE, H. (Hrsg.) ; BEYGELZIMER, A. (Hrsg.) ; ALCHÉ-BUC, F. d'(Hrsg.) ; FOX, E. (Hrsg.) ; GARNETT, R. (Hrsg.): Advances in Neural Information Processing Systems Bd. 32, Curran Associates, Inc., 2019. – URL https://proceedings.neurips.cc/paper/2019/file/d97d404b6119214e4a7018391195240a-Paper.pdf
  18. 18.GRILL, Jean-Bastien ; STRUB, Florian ; ALTCHÉ, Florent ; TALLEC, Corentin ; RICHEMOND, Pierre ; BUCHATSKAYA, Elena ; DOERSCH, Carl ; AVILA PIRES, Bernardo ; GUO, Zhaohan ; GHESHLAGHI AZAR, Mohammad u. a.: Bootstrap your own latent-a new approach to self-supervised learning. In: Advances in neural information processing systems 33 (2020), S. 21271–21284
  19. 19.HAVRYLOV, Serhii ; TITOV, Ivan: Emergence of language with multi-agent games: Learning to communicate with sequences of symbols. In: Advances in neural information processing systems, 2017, S. 2149–2159
  20. 20.HIGGINS, Irina ; MATTHEY, Loic ; PAL, Arka ; BURGESS, Christopher ; GLOROT, Xavier ; BOTVINICK, Matthew ; MOHAMED, Shakir ; LERCHNER, Alexander: β-VAE: Learning basic visual concepts with a constrained variational framework. In: International Conference on Learning Representations, 2017
  21. 21.HIGGINS, Irina ; SONNERAT, Nicolas ; MATTHEY, Loic ; PAL, Arka ; BURGESS, Christopher P. ; BOŠNJAK, Matko ; SHANAHAN, Murray ; BOTVINICK, Matthew ; HASSABIS, Demis ; LERCHNER, Alexander: SCAN: Learning Hierarchical Compositional Visual Concepts. In: International Conference on Learning Representations, 2018
  22. 22.HOFFMAN, Donald D. ; RICHARDS, Whitman A.: Parts of recognition. In: Cognition 18 (1984), Nr. 1-3, S. 65–96
  23. 23.JANG, Eric ; GU, Shixiang ; POOLE, Ben: Categorical reparameterization with gumbel-softmax. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=rkE3y85ee, 2017
  24. 24.KIM, Hyunjik ; MNIH, Andriy: Disentangling by factorising. In: International Conference on Machine Learning PMLR (Veranst.), 2018, S. 2649–2658
  25. 25.KINGMA, Diederik P. ; WELLING, Max: Auto-encoding variational bayes. In: arXiv preprint arXiv:1312.6114 (2013)
  26. 26.KINGMA, Durk P. ; MOHAMED, Shakir ; JIMENEZ REZENDE, Danilo ; WELLING, Max: Semi-supervised learning with deep generative models. In: Advances in neural information processing systems 27 (2014)
  27. 27.KUMAR, Abhishek ; SATTIGERI, Prasanna ; BALAKRISHNAN, Avinash: VARIATIONAL INFERENCE OF DISENTANGLED LATENT CONCEPTS FROM UNLABELED OBSERVATIONS. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=H1kG7GZAW, 2018
  28. 28.LAKE, Brenden M. ; ULLMAN, Tomer D. ; TENENBAUM, Joshua B. ; GERSHMAN, Samuel J.: Building machines that learn and think like people. In: Behavioral and brain sciences 40 (2017)
  29. 29.LAZARIDOU, Angeliki ; PHAM, Nghia T. ; BARONI, Marco: Towards multi-agent communication-based language learning. In: arXiv preprint arXiv:1605.07133 (2016)
  30. 30.LIN, Zhixuan ; WU, Yi-Fu ; PERI, Skand V. ; SUN, Weihao ; SINGH, Gautam ; DENG, Fei ; JIANG, Jindong ; AHN, Sungjin: SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=rkl03ySYDH, 2020
  31. 31.LOCATELLO, Francesco ; BAUER, Stefan ; LUCIC, Mario ; RAETSCH, Gunnar ; GELLY, Sylvain ; SCHÖLKOPF, Bernhard ; BACHEM, Olivier: Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations. In: International Conference on Machine Learning, 2019, S. 4114–4124
  32. 32.MADDISON, Chris J. ; MNIH, Andriy ; TEH, Yee W.: The concrete distribution: A continuous relaxation of discrete random variables. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=S1jE5L5gl, 2017
  33. 33.MARCUS, Gary: Deep learning: A critical appraisal. In: arXiv preprint arXiv:1801.00631 (2018)
  34. 34.MATHIEU, Emile ; RAINFORTH, Tom ; SIDDHARTH, Nana ; TEH, Yee W.: Disentangling disentanglement in variational autoencoders. In: International Conference on Machine Learning PMLR (Veranst.), 2019, S. 4402–4412
  35. 35.MATTHEY, Loic ; HIGGINS, Irina ; HASSABIS, Demis ; LERCHNER, Alexander: dSprites: Disentanglement testing Sprites dataset. https://github.com/deepmind/dsprites-dataset/. 2017
  36. 36.MONTERO, Milton L. ; LUDWIG, Casimir J. ; COSTA, Rui P. ; MALHOTRA, Gaurav ; BOWERS, Jeffrey: The role of Disentanglement in Generalisation. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=qbH974jKUVy, 2021
  37. 37.OTT, Patrick ; EVERINGHAM, Mark: Shared parts for deformable part-based models. In: CVPR 2011 IEEE (Veranst.), 2011, S. 1513–1520
  38. 38.PANDEY, Megha ; LAZEBNIK, Svetlana: Scene recognition and weakly supervised object localization with deformable part-based models. In: 2011 International Conference on Computer Vision IEEE (Veranst.), 2011, S. 1307–1314
  39. 39.PURUSHWALKAM, Senthil ; NICKEL, Maximilian ; GUPTA, Abhinav ; RANZATO, Marc’Aurelio: Task-driven modular networks for zero-shot compositional learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, S. 3593–3602
  40. 40.REN, Yi ; GUO, Shangmin ; LABEAU, Matthieu ; COHEN, Shay B. ; KIRBY, Simon: Compositional languages emerge in a neural iterated learning model. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=HkePNpVKPB, 2020
  41. 41.SCHOTT, Lukas ; KÜGELGEN, Julius V. ; TRÄUBLE, Frederik ; GEHLER, Peter V. ; RUSSELL, Chris ; BETHGE, Matthias ; SCHÖLKOPF, Bernhard ; LOCATELLO, Francesco ; BRENDEL, Wieland: Visual Representation Learning Does Not Generalize Strongly Within the Same Domain. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=9RUHPlladgh, 2022
  42. 42.STONE, Austin ; WANG, Huayan ; STARK, Michael ; LIU, Yi ; SCOTT PHOENIX, D ; GEORGE, Dileep: Teaching compositionality to cnns. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, S. 5058–5067
  43. 43.SUTER, Raphael ; MILADINOVIC, Djordje ; SCHÖLKOPF, Bernhard ; BAUER, Stefan: Robustly disentangled causal mechanisms: Validating deep representations for interventional robustness. In: International Conference on Machine Learning PMLR (Veranst.), 2019, S. 6056–6065
  44. 44.THRUSH, Tristan ; JIANG, Ryan ; BARTOLO, Max ; SINGH, Amanpreet ; WILLIAMS, Adina ; KIELA, Douwe ; ROSS, Candace: Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality. In: arXiv preprint arXiv:2204.03162 (2022)
  45. 45.TOKMAKOV, Pavel ; WANG, Yu-Xiong ; HEBERT, Martial: Learning compositional representations for few-shot recognition. In: Proceedings of the IEEE International Conference on Computer Vision, 2019, S. 6372–6381
  46. 46.VAN DEN OORD, Aaron ; VINYALS, Oriol u. a.: Neural discrete representation learning. In: Advances in neural information processing systems 30 (2017)
  47. 47.VASWANI, Ashish ; SHAZEER, Noam ; PARMAR, Niki ; USZKOREIT, Jakob ; JONES, Llion ; GOMEZ, Aidan N. ; KAISER, Łukasz ; POLOSUKHIN, Illia: Attention is all you need. In: Advances in neural information processing systems (2017)
  48. 48.YAO, Shunyu ; YU, Mo ; ZHANG, Yang ; NARASIMHAN, Karthik R. ; TENENBAUM, Joshua B. ; GAN, Chuang: Linking Emergent and Natural Languages via Corpus Transfer. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=49A1Y6tRhaq, 2022
  49. 49.YUN, Tian ; BHALLA, Usha ; PAVLICK, Ellie ; SUN, Chen: Do Vision-Language Pretrained Models Learn Primitive Concepts? In: arXiv preprint arXiv:2203.17271 (2022)
  50. 50.ZHAO, Shengjia ; REN, Hongyu ; YUAN, Arianna ; SONG, Jiaming ; GOODMAN, Noah ; ERMON, Stefano: Bias and generalization in deep generative models: An empirical study. In: Advances in Neural Information Processing Systems 31 (2018)

Citation

MLA
Xu, Z., et al. “Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 25074–87, https://proceedings.neurips.cc/paper_files/paper/2022/file/9f9ecbf4062842df17ec3f4ea3ad7f54-Paper-Conference.pdf.
APA
Xu, Z., Niethammer, M., & Raffel, C. (2022). Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language. Advances in Neural Information Processing Systems, 35, 25074–25087. https://proceedings.neurips.cc/paper_files/paper/2022/file/9f9ecbf4062842df17ec3f4ea3ad7f54-Paper-Conference.pdf
Chicago
Xu, Z., M. Niethammer, and C. Raffel. 2022. “Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language”. Advances in Neural Information Processing Systems 35: 25074–87. https://proceedings.neurips.cc/paper_files/paper/2022/file/9f9ecbf4062842df17ec3f4ea3ad7f54-Paper-Conference.pdf.
Harvard
Xu, Z., Niethammer, M. and Raffel, C. (2022) “Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 25074–25087. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/9f9ecbf4062842df17ec3f4ea3ad7f54-Paper-Conference.pdf.
Vancouver
1. Xu Z, Niethammer M, Raffel C (2022) Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 25074–25087

BibTeX

@inproceedings{xu2022compositional,
  title = {Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language},
  author = {Xu, Zhenlin and Niethammer, Marc and Raffel, Colin},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {25074-25087},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/9f9ecbf4062842df17ec3f4ea3ad7f54-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors