Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language
Zhenlin XuMarc NiethammerColin Raffel
Demonstrates through systematic evaluation that emergent language models yield superior compositional generalization compared to standard disentangled representation methods like β-VAEs, which unexpectedly hurt generalization as disentanglement pressure increases.
Modern artificial intelligence systems struggle with compositional generalization, which is the ability to recognize or produce unseen combinations of familiar elementary concepts, such as identifying a known shape in a previously unseen color. Because manual annotation of all possible attribute combinations is prohibitively expensive, machine learning researchers frequently rely on unsupervised learning algorithms designed with inductive biases toward compositionality. The article evaluates how effectively these unsupervised representations enable downstream models to generalize to novel attribute combinations when trained with limited supervision.
The article systematically evaluated three unsupervised algorithms: two standard disentanglement approaches (beta-variational autoencoders and beta-total correlation variational autoencoders) and emergent language models, which communicate visual information via sequences of discrete tokens. The evaluation protocol trained each model on unlabeled image sets (dSprites and MPI3D-Real) with a 1:9 train-to-test split, ensuring the test set contained only unseen combinations of generative factors. The researchers then froze the learned representations and trained simple linear readout models using very few labeled samples (between 100 and 1,000) to predict underlying visual factors on the test data.
The analysis yielded three principal findings. First, extracting representations from the intermediate layers immediately before or after the bottleneck achieved superior generalization compared to using the bottleneck latent variables themselves. Second, standard disentanglement and compositionality metrics showed little or no positive correlation with actual downstream generalization; in fact, increasing disentanglement pressure systematically degraded downstream generalization. Third, emergent language models consistently demonstrated the strongest compositional generalization across both classification and regression tasks, maintaining robust performance even when unsupervised training data was reduced to 5% and labeled samples were highly restricted.
These findings indicate that existing benchmarks and disentanglement metrics may be misaligned with the real-world goal of generalization. Optimizing models strictly for human-interpretable factor separation can impair practical transferability, whereas language-like discrete communication bottlenecks offer a more resilient and label-efficient path for downstream deployment.
Organizations developing computer vision and representation learning pipelines should prioritize emergent language mechanisms and discrete communication bottlenecks over traditional disentanglement methods for tasks requiring out-of-distribution combinatorial reasoning. Furthermore, teams should evaluate feature representations across intermediate layers rather than restricting downstream tasks to bottleneck representations. Future efforts should validate these findings on complex multi-object datasets, investigate advanced architectures such as Transformers, and explore alternative pre-training tasks beyond basic image reconstruction.
- Paper: Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations, Francesco Locatello et al. (2018). This seminal study proves theoretical limits and challenges core assumptions about the utility of unsupervised disentangled representations for downstream tasks, forming the direct conceptual and empirical motivation for re-evaluating unsupervised compositional generalization.
- Paper: beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework, Irina Higgins et al. (2016). It introduces the beta-VAE framework and inductive bias towards disentangling generative factors, which serves as one of the primary unsupervised baseline methods evaluated in the source paper.
- Paper: Isolating Sources of Disentanglement in Variational Autoencoders, Ricky T. Q. Chen et al. (2018). It establishes the beta-TCVAE algorithm and total correlation penalty for isolating explanatory factors, which is another central disentanglement approach tested and analyzed in the source.
- Paper: Disentangling by Factorising, Hyunjik Kim et al. (2018). It develops FactorVAE and formulates total correlation minimization for factor separation, establishing key concepts and metrics evaluated in the source paper's benchmark.
- Paper: Relational inductive biases, deep learning, and graph networks, Peter W. Battaglia et al. (2018). It provides foundational theory on inductive biases and combinatorial generalization across neural architectures, motivating the pursuit of compositional representation learning.
- Paper: Representation Learning: A Review and New Perspectives, Yoshua Bengio et al. (2012). This foundational review defines the core objectives and theoretical priors of representation learning, notably the disentanglement of underlying explanatory factors.
- Paper: Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning, Xiangyu Li et al. (2022). It applies compositional representation principles to zero-shot recognition by designing dedicated contrastive encoders to decouple object and state representations for novel concept combinations.
- Paper: T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation, Kaiyi Huang et al. (2023). It extends the evaluation of compositional generalization from unsupervised visual factor readouts to open-world compositional text-to-image generation and complex prompt-binding benchmarks.
- Paper: Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMC, Yilun Du et al. (2023). It tackles compositional reasoning in generative modeling by developing MCMC sampling strategies that compose independent pre-trained concepts without retraining.
- Paper: InfoDiffusion: Representation Learning Using Information Maximizing Diffusion Models, Yingheng Wang et al. (2023). It advances unsupervised representation learning by combining diffusion models with auxiliary latent variables to simultaneously achieve factor disentanglement and high-fidelity generation.
- Paper: RoboDreamer: Learning Compositional World Models for Robot Imagination, Siyuan Zhou et al. (2024). It builds upon compositional representation learning by decomposing actions and spatial relations to enable zero-shot compositional generalization in robotic world models.
