Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning
Xiangyu LiXu YangKun WeiCheng DengMuli Yang
Proposes a Siamese Contrastive Embedding Network alongside a State Transition Module to disentangle state and object prototypes through contrastive learning, achieving state-of-the-art performance in recognizing unseen visual compositions across multiple benchmark datasets.
Modern computer vision systems often struggle to recognize novel combinations of known concepts, such as identifying a "sliced apple" when the system has only encountered whole apples and other sliced fruits during training. This capability, known as compositional zero-shot learning, is essential for building scalable artificial intelligence that can interpret unfamiliar real-world scenes without requiring exhaustive training data for every possible variation. The primary challenge stems from visual entanglement: the visual appearance of an attribute or state changes drastically depending on the object it modifies, creating a significant performance gap when models encounter new combinations.
The article aims to evaluate and demonstrate a novel framework called the Siamese Contrastive Embedding Network, designed to reliably recognize both previously seen and entirely unseen state-object compositions. The objective is to decouple object and state representations while generating realistic synthetic training examples to improve overall generalization.
The authors develop an approach comprising two dedicated encoders that project image features into separate contrastive spaces—one focusing purely on the state and the other on the object. To prevent the model from confusing entangled features, they structure positive and negative sample databases that isolate attributes during training. Additionally, they introduce a state transition module that pairs a generator with an adversarial discriminator to synthesize plausible, novel compositions (such as creating virtual instances of uncommon combinations) while filtering out nonsensical pairings. The framework was evaluated across three standard benchmark datasets: MIT-States (comprising over 53,000 images across 1,962 concepts), UT-Zappos (comprising over 50,000 shoe images), and the extensive C-GQA dataset (encompassing more than 9,500 concepts).
The experimental findings show that the proposed framework consistently outperforms existing state-of-the-art methods across all benchmarks. On the UT-Zappos dataset, the system increased the Area Under the Curve metric from 28.7% to 32.0% and achieved a balanced harmonic mean accuracy of 47.8%, representing an improvement of approximately 4.5 percentage points over previous top models. On MIT-States, the model reached a top Area Under the Curve score of 5.3% and lifted the harmonic mean from 17.2% to 18.4%. On the large-scale C-GQA benchmark, it achieved leading performance with a 5.5% Area Under the Curve score alongside the highest individual state (28.1%) and object (32.8%) recognition accuracies. Ablation analyses confirmed that combining contrastive embedding spaces with synthetic sample generation yielded significantly better performance than using either component alone.
These results demonstrate that explicitly separating state and object representations, reinforced with synthetically generated compositions, effectively bridges the domain gap between familiar and novel visual concepts. For decision-makers and technical leaders, this approach reduces the data acquisition costs and operational risks associated with deploying visual recognition models in dynamic, unconstrained environments where rare or novel combinations frequently appear.
Organizations developing automated visual inspection, categorization, or search systems should consider adopting decoupled representation techniques and adversarial synthetic data generation to enhance model flexibility. Future efforts should focus on transitioning evaluation protocols to multi-label frameworks, as real-world objects often possess multiple valid attributes simultaneously (such as texture, color, and age) that single-label benchmarks penalize as classification errors.
The findings are supported by consistent empirical improvements across three diverse datasets; however, some limitations remain. Performance is inherently bounded when negative contrastive databases omit relevant categories, and single-label ground-truth annotations occasionally lead to misleading error classifications. Practitioners should account for these dataset boundary conditions when applying the model to complex, multi-attribute operating environments.
- Paper: Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly, Yongqin Xian et al. (2017). This benchmark paper establishes the standardized evaluation protocols, data splits, and generalized zero-shot learning frameworks upon which compositional zero-shot learning methodologies rely.
- Paper: An embarrassingly simple approach to zero-shot learning, Bernardino Romera-Paredes et al. (2015). It provides foundational principles for mapping visual features into semantic attribute spaces to enable zero-shot transfer across unseen concepts.
- Paper: Exploring Simple Siamese Representation Learning, Xinlei Chen et al. (2021). It introduces essential mechanisms for Siamese representation learning that inform Siamese embedding designs without collapsing representations.
- Paper: Momentum Contrast for Unsupervised Visual Representation Learning, Kaiming He et al. (2020). It develops the core contrastive representation learning paradigms and dictionary look-up formulations utilized to structure metric embedding spaces.
- Paper: Supervised Contrastive Learning, Prannay Khosla et al. (2020). It establishes supervised contrastive learning objectives that motivate leveraging discrete semantic labels and prototypes within contrastive feature spaces.
- Paper: Learning to Compare: Relation Network for Few-Shot Learning, Flood Sung et al. (2017). It lays the groundwork for episodic meta-learning and metric-based relation networks for zero-shot and few-shot visual classification.
- Paper: Compositional Generalization in Unsupervised Compositional Representation Learning: A Study on Disentanglement and Emergent Language, Zhenlin Xu et al. (2022). It investigates how unsupervised disentangled representation learning translates to compositional generalization across novel factor combinations.
- Paper: Learning Transferable Human-Object Interaction Detector with Natural Language Supervision, Suchen Wang et al. (2022). It extends compositional visual learning principles to detect unseen human-object interactions using joint visual-and-textual natural language supervision.
- Paper: T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation, Kaiyi Huang et al. (2023). It broadens the study of visual-semantic compositionality from discriminative zero-shot recognition to evaluating complex attribute and object binding in generative text-to-image models.
- Paper: MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis, Dewei Zhou et al. (2024). It tackles the inverse problem of disentangling and binding object-attribute compositions spatially during multi-instance image synthesis.
