Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning

Xiangyu LiXu YangKun WeiCheng DengMuli Yang

article2022CVPR99 citations

Proposes a Siamese Contrastive Embedding Network alongside a State Transition Module to disentangle state and object prototypes through contrastive learning, achieving state-of-the-art performance in recognizing unseen visual compositions across multiple benchmark datasets.

Listen

Modern computer vision systems often struggle to recognize novel combinations of known concepts, such as identifying a "sliced apple" when the system has only encountered whole apples and other sliced fruits during training. This capability, known as compositional zero-shot learning, is essential for building scalable artificial intelligence that can interpret unfamiliar real-world scenes without requiring exhaustive training data for every possible variation. The primary challenge stems from visual entanglement: the visual appearance of an attribute or state changes drastically depending on the object it modifies, creating a significant performance gap when models encounter new combinations.

The article aims to evaluate and demonstrate a novel framework called the Siamese Contrastive Embedding Network, designed to reliably recognize both previously seen and entirely unseen state-object compositions. The objective is to decouple object and state representations while generating realistic synthetic training examples to improve overall generalization.

The authors develop an approach comprising two dedicated encoders that project image features into separate contrastive spaces—one focusing purely on the state and the other on the object. To prevent the model from confusing entangled features, they structure positive and negative sample databases that isolate attributes during training. Additionally, they introduce a state transition module that pairs a generator with an adversarial discriminator to synthesize plausible, novel compositions (such as creating virtual instances of uncommon combinations) while filtering out nonsensical pairings. The framework was evaluated across three standard benchmark datasets: MIT-States (comprising over 53,000 images across 1,962 concepts), UT-Zappos (comprising over 50,000 shoe images), and the extensive C-GQA dataset (encompassing more than 9,500 concepts).

The experimental findings show that the proposed framework consistently outperforms existing state-of-the-art methods across all benchmarks. On the UT-Zappos dataset, the system increased the Area Under the Curve metric from 28.7% to 32.0% and achieved a balanced harmonic mean accuracy of 47.8%, representing an improvement of approximately 4.5 percentage points over previous top models. On MIT-States, the model reached a top Area Under the Curve score of 5.3% and lifted the harmonic mean from 17.2% to 18.4%. On the large-scale C-GQA benchmark, it achieved leading performance with a 5.5% Area Under the Curve score alongside the highest individual state (28.1%) and object (32.8%) recognition accuracies. Ablation analyses confirmed that combining contrastive embedding spaces with synthetic sample generation yielded significantly better performance than using either component alone.

These results demonstrate that explicitly separating state and object representations, reinforced with synthetically generated compositions, effectively bridges the domain gap between familiar and novel visual concepts. For decision-makers and technical leaders, this approach reduces the data acquisition costs and operational risks associated with deploying visual recognition models in dynamic, unconstrained environments where rare or novel combinations frequently appear.

Organizations developing automated visual inspection, categorization, or search systems should consider adopting decoupled representation techniques and adversarial synthetic data generation to enhance model flexibility. Future efforts should focus on transitioning evaluation protocols to multi-label frameworks, as real-world objects often possess multiple valid attributes simultaneously (such as texture, color, and age) that single-label benchmarks penalize as classification errors.

The findings are supported by consistent empirical improvements across three diverse datasets; however, some limitations remain. Performance is inherently bounded when negative contrastive databases omit relevant categories, and single-label ground-truth annotations occasionally lead to misleading error classifications. Practitioners should account for these dataset boundary conditions when applying the model to complex, multi-attribute operating environments.

arXiv: 2206.14475
Cover for Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning

Abstract

Compositional Zero-Shot Learning (CZSL) aims to recognize unseen compositions formed from seen state and object during training. Since the same state may be various in the visual appearance while entangled with different objects, CZSL is still a challenging task. Some methods recognize state and object with two trained classifiers, ignoring the impact of the interaction between object and state; the other methods try to learn the joint representation of the state-object compositions, leading to the domain gap between seen and unseen composition sets. In this paper, we propose a novel Siamese Contrastive Embedding Network (SCEN)1 for unseen composition recognition. Considering the entanglement between state and object, we embed the visual feature into a Siamese Contrastive Space to capture prototypes of them separately, alleviating the interaction between state and object. In addition, we design a State Transition Module (STM) to increase the diversity of training compositions, improving the robustness of the recognition model. Extensive experiments indicate that our method significantly outperforms the state-of-the-art approaches on three challenging benchmark datasets, including the recent proposed C-QGA dataset.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Approach
  • 3.1. Problem Definition
  • 3.2. Siamese Contrastive Embedding Network
  • 3.3. Inference
  • 4. Experiment
  • 4.1. Experimental Setup
  • 4.2. Comparison with State-of-the-Arts
  • 4.3. Ablation Study
  • 4.4. Qualitative Results
  • 4.5. Hyper-Parameter Analysis
  • 5. Conclusion
  • 6. Acknowledgements
  • References

Knowls

  1. Knowl 1 — Generalized compositional zero-shot learning formulation

    definition

    Compositional Zero-Shot Learning (CZSL) represents each image label as a pair of a state and an object. Let AA be the set of states, OO the set of objects, and C=A×O={(a,o)∣a∈A,o∈O}C=A\times O=\{(a,o)\mid a\in A,o\in O\} the complete set of possible compositions. The training set is Dtr={(x,c)∣x∈Xs,c∈Cs}D_{\mathrm{tr}}=\{(x,c)\mid x\in X_s,c\in C_s\}, where XsX_s is the training image set and Cs⊂CC_s\subset C contains the compositions observed during training. The unseen compositions are Cu=C∖CsC_u=C\setminus C_s.

    The paper uses the generalized CZSL setting: test images may belong to either CsC_s or CuC_u, so the model must learn a mapping from an image to a composition in Cs∪CuC_s\cup C_u, rather than predicting only unseen compositions. This setting is difficult because the same state can have different visual appearances when combined with different objects.

  2. Knowl 2 — Siamese Contrastive Embedding Network

    model/method

    The proposed Siamese Contrastive Embedding Network (SCEN) separates state and object information from an image instead of using a single entangled representation. A feature extractor FCF_C maps an input image xx to a visual feature v=FC(x)v=F_C(x). Two independent encoders then produce state and object prototypes:

    hs=Es(v),ho=Eo(v),h_s=E_s(v),\qquad h_o=E_o(v),

    where EsE_s is the State-Specific Encoder, EoE_o is the Object-Specific Encoder, hsh_s is the state prototype, and hoh_o is the object prototype. The two prototypes are trained in separate state and object contrastive embedding spaces. The architecture diagram on page 3 shows the two encoder branches and their shared use of irrelevant samples as negatives. The intended effect is to make hsh_s discriminative for the state while making hoh_o discriminative for the object, reducing the adverse influence of state-object interaction and improving transfer to unseen compositions.

  3. Knowl 3 — State-constant, object-constant, and irrelevant databases

    definition

    For a training image whose composition is (a^,o^)∈Cs(\hat a,\hat o)\in C_s, SCEN constructs three composition databases from the seen training compositions. The object-constant database contains compositions with the same object and varying states, the state-constant database contains compositions with the same state and varying objects, and the irrelevant database contains compositions with both primitives different:

    Do={(a,o)∣o=o^, (a,o)∈Cs},D_o=\{(a,o)\mid o=\hat o,\ (a,o)\in C_s\}, Ds={(a,o)∣a=a^, (a,o)∈Cs},D_s=\{(a,o)\mid a=\hat a,\ (a,o)\in C_s\}, Dir={(a,o)∣a≠a^, o≠o^, (a,o)∈Cs}.D_{\mathrm{ir}}=\{(a,o)\mid a\ne\hat a,\ o\ne\hat o,\ (a,o)\in C_s\}.

    Images associated with DsD_s provide positive examples for state-prototype learning, images associated with DoD_o provide positive examples for object-prototype learning, and images associated with DirD_{\mathrm{ir}} provide negative examples for both branches. The same irrelevant samples are shared between the two contrastive spaces so that the state and object encoders receive similarly balanced negative supervision.

  4. Knowl 4 — Siamese contrastive and classification objectives

    equation

    SCEN uses separate contrastive losses to train the state and object prototypes. Let hsh_s and hoh_o be the prototypes of a current image. Let hs+h_s^+ be a state prototype encoded from a sample in DsD_s, ho+h_o^+ an object prototype encoded from a sample in DoD_o, and hs,i−h_{s,i}^- and ho,i−h_{o,i}^- the prototypes of the iith irrelevant negative sample from DirD_{\mathrm{ir}}. With KK negative samples and temperatures τs>0\tau_s>0 and τo>0\tau_o>0, the losses are

    Lscl=−log⁡exp⁡(hsThs+/τs)exp⁡(hsThs+/τs)+∑i=1Kexp⁡(hsThs,i−/τs),\mathcal L_{\mathrm{scl}}=-\log\frac{\exp(h_s^{\mathsf T}h_s^+/\tau_s)}{\exp(h_s^{\mathsf T}h_s^+/\tau_s)+\sum_{i=1}^{K}\exp(h_s^{\mathsf T}h_{s,i}^-/\tau_s)}, Locl=−log⁡exp⁡(hoTho+/τo)exp⁡(hoTho+/τo)+∑i=1Kexp⁡(hoTho,i−/τo).\mathcal L_{\mathrm{ocl}}=-\log\frac{\exp(h_o^{\mathsf T}h_o^+/\tau_o)}{\exp(h_o^{\mathsf T}h_o^+/\tau_o)+\sum_{i=1}^{K}\exp(h_o^{\mathsf T}h_{o,i}^-/\tau_o)}.

    The positive pair is pulled together in the relevant contrastive space, while each irrelevant negative is pushed away. Fully connected classifiers CaC_a and CoC_o classify the state and object prototypes using cross-entropy:

    Lcls=Ca(hs,a)+Co(ho,o).\mathcal L_{\mathrm{cls}}=C_a(h_s,a)+C_o(h_o,o).

    The complete Siamese Contrastive Space loss is

    Lcts=Lscl+Locl+Lcls.\mathcal L_{\mathrm{cts}}=\mathcal L_{\mathrm{scl}}+\mathcal L_{\mathrm{ocl}}+\mathcal L_{\mathrm{cls}}.

    Here aa and oo are the ground-truth state and object of the current training image. The contrastive terms enforce prototype discrimination, while the classification terms preserve state and object information needed for composition recognition.

  5. Knowl 5 — State Transition Module for virtual compositions

    model/method

    The State Transition Module (STM) augments training with virtual compositions intended to reduce the distribution gap between seen and unseen compositions. A generator GG receives an object prototype hoh_o and a state prototype h~s\tilde h_s and produces a virtual image or visual feature G(h~s,ho)G(\tilde h_s,h_o). The state and object can be drawn from combinations that do not occur in the training set, such as transferring a state observed with one object to another object.

    Because arbitrary state-object combinations may be implausible, a discriminator DD distinguishes real training images xa,ox_{a,o} from generated samples. The adversarial objective is

    min⁡G,Es,Eomax⁡DV(G,D)=Exa,o[log⁡D(xa,o)]+Eh~s,ho[log⁡(1−D(G(h~s,ho)))].\min_{G,E_s,E_o}\max_D V(G,D)=\mathbb E_{x_{a,o}}[\log D(x_{a,o})]+\mathbb E_{\tilde h_s,h_o}[\log(1-D(G(\tilde h_s,h_o)))].

    The state and object encoders are also applied to each generated sample. If a~\tilde a is the state label associated with h~s\tilde h_s and oo is the object label associated with hoh_o, the reclassification constraint is

    Lclsre=Ca(Es(G(h~s,ho)),a~)+Co(Eo(G(h~s,ho)),o).\mathcal L_{\mathrm{clsre}}=C_a(E_s(G(\tilde h_s,h_o)),\tilde a)+C_o(E_o(G(\tilde h_s,h_o)),o).

    The STM objective is Lstm=max⁡Dmin⁡G,Es,EoV(G,D)+Lclsre\mathcal L_{\mathrm{stm}}=\max_D\min_{G,E_s,E_o}V(G,D)+\mathcal L_{\mathrm{clsre}}. Thus, the generator is encouraged to produce realistic virtual compositions whose re-encoded state and object prototypes remain classifiable. The STM workflow shown on page 4 combines state and object prototypes, adversarial discrimination, and reclassification of the generated sample.

  6. Knowl 6 — Joint SCEN training objective

    equation

    The complete SCEN model balances prototype learning and virtual-composition augmentation with

    Ltotal=αLcts+βLstm,\mathcal L_{\mathrm{total}}=\alpha\mathcal L_{\mathrm{cts}}+\beta\mathcal L_{\mathrm{stm}},

    where Lcts\mathcal L_{\mathrm{cts}} is the Siamese contrastive-space loss, Lstm\mathcal L_{\mathrm{stm}} is the State Transition Module loss, and α,β\alpha,\beta are scalar weighting coefficients. The contrastive component learns discriminative state and object representations from real training samples, whereas the STM component exposes the encoders to realistic virtual state-object combinations.

  7. Knowl 7 — Composition prediction at inference

    model/method

    At test time, SCEN extracts a visual feature v=FC(x)v=F_C(x) from an image xx, computes the state and object prototypes hs=Es(v)h_s=E_s(v) and ho=Eo(v)h_o=E_o(v), and scores every candidate composition (a,o)(a,o) in the generalized prediction space Cs∪CuC_s\cup C_u. The predicted composition is

    c^=(a^,o^)=arg⁡max⁡(a,o)∈Cs∪Cup(x∣Es(v),Eo(v),a,o).\hat c=(\hat a,\hat o)=\arg\max_{(a,o)\in C_s\cup C_u}p\bigl(x\mid E_s(v),E_o(v),a,o\bigr).

    The candidate with the highest joint state-object likelihood is returned, allowing the model to predict both seen and unseen compositions.

  8. Knowl 8 — Datasets, evaluation protocol, and implementation

    experimental setup

    SCEN was evaluated on MIT-States, UT-Zappos, and C-GQA under generalized CZSL. The dataset statistics reported in the paper are:

    Could not parse LaTeX table

    Here ss and oo are the numbers of states and objects, csc_s and cuc_u are seen and unseen compositions, and ii is the number of images. Performance is measured by accuracy on seen compositions, accuracy on unseen compositions, their harmonic mean (HM), and the area under the seen-unseen accuracy curve (AUC) as the calibration bias changes.

    For each image, a 1024-dimensional ResNet-18 feature pretrained on ImageNet is used. The state and object encoders each output 300-dimensional features through two fully connected layers with ReLU activation. The model is implemented in PyTorch and optimized with Adam using learning rate 0.000040.00004, batch size 128128, and K=10K=10 negative samples. Training used 800 epochs for MIT-States, 500 for UT-Zappos, and 1000 for C-GQA, taking approximately 3, 1, and 4 hours, respectively, on an NVIDIA GTX 1080Ti GPU. Hyperparameter analysis fixed α=0.1\alpha=0.1 and selected β=0.5\beta=0.5 for MIT-States and β=0.1\beta=0.1 for both UT-Zappos and C-GQA.

  9. Knowl 9 — State-of-the-art generalized CZSL performance

    data/table

    SCEN's generalized CZSL results, reported on page 7, are shown below. All entries are percentages. AUC is reported on validation and test splits; HM is the harmonic mean of seen and unseen composition accuracy; ss and oo are state and object prediction accuracies.

    Could not parse LaTeX table

    SCEN obtains the best test AUC among the compared methods on all three datasets: 5.3% on MIT-States, 32.0% on UT-Zappos, and 5.5% on C-GQA. Its HM improves over the strongest reported prior HM from 17.2% to 18.4% on MIT-States, from 43.1% to 47.8% on UT-Zappos, and from 17.2% to 17.5% on C-GQA. The results support the paper's claim that separately learned prototypes together with state-transition augmentation improve recognition of unseen compositions.

  10. Knowl 10 — Ablation evidence for contrastive learning and state transition

    empirical result

    The ablation study compares a base model with separate state and object classifiers against versions augmented with the Siamese contrastive loss Lcts\mathcal L_{\mathrm{cts}}, the State Transition Module loss Lstm\mathcal L_{\mathrm{stm}}, or both. The complete model gives the strongest combined performance on both datasets.

    Could not parse LaTeX table

    Adding either component improves over the base model, while combining both yields the best HM and AUC. This supports the distinct roles claimed for the components: contrastive learning improves discriminative state/object prototype extraction, and state transition augmentation improves robustness to the seen-to-unseen composition shift.

  11. Knowl 11 — Reported failure modes and limitations

    limitation

    The paper's qualitative analysis identifies two limitations of SCEN. First, the irrelevant database DirD_{\mathrm{ir}} is sampled with a finite number of negatives, so it may omit visually or semantically confusing pairs. The authors associate such omissions with errors such as classifying slippers as sandals or boots. Second, CZSL datasets can assign only one composition label to an image even when several state descriptions are visually valid; for example, an image may exhibit both texture and age while the annotation specifies only age. Consequently, a prediction that captures a valid but unannotated factor is counted as incorrect. These issues limit the interpretation of compositional accuracy and motivate richer negative sampling and multi-label annotations.

Coverage note — The paper's individual qualitative top-3 examples and training-time plots are not reproduced separately because they illustrate the benchmark and ablation findings without adding a distinct load-bearing method or result; the reported failure modes are retained.

References

  1. 1.Yuval Atzmon, Felix Kreuk, Uri Shalit, and Gal Chechik. A causal view of compositional zero-shot recognition. arXiv preprint arXiv:2006.14610, 2020. 2
  2. 2.Irving Biederman. Recognition-by-components: a theory of human image understanding. Psychological review, 94(2):115, 1987. 2
  3. 3.Wei-Lun Chao, Soravit Changpinyo, Boqing Gong, and Fei Sha. An empirical study and analysis of generalized zero-shot learning for object recognition in the wild. In European conference on computer vision, pages 52–68. Springer, 2016. 6
  4. 4.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020. 2
  5. 5.Wuyang Chen, Zhiding Yu, Shalini De Mello, Sifei Liu, Jose M Alvarez, Zhangyang Wang, and Anima Anandkumar. Contrastive syn-to-real generalization. arXiv preprint arXiv:2104.02290, 2021. 2
  6. 6.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in neural information processing systems, 27, 2014. 5
  7. 7.Yanan Gu, Cheng Deng, and Kun Wei. Class-incremental instance segmentation via multi-teacher networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1478–1486, 2021. 2
  8. 8.Michael Gutmann and Aapo Hyvarinen. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, pages 297–304. JMLR Workshop and Conference Proceedings, 2010. 2
  9. 9.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9729–9738, 2020. 2
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 6
  11. 11.Donald D Hoffman and Whitman A Richards. Parts of recognition. Cognition, 18(1-3):65–96, 1984. 2
  12. 12.Phillip Isola, Joseph J Lim, and Edward H Adelson. Discovering states and transformations in image collections. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1383–1391, 2015. 5, 7
  13. 13.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning, 2021. 2
  14. 14.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 6
  15. 15.Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling. Attribute-based classification for zero-shot visual object categorization. IEEE transactions on pattern analysis and machine intelligence, 36(3):453–465, 2013. 2
  16. 16.Xiangyu Li, Zhe Xu, Kun Wei, and Cheng Deng. Generalized zero-shot learning via disentangled representation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1966–1974, 2021. 2
  17. 17.Yong-Lu Li, Yue Xu, Xiaohan Mao, and Cewu Lu. Symmetry and group in attribute-object compositions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11316–11325, 2020. 1, 2, 7
  18. 18.Massimiliano Mancini, Muhammad Ferjad Naeem, Yongqin Xian, and Zeynep Akata. Open world compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5222–5230, 2021. 7
  19. 19.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems, pages 3111–3119, 2013. 2
  20. 20.Ishan Misra, Abhinav Gupta, and Martial Hebert. From red wine to red tomato: Composition with context. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1792–1801, 2017. 1, 2, 3, 7
  21. 21.Andriy Mnih and Koray Kavukcuoglu. Learning word embeddings efficiently with noise-contrastive estimation. Advances in neural information processing systems, 26:2265–2273, 2013. 2
  22. 22.Muhammad Ferjad Naeem, Yongqin Xian, Federico Tombari, and Zeynep Akata. Learning graph embeddings for compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 953–962, 2021. 2, 5, 6, 7
  23. 23.Tushar Nagarajan and Kristen Grauman. Attributes as operators: factorizing unseen attribute-object compositions. In Proceedings of the European Conference on Computer Vision (ECCV), pages 169–185, 2018. 1, 2, 7
  24. 24.Zhixiong Nan, Yang Liu, Nanning Zheng, and Song-Chun Zhu. Recognizing unseen attribute-object pair with generative model. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 8811–8818, 2019. 1
  25. 25.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In Advances in neural information processing systems, pages 8026–8037, 2019. 6
  26. 26.Senthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, and Marc’Aurelio Ranzato. Task-driven modular networks for zero-shot compositional learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3593–3602, 2019. 1, 2, 3, 6, 7
  27. 27.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015. 6
  28. 28.Xin Wang, Fisher Yu, Trevor Darrell, and Joseph E Gonzalez. Task-aware feature generation for zero-shot compositional learning. arXiv preprint arXiv:1906.04854, 2019. 2
  29. 29.Kun Wei, Cheng Deng, and Xu Yang. Lifelong zero-shot learning. In IJCAI, pages 551–557, 2020. 2
  30. 30.Kun Wei, Muli Yang, Hao Wang, Cheng Deng, and Xianglong Liu. Adversarial fine-grained composition learning for unseen attribute-object recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3741–3749, 2019. 2
  31. 31.Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata. Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly. IEEE transactions on pattern analysis and machine intelligence, 41(9):2251–2265, 2018. 3
  32. 32.Muli Yang, Cheng Deng, Junchi Yan, Xianglong Liu, and Dacheng Tao. Learning unseen concepts via hierarchical decomposition and composition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10248–10256, 2020. 2
  33. 33.Aron Yu and Kristen Grauman. Fine-grained visual comparisons with local learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 192–199, 2014. 5, 7
  34. 34.Han Zhao, Xu Yang, Zhenru Wang, Erkun Yang, and Cheng Deng. Graph debiased contrastive learning with joint representation clustering. In Proc. IJCAI, pages 3434–3440, 2021. 2

Citation

MLA
Li, X., et al. “Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning”. arXiv, 2022, http://arxiv.org/abs/2206.14475v1.
APA
Li, X., Yang, X., Wei, K., Deng, C., & Yang, M. (2022). Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning. arXiv. http://arxiv.org/abs/2206.14475v1
Chicago
Li, X., X. Yang, K. Wei, C. Deng, and M. Yang. 2022. “Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning”. arXiv. http://arxiv.org/abs/2206.14475v1.
Harvard
Li, X. et al. (2022) “Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2206.14475v1.
Vancouver
1. Li X, Yang X, Wei K, Deng C, Yang M (2022) Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning. arXiv

BibTeX

@article{li2022siamese,
  title = {Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning},
  author = {Li, Xiangyu and Yang, Xu and Wei, Kun and Deng, Cheng and Yang, Muli},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2206.14475v1},
  eprint = {2206.14475}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE