Primitive Generation and Semantic-Related Alignment for Universal Zero-Shot Segmentation

Shuting HeHenghui DingWei Jiang

article2023CVPR54 citations

Proposes PADing, a unified universal zero-shot segmentation framework that bridges the cross-modal domain gap by assembling learned fine-grained primitives to synthesize unseen visual features and aligning their semantic-related components with linguistic class relationships.

Listen

Standard deep learning models for image segmentation demand vast volumes of manually annotated training images, making the recognition of novel or previously unseen visual categories both cost-prohibitive and labor-intensive. Zero-shot learning offers a pathway to identify objects without explicit training samples by leveraging language-based semantic knowledge, but adapting this capability to fine-grained visual tasks such as panoptic, instance, and semantic segmentation remains difficult due to strong biases toward seen classes and feature granularity mismatches between vision and language.

The article introduces and evaluates a unified framework called PADing (Primitive generation with collaborative relationship Alignment and feature Disentanglement learning) designed to perform universal zero-shot image segmentation across semantic, instance, and panoptic tasks.

The evaluated approach operates at the object level by decoupling mask generation from classification. Credibility is established through extensive experiments on standard Microsoft COCO benchmarks using a ResNet-50 visual backbone paired with semantic text representations like CLIP and word2vec embeddings. The methodology employs a specialized Transformer-based generative model that combines fine-grained learned visual units (primitives) to synthesize training features for unseen classes, while simultaneously separating visual representations into language-relevant and language-unrelated components to enforce category relationship alignments without distorting visual details.

The experimental findings demonstrate significant performance gains across all zero-shot benchmarks. First, the primitive-based generator combined with alignment and feature disentanglement achieved a harmonic mean panoptic quality of 22.3% on zero-shot panoptic segmentation, outperforming standard generative baselines (8.7%) and projection models (0.0%). Second, testing showed that increasing the number of learned primitives up to an optimal count of 400 yielded a 4.2% absolute gain in panoptic quality before leveling off. Third, on zero-shot semantic segmentation using the COCO-Stuff benchmark, the framework achieved a harmonic mean intersection-over-union of 30.7%, surpassing previous state-of-the-art methods such as ZegFormer (27.2%) even while utilizing a smaller backbone architecture. Fourth, on generalized zero-shot instance segmentation benchmarks, the system outperformed prior leading models by 7.20% in harmonic mean average precision on the standard 48-seen/17-unseen class split.

These results confirm that addressing feature granularity differences and isolating language-unrelated visual noise are critical for transferring linguistic knowledge into visual models. For organizations deploying computer vision, this approach lowers the risk, development time, and financial cost associated with continuous data re-annotation when expanding systems to recognize new visual categories.

Practitioners seeking to adopt or build upon this system should implement object-level feature synthesis rather than pixel-level generation and structure their generative components around fine-grained attribute primitives. Organizations should also consider incorporating feature disentanglement before enforcing cross-modal alignment to preserve essential visual distinctions.

The evaluations were conducted under controlled generalized zero-shot conditions on standard dataset splits, meaning confidence is high for comparable visual segmentation tasks. However, practitioners should exercise caution when deploying the framework in unconstrained, open-domain environments where class overlaps and vocabulary definitions may diverge from the benchmark settings.

arXiv: 2306.11087
Cover for Primitive Generation and Semantic-Related Alignment for Universal Zero-Shot Segmentation

Abstract

We study universal zero-shot segmentation in this work to achieve panoptic, instance, and semantic segmentation for novel categories without any training samples. Such zero-shot segmentation ability relies on inter-class relationships in semantic space to transfer the visual knowledge learned from seen categories to unseen ones. Thus, it is desired to well bridge semantic and visual spaces and apply the semantic relationships to visual feature learning. We introduce a generative model to synthesize features for unseen categories, which links semantic and visual spaces as well as addresses the issue of lack of unseen training data. Furthermore, to mitigate the domain gap between semantic and visual spaces, firstly, we enhance the vanilla generator with learned primitives, each of which contains fine-grained attributes related to categories, and synthesize unseen features by selectively assembling these primitives. Secondly, we propose to disentangle the visual feature into the semantic-related part and the semantic-unrelated part that contains useful visual classification clues but is less relevant to semantic representation. The inter-class relationships of semantic-related visual features are then required to be aligned with those in semantic space, thereby transferring semantic knowledge to visual feature learning. The proposed approach achieves impressively state-of-the-art performance on zero-shot panoptic segmentation, instance segmentation, and semantic segmentation.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Task Formulation
  • 3.2. Primitive Cross-Modal Generation
  • 3.3. Semantic-Visual Relationship Alignment
  • 3.4. Training Objective
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Zero-Shot Panoptic Segmentation Task
  • 4.3. Ablation Study
  • 4.4. Comparison with State-of-the-art ZSS Methods
  • 4.5. Comparison with State-of-the-art ZSI Methods
  • 4.6. Qualitative Results
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Universal Zero-Shot Segmentation Formulation

    definition

    Universal zero-shot segmentation unifies zero-shot panoptic segmentation (ZSP), zero-shot instance segmentation (ZSI), and zero-shot semantic segmentation (ZSS) under a generalized zero-shot learning (GZSL) paradigm.

    Let X={Xs,Xu}\mathcal{X} = \{\mathcal{X}^s, \mathcal{X}^u\} denote the visual feature space and A={As,Au}\mathcal{A} = \{\mathcal{A}^s, \mathcal{A}^u\} denote the category semantic embedding space, where superscripts ss and uu designate NsN^s seen categories and NuN^u unseen categories, respectively. The label sets are disjoint, Ys∩Yu=∅\mathcal{Y}^s \cap \mathcal{Y}^u = \emptyset. The visual model is trained exclusively on images containing seen classes Ys\mathcal{Y}^s, with all images containing any unseen class pixels removed to prevent information leakage. At test time, the model must segment and classify mask proposals into the combined set of seen and unseen classes Y=Ys∪Yu\mathcal{Y} = \mathcal{Y}^s \cup \mathcal{Y}^u.

    Performance is evaluated using the harmonic mean (HM) of seen and unseen performance metrics: HM=2×Pseen×PunseenPseen+Punseen\mathrm{HM} = \frac{2 \times P_{\text{seen}} \times P_{\text{unseen}}}{P_{\text{seen}} + P_{\text{unseen}}} where PseenP_{\text{seen}} and PunseenP_{\text{unseen}} represent task-specific metrics: Panoptic Quality (PQ) for panoptic segmentation, mean Average Precision (mAP) at IoU=0.5\text{IoU}=0.5 for instance segmentation, or mean Intersection-over-Union (mIoU) for semantic segmentation.

  2. Knowl 2 — Primitive Cross-Modal Generator

    model/method

    The Primitive Cross-Modal Generator synthesizes visual class embeddings X′={Xs′,Xu′}\mathcal{X}' = \{\mathcal{X}^{s\prime}, \mathcal{X}^{u\prime}\} for categories using their semantic text embeddings A={As,Au}\mathcal{A} = \{\mathcal{A}^s, \mathcal{A}^u\}. Instead of directly mapping low-granularity semantic embeddings to high-granularity visual features, it composes visual representations from a learnable bank of fine-grained attribute primitives P={pi}i=1N\mathcal{P} = \{p_i\}_{i=1}^N, where pi∈Rdkp_i \in \mathbb{R}^{d_k} and NN is the number of primitives (default N=400N=400).

    A self-attention operation is first applied over P\mathcal{P} to model primitive interdependencies. Linear projections ωK\omega_K and ωV\omega_V map P\mathcal{P} into keys KK and values VV. Semantic embeddings A\mathcal{A} are used as queries QQ to dynamically aggregate primitives via cross-attention: X′=ω1(softmax(QKTdk)V+A+Z)\mathcal{X}' = \omega_1\left(\mathrm{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V + \mathcal{A} + \mathcal{Z}\right) where Z∼N(0,I)\mathcal{Z} \sim \mathcal{N}(0, I) provides stochastic noise, and ω1\omega_1 is a linear projection layer.

    The generator is trained on seen categories using maximum mean discrepancy (MMD) loss LG\mathcal{L}_\mathcal{G} between real seen visual features Xs\mathcal{X}^s and synthesized seen features Xs′\mathcal{X}^{s\prime}: LG=∑f,f˙∈Xsk(f,f˙)+∑f′,f˙′∈Xs′k(f′,f˙′)−2∑f∈Xs∑f′∈Xs′k(f,f′)\mathcal{L}_\mathcal{G} = \sum_{f, \dot{f} \in \mathcal{X}^s} k(f, \dot{f}) + \sum_{f', \dot{f}' \in \mathcal{X}^{s\prime}} k(f', \dot{f}') - 2 \sum_{f \in \mathcal{X}^s} \sum_{f' \in \mathcal{X}^{s\prime}} k(f, f') where k(f,f′)=exp⁡(−12σ2∥f−f′∥2)k(f, f') = \exp\left(-\frac{1}{2\sigma^2} \|f - f'\|^2\right) is a Gaussian kernel evaluated across bandwidths σ∈{2,5,10,20,40,60}\sigma \in \{2, 5, 10, 20, 40, 60\}.

  3. Knowl 3 — Visual Feature Disentanglement Module

    model/method

    To prevent visual features that lack textual semantic relevance from distorting semantic alignment, class embeddings xi∈Xx_i \in \mathcal{X} (from the visual backbone or primitive generator) are decomposed into a semantic-related feature x^i=ER(xi)\hat{x}_i = E_R(x_i) and a semantic-unrelated feature x¨i=EU(xi)\ddot{x}_i = E_U(x_i) via single-hidden-layer MLP encoders ERE_R and EUE_U.

    The semantic-related encoder ERE_R is trained using a cross-entropy loss LR\mathcal{L}_\mathbb{R} against class semantic embeddings A={a1,…,aNs+Nu}\mathcal{A} = \{a_1, \dots, a_{N^s + N^u}\}: LR=−∑i∑kI([xi]=k)log⁡exp⁡(x^iak/τ)∑jexp⁡(x^iaj/τ)\mathcal{L}_\mathbb{R} = -\sum_{i}\sum_{k} \mathbb{I}([x_i] = k) \log \frac{\exp(\hat{x}_i a_k / \tau)}{\sum_{j} \exp(\hat{x}_i a_j / \tau)} where [xi][x_i] is the ground-truth class index, I(⋅)\mathbb{I}(\cdot) is the indicator function, and τ=0.1\tau = 0.1 is temperature.

    The semantic-unrelated encoder EUE_U is constrained to match a standard normal prior via KL divergence loss LU\mathcal{L}_\mathbb{U}: LU=∑iDKL[x¨i∥N(0,I)]\mathcal{L}_\mathbb{U} = \sum_{i} D_{KL}[\ddot{x}_i \parallel \mathcal{N}(0, I)]

    A decoder DD (two stacked MLP layers) reconstructs the original visual feature xix_i under ℓ1\ell_1 loss: Lrecon=∑i∥xi−D(x^i,x¨i)∥1\mathcal{L}_{recon} = \sum_{i} \|x_i - D(\hat{x}_i, \ddot{x}_i)\|_1

    The composite feature disentanglement loss is: LD=LR+LU+Lrecon\mathcal{L}_\mathbb{D} = \mathcal{L}_\mathbb{R} + \mathcal{L}_\mathbb{U} + \mathcal{L}_{recon}

  4. Knowl 4 — Semantic-Visual Relationship Alignment

    model/method

    Semantic-Visual Relationship Alignment transfers relational category knowledge from semantic space A\mathcal{A} to visual feature space by constraining the pairwise cosine similarities of disentangled semantic-related visual features x^i,x^j\hat{x}_i, \hat{x}_j to match the pairwise similarities of their ground-truth class semantic embeddings a[x^i],a[x^j]a_{[\hat{x}_i]}, a_{[\hat{x}_j]} via KL divergence: LA=DKL[x^iTx^j∥x^i∥∥x^j∥/τ  ∥  a[x^i]Ta[x^j]∥a[x^i]∥∥a[x^j]∥/τ]\mathcal{L}_\mathbb{A} = D_{KL}\left[ \frac{\hat{x}_i^T \hat{x}_j}{\|\hat{x}_i\| \|\hat{x}_j\|} / \tau \;\Bigg\|\; \frac{a_{[\hat{x}_i]}^T a_{[\hat{x}_j]}}{\|a_{[\hat{x}_i]}\| \|a_{[\hat{x}_j]}\|} / \tau \right] where [x^i][\hat{x}_i] is the class index, and τ=0.1\tau = 0.1 is the distribution temperature parameter.

    This alignment operates across two modes:

    1. Intra-group alignment: Both x^i\hat{x}_i and x^j\hat{x}_j are from the same group (e.g., both seen), refining representation clustering according to semantic distances.
    2. Inter-group alignment: One feature is from seen classes (real or synthetic) and the other is from unseen classes (synthetic only), transferring relational structure to unseen categories and regularizing generator outputs.
  5. Knowl 5 — Universal Zero-Shot Segmentation Training Algorithm

    algorithm

    The PADing training procedure trains a universal segmentation backbone on seen classes, optimizes the primitive generator with collaborative disentanglement and relationship alignment, and fine-tunes the mask classifier on joint real seen and synthetic unseen visual features.

    Input: Training images I_s containing only seen categories, semantic text embeddings A = {A^s, A^u}.
    Output: Trained universal segmentation model capable of seen and unseen classification.
    1. Pre-train visual backbone (Mask2Former with ResNet-50) using fully supervised annotations of seen classes.
    2. Forward seen images I_s through backbone to obtain class-agnostic masks M^s and visual class embeddings X^s.
    3. Train Primitive Generator, Encoders E_R, E_U, and Decoder D using loss L_total = L_G + lambda * (L_D + L_A):
        a. Sample noise Z ~ N(0, I) and pass A through Primitive Generator to obtain synthetic visual embeddings X' = {X^{s\prime}, X^{u\prime}}.
        b. Supervise X^{s\prime} with real visual features X^s via MMD loss L_G.
        c. Disentangle features x_i in {X^s, X'} into semantic-related x_hat_i = E_R(x_i) and semantic-unrelated x_ddot_i = E_U(x_i).
        d. Compute disentanglement loss L_D = L_R + L_U + L_recon.
        e. Align similarity distributions of x_hat_i with A via relationship alignment loss L_A.
        f. Update generator and disentanglement parameters using Adam (lr = 2e-4, loss weight lambda = 0.002).
    4. Synthesize unseen class embeddings X^{u\prime} using trained generator and unseen text embeddings A^u.
    5. Fine-tune the classification layer of the visual backbone on X^s union X^{u\prime} using SGD (lr = 1e-3, momentum = 0.9, weight decay = 5e-4).
  6. Knowl 6 — MS-COCO Zero-Shot Panoptic Segmentation Benchmark Setup

    experimental setup

    The Zero-Shot Panoptic Segmentation (ZSP) benchmark is derived from MSCOCO 2017 (133 total classes: 80 thing and 53 stuff). Unseen categories are established by matching classes with SPNet's 15 unseen classes in COCO-Stuff that do not overlap with ImageNet classification classes, yielding 14 unseen categories in COCO Panoptic: {cow, giraffe, suitcase, frisbee, skateboard, carrot, scissors, cardboard, sky-other-merged, grass-merged, playingfield, river, road, tree-merged}. The remaining 119 categories are seen.

    To prevent information leakage, all training images containing any pixels of the 14 unseen classes are removed from the training split, resulting in a strictly clean training set of 45,617 seen images. The full MSCOCO 2017 validation set (5,000 images) is used for evaluation under the Generalized Zero-Shot Learning (GZSL) setting across thing and stuff classes.

  7. Knowl 7 — Ablation of PADing Components across Zero-Shot Segmentation Tasks

    data/table

    Ablation experiments evaluate the contribution of the Primitive Generator (P), Feature Disentanglement (D), and Relationship Alignment (A) against a Projection baseline and a standard Generative Moment Matching Network (GMMN, G). All models use Mask2Former with a ResNet-50 backbone trained on the MS-COCO zero-shot panoptic split.

    Seen Unseen HM
    Method G/P A D PQ SQ RQ PQ SQ RQ PQ SQ RQ
    Supervised ✗ ✗ ✗ 43.2 80.8 51.6 0.0 0.0 0.0 - - -
    Projection ✗ ✗ ✗ 43.3 80.9 51.7 0.0 0.0 0.0 0.0 0.0 0.0
    Baseline (GMMN) G ✗ ✗ 40.1 80.1 48.2 4.9 48.5 5.8 8.7 60.4 10.3
    P-only P ✗ ✗ 38.9 79.9 46.2 11.5 56.5 13.8 17.7 66.1 21.2
    PA P ✓ ✗ 38.4 79.2 45.8 13.8 52.7 16.4 20.3 63.2 24.1
    PADing P ✓ ✓ 41.5 80.6 49.7 15.3 72.8 18.4 22.3 76.5 26.8
    Object Detection (ZSD) Instance Seg. (ZSI) Semantic Seg. (ZSS)
    Method G/P A D Seen Unseen HM Seen Unseen HM Seen Unseen HM
    Supervised ✗ ✗ ✗ 53.2 0.0 - 53.9 0.0 - 51.2 0.0 -
    Projection ✗ ✗ ✗ 53.2 12.6 20.4 54.0 12.5 20.3 50.8 1.2 2.3
    Baseline G ✗ ✗ 52.2 16.5 25.0 52.8 16.2 24.7 50.8 11.6 18.8
    P-only P ✗ ✗ 52.0 18.5 27.2 52.5 18.4 27.2 50.4 16.7 25.0
    PA P ✓ ✗ 51.9 18.8 27.6 52.3 18.6 27.4 50.2 16.0 24.2
    PADing P ✓ ✓ 52.1 19.6 28.4 52.6 19.2 28.1 50.5 18.5 27.0

    The primitive generator alone improves HM-PQ from 8.7% to 17.7% over GMMN. Adding relationship alignment increases HM-PQ to 20.3%, and combining it with feature disentanglement brings HM-PQ to 22.3% and unseen PQ to 15.3%.

  8. Knowl 8 — Effect of Primitive Bank Size on Zero-Shot Segmentation Quality

    data/table

    The number of learned attribute primitives in the Primitive Generator controls visual representation granularity. The table details unseen Panoptic Quality (PQ) on MSCOCO as the primitive count varies from 100 to 700:

    #Primitives 100 200 300 400 500 600 700
    PQ (%) 7.3 9.1 10.2 11.5 11.5 11.3 11.2

    Performance increases by 4.2% PQ when scaling primitives from 100 to 400. Above 500 primitives, performance slightly declines; 400 is chosen as the default primitive bank size.

  9. Knowl 9 — Zero-Shot Semantic Segmentation Comparison on COCO-Stuff

    data/table

    Comparison of zero-shot semantic segmentation methods on the COCO-Stuff dataset (171 classes under GZSL without self-training or crop-mask preprocessing). PADing uses a ResNet-50 backbone while competing methods employ ResNet-101.

    Method Embedding Seen IoU (%) Unseen IoU (%) HM IoU (%)
    SPNet Word2vec 35.2 8.7 14.0
    ZS3 Word2vec 34.7 9.5 15.0
    CaGNet Word2vec 33.5 12.2 18.2
    SIGN Word2vec 32.3 15.5 20.9
    Zsseg-seg CLIP 38.7 4.9 8.7
    ZegFormer-seg CLIP 37.4 21.4 27.2
    PADing (ours) CLIP 40.4 24.8 30.7

    PADing outperforms the previous best method, ZegFormer-seg, by 3.5% HM-IoU and 3.4% unseen-IoU despite using a lighter backbone (ResNet-50 vs. ResNet-101).

  10. Knowl 10 — Generalized Zero-Shot Instance Segmentation on MSCOCO

    data/table

    Performance comparison on MSCOCO Generalized Zero-Shot Instance Segmentation (GZSI) across two standard seen/unseen splits (48/17 and 65/15) using word2vec semantic embeddings.

    Seen Unseen HM
    Split Method mAP Recall mAP Recall mAP Recall
    48/17 ZSI 43.0 64.4 3.6 44.9 6.7 52.9
    PADing (ours) 53.0 75.1 8.0 47.5 13.9 58.2
    65/15 ZSI 35.7 62.5 10.4 49.9 16.2 55.5
    PADing (ours) 41.8 73.2 13.9 51.3 20.9 60.3

    Metrics are reported at IoU=0.5\text{IoU} = 0.5. PADing with ResNet-50 surpasses ZSI (which uses ResNet-101) by +7.20% HM-mAP and +5.27% HM-Recall on the 48/17 split, and by +4.70% HM-mAP and +4.80% HM-Recall on the 65/15 split.

Coverage note — Qualitative t-SNE distribution plots and qualitative mask prediction figures were omitted in favor of quantitative experimental tables and formal algorithmic and mathematical definitions that comprehensively describe the core contribution.

References

  1. 1.Zeynep Akata, Scott Reed, Daniel Walter, Honglak Lee, and Bernt Schiele. Evaluation of output embeddings for fine-grained image classification. In IEEE Conf. Comput. Vis. Pattern Recog., 2015. 2
  2. 2.Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran. Zero-shot object detection. In Eur. Conf. Comput. Vis., 2018. 6
  3. 3.Supritam Bhattacharjee, Devraj Mandal, and Soma Biswas. Autoencoder based novelty detection for generalized zero shot learning. In ICIP. IEEE, 2019. 2
  4. 4.Maxime Bucher, Tuan-Hung Vu, Matthieu Cord, and Patrick Pérez. Zero-shot semantic segmentation. Adv. Neural Inform. Process. Syst., 32, 2019. 1, 2, 3, 4, 6, 8
  5. 5.Soravit Changpinyo, Wei-Lun Chao, Boqing Gong, and Fei Sha. Classifier and exemplar synthesis for zero-shot learning. Int. J. Comput. Vis., 128(1), 2020. 2
  6. 6.Wei-Lun Chao, Soravit Changpinyo, Boqing Gong, and Fei Sha. An empirical study and analysis of generalized zero-shot learning for object recognition in the wild. In Eur. Conf. Comput. Vis. Springer, 2016. 2, 3
  7. 7.Junjie Chen, Li Niu, Liu Liu, and Liqing Zhang. Weak-shot fine-grained classification via similarity transfer. NeurIPS, 2021. 4
  8. 8.Junjie Chen, Li Niu, Siyuan Zhou, Jianlou Si, Chen Qian, and Liqing Zhang. Weak-shot semantic segmentation via dual similarity transfer. NeurIPS, 2022. 2
  9. 9.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell., 40(4), 2017. 1, 2
  10. 10.Zhi Chen, Yadan Luo, Ruihong Qiu, Sen Wang, Zi Huang, Jingjing Li, and Zheng Zhang. Semantics disentangling for generalized zero-shot learning. In Int. Conf. Comput. Vis., 2021. 2
  11. 11.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 1, 2, 6, 8
  12. 12.Jiaxin Cheng, Soumyaroop Nandi, Prem Natarajan, and Wael Abd-Almageed. Sign: Spatial-information incorporated generative network for generalized zero-shot semantic segmentation. In Int. Conf. Comput. Vis., pages 9556–9566, October 2021. 2, 8
  13. 13.Debasmit Das and CS George Lee. Zero-shot image recognition using relational matching, adaptation and calibration. In IJCNN. IEEE, 2019. 2
  14. 14.Henghui Ding, Xudong Jiang, Ai Qun Liu, Nadia Magnenat Thalmann, and Gang Wang. Boundary-aware feature propagation for scene segmentation. In Int. Conf. Comput. Vis., 2019. 2
  15. 15.Henghui Ding, Xudong Jiang, Bing Shuai, Ai Qun Liu, and Gang Wang. Context contrasted feature and gated multi-scale aggregation for scene segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2018. 2
  16. 16.Henghui Ding, Xudong Jiang, Bing Shuai, Ai Qun Liu, and Gang Wang. Semantic correlation promoted shape-variant context for segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2019. 2
  17. 17.Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang. Vision-language transformer and query generation for referring segmentation. In Int. Conf. Comput. Vis., 2021. 2
  18. 18.Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang. Vlt: Vision-language transformer and query generation for referring segmentation. IEEE Trans. Pattern Anal. Mach. Intell., 2022. 2
  19. 19.Jian Ding, Nan Xue, Gui-Song Xia, and Dengxin Dai. Decoupling zero-shot semantic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 1, 2, 6, 8
  20. 20.Georgiana Dinu, Angeliki Lazaridou, and Marco Baroni. Improving zero-shot learning by mitigating the hubness problem. arXiv preprint arXiv:1412.6568, 2014. 2
  21. 21.Mark Everingham, SM Ali Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge: A retrospective. International journal of computer vision, 111(1), 2015. 6
  22. 22.Rafael Felix, Ben Harwood, Michele Sasdelli, and Gustavo Carneiro. Generalised zero-shot learning with domain classification in a joint semantic and visual space. In DICTA. IEEE, 2019. 2
  23. 23.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Communications of the ACM, 63(11), 2020. 3
  24. 24.Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary object detection via vision and language knowledge distillation. arXiv preprint arXiv:2104.13921, 2021. 6
  25. 25.Zhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao, and Liqing Zhang. Context-aware feature generation for zero-shot semantic segmentation. In ACM Int. Conf. Multimedia, 2020. 1, 2, 3, 4, 8
  26. 26.Yuchen Guo, Guiguang Ding, Jungong Han, Xiaohan Ding, Sicheng Zhao, Zheng Wang, Chenggang Yan, and Qionghai Dai. Dual-view ranking with hardness assessment for zero-shot learning. In AAAI, 2019. 2
  27. 27.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Int. Conf. Comput. Vis., 2017. 1, 2
  28. 28.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 1, 6
  29. 29.Shuting He, Henghui Ding, and Wei Jiang. Semantic-promoted debiasing and background disambiguation for zero-shot instance segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 2
  30. 30.Ping Hu, Stan Sclaroff, and Kate Saenko. Uncertainty-aware learning for zero-shot semantic segmentation. Adv. Neural Inform. Process. Syst., 33, 2020. 2
  31. 31.Dat Huynh and Ehsan Elhamifar. Fine-grained generalized zero-shot learning via dense attribute-based attention. In IEEE Conf. Comput. Vis. Pattern Recog., 2020. 2
  32. 32.Dat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu, and Ehsan Elhamifar. Open-vocabulary instance segmentation via robust cross-modal pseudo-labeling. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 3
  33. 33.Pichai Kankuekul, Aram Kawewong, Sirinart Tangruamsub, and Osamu Hasegawa. Online incremental attribute-based zero-shot learning. In IEEE Conf. Comput. Vis. Pattern Recog. IEEE, 2012. 2
  34. 34.Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollár. Panoptic feature pyramid networks. In IEEE Conf. Comput. Vis. Pattern Recog., 2019. 1, 2
  35. 35.Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollár. Panoptic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2019. 1, 2, 6
  36. 36.Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling. Learning to detect unseen object classes by between-class attribute transfer. In IEEE Conf. Comput. Vis. Pattern Recog. IEEE, 2009. 1, 2
  37. 37.Hsin-Ying Lee, Hung-Yu Tseng, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang. Diverse image-to-image translation via disentangled representations. In Eur. Conf. Comput. Vis., 2018. 5
  38. 38.Peike Li, Yunchao Wei, and Yi Yang. Consistent structural relation learning for zero-shot segmentation. Adv. Neural Inform. Process. Syst., 33, 2020. 1, 2, 3, 4
  39. 39.Yan Li, Zhen Jia, Junge Zhang, Kaiqi Huang, and Tieniu Tan. Deep semantic structural constraints for zero-shot learning. In AAAI, 2018. 2
  40. 40.Yujia Li, Kevin Swersky, and Rich Zemel. Generative moment matching networks. In ICML, 2015. 3, 4
  41. 41.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV. Springer, 2014. 6
  42. 42.Chang Liu, Henghui Ding, and Xudong Jiang. Gres: Generalized referring expression segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 2
  43. 43.Chang Liu, Xudong Jiang, and Henghui Ding. Instance-specific feature propagation for referring segmentation. IEEE Trans. Multimedia, 2022. 2
  44. 44.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2015. 1, 2
  45. 45.Fengmao Lv, Haiyang Liu, Yichen Wang, Jiayi Zhao, and Guowu Yang. Learning unbiased zero-shot semantic segmentation networks via transductive transfer. IEEE Signal Processing Letters, 27, 2020. 2
  46. 46.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. Distributed representations of words and phrases and their compositionality. In Adv. Neural Inform. Process. Syst., 2013. 6
  47. 47.Mark Palatucci, Dean Pomerleau, Geoffrey E Hinton, and Tom M Mitchell. Zero-shot learning with semantic output codes. Adv. Neural Inform. Process. Syst., 22, 2009. 1, 2
  48. 48.Giuseppe Pastore, Fabio Cermelli, Yongqin Xian, Massimiliano Mancini, Zeynep Akata, and Barbara Caputo. A closer look at self-training for zero-label semantic segmentation, 2021. 2
  49. 49.Farhad Pourpanah, Moloud Abdar, Yuxuan Luo, Xinlei Zhou, Ran Wang, Chee Peng Lim, and Xi-Zhao Wang. A review of generalized zero-shot learning methods. arXiv preprint arXiv:2011.08641, 2020. 1
  50. 50.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In Proc. Int. Conf. Mach. Learn. PMLR, 2021. 6
  51. 51.Walter J Scheirer, Anderson de Rezende Rocha, Archana Sapkota, and Terrance E Boult. Toward open set recognition. IEEE Trans. Pattern Anal. Mach. Intell., 35(7), 2012. 2
  52. 52.Bin Tong, Chao Wang, Martin Klinkigt, Yoshiyuki Kobayashi, and Yuuichi Nonaka. Hierarchical disentanglement of discriminative latent features for zero-shot learning. In IEEE Conf. Comput. Vis. Pattern Recog., 2019. 2
  53. 53.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008. 7
  54. 54.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Adv. Neural Inform. Process. Syst., 2017. 1
  55. 55.Maunil R Vyas, Hemanth Venkateswara, and Sethuraman Panchanathan. Leveraging seen and unseen semantic relationships for generative zero-shot learning. In Eur. Conf. Comput. Vis. Springer, 2020. 4
  56. 56.Xinlong Wang, Tao Kong, Chunhua Shen, Yuning Jiang, and Lei Li. Solo: Segmenting objects by locations. In Eur. Conf. Comput. Vis. Springer, 2020. 2
  57. 57.Yongqin Xian, Subhabrata Choudhury, Yang He, Bernt Schiele, and Zeynep Akata. Semantic projection network for zero-and few-label semantic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2019. 1, 2, 6, 8
  58. 58.Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, and Xiang Bai. A simple baseline for open-vocabulary semantic segmentation with pre-trained vision-language model. In Eur. Conf. Comput. Vis. Springer, 2022. 2, 8
  59. 59.Felix X Yu, Liangliang Cao, Rogerio S Feris, John R Smith, and Shih-Fu Chang. Designing category-level attributes for discriminative visual recognition. In IEEE Conf. Comput. Vis. Pattern Recog., 2013. 2
  60. 60.Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang. Open-vocabulary object detection using captions. In IEEE Conf. Comput. Vis. Pattern Recog., 2021. 3
  61. 61.Hui Zhang and Henghui Ding. Prototypical matching and open set rejection for zero-shot semantic segmentation. In Int. Conf. Comput. Vis., 2021. 1, 2, 4
  62. 62.Ziming Zhang and Venkatesh Saligrama. Zero-shot learning via joint latent similarity embedding. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 2
  63. 63.Ye Zheng, Jiahong Wu, Yongqiang Qin, Faen Zhang, and Li Cui. Zero-shot instance segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2021. 1, 2, 4, 6, 8

Citation

MLA
He, S., et al. “Primitive Generation and Semantic-related Alignment for Universal Zero-Shot Segmentation”. arXiv, 2023, http://arxiv.org/abs/2306.11087v1.
APA
He, S., Ding, H., & Jiang, W. (2023). Primitive Generation and Semantic-related Alignment for Universal Zero-Shot Segmentation. arXiv. http://arxiv.org/abs/2306.11087v1
Chicago
He, S., H. Ding, and W. Jiang. 2023. “Primitive Generation and Semantic-related Alignment for Universal Zero-Shot Segmentation”. arXiv. http://arxiv.org/abs/2306.11087v1.
Harvard
He, S., Ding, H. and Jiang, W. (2023) “Primitive Generation and Semantic-related Alignment for Universal Zero-Shot Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.11087v1.
Vancouver
1. He S, Ding H, Jiang W (2023) Primitive Generation and Semantic-related Alignment for Universal Zero-Shot Segmentation. arXiv

BibTeX

@article{he2023primitive,
  title = {Primitive Generation and Semantic-related Alignment for Universal Zero-Shot Segmentation},
  author = {He, Shuting and Ding, Henghui and Jiang, Wei},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.11087v1},
  eprint = {2306.11087}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE