Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Sunan HeTaian GuoTao DaiRuizhi QiaoXiujun ShuBo RenShu-Tao Xia

article2023AAAI81 citations

Proposes a multi-modal knowledge transfer framework that adapts vision-language pre-trained models via knowledge distillation, prompt tuning, and a two-stream feature extractor to recognize unseen object labels in multi-label classification.

Listen

Real-world computer vision systems for applications such as autonomous driving, surveillance, and automated scene understanding must recognize thousands of diverse concepts, including categories never encountered during training. Traditional multi-label zero-shot recognition approaches attempt to generalize to unseen categories by relying on text-only language models. However, because these systems lack visual context, they struggle to capture visual consistency across concepts and perform poorly when evaluating complex or descriptive multi-word phrases.

The article evaluates an open-vocabulary framework designed to recognize both seen and unseen multi-word visual labels by transferring rich multi-modal knowledge from pre-trained vision-and-language models. The primary objective is to demonstrate that aligning visual image representations with multi-modal pre-trained text embeddings significantly improves multi-label classification accuracy across open vocabularies.

The authors develop the Multi-Modal Knowledge Transfer framework, which integrates a Vision Transformer backbone with the pre-trained CLIP vision-and-language model. The architecture employs a two-stream feature extraction module to process both broad global image context and granular local image patches. To ensure effective transfer, the system uses knowledge distillation to align global image representations with pre-trained visual embeddings, paired with continuous prompt tuning to optimize text label embeddings. The framework was evaluated on two benchmark datasets: NUS-WIDE, containing over 269,000 images and 1,006 label categories, and Open Images, spanning over 9 million training images and thousands of complex categories.

The experimental findings show significant performance improvements over prior state-of-the-art models. On the NUS-WIDE zero-shot benchmark, the proposed framework achieved a mean average precision of 37.6%, outperforming the previous leading method by an absolute margin of 11.7% and a direct fine-tuned baseline by 7.1%. In generalized zero-shot testing, which requires classifying seen and unseen labels simultaneously, the method improved precision from 12.1% to 18.3%. On the large-scale Open Images dataset, the framework achieved a zero-shot mean average precision of 68.1% and a weighted precision of 89.2%, outperforming existing zero-shot baselines by 2.5% and 16.3%, respectively. Ablation analyses confirmed that combining global distillation with local patch analysis and prompt tuning produced superior noise resistance and higher predictive accuracy than any component used in isolation.

These results indicate that leveraging joint vision-and-language pre-training substantially reduces the performance penalty typically associated with unseen categories. By effectively handling arbitrary, multi-word descriptive queries without requiring task-specific manual retraining, this approach offers an efficient path to deploying adaptable vision systems. This flexibility lowers ongoing data annotation costs, speeds up deployment timelines for new visual categories, and reduces the risk of classification failures in dynamic real-world environments.

Organizations developing large-scale image tagging, content moderation, or visual surveillance pipelines should consider adopting vision-language distillation frameworks rather than purely text-based zero-shot architectures. Implementation teams should balance local and global feature extraction parameters, as localized analysis improves the detection of small objects but requires calibration to avoid noise sensitivity. Further research and validation should focus on extending this multi-modal transfer methodology to complex video streams and dense real-time robotic perception tasks.

Confidence in these findings is supported by consistent improvements across multiple large-scale public benchmarks. However, stakeholders should note that performance remains sensitive to hyperparameter tuning, specifically the balance between knowledge distillation weights and local pooling configurations. Additionally, overall accuracy remains dependent on the underlying semantic coverage of the pre-trained vision-language foundation model.

Cover for Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Abstract

Real-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowledge by a pre-trained textual label embedding (e.g., GloVe). However, such methods only exploit single-modal knowledge from a language model, while ignoring the rich semantic information inherent in image-text pairs. Instead, recently developed open-vocabulary (OV) based methods succeed in exploiting such information of image-text pairs in object detection, and achieve impressive performance. Inspired by the success of OV-based methods, we propose a novel open-vocabulary framework, named multi-modal knowledge transfer (MKT), for multi-label classification. Specifically, our method exploits multi-modal knowledge of image-text pairs based on a vision and language pre-training (VLP) model. To facilitate transferring the image-text matching ability of VLP model, knowledge distillation is employed to guarantee the consistency of image and label embeddings, along with prompt tuning to further update the label embeddings. To further enable the recognition of multiple objects, a simple but effective two-stream module is developed to capture both local and global features. Extensive experimental results show that our method significantly outperforms state-of-the-art methods on public benchmark datasets.

Table of Contents

  • Introduction
  • Related Works Multi-Label Zero-Shot Learning
  • Open-Vocabulary Classification
  • Multi-modal Knowledge Transfer
  • Preliminary
  • The Overall Framework
  • Vision Transformer with Two-Stream Module
  • Knowledge Distillation for Alignment
  • Prompt Tuning for Label Embedding
  • Loss Functions
  • Experiments
  • Experiments Setup
  • State-of-the-art Comparison
  • Ablation Studies
  • Qualitative Assessment
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Multi-modal knowledge transfer for open-vocabulary multi-label classification

    model/method

    The paper introduces multi-modal knowledge transfer (MKT), an open-vocabulary multi-label classifier that transfers image–text matching knowledge from a vision-and-language pre-training (VLP) model to multi-label recognition. MKT combines a trainable Vision Transformer image encoder, a frozen CLIP image encoder used as a teacher, and a CLIP text encoder that produces embeddings for arbitrary label names, including labels absent from supervised training images.

    MKT improves transfer in three ways: (1) knowledge distillation aligns the trainable image representation with the CLIP image representation; (2) prompt tuning adapts CLIP label embeddings to multi-label classification without fine-tuning the entire text encoder; and (3) a two-stream prediction module combines global image features with local patch features. The architecture diagram on page 3 depicts this process: image and patch features are mapped into the same embedding space as label texts, similarities are computed for each label, and global and local scores are combined to produce multi-label predictions.

  2. Knowl 2 — Global–local two-stream label scoring

    model/method

    MKT uses a Vision Transformer to produce one global class-token feature oclso_{\mathrm{cls}} and NN local patch features opatch=[o1,…,oN]o_{\mathrm{patch}}=[o_1,\ldots,o_N]. A trainable global head ΘG\Theta_G and local head ΘL\Theta_L map these features into the DeD_e-dimensional label-embedding space:

    ecls=ΘG(ocls),er=ΘL(or),r=1,…,N.e_{\mathrm{cls}}=\Theta_G(o_{\mathrm{cls}}),\qquad e_r=\Theta_L(o_r),\quad r=1,\ldots,N.

    For candidate label ii with embedding zi∈RDez_i\in\mathbb{R}^{D_e}, the prediction score for an image is

    si=⟨zi,ecls⟩+TopMean⁡k(⟨zi,e1⟩,…,⟨zi,eN⟩),s_i=\langle z_i,e_{\mathrm{cls}}\rangle+\operatorname{TopMean}_k\left(\langle z_i,e_1\rangle,\ldots,\langle z_i,e_N\rangle\right),

    where ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle is the inner product and TopMean⁡k\operatorname{TopMean}_k averages the kk largest local similarity scores. Thus, the global stream supplies a general image-level response, while the local stream allows a label associated with only a few image regions to obtain a high score. The two-stream module is a central MKT contribution; the implementation uses two linear layers for the local head and one linear projection for the global head.

  3. Knowl 3 — Knowledge distillation and prompt tuning

    model/method

    MKT distills the image representation of a frozen CLIP image encoder into the trainable Vision Transformer. Let xx be an input image, ΦICLIP(x)\Phi_I^{\mathrm{CLIP}}(x) be the CLIP image-encoder output, odist=ΦICLIP(x)o_{\mathrm{dist}}=\Phi_I^{\mathrm{CLIP}}(x), and oclso_{\mathrm{cls}} be the student Vision Transformer’s global feature. The distillation loss is the elementwise ℓ1\ell_1 distance

    Ldist=∥ΦICLIP(x)−ocls∥1=∥odist−ocls∥1.\mathcal{L}_{\mathrm{dist}}=\left\|\Phi_I^{\mathrm{CLIP}}(x)-o_{\mathrm{cls}}\right\|_1=\left\|o_{\mathrm{dist}}-o_{\mathrm{cls}}\right\|_1.

    Distillation is applied to the global feature because both representations correspond to image-level class-token features, whereas local patch features should remain different enough to represent multiple objects.

    For label embedding, MKT inserts each label into the prompt 'There is a {label} in the scene' and encodes the resulting sentence with the CLIP text encoder. During prompt tuning, the CLIP text encoder and all other model parameters remain fixed while only the continuous context embedding in the prompt is optimized. Training has two stages: stage 1 optimizes the ranking loss plus distillation, Lstage1=Lrank+λLdist\mathcal{L}_{\mathrm{stage1}}=\mathcal{L}_{\mathrm{rank}}+\lambda\mathcal{L}_{\mathrm{dist}}, where λ≥0\lambda\geq 0 is the distillation weight; stage 2 optimizes only Lrank\mathcal{L}_{\mathrm{rank}} while tuning the context embedding.

  4. Knowl 4 — Zero-shot and generalized zero-shot multi-label setting

    definition

    The task partitions the label vocabulary into disjoint seen and unseen sets, YS\mathcal{Y}^{S} and YU\mathcal{Y}^{U}. Training consists of image–label pairs (xj,yj)(x_j,y_j), where xjx_j is an image and yj⊆YSy_j\subseteq\mathcal{Y}^{S} is the set of seen labels present in that image; unseen labels have no training images.

    In the zero-shot learning (ZSL) setting, the classifier must identify relevant unseen labels using fZSL:X→YUf_{\mathrm{ZSL}}:\mathcal{X}\rightarrow\mathcal{Y}^{U}. In generalized zero-shot learning (GZSL), the classifier must identify both seen and unseen labels using fGZSL:X→YS∪YUf_{\mathrm{GZSL}}:\mathcal{X}\rightarrow\mathcal{Y}^{S}\cup\mathcal{Y}^{U}. MKT addresses both settings by generating candidate-label embeddings from CLIP text prompts rather than restricting the classifier to a fixed learned output layer.

  5. Knowl 5 — Ranking objective for multi-label prediction

    equation

    For training image xjx_j with ground-truth seen-label set yj⊆YSy_j\subseteq\mathcal{Y}^{S}, MKT trains its prediction scores with a unit-margin pairwise ranking loss:

    Lrank=∑j∑p∈yj∑n∈YS∖yjmax⁡(1+sj,n−sj,p,0),\mathcal{L}_{\mathrm{rank}}=\sum_j\sum_{p\in y_j}\sum_{n\in\mathcal{Y}^{S}\setminus y_j}\max\left(1+s_{j,n}-s_{j,p},0\right),

    where sj,ps_{j,p} and sj,ns_{j,n} are the two-stream scores for positive label pp and negative label nn on image xjx_j. The loss encourages every ground-truth label to score at least one unit higher than every non-ground-truth seen label. In stage 1, this ranking loss is combined with the CLIP-image distillation loss; in stage 2, it is the sole optimization objective for prompt-context adaptation.

  6. Knowl 6 — Benchmark protocol and implementation configuration

    experimental setup

    MKT is evaluated on NUS-WIDE and Open Images v4 under ZSL and GZSL protocols. NUS-WIDE contains 81 human-verified labels treated as unseen and 925 Flickr-user-tag labels treated as seen; the official split provides 161,789 training images and 107,859 test images. Open Images v4 contains approximately 9 million training images and 125,456 test images; 7,186 labels with more than 100 training images are treated as seen, and the 400 most frequent test labels absent from training are treated as unseen.

    The evaluation uses mean average precision (mAP), which measures label ranking across images, and top-KK F1, which measures label ranking within each image. The paper uses K∈{3,5}K\in\{3,5\} on NUS-WIDE and K∈{10,20}K\in\{10,20\} on Open Images; Open Images additionally reports sample-count-weighted mAP (WmAP).

    The image backbone is ImageNet-1K-pretrained ViT-B/16, and the VLP model is CLIP with a ViT-B/16 image encoder. At 224×224224\times224 resolution, the backbone produces 14×14=19614\times14=196 patch tokens. MKT uses k=18k=18 for local top-kk pooling and λ=1\lambda=1. Stage 1 uses AdamW with learning rate 0.0010.001, weight decay 0.0050.005, batch size 128, and 20 epochs on NUS-WIDE or 4 epochs on Open Images. Stage 2 uses learning rate 0.000030.00003, batch size 16, and 10 epochs on NUS-WIDE or 2 epochs on Open Images.

  7. Knowl 7 — State-of-the-art comparison on NUS-WIDE and Open Images

    data/table

    The main comparison evaluates prior ML-ZSL methods, a fine-tuned CLIP baseline (CLIP-FT), and MKT. Values below are reported percentages; NUS-WIDE uses F1 at K=3,5K=3,5, while Open Images uses F1 at K=10,20K=10,20 and also reports WmAP. The comparison demonstrates that MKT achieves the strongest overall results, especially for unseen-label recognition.

    • LESA (M=10): NUS-WIDE ZSL: F1(3) 31.631.6, F1(5) 28.728.7, mAP 19.419.4; GZSL: 14.414.4, 16.816.8, 5.65.6. Open Images ZSL: F1(10) 1.41.4, F1(20) 1.01.0, mAP 41.741.7, WmAP not reported; GZSL: 17.417.4, 14.314.3, 45.445.4, WmAP not reported.
    • ZS-SDL: NUS-WIDE ZSL: 30.530.5, 27.827.8, 25.925.9; GZSL: 18.518.5, 21.021.0, 12.112.1. Open Images ZSL: 10.710.7, 8.38.3, 62.962.9; GZSL: 37.837.8, 32.932.9, 75.375.3.
    • BiAM: NUS-WIDE ZSL: 32.732.7, 29.829.8, 25.925.9; GZSL: 15.415.4, 18.218.2, 9.49.4. Open Images ZSL: 7.07.0, 5.55.5, mAP 65.665.6, WmAP 72.972.9; GZSL: 14.814.8, 9.79.7, mAP 81.781.7, WmAP 85.085.0.
    • CLIP-FT: NUS-WIDE ZSL: 23.523.5, 21.721.7, 30.530.5; GZSL: 20.320.3, 23.223.2, 16.816.8. Open Images ZSL: 19.119.1, 11.111.1, mAP 66.266.2, WmAP 88.288.2; GZSL: 40.240.2, 35.435.4, mAP 77.577.5, WmAP 85.985.9.
    • MKT: NUS-WIDE ZSL: F1(3) 34.134.1, F1(5) 31.131.1, mAP 37.637.6; GZSL: 22.022.0, 25.425.4, 18.318.3. Open Images ZSL: F1(10) 19.719.7, F1(20) 11.411.4, mAP 68.168.1, WmAP 89.289.2; GZSL: 40.540.5, 35.435.4, mAP 81.481.4, WmAP 89.889.8.

    Relative to the previous best ML-ZSL results, MKT improves NUS-WIDE ZSL mAP over BiAM by 11.7 percentage points and NUS-WIDE GZSL mAP over ZS-SDL by 6.5 points. On Open Images, MKT improves ZSL F1 over ZS-SDL by 9.0 points at K=10K=10 and 2.7 points at K=20K=20, and improves GZSL F1 by 3.1 and 2.5 points, respectively.

  8. Knowl 8 — Ablation of distillation and prompt tuning

    data/table

    On NUS-WIDE, the paper compares four MKT training schemes. Each row reports mAP, F1 at K=3K=3, and F1 at K=5K=5.

    • Without distillation and without prompt tuning: ZSL (32.4,29.4,26.5)(32.4,29.4,26.5); GZSL (16.8,21.0,24.0)(16.8,21.0,24.0).
    • With distillation but without prompt tuning: ZSL (37.3,32.5,29.5)(37.3,32.5,29.5); GZSL (18.2,21.7,24.9)(18.2,21.7,24.9).
    • Without distillation but with prompt tuning: ZSL (32.5,29.5,26.4)(32.5,29.5,26.4); GZSL (16.8,21.1,24.1)(16.8,21.1,24.1).
    • With both distillation and prompt tuning, corresponding to full MKT: ZSL (37.6,34.1,31.1)(37.6,34.1,31.1); GZSL (18.3,22.0,25.4)(18.3,22.0,25.4).

    Knowledge distillation produces the largest isolated improvement, consistent with its role in aligning image and label representations and reducing overfitting to seen labels. Prompt tuning gives an additional improvement when combined with distillation, which the paper attributes to learning context embeddings that better reflect visual information relevant to classification.

  9. Knowl 9 — VLP label embeddings outperform language-only embeddings

    empirical result

    The paper isolates the effect of label representation by comparing GloVe word embeddings with CLIP text embeddings while disabling both knowledge distillation and prompt tuning. On NUS-WIDE, the GloVe model obtains ZSL mAP 27.127.1, F1(3) 22.822.8, and F1(5) 21.421.4, together with GZSL mAP 16.116.1, F1(3) 20.620.6, and F1(5) 23.423.4. The otherwise identical CLIP-embedding model obtains ZSL mAP 32.432.4, F1(3) 29.429.4, and F1(5) 26.526.5, together with GZSL mAP 16.816.8, F1(3) 21.021.0, and F1(5) 24.024.0.

    A label-retrieval experiment uses 62 common NUS-WIDE labels grouped into 14 visually and semantically similar categories. After normalization, the Top-3 retrieval accuracy is 28.49%28.49\% for GloVe, 21.51%21.51\% for BERT, 66.13%66.13\% for CLIP, and 71.51%71.51\% for prompt-tuned CLIP. For example, for the query 'Girls', CLIP retrieves 'Man', 'Kid', and 'School', while prompt-tuned CLIP retrieves 'Man', 'Kid', and 'Person'; for 'Airport', CLIP retrieves 'Airplane', 'Plane', and 'Aircraft', while prompt-tuned CLIP retrieves 'Airplane', 'Aircraft', and 'Plane'. These results support the paper’s claim that VLP embeddings capture visual as well as linguistic consistency among labels.

  10. Knowl 10 — Complementarity of local and global heads and hyperparameter behavior

    data/table

    On NUS-WIDE, the two-stream module is compared with single-head variants. The local-head-only model obtains ZSL mAP 29.129.1, F1(3) 29.929.9, and F1(5) 27.227.2, and GZSL mAP 15.715.7, F1(3) 20.820.8, and F1(5) 23.823.8. The global-head-only model obtains ZSL mAP 30.330.3, F1(3) 23.323.3, and F1(5) 21.421.4, and GZSL mAP 15.515.5, F1(3) 19.419.4, and F1(5) 22.122.1. Combining both heads obtains ZSL mAP 32.432.4, F1(3) 29.429.4, and F1(5) 26.526.5, and GZSL mAP 16.816.8, F1(3) 21.021.0, and F1(5) 24.024.0.

    The global head is more general and tends to yield stronger mAP, whereas the local head is more discriminative and tends to yield stronger ZSL F1 but is more sensitive to noisy high scores. Combining both heads improves resistance to such noise while preserving local discrimination. Hyperparameter studies report that ZSL performance improves as the distillation weight λ\lambda is reduced below 1, but drops when λ\lambda becomes larger than 2; the paper therefore uses λ=1\lambda=1. For the local top-kk pooling parameter on GZSL, F1 is highest at k=18k=18. Smaller kk values are sensitive to local noise, whereas very large kk values become less discriminative; mAP generally increases as kk increases because the pooled score becomes more stable.

  11. Knowl 11 — Qualitative predictions and localization behavior

    empirical result

    The qualitative examples on page 7 show that MKT produces more diverse predictions than CLIP and more semantically coherent predictions than the comparison multi-label model BiAM. In a plane example, MKT assigns similar high scores to the synonymous labels 'plane', 'airplane', and 'aircraft', illustrating the use of VLP label relationships. In other examples, MKT identifies combinations such as flowers, sky, clouds, grass, and plants, or birds, animals, sky, planes, and clouds, rather than concentrating on a narrow set of generic labels.

    Grad-CAM comparisons also show more precise localization. For an image containing a boat, MKT focuses on the boat region, whereas BiAM assigns substantial attention to larger irrelevant areas. These visual results support the intended division of labor between the local feature stream, which captures object-specific regions, and the VLP label embeddings, which encode semantic and visual relationships among arbitrary labels.

Coverage note — No substantial contributed material was omitted; the standard Vision Transformer block recurrence and individual visual panels were treated as supporting implementation or qualitative detail rather than separate load-bearing contributions.

References

  1. 1.Ben-Cohen, A.; Zamir, N.; Ben-Baruch, E.; Friedman, I.; and Zelnik-Manor, L. 2021. Semantic Diversity Learning for Zero-Shot Multi-Label Classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 640–650.
  2. 2.Chen, Y.-C.; Li, L.; Yu, L.; Kholy, A. E.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. J. 2020. UNITER: Learning UNiversal Image-TExt Representations. In European Conference on Computer Vision (ECCV 2020).
  3. 3.Chen, Z.-M.; Wei, X.-S.; Wang, P.; and Guo, Y. 2019. Multi-Label Image Recognition With Graph Convolutional Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  4. 4.Cheng, X.; Lin, H.; Wu, X.; Shen, D.; Yang, F.; Liu, H.; and Shi, N. 2022. Mltr: Multi-label classification with transformer. In 2022 IEEE International Conference on Multimedia and Expo (ICME), 1–6. IEEE.
  5. 5.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. N. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186.
  6. 6.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR 2021: The Ninth International Conference on Learning Representations.
  7. 7.Du, Y.; Wei, F.; Zhang, Z.; Shi, M.; Gao, Y.; and Li, G. 2022. Learning to prompt for open-vocabulary object detection with vision-language model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 14084–14093.
  8. 8.Ghiasi, G.; Gu, X.; Cui, Y.; and Lin, T.-Y. 2022. Scaling open-vocabulary image segmentation with image-level labels. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXVI, 540–557. Springer.
  9. 9.Gong, Y.; Jia, Y.; Leung, T.; Toshev, A.; and Ioffe, S. 2014. Deep Convolutional Ranking for Multilabel Image Annotation. In ICLR 2014 : International Conference on Learning Representations (ICLR) 2014.
  10. 10.Gu, X.; Lin, T.; Kuo, W.; and Cui, Y. 2022. Open-vocabulary Object Detection via Vision and Language Knowledge Distillation. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  11. 11.Gupta, A.; Narayan, S.; Khan, S.; Khan, F. S.; Shao, L.; and van de Weijer, J. 2021. Generative Multi-Label Zero-Shot Learning. arXiv preprint arXiv:2101.11606.
  12. 12.Hinton, G.; Vinyals, O.; Dean, J.; et al. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7).
  13. 13.Huynh, D.; and Elhamifar, E. 2020. A Shared Multi-Attention Framework for Multi-Label Zero-Shot Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  14. 14.Huynh, D.; Kuen, J.; Lin, Z.; Gu, J.; and Elhamifar, E. 2022. Open-vocabulary instance segmentation via robust cross-modal pseudo-labeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7020–7031.
  15. 15.Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.-T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.-H.; Li, Z.; and Duerig, T. 2021. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. In Meila, M.; and Zhang, T., eds., Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, 4904–4916. PMLR.
  16. 16.Kim, W.; Son, B.; and Kim, I. 2021. ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. In ICML 2021: 38th International Conference on Machine Learning, 5583–5594.
  17. 17.Lampert, C. H.; Nickisch, H.; and Harmeling, S. 2009. Learning to detect unseen object classes by between-class attribute transfer. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, 951–958.
  18. 18.Lampert, C. H.; Nickisch, H.; and Harmeling, S. 2014. Attribute-Based Classification for Zero-Shot Visual Object Categorization. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(3): 453–465.
  19. 19.Lanchantin, J.; Wang, T.; Ordonez, V.; and Qi, Y. 2021. General Multi-label Image Classification with Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16478–16488.
  20. 20.Lee, C.-W.; Fang, W.; Yeh, C.-K.; and Wang, Y.-C. F. 2018. Multi-Label Zero-Shot Learning With Structured Knowledge Graphs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  21. 21.Li, G.; Duan, N.; Fang, Y.; Gong, M.; and Jiang, D. 2020. Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training. Proceedings of the AAAI Conference on Artificial Intelligence, 34(07): 11336–11344.
  22. 22.Li, Q.; Qiao, M.; Bian, W.; and Tao, D. 2016. Conditional Graphical Lasso for Multi-Label Image Classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  23. 23.Li, X.; Yin, X.; Li, C.; Zhang, P.; Hu, X.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; Choi, Y.; and Gao, J. 2020. Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks. In European Conference on Computer Vision, 121–137.
  24. 24.Li, X. L.; and Liang, P. 2021. Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Zong, C.; Xia, F.; Li, W.; and Navigli, R., eds., Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, 4582–4597. Association for Computational Linguistics.
  25. 25.Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32.
  26. 26.Ma, Z.; Luo, G.; Gao, J.; Li, L.; Chen, Y.; Wang, S.; Zhang, C.; and Hu, W. 2022. Open-vocabulary one-stage detection with hierarchical visual-language knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 14074–14083.
  27. 27.Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G. S.; and Dean, J. 2013. Distributed Representations of Words and Phrases and their Compositionality. In Advances in Neural Information Processing Systems 26, volume 26, 3111–3119.
  28. 28.Narayan, S.; Gupta, A.; Khan, S.; Khan, F. S.; Shao, L.; and Shah, M. 2021. Discriminative Region-Based Multi-Label Zero-Shot Learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 8731–8740.
  29. 29.Pennington, J.; Socher, R.; and Manning, C. 2014. Glove: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1532–1543.
  30. 30.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models From Natural Language Supervision. In ICML 2021: 38th International Conference on Machine Learning, 8748–8763.
  31. 31.Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8): 9.
  32. 32.Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research, 21(140): 1–67.
  33. 33.Read, J.; Pfahringer, B.; Holmes, G.; and Frank, E. 2011. Classifier chains for multi-label classification. Machine learning, 85(3): 333–359.
  34. 34.Tsoumakas, G.; and Katakis, I. 2007. Multi-label classification: An overview. International Journal of Data Warehousing and Mining (IJDWM), 3(3): 1–13.
  35. 35.Wang, J.; Yang, Y.; Mao, J.; Huang, Z.; Huang, C.; and Xu, W. 2016. CNN-RNN: A Unified Framework for Multi-Label Image Classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  36. 36.Wang, Z.; Chen, T.; Li, G.; Xu, R.; and Lin, L. 2017. Multi-Label Image Recognition by Recurrently Discovering Attentional Regions. In Proceedings of the IEEE International Conference on Computer Vision (ICCV).
  37. 37.Xian, Y.; Schiele, B.; and Akata, Z. 2017. Zero-Shot Learning - the Good, the Bad and the Ugly. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  38. 38.Zang, Y.; Li, W.; Zhou, K.; Huang, C.; and Loy, C. C. 2022. Open-Vocabulary DETR with Conditional Matching. In Avidan, S.; Brostow, G. J.; Cisse, M.; Farinella, G. M.; and ´Hassner, T., eds., Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part IX, volume 13669 of Lecture Notes in Computer Science, 106–122. Springer.
  39. 39.Zareian, A.; Rosa, K. D.; Hu, D. H.; and Chang, S.-F. 2021. Open-Vocabulary Object Detection Using Captions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14393–14402.
  40. 40.Zhang, Y.; Gong, B.; and Shah, M. 2016. Fast Zero-Shot Image Tagging. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 5985–5994.
  41. 41.Zhang, Z.; and Saligrama, V. 2015. Zero-Shot Learning via Semantic Similarity Embedding. In Proceedings of the IEEE International Conference on Computer Vision (ICCV).
  42. 42.Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2022. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9): 2337–2348.
  43. 43.Zhu, F.; Li, H.; Ouyang, W.; Yu, N.; and Wang, X. 2017. Learning Spatial Regularization With Image-Level Supervisions for Multi-Label Image Classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).

Citation

MLA
He, S., et al. “Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer”. arXiv, 2022, http://arxiv.org/abs/2207.01887v2.
APA
He, S., Guo, T., Dai, T., Qiao, R., Ren, B., & Xia, S.-T. (2022). Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer. arXiv. http://arxiv.org/abs/2207.01887v2
Chicago
He, S., T. Guo, T. Dai, R. Qiao, B. Ren, and S.-T. Xia. 2022. “Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer”. arXiv. http://arxiv.org/abs/2207.01887v2.
Harvard
He, S. et al. (2022) “Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2207.01887v2.
Vancouver
1. He S, Guo T, Dai T, Qiao R, Ren B, Xia S-T (2022) Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer. arXiv

BibTeX

@article{he2022open,
  title = {Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer},
  author = {He, Sunan and Guo, Taian and Dai, Tao and Qiao, Ruizhi and Ren, Bo and Xia, Shu-Tao},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2207.01887v2},
  eprint = {2207.01887}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF