Texts as Images in Prompt Tuning for Multi-Label Image Recognition

Zixian GuoBowen DongZhilong JiJinfeng BaiYiwen GuoWangmeng Zuo

article2023CVPR110 citations

Proposes a prompt tuning framework that trains vision-language models for multi-label image recognition using easily accessible text descriptions instead of labeled images, achieving strong classification performance without visual training data.

Listen

Adapting large vision-language artificial intelligence models to recognize multiple objects within a single image typically requires extensive collections of labeled training pictures. Acquiring and manually annotating these image sets is labor-intensive, costly, and frequently impractical in specialized or rapidly evolving domains. While prompt tuning offers a lightweight method to customize models without updating their entire architecture, current approaches still rely on access to annotated visual data. This reliance creates an operational bottleneck whenever image data is scarce or expensive to label.

The article demonstrates that freely available text descriptions can substitute for images during prompt tuning for multi-label image recognition. The authors evaluate this Text-as-Image prompting strategy across diverse benchmark datasets, showing that text alone can effectively train models to identify visual objects.

The approach leverages vision-language models such as CLIP, whose training aligns images and text into a shared feature space. The researchers collect descriptive sentences from standard text sources and run them through a basic noun filter to map synonyms into target object categories, automatically creating category labels without manual image annotation. To handle scenes containing multiple objects, the authors implement double-grained prompt tuning (TaI-DPT), which pairs global sentence-level prompts for overall context with local word-level prompts to capture specific regional objects. These learned prompts are trained using a ranking loss and subsequently deployed to classify actual images across three major benchmarks: VOC2007, MS-COCO, and NUS-WIDE.

The experimental findings show substantial improvements over baseline zero-shot models and conventional few-shot approaches. First, without using any training images, the proposed method outperforms zero-shot CLIP by 9.8% mean average precision on VOC2007, 13.8% on MS-COCO, and 8.5% on NUS-WIDE. Second, the zero-shot text-prompted model matches or exceeds the accuracy of existing methods that were trained on up to 16 labeled images per category. Third, combining text-learned prompts with traditional image-learned prompts consistently enhances overall accuracy in limited-data and partially labeled settings, demonstrating that textual and visual supervision offer complementary advantages.

These results demonstrate that organizations can deploy high-performing image recognition systems without undertaking costly and slow image-gathering campaigns. Bypassing manual image annotation reduces operational overhead, accelerates deployment schedules, and enables rapid model adaptation for emerging visual categories. Furthermore, because text prompts can seamlessly integrate with existing image-based frameworks, teams can improve their current computer vision systems without redesigning underlying pipelines.

Organizations operating in data-constrained visual environments should consider text-driven prompting as an immediate, low-cost baseline. When labeled images are available, teams should ensemble image-based prompts with text-based prompts to maximize classification accuracy. Future initiatives should focus on scaling the approach to larger web-crawled text corpora and testing performance across broader enterprise domains.

The primary limitation of this study lies in its reliance on simple noun filtering, which can miss complex phrasing, paraphrases, or misspellings in free-form language. Additionally, performance gains were less pronounced when text captions were drawn strictly from the target image set rather than broader external language sources. Nevertheless, confidence in the findings is high, supported by consistent gains across multiple standardized benchmarks and rigorous ablation testing.

arXiv: 2211.12739

No sufficiently relevant recommendations were found.

Cover for Texts as Images in Prompt Tuning for Multi-Label Image Recognition

Abstract

Prompt tuning has been employed as an efficient way to adapt large vision-language pre-trained models (e.g. CLIP) to various downstream tasks in data-limited or label-limited settings. Nonetheless, visual data (e.g., images) is by default prerequisite for learning prompts in existing methods. In this work, we advocate that the effectiveness of image-text contrastive learning in aligning the two modalities (for training CLIP) further makes it feasible to treat texts as images for prompt tuning and introduce TaI prompting. In contrast to the visual data, text descriptions are easy to collect, and their class labels can be directly derived. Particularly, we apply TaI prompting to multi-label image recognition, where sentences in the wild serve as alternatives to images for prompt tuning. Moreover, with TaI, dual-grained prompt tuning (TaI-DPT) is further presented to extract both coarse-grained and fine-grained embeddings for enhancing the multi-label recognition performance. Experimental results show that our proposed TaI-DPT outperforms zero-shot CLIP by a large margin on multiple benchmarks, e.g., MS-COCO, VOC2007, and NUS-WIDE, while it can be combined with existing methods of prompting from images to improve recognition performance further. The code is released at https://github.com/guozix/TaI-DPT.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Multi-Label Image Recognition
  • 2.2. Prompt Tuning for Vision-Language Models
  • 3. Proposed Method
  • 3.1. Overview of Our Method
  • 3.2. Preparation of Text Descriptions
  • 3.3. Text-as-Image for Dual-grained Prompt Tuning
  • 3.4. Learning Objective
  • 3.5. Incorporating with Prompting from Images
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Comparison with Zero-Shot Methods
  • 4.3. Comparison with Few-Shot Methods
  • 4.4. Integration with Partially Labeled Methods
  • 4.5. Ablation Study
  • 5. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Text-as-Image prompt tuning

    model/method

    Text-as-Image (TaI) prompting adapts a frozen CLIP-style vision-language model to multi-label image recognition without using labeled downstream images during prompt training. Given free-form text descriptions and a target class vocabulary, TaI uses the text encoder to encode both learnable class prompts and the descriptions, derives multi-label pseudo-labels from class names appearing in the descriptions, and optimizes only the prompt embeddings. At inference time, the learned text prompts are encoded into class embeddings, while the input modality is changed from text descriptions to test images encoded by the frozen image encoder. The method relies on the image-text alignment learned during CLIP pre-training: text descriptions of objects are treated as alternatives to image examples in prompt learning.

  2. Knowl 2 — Noun-filtered pseudo-label construction

    model/method

    For a target multi-label dataset with class set S={s1,…,sC}S=\{s_1,\ldots,s_C\}, TaI constructs training data from captions or localized narratives while discarding the paired images and their original labels. A synonym dictionary DD contains each class name and manually specified alternatives, such as dog/pup/puppy, person/people/human, bicycle/bike, and car/taxi. Each text description is tokenized and lemmatized with NLTK; it is retained only if at least one token matches a synonym in DD. For every retained description, the matched classes form the positive set and all unmatched classes form the negative set, producing a binary pseudo-label vector in the target dataset’s class order. The procedure is reproducible and guarantees at least one positive class per training text, but its labels can be inaccurate when free-form language uses paraphrases, misspellings, or concepts not covered by the synonym dictionary.

  3. Knowl 3 — Double-grained prompt construction

    model/method

    TaI-DPT uses separate global and local prompts for each of the CC target classes. For class ii, let sis_i be the fixed word embedding of the class name, let MM be the number of learnable context embeddings, and let vjG,vjL∈Rdv_j^G,v_j^L\in\mathbb{R}^d be learnable word embeddings of the same dimension dd as CLIP vocabulary embeddings. The prompts are

    tiG=[v1G,…,vMG,si],tiL=[v1L,…,vML,si].t_i^G=[v_1^G,\ldots,v_M^G,s_i],\qquad t_i^L=[v_1^L,\ldots,v_M^L,s_i].

    A frozen CLIP text encoder Enc⁡T\operatorname{Enc}_T maps them to global and local class embeddings Gi=Enc⁡T(tiG)G_i=\operatorname{Enc}_T(t_i^G) and Li=Enc⁡T(tiL)L_i=\operatorname{Enc}_T(t_i^L). The global branch is intended to represent an entire sentence or image, whereas the local branch is intended to match individual text tokens or spatial image regions. Qualitative localization results show that the learned local embeddings correlate with words describing their classes and focus on corresponding object regions, despite the local features not receiving explicit fine-grained supervision during CLIP pre-training.

  4. Knowl 4 — Global and local similarity scoring

    equation

    For a training text description rr, the frozen text encoder produces a global feature h∈RDh\in\mathbb{R}^D and a sequence of N2N_2 token features H∈RN2×DH\in\mathbb{R}^{N_2\times D}. For a test image xx, the frozen image encoder produces a global feature f∈RDf\in\mathbb{R}^D and N1N_1 flattened dense spatial features F∈RN1×DF\in\mathbb{R}^{N_1\times D}. Using cosine similarity ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle, the class-ii global and token/region similarities are

    pi=⟨u,Gi⟩,Pij=⟨Uj,Li⟩,p_i=\langle u,G_i\rangle,\qquad P_{ij}=\langle U_j,L_i\rangle,

    where (u,U)=(h,H)(u,U)=(h,H) during prompt training and (u,U)=(f,F)(u,U)=(f,F) during image testing. The local similarities are aggregated with a softmax weighting controlled by a positive temperature τs\tau_s:

    pi′=∑j=1Nexp⁡(Pij/τs)∑k=1Nexp⁡(Pik/τs)Pij,p_i'=\sum_{j=1}^{N}\frac{\exp(P_{ij}/\tau_s)}{\sum_{k=1}^{N}\exp(P_{ik}/\tau_s)}P_{ij},

    where N=N2N=N_2 for text tokens and N=N1N=N_1 for image regions. At test time, the global score pip_i and aggregated local score pi′p_i' are fused for each class, using element-wise score addition in the proposed pipeline.

  5. Knowl 5 — Ranking-loss training objective

    equation

    Let c+c^+ be the set of positive classes and c−c^- the set of negative classes in the noun-filtered pseudo-label vector for a training description. TaI-DPT minimizes a sum of global and local pairwise ranking losses:

    L=Lglobal+Llocal,\mathcal{L}=\mathcal{L}_{\mathrm{global}}+\mathcal{L}_{\mathrm{local}},

    Lglobal=∑i∈c+∑j∈c−max⁡(0,m−pi+pj),\mathcal{L}_{\mathrm{global}}=\sum_{i\in c^+}\sum_{j\in c^-}\max(0,m-p_i+p_j),

    Llocal=∑i∈c+∑j∈c−max⁡(0,m−pi′+pj′),\mathcal{L}_{\mathrm{local}}=\sum_{i\in c^+}\sum_{j\in c^-}\max(0,m-p_i'+p_j'),

    where pip_i and pi′p_i' are the global and aggregated local similarities for class ii, and m>0m>0 is the margin. The experiments use m=1m=1. The text encoders remain frozen and only the global and local prompt embeddings are optimized. The authors use ranking loss instead of directly applying sigmoid-based binary cross-entropy because direct probability optimization produced a gap between text-training behavior and image-testing behavior, which they attribute to the vision-language modality gap.

  6. Knowl 6 — Score-level integration with image-trained prompts

    model/method

    TaI-DPT can be combined with prompt-tuning methods trained from images by fusing their per-class prediction scores with TaI-DPT scores through a weighted sum. With a few annotated images, the image-trained method can be CoOp; with partially labeled images, it can be DualCoOp. Score fusion is used rather than averaging class embeddings because the component methods may use different image encoders, such as CLIP ResNet-50 for TaI-DPT and ResNet-101 in the partial-label DualCoOp setting. This makes the integration independent of the particular image-encoder architecture and allows text-derived and image-derived prompt information to contribute jointly.

  7. Knowl 7 — Training and evaluation configuration

    experimental setup

    The experiments use CLIP ResNet-50 as the frozen image encoder and the CLIP Transformer as the frozen text encoder. Global and local prompts are shared across classes and datasets; each contains M=16M=16 learnable context embeddings initialized independently from N(0,0.02)\mathcal{N}(0,0.02). Prompt parameters are optimized with SGD for 20 epochs. The initial learning rates are 10−410^{-4} for MS-COCO, 10−410^{-4} for VOC2007, and 10−310^{-3} for NUS-WIDE, followed by cosine annealing. The global and local similarity scores are scaled by 4, and the local aggregation temperature is τs=0.02\tau_s=0.02. Evaluation uses VOC2007 with 20 classes and 5,011/4,952 train/test images, MS-COCO with 80 classes and 82,081/40,504 train/test images, and NUS-WIDE with 81 concepts, 161,789 training images, and 107,859 test images. In zero-shot experiments, no images from the target training sets are used: MS-COCO captions supply text for VOC2007 and MS-COCO, while OpenImages localized narratives supply broader text coverage for NUS-WIDE.

  8. Knowl 8 — Zero-shot multi-label recognition gains

    data/table

    The zero-shot evaluation compares standard CLIP prompting (ZSCLIP) and TaI prompting, with and without double-grained prompt tuning. The metric is mean average precision (mAP); no labeled target-dataset images are used to learn TaI prompts.

    Method DPT VOC2007 MS-COCO NUS-WIDE
    ZSCLIP ×\times 76.2 47.3 36.4
    ZSCLIP ✓\checkmark 77.3 49.7 37.4
    TaI ×\times 86.0 61.1 44.9
    TaI ✓\checkmark 88.3 65.1 46.5

    TaI without local prompting improves over ordinary ZSCLIP by 9.8, 13.8, and 8.5 mAP points on VOC2007, MS-COCO, and NUS-WIDE, respectively. Adding the local branch produces the best scores on all three datasets, showing that text-trained fine-grained prompt embeddings improve image-side multi-label recognition.

  9. Knowl 9 — Few-shot comparison on novel classes

    empirical result

    On MS-COCO with 16 novel classes, zero-shot TaI-DPT reaches 59.2 mAP, exceeding the reported 1-shot results of LaSO and ML-FSL and exceeding the corresponding zero-shot CLIP baselines used by CoOp and Tip-Adapter. The comparison is:

    Method 0-shot 1-shot 5-shot
    LaSO - 45.3 58.1
    ML-FSL - 54.4 63.6
    CoOp 40.2 ZSCLIP) 46.9 55.6
    Tip-Adapter 40.2 ZSCLIP) 53.8 59.7
    TaI-DPT 59.2 - -

    In a separate setting where every class is treated as novel, TaI-DPT is reported to achieve results comparable to CoOp trained with 16 labeled images per class. Combining TaI-DPT with CoOp-DPT through score ensembling further improves recognition over either prompt source across the evaluated shot counts, indicating complementary information in text and images.

  10. Knowl 10 — Improvement under partial image labels

    data/table

    Partially labeled training data are created by randomly masking labels from fully annotated images, while inference evaluates all categories. TaI-DPT is fused with a reproduced DualCoOp model, denoted DualCoOp*, and improves the reproduced model’s average mAP on all three datasets. The complete results are:

    Dataset Method 10% 20% 30% 40% 50% 60% 70% 80% 90% Avg.
    MS-COCO SARB 71.2 75.0 77.1 78.3 78.9 79.6 79.8 80.5 80.5 77.9
    MS-COCO DualCoOp 78.7 80.9 81.7 82.0 82.5 82.7 82.8 83.0 83.1 81.9
    MS-COCO DualCoOp* 81.0 82.3 82.9 83.4 83.5 83.9 84.0 84.1 84.3 83.3
    MS-COCO +TaI-DPT 81.5 82.6 83.3 83.7 83.9 84.0 84.2 84.4 84.5 83.6
    PascalVOC 2007 SARB 83.5 88.6 90.7 91.4 91.9 92.2 92.6 92.8 92.9 90.7
    PascalVOC 2007 DualCoOp 90.3 92.2 92.8 93.3 93.6 93.9 94.0 94.1 94.2 93.2
    PascalVOC 2007 DualCoOp* 91.4 93.8 93.8 94.3 94.6 94.7 94.8 94.9 94.9 94.1
    PascalVOC 2007 +TaI-DPT 93.3 94.6 94.8 94.9 95.1 95.0 95.1 95.3 95.5 94.8
    NUS-WIDE DualCoOp* 54.0 56.2 56.9 57.4 57.9 57.9 57.6 58.2 58.8 57.2
    NUS-WIDE +TaI-DPT 56.4 57.9 57.8 58.1 58.5 58.8 58.6 59.1 59.4 58.3

    The average mAP increases from 83.3 to 83.6 on MS-COCO, from 94.1 to 94.8 on PascalVOC 2007, and from 57.2 to 58.3 on NUS-WIDE after adding TaI-DPT. The gains are especially notable when the text source is external to the image dataset; the benefit is smaller on MS-COCO because its captions are derived from the same images used by the image-trained method.

Coverage note — The ablation showing that performance rises as the number of VOC2007 text descriptions increases, up to a pool of 66,087 texts, is omitted as a secondary diagnostic; the qualitative evidence that local prompts attend to class-relevant words and image regions is incorporated into the double-grained prompt knowl.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022. 1
  2. 2.Amit Alfassy, Leonid Karlinsky, Amit Aides, Joseph Shtok, Sivan Harary, Rogerio Feris, Raja Giryes, and Alex M Bronstein. Laso: Label-set operations networks for multi-label few-shot learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6548–6557, 2019. 2, 7
  3. 3.Emanuel Ben-Baruch, Tal Ridnik, Nadav Zamir, Asaf Noy, Itamar Friedman, Matan Protter, and Lihi Zelnik-Manor. Asymmetric loss for multi-label classification, 2021. 2
  4. 4.Steven Bird, Ewan Klein, and Edward Loper. Natural language processing with Python: analyzing text with the natural language toolkit. ” O’Reilly Media, Inc.”, 2009. 4
  5. 5.Tianshui Chen, Tao Pu, Hefeng Wu, Yuan Xie, and Liang Lin. Structured semantic transfer for multi-label recognition with partial labels. In Proceedings of the AAAI conference on artificial intelligence, volume 36, pages 339–346, 2022. 2, 8
  6. 6.Tianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu, and Liang Lin. Learning semantic-specific graph representation for multi-label image recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pages 522–531, 2019. 2, 6
  7. 7.Zhao-Min Chen, Xiu-Shen Wei, Xin Jin, and Yanwen Guo. Multi-label image recognition with joint class-aware map disentangling and label correlation embedding. In ICME 2019. 2
  8. 8.Zhao-Min Chen, Xiu-Shen Wei, Peng Wang, and Yanwen Guo. Multi-label image recognition with graph convolutional networks. In CVPR, 2019. 2, 6
  9. 9.Tat-Seng Chua, Jinhui Tang, Richang Hong, Haojie Li, Zhiping Luo, and Yantao Zheng. Nus-wide: a real-world web image database from national university of singapore. In Proceedings of the ACM international conference on image and video retrieval, pages 1–9, 2009. 2, 6, 7
  10. 10.Thibaut Durand, Nazanin Mehrasa, and Greg Mori. Learning a deep convnet for multi-label classification with partial labels. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 647–657, 2019. 2
  11. 11.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010. 2, 6, 7
  12. 12.Bin-Bin Gao and Hong-Yu Zhou. Multi-label image recognition with multi-class attentional regions. arXiv preprint arXiv:2007.01755, 2020. 2
  13. 13.Chunjiang Ge, Rui Huang, Mixue Xie, Zihang Lai, Shiji Song, Shuang Li, and Gao Huang. Domain adaptation via prompt learning. arXiv preprint arXiv:2202.06687, 2022. 3
  14. 14.Yunchao Gong, Yangqing Jia, Thomas Leung, Alexander Toshev, and Sergey Ioffe. Deep convolutional ranking for multilabel image annotation. arXiv preprint arXiv:1312.4894, 2013. 2, 6
  15. 15.Shiyi He, Chang Xu, Tianyu Guo, Chao Xu, and Dacheng Tao. Reinforced multi-label image classification by exploring curriculum. In AAAI, 2018. 2
  16. 16.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904–4916. PMLR, 2021. 1
  17. 17.Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. arXiv preprint arXiv:2203.12119, 2022. 3
  18. 18.Ivan Krasin, Tom Duerig, Neil Alldrin, Vittorio Ferrari, Sami Abu-El-Haija, Alina Kuznetsova, Hassan Rom, Jasper Uijlings, Stefan Popov, Shahab Kamali, Matteo Malloci, Jordi Pont-Tuset, Andreas Veit, Serge Belongie, Victor Gomes, Abhinav Gupta, Chen Sun, Gal Chechik, David Cai, Zheyun Feng, Dhyanesh Narayanan, and Kevin Murphy. Openimages: A public dataset for large-scale multi-label and multi-class image classification. Dataset available from https://storage.googleapis.com/openimages/web/index.html, 2017. 2, 4, 6
  19. 19.Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan. Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. arXiv preprint arXiv:2110.05208, 2021. 1
  20. 20.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 2, 4, 6, 7
  21. 21.Luchen Liu, Sheng Guo, Weilin Huang, and Matthew Scott. Decoupling category-wise independence and relevance with self-attention for multi-label image classification. In ICASSP 2019, 05. 2
  22. 22.Yuning Lu, Jianzhuang Liu, Yonggang Zhang, Yajing Liu, and Xinmei Tian. Prompt distribution learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5206–5215, 2022. 3
  23. 23.Tao Pu, Tianshui Chen, Hefeng Wu, and Liang Lin. Semantic-aware representation blending for multi-label image recognition with partial labels. arXiv preprint arXiv:2203.02172, 2022. 2, 8
  24. 24.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021. 1, 2, 4, 6, 7, 8
  25. 25.Tal Ridnik, Emanuel Ben-Baruch, Nadav Zamir, Asaf Noy, Itamar Friedman, Matan Protter, and Lihi Zelnik-Manor. Asymmetric loss for multi-label classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 82–91, 2021. 6
  26. 26.Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, and Chaowei Xiao. Test-time prompt tuning for zero-shot generalization in vision-language models. arXiv preprint arXiv:2209.07511, 2022. 3
  27. 27.Christian Simon, Piotr Koniusz, and Mehrtash Harandi. Meta-learning for multi-label few-shot classification. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 3951–3960, 2022. 2, 7
  28. 28.Ximeng Sun, Ping Hu, and Kate Saenko. Dualcoop: Fast adaptation to multi-label recognition with limited annotations. arXiv preprint arXiv:2206.09541, 2022. 1, 2, 3, 5, 6, 8
  29. 29.Ashwin Vaswani, Gaurav Aggarwal, Praneeth Netrapalli, and Narayan G Hegde. All mistakes are not equal: Comprehensive hierarchy aware multi-label predictions (champ). arXiv preprint arXiv:2206.08653, 2022. 2
  30. 30.Jiang Wang, Yi Yang, Junhua Mao, Zhiheng Huang, Chang Huang, and Wei Xu. Cnn-rnn: A unified framework for multi-label image classification. In CVPR, 2016. 2
  31. 31.Yangtao Wang, Yanzhao Xie, Yu Liu, Ke Zhou, and Xiaocui Li. Fast graph convolution network based multi-label image recognition via cross-modal fusion. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, pages 1575–1584, 2020. 2
  32. 32.Yunchao Wei, Wei Xia, Min Lin, Junshi Huang, Bingbing Ni, Jian Dong, Yao Zhao, and Shuicheng Yan. Hcp: A flexible cnn framework for multi-label image classification. IEEE transactions on pattern analysis and machine intelligence, 38, 2015. 2
  33. 33.Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. Filip: Fine-grained interactive language-image pre-training. arXiv preprint arXiv:2111.07783, 2021. 1
  34. 34.Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun. Cpt: Colorful prompt tuning for pre-trained vision-language models. arXiv preprint arXiv:2109.11797, 2021. 3
  35. 35.Jin Ye, Junjun He, Xiaojiang Peng, Wenhao Wu, and Yu Qiao. Attention-driven dynamic graph convolutional network for multi-label image recognition. In ECCV, 2020. 2
  36. 36.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021. 1
  37. 37.Junjie Zhang, Qi Wu, Chunhua Shen, Jian Zhang, and Jianfeng Lu. Multilabel image classification with regional latent semantic dependencies. IEEE Transactions on Multimedia, 20, 2018. 2
  38. 38.Renrui Zhang, Zhang Wei, Rongyao Fang, Peng Gao, Kunchang Li, Jifeng Dai, Yu Qiao, and Hongsheng Li. Tip-adapter: Training-free adaption of clip for few-shot classification. arXiv preprint arXiv:2207.09519, 2022. 7
  39. 39.Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. Neural prompt search. arXiv preprint arXiv:2206.04673, 2022. 3
  40. 40.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16816–16825, 2022. 1, 3
  41. 41.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9):2337–2348, 2022. 1, 3, 4, 6, 7
  42. 42.Beier Zhu, Yulei Niu, Yucheng Han, Yue Wu, and Hanwang Zhang. Prompt-aligned gradient for prompt tuning. arXiv preprint arXiv:2205.14865, 2022. 3

Citation

MLA
Guo, Z., et al. “Texts as Images in Prompt Tuning for Multi-Label Image Recognition”. arXiv, 2022, http://arxiv.org/abs/2211.12739v2.
APA
Guo, Z., Dong, B., Ji, Z., Bai, J., Guo, Y., & Zuo, W. (2022). Texts as Images in Prompt Tuning for Multi-Label Image Recognition. arXiv. http://arxiv.org/abs/2211.12739v2
Chicago
Guo, Z., B. Dong, Z. Ji, J. Bai, Y. Guo, and W. Zuo. 2022. “Texts as Images in Prompt Tuning for Multi-Label Image Recognition”. arXiv. http://arxiv.org/abs/2211.12739v2.
Harvard
Guo, Z. et al. (2022) “Texts as Images in Prompt Tuning for Multi-Label Image Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.12739v2.
Vancouver
1. Guo Z, Dong B, Ji Z, Bai J, Guo Y, Zuo W (2022) Texts as Images in Prompt Tuning for Multi-Label Image Recognition. arXiv

BibTeX

@article{guo2022texts,
  title = {Texts as Images in Prompt Tuning for Multi-Label Image Recognition},
  author = {Guo, Zixian and Dong, Bowen and Ji, Zhilong and Bai, Jinfeng and Guo, Yiwen and Zuo, Wangmeng},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.12739v2},
  eprint = {2211.12739}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE