Texts as Images in Prompt Tuning for Multi-Label Image Recognition
Zixian GuoBowen DongZhilong JiJinfeng BaiYiwen GuoWangmeng Zuo
Proposes a prompt tuning framework that trains vision-language models for multi-label image recognition using easily accessible text descriptions instead of labeled images, achieving strong classification performance without visual training data.
Adapting large vision-language artificial intelligence models to recognize multiple objects within a single image typically requires extensive collections of labeled training pictures. Acquiring and manually annotating these image sets is labor-intensive, costly, and frequently impractical in specialized or rapidly evolving domains. While prompt tuning offers a lightweight method to customize models without updating their entire architecture, current approaches still rely on access to annotated visual data. This reliance creates an operational bottleneck whenever image data is scarce or expensive to label.
The article demonstrates that freely available text descriptions can substitute for images during prompt tuning for multi-label image recognition. The authors evaluate this Text-as-Image prompting strategy across diverse benchmark datasets, showing that text alone can effectively train models to identify visual objects.
The approach leverages vision-language models such as CLIP, whose training aligns images and text into a shared feature space. The researchers collect descriptive sentences from standard text sources and run them through a basic noun filter to map synonyms into target object categories, automatically creating category labels without manual image annotation. To handle scenes containing multiple objects, the authors implement double-grained prompt tuning (TaI-DPT), which pairs global sentence-level prompts for overall context with local word-level prompts to capture specific regional objects. These learned prompts are trained using a ranking loss and subsequently deployed to classify actual images across three major benchmarks: VOC2007, MS-COCO, and NUS-WIDE.
The experimental findings show substantial improvements over baseline zero-shot models and conventional few-shot approaches. First, without using any training images, the proposed method outperforms zero-shot CLIP by 9.8% mean average precision on VOC2007, 13.8% on MS-COCO, and 8.5% on NUS-WIDE. Second, the zero-shot text-prompted model matches or exceeds the accuracy of existing methods that were trained on up to 16 labeled images per category. Third, combining text-learned prompts with traditional image-learned prompts consistently enhances overall accuracy in limited-data and partially labeled settings, demonstrating that textual and visual supervision offer complementary advantages.
These results demonstrate that organizations can deploy high-performing image recognition systems without undertaking costly and slow image-gathering campaigns. Bypassing manual image annotation reduces operational overhead, accelerates deployment schedules, and enables rapid model adaptation for emerging visual categories. Furthermore, because text prompts can seamlessly integrate with existing image-based frameworks, teams can improve their current computer vision systems without redesigning underlying pipelines.
Organizations operating in data-constrained visual environments should consider text-driven prompting as an immediate, low-cost baseline. When labeled images are available, teams should ensemble image-based prompts with text-based prompts to maximize classification accuracy. Future initiatives should focus on scaling the approach to larger web-crawled text corpora and testing performance across broader enterprise domains.
The primary limitation of this study lies in its reliance on simple noun filtering, which can miss complex phrasing, paraphrases, or misspellings in free-form language. Additionally, performance gains were less pronounced when text captions were drawn strictly from the target image set rather than broader external language sources. Nevertheless, confidence in the findings is high, supported by consistent gains across multiple standardized benchmarks and rigorous ablation testing.
- Paper: DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations, Ximeng Sun et al. (2022). DualCoOp introduces prompt tuning for multi-label recognition on vision-language models like CLIP, establishing the baseline framework that this work adapts by substituting training images with text.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). Context Optimization (CoOp) pioneered continuous prompt tuning for pre-trained vision-language models, forming the core algorithmic foundation upon which subsequent prompt-learning methods build.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. CLIP provides the underlying pre-trained multimodal joint embedding space that aligns vision and language, which enables text descriptions to serve as direct proxies for images.
- Paper: Conditional Prompt Learning for Vision-Language Models, Kaiyang Zhou et al. (2022). CoCoOp establishes conditioning strategies to address generalization limits in static prompt tuning, which provides important context for dual-grained and multi-label prompt design.
- Paper: CLIP-Adapter: Better Vision-Language Models with Feature Adapters, Peng Gao et al. (2021). CLIP-Adapter introduces parameter-efficient feature adaptation for frozen vision-language encoders, representing a key alternative parameter-efficient transfer paradigm.
- Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). Visual Prompt Tuning establishes how prompt tokens can be injected directly into visual transformer architectures as a lightweight adaptation strategy.
No sufficiently relevant recommendations were found.
