Built independently by an author, for readers. Read the story and support ChapterPal

keyword

pre-trained CLIP models

Pre-trained CLIP models are multimodal vision-language neural networks that have been trained on large-scale datasets of paired images and natural language descriptions using contrastive learning. Consisting of joint image and text encoders optimized to map visual and linguistic concepts into a shared embedding space, these models align visual features directly with textual semantics. Because they learn rich, open-vocabulary representations without relying on rigid, predefined category labels, pre-trained CLIP models exhibit strong generalization capabilities and can perform zero-shot visual classification. They frequently serve as foundational backbones that can be adapted, fine-tuned, or integrated into broader architectures for diverse downstream tasks, including image retrieval, object detection, and open-vocabulary image segmentation.

3 items

DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting

DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting

Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, Jiwen Lu

OrganizationsPhigent RoboticsTsinghua University

Why you should read this

Proposes a model-agnostic framework that adapts pre-trained vision-language models to dense prediction tasks like semantic segmentation and object detection by converting image-text matching into pixel-text score maps guided by context-aware prompting.

Recent progress has shown that large-scale pre-training using contrastive image-text pairs can be a promising alternative for high-quality visual representation learning from natural language supervision. Benefiting from a broader source of supervision, this new paradigm exhibits impressive transferability to downstream classification tasks and datasets. However, the problem of transferring the knowledge learned from image-text pairs to more complex dense prediction tasks has barely been visited. In this work, we present a new framework for dense prediction by implicitly and explicitly leveraging the pre-trained knowledge from CLIP. Specifically, we convert the original image-text matching problem in CLIP to a pixel-text matching problem and use the pixel-text score maps to guide the learning of dense prediction models. By further using the contextual information from the image to prompt the language model, we are able to facilitate our model to better exploit the pre-trained knowledge. Our method is model-agnostic, which can be applied to arbitrary dense prediction systems and various pre-trained visual backbones including both CLIP models and ImageNet pre-trained models. Extensive experiments demonstrate the superior performance of our methods on semantic segmentation, object detection, and instance segmentation tasks. Code is available at https://github.com/raoyongming/DenseCLIP.

Added

2026-10-05

Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models

Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models

Yabin Zhang, Wenjie Zhu, Hui Tang, Zhiyuan Ma, Kaiyang Zhou, Lei Zhang

OrganizationsHong Kong Baptist UniversityHong Kong Polytechnic UniversityOPPOThe Hong Kong University of Science and Technology

Why you should read this

Proposes Dual Memory Networks, a unified framework combining static training caches with dynamic test-time memory to adapt pre-trained vision-language models across zero-shot, few-shot, and training-free settings without relying on external data.

With the emergence of pre-trained vision-language models like CLIP, how to adapt them to various downstream classification tasks has garnered significant attention in recent research. The adaptation strategies can be typically categorized into three paradigms: zero-shot adaptation, few-shot adaptation, and the recently-proposed training-free few-shot adaptation. Most existing approaches are tailored for a specific setting and can only cater to one or two of these paradigms. In this paper, we introduce a versatile adaptation approach that can effectively work under all three settings. Specifically, we propose the dual memory networks that comprise dynamic and static memory components. The static memory caches training data knowledge, enabling training-free few-shot adaptation, while the dynamic memory preserves historical test features online during the testing process, allowing for the exploration of additional data insights beyond the training set. This novel capability enhances model performance in the few-shot setting and enables model usability in the absence of training data. The two memory networks employ the same flexible memory interactive strategy, which can operate in a training-free mode and can be further enhanced by incorporating learnable projection layers. Our approach is tested across 11 datasets under the three task settings. Remarkably, in the zero-shot scenario, it outperforms existing methods by over 3% and even shows superior results against methods utilizing external training data. Additionally, our method exhibits robust performance against natural distribution shifts. Codes are available at https://github.com/YBZh/DMN.

Added

2026-09-26

CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor

CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor

Shuyang Sun, Runjia Li, Philip Torr, Xiuye Gu, Siyang Li

OrganizationsGoogleUniversity of Oxford

Why you should read this

Presents a training-free recurrent framework that leverages frozen CLIP models to iteratively refine mask proposals and filter non-existent text queries, setting new state-of-the-art benchmarks in zero-shot open-vocabulary semantic and referring segmentation without sacrificing vocabulary breadth.

Existing open-vocabulary image segmentation methods require a fine-tuning step on mask labels and/or image-text datasets. Mask labels are labor-intensive, which limits the number of categories in segmentation datasets. Consequently, the vocabulary capacity of pre-trained VLMs is severely reduced after fine-tuning. However, without fine-tuning, VLMs trained under weak image-text supervision tend to make suboptimal mask predictions. To alleviate these issues, we introduce a novel recurrent framework that progressively filters out irrelevant texts and enhances mask quality without training efforts. The recurrent unit is a two-stage segmenter built upon a frozen VLM. Thus, our model retains the VLM’s broad vocabulary space and equips it with segmentation ability. Experiments show that our method outperforms not only the training-free counterparts, but also those fine-tuned with millions of data samples, and sets the new state-of-the-art records for both zero-shot semantic and referring segmentation. Concretely, we improve the current record by 28.8, 16.0, and 6.9 mIoU on Pascal VOC, COCO Object, and Pascal Context.

Added

2026-09-26