Built independently by an author, for readers. Read the story and support ChapterPal

keyword

vision-language pretraining

Vision-language pretraining is a machine learning paradigm in which computational models are trained on large-scale collections of paired or correlated visual and textual data to learn shared cross-modal representations before being adapted to specific downstream applications. By employing training objectives such as contrastive learning, masked multimodal modeling, and prefix language modeling, this process aligns visual inputs—such as whole images, localized image regions, or multidimensional spatial data—with natural language concepts in a joint feature space. The resulting generalized representations allow models to transfer effectively to a wide range of tasks, including visual question answering, image-text retrieval, image captioning, and zero-shot or few-shot visual recognition.

6 items

SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, Yuan Cao

OrganizationsCarnegie Mellon UniversityGoogleUniversity of Washington

Why you should read this

Introduces SimVLM, a simplified vision-language model trained end-to-end on weakly supervised data using a single prefix language modeling objective, achieving state-of-the-art benchmark performance and strong zero-shot multimodal capabilities without requiring expensive object-level annotations.

With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations including clean image captions and regional labels limits the scalability of existing approaches, and complicates the pretraining procedure with the introduction of multiple dataset-specific objectives. In this work, we relax these constraints and present a minimalist pretraining framework, named Simple Visual Language Model (SimVLM). Unlike prior work, SimVLM reduces the training complexity by exploiting large-scale weak supervision, and is trained end-to-end with a single prefix language modeling objective. Without utilizing extra data or task-specific customization, the resulting model significantly outperforms previous pretraining methods and achieves new state-of-the-art results on a wide range of discriminative and generative vision-language benchmarks, including VQA (+3.74% vqa-score), NLVR2 (+1.17% accuracy), SNLI-VE (+1.37% accuracy) and image captioning tasks (+10.1% average CIDEr score). Furthermore, we demonstrate that SimVLM acquires strong generalization and transfer ability, enabling zero-shot behavior including open-ended visual question answering and cross-modality transfer.

Added

2026-10-05

Robust Cross-Modal Representation Learning with Progressive Self-Distillation

Robust Cross-Modal Representation Learning with Progressive Self-Distillation

Alex Andonian, Shixing Chen, Raffay Hamid

OrganizationsAmazonMassachusetts Institute of Technology

Why you should read this

Proposes a progressive self-distillation framework that replaces rigid one-to-one pairings in vision-language pretraining with dynamic soft-alignment targets, consistently outperforming CLIP across zero-shot, transfer, and retrieval benchmarks without extra computational overhead.

The learning objective of vision-language approach of CLIP [63] does not effectively account for the noisy many-to-many correspondences found in web-harvested image captioning datasets, which contributes to its compute and data inefficiency. To address this challenge, we introduce a novel training framework based on cross-modal contrastive learning that uses progressive self-distillation and soft image-text alignments to more efficiently learn robust representations from noisy data. Our model distills its own knowledge to dynamically generate soft-alignment targets for a subset of images and captions in every minibatch, which are then used to update its parameters. Extensive evaluation across 14 benchmark datasets shows that our method consistently outperforms its CLIP counterpart in multiple settings, including: (a) zero-shot classification, (b) linear probe transfer, and (c) image-text retrieval, without incurring extra computational cost. Analysis using an ImageNet-based robustness test-bed [70] reveals that our method offers better effective robustness to natural distribution shifts compared to both ImageNet-trained models and CLIP itself. Lastly, pretraining with datasets spanning two orders of magnitude in size shows that our improvements over CLIP tend to scale with number of training examples.

Added

2026-09-26

HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention

HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention

Shijie Geng, Jianbo Yuan, Yu Tian, Yuxiao Chen, Yongfeng Zhang

OrganizationsByteDanceRutgers University

Why you should read this

Introduces HiCLIP to integrate hierarchy-aware attention into contrastive vision-language pretraining, enabling the unsupervised discovery of multi-level semantics across images and text to improve cross-modal alignment on downstream tasks.

The success of large-scale contrastive vision-language pretraining (CLIP) has benefited both visual recognition and multimodal content understanding. The concise design brings CLIP the advantage in inference efficiency against other vision-language models with heavier cross-attention fusion layers, making it a popular choice for a wide spectrum of downstream tasks. However, CLIP does not explicitly capture the hierarchical nature of high-level and fine-grained semantics conveyed in images and texts, which is arguably critical to vision-language understanding and reasoning. To this end, we equip both the visual and language branches in CLIP with hierarchy-aware attentions, namely Hierarchy-aware CLIP (HiCLIP), to progressively discover semantic hierarchies layer-by-layer from both images and texts in an unsupervised manner. As a result, such hierarchical aggregation significantly improves the cross-modal alignment. To demonstrate the advantages of HiCLIP, we conduct qualitative analysis on its unsupervised hierarchy induction during inference, as well as extensive quantitative experiments on both visual recognition and vision-language downstream tasks.

Added

2026-09-26

CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data

CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data

Yihan Zeng, Chenhan Jiang, Jiageng Mao, Jianhua Han, Chaoqiang Ye, Qingqiu Huang, Dit-Yan Yeung, Zhen Yang, Xiaodan Liang, Hang Xu

OrganizationsHuaweiSun Yat-sen UniversityThe Chinese University of Hong KongThe Hong Kong University of Science and Technology

Why you should read this

Introduces a cross-modal pretraining framework that aligns real-world 3D point clouds with text and images using automatically generated triplet proxies, enabling strong zero-shot and few-shot 3D recognition without relying on intermediate 2D projection losses.

Contrastive Language-Image Pre-training, benefiting from large-scale unlabeled text-image pairs, has demonstrated great performance in open-world vision understanding tasks. However, due to the limited Text-3D data pairs, adapting the success of 2D Vision-Language Models (VLM) to the 3D space remains an open problem. Existing works that leverage VLM for 3D understanding generally resort to constructing intermediate 2D representations for the 3D data, but at the cost of losing 3D geometry information. To take a step toward open-world 3D vision understanding, we propose Contrastive Language-Image-Point Cloud Pretraining (CLIP²) to directly learn the transferable 3D point cloud representation in realistic scenarios with a novel proxy alignment mechanism. Specifically, we exploit naturally-existed correspondences in 2D and 3D scenarios, and build well-aligned and instance-based text-image-point proxies from those complex scenarios. On top of that, we propose a cross-modal contrastive objective to learn semantic and instance-level aligned point cloud representation. Experimental results on both indoor and outdoor scenarios show that our learned 3D representation has great transfer ability in downstream tasks, including zero-shot and few-shot 3D recognition, which boosts the state-of-the-art methods by large margins. Furthermore, we provide analyses of the capability of different representations in real scenarios and present the optional ensemble scheme.

Added

2026-09-26

RegionCLIP: Region-based Language-Image Pretraining

RegionCLIP: Region-based Language-Image Pretraining

Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, Jianfeng Gao

OrganizationsMicrosoftUniversity of California, Los AngelesUniversity of Wisconsin Madison

Why you should read this

Introduces RegionCLIP, a pretraining framework that bootstraps from CLIP to align local image regions with synthesized textual concepts, significantly advancing open-vocabulary and zero-shot object detection on COCO and LVIS.

Contrastive language-image pretraining (CLIP) using image-text pairs has achieved impressive results on image classification in both zero-shot and transfer learning settings. However, we show that directly applying such models to recognize image regions for object detection leads to unsatisfactory performance due to a major domain shift: CLIP was trained to match an image as a whole to a text description, without capturing the fine-grained alignment between image regions and text spans. To mitigate this issue, we propose a new method called RegionCLIP that significantly extends CLIP to learn region-level visual representations, thus enabling fine-grained alignment between image regions and textual concepts. Our method leverages a CLIP model to match image regions with template captions, and then pretrains our model to align these region-text pairs in the feature space. When transferring our pretrained model to the open-vocabulary object detection task, our method outperforms the state of the art by 3.8 AP50 and 2.2 AP for novel categories on COCO and LVIS datasets, respectively. Further, the learned region representations support zero-shot inference for object detection, showing promising results on both COCO and LVIS datasets. Our code is available at https://github.com/microsoft/RegionCLIP.

Added

2026-09-26

DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations

DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations

Ximeng Sun, Ping Hu, Kate Saenko

OrganizationsBoston UniversityMIT-IBM Watson AI Lab

Why you should read this

Introduces DualCoOp, a lightweight prompt-learning framework for pretrained vision-language models that optimizes pairs of positive and negative context prompts to efficiently adapt CLIP for both partial-label and zero-shot multi-label image recognition.

Solving multi-label recognition (MLR) for images in the low-label regime is a challenging task that has many real-world applications. Recent work learns an alignment between textual and visual spaces to compensate for insufficient image labels, but loses accuracy because of the limited amount of available MLR annotations. In this work, we utilize the strong alignment of textual and visual features pretrained with millions of auxiliary image-text pairs and propose Dual Context Optimization (DualCoOp) as a unified framework for partial-label MLR and zero-shot MLR. DualCoOp encodes positive and negative contexts with class names as part of the linguistic input (i.e. prompts). Since DualCoOp only introduces a very light learnable overhead upon the pretrained vision-language framework, it can quickly adapt to multi-label recognition tasks that have limited annotations and even unseen classes. Experiments on standard multi-label recognition benchmarks across two challenging low-label settings demonstrate the advantages of our approach over state-of-the-art methods. Project page: https://cs-people.bu.edu/sunxm/DualCoOp/project.html

Added

2026-09-26