Built independently by an author, for readers. Read the story and support ChapterPal

keyword

natural language supervision

Natural language supervision is a machine learning training paradigm where models learn visual or perceptual representations directly from free-form text, descriptions, or captions rather than from fixed, predefined categorical labels. By mapping paired multimodal data, such as images and corresponding text, into a shared semantic embedding space, this approach enables systems to capture nuanced relationships across a broad conceptual vocabulary. Because the supervisory signal relies on expressive language instead of closed-set category indices, models trained with natural language supervision can generalize to unseen classes and perform zero-shot or open-vocabulary recognition across various downstream tasks without requiring task-specific retraining or manual class annotations.

3 items

Learning Transferable Human-Object Interaction Detector with Natural Language Supervision

Learning Transferable Human-Object Interaction Detector with Natural Language Supervision

Suchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan, Kim-Hui Yap, Junsong Yuan

OrganizationsNanyang Technological UniversityTsinghua UniversityUniversity at Buffalo

Why you should read this

Proposes a vision-language framework for zero-shot human-object interaction detection that aligns Vision Transformer-extracted interaction tokens with learned text prompts to recognize unseen action-object combinations without predefining interaction classes.

It is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors often treat interactions as discrete labels and learn a classifier according to a predetermined category space. This is inherently inapt for detecting unseen interactions which are out of the predefined categories. Conversely, we treat independent HOI labels as the natural language supervision of interactions and embed them into a joint visual-and-text space to capture their correlations. More specifically, we propose a new HOI visual encoder to detect the interacting humans and objects, and map them to a joint feature space to perform interaction recognition. Our visual encoder is instantiated as a Vision Transformer with new learnable HOI tokens and a sequence parser to generate unique HOI predictions. It distills and leverages the transferable knowledge from the pretrained CLIP model to perform the zero-shot interaction detection. Experiments on two datasets, SWIG-HOI and HICO-DET, validate that our proposed method can achieve a notable mAP improvement on detecting both seen and unseen HOIs. Our code is available at https://github.com/scwangdyd/promting_hoi.

Added

2026-09-26

Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision

Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision

Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng, Yi Wang, Yu Qiao, Weidi Xie

OrganizationsFudan UniversityShanghai Artificial Intelligence LaboratoryShanghai Jiao Tong University

Why you should read this

Presents OVSegmentor, a vision-language framework that learns open-vocabulary semantic segmentation directly from web-scale image-caption pairs without manual mask annotations by using slot-attention group tokens and proxy tasks for masked entity completion and cross-image consistency.

This paper considers the problem of open-vocabulary semantic segmentation (OVS), that aims to segment objects of arbitrary classes beyond a pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor, which only exploits web-crawled image-text pairs for pre-training without using any mask annotations. OVSegmentor assembles the image pixels into a set of learnable group tokens via a slot-attention based binding module, then aligns the group tokens to corresponding caption embeddings. Second, we propose two proxy tasks for training, namely masked entity completion and cross-image mask consistency. The former aims to infer all masked entities in the caption given group tokens, that enables the model to learn fine-grained alignment between visual groups and text entities. The latter enforces consistent mask predictions between images that contain shared entities, encouraging the model to learn visual invariance. Third, we construct CC4M dataset for pre-training by filtering CC12M with frequently appeared entities, which significantly improves training efficiency. Fourth, we perform zero-shot transfer on four benchmark datasets, PASCAL VOC, PASCAL Context, COCO Object, and ADE20K. OVSegmentor achieves superior results over state-of-the-art approaches on PASCAL VOC using only 3% data (4M vs 134M) for pre-training.

Added

2026-09-26

Learning Transferable Visual Models From Natural Language Supervision

Learning Transferable Visual Models From Natural Language Supervision

Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, Ilya Sutskever

OrganizationsOpenAI

Why you should read this

Proposes a contrastive pre-training strategy on 400 million image-text pairs that enables zero-shot visual classification across diverse benchmarks, matching fully supervised ResNet-50 performance on ImageNet without task-specific training data.

State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is needed to specify any other visual concept. Learning directly from raw text about images is a promising alternative which leverages a much broader source of supervision. We demonstrate that the simple pre-training task of predicting which caption goes with which image is an efficient and scalable way to learn SOTA image representations from scratch on a dataset of 400 million (image, text) pairs collected from the internet. After pre-training, natural language is used to reference learned visual concepts (or describe new ones) enabling zero-shot transfer of the model to downstream tasks. We study the performance of this approach by benchmarking on over 30 different existing computer vision datasets, spanning tasks such as OCR, action recognition in videos, geo-localization, and many types of fine-grained object classification. The model transfers non-trivially to most tasks and is often competitive with a fully supervised baseline without the need for any dataset specific training. For instance, we match the accuracy of the original ResNet-50 on ImageNet zero-shot without needing to use any of the 1.28 million training examples it was trained on. We release our code and pre-trained model weights at this https URL.

Added

2026-08-12

License

Published with permission