Built independently by an author, for readers. Read the story and support ChapterPal

keyword

compositional learning

Compositional learning is a machine learning paradigm where a system learns to recognize and process complex concepts by breaking them down into basic constituent components and learning how those components combine. Instead of treating every unique configuration or interaction as an isolated class, the model learns separate representations for primitive elements, such as objects, actions, or visual attributes. This structural decomposition enables the system to generalize across combinatorial variations and recognize novel, unseen combinations of familiar components during inference. Consequently, compositional learning significantly improves sample efficiency and facilitates zero-shot generalization in computer vision, natural language processing, and multimodal tasks where collecting exhaustive training data for every possible combination is impractical.

1 item

Learning Transferable Human-Object Interaction Detector with Natural Language Supervision

Learning Transferable Human-Object Interaction Detector with Natural Language Supervision

Suchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan, Kim-Hui Yap, Junsong Yuan

OrganizationsNanyang Technological UniversityTsinghua UniversityUniversity at Buffalo

Why you should read this

Proposes a vision-language framework for zero-shot human-object interaction detection that aligns Vision Transformer-extracted interaction tokens with learned text prompts to recognize unseen action-object combinations without predefining interaction classes.

It is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors often treat interactions as discrete labels and learn a classifier according to a predetermined category space. This is inherently inapt for detecting unseen interactions which are out of the predefined categories. Conversely, we treat independent HOI labels as the natural language supervision of interactions and embed them into a joint visual-and-text space to capture their correlations. More specifically, we propose a new HOI visual encoder to detect the interacting humans and objects, and map them to a joint feature space to perform interaction recognition. Our visual encoder is instantiated as a Vision Transformer with new learnable HOI tokens and a sequence parser to generate unique HOI predictions. It distills and leverages the transferable knowledge from the pretrained CLIP model to perform the zero-shot interaction detection. Experiments on two datasets, SWIG-HOI and HICO-DET, validate that our proposed method can achieve a notable mAP improvement on detecting both seen and unseen HOIs. Our code is available at https://github.com/scwangdyd/promting_hoi.

Added

2026-09-26