Built independently by an author, for readers. Read the story and support ChapterPal

keyword

human-object interaction detection

Human-object interaction detection is a computer vision task that involves identifying the spatial locations of people and objects in an image and recognizing the specific actions or relational behaviors occurring between them. The objective is to comprehensively understand visual scenes by detecting interacting human-object pairs and classifying their relationships, typically represented as structured triplets consisting of a human, an action verb, and an object entity. Beyond standard object detection, the task requires relational and contextual reasoning to determine whether an interaction exists and to correctly categorize complex interactions across diverse open-world and closed-world visual contexts, supporting broader applications in scene understanding, robotics, and vision-language interpretation.

4 items

Learning Transferable Human-Object Interaction Detector with Natural Language Supervision

Learning Transferable Human-Object Interaction Detector with Natural Language Supervision

Suchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan, Kim-Hui Yap, Junsong Yuan

OrganizationsNanyang Technological UniversityTsinghua UniversityUniversity at Buffalo

Why you should read this

Proposes a vision-language framework for zero-shot human-object interaction detection that aligns Vision Transformer-extracted interaction tokens with learned text prompts to recognize unseen action-object combinations without predefining interaction classes.

It is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors often treat interactions as discrete labels and learn a classifier according to a predetermined category space. This is inherently inapt for detecting unseen interactions which are out of the predefined categories. Conversely, we treat independent HOI labels as the natural language supervision of interactions and embed them into a joint visual-and-text space to capture their correlations. More specifically, we propose a new HOI visual encoder to detect the interacting humans and objects, and map them to a joint feature space to perform interaction recognition. Our visual encoder is instantiated as a Vision Transformer with new learnable HOI tokens and a sequence parser to generate unique HOI predictions. It distills and leverages the transferable knowledge from the pretrained CLIP model to perform the zero-shot interaction detection. Experiments on two datasets, SWIG-HOI and HICO-DET, validate that our proposed method can achieve a notable mAP improvement on detecting both seen and unseen HOIs. Our code is available at https://github.com/scwangdyd/promting_hoi.

Added

2026-09-26

Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models

Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models

Yichao Cao, Qingfei Tang, Xiu Su, Song Chen, Shan You, Xiaobo Lu, Chang Xu

OrganizationsNanjing Enbo Tech.SenseTimeSoutheast UniversityUniversity of Sydney

Why you should read this

Proposes UniHOI, a framework that integrates vision-language foundation models with language-model-generated knowledge via spatial prompt learning to advance open-world and zero-shot human-object interaction detection.

Human-object interaction (HOI) detection aims to comprehend the intricate relationships between humans and objects, predicting < human, action, object > triplets, and serving as the foundation for numerous computer vision tasks. The complexity and diversity of human-object interactions in the real world, however, pose significant challenges for both annotation and recognition, particularly in recognizing interactions within an open world context. This study explores the universal interaction recognition in an open-world setting through the use of Vision-Language (VL) foundation models and large language models (LLMs). The proposed method is dubbed as UniHOI. We conduct a deep analysis of the three hierarchical features inherent in visual HOI detectors and propose a method for high-level relation extraction aimed at VL foundation models, which we call HO prompt-based learning. Our design includes an HO Prompt-guided Decoder (HOPD), facilitates the association of high-level relation representations in the foundation model with various HO pairs within the image. Furthermore, we utilize a LLM (i.e. GPT) for interaction interpretation, generating a richer linguistic understanding for complex HOIs. For open-category interaction recognition, our method supports either of two input types: interaction phrase or interpretive sentence. Our efficient architecture design and learning methods effectively unleash the potential of the VL foundation models and LLMs, allowing UniHOI to surpass all existing methods with a substantial margin, under both supervised and zero-shot settings. The code and pre-trained weights are available at: https://github.com/Caoyichao/UniHOI.

Added

2026-09-26

End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge Distillation

End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge Distillation

Mingrui Wu, Jiaxin Gu, Yunhang Shen, Mingbao Lin, Chao Chen, Xiaoshuai Sun

Why you should read this

Proposes an end-to-end zero-shot human-object interaction detection framework that distills region-level vision-language knowledge from CLIP and uses a two-stage bipartite matching algorithm to detect unseen interactions effectively.

Most existing Human-Object Interaction (HOI) Detection methods rely heavily on full annotations with predefined HOI categories, which is limited in diversity and costly to scale further. We aim at advancing zero-shot HOI detection to detect both seen and unseen HOIs simultaneously. The fundamental challenges are to discover potential human-object pairs and identify novel HOI categories. To overcome the above challenges, we propose a novel End-to-end zero-shot HOI Detection (EoID) framework via vision-language knowledge distillation. We first design an Interactive Score module combined with a Two-stage Bipartite Matching algorithm to achieve interaction distinguishment for human-object pairs in an action-agnostic manner. Then we transfer the distribution of action probability from the pretrained vision-language teacher as well as the seen ground truth to the HOI model to attain zero-shot HOI classification. Extensive experiments on HICO-Det dataset demonstrate that our model discovers potential interactive pairs and enables the recognition of unseen HOIs. Finally, our EoID outperforms the previous SOTAs under various zero-shot settings. Moreover, our method is generalizable to large-scale object detection data to further scale up the action sets. The source code is available at: https://github.com/mrwu-mac/EoID.

Added

2026-09-26