Learning Transferable Human-Object Interaction Detector with Natural Language Supervision
Suchen WangYueqi DuanHenghui DingYap-Peng TanKim-Hui YapJunsong Yuan
Proposes a vision-language framework for zero-shot human-object interaction detection that aligns Vision Transformer-extracted interaction tokens with learned text prompts to recognize unseen action-object combinations without predefining interaction classes.
Human-object interaction detection is crucial for computer vision applications that analyze human activities, intentions, and behaviors. However, real-world interactions consist of thousands of combinations between actions and objects, making it impractical to capture and label every possible combination in training datasets. Conventional detectors rely on discrete classifiers with predefined category lists, which severely restricts their ability to identify new, unobserved combinations without explicit prior knowledge.
The main objective of the article is to demonstrate a transferable, one-stage human-object interaction detector that can reliably identify both seen and unseen interactions without relying on predetermined interaction lists.
To achieve this, the authors develop a detector named THID (Transferable Human-object Interaction Detector) that reframes detection into an instance-level visual-to-text matching problem. The architecture integrates a Vision Transformer visual encoder with a frozen, pretrained text encoder from the Contrastive Language-Image Pretraining (CLIP) foundation model. The visual network incorporates specialized learnable tokens and an autoregressive sequence parser to distinguish multiple distinct interactions in an image. Instead of rigid manual templates, learnable text prompts automatically construct interaction descriptions, and a tailored instance-level loss function aligns visual features with text embeddings while avoiding feature distortion. The method was evaluated on two widely used benchmark datasets: HICO-DET, under a simulated zero-shot split of 600 categories, and the large-vocabulary SWIG-HOI dataset containing 400 actions and 1,000 objects.
Key findings show that THID achieves state-of-the-art performance, especially for unseen interactions. On the SWIG-HOI dataset, THID improved mean Average Precision by 3.83 points (a relative improvement of over 60%) on unseen categories and by 2.14 points overall compared to previous leading methods. On HICO-DET, it increased unseen interaction performance by 2.37 mean Average Precision points while maintaining competitive accuracy on previously seen interactions. Ablation studies demonstrated that the sequential token parser and learnable language prompts contributed significantly to performance, raising unseen accuracy on SWIG-HOI from 6.27 to 10.04 mean Average Precision points compared to baseline configurations.
These findings indicate that language-supervised vision modeling enables artificial intelligence systems to recognize open-vocabulary interactions flexibly. This capability reduces the high costs and operational risks of manually annotating rare interaction data in deployment pipelines, allowing perception models to generalize to unpredictable real-world environments more robustly than rigid, discrete classification systems.
Future technical efforts should focus on integrating multi-scale image processing and adaptive resolution mechanisms into the detector. Organizations seeking to deploy open-vocabulary vision systems should consider piloting transferable vision-language architectures, balancing the benefits of high recognition accuracy against computational localization trade-offs.
The primary limitation noted in the article is the reliance on a fixed visual input resolution of 224 by 224 pixels, which leads to lower precision when localizing bounding boxes for small objects compared to specialized multi-scale object detectors. While confidence in the model's semantic recognition and zero-shot transfer capabilities is high based on cross-benchmark validation, stakeholders should exercise caution in applications that demand ultra-precise physical localization of small items.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. Introduces the CLIP foundation model and natural language supervision paradigm that THID integrates and adapts for visual-to-text interaction matching.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Establishes the Vision Transformer architecture that serves as the visual encoder backbone in THID's human-object interaction framework.
- Paper: End-to-End Object Detection with Transformers, Nicolas Carion et al. (2020). Pioneers end-to-end set prediction and learned query parsing in vision transformers, laying the conceptual groundwork for THID's token parser.
- Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). Demonstrates open-vocabulary detection using vision-and-language knowledge distillation from foundation models, directly informing THID's zero-shot interaction strategy.
- Paper: Scene Graph Generation by Iterative Message Passing, Danfei Xu et al. (2017). Provides foundational concepts for structured relationship modeling and message passing across human and object entities in visual scenes.
- Paper: CLIP-Adapter: Better Vision-Language Models with Feature Adapters, Peng Gao et al. (2021). Introduces lightweight adaptation and prompt mechanisms for CLIP representations, which THID builds on through learnable text prompt generation.
- Paper: Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models, Yichao Cao et al. (2023). Extends open-vocabulary HOI detection beyond THID by incorporating spatial prompt learning and multimodal large language model queries into a universal HOI framework.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). Provides a comprehensive survey contextualizing how vision-language models and transfer techniques, such as those introduced in THID, generalize across diverse visual tasks.
- Paper: Hierarchical Open-vocabulary Universal Image Segmentation, Xudong Wang et al. (2023). Generalizes open-vocabulary visual relation modeling into hierarchical universal image segmentation and multi-granularity entity reasoning.
- Paper: Unified Human-Scene Interaction via Prompted Chain-of-Contacts, Zeqi Xiao et al. (2024). Applies language-driven human-object interaction understanding to embodied 3D physics simulations and chain-of-contact generation.
