Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models
Yichao CaoQingfei TangXiu SuSong ChenShan YouXiaobo LuChang Xu
Proposes UniHOI, a framework that integrates vision-language foundation models with language-model-generated knowledge via spatial prompt learning to advance open-world and zero-shot human-object interaction detection.
Modern computer vision systems struggle to reliably detect human-object interactions in real-world environments because traditional methods rely on closed-set categories and expensive manual annotations. While vision-and-language models offer broader semantic capabilities, prior approaches transfer cross-modal knowledge too narrowly. This limits their scalability, adaptability to unseen categories, and capacity to comprehend nuanced human activities.
The article introduces and evaluates UniHOI, a universal framework that integrates vision-language foundation models and large language models into human-object interaction detection. The goal is to accurately recognize both standard and open-world interactive relationships using flexible textual inputs.
The researchers designed an architecture structured around a three-tier visual feature hierarchy: basic visual extraction, instance-level detection, and high-level relation modeling. To link spatial instances with broad semantic knowledge, they introduced a specialized spatial prompt-guided decoder that queries pre-trained foundation models such as BLIP-2. Furthermore, they used a large language model to retrieve descriptive, human-like explanations for complex interaction categories. The framework was evaluated against established industry benchmarks, specifically the HICO-DET and V-COCO datasets, across both fully supervised and zero-shot configurations.
The experimental findings show significant performance gains across all major metrics. First, UniHOI achieved a state-of-the-art 40.06 mean average precision on HICO-DET with a standard ResNet-50 backbone, outperforming the leading baseline by 6.31 points. Second, the model set new performance marks on the V-COCO benchmark, scoring 68.05 and 70.82 across primary role evaluations. Third, under zero-shot testing for unseen objects and unseen verbs, the model exceeded prior state-of-the-art results by margins up to 9.21 points, matching or outperforming earlier fully supervised detectors. Finally, ablation tests confirmed that both the spatial prompt-guided decoder and text-based knowledge retrieval contributed substantial standalone performance gains.
These results demonstrate that grounding vision detectors with broad multimodal foundation models significantly lowers the cost and effort of manual dataset expansion. The ability to recognize interactions from natural language descriptions reduces the operational overhead of retraining models for novel tasks, enabling faster and safer deployment across automated monitoring, robotics, and interactive vision systems.
Organizations developing or deploying vision systems should adopt prompt-guided architectures that integrate multimodal foundation models rather than relying strictly on closed-set visual detectors. Teams should also evaluate incorporating large language models for automated knowledge retrieval to boost zero-shot recognition. Before broad deployment, practitioners must assess computational trade-offs, as rich foundation models like BLIP-2 provide superior accuracy over smaller options like CLIP but require greater memory and processing power during inference.
- Paper: Learning Transferable Human-Object Interaction Detector with Natural Language Supervision, Suchen Wang et al. (2022). This work directly establishes the paradigm of using natural language supervision and joint visual-text embeddings to detect unseen human-object interactions, serving as a foundational prerequisite for UniHOI's open-world formulation.
- Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). ViLD pioneered the distillation of pre-trained vision-language foundation models into region-level detectors for open-vocabulary visual recognition, which UniHOI adapts to interaction detection.
- Paper: Scene Graph Generation by Iterative Message Passing, Danfei Xu et al. (2017). It introduces the core relational modeling concepts for grounding visual entities and predicting structured pairwise relationships from contextual image features.
- Paper: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Junnan Li et al. (2022). UniHOI directly employs BLIP-2 architectures and representations as its querying foundation model, making BLIP's multimodal vision-language framework an essential architectural antecedent.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). It provides foundational methodology for querying large language models to extract structured semantic and commonsense knowledge for understanding physical human actions.
- Paper: DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations, Ximeng Sun et al. (2022). It demonstrates how to leverage prompt learning and spatial region aggregation on vision-language backbones for multi-label recognition under limited annotation.
- Paper: Unified Human-Scene Interaction via Prompted Chain-of-Contacts, Zeqi Xiao et al. (2024). UniHSI builds upon open-vocabulary human-object interaction reasoning by using language models to translate text instructions into physically grounded chains of human-scene contact sequences.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). This work extends visual-textual affordance and interaction representations into actionable physical robotic execution via vision-language-action policies.
- Paper: UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent, Jianke Zhang et al. (2025). UP-VLA applies the principles of multimodal spatial-relational understanding directly to future visual prediction and generalist robotic manipulation.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). CoT-VLA extends multi-modal interaction and spatial reasoning by generating explicit visual subgoals for sequential action planning.
