Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models

Yichao CaoQingfei TangXiu SuSong ChenShan YouXiaobo LuChang Xu

article2023NeurIPS64 citations

Proposes UniHOI, a framework that integrates vision-language foundation models with language-model-generated knowledge via spatial prompt learning to advance open-world and zero-shot human-object interaction detection.

Listen

Modern computer vision systems struggle to reliably detect human-object interactions in real-world environments because traditional methods rely on closed-set categories and expensive manual annotations. While vision-and-language models offer broader semantic capabilities, prior approaches transfer cross-modal knowledge too narrowly. This limits their scalability, adaptability to unseen categories, and capacity to comprehend nuanced human activities.

The article introduces and evaluates UniHOI, a universal framework that integrates vision-language foundation models and large language models into human-object interaction detection. The goal is to accurately recognize both standard and open-world interactive relationships using flexible textual inputs.

The researchers designed an architecture structured around a three-tier visual feature hierarchy: basic visual extraction, instance-level detection, and high-level relation modeling. To link spatial instances with broad semantic knowledge, they introduced a specialized spatial prompt-guided decoder that queries pre-trained foundation models such as BLIP-2. Furthermore, they used a large language model to retrieve descriptive, human-like explanations for complex interaction categories. The framework was evaluated against established industry benchmarks, specifically the HICO-DET and V-COCO datasets, across both fully supervised and zero-shot configurations.

The experimental findings show significant performance gains across all major metrics. First, UniHOI achieved a state-of-the-art 40.06 mean average precision on HICO-DET with a standard ResNet-50 backbone, outperforming the leading baseline by 6.31 points. Second, the model set new performance marks on the V-COCO benchmark, scoring 68.05 and 70.82 across primary role evaluations. Third, under zero-shot testing for unseen objects and unseen verbs, the model exceeded prior state-of-the-art results by margins up to 9.21 points, matching or outperforming earlier fully supervised detectors. Finally, ablation tests confirmed that both the spatial prompt-guided decoder and text-based knowledge retrieval contributed substantial standalone performance gains.

These results demonstrate that grounding vision detectors with broad multimodal foundation models significantly lowers the cost and effort of manual dataset expansion. The ability to recognize interactions from natural language descriptions reduces the operational overhead of retraining models for novel tasks, enabling faster and safer deployment across automated monitoring, robotics, and interactive vision systems.

Organizations developing or deploying vision systems should adopt prompt-guided architectures that integrate multimodal foundation models rather than relying strictly on closed-set visual detectors. Teams should also evaluate incorporating large language models for automated knowledge retrieval to boost zero-shot recognition. Before broad deployment, practitioners must assess computational trade-offs, as rich foundation models like BLIP-2 provide superior accuracy over smaller options like CLIP but require greater memory and processing power during inference.

arXiv: 2311.03799
Cover for Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models

Abstract

Human-object interaction (HOI) detection aims to comprehend the intricate relationships between humans and objects, predicting < human, action, object > triplets, and serving as the foundation for numerous computer vision tasks. The complexity and diversity of human-object interactions in the real world, however, pose significant challenges for both annotation and recognition, particularly in recognizing interactions within an open world context. This study explores the universal interaction recognition in an open-world setting through the use of Vision-Language (VL) foundation models and large language models (LLMs). The proposed method is dubbed as UniHOI. We conduct a deep analysis of the three hierarchical features inherent in visual HOI detectors and propose a method for high-level relation extraction aimed at VL foundation models, which we call HO prompt-based learning. Our design includes an HO Prompt-guided Decoder (HOPD), facilitates the association of high-level relation representations in the foundation model with various HO pairs within the image. Furthermore, we utilize a LLM (i.e. GPT) for interaction interpretation, generating a richer linguistic understanding for complex HOIs. For open-category interaction recognition, our method supports either of two input types: interaction phrase or interpretive sentence. Our efficient architecture design and learning methods effectively unleash the potential of the VL foundation models and LLMs, allowing UniHOI to surpass all existing methods with a substantial margin, under both supervised and zero-shot settings. The code and pre-trained weights are available at: https://github.com/Caoyichao/UniHOI.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Overview
  • 3.2 HO Spatial Prompts Generation
  • 3.3 Prompting Foundation Models for HOI Modeling
  • 3.4 Knowledge Retrieval for HOI Reasoning in Open World
  • 4 Experiments
  • 4.1 Implementation Details
  • 4.2 HOI Detection in the Closed World
  • 4.3 Comparisons with Methods that Utilize Extra Information
  • 4.4 Effectiveness for Zero-Shot HOI Detection
  • 4.5 HOI Detection in the Wild
  • 4.6 Ablation Studies
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — UniHOI Architecture and Three-Tier Feature Hierarchy

    model/method

    The UniHOI framework formulates Human-Object Interaction (HOI) detection through a three-tier visual feature hierarchy to bridge visual detectors, Vision-Language (VL) foundation models, and Large Language Models (LLMs):

    1. Basic Visual Feature Extraction: A visual backbone (e.g., CNN) processes an input image I∈RH×W×CI \in \mathbb{R}^{H \times W \times C} into feature maps, which are projected and augmented with position embeddings to form sequence tokens Xp∈RNv×DvX^p \in \mathbb{R}^{N_v \times D_v}, followed by a Transformer encoder to compute self-attended feature memory M∈RNv×DvM \in \mathbb{R}^{N_v \times D_v}.
    2. Instance-Level Feature Learning: An instance decoder consumes human queries QhQ^h, object queries QoQ^o, and position-guided embeddings QgQ^g alongside memory MM to produce spatial token pairs Pho=[Ph,Po]∈R2Nq×DvP^{ho} = [P^h, P^o] \in \mathbb{R}^{2N_q \times D_v} as well as bounding boxes Bh,BoB^h, B^o and object class predictions CoC^o.
    3. High-Level Relationship Modeling: Spatial prompts derived from the HO pair tokens, (Ph+Po)/2(P^h + P^o)/2, are fed into two parallel decoders:
      • An HO Prompt-guided Decoder (HOPD) that queries high-level multi-modal features XqX^q extracted from a frozen VL foundation model (e.g., BLIP-2 Q-Former), outputting foundation interaction tokens VfV^f.
      • A standard Interaction Decoder that queries the visual memory MM, outputting visual interaction tokens ViV^i.

    The combined representation [Vf;Vi][V^f; V^i] is fed to a feed-forward network (FFN) to predict interaction categories via similarity matching against text embeddings (either category phrases or LLM-generated descriptive texts).

  2. Knowl 2 — Human-Object Spatial Prompts Generation

    model/method

    In the instance-level stage of UniHOI, spatial location prompts for human-object (HO) pairs are generated as follows:

    Given an input image I∈RH×W×CI \in \mathbb{R}^{H \times W \times C}, a CNN backbone generates a 2D feature map Xv∈Rh×w×cX^v \in \mathbb{R}^{h \times w \times c}. A 1×11 \times 1 projection convolution compresses XvX^v, which is then flattened into NvN_v patch embeddings {x1v,x2v,…,xNvv}\{x_1^v, x_2^v, \dots, x_{N_v}^v\}. These are mapped to dimension DvD_v via a linear transformation Ev∈Rc×DvE^v \in \mathbb{R}^{c \times D_v} and added to learnable positional embeddings Eposv∈RNv×DvE_{pos}^v \in \mathbb{R}^{N_v \times D_v}:

    Xp=[x1vEv;x2vEv;… ;xNvvEv]+Eposv∈RNv×DvX^p = [x_1^v E^v; x_2^v E^v; \dots; x_{N_v}^v E^v] + E_{pos}^v \in \mathbb{R}^{N_v \times D_v}

    A Transformer encoder ΦθIE\Phi_{\theta_{IE}} composed of NN layers produces memory representations M=ΦθIE(Xp)∈RNv×DvM = \Phi_{\theta_{IE}}(X^p) \in \mathbb{R}^{N_v \times D_v}.

    An instance decoder ΦθID\Phi_{\theta_{ID}} processes MM using human query set Qh∈RNq×DvQ^h \in \mathbb{R}^{N_q \times D_v}, object query set Qo∈RNq×DvQ^o \in \mathbb{R}^{N_q \times D_v}, and a shared position-guided embedding Qg∈RNq×DvQ^g \in \mathbb{R}^{N_q \times D_v} (which binds human and object queries at the same query index):

    Pho=[Ph,Po]=ΦθID(M,[Qh+Qg,Qo+Qg])∈R2Nq×DvP^{ho} = [P^h, P^o] = \Phi_{\theta_{ID}}(M, [Q^h + Q^g, Q^o + Q^g]) \in \mathbb{R}^{2N_q \times D_v}

    Instance detection feed-forward networks FFNs\text{FFN}_s compute the human bounding boxes Bh∈RNq×4B^h \in \mathbb{R}^{N_q \times 4}, object bounding boxes Bo∈RNq×4B^o \in \mathbb{R}^{N_q \times 4}, and object classification logits Co∈RNq×NcC^o \in \mathbb{R}^{N_q \times N_c}:

    [Bh;Bo;Co]=FFNs([Ph,Po])[B^h; B^o; C^o] = \text{FFN}_s([P^h, P^o])

    The resulting token pairs [Ph,Po][P^h, P^o] provide spatial position cues used to prompt the relationship decoding stage.

  3. Knowl 3 — HO Prompt-Guided Decoder (HOPD) for Foundation Models

    model/method

    To extract high-level relationship representations specific to individual human-object pairs from Vision-Language (VL) foundation models (e.g., BLIP-2), UniHOI employs the HO Prompt-guided Decoder (HOPD), denoted ΦθP\Phi_{\theta_P}:

    1. Foundation Feature Extraction: The input image II is downsampled to I′I' of size W′×H′W' \times H' and processed by the foundation image encoder ΦθI\Phi_{\theta_I} to produce a feature map Xf∈Rw′×h′×c′X^f \in \mathbb{R}^{w' \times h' \times c'}. The Q-Former ΦθQ\Phi_{\theta_Q} with queries QfQ^f extracts high-level visual-semantic tokens:

    Xq=ΦθF(I′)=ΦθI∘ΦθQ(I′,Qf)∈RNf×DfX^q = \Phi_{\theta_F}(I') = \Phi_{\theta_I} \circ \Phi_{\theta_Q}(I', Q^f) \in \mathbb{R}^{N_f \times D_f}

    1. Spatial Prompt Decoding: The pair-level spatial prompts (Ph+Po)/2(P^h + P^o)/2 (where Ph,Po∈RNq×DvP^h, P^o \in \mathbb{R}^{N_q \times D_v} are instance decoder tokens) act as queries in HOPD. These prompts interact with one another via self-attention layers and attend to the frozen foundation tokens XqX^q via cross-attention layers (inserted every other transformer block in HOPD):

    Vf=ΦθP(Ph+Po2,Xq)V^f = \Phi_{\theta_P}\left(\frac{P^h + P^o}{2}, X^q\right)

    1. Visual Interaction Decoding & Fusion: Concurrently, an interaction decoder ΦθIN\Phi_{\theta_{IN}} decodes relation cues directly from the visual detector memory MM:

    Vi=ΦθIN(Ph+Po2,M)V^i = \Phi_{\theta_{IN}}\left(\frac{P^h + P^o}{2}, M\right)

    The foundation relation tokens VfV^f and visual relation tokens ViV^i are concatenated and passed through an FFN to produce the final HOI prediction.

  4. Knowl 4 — Knowledge Retrieval via LLMs (GPTs-as-KBs) for Open-World HOI

    model/method

    UniHOI incorporates Large Language Models (LLMs, e.g., ChatGPT/GPT-4) as implicit knowledge bases (GPTs-as-KBs) to enable open-world, open-vocabulary, and zero-shot HOI recognition beyond fixed word embeddings.

    Instead of relying solely on brief discrete triplet phrase annotations Ti=⟨human,action,object⟩T_i = \langle\text{human}, \text{action}, \text{object}\rangle (e.g., "Human ride bicycle"), UniHOI constructs prompt queries to the LLM formatted as: "Knowledge retrieve for verb_object, limited to N words"\text{"Knowledge retrieve for } \textit{verb\_object}\text{, limited to } N \text{ words"}

    The LLM outputs rich, multi-sentence physical and semantic descriptions KiK_i specifying mechanics, posture, body balance, hand/foot actions, motion trajectories, and context (e.g., describing foot placement, handlebar steering, and pedal propulsion for bicycle riding).

    During zero-shot and open-vocabulary inference, candidate interactions can be supplied as either concise phrases or descriptive paragraphs. The text encoder converts these descriptive texts into semantic embeddings, and the model classifies candidate HO pairs by computing cosine similarity between visual-relational tokens and the descriptive text embeddings.

  5. Knowl 5 — HICO-DET Benchmark Evaluation in Closed-World Setting

    data/table

    UniHOI was evaluated on the HICO-DET benchmark across the Default and Known Objects evaluation settings. Performance is measured using mean Average Precision (mAP, %) across Full (600 HOI categories), Rare (138 categories with <10<10 training instances), and Non-rare (462 categories) splits.

    Default Setting Known Objects Setting
    Method Backbone Full Rare Non-rare Full Rare Non-rare
    Two-stage Methods:
    ATL ResNet-50 23.81 17.43 27.42 27.38 22.09 28.96
    UPT ResNet-50 31.66 25.90 33.36 35.05 29.27 36.77
    UPT ResNet-101 32.31 28.55 33.44 35.65 31.60 36.86
    ViPLO-s ViT-B/32 34.95 33.83 35.28 38.15 36.77 38.56
    ViPLO-l ViT-B/16 37.22 35.45 37.75 40.61 38.82 41.15
    One-stage Methods:
    QPIC ResNet-101 29.90 23.92 31.69 32.38 26.06 34.27
    CDN-S ResNet-50 31.44 27.39 32.64 34.09 29.63 35.42
    CDN-L ResNet-101 32.07 27.19 33.53 34.79 29.48 36.38
    GEN-VLKT-s ResNet-50 33.75 29.25 35.10 36.78 32.75 37.99
    GEN-VLKT-m ResNet-101 34.78 31.50 35.77 38.07 34.94 39.01
    GEN-VLKT-l ResNet-101 34.96 31.18 36.08 38.22 34.36 39.37
    HOICLIP ResNet-101 34.69 31.12 35.74 37.61 34.47 38.54
    Xie et al. (large) ResNet-101 36.03 33.16 36.89 38.82 35.51 39.81
    UniHOI-s (w/ BLIP2) ResNet-50 40.06 39.91 40.11 42.20 42.60 42.08
    UniHOI-m (w/ BLIP2) ResNet-101 40.74 40.03 40.95 42.96 42.86 42.98
    UniHOI-l (w/ BLIP2) ResNet-101 40.95 40.27 41.32 43.26 43.12 43.25

    UniHOI-s outperforms the baseline GEN-VLKT-s by +6.31 mAP on Full and +10.66 mAP on Rare in the Default Setting. UniHOI-l achieves 40.95 mAP (Default Full) and 43.26 mAP (Known Objects Full), establishing superior performance across all categories.

  6. Knowl 6 — V-COCO Benchmark Evaluation in Closed-World Setting

    data/table

    Performance of UniHOI on the V-COCO dataset evaluated using role mean Average Precision under Scenario 1 (AProle#1AP_{role}^{\#1}, which requires predicting empty bounding boxes for non-interacted objects) and Scenario 2 (AProle#2AP_{role}^{\#2}, which ignores empty objects):

    Method AProle#1AP_{role}^{\#1} AProle#2AP_{role}^{\#2}
    VSGNet 51.8 57.0
    IDN 53.3 60.3
    UPT 60.7 66.2
    ViPLO-s 60.9 66.6
    ViPLO-l 62.2 68.0
    HOTR 55.2 64.4
    QPIC 58.8 61.0
    CDN 63.91 65.89
    Liu et al. 63.0 65.2
    GEN-VLKT-s 62.41 64.46
    GEN-VLKT-m 63.28 65.58
    GEN-VLKT-l 63.58 65.93
    HOICLIP 63.50 64.80
    Xie et al. (large) 66.50 69.90
    UniHOI-s (w/ BLIP2) 65.58 (+3.17) 68.27 (+3.81)
    UniHOI-m (w/ BLIP2) 67.95 (+4.67) 70.61 (+5.03)
    UniHOI-l (w/ BLIP2) 68.05 (+4.47) 70.82 (+4.89)

    UniHOI models outperform all prior one-stage and two-stage methods, with UniHOI-l achieving 68.05 AProle#1AP_{role}^{\#1} and 70.82 AProle#2AP_{role}^{\#2}.

  7. Knowl 7 — Zero-Shot HOI Detection Performance on HICO-DET

    data/table

    Zero-shot HOI detection performance was evaluated on HICO-DET across four settings:

    • RF-UC (Rare First Unseen Composition): 120 rare HOI classes held out as unseen.
    • NF-UC (Non-rare First Unseen Composition): 120 non-rare HOI classes held out as unseen.
    • UO (Unseen Object): 12 object categories (100 HOI classes) held out during training.
    • UV (Unseen Verb): 20 verb categories (84 HOI classes) held out during training.
    Method Type Unseen Seen Full
    Shen et al. UC 5.62 - 6.26
    FG UC 10.93 12.60 12.26
    ATL UC 16.99 20.51 19.81
    VCL RF-UC 10.06 24.28 21.43
    ATL RF-UC 9.18 24.67 21.57
    FCL RF-UC 13.16 24.23 22.01
    GEN-VLKT RF-UC 21.36 32.91 30.56
    UniHOI-s (w/ BLIP2) RF-UC 28.68 (+7.32) 33.16 (+0.25) 32.27 (+1.71)
    VCL NF-UC 16.22 18.52 18.06
    ATL NF-UC 18.25 18.78 18.67
    FCL NF-UC 18.66 19.55 19.37
    GEN-VLKT NF-UC 25.05 23.38 23.71
    UniHOI-s (w/ BLIP2) NF-UC 28.45 (+3.40) 32.63 (+9.25) 31.79 (+8.08)
    GEN-VLKT UO 10.51 28.92 25.63
    UniHOI-s (w/ BLIP2) UO 19.72 (+9.21) 34.76 (+5.84) 31.56 (+5.93)
    GEN-VLKT UV 20.96 30.23 28.74
    UniHOI-s (w/ BLIP2) UV 26.05 (+5.09) 36.78 (+6.55) 34.68 (+5.94)

    UniHOI-s improves Unseen mAP by +7.32 on RF-UC, +3.40 on NF-UC, +9.21 on UO, and +5.09 on UV compared to GEN-VLKT.

  8. Knowl 8 — Ablation Analysis of UniHOI Components

    data/table

    Ablation experiments evaluate the sequential addition of core components starting from the baseline (GEN-VLKT-s) on V-COCO (supervised) and HICO-DET Unseen Verb (UV zero-shot):

    V-COCO HICO-DET (UV)
    Configuration AProle#1AP_{role}^{\#1} AProle#2AP_{role}^{\#2} Unseen Seen Full
    Baseline (GEN-VLKT-s) 62.41 64.46 20.96 30.23 28.74
    + VL Foundation Model (Sequential BLIP-2) 62.91 64.83 21.57 31.62 29.85
    + HOPD (HO Prompt-guided Decoder) 65.58 68.27 26.05 36.78 34.68
    + Knowledge Retrieval (GPT descriptions) 66.74 69.31 27.41 37.82 35.89

    Key takeaways:

    1. Merely appending BLIP-2 features via learned queries yields minor gains (+0.50 AProle#1AP_{role}^{\#1} on V-COCO, +0.61 Unseen mAP on HICO-DET UV).
    2. Introducing spatial prompting via HOPD provides a large performance boost (+2.67 AProle#1AP_{role}^{\#1} on V-COCO, +4.48 Unseen mAP on HICO-DET UV), proving that explicit HO spatial prompts are needed to query pair-specific relations from the foundation model.
    3. Incorporating LLM knowledge retrieval further increases Unseen mAP by +1.36 and V-COCO AProle#1AP_{role}^{\#1} by +1.16.
  9. Knowl 9 — Comparison of Foundation Models in UniHOI: BLIP-2 vs CLIP

    data/table

    A comparative analysis of UniHOI using CLIP versus BLIP-2 as the vision-language foundation backbone evaluated on V-COCO across model sizes (small, medium, large):

    Method AProle#1AP_{role}^{\#1} AProle#2AP_{role}^{\#2}
    GEN-VLKT-s 62.41 64.46
    GEN-VLKT-m 63.28 65.58
    GEN-VLKT-l 63.58 65.93
    UniHOI-s (w/ CLIP) 62.92 (+0.51) 65.67 (+1.21)
    UniHOI-m (w/ CLIP) 64.85 (+1.57) 67.62 (+2.04)
    UniHOI-l (w/ CLIP) 64.93 (+1.35) 67.86 (+1.93)
    UniHOI-s (w/ BLIP2) 65.58 (+3.17) 68.27 (+3.81)
    UniHOI-m (w/ BLIP2) 67.95 (+4.67) 70.61 (+5.03)
    UniHOI-l (w/ BLIP2) 68.05 (+4.47) 70.82 (+4.89)

    While CLIP provides moderate performance gains over GEN-VLKT (+0.51 to +1.57 on AProle#1AP_{role}^{\#1}), BLIP-2 delivers significantly larger gains (+3.17 to +4.67 on AProle#1AP_{role}^{\#1}). This difference is attributed to CLIP's global contrastive image-level training yielding lower-dimensional, less detailed relational representations compared to BLIP-2 Q-Former tokens.

Coverage note — None was omitted; all main architectural components, formulations, benchmark evaluations, zero-shot splits, and ablation analyses are included.

References

  1. 1.Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al.: Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198 (2022) 2, 3
  2. 2.Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., Van Den Hengel, A.: Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3674–3683 (2018) 3
  3. 3.Bansal, A., Rambhatla, S.S., Shrivastava, A., Chellappa, R.: Detecting human-object interactions via functional generalization. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 10460–10469 (2020) 8
  4. 4.Cao, Y., Su, X., Tang, Q., You, S., Lu, X., Xu, C.: Searching for better spatio-temporal alignment in few-shot action recognition. Advances in Neural Information Processing Systems 35, 21429–21441 (2022) 3
  5. 5.Cao, Y., Tang, Q., Yang, F., Su, X., You, S., Lu, X., Xu, C.: Re-mine, learn and reason: Exploring the cross-modal semantic correlations for language-guided hoi detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 23492–23503 (2023) 2
  6. 6.Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: European conference on computer vision. pp. 213–229. Springer (2020) 4, 5, 15
  7. 7.Chao, Y.W., Liu, Y., Liu, X., Zeng, H., Deng, J.: Learning to detect human-object interactions. In: 2018 ieee winter conference on applications of computer vision (wacv). pp. 381–389. IEEE (2018) 1, 3, 6, 7, 8, 16
  8. 8.Chen, M., Liao, Y., Liu, S., Chen, Z., Wang, F., Qian, C.: Reformulating hoi detection as adaptive set prediction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9004–9013 (2021) 1, 3, 7
  9. 9.Chen, Z., Li, G., Wan, X.: Align, reason and learn: Enhancing medical vision-and-language pre-training with knowledge. In: Proceedings of the 30th ACM International Conference on Multimedia. pp. 5152–5161 (2022) 3
  10. 10.Cheng, M., Sun, Y., Wang, L., Zhu, X., Yao, K., Chen, J., Song, G., Han, J., Liu, J., Ding, E., Wang, J.: Vista: Vision and scene text aggregation for cross-modal retrieval. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 5184–5193 (June 2022) 3
  11. 11.Dzabraev, M., Kalashnikov, M., Komkov, S., Petiushko, A.: Mdmmt: Multidomain multimodal transformer for video retrieval. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3354–3363 (2021) 3
  12. 12.Gao, C., Xu, J., Zou, Y., Huang, J.B.: Drg: Dual relation graph for human-object interaction detection. In: European Conference on Computer Vision. pp. 696–712. Springer (2020) 1, 3, 7, 8
  13. 13.Gao, C., Zou, Y., Huang, J.B.: ican: Instance-centric attention network for human-object interaction detection. arXiv preprint arXiv:1808.10437 (2018) 3
  14. 14.Gkioxari, G., Girshick, R., Dollár, P., He, K.: Detecting and recognizing human-object interactions. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 8359–8367 (2018) 1, 3
  15. 15.Gupta, S., Malik, J.: Visual semantic role labeling. arXiv preprint arXiv:1505.04474 (2015) 1, 6, 9, 10, 14, 15
  16. 16.Hou, Z., Peng, X., Qiao, Y., Tao, D.: Visual compositional learning for human-object interaction detection. In: European Conference on Computer Vision. pp. 584–600. Springer (2020) 7, 8
  17. 17.Hou, Z., Yu, B., Qiao, Y., Peng, X., Tao, D.: Affordance transfer learning for human-object interaction detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 495–504 (2021) 7, 8
  18. 18.Hou, Z., Yu, B., Qiao, Y., Peng, X., Tao, D.: Detecting human-object interaction via fabricated compositional learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 14646–14655 (2021) 1, 3, 7, 8
  19. 19.Iftekhar, A., Chen, H., Kundu, K., Li, X., Tighe, J., Modolo, D.: What to look at and where: Semantic and spatial refined transformer for detecting human-object interactions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 5353–5363 (2022) 2, 3, 7
  20. 20.Kim, B., Choi, T., Kang, J., Kim, H.J.: Uniondet: Union-level detector towards real-time human-object interaction detection. In: European Conference on Computer Vision. pp. 498–514. Springer (2020) 3
  21. 21.Kim, B., Lee, J., Kang, J., Kim, E.S., Kim, H.J.: Hotr: End-to-end human-object interaction detection with transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 74–83 (2021) 8
  22. 22.Kim, S., Jung, D., Cho, M.: Relational context learning for human-object interaction detection. arXiv preprint arXiv:2304.04997 (2023) 7
  23. 23.Li, J., Li, D., Savarese, S., Hoi, S.: Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597 (2023) 5, 6, 10, 13
  24. 24.Li, L.H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.N., et al.: Grounded language-image pre-training. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10965–10975 (2022) 3
  25. 25.Li, Y.L., Liu, X., Lu, H., Wang, S., Liu, J., Li, J., Lu, C.: Detailed 2d-3d joint representation for human-object interaction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10166–10175 (2020) 7
  26. 26.Li, Y.L., Liu, X., Wu, X., Li, Y., Lu, C.: Hoi analysis: Integrating and decomposing human-object interaction. Advances in Neural Information Processing Systems 33, 5011–5022 (2020) 7, 8
  27. 27.Li, Y.L., Zhou, S., Huang, X., Xu, L., Ma, Z., Fang, H.S., Wang, Y., Lu, C.: Transferable interactiveness knowledge for human-object interaction detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3585–3594 (2019) 1, 3, 16
  28. 28.Li, Z., Zou, C., Zhao, Y., Li, B., Zhong, S.: Improving human-object interaction detection via phrase learning and label composition. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 36, pp. 1509–1517 (2022) 2, 3, 8, 13, 14
  29. 29.Liao, Y., Liu, S., Wang, F., Chen, Y., Qian, C., Feng, J.: Ppdm: Parallel point detection and matching for real-time human-object interaction detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 482–490 (2020) 1, 3, 7
  30. 30.Liao, Y., Zhang, A., Lu, M., Wang, Y., Li, X., Liu, S.: Gen-vlkt: Simplify association and enhance interaction understanding for hoi detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 20123–20132 (2022) 3, 5, 6, 7, 8, 10, 13, 14, 15
  31. 31.Liu, X., Li, Y.L., Wu, X., Tai, Y.W., Lu, C., Tang, C.K.: Interactiveness field in human-object interactions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 20113–20122 (2022) 3, 7, 8
  32. 32.Liu, Y., Chen, Q., Zisserman, A.: Amplifying key cues for human-object-interaction detection. In: European Conference on Computer Vision. pp. 248–265. Springer (2020) 7, 8
  33. 33.Liu, Y., Yuan, J., Chen, C.W.: Consnet: Learning consistency graph for zero-shot human-object interaction detection. In: Proceedings of the 28th ACM International Conference on Multimedia. pp. 4235–4243 (2020) 8
  34. 34.Mu, N., Kirillov, A., Wagner, D., Xie, S.: Slip: Self-supervision meets language-image pre-training. arXiv preprint arXiv:2112.12750 (2021) 3
  35. 35.Ning, S., Qiu, L., Liu, Y., He, X.: Hoiclip: Efficient knowledge transfer for hoi detection with vision-language models. arXiv preprint arXiv:2303.15786 (2023) 6, 7, 8, 14
  36. 36.OpenAI: Gpt-4 technical report (2023) 3, 9
  37. 37.Park, J., Park, J.W., Lee, J.S.: Viplo: Vision transformer based pose-conditioned self-loop graph for human-object interaction detection. arXiv preprint arXiv:2304.08114 (2023) 6, 7, 8
  38. 38.Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning. pp. 8748–8763. PMLR (2021) 3, 10, 13
  39. 39.Shen, L., Yeung, S., Hoffman, J., Mori, G., Fei-Fei, L.: Scaling human-object interaction recognition through zero-shot learning. In: 2018 IEEE Winter Conference on Applications of Computer Vision (WACV). pp. 1568–1576. IEEE (2018) 8
  40. 40.Tamura, M., Ohashi, H., Yoshinaga, T.: Qpic: Query-based pairwise human-object interaction detection with image-wide contextual information. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10410–10419 (2021) 7, 8, 15
  41. 41.Tsimpoukelli, M., Menick, J.L., Cabi, S., Eslami, S., Vinyals, O., Hill, F.: Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems 34, 200–212 (2021) 3
  42. 42.Ulutan, O., Iftekhar, A., Manjunath, B.S.: Vsgnet: Spatial attention network for detecting human object interactions using graph convolutions. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 13617–13626 (2020) 7, 8
  43. 43.Wang, S., Duan, Y., Ding, H., Tan, Y.P., Yap, K.H., Yuan, J.: Learning transferable human-object interaction detector with natural language supervision. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 939–948 (2022) 2, 3
  44. 44.Xie, C., Zeng, F., Hu, Y., Liang, S., Wei, Y.: Category query learning for human-object interaction classification. arXiv preprint arXiv:2303.14005 (2023) 7, 8
  45. 45.Yuan, H., Jiang, J., Albanie, S., Feng, T., Huang, Z., Ni, D., Tang, M.: Rlip: Relational language-image pre-training for human-object interaction detection. arXiv preprint arXiv:2209.01814 (2022) 1, 2, 3, 7, 8
  46. 46.Yuan, H., Wang, M., Ni, D., Xu, L.: Detecting human-object interactions with object-guided cross-modal calibrated semantics. arXiv preprint arXiv:2202.00259 (2022) 3, 8
  47. 47.Zhang, A., Liao, Y., Liu, S., Lu, M., Wang, Y., Gao, C., Li, X.: Mining the benefits of two-stage and one-stage hoi detection. Advances in Neural Information Processing Systems 34, 17209–17220 (2021) 3, 7, 8
  48. 48.Zhang, F.Z., Campbell, D., Gould, S.: Spatially conditioned graphs for detecting human-object interactions. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 13319–13327 (2021) 7
  49. 49.Zhang, F.Z., Campbell, D., Gould, S.: Efficient two-stage detection of human-object interactions with a novel unary-pairwise transformer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 20104–20112 (2022) 1, 3, 7, 8
  50. 50.Zhang, H., Zhang, P., Hu, X., Chen, Y.C., Li, L.H., Dai, X., Wang, L., Yuan, L., Hwang, J.N., Gao, J.: Glipv2: Unifying localization and vision-language understanding. arXiv preprint arXiv:2206.05836 (2022) 3
  51. 51.Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X.V., et al.: Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068 (2022) 13
  52. 52.Zhong, X., Ding, C., Qu, X., Tao, D.: Polysemy deciphering network for robust human–object interaction detection. International Journal of Computer Vision 129(6), 1910–1929 (2021) 2, 3, 8
  53. 53.Zhong, X., Qu, X., Ding, C., Tao, D.: Glance and gaze: Inferring action-aware points for one-stage human-object interaction detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 13234–13243 (2021) 3
  54. 54.Zou, C., Wang, B., Hu, Y., Liu, J., Wu, Q., Zhao, Y., Li, B., Zhang, C., Zhang, C., Wei, Y., et al.: End-to-end human object interaction detection with hoi transformer. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 11825–11834 (2021) 7

Citation

MLA
Cao, Y., et al. “Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 739–51, https://proceedings.neurips.cc/paper_files/paper/2023/file/02687e7b22abc64e651be8da74ec610e-Paper-Conference.pdf.
APA
Cao, Y., Tang, Q., Su, X., Chen, S., You, S., Lu, X., & Xu, C. (2023). Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models. Advances in Neural Information Processing Systems, 36, 739–751. https://proceedings.neurips.cc/paper_files/paper/2023/file/02687e7b22abc64e651be8da74ec610e-Paper-Conference.pdf
Chicago
Cao, Y., Q. Tang, X. Su, et al. 2023. “Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models”. Advances in Neural Information Processing Systems 36: 739–51. https://proceedings.neurips.cc/paper_files/paper/2023/file/02687e7b22abc64e651be8da74ec610e-Paper-Conference.pdf.
Harvard
Cao, Y. et al. (2023) “Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 739–751. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/02687e7b22abc64e651be8da74ec610e-Paper-Conference.pdf.
Vancouver
1. Cao Y, Tang Q, Su X, Chen S, You S, Lu X, Xu C (2023) Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 739–751

BibTeX

@inproceedings{cao2023detecting,
  title = {Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models},
  author = {Cao, Yichao and Tang, Qingfei and Su, Xiu and Chen, Song and You, Shan and Lu, Xiaobo and Xu, Chang},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {739-751},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/02687e7b22abc64e651be8da74ec610e-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors