Learning Transferable Human-Object Interaction Detector with Natural Language Supervision

Suchen WangYueqi DuanHenghui DingYap-Peng TanKim-Hui YapJunsong Yuan

article2022CVPR87 citations

Proposes a vision-language framework for zero-shot human-object interaction detection that aligns Vision Transformer-extracted interaction tokens with learned text prompts to recognize unseen action-object combinations without predefining interaction classes.

Listen

Human-object interaction detection is crucial for computer vision applications that analyze human activities, intentions, and behaviors. However, real-world interactions consist of thousands of combinations between actions and objects, making it impractical to capture and label every possible combination in training datasets. Conventional detectors rely on discrete classifiers with predefined category lists, which severely restricts their ability to identify new, unobserved combinations without explicit prior knowledge.

The main objective of the article is to demonstrate a transferable, one-stage human-object interaction detector that can reliably identify both seen and unseen interactions without relying on predetermined interaction lists.

To achieve this, the authors develop a detector named THID (Transferable Human-object Interaction Detector) that reframes detection into an instance-level visual-to-text matching problem. The architecture integrates a Vision Transformer visual encoder with a frozen, pretrained text encoder from the Contrastive Language-Image Pretraining (CLIP) foundation model. The visual network incorporates specialized learnable tokens and an autoregressive sequence parser to distinguish multiple distinct interactions in an image. Instead of rigid manual templates, learnable text prompts automatically construct interaction descriptions, and a tailored instance-level loss function aligns visual features with text embeddings while avoiding feature distortion. The method was evaluated on two widely used benchmark datasets: HICO-DET, under a simulated zero-shot split of 600 categories, and the large-vocabulary SWIG-HOI dataset containing 400 actions and 1,000 objects.

Key findings show that THID achieves state-of-the-art performance, especially for unseen interactions. On the SWIG-HOI dataset, THID improved mean Average Precision by 3.83 points (a relative improvement of over 60%) on unseen categories and by 2.14 points overall compared to previous leading methods. On HICO-DET, it increased unseen interaction performance by 2.37 mean Average Precision points while maintaining competitive accuracy on previously seen interactions. Ablation studies demonstrated that the sequential token parser and learnable language prompts contributed significantly to performance, raising unseen accuracy on SWIG-HOI from 6.27 to 10.04 mean Average Precision points compared to baseline configurations.

These findings indicate that language-supervised vision modeling enables artificial intelligence systems to recognize open-vocabulary interactions flexibly. This capability reduces the high costs and operational risks of manually annotating rare interaction data in deployment pipelines, allowing perception models to generalize to unpredictable real-world environments more robustly than rigid, discrete classification systems.

Future technical efforts should focus on integrating multi-scale image processing and adaptive resolution mechanisms into the detector. Organizations seeking to deploy open-vocabulary vision systems should consider piloting transferable vision-language architectures, balancing the benefits of high recognition accuracy against computational localization trade-offs.

The primary limitation noted in the article is the reliance on a fixed visual input resolution of 224 by 224 pixels, which leads to lower precision when localizing bounding boxes for small objects compared to specialized multi-scale object detectors. While confidence in the model's semantic recognition and zero-shot transfer capabilities is high based on cross-benchmark validation, stakeholders should exercise caution in applications that demand ultra-precise physical localization of small items.

Cover for Learning Transferable Human-Object Interaction Detector with Natural Language Supervision

Abstract

It is difficult to construct a data collection including all possible combinations of human actions and interacting objects due to the combinatorial nature of human-object interactions (HOI). In this work, we aim to develop a transferable HOI detector for unseen interactions. Existing HOI detectors often treat interactions as discrete labels and learn a classifier according to a predetermined category space. This is inherently inapt for detecting unseen interactions which are out of the predefined categories. Conversely, we treat independent HOI labels as the natural language supervision of interactions and embed them into a joint visual-and-text space to capture their correlations. More specifically, we propose a new HOI visual encoder to detect the interacting humans and objects, and map them to a joint feature space to perform interaction recognition. Our visual encoder is instantiated as a Vision Transformer with new learnable HOI tokens and a sequence parser to generate unique HOI predictions. It distills and leverages the transferable knowledge from the pretrained CLIP model to perform the zero-shot interaction detection. Experiments on two datasets, SWIG-HOI and HICO-DET, validate that our proposed method can achieve a notable mAP improvement on detecting both seen and unseen HOIs. Our code is available at https://github.com/scwangdyd/promting_hoi.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Preliminary: Visual-and-Text Modeling
  • 3.2. The Proposed Method
  • 3.2.1 HOI Visual Encoder
  • 3.2.2 Text Encoding
  • 3.2.3 Training
  • 4. Experiments
  • 4.1. Ablation Studies
  • 4.1.1 Vision Transformer with HOI tokens
  • 4.1.2 Sequential HOI Parser
  • 4.1.3 Text Description of Interactions
  • 4.1.4 Unsymmetrical classification loss
  • 4.2. Comparison on HICO-DET dataset
  • 4.3. Comparison on SWIG-HOI dataset
  • 4.4. Limitations
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Transferable Human-Object Interaction Detector Architecture

    model/method

    The Transferable Human-Object Interaction Detector (THID) is a one-stage architecture that frames human-object interaction (HOI) detection as visual-to-text semantic matching, leveraging pretrained vision-language representations from CLIP without requiring predefined unseen interaction lists.

    Given an RGB image I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3}, THID processes the visual and textual modalities as follows:

    1. HOI Visual Encoder: Instantiated as a Vision Transformer (ViT-B/16) backbone augmented with MM learnable interaction query tokens H0=[h10;… ;hM0]∈RM×D\boldsymbol{H}_0 = [\boldsymbol{h}_1^0; \dots; \boldsymbol{h}_M^0] \in \mathbb{R}^{M \times D} and an autoregressive HOI sequence parser. The encoder outputs MM final interaction representations {hiL}i=1M⊂RD\{\boldsymbol{h}_i^L\}_{i=1}^M \subset \mathbb{R}^D.
    2. Prediction Heads:
      • Linear Projection Head: Fproj(h):RD→RD\mathcal{F}_{\text{proj}}(\boldsymbol{h}): \mathbb{R}^D \to \mathbb{R}^D maps each visual interaction embedding h\boldsymbol{h} into the shared visual-and-text embedding space.
      • Bounding Box Regressor: A class-agnostic head Fbbox(h):RD→R9\mathcal{F}_{\text{bbox}}(\boldsymbol{h}): \mathbb{R}^D \to \mathbb{R}^9 outputs [c^,b^p,b^o][\hat{c}, \hat{\boldsymbol{b}}_p, \hat{\boldsymbol{b}}_o], where c^∈[0,1]\hat{c} \in [0, 1] is a confidence score predicting whether the candidate contains an active interaction, and b^p,b^o∈R4\hat{\boldsymbol{b}}_p, \hat{\boldsymbol{b}}_o \in \mathbb{R}^4 denote the normalized coordinates of the interacting human and object bounding boxes, respectively.
    3. Text Encoder: A frozen pretrained CLIP text encoder transforms structured text prompts into semantic category vectors {sj}⊂RD\{\boldsymbol{s}_j\} \subset \mathbb{R}^D.
    4. Interaction Recognition: The interaction class for candidate token ii is determined by computing the cosine similarity hi⊤sj\boldsymbol{h}_i^\top \boldsymbol{s}_j against text embeddings in the joint semantic space.
  2. Knowl 2 — Layer-wise HOI Feature Extraction with Dedicated HOI Tokens

    model/method

    To extract instance-level interaction representations from an image I∈RH×W×3I \in \mathbb{R}^{H \times W \times 3} using a Vision Transformer (ViT), the input is partitioned into N=(H/P)×(W/P)N = (H/P) \times (W/P) non-overlapping patches of size P×PP \times P and projected with positional embeddings to form patch sequence X0[1:]=[x10;x20;… ;xN0]∈RN×D\boldsymbol{X}_0^{[1:]} = [\boldsymbol{x}_1^0; \boldsymbol{x}_2^0; \dots; \boldsymbol{x}_N^0] \in \mathbb{R}^{N \times D}, prepended with a learnable classification token z00∈RD\boldsymbol{z}_0^0 \in \mathbb{R}^D such that X0=[z00;X0[1:]]\boldsymbol{X}_0 = [\boldsymbol{z}_0^0; \boldsymbol{X}_0^{[1:]}].

    In addition to the standard patch sequence, MM learnable interaction query tokens H0=[h10;h20;… ;hM0]∈RM×D\boldsymbol{H}_0 = [\boldsymbol{h}_1^0; \boldsymbol{h}_2^0; \dots; \boldsymbol{h}_M^0] \in \mathbb{R}^{M \times D} are introduced to extract distinct HOI instances. At each Transformer layer ℓ∈{1,…,L}\ell \in \{1, \dots, L\}:

    1. Image patch and classification tokens are updated via standard multi-head self-attention (MHA): Xℓ′=MHA(LN(Xℓ−1))+Xℓ−1\boldsymbol{X}'_\ell = \text{MHA}(\text{LN}(\boldsymbol{X}_{\ell-1})) + \boldsymbol{X}_{\ell-1} Xℓ=MLP(LN(Xℓ′))+Xℓ′\boldsymbol{X}_\ell = \text{MLP}(\text{LN}(\boldsymbol{X}'_\ell)) + \boldsymbol{X}'_\ell

    2. The HOI tokens Hℓ−1\boldsymbol{H}_{\ell-1} aggregate spatial context by querying exclusively over the image patch embeddings Xℓ−1[1:]\boldsymbol{X}_{\ell-1}^{[1:]}, explicitly masking out the global classification token z0ℓ−1\boldsymbol{z}_0^{\ell-1}: Hℓ′=MHA(LN(Hℓ−1),LN(Xℓ−1[1:]),LN(Xℓ−1[1:]))+Hℓ−1\boldsymbol{H}'_\ell = \text{MHA}\big(\text{LN}(\boldsymbol{H}_{\ell-1}), \text{LN}(\boldsymbol{X}_{\ell-1}^{[1:]}), \text{LN}(\boldsymbol{X}_{\ell-1}^{[1:]})\big) + \boldsymbol{H}_{\ell-1} Hℓ=MLP(LN(Hℓ′))+Hℓ′\boldsymbol{H}_\ell = \text{MLP}(\text{LN}(\boldsymbol{H}'_\ell)) + \boldsymbol{H}'_\ell where LN(⋅)\text{LN}(\cdot) denotes Layer Normalization and MLP(⋅)\text{MLP}(\cdot) is a two-layer perceptron. Masking the [CLS][\text{CLS}] token prevents the HOI tokens from merely replicating image-level summaries rather than localizing specific interaction regions.

  3. Knowl 3 — Sequential HOI Parser for Unique Interaction Disambiguation

    model/method

    When processing multiple interaction queries simultaneously, parallel attention causes different [HOI][\text{HOI}] tokens to collapse onto the most salient interaction in an image. To enforce unique instance detection, a sequential parser module updates each [HOI][\text{HOI}] token conditionally on the global [CLS][\text{CLS}] token and all preceding [HOI][\text{HOI}] tokens at every layer ℓ\ell:

    hiℓ:=F(hiℓ∣z0ℓ,h1ℓ,…,hi−1ℓ)\boldsymbol{h}_i^\ell := \mathcal{F}(\boldsymbol{h}_i^\ell \mid \boldsymbol{z}_0^\ell, \boldsymbol{h}_1^\ell, \dots, \boldsymbol{h}_{i-1}^\ell)

    where z0ℓ∈RD\boldsymbol{z}_0^\ell \in \mathbb{R}^D is the [CLS][\text{CLS}] token at layer ℓ\ell, and hiℓ∈RD\boldsymbol{h}_i^\ell \in \mathbb{R}^D is the feature of the ii-th [HOI][\text{HOI}] token (i∈{1,…,M}i \in \{1, \dots, M\}).

    The parser is implemented as a Multi-Head Attention (MHA) block configured with an autoregressive causal attention mask. This mask restricts the attention field of the ii-th token so that it only attends to z0ℓ\boldsymbol{z}_0^\ell and {hkℓ}k=1i−1\{\boldsymbol{h}_k^\ell\}_{k=1}^{i-1}. Consequently, the primary token h1ℓ\boldsymbol{h}_1^\ell anchors to the most prominent interaction informed by the global context z0ℓ\boldsymbol{z}_0^\ell, while subsequent tokens hiℓ\boldsymbol{h}_i^\ell (i>1i > 1) condition on earlier detections to localize distinct, non-overlapping interaction pairs across the image.

  4. Knowl 4 — Prompt Learning with Learnable Context Tokens for Natural Language HOI Supervision

    model/method

    To construct textual representations for diverse human-object interaction pairs (action,object)(\text{action}, \text{object}) without relying on rigid manual sentence templates (e.g., "a photo of a person [ACT] [OBJ]"), prompt learning is applied directly in the text encoder.

    The input prompt for an action aa and object oo is formed by concatenating learnable continuous vectors with the category name tokens:

    Prompt(a,o)=[PREFIX]1[PREFIX]2…[PREFIX]Npre [ACT] [CONJUN]1…[CONJUN]Nconj [OBJ]\text{Prompt}(a, o) = [\text{PREFIX}]_1 [\text{PREFIX}]_2 \dots [\text{PREFIX}]_{N_{\text{pre}}} \, [\text{ACT}] \, [\text{CONJUN}]_1 \dots [\text{CONJUN}]_{N_{\text{conj}}} \, [\text{OBJ}]

    where:

    • [ACT][\text{ACT}] and [OBJ][\text{OBJ}] denote the tokenized category names of the human action and object.
    • [PREFIX]k∈RDtext[\text{PREFIX}]_k \in \mathbb{R}^{D_{\text{text}}} (k=1,…,Nprek = 1, \dots, N_{\text{pre}}, with Npre=8N_{\text{pre}} = 8) are learnable prefix context tokens replacing fixed phrases such as "a photo of a person".
    • [CONJUN]m∈RDtext[\text{CONJUN}]_m \in \mathbb{R}^{D_{\text{text}}} (m=1,…,Nconjm = 1, \dots, N_{\text{conj}}, with Nconj=4N_{\text{conj}} = 4) are learnable conjunction tokens that dynamically parameterize grammatical transitions between verbs and nouns.

    A learnable [EOS][\text{EOS}] token is appended to the sequence. The output of the frozen CLIP text encoder at [EOS][\text{EOS}] serves as the final interaction text representation s∈RD\boldsymbol{s} \in \mathbb{R}^D.

  5. Knowl 5 — Instance-Level Hold-Out Contrastive Loss for HOI Alignment

    equation

    For instance-level visual-text alignment in human-object interaction detection, multiple visual detections in an image may correspond to the same textual interaction category. Standard symmetric cross-entropy penalizes valid duplicate detections; hence, an asymmetric hold-out contrastive loss is used.

    Given predicted visual embedding h^i∈RD\hat{\boldsymbol{h}}_i \in \mathbb{R}^D assigned to target index ϕi\phi_i with text representation sϕi∈RD\boldsymbol{s}_{\phi_i} \in \mathbb{R}^D, visual-to-text loss Lv2t\mathcal{L}_{v2t} and text-to-visual loss Lt2v\mathcal{L}_{t2v} with learned temperature τ\tau are defined as:

    Lh(h^i,sϕi)=Lv2t+Lt2v\mathcal{L}_h(\hat{\boldsymbol{h}}_i, \boldsymbol{s}_{\phi_i}) = \mathcal{L}_{v2t} + \mathcal{L}_{t2v}

    Lv2t=−log⁡exp⁡(h^i⊤sϕi/τ)∑jexp⁡(h^i⊤sj/τ)\mathcal{L}_{v2t} = - \log \frac{\exp(\hat{\boldsymbol{h}}_i^\top \boldsymbol{s}_{\phi_i} / \tau)}{\sum_{j} \exp(\hat{\boldsymbol{h}}_i^\top \boldsymbol{s}_j / \tau)}

    Lt2v=−log⁡exp⁡(h^i⊤sϕi/τ)∑j: ϕj≠ϕiexp⁡(h^j⊤sϕi/τ)\mathcal{L}_{t2v} = - \log \frac{\exp(\hat{\boldsymbol{h}}_i^\top \boldsymbol{s}_{\phi_i} / \tau)}{\sum_{j:\, \phi_j \neq \phi_i} \exp(\hat{\boldsymbol{h}}_j^\top \boldsymbol{s}_{\phi_i} / \tau)}

    where ∑j\sum_j in Lv2t\mathcal{L}_{v2t} sums over all interaction text categories in the dictionary, and the denominator of Lt2v\mathcal{L}_{t2v} sums over visual prediction tokens jj whose assigned target index differs from ϕi\phi_i (holding out duplicate positive instances of the same interaction label in the same image).

  6. Knowl 6 — Bipartite Matching and End-to-End Optimization for HOI Detection

    model/method

    To train the one-stage HOI detector, a bipartite matching problem is solved per image between MM predicted candidate tuples {(h^i,c^i,b^pi,b^oi)}i=1M\{(\hat{\boldsymbol{h}}_i, \hat{c}_i, \hat{\boldsymbol{b}}_p^i, \hat{\boldsymbol{b}}_o^i)\}_{i=1}^M and KK ground-truth targets {(sk,bpk,bok)}k=1K\{(\boldsymbol{s}_k, \boldsymbol{b}_p^k, \boldsymbol{b}_o^k)\}_{k=1}^K, where K≤MK \le M.

    An optimal assignment ϕ∗=[ϕ1∗,…,ϕM∗]\phi^* = [\phi_1^*, \dots, \phi_M^*] (with ϕi∈{0,1,…,K}\phi_i \in \{0, 1, \dots, K\}, where ϕi=0\phi_i = 0 indicates no target assigned) is obtained by minimizing the matching cost:

    ϕ∗=arg⁡min⁡ϕ∑i=1M(Lm(i,ϕi)−c^i)\phi^* = \arg\min_\phi \sum_{i=1}^M \Big( \mathcal{L}_m(i, \phi_i) - \hat{c}_i \Big)

    Lm(i,ϕi)=Lb(b^pi,bpϕi)+Lb(b^oi,boϕi)+Lh(h^i,sϕi)\mathcal{L}_m(i, \phi_i) = \mathcal{L}_b(\hat{\boldsymbol{b}}_p^i, \boldsymbol{b}_p^{\phi_i}) + \mathcal{L}_b(\hat{\boldsymbol{b}}_o^i, \boldsymbol{b}_o^{\phi_i}) + \mathcal{L}_h(\hat{\boldsymbol{h}}_i, \boldsymbol{s}_{\phi_i})

    where Lb\mathcal{L}_b is the bounding box regression loss consisting of ℓ1\ell_1 loss and Generalized IoU (GIoU) loss on human and object bounding boxes, and Lh\mathcal{L}_h is the visual-text alignment loss.

    After matching, target confidence labels are assigned as ci=1c_i = 1 if ϕi∗≠0\phi_i^* \neq 0, and ci=0c_i = 0 otherwise. The parameters of the newly added modules are updated by minimizing:

    Ltotal=∑i=1MLm(i,ϕi∗)+Lc(c^i,ci)\mathcal{L}_{\text{total}} = \sum_{i=1}^M \mathcal{L}_m(i, \phi_i^*) + \mathcal{L}_c(\hat{c}_i, c_i)

    where Lc\mathcal{L}_c is binary cross-entropy loss. Pretrained CLIP parameters remain frozen during training.

  7. Knowl 7 — Zero-Shot HOI Detection Performance on HICO-DET

    data/table

    Under the generalized zero-shot setting on HICO-DET (holding out 120 rare interaction categories as unseen during training out of 600 total categories formed by 117 actions and 80 objects), THID outperforms prior composition-based zero-shot methods without requiring pre-specified unseen interaction lists or sample synthesis during training.

    Method One-stage Unseen mAP Seen mAP Full mAP
    Shen et al. (WACV 2018) 5.62 - 6.26
    FG (AAAI 2020) 10.93 12.60 12.26
    VCL (ECCV 2020) 10.06 24.28 21.43
    ATL (CVPR 2021) 9.18 24.67 21.57
    FCL (CVPR 2021) 13.16 24.23 22.01
    THID (Ours) ✓ 15.53 24.32 22.96

    THID achieves 15.53 mAP on unseen categories, a 2.37 mAP gain over the previous top method (FCL at 13.16 mAP), while preserving competitive performance on seen interactions (24.32 mAP vs. 24.67 mAP for ATL).

  8. Knowl 8 — Large-Vocabulary HOI Detection on SWIG-HOI

    data/table

    On the SWIG-HOI dataset, which contains 400 human actions and 1,000 object categories with ~1.8k naturally occurring unseen interactions out of ~5.5k test interactions across ~14k images, THID achieves state-of-the-art performance across all frequency splits.

    Method Non-rare mAP Rare mAP Unseen mAP Full mAP
    JSR (ECCV 2020) 10.01 6.10 2.34 6.08
    CHOID (ICCV 2021) 10.93 6.63 2.64 6.64
    QPIC + Proj (CVPR 2021) 16.95 10.84 6.21 11.12
    THID (Ours) 17.67 12.82 10.04 13.26

    Compared to a one-stage Transformer baseline adapted to CLIP via an MLP projection (QPIC + Proj), THID increases overall Full mAP by 2.14 points (from 11.12 to 13.26) and substantially improves generalization on unseen interactions by 3.83 mAP (from 6.21 to 10.04).

  9. Knowl 9 — Ablation of Sequential Attention Parsing and Global Prior Conditioning

    data/table

    An ablation study on SWIG-HOI evaluates the impact of the sequential parsing constraint and the inclusion of the global [CLS][\text{CLS}] token in the sequence parser module.

    Architecture Sequential Uses [CLS] Non-rare mAP Rare mAP Unseen mAP
    ViT-B/16 (No parser) 12.26 7.90 6.27
    + SetParser (Ablated CLS) 15.14 9.18 6.02
    + SetParser ✓ 16.93 10.09 7.07
    + SeqParser (Ablated CLS) ✓ 15.78 11.23 8.43
    + SeqParser (Full) ✓ ✓ 17.67 12.82 10.04

    Key observations:

    1. Removing the parser entirely causes unseen mAP to drop from 10.04 to 6.27.
    2. Enforcing sequential attention constraints (SeqParser vs. SetParser with [CLS][\text{CLS}]) improves unseen mAP by 2.97 points (10.04 vs. 7.07), demonstrating that order-dependent parsing prevents duplicate detections and encourages complementary interaction discovery.
    3. Conditioning the sequence parser on the global [CLS][\text{CLS}] token adds +1.61 mAP on unseen interactions (10.04 vs. 8.43) by providing image-wide contextual priors.
  10. Knowl 10 — Ablation of Prompt Formatting and Loss Formulations for HOI Alignment

    data/table

    Ablations on SWIG-HOI compare text prompt designs (fixed templates, definition-augmented templates, and learnable context tokens) as well as text-to-visual loss formulations.

    Text Prompt Formulation Non-rare mAP Rare mAP Unseen mAP
    Ä photo of a person [ACT] [OBJ]" 11.98 7.11 6.96
    Ä photo of a person [ACT] [OBJ]. [OBJ_DEF]" 12.22 8.66 7.70
    Ä photo of a person [ACT] [OBJ]. [ACT_DEF]" 14.97 9.62 8.22
    Learnable [Prefix] [ACT] [Conjun] [OBJ] 17.67 12.82 10.04
    Loss Formulation Non-rare mAP Rare mAP Unseen mAP
    Ablation of text-to-visual loss 15.81 11.83 9.30
    Multi-label soft margin loss 13.98 8.01 6.02
    Hold-out cross entropy loss 17.67 12.82 10.04

    Learnable prompt tokens outperform rigid templates by +3.08 mAP on unseen categories. For loss formulation, hold-out cross-entropy achieves 10.04 mAP on unseen interactions, outperforming multi-label soft margin loss (6.02 mAP) which disrupts CLIP's pretrained feature alignment.

  11. Knowl 11 — Fixed Input Resolution and Small Object Localization Limitation

    limitation

    The visual encoder is constrained to a fixed input image resolution of 224×224224 \times 224 pixels inherited from the pretrained ViT-B/16 architecture. This fixed input size restricts the detector's ability to localize small objects accurately compared to multi-scale HOI architectures that leverage feature pyramid networks and variable-resolution multi-scale training.

Coverage note — Preliminary heuristic baseline numbers (Table 1) and qualitative spatial heatmaps (Figure 4) were omitted in favor of the formal method definitions, complete benchmark evaluations, and quantitative ablation studies.

References

  1. 1.Jimmy Ba, J. Kiros, and Geoffrey E. Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  2. 2.Ankan Bansal, Sai Saketh Rambhatla, Abhinav Shrivastava, and Rama Chellappa. Detecting human-object interactions via functional generalization. In AAAI, 2020.
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020.
  4. 4.Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng. Learning to detect human-object interactions. In WACV, 2018.
  5. 5.Mingfei Chen, Yue Liao, Si Liu, Zhiyuan Chen, Fei Wang, and Chen Qian. Reformulating hoi detection as adaptive set prediction. In CVPR, 2021.
  6. 6.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Scaling egocentric vision: The epic-kitchens dataset. In ECCV, 2018.
  7. 7.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  8. 8.Hao-Shu Fang, Jinkun Cao, Yu-Wing Tai, and Cewu Lu. Pairwise body-part attention for recognizing human-object interactions. In ECCV, 2018.
  9. 9.Hao-Shu Fang, Yichen Xie, Dian Shao, and Cewu Lu. Dirv: Dense interaction region voting for end-to-end human-object interaction detection. In AAAI, 2021.
  10. 10.David F Fouhey, Wei-cheng Kuo, Alexei A Efros, and Jitendra Malik. From lifestyle vlogs to everyday interactions. In CVPR, 2018.
  11. 11.Chen Gao, Jiarui Xu, Yuliang Zou, and Jia-Bin Huang. Drg: Dual relation graph for human-object interaction detection. In ECCV, 2020.
  12. 12.Georgia Gkioxari, Ross Girshick, Piotr Dollár, and Kaiming He. Detecting and recognizing human-object interactions. In CVPR, 2018.
  13. 13.Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary object detection via vision and language knowledge distillation. arXiv preprint arXiv: 2104.13921, 2021.
  14. 14.Tanmay Gupta, Alexander Schwing, and Derek Hoiem. No-frills human-object interaction detection: Factorization, layout encodings, and training techniques. In ICCV, 2019.
  15. 15.Zhi Hou, Xiaojiang Peng, Yu Qiao, and Dacheng Tao. Visual compositional learning for human-object interaction detection. In ECCV, 2020.
  16. 16.Zhi Hou, Baosheng Yu, Yu Qiao, Xiaojiang Peng, and Dacheng Tao. Affordance transfer learning for human-object interaction detection. In CVPR, 2021.
  17. 17.Zhi Hou, Baosheng Yu, Yu Qiao, Xiaojiang Peng, and Dacheng Tao. Detecting human-object interaction via fabricated compositional learning. In CVPR, 2021.
  18. 18.Dat Huynh and Ehsan Elhamifar. Interaction compass: Multi-label zero-shot learning of human-object interactions via spatial relations. In ICCV, 2021.
  19. 19.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. arXiv preprint arXiv: 2102.05918, 2021.
  20. 20.Woojeong Jin, Yu Cheng, Yelong Shen, Weizhu Chen, and Xiang Ren. A good prompt is worth millions of parameters? low-resource prompt-based learning for vision-language models. arXiv preprint arXiv: 2110.08484, 2021.
  21. 21.Keizo Kato, Yin Li, and Abhinav Gupta. Compositional learning for human object interaction. In ECCV, 2018.
  22. 22.Bumsoo Kim, Taeho Choi, Jaewoo Kang, and Hyunwoo J. Kim. Uniondet: Union-level detector towards real-time human-object interaction detection. In ECCV, 2020.
  23. 23.Bumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim, and Hyunwoo J. Kim. Hotr: End-to-end human-object interaction detection with transformers. In CVPR, 2021.
  24. 24.Shuang Li, Yilun Du, Antonio Torralba, Josef Sivic, and Bryan Russell. Weakly supervised human-object interaction detection in video via contrastive spatiotemporal regions. In ICCV, 2021.
  25. 25.Yong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li, and Cewu Lu. Hoi analysis: Integrating and decomposing human-object interaction. In NeurIPS, 2020.
  26. 26.Yong-Lu Li, Liang Xu, Xinpeng Liu, Xijie Huang, Yue Xu, Shiyi Wang, Hao-Shu Fang, Ze Ma, Mingyang Chen, and Cewu Lu. Pastanet: Toward human activity knowledge engine. In CVPR, 2020.
  27. 27.Yong-Lu Li, Siyuan Zhou, Xijie Huang, Liang Xu, Ze Ma, Hao-Shu Fang, Yanfeng Wang, and Cewu Lu. Transferable interactiveness knowledge for human-object interaction detection. In CVPR, 2019.
  28. 28.Yue Liao, Si Liu, Fei Wang, Yanjie Chen, Chen Qian, and Jiashi Feng. Ppdm: Parallel point detection and matching for real-time human-object interaction detection. In CVPR, 2020.
  29. 29.Ye Liu, Junsong Yuan, and Chang Wen Chen. Consnet: Learning consistency graph for zero-shot human-object interaction detection. In ACM Multimedia, 2020.
  30. 30.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019.
  31. 31.Romero Morais, Vuong Le, Svetha Venkatesh, and Truyen Tran. Learning asynchronous and sparse human-object interaction in videos. In CVPR, 2021.
  32. 32.Julia Peyre, Ivan Laptev, Cordelia Schmid, and Josef Sivic. Detecting unseen visual relations using analogies. In ICCV, 2019.
  33. 33.Sarah Pratt, Mark Yatskar, Luca Weihs, Ali Farhadi, and Aniruddha Kembhavi. Grounded situation recognition. In ECCV, 2020.
  34. 34.Siyuan Qi, Wenguan Wang, Baoxiong Jia, Jianbing Shen, and Song-Chun Zhu. Learning human-object interactions by graph parsing neural networks. In ECCV, 2018.
  35. 35.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. arXiv preprint arXiv: 2103.00020, 2021.
  36. 36.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015.
  37. 37.Dandan Shan, Jiaqi Geng, Michelle Shu, and David Fouhey. Understanding human hands in contact at internet scale. In CVPR, 2020.
  38. 38.Liyue Shen, Serena Yeung, Judy Hoffman, Greg Mori, and Fei Fei Li. Scaling human-object interaction recognition through zero-shot learning. In WACV, 2018.
  39. 39.Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. How much can clip benefit vision-and-language tasks? arXiv preprint arXiv: 2107.06383, 2021.
  40. 40.Masato Tamura, Hiroki Ohashi, and Tomoaki Yoshinaga. Qpic: Query-based pairwise human-object interaction detection with image-wide contextual information. In CVPR, 2021.
  41. 41.Zhigang Tu, Hongyan Li, Dejun Zhang, Justin Dauwels, Baoxin Li, and Junsong Yuan. Action-stage emphasized spatiotemporal vlad for video action recognition. IEEE Transactions on Image Processing, 28(6):2799–2812, 2019.
  42. 42.Zhigang Tu, Wei Xie, Justin Dauwels, Baoxin Li, and Junsong Yuan. Semantic cues enhanced multimodality multistream cnn for action recognition. IEEE Transactions on Circuits and Systems for Video Technology, 29(5):1423–1437, 2019.
  43. 43.Oytun Ulutan, A S M Iftekhar, and B. S. Manjunath. Vsgnet: Spatial attention network for detecting human object interactions using graph convolutions. In CVPR, 2020.
  44. 44.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017.
  45. 45.Bo Wan, Desen Zhou, Yongfei Liu, Rongjie Li, and Xuming He. Pose-aware multi-level feature network for human object interaction detection. In ICCV, 2019.
  46. 46.Mengmeng Wang, Jiazheng Xing, and Yong Liu. Actionclip: A new paradigm for video action recognition, 2021.
  47. 47.Suchen Wang, Kim-Hui Yap, Henghui Ding, Jiyan Wu, Junsong Yuan, and Yap-Peng Tan. Discovering human interactions with large-vocabulary objects via query and multi-scale detection. In ICCV, 2021.
  48. 48.Suchen Wang, Kim-Hui Yap, Junsong Yuan, and Yap-Peng Tan. Discovering human interactions with novel objects via zero-shot learning. In CVPR, 2020.
  49. 49.Tiancai Wang, Rao Muhammad Anwer, Muhammad Haris Khan, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao, and Jorma Laaksonen. Deep contextual attention for human-object interaction detection. In ICCV, 2019.
  50. 50.Tiancai Wang, Tong Yang, Martin Danelljan, Fahad Shahbaz Khan, Xiangyu Zhang, and Jian Sun. Learning human-object interaction detection using interaction points. In CVPR, 2020.
  51. 51.Tete Xiao, Quanfu Fan, Dan Gutfreund, Mathew Monfort, Aude Oliva, and Bolei Zhou. Reasoning about human-object interactions through dual attention networks. In ICCV, 2019.
  52. 52.Bingjie Xu, Yongkang Wong, Junnan Li, Qi Zhao, and Mohan S. Kankanhalli. Learning to detect human-object interactions with knowledge. In CVPR, 2019.
  53. 53.Frederic Z. Zhang, Dylan Campbell, and Stephen Gould. Spatially conditioned graphs for detecting human-object interactions. In ICCV, 2021.
  54. 54.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. arXiv preprint arXiv: 2109.01134, 2021.
  55. 55.Penghao Zhou and Mingmin Chi. Relation parsing neural network for human-object interaction detection. In ICCV, 2019.
  56. 56.Tianfei Zhou, Wenguan Wang, Siyuan Qi, Haibin Ling, and Jianbing Shen. Cascaded human-object interaction recognition. In CVPR, 2020.
  57. 57.Cheng Zou, Bohan Wang, Yue Hu, Junqi Liu, Qian Wu, Yu Zhao, Boxun Li, Chenguang Zhang, Chi Zhang, Yichen Wei, and Jian Sun. End-to-end human object interaction detection with hoi transformer. In CVPR, 2021.

Citation

MLA
Wang, S., et al. “Learning Transferable Human-Object Interaction Detector with Natural Language Supervision”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 929–38, https://doi.org/10.1109/CVPR52688.2022.00101.
APA
Wang, S., Duan, Y., Ding, H., Tan, Y.-P., Yap, K.-H., & Yuan, J. (2022). Learning Transferable Human-Object Interaction Detector with Natural Language Supervision. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 929–938. https://doi.org/10.1109/CVPR52688.2022.00101
Chicago
Wang, S., Y. Duan, H. Ding, Y.-P. Tan, K.-H. Yap, and J. Yuan. 2022. “Learning Transferable Human-Object Interaction Detector with Natural Language Supervision”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 929–38. https://doi.org/10.1109/CVPR52688.2022.00101.
Harvard
Wang, S. et al. (2022) “Learning Transferable Human-Object Interaction Detector with Natural Language Supervision”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 929–938. Available at: https://doi.org/10.1109/CVPR52688.2022.00101.
Vancouver
1. Wang S, Duan Y, Ding H, Tan Y-P, Yap K-H, Yuan J (2022) Learning Transferable Human-Object Interaction Detector with Natural Language Supervision. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 929–938

BibTeX

@inproceedings{Wang_2022, title={Learning Transferable Human-Object Interaction Detector with Natural Language Supervision}, url={http://dx.doi.org/10.1109/CVPR52688.2022.00101}, DOI={10.1109/cvpr52688.2022.00101}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Wang, Suchen and Duan, Yueqi and Ding, Henghui and Tan, Yap-Peng and Yap, Kim-Hui and Yuan, Junsong}, year={2022}, month=June, pages={929–938} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE