Relational Context Learning for Human-Object Interaction Detection

Sanghyun KimDeunsol JungMinsu Cho

article2023CVPR75 citations

Proposes a multiplex relation network with a three-branch transformer architecture that systematically exchanges unary, pairwise, and ternary context across human, object, and interaction tokens to advance state-of-the-art detection on HICO-DET and V-COCO benchmarks.

Listen

Understanding human activities in visual data requires accurately identifying people, objects, and the relationships between them. This capability is essential for advanced computer vision applications such as automated video surveillance, action recognition, image search, and assistive captioning. While recent artificial intelligence systems rely on transformer networks to detect human-object interactions, current architectures struggle to balance task specialization with relationship understanding. Single-branch networks attempt all detection sub-tasks simultaneously, losing specificity, while multi-branch systems separate human-object localization from interaction classification without exchanging sufficient relational context.

The article demonstrates a novel framework called the Multiplex Relation Network, designed to overcome these limitations. The objective of the article is to establish an architecture that learns specialized representations for individual detection tasks while enabling rich, bidirectional relational reasoning across them.

The authors develop a three-branch architecture that assigns dedicated transformer decoders to human detection, object detection, and interaction classification. To connect these branches, the framework uses a multiplex relation module that progressively integrates individual, paired, and three-way visual relationships into a unified context representation. An attentive fusion module then selects and routes the necessary relationship cues back to each task-specific branch. The approach was evaluated through extensive benchmarking on standard public datasets, specifically the 38,118-training-image HICO-DET benchmark and the V-COCO dataset, comparing performance against existing one-stage and two-stage methods.

The key findings show that the proposed framework consistently surpasses prior state-of-the-art methods across all standard evaluation settings. On the HICO-DET benchmark, the model achieves a mean average precision of 32.87 on the full default setting, exceeding previous transformer and graph-based models without requiring auxiliary inputs like human pose estimation or linguistic features. On the V-COCO dataset, it achieves leading scores of 68.8 and 71.0 across both standard evaluation scenarios. Ablation experiments demonstrate that incorporating the complete relational context—combining single, pairwise, and three-way relations—boosts detection performance by roughly 6 percentage points compared to isolated branches. Furthermore, maintaining fully separated parameters for human and object decoders improves accuracy by about 2 percentage points over shared-parameter baselines, confirming the necessity of treating human and object detection as distinct sub-tasks.

These results demonstrate that high-order relational modeling directly addresses the core trade-off in visual interaction detection between task specialization and context sharing. Organizations building computer vision pipelines can achieve higher accuracy and robustness without relying on complex, multi-modal hand-crafted feature pipelines. Decision-makers in AI development should consider adopting three-branch contextualized architectures over conventional single-decoder or weakly coupled two-branch models.

While the empirical results on standard benchmarks are strong and demonstrate high statistical confidence, the system's reliance on fixed query counts and standard object detection backbones means real-time deployment constraints and generalization to non-standard, open-world settings remain areas for future validation.

arXiv: 2304.04997
Cover for Relational Context Learning for Human-Object Interaction Detection

Abstract

Recent state-of-the-art methods for HOI detection typically build on transformer architectures with two decoder branches, one for human-object pair detection and the other for interaction classification. Such disentangled transformers, however, may suffer from insufficient context exchange between the branches and lead to a lack of context information for relational reasoning, which is critical in discovering HOI instances. In this work, we propose the multiplex relation network (MUREN) that performs rich context exchange between three decoder branches using unary, pairwise, and ternary relations of human, object, and interaction tokens. The proposed method learns comprehensive relational contexts for discovering HOI instances, achieving state-of-the-art performance on two standard benchmarks for HOI detection, HICO-DET and V-COCO.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. CNN-based HOI Methods
  • 2.2. Transformer-based HOI Methods
  • 3. Problem Definition
  • 4. Method
  • 4.1. Image Encoding
  • 4.2. HOI Token Decoding
  • 4.3. Relational Contextualization
  • 4.4. Attentive Fusion
  • 4.5. Training Objective
  • 4.6. Inference
  • 5. Experiments
  • 5.1. Datasets and Metrics
  • 5.2. Implementation Details
  • 5.3. Comparison with State-of-the-Art
  • 5.4. Ablation Study
  • 5.5. Qualitative Results
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Three-branch MUREN architecture

    model/method

    MUREN is a one-stage transformer architecture for detecting human–object interactions (HOIs) by exchanging relational context among three task-specific decoder branches. Given an image, a CNN backbone followed by a 1×11\times1 convolution, positional encoding, flattening, and a transformer encoder produces image tokens X∈RT×D\mathbf{X}\in\mathbb{R}^{T\times D}, where TT is the number of image tokens and DD is their channel dimension.

    The model has separate human, object, and interaction branches, indexed by τ∈{H,O,I}\tau\in\{\mathrm{H},\mathrm{O},\mathrm{I}\}. Each branch contains LL repeated layers and starts from NN learnable query tokens Q(0)τ∈RN×D\mathbf{Q}^{\tau}_{(0)}\in\mathbb{R}^{N\times D}. At layer ll, a branch-specific transformer decoder cross-attends to the image tokens and produces task-specific tokens:

    F(l)τ=Dec⁡(l)τ ⁣(Q(l−1)τ,X),\mathbf{F}^{\tau}_{(l)}=\operatorname{Dec}^{\tau}_{(l)}\!\left(\mathbf{Q}^{\tau}_{(l-1)},\mathbf{X}\right),

    where F(l)τ={f(l),iτ}i=1N\mathbf{F}^{\tau}_{(l)}=\{\mathbf{f}^{\tau}_{(l),i}\}_{i=1}^{N} and Dec⁡(l)τ\operatorname{Dec}^{\tau}_{(l)} is a transformer decoder layer. The human branch predicts human boxes, the object branch predicts object boxes and object classes, and the interaction branch predicts interaction classes. For every aligned query index ii, the three task-specific tokens are processed by the multiplex relation embedding module (MURE), which produces relational context, and an attentive fusion module injects selected relational context back into every branch before the next layer. The final tokens from the three branches are converted into HOI predictions by feed-forward networks.

  2. Knowl 2 — Multiplex relation embedding for unary, pairwise, and ternary context

    model/method

    The multiplex relation embedding module (MURE) constructs a relational context vector for each aligned query i∈{1,…,N}i\in\{1,\ldots,N\} from the human, object, and interaction task-specific features fiH,fiO,fiI∈RD\mathbf{f}_i^{\mathrm{H}},\mathbf{f}_i^{\mathrm{O}},\mathbf{f}_i^{\mathrm{I}}\in\mathbb{R}^{D}. It combines three orders of context: unary context for individual task features, pairwise context for human–object, human–interaction, and object–interaction pairs, and ternary context for the complete HOI triplet.

    First, MURE constructs a ternary feature with a learned multilayer perceptron (MLP):

    fiHOI=MLP⁡([fiH;fiO;fiI]),\mathbf{f}^{\mathrm{HOI}}_i=\operatorname{MLP}\left([\mathbf{f}^{\mathrm{H}}_i;\mathbf{f}^{\mathrm{O}}_i;\mathbf{f}^{\mathrm{I}}_i]\right),

    where [ ; ][\,;\,] denotes concatenation. Unary context is computed by self-attention over the three task features and embedded into the ternary feature by cross-attention:

    Ui=SelfAttn⁡({fiH,fiO,fiI}),\mathbf{U}_i=\operatorname{SelfAttn}\left(\{\mathbf{f}^{\mathrm{H}}_i,\mathbf{f}^{\mathrm{O}}_i,\mathbf{f}^{\mathrm{I}}_i\}\right), f~iHOI=CrossAttn⁡(fiHOI,Ui).\widetilde{\mathbf{f}}^{\mathrm{HOI}}_i=\operatorname{CrossAttn}\left(\mathbf{f}^{\mathrm{HOI}}_i,\mathbf{U}_i\right).

    MURE forms three pairwise features using separate MLPs:

    fiHO=MLP⁡([fiH;fiO]),fiHI=MLP⁡([fiH;fiI]),fiOI=MLP⁡([fiO;fiI]).\mathbf{f}^{\mathrm{HO}}_i=\operatorname{MLP}([\mathbf{f}^{\mathrm{H}}_i;\mathbf{f}^{\mathrm{O}}_i]),\quad \mathbf{f}^{\mathrm{HI}}_i=\operatorname{MLP}([\mathbf{f}^{\mathrm{H}}_i;\mathbf{f}^{\mathrm{I}}_i]),\quad \mathbf{f}^{\mathrm{OI}}_i=\operatorname{MLP}([\mathbf{f}^{\mathrm{O}}_i;\mathbf{f}^{\mathrm{I}}_i]).

    Pairwise context is self-attended and then embedded into the unary-enriched ternary feature:

    Pi=SelfAttn⁡({fiHO,fiHI,fiOI}),\mathbf{P}_i=\operatorname{SelfAttn}\left(\{\mathbf{f}^{\mathrm{HO}}_i,\mathbf{f}^{\mathrm{HI}}_i,\mathbf{f}^{\mathrm{OI}}_i\}\right), f^iHOI=CrossAttn⁡(f~iHOI,Pi).\widehat{\mathbf{f}}^{\mathrm{HOI}}_i=\operatorname{CrossAttn}\left(\widetilde{\mathbf{f}}^{\mathrm{HOI}}_i,\mathbf{P}_i\right).

    Finally, MURE cross-attends to the image tokens to produce the multiplex relation context mi\mathbf{m}_i:

    mi=CrossAttn⁡(f^iHOI,X).\mathbf{m}_i=\operatorname{CrossAttn}\left(\widehat{\mathbf{f}}^{\mathrm{HOI}}_i,\mathbf{X}\right).

    The MLPs operate on tuples of multiple inputs, so the resulting high-order features can model structural relations jointly rather than only summing independently computed unary features.

  3. Knowl 3 — Task-conditioned attentive fusion

    model/method

    MUREN propagates the multiplex relation context m(l),i∈RD\mathbf{m}_{(l),i}\in\mathbb{R}^{D} back to each task-specific token f(l),iτ∈RD\mathbf{f}^{\tau}_{(l),i}\in\mathbb{R}^{D} using a task-conditioned channel-attention mechanism. For branch τ∈{H,O,I}\tau\in\{\mathrm{H},\mathrm{O},\mathrm{I}\}, query ii, and layer ll, the channel gate is

    α(l),iτ=σ ⁣(MLP⁡([f(l),iτ;m(l),i])),\boldsymbol{\alpha}_{(l),i}^{\tau}=\sigma\!\left(\operatorname{MLP}\left([\mathbf{f}^{\tau}_{(l),i};\mathbf{m}_{(l),i}]\right)\right),

    where σ\sigma is the element-wise sigmoid function and α(l),iτ∈(0,1)D\boldsymbol{\alpha}_{(l),i}^{\tau}\in(0,1)^D. The refined branch token is

    q(l),iτ=f(l),iτ+α(l),iτ⊙MLP⁡([f(l),iτ;m(l),i]),\mathbf{q}^{\tau}_{(l),i}=\mathbf{f}^{\tau}_{(l),i}+\boldsymbol{\alpha}_{(l),i}^{\tau}\odot\operatorname{MLP}\left([\mathbf{f}^{\tau}_{(l),i};\mathbf{m}_{(l),i}]\right),

    where ⊙\odot is element-wise multiplication. Because the transformation and channel gate are conditioned on the receiving task token, the human, object, and interaction branches can select different portions of the same multiplex relation context. The refined tokens q(l),iτ\mathbf{q}^{\tau}_{(l),i} become the inputs to the next decoder layer.

  4. Knowl 4 — HOI prediction, training objective, and inference

    algorithm

    For each query i∈{1,…,N}i\in\{1,\ldots,N\}, MUREN uses the final human token q(L),iH\mathbf{q}^{\mathrm{H}}_{(L),i}, object token q(L),iO\mathbf{q}^{\mathrm{O}}_{(L),i}, and interaction token q(L),iI\mathbf{q}^{\mathrm{I}}_{(L),i} to predict a human box, object box, object-class distribution, and interaction-class distribution:

    biH=FFN⁡hbox(q(L),iH)∈R4,biO=FFN⁡obox(q(L),iO)∈R4,\mathbf{b}^{\mathrm{H}}_i=\operatorname{FFN}_{\mathrm{hbox}}(\mathbf{q}^{\mathrm{H}}_{(L),i})\in\mathbb{R}^{4},\qquad \mathbf{b}^{\mathrm{O}}_i=\operatorname{FFN}_{\mathrm{obox}}(\mathbf{q}^{\mathrm{O}}_{(L),i})\in\mathbb{R}^{4}, piO=softmax⁡(FFN⁡oc(q(L),iO))∈R∣O∣,piI=sigmoid⁡(FFN⁡ic(q(L),iI))∈R∣I∣,\mathbf{p}^{\mathrm{O}}_i=\operatorname{softmax}(\operatorname{FFN}_{\mathrm{oc}}(\mathbf{q}^{\mathrm{O}}_{(L),i}))\in\mathbb{R}^{|\mathcal{O}|}, \qquad \mathbf{p}^{\mathrm{I}}_i=\operatorname{sigmoid}(\operatorname{FFN}_{\mathrm{ic}}(\mathbf{q}^{\mathrm{I}}_{(L),i}))\in\mathbb{R}^{|\mathcal{I}|},

    where O\mathcal{O} and I\mathcal{I} are the object and interaction label sets, respectively. The four-dimensional box vectors use the paper's box parameterization.

    Training uses Hungarian matching to assign predicted queries to ground-truth HOIs. The matched predictions are optimized with a weighted sum of L1L_1 box loss, generalized-IoU box loss, object cross-entropy loss, and interaction focal loss:

    L=λL1LL1+λGIoULGIoU+λocLoc+λicLic.\mathcal{L}=\lambda_{\mathrm{L1}}\mathcal{L}_{\mathrm{L1}}+\lambda_{\mathrm{GIoU}}\mathcal{L}_{\mathrm{GIoU}}+\lambda_{\mathrm{oc}}\mathcal{L}_{\mathrm{oc}}+\lambda_{\mathrm{ic}}\mathcal{L}_{\mathrm{ic}}.

    The same prediction heads are attached to every decoder layer, and the same loss is applied at intermediate layers as auxiliary supervision. At inference, the object class for query ii is selected as ji′=arg⁡max⁡jpi,jOj'_i=\arg\max_j p^{\mathrm{O}}_{i,j}. Each candidate interaction class t∈It\in\mathcal{I} receives the score pi,ji′Opi,tIp^{\mathrm{O}}_{i,j'_i}p^{\mathrm{I}}_{i,t}, and the highest-scoring candidates are retained as the final top-kk HOI detections.

  5. Knowl 5 — HICO-DET benchmark performance

    data/table

    On HICO-DET, MUREN is evaluated on 38,118 training images and 9,658 test images containing 80 object classes, 117 interaction classes, and 600 HOI classes. Full, Rare, and Non-Rare contain 600, 138, and 462 HOI classes, respectively. Default evaluation computes average precision over all test images, whereas Known Object evaluation restricts each HOI class to images containing its object. The following values compare MUREN with the strongest previously reported value among the compared methods for each metric; all entries are mAP values as reported by the paper.

    Could not parse LaTeX table

    MUREN is best among the compared methods in five of the six reported splits and is second in Known Object Rare, where its 30.88 mAP is below the previous best 31.43. It uses only appearance features with an R50 backbone, while several competing methods use additional spatial, linguistic, pose, or multiscale features, or deeper backbones.

  6. Knowl 6 — V-COCO benchmark performance

    data/table

    On V-COCO, MUREN is evaluated on 5,400 training images and 4,946 test images with 80 object classes and 29 action classes. Scenario 1 requires the model to output [0,0,0,0][0,0,0,0] for an occluded object's box, while Scenario 2 ignores the predicted occluded-object box when computing role average precision. MUREN achieves the following role average precision values compared with the strongest previously reported values among the compared methods.

    Could not parse LaTeX table

    The best previous values are 66.2 for Scenario 1 and 70.7 for Scenario 2. MUREN improves on them by 2.6 and 0.3 percentage points, respectively, establishing the strongest reported performance in both V-COCO evaluation scenarios used by the paper.

  7. Knowl 7 — Complementarity of ternary, unary, and pairwise context

    empirical result

    An ablation on V-COCO measures the effect of adding each relation-context order to a baseline with no relational context exchange. The baseline and all variants use the same MUREN framework; the check marks indicate which context types are used for exchange. The reported metrics are role average precision for Scenario 1 and Scenario 2.

    Could not parse LaTeX table

    Ternary context alone improves the baseline by 4.55 and 4.22 percentage points in the two scenarios, showing the value of holistic HOI information. Adding either unary or pairwise context provides further gains, and using all three contexts gives the largest improvement: 6.23 points in Scenario 1 and 5.86 points in Scenario 2. The results support the paper's claim that fine-grained unary and pairwise information complements the holistic ternary representation.

  8. Knowl 8 — Relational-context propagation and attentive fusion ablations

    empirical result

    MUREN was ablated by changing which task branches receive the multiplex relation context. On V-COCO, the no-propagation baseline scores 62.52 and 65.14 in Scenarios 1 and 2. Propagating context to the human branch alone gives 64.44 and 66.62; to the object branch alone gives 63.66 and 66.00; to both detection branches gives 65.29 and 67.50; and to the interaction branch alone gives 65.71 and 67.91. Propagating context to all three branches produces 68.75 and 71.00. The interaction branch benefits particularly strongly, but complete three-way exchange is best.

    The attentive fusion module was separately decomposed into task-conditioned transformation of the multiplex context and channel attention. The values below are V-COCO role AP scores.

    Could not parse LaTeX table

    Either attentive-fusion component improves over simple element-wise addition of task tokens and multiplex context, while using both is best. This demonstrates that each subtask benefits from selecting context conditioned on its own representation rather than receiving an undifferentiated context vector.

  9. Knowl 9 — Disentangling human and object branches

    empirical result

    MUREN keeps the human and object decoder branches separate because human and object roles require different representations. The human is an active HOI participant whose pose, clothing, and other attributes may need specialized features, whereas the object is comparatively passive. The paper tests sharing parameters between corresponding human and object layers across kk decoder layers on V-COCO.

    Could not parse LaTeX table

    MUREN-(0) shares no corresponding layers and is the full model. MUREN-(3) shares three layers, and MUREN-(6) shares all six layers. The fully shared variant loses 2.2 and 1.9 percentage points relative to no sharing in Scenarios 1 and 2. A parameter-matched variant, MUREN†\mathrm{MUREN}^{\dagger}, changes the number of layers rather than sharing the branches and still performs below MUREN-(0), supporting the need for disentangled human and object processing rather than merely attributing the gain to parameter count.

  10. Knowl 10 — Training and implementation configuration

    experimental setup

    The evaluated MUREN model uses a ResNet-50 CNN backbone followed by a six-layer transformer encoder and L=6L=6 layers in each of the three decoder branches. It uses N=64N=64 queries on HICO-DET and N=100N=100 queries on V-COCO. The loss weights (λL1,λGIoU,λoc,λic)(\lambda_{\mathrm{L1}},\lambda_{\mathrm{GIoU}},\lambda_{\mathrm{oc}},\lambda_{\mathrm{ic}}) are (2.5,1,1,1)(2.5,1,1,1).

    The network is initialized from a DETR model pretrained on MS-COCO and optimized with AdamW using weight decay 10−410^{-4}. The initial learning rate is 10−510^{-5} for the CNN backbone and 10−410^{-4} for the other components. Models are trained for 100 epochs with batch size 16 on four RTX 3090 GPUs. For V-COCO, the CNN backbone is frozen to reduce overfitting and the learning rate is set to 4×10−54\times10^{-5}.

    The HICO-DET evaluation reports mAP under Default and Known Object settings, each split into Full, Rare, and Non-Rare classes. The V-COCO evaluation reports role average precision under the two occlusion-handling scenarios described in the benchmark definition.

Coverage note — The qualitative attention visualizations are omitted because they provide illustrative evidence about branch focus and MURE coverage but no additional quantitative or load-bearing method result.

References

  1. 1.Carlo Bretti and Pascal Mettes. Zero-shot action recognition from diverse object-scene compositions. arXiv preprint arXiv:2110.13479, 2021.
  2. 2.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020.
  3. 3.Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng. Learning to detect human-object interactions. In 2018 ieee winter conference on applications of computer vision (wacv), pages 381–389. IEEE, 2018.
  4. 4.Mingfei Chen, Yue Liao, Si Liu, Zhiyuan Chen, Fei Wang, and Chen Qian. Reformulating hoi detection as adaptive set prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9004–9013, 2021.
  5. 5.Qi Dong, Zhuowen Tu, Haofu Liao, Yuting Zhang, Vijay Mahadevan, and Stefano Soatto. Visual relationship detection using part-and-sum transformers with composite queries. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3550–3559, 2021.
  6. 6.Hao-Shu Fang, Yichen Xie, Dian Shao, and Cewu Lu. Dirv: Dense interaction region voting for end-to-end human-object interaction detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1291–1299, 2021.
  7. 7.Chen Gao, Jiarui Xu, Yuliang Zou, and Jia-Bin Huang. Drg: Dual relation graph for human-object interaction detection. In European Conference on Computer Vision, pages 696–712. Springer, 2020.
  8. 8.Chen Gao, Yuliang Zou, and Jia-Bin Huang. ican: Instance-centric attention network for human-object interaction detection. arXiv preprint arXiv:1808.10437, 2018.
  9. 9.Albert Gordo and Diane Larlus. Beyond instance-level image retrieval: Leveraging captions to learn a global visual representation for semantic retrieval. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6589–6598, 2017.
  10. 10.Saurabh Gupta and Jitendra Malik. Visual semantic role labeling. arXiv preprint arXiv:1505.04474, 2015.
  11. 11.Tanmay Gupta, Alexander Schwing, and Derek Hoiem. No-frills human-object interaction detection: Factorization, layout encodings, and training techniques. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9677–9685, 2019.
  12. 12.Simao Herdade, Armin Kappeler, Kofi Boakye, and Joao Soares. Image captioning: Transforming objects into words. Advances in Neural Information Processing Systems, 32, 2019.
  13. 13.Zhi Hou, Xiaojiang Peng, Yu Qiao, and Dacheng Tao. Visual compositional learning for human-object interaction detection. In European Conference on Computer Vision, pages 584–600. Springer, 2020.
  14. 14.Bumsoo Kim, Taeho Choi, Jaewoo Kang, and Hyunwoo J Kim. Uniondet: Union-level detector towards real-time human-object interaction detection. In European Conference on Computer Vision, pages 498–514. Springer, 2020.
  15. 15.Bumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim, and Hyunwoo J Kim. Hotr: End-to-end human-object interaction detection with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 74–83, 2021.
  16. 16.Bumsoo Kim, Jonghwan Mun, Kyoung-Woon On, Minchul Shin, Junhyun Lee, and Eun-Sol Kim. Mstr: Multi-scale transformer for end-to-end human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19578–19587, 2022.
  17. 17.Harold W Kuhn. The hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97, 1955.
  18. 18.Yong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li, and Cewu Lu. Hoi analysis: Integrating and decomposing human-object interaction. Advances in Neural Information Processing Systems, 33:5011–5022, 2020.
  19. 19.Yong-Lu Li, Siyuan Zhou, Xijie Huang, Liang Xu, Ze Ma, Hao-Shu Fang, Yanfeng Wang, and Cewu Lu. Transferable interactiveness knowledge for human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3585–3594, 2019.
  20. 20.Yue Liao, Si Liu, Fei Wang, Yanjie Chen, Chen Qian, and Jiashi Feng. Ppdm: Parallel point detection and matching for real-time human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 482–490, 2020.
  21. 21.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, pages 2980–2988, 2017.
  22. 22.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  23. 23.Ye Liu, Junsong Yuan, and Chang Wen Chen. Consnet: Learning consistency graph for zero-shot human-object interaction detection. In Proceedings of the 28th ACM International Conference on Multimedia, pages 4235–4243, 2020.
  24. 24.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  25. 25.Gyeongsik Moon, Heeseung Kwon, Kyoung Mu Lee, and Minsu Cho. Integralaction: Pose-driven feature integration for robust human action recognition in videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3339–3348, 2021.
  26. 26.Siyuan Qi, Wenguan Wang, Baoxiong Jia, Jianbing Shen, and Song-Chun Zhu. Learning human-object interactions by graph parsing neural networks. In Proceedings of the European conference on computer vision (ECCV), pages 401–417, 2018.
  27. 27.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28, 2015.
  28. 28.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 658–666, 2019.
  29. 29.Masato Tamura, Hiroki Ohashi, and Tomoaki Yoshinaga. Qpic: Query-based pairwise human-object interaction detection with image-wide contextual information. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10410–10419, 2021.
  30. 30.Oytun Ulutan, ASM Iftekhar, and Bangalore S Manjunath. Vsgnet: Spatial attention network for detecting human object interactions using graph convolutions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 13617–13626, 2020.
  31. 31.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  32. 32.Hai Wang, Wei-shi Zheng, and Ling Yingbiao. Contextual heterogeneous graph network for human-object interaction detection. In European Conference on Computer Vision, pages 248–264. Springer, 2020.
  33. 33.Hui Wu, Min Wang, Wengang Zhou, Houqiang Li, and Qi Tian. Contextual similarity distillation for asymmetric image retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9489–9498, June 2022.
  34. 34.Mingrui Wu, Xuying Zhang, Xiaoshuai Sun, Yiyi Zhou, Chao Chen, Jiaxin Gu, Xing Sun, and Rongrong Ji. Difnet: Boosting visual information flow for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18020–18029, June 2022.
  35. 35.Bingjie Xu, Yongkang Wong, Junnan Li, Qi Zhao, and Mohan S Kankanhalli. Learning to detect human-object interactions with knowledge. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019.
  36. 36.Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. Exploring visual relationship for image captioning. In Proceedings of the European conference on computer vision (ECCV), pages 684–699, 2018.
  37. 37.Sangwoong Yoon, Woo Young Kang, Sungwook Jeon, SeongEun Lee, Changjin Han, Jonghun Park, and Eun-Sol Kim. Image-to-image retrieval by learning similarity between scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 10718–10726, 2021.
  38. 38.Aixi Zhang, Yue Liao, Si Liu, Miao Lu, Yongliang Wang, Chen Gao, and Xiaobo Li. Mining the benefits of two-stage and one-stage hoi detection. Advances in Neural Information Processing Systems, 34:17209–17220, 2021.
  39. 39.Frederic Z Zhang, Dylan Campbell, and Stephen Gould. Spatially conditioned graphs for detecting human-object interactions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13319–13327, 2021.
  40. 40.Frederic Z Zhang, Dylan Campbell, and Stephen Gould. Efficient two-stage detection of human-object interactions with a novel unary-pairwise transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20104–20112, 2022.
  41. 41.Yong Zhang, Yingwei Pan, Ting Yao, Rui Huang, Tao Mei, and Chang-Wen Chen. Exploring structure-aware transformer over interaction proposals for human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19548–19557, 2022.
  42. 42.Yubo Zhang, Pavel Tokmakov, Martial Hebert, and Cordelia Schmid. A structured model for action detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9975–9984, 2019.
  43. 43.Xubin Zhong, Xian Qu, Changxing Ding, and Dacheng Tao. Glance and gaze: Inferring action-aware points for one-stage human-object interaction detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13234–13243, 2021.
  44. 44.Desen Zhou, Zhichao Liu, Jian Wang, Leshan Wang, Tao Hu, Errui Ding, and Jingdong Wang. Human-object interaction detection via disentangled transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19568–19577, 2022.
  45. 45.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020.
  46. 46.Cheng Zou, Bohan Wang, Yue Hu, Junqi Liu, Qian Wu, Yu Zhao, Boxun Li, Chenguang Zhang, Chi Zhang, Yichen Wei, et al. End-to-end human object interaction detection with hoi transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11825–11834, 2021.
  47. 47.Cheng Zou, Bohan Wang, Yue Hu, Junqi Liu, Qian Wu, Yu Zhao, Boxun Li, Chenguang Zhang, Chi Zhang, Yichen Wei, et al. End-to-end human object interaction detection with hoi transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11825–11834, 2021.

Citation

MLA
Kim, S., et al. “Relational Context Learning for Human-Object Interaction Detection”. arXiv, 2023, http://arxiv.org/abs/2304.04997v1.
APA
Kim, S., Jung, D., & Cho, M. (2023). Relational Context Learning for Human-Object Interaction Detection. arXiv. http://arxiv.org/abs/2304.04997v1
Chicago
Kim, S., D. Jung, and M. Cho. 2023. “Relational Context Learning for Human-Object Interaction Detection”. arXiv. http://arxiv.org/abs/2304.04997v1.
Harvard
Kim, S., Jung, D. and Cho, M. (2023) “Relational Context Learning for Human-Object Interaction Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.04997v1.
Vancouver
1. Kim S, Jung D, Cho M (2023) Relational Context Learning for Human-Object Interaction Detection. arXiv

BibTeX

@article{kim2023relational,
  title = {Relational Context Learning for Human-Object Interaction Detection},
  author = {Kim, Sanghyun and Jung, Deunsol and Cho, Minsu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.04997v1},
  eprint = {2304.04997}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE