Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection

Xian QuChangxing DingXingao LiXubin ZhongDacheng Tao

article2022CVPR61 citations

Proposes a knowledge distillation framework that guides transformer decoders using ground-truth oracle queries alongside a context-consistent image stitching augmentation, substantially accelerating training convergence and setting state-of-the-art accuracy in human-object interaction detection without adding inference overhead.

Listen

Human-Object Interaction detection aims to identify people, the objects they interact with, and the specific actions linking them within an image. This capability is critical for computer vision applications such as scene understanding, healthcare monitoring, and autonomous systems. While modern transformer-based architectures have improved performance by capturing global visual context, they face two key bottlenecks: ambiguous queries that slow training and degrade feature representation, and a scarcity of training images containing multiple labeled interaction pairs, which restricts the model's ability to process complex visual scenes.

The article develops and evaluates two complementary solutions: a knowledge distillation framework termed Distillation using Oracle Queries (DOQ) and an online data augmentation technique called Context-Consistent Stitching (CCS). Together, these methods aim to improve the accuracy, training speed, and scalability of transformer-based interaction detection without adding computational overhead during practical deployment.

The evaluation was conducted using standardized computer vision benchmarks, including HICO-DET (over 47,000 images), HOI-A (over 38,000 images), and V-COCO (over 10,000 images). The authors paired the proposed training framework with standard baseline architectures like QPIC, HOTR, and CDN across multiple convolutional backbones. The approach shares network parameters between a teacher model guided by ground-truth "oracle" positions and text embeddings and a student model learning to match the teacher's focus. Concurrently, the data augmentation pipeline automatically synthesizes complex scenes by stitching together contextually matched image regions containing labeled interactions.

The findings establish that the proposed framework consistently outperforms existing state-of-the-art methods across all tested benchmarks. When applied to the baseline model, the framework improved overall detection accuracy by roughly 2.5 to 2.8 percentage points on major benchmarks and achieved a notable gain of nearly 5.0 percentage points on rare interaction categories. Furthermore, the distillation approach substantially accelerated model training convergence, requiring only about one-third of the training epochs of the standard baseline to achieve superior accuracy while leaving deployment inference costs completely unchanged. Ablation experiments also confirmed that preserving contextual consistency when creating synthesized images is essential, as removing context alignment led to measurable performance drops.

These results demonstrate that providing explicit spatial and semantic guidance during training overcomes the slow convergence typical of transformer models without sacrificing their global reasoning capabilities. In operational settings, this translates to lower training compute costs, faster model iteration cycles, and higher reliability in detecting uncommon interactions. Because DOQ and CCS are modular and portable, engineering teams can readily integrate them into existing transformer pipelines without rearchitecting deployed systems. A noted limitation is that the framework does not reduce the substantial runtime memory footprint inherent to transformer attention mechanisms. Future research and development should prioritize memory-efficient attention designs to make high-performing interaction models more feasible for resource-constrained edge hardware.

Cover for Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection

Abstract

Transformer-based methods have achieved great success in the field of human-object interaction (HOI) detection. However, these models tend to adopt semantically ambiguous queries, which lowers the transformer's representation learning power. Moreover, there are a very limited number of labeled human-object pairs for most images in existing datasets, which constrains the transformer's set prediction power. To handle the first problem, we propose an efficient knowledge distillation model, named Distillation using Oracle Queries (DOQ), which shares parameters between teacher and student networks. The teacher network adopts oracle queries that are semantically clear and generates high-quality decoder embeddings. By mimicking both the attention maps and decoder embeddings of the teacher network, the representation learning power of the student network is significantly promoted. To address the second problem, we introduce an efficient data augmentation method, named Context-Consistent Stitching (CCS), which generates complicated images online. Each new image is obtained by stitching labeled human-object pairs cropped from multiple training images. By selecting source images with similar context, the new synthesized image is made visually realistic. Our methods significantly promote both the accuracy and training efficiency of transformer-based HOI detection models. Experimental results show that our proposed approach consistently outperforms state-of-the-art methods on three benchmarks: HICO-DET, HOI-A, and V-COCO. Code is available at: https://github.com/SherlockHolmes221/DOQ.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methods
  • 3.1. QPIC Revisited
  • 3.2. Distillation Using Oracle Queries
  • 3.3. Context-Consistent Stitching
  • 4.2. Implementation Details
  • 4.3. Ablation Study
  • 4.4. Comparisons with Variants of DOQ and CCS
  • 4.5. Comparisons with State-of-the-Art Methods
  • 4.6. Qualitative Comparisons
  • 5. Conclusion and Limitations
  • References

Knowls

  1. Knowl 1 — Shared-parameter distillation with oracle queries

    model/method

    Distillation Using Oracle Queries (DOQ) augments a transformer-based HOI detector, instantiated with QPIC, with a training-only teacher network. For an image II, a CNN backbone produces visual features F∈RC×H×WF\in\mathbb{R}^{C\times H\times W}, positional encodings PP are added, and a transformer encoder produces EE. A decoder maps HOI queries QQ and initial decoder embeddings D0D_0 to output embeddings:

    D=fdec(Q,D0,E,P).D=f_{\mathrm{dec}}(Q,D_0,E,P).

    The student uses the original QPIC inputs: NqN_q image-independent learnable queries Q={qi∈Rd}i=1NqQ=\{q_i\in\mathbb{R}^d\}_{i=1}^{N_q} and zero initial embeddings D0D_0. The teacher instead uses ground-truth-derived oracle queries and object-semantic initial embeddings. The CNN backbone, transformer encoder and decoder, and four interaction heads are shared between teacher and student; the heads predict human boxes, object boxes, object categories, and interaction categories. The student is trained to imitate the teacher's decoder representations and cross-attention maps. The architecture diagram on page 3 depicts this parameter sharing and shows that the teacher is removed at inference, so DOQ adds no inference-time computation.

  2. Knowl 2 — Construction of oracle queries and semantic decoder inputs

    equation

    For each labeled human-object pair ii in a training image, DOQ forms a 12-dimensional spatial feature hi∈R12h_i\in\mathbb{R}^{12} from the human and object bounding boxes:

    hi=[xis,yis,wis,his,xio,yio,wio,hio,xis−xio,yis−yio,wishis,wiohio]T.h_i=[x_i^s,y_i^s,w_i^s,h_i^s,x_i^o,y_i^o,w_i^o,h_i^o,x_i^s-x_i^o,y_i^s-y_i^o,w_i^s h_i^s,w_i^o h_i^o]^\mathsf{T}.

    Here x,yx,y are box-center coordinates, w,hw,h are box widths and heights, superscripts ss and oo denote the human and object, and the final two entries are their box areas. For the NtqN_{tq} labeled pairs in an image, a two-layer ReLU feed-forward network Fq:R12→RdF_q:\mathbb{R}^{12}\rightarrow\mathbb{R}^{d} followed by elementwise tanh⁡\tanh produces the teacher's oracle queries:

    Qt=tanh⁡(Fq(Ht)),Ht={hi}i=1Ntq.Q_t=\tanh(F_q(H_t)),\qquad H_t=\{h_i\}_{i=1}^{N_{tq}}.

    For each pair, the teacher also receives the word embedding wi∈R512w_i\in\mathbb{R}^{512} of its ground-truth object category. A second two-layer ReLU feed-forward network FwF_w maps these embeddings to the decoder dimension:

    Dt0=Fw(Wt),Wt={wi}i=1Ntq.D_t^0=F_w(W_t),\qquad W_t=\{w_i\}_{i=1}^{N_{tq}}.

    The teacher decoder therefore computes Dt=fdec(Qt,Dt0,E,P)D_t=f_{\mathrm{dec}}(Q_t,D_t^0,E,P). The oracle queries provide pair-specific absolute and relative spatial information, while the initial decoder embeddings provide object-category semantics; the tanh⁡\tanh normalization makes the query amplitudes compatible with the positional encodings.

  3. Knowl 3 — Representation and attention-map distillation objective

    equation

    DOQ establishes teacher-student correspondences by bipartite matching the student's HOI predictions to ground-truth human-object pairs. The student decoder embeddings are reordered according to these matches and denoted Ds={dis}i=1NtqD_s=\{d_i^s\}_{i=1}^{N_{tq}}; each teacher embedding ditd_i^t already corresponds to the same ground-truth pair because its query and initial embedding are oracle inputs. The distillation loss is

    Ldis=α1Lcos+α2LKL,L_{\mathrm{dis}}=\alpha_1L_{\mathrm{cos}}+\alpha_2L_{\mathrm{KL}},

    where cosine-distance distillation aligns the final decoder embeddings:

    Lcos=1Ntq∑i=1Ntq(1−(dit)Tdis∥dit∥2∥dis∥2).L_{\mathrm{cos}}=\frac{1}{N_{tq}}\sum_{i=1}^{N_{tq}}\left(1-\frac{(d_i^t)^\mathsf{T}d_i^s}{\lVert d_i^t\rVert_2\lVert d_i^s\rVert_2}\right).

    Let At,ijA_{t,i}^j and As,ijA_{s,i}^j be the teacher and student cross-attention maps for matched pair ii at decoder layer jj, averaged over attention heads, and let ll be the number of decoder layers. The attention loss is applied to the last half of the layers:

    LKL=2Ntql∑j=l/2+1l∑i=1NtqAt,ij(ln⁡At,ij−ln⁡As,ij).L_{\mathrm{KL}}=\frac{2}{N_{tq}l}\sum_{j=l/2+1}^{l}\sum_{i=1}^{N_{tq}}A_{t,i}^j\left(\ln A_{t,i}^j-\ln A_{s,i}^j\right).

    The paper uses l=6l=6, α1=1\alpha_1=1, and α2=10\alpha_2=10. This supervision transfers both the teacher's pair-specific representations and its image-wide spatial attention patterns to the student.

  4. Knowl 4 — Context-Consistent Stitching augmentation

    algorithm

    Context-Consistent Stitching (CCS) creates training images containing more labeled human-object pairs while preserving scene compatibility. The paper's stitching examples and pair-count visualization on page 5 show the intended effect: several pair regions from visually related images are combined into one image, shifting the long-tailed pair-count distribution toward more crowded examples.

    Input: Training image I with labeled human-object pairs; replacement probability gamma; neighbor count K
    Output: Training image and automatically transformed HOI annotations
    With probability gamma, leave I unchanged; otherwise continue
    Use an off-the-shelf scene classification model to compute scene features for all training images
    Use offline scene-feature distances to obtain the K nearest neighbors of I
    Randomly sample three neighbors of I
    For each of I and the three sampled neighbors:
        Randomly select one labeled human-object pair
        Crop the union bounding region of the selected pair
        If the selected region overlaps another labeled pair in the same source image:
            Extend the crop to include every overlapping pair
    Tightly stitch the four cropped regions together
    Resize the stitched image to a size similar to I
    Transform each included pair's boxes and interaction annotations to its new image coordinates
    Return the stitched image and transformed annotations

    The implementation uses K=15K=15 and γ=0.25\gamma=0.25. Unlike ordinary copy-paste augmentation, CCS selects source images with similar scene context, uses regions from multiple images to increase pair diversity, and deliberately increases the number of labeled human-object pairs in the synthesized image.

  5. Knowl 5 — Joint teacher-student training and inference objective

    model/method

    DOQ trains the teacher and student jointly rather than pretraining a separate larger teacher. For k∈{t,s}k\in\{t,s\}, the supervised loss of network kk is

    Lk=λbLkb+λuLku+λcLkc+λaLka,L_k=\lambda_bL_{kb}+\lambda_uL_{ku}+\lambda_cL_{kc}+\lambda_aL_{ka},

    where LkbL_{kb} is the human/object box L1L_1 regression loss, LkuL_{ku} is the generalized-IoU box loss, LkcL_{kc} is object-category cross-entropy, and LkaL_{ka} is interaction-category focal loss. The complete training objective is

    L=Lt+Ls+Ldis,L=L_t+L_s+L_{\mathrm{dis}},

    with λb=2.5\lambda_b=2.5 and λu=λc=λa=1\lambda_u=\lambda_c=\lambda_a=1. The teacher receives ground-truth pair geometry and object-word embeddings, whereas the student receives learnable queries and zero initial embeddings. During inference only the student-side shared network and detection heads are evaluated; oracle inputs, teacher activations, and distillation losses are absent.

  6. Knowl 6 — Training configuration and evaluation benchmarks

    experimental setup

    The experiments evaluate DOQ and CCS on HICO-DET, HOI-A, and V-COCO. HICO-DET has 47,776 images (38,118 training and 9,658 test), 80 object categories, 117 interaction categories, and 600 HOI categories, including 138 rare categories with fewer than 10 training examples; it is evaluated by mAP in Default and Known-Object modes. HOI-A has 38,629 images (29,842 training and 8,787 test), 11 object categories, and 10 interaction categories, and uses mAP. V-COCO has 10,346 images (5,400 training and 4,946 test), 80 object categories, and 26 interaction categories, and uses Scenario-1 mAProle\mathrm{mAP}_{\mathrm{role}}.

    The implementation uses ResNet-50 or ResNet-101 backbones, AdamW, batch size 16 across 8 GPUs, initial learning rate 10−410^{-4}, a tenfold learning-rate reduction after 60 epochs, and 80 total epochs. The detector is initialized from DETR trained on MS-COCO. The number of student queries is Nq=100N_q=100, decoder dimension is d=256d=256, and CLIP object-word embeddings have dimension 512. Unless otherwise specified, ablations use ResNet-50; DOQ uses α1=1\alpha_1=1 and α2=10\alpha_2=10, while CCS uses K=15K=15 and γ=0.25\gamma=0.25.

  7. Knowl 7 — Component ablations on HICO-DET and HOI-A

    data/table

    The ablation study uses QPIC with a ResNet-50 backbone and reports mAP on HICO-DET Default Mode and HOI-A. It separates parameter sharing, embedding distillation, attention-map distillation, and CCS. The results show that each component contributes, while the full combination is strongest.

    Could not parse LaTeX table

    Here MTL denotes the shared teacher-student architecture without the distillation loss. Relative to QPIC, CCS alone improves mAP by 1.69 points on HICO-DET and 1.35 points on HOI-A; the complete method reaches 31.55 and 76.87, respectively.

  8. Knowl 8 — HICO-DET performance and portability to other detectors

    data/table

    In HICO-DET Default Mode, DOQ and CCS improve several transformer-based HOI detectors. The table reports mAP for all, rare, and non-rare HOI categories; all values are percentages as reported by the paper.

    Could not parse LaTeX table

    With ResNet-50, the complete method improves QPIC by 2.48 mAP points overall, 4.90 points on rare categories, and 1.76 points on non-rare categories. Applying the same plug-and-play components to HOTR and CDN-S improves their full-category scores by 2.51 and 1.84 points, respectively, demonstrating that the method is not tied to QPIC's architecture.

  9. Knowl 9 — Cross-dataset accuracy and convergence analysis

    empirical result

    The complete method with QPIC obtains 76.87 mAP on HOI-A, compared with 74.10 for QPIC using the same ResNet-50 backbone, a gain of 2.77 points. On V-COCO Scenario 1, it obtains 63.5 mAProle63.5\ \mathrm{mAP}_{\mathrm{role}}, exceeding QPIC's 58.8 and CDN-S's 61.7.

    The spatial-prior comparison on HICO-DET Default Mode shows that oracle-query supervision is also more effective for HOI detection than estimated-location priors:

    Could not parse LaTeX table

    The convergence plot on page 1 and the spatial-prior results show that DOQ reaches higher accuracy with fewer training epochs. The qualitative attention visualization on page 8 further shows that DOQ's decoder maps highlight relevant pixels across the whole image, rather than concentrating only on estimated human-object locations; this is consistent with the method's emphasis on contextual cues.

  10. Knowl 10 — Remaining computational limitation

    limitation

    The paper does not reduce the memory cost of the self-attention and cross-attention operations in the underlying transformer. DOQ removes the teacher during inference, so it avoids additional teacher computation at test time, but the shared student transformer still has the original attention-memory burden. The authors identify a less memory-intensive transformer-based HOI detector as future work.

Coverage note — The exhaustive literature-comparison rows, supplementary CCS probability sweep, and additional qualitative examples were omitted because the key quantitative comparisons, ablations, convergence analysis, and stated limitation are already represented.

References

  1. 1.H. Zhao, R. Wildes. Spatiotemporal feature residual propagation for action prediction. In ICCV, 2019. 1
  2. 2.Y. Kong, Z. Tao, Y. Fu. Deep sequential context networks for action prediction. In CVPR, 2017 1
  3. 3.X. Lin, C. Ding, J. Zeng, D. Tao. Gps-net: Graph property sensing network for scene graph generation. In CVPR, 2020. 1
  4. 4.M. Suhail, A. Mittal, B. Siddiquie, C. Broaddus, J. Eledath, G. Medioni, L. Sigal. Energy-Based Learning for Scene Graph Generation. In CVPR, 2021. 1
  5. 5.K. Tang, H. Zhang, B. Wu, W. Luo, W. Liu. Learning to compose dynamic tree structures for visual contexts. In CVPR, 2019. 1
  6. 6.L. Chen, X. Yan, J. Xiao, H. Zhang, S. Pu, Y. Zhuang. Counterfactual samples synthesizing for robust visual question answering. In CVPR, 2020. 1
  7. 7.J. Yim, D. Joo, J. Bae, J. Kim. A gift from knowledge distillation: Fast optimization, network minimization and transfer learning. In CVPR, 2017. 5
  8. 8.F. Tung, G. Mori. Similarity-preserving knowledge distillation. In ICCV, 2019. 5
  9. 9.P. Chen, S. Liu, H. Zhao, J. Jia. Distilling Knowledge via Knowledge Review. In CVPR, 2021. 5
  10. 10.X. Zhu, W. Su, L. Lu, B. Li, X. Wang, J. Dai. Deformable DETR: Deformable Transformers for End-to-End Object Detection. In ICLR, 2020. 3
  11. 11.N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, S. Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 2, 3, 4, 6
  12. 12.P. Gao, M. Zheng, X. Wang, J. Dai, H. Li. Fast Convergence of DETR With Spatially Modulated Co-Attention. In ICCV, 2021. 3, 8
  13. 13.D. Meng, X. Chen, Z. Fan, G. Zeng, H. Li, Y. Yuan, L. Sun, J. Wang. Conditional DETR for Fast Training Convergence. In ICCV, 2021. 3, 4
  14. 14.X. Dai, Y. Chen, J. Yang, P. Zhang, L. Yuan, L. Zhang. Dynamic DETR: End-to-End Object Detection With Dynamic Attention. In ICCV, 2021. 3, 8
  15. 15.X. Jiang, Z. Chen, Z. Wang, E. Zhou, C. Yuan. Guiding Query Position and Performing Similar Attention for Transformer-Based Detection Heads. arXiv:2108.09691, 2021. 3
  16. 16.C. Gao, Y. Zou, J. Huang. iCAN: Instance-Centric Attention Network for Human-Object Interaction Detection. In BMVC, 2018. 7
  17. 17.Y. Li, S. Zhou, X. Huang, L. Xu, Z. Ma, H. Fang, Y. Wang, C. Lu. Transferable interactiveness knowledge for human-object interaction detection. In CVPR, 2019. 7
  18. 18.T. He, L. Gao, J. Song, Y. Li. Exploiting Scene Graphs for Human-Object Interaction Detection. In ICCV, 2021. 7, 8
  19. 19.Q. Dong, Z. Tu, H. Liao, Y. Zhang, V. Mahadevan, S. Soatto. Visual Relationship Detection Using Part-and-Sum Transformers With Composite Queries. In ICCV, 2021. 2, 7
  20. 20.F. Zhang, D. Campbell, S. Gould. Spatially Conditioned Graphs for Detecting Human-Object Interactions. In ICCV, 2021. 7, 8
  21. 21.S. Wang, K. Yap, H. Ding, J. Wu, J. Yuan, Y. Tan. Discovering Human Interactions With Large-Vocabulary Objects via Query and Multi-Scale Detection. In ICCV, 2021. 2
  22. 22.X. Zhong, X. Qu, C. Ding, D. Tao. Glance and Gaze: Inferring Action-aware Points for One-Stage Human-Object Interaction Detection. In CVPR, 2021. 2, 7, 8
  23. 23.M. Tamura, H. Ohashi, T. Yoshinaga. QPIC: Query-Based Pairwise Human-Object Interaction Detection with Image-Wide Contextual Information. In CVPR, 2021. 1, 2, 3, 4, 5, 6, 7, 8
  24. 24.M. Chen, Y. Liao, S. Liu, Z. Chen, F. Wang, C. Qian. Reformulating hoi detection as adaptive set prediction. In CVPR, 2021. 1, 2, 3, 7, 8
  25. 25.C. Zou, B. Wang, Y. Hu, J. Liu, Q. Wu, Y. Zhao, B. Li, C. Zhang, C. Zhang, Y. Wei, J. Sun. End-to-End Human Object Interaction Detection with HOI Transformer. In CVPR, 2021. 1, 2, 3, 4, 7, 8
  26. 26.B. Kim, J. Lee, J. Kang, E. Kim, H. Kim. HOTR: End-to-End Human-Object Interaction Detection with Transformers. In CVPR, 2021. 1, 2, 3, 4, 6, 7, 8
  27. 27.Z. Hou, B. Yu, Y. Qiao, X. Peng, D. Tao. Detecting human-object interaction via fabricated compositional learning. In CVPR, 2021. 8
  28. 28.Z. Hou, B. Yu, Y. Qiao, X. Peng, D. Tao. Affordance Transfer Learning for Human-Object Interaction Detection. In CVPR, 2021. 2
  29. 29.Y. Liao, S. Liu, F. Wang, Y. Chen, C. Qian, J. Feng. Ppdm: Parallel point detection and matching for real-time human-object interaction detection. In CVPR, 2020. 2, 5, 6, 7
  30. 30.T. Wang, T. Yang, M. Danelljan, F. Khan, X. Zhang, J. Sun. Learning human-object interaction detection using interaction points. In CVPR, 2020. 2, 7
  31. 31.T. Zhou, W. Wang, S. Qi, H. Ling, J. Shen. Cascaded human-object interaction recognition. In CVPR, 2020. 2, 7
  32. 32.O. Ulutan, A. Iftekhar, B. Manjunath. Vsgnet: Spatial attention network for detecting human object interactions using graph convolutions. In CVPR, 2020. 2, 8
  33. 33.Y. Li, X. Liu, H. Lu, S. Wang, J. Liu, J. Li, C. Lu. Detailed 2d-3d joint representation for human-object interaction. In CVPR, 2020. 2, 7
  34. 34.S. Wang, K. Yap, J. Yuan, Y. Tan. Discovering human interactions with novel objects via zero-shot learning. In CVPR, 2020. 2
  35. 35.Y. Li, L. Xu, X. Liu, X. Huang, Y. Xu, S. Wang, H. Fang, Z. Ma, M. Chen, C. Lu. Pastanet: Toward human activity knowledge engine. In CVPR, 2020. 2, 7
  36. 36.Y. Li, X. Liu, X. Wu, Y. Li, C. Lu. HOI Analysis: Integrating and Decomposing Human-Object Interaction. In NeurIPS, 2020. 7
  37. 37.A. Zhang, Y. Liao, S. Liu, M. Lu, Y. Wang, C. Gao, X. Li. Mining the Benefits of Two-stage and One-stage HOI Detection. In NeurIPS, 2021. 3, 6, 7, 8
  38. 38.H. Fang, Y. Xie, D. Shao, C. Lu. DIRV: Dense Interaction Region Voting for End-to-End Human-Object Interaction Detection. In AAAI, 2021. 8
  39. 39.Y. Liu, J. Yuan, C. Chen. Consnet: Learning consistency graph for zero-shot human-object interaction detection. In ACM MM, 2020. 2, 7
  40. 40.B. Kim, T. Choi, J. Kang, H. Kim. Uniondet: Union-level detector towards real-time human-object interaction detection. In ECCV, 2020. 2
  41. 41.Z. Hou, X. Peng, Y. Qiao, D. Tao. Visual compositional learning for human-object interaction detection. In ECCV, 2020. 2
  42. 42.X. Zhong, C. Ding, X. Qu, D. Tao. Polysemy deciphering network for human-object interaction detection. In ECCV, 2020. 2, 4
  43. 43.X. Zhong, C. Ding, X. Qu, D. Tao. Polysemy Deciphering Network for Robust Human–Object Interaction Detection. In IJCV, 2021. 2
  44. 44.H. Wang, W. Zheng, L. Yingbiao. Contextual heterogeneous graph network for human-object interaction detection. In ECCV, 2020. 2
  45. 45.Y. Liu, Q. Chen, A. Zisserman. Amplifying key cues for human-object-interaction detection. In ECCV, 2020. 2, 8
  46. 46.D. Kim, X. Sun, J. Choi, S. Lin, I. Kweon. Detecting human-object interactions with action co-occurrence priors. In ECCV, 2020. 2
  47. 47.C. Gao, J. Xu, Y. Zou, J. Huang. Drg: Dual relation graph for human-object interaction detection. In ECCV, 2020. 7
  48. 48.C. Yu-Wei, L. Yunfan, L. Xieyang, Z. Huayi, D. Jia. Learning to Detect Human-Object Interactions. In WACV, 2018. 1, 2, 5, 6
  49. 49.S. Gupta, J. Malik. Visual Semantic Role Labeling. arXiv:1505.04474, 2015. 2, 6
  50. 50.T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, Microsoft COCO: Common objects in context. In ECCV, 2014. 6
  51. 51.K. He, X. Zhang, S. Ren, J. Sun. Deep residual learning for image recognition. In CVPR, 2016. 6
  52. 52.I. Loshchilov, F. Hutter. Decoupled weight decay regularization. In ICLR, 2019. 6
  53. 53.T. Mikolov, T. Sutskever, K. Chen, G. Corrado, J. Dean. Distributed representations of words and phrases and their compositionality. In NeurIPS, 2013. 7
  54. 54.A. Radford, J. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger. Learning transferable visual models from natural language supervision. In ICML, 2021. 6, 7
  55. 55.B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, A. Torralba. Places: A 10 million Image Database for Scene Recognition. In IEEE TPAMI, 2017. 5
  56. 56.G. Ghiasi, Y. Cui, A. Srinivas, R. Qian, T. Lin, E. Cubuk, Q. Le, B. Zoph. Simple Copy-Paste Is a Strong Data Augmentation Method for Instance Segmentation. In CVPR, 2021. 5
  57. 57.D. Dwibedi, I. Misra, M. Hebert. Cut, Paste and Learn: Surprisingly Easy Synthesis for Instance Detection. In ICCV, 2017. 5
  58. 58.N. Dvornik, J. Mairal, C. Schmid. Modeling Visual Context is Key to Augmenting Object Detection Datasets. In ECCV, 2018. 5
  59. 59.H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, S. Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In CVPR, 2019. 5
  60. 60.T. Lin, P. Goyal, R. Girshick, K. He, P. Dollar. Focal loss for dense object detection. In ICCV, 2017. 5
  61. 61.H. Kuhn. The Hungarian method for the assignment problem. In Naval research logistics quarterly, 1955. 4
  62. 62.Pic leaderboard. http://www.picdataset.com/challenge/leaderboard/hoi2019, 2019. 7

Citation

MLA
Qu, X., et al. “Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 19536–45, https://doi.org/10.1109/CVPR52688.2022.01895.
APA
Qu, X., Ding, C., Li, X., Zhong, X., & Tao, D. (2022). Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19536–19545. https://doi.org/10.1109/CVPR52688.2022.01895
Chicago
Qu, X., C. Ding, X. Li, X. Zhong, and D. Tao. 2022. “Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19536–45. https://doi.org/10.1109/CVPR52688.2022.01895.
Harvard
Qu, X. et al. (2022) “Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 19536–19545. Available at: https://doi.org/10.1109/CVPR52688.2022.01895.
Vancouver
1. Qu X, Ding C, Li X, Zhong X, Tao D (2022) Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 19536–19545

BibTeX

@inproceedings{Qu_2022, title={Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection}, url={http://dx.doi.org/10.1109/CVPR52688.2022.01895}, DOI={10.1109/cvpr52688.2022.01895}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Qu, Xian and Ding, Changxing and Li, Xingao and Zhong, Xubin and Tao, Dacheng}, year={2022}, month=June, pages={19536–19545} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE