Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection
Xian QuChangxing DingXingao LiXubin ZhongDacheng Tao
Proposes a knowledge distillation framework that guides transformer decoders using ground-truth oracle queries alongside a context-consistent image stitching augmentation, substantially accelerating training convergence and setting state-of-the-art accuracy in human-object interaction detection without adding inference overhead.
Human-Object Interaction detection aims to identify people, the objects they interact with, and the specific actions linking them within an image. This capability is critical for computer vision applications such as scene understanding, healthcare monitoring, and autonomous systems. While modern transformer-based architectures have improved performance by capturing global visual context, they face two key bottlenecks: ambiguous queries that slow training and degrade feature representation, and a scarcity of training images containing multiple labeled interaction pairs, which restricts the model's ability to process complex visual scenes.
The article develops and evaluates two complementary solutions: a knowledge distillation framework termed Distillation using Oracle Queries (DOQ) and an online data augmentation technique called Context-Consistent Stitching (CCS). Together, these methods aim to improve the accuracy, training speed, and scalability of transformer-based interaction detection without adding computational overhead during practical deployment.
The evaluation was conducted using standardized computer vision benchmarks, including HICO-DET (over 47,000 images), HOI-A (over 38,000 images), and V-COCO (over 10,000 images). The authors paired the proposed training framework with standard baseline architectures like QPIC, HOTR, and CDN across multiple convolutional backbones. The approach shares network parameters between a teacher model guided by ground-truth "oracle" positions and text embeddings and a student model learning to match the teacher's focus. Concurrently, the data augmentation pipeline automatically synthesizes complex scenes by stitching together contextually matched image regions containing labeled interactions.
The findings establish that the proposed framework consistently outperforms existing state-of-the-art methods across all tested benchmarks. When applied to the baseline model, the framework improved overall detection accuracy by roughly 2.5 to 2.8 percentage points on major benchmarks and achieved a notable gain of nearly 5.0 percentage points on rare interaction categories. Furthermore, the distillation approach substantially accelerated model training convergence, requiring only about one-third of the training epochs of the standard baseline to achieve superior accuracy while leaving deployment inference costs completely unchanged. Ablation experiments also confirmed that preserving contextual consistency when creating synthesized images is essential, as removing context alignment led to measurable performance drops.
These results demonstrate that providing explicit spatial and semantic guidance during training overcomes the slow convergence typical of transformer models without sacrificing their global reasoning capabilities. In operational settings, this translates to lower training compute costs, faster model iteration cycles, and higher reliability in detecting uncommon interactions. Because DOQ and CCS are modular and portable, engineering teams can readily integrate them into existing transformer pipelines without rearchitecting deployed systems. A noted limitation is that the framework does not reduce the substantial runtime memory footprint inherent to transformer attention mechanisms. Future research and development should prioritize memory-efficient attention designs to make high-performing interaction models more feasible for resource-constrained edge hardware.
- Paper: End-to-End Object Detection with Transformers, Nicolas Carion et al. (2020). This foundational work establishes the query-based transformer architecture and Hungarian matching for set prediction that DOQ adapts and distills for human-object interaction detection.
- Paper: DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR, Shilong Liu et al. (2022). It provides critical background on addressing query ambiguity and slow convergence in detection transformers by refining decoder queries.
- Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). It introduces deformable attention mechanisms that underpin efficient multi-scale query processing in modern transformer-based detection architectures.
- Paper: Relational Knowledge Distillation, Wonpyo Park et al. (2019). It outlines relational knowledge distillation principles that motivate DOQ's cross-network embedding and attention map mimicry.
- Paper: Relation Networks for Object Detection, Han Hu et al. (2017). It introduces relation modeling for detecting interactions among co-occurring visual entities, establishing the core problem setting DOQ addresses.
- Paper: Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models, Yichao Cao et al. (2023). UniHOI extends transformer-based human-object interaction detection beyond closed-set paired supervision to open-world scenarios using vision-language foundation models and spatial prompts.
- Paper: DETRs with Hybrid Matching, Ding Jia et al. (2023). H-DETR further explores query-matching supervision by introducing auxiliary one-to-many matching branches to improve transformer representation learning.
- Paper: Enhanced Training of Query-Based Object Detection via Selective Query Recollection, Fangyi Chen et al. (2023). This work advances query representation learning by addressing intermediate decoding mispredictions via selective query recollection.
- Paper: DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection, Hao Zhang et al. (2023). DINO generalizes transformer query refinement with denoising anchor boxes, advancing the broader detection framework used in set-prediction models.
