Enhanced Training of Query-Based Object Detection via Selective Query Recollection
Fangyi ChenHan ZhangKai HuYu-Kai HuangChenchen ZhuMarios Savvides
Proposes Selective Query Recollection, a training strategy that routes intermediate queries directly to later decoder stages to prevent cascading errors and boost detection accuracy across query-based detectors without increasing inference cost.
Modern artificial intelligence systems for visual object detection increasingly rely on query-based transformer models. These models process image data through multiple sequential decoding stages to iteratively refine object locations and categories. While the final stage is assumed to produce the most accurate predictions, empirical analysis reveals a critical flaw: models frequently generate their best predictions at intermediate stages and degrade them by the final stage. In fact, true positive predictions degrade between 27% and 51% of the time, while false alarms worsen in over 50% of evaluated cases. Replacing final predictions with optimal intermediate outputs reveals that existing architectures leave 7.2 to 10.7 average precision points unrealized due to two root causes: standard training distributes supervision uniformly rather than emphasizing the critical later stages, and sequential pipelines cause errors from earlier stages to cascade forward.
The article develops and evaluates a targeted training framework called Selective Query Recollection to overcome these structural issues. Instead of passing representations strictly one stage at a time, the framework stores intermediate representations and selectively routes outputs from the preceding two stages directly into downstream stages during training. This architecture naturally structures supervision so that later stages receive progressively more feedback following a Fibonacci sequence, while insulating downstream layers from single-stage cascading errors. Because this mechanism operates strictly as a training strategy, the final runtime inference pipeline and operational latency remain completely unchanged.
Evaluation on the standard Microsoft Common Objects in Context benchmark across major query-based architectures—including Adamixer, DAB-DETR, and Deformable-DETR—demonstrates consistent improvements of 1.4 to 2.8 average precision points across diverse backbones and training schedules. Comparative testing shows that selectively recycling representations from the two nearest preceding stages outperforms indiscriminately collecting all past representations, while cutting the additional computational overhead in half. Furthermore, the performance gains cannot be replicated simply by adding extra query groups, re-weighting loss functions, or training baseline models for longer schedules.
These findings establish that training-stage supervision balance and non-sequential query routing are crucial for query-based vision models. For engineering and deployment teams, adopting Selective Query Recollection delivers substantial accuracy gains at zero runtime latency or hardware inference cost. The primary operational trade-off is an increase in initial model training time, ranging from under 10% on large vision backbones up to 57% on lighter networks. Organizations training transformer-based object detectors should integrate Selective Query Recollection into their development pipelines, using the second stage as an optimal starting configuration to balance computational efficiency with model accuracy.
- Paper: DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR, Shilong Liu et al. (2022). Introduces explicit dynamic anchor box formulation for DETR queries, serving as one of the primary query-based detector baselines improved by SQR.
- Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). Presents multi-scale deformable attention for DETR decoding stages, providing a foundational baseline architecture tested directly in this paper.
- Paper: End-to-End Object Detection with Transformers, Nicolas Carion et al. (2020). Establishes the foundational transformer-based end-to-end set prediction framework using object queries that this work fundamentally seeks to enhance.
- Paper: Sparse R-CNN: End-to-End Object Detection with Learnable Proposals, Peize Sun et al. (2020). Develops iterative stage-by-stage query and proposal refinement in end-to-end detection, which exemplifies the multi-stage decoding dynamics examined by SQR.
- Paper: Cascade R-CNN: Delving Into High Quality Object Detection, Zhaowei Cai et al. (2017). Pioneers the multi-stage cascaded decoding paradigm in object detection where successive stages refine intermediate hypothesis representations.
- Paper: DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection, Hao Zhang et al. (2023). Builds on query-based transformer detectors by introducing contrastive denoising and improved query initialization to accelerate training convergence and accuracy.
- Paper: DETRs with Hybrid Matching, Ding Jia et al. (2023). Addresses positive sample training deficiency in DETR decoders through a hybrid matching scheme to enhance query supervision.
- Paper: DETRs Beat YOLOs on Real-time Object Detection, Yian Zhao et al. (2024). Extends query selection and decoder efficiency to develop a real-time end-to-end Transformer detector.
