Enhanced Training of Query-Based Object Detection via Selective Query Recollection

Fangyi ChenHan ZhangKai HuYu-Kai HuangChenchen ZhuMarios Savvides

article2023CVPR64 citations

Proposes Selective Query Recollection, a training strategy that routes intermediate queries directly to later decoder stages to prevent cascading errors and boost detection accuracy across query-based detectors without increasing inference cost.

Listen

Modern artificial intelligence systems for visual object detection increasingly rely on query-based transformer models. These models process image data through multiple sequential decoding stages to iteratively refine object locations and categories. While the final stage is assumed to produce the most accurate predictions, empirical analysis reveals a critical flaw: models frequently generate their best predictions at intermediate stages and degrade them by the final stage. In fact, true positive predictions degrade between 27% and 51% of the time, while false alarms worsen in over 50% of evaluated cases. Replacing final predictions with optimal intermediate outputs reveals that existing architectures leave 7.2 to 10.7 average precision points unrealized due to two root causes: standard training distributes supervision uniformly rather than emphasizing the critical later stages, and sequential pipelines cause errors from earlier stages to cascade forward.

The article develops and evaluates a targeted training framework called Selective Query Recollection to overcome these structural issues. Instead of passing representations strictly one stage at a time, the framework stores intermediate representations and selectively routes outputs from the preceding two stages directly into downstream stages during training. This architecture naturally structures supervision so that later stages receive progressively more feedback following a Fibonacci sequence, while insulating downstream layers from single-stage cascading errors. Because this mechanism operates strictly as a training strategy, the final runtime inference pipeline and operational latency remain completely unchanged.

Evaluation on the standard Microsoft Common Objects in Context benchmark across major query-based architectures—including Adamixer, DAB-DETR, and Deformable-DETR—demonstrates consistent improvements of 1.4 to 2.8 average precision points across diverse backbones and training schedules. Comparative testing shows that selectively recycling representations from the two nearest preceding stages outperforms indiscriminately collecting all past representations, while cutting the additional computational overhead in half. Furthermore, the performance gains cannot be replicated simply by adding extra query groups, re-weighting loss functions, or training baseline models for longer schedules.

These findings establish that training-stage supervision balance and non-sequential query routing are crucial for query-based vision models. For engineering and deployment teams, adopting Selective Query Recollection delivers substantial accuracy gains at zero runtime latency or hardware inference cost. The primary operational trade-off is an increase in initial model training time, ranging from under 10% on large vision backbones up to 57% on lighter networks. Organizations training transformer-based object detectors should integrate Selective Query Recollection into their development pipelines, using the second stage as an optimal starting configuration to balance computational efficiency with model accuracy.

arXiv: 2212.07593
Cover for Enhanced Training of Query-Based Object Detection via Selective Query Recollection

Abstract

This paper investigates a phenomenon where query-based object detectors mispredict at the last decoding stage while predicting correctly at an intermediate stage. We review the training process and attribute the overlooked phenomenon to two limitations: lack of training emphasis and cascading errors from decoding sequence. We design and present Selective Query Recollection (SQR), a simple and effective training strategy for query-based object detectors. It cumulatively collects intermediate queries as decoding stages go deeper and selectively forwards the queries to the downstream stages aside from the sequential structure. Such-wise, SQR places training emphasis on later stages and allows later stages to work with intermediate queries from earlier stages directly. SQR can be easily plugged into various query-based object detectors and significantly enhances their performance while leaving the inference pipeline unchanged. As a result, we apply SQR on Adamixer, DAB-DETR, and Deformable-DETR across various settings (backbone, number of queries, schedule) and consistently brings 1.4 ~ 2.8 AP improvement. Code is available at https://github.com/Fangyi-Chen/SQR

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Training Strategy for Object Detection
  • 2.2. Query-Based Object Detection
  • 3. Motivation
  • 4. Query Recollection
  • 4.1. Expectancy
  • 4.2. Dense Query Recollection
  • 4.3. Selective Query Recollection
  • 5. Experiments
  • 5.1. Ablation Study
  • 5.2. Relation with Increased Number of Supervision
  • 5.3. Training Emphasis via Re-weighted Loss
  • 5.4. Relation with Stochastic Depth
  • 5.5. Training efficiency
  • 5.6. Comparison with State-of-the-art
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Intermediate decoding stages can outperform the final stage

    empirical result

    A query-based detector starts with learned queries qi0q_i^0 and updates them through SS sequential decoding stages. For query index ii, stage ss produces

    qis=Ds(qis−1,qs−1,x)=qis−1+(A∘F)(qis−1,qs−1,x),q_i^s=D^s(q_i^{s-1},q^{s-1},x)=q_i^{s-1}+ (A\circ F)(q_i^{s-1},q^{s-1},x),

    where qs={qis}i=1nq^s=\{q_i^s\}_{i=1}^n is the set of nn queries at stage ss, xx denotes encoded image features, DsD^s is the stage-ss decoder, and A∘FA\circ F denotes the self-attention, cross-attention, and feed-forward operations. Each query predicts a class and bounding box through classification and regression heads:

    Pis=(MLP⁡cls(qis),MLP⁡reg(qis)).P_i^s=\bigl(\operatorname{MLP}_{\mathrm{cls}}(q_i^s),\operatorname{MLP}_{\mathrm{reg}}(q_i^s)\bigr).

    Although average precision generally increases with depth, the paper finds that the last-stage prediction is not always the best prediction for an individual query. A prediction is counted as a true positive for a ground-truth object when its category is correct, its category score is highest among competing predictions, and its IoU with the object exceeds the selected threshold. If the final-stage prediction Pi6P_i^6 is a true positive, the true-positive fading rate measures how often an earlier prediction for the same query is both higher-IoU and higher-confidence; if Pi6P_i^6 is a false positive, the false-positive exacerbation rate measures how often an earlier prediction is also false positive but has a lower category score.

    On COCO validation, the stage-wise AP values for six-stage Deformable DETR with 300 queries were 38.4,42.2,43.7,44.2,44.4,44.538.4,42.2,43.7,44.2,44.4,44.5 at stages 1 through 6, while Adamixer with 100 queries obtained 15.1,30.3,37.7,40.6,42.1,42.515.1,30.3,37.7,40.6,42.1,42.5. Nevertheless, the final-stage error statistics were substantial: Deformable DETR had true-positive fading rates of 51.4%51.4\% and 49.5%49.5\% at IoU thresholds 0.500.50 and 0.750.75, with false-positive exacerbation rates of 55.7%55.7\% and 55.9%55.9\%; Adamixer had corresponding rates of 28.6%28.6\%, 26.7%26.7\%, 50.8%50.8\%, and 51.2%51.2\%. The page-2 stage-wise visualization illustrates the effect: a traffic-light confidence falls from 0.410.41 at stage 1 to 0.210.21 at stage 5, and a remote misclassified as a cell phone rises in confidence from 0.260.26 to 0.420.42 across later stages.

    Replacing each final-stage prediction by the best prediction among stages 1 through 6 raises AP from 44.544.5 to 51.751.7 for Deformable DETR and from 42.542.5 to 53.353.3 for Adamixer, showing the potential lost by forcing the final stage to be the sole output.

  2. Knowl 2 — Selective Query Recollection addresses unequal stage responsibility and cascading errors

    model/method

    The paper identifies two training limitations in sequential query decoders. First, every decoding stage receives analogous supervision even though later stages are responsible for the deployed prediction; this under-emphasizes later stages. Second, every query refinement is passed to the next stage, so a harmful intermediate refinement is cascaded while an earlier, better query is discarded. Query Recollection (QR) changes training by retaining intermediate query states and independently applying matching and losses to multiple query pathways at downstream stages.

    QR is training-only: during inference, the detector retains the original sequential pathway from the initial query through all decoder stages. Thus, QR is intended to increase later-stage supervision and expose later stages to queries produced at earlier points without changing the number of inference queries, decoder stages, or inference computation.

  3. Knowl 3 — Dense Query Recollection creates exponentially increasing supervision

    algorithm

    Dense Query Recollection (DQR) collects every intermediate query and sends every collected query through every subsequent decoder stage. Let q0q^0 be the initial set of queries and let DsD^s be the transformation performed by decoder stage ss. The collection after stage ss is

    C0={q0},C^0=\{q^0\}, Cs={Ds(q):q∈Cs−1}∪Cs−1.C^s=\{D^s(q):q\in C^{s-1}\}\cup C^{s-1}.

    At stage ss, each newly transformed query in {Ds(q):q∈Cs−1}\{D^s(q):q\in C^{s-1}\} is independently matched to the ground truth with Hungarian assignment and receives the detector loss; retained queries are carried forward but are not counted as newly generated supervision at that stage. Consequently, the collection size doubles at every stage, and the number of newly supervised query pathways across stages 1 through 6 is (1,2,4,8,16,32)(1,2,4,8,16,32). Every later stage can directly process queries originating from all earlier stages.

    During inference, DQR is disabled and only the ordinary six-stage pathway is used. DQR therefore changes training supervision and connectivity while leaving the inference pipeline unchanged.

  4. Knowl 4 — Selective Query Recollection uses nearby intermediate queries with Fibonacci supervision

    algorithm

    Selective Query Recollection (SQR) reduces DQR's cost by retaining only query pathways from the two nearest preceding stages. With CsC^s denoting the query collection after stage ss and DsD^s the transformation of decoder stage ss, SQR initializes

    C0={q0},C1={q0,q0−1},C^0=\{q^0\},\qquad C^1=\{q^0,q^{0-1}\},

    and for s≥2s\ge 2 forms

    Cs={Ds(q):q∈Cs−1}∪{Ds−1(q):q∈Cs−2}.C^s=\{D^s(q):q\in C^{s-1}\}\cup\{D^{s-1}(q):q\in C^{s-2}\}.

    The first term processes the previous collection through the current stage, while the second term injects queries produced by the immediately preceding stage from the collection before it. Each newly produced query pathway receives an independent Hungarian assignment and detection loss. The number of supervised pathways at stages 1 through 6 is therefore (1,2,3,5,8,13)(1,2,3,5,8,13), a Fibonacci sequence, rather than DQR's (1,2,4,8,16,32)(1,2,4,8,16,32).

    SQR can start recollection at a later stage to reduce cost. Starting at stage 2 gives supervision counts (1,1,2,3,5,8)(1,1,2,3,5,8), and starting at stage 3 gives (1,1,1,2,3,5)(1,1,1,2,3,5). The starting stage is a training hyperparameter. In all variants, inference uses only the original sequential pathway.

  5. Knowl 5 — Adjacent stages provide most useful alternative predictions for recollection

    empirical result

    The paper selects the two preceding stages for SQR by measuring which intermediate stages most often improve on the final-stage prediction. On Adamixer, the individual-stage true-positive fading and false-positive exacerbation rates were:

    stage 1 2 3 4 5
    TP fading rate (%) 1.2 4.4 8.5 12.4 16.9
    FP exacerbation rate (%) 14.3 18.6 20.8 24.5 30.2

    For groups of stages, the rates were 11.2%11.2\% and 32.4%32.4\% for stages 1--3, 23.9%23.9\% and 40.8%40.8\% for stages 4--5, 26.9%26.9\% and 45.3%45.3\% for stages 3--5, 28.3%28.3\% and 48.5%48.5\% for stages 2--5, and 28.6%28.6\% and 50.8%50.8\% for stages 1--5, respectively. Thus, stages 4 and 5, which are adjacent to the final stage 6, account for most of the useful alternatives, whereas stages 1--3 contribute substantially less. SQR consequently recollects queries from the immediately preceding and second-preceding stages instead of forwarding all earlier queries. The authors also report that forwarding a query across too many stages can introduce a large learning gap and cause its loss to dominate training, explaining why the selective scheme can outperform dense recollection.

  6. Knowl 6 — SQR improves the Adamixer baseline with less training cost than dense recollection

    data/table

    On COCO validation, Adamixer-R50 trained for the standard 12-epoch schedule obtained the following results:

    method AP AP_50 AP_75 AP_S AP_M AP_L
    Baseline 42.5 61.5 45.6 24.6 45.1 59.2
    DQR 44.2 62.8 47.9 26.7 46.9 60.5
    SQR 44.4 63.2 47.8 25.7 47.4 60.2

    Relative to the baseline, DQR improves AP by 1.71.7 points and SQR by 1.91.9 points. DQR requires 2.24×2.24\times the baseline training time, whereas SQR starting at stage 1 requires 1.57×1.57\times, SQR starting at stage 2 requires 1.34×1.34\times, SQR starting at stage 3 requires 1.18×1.18\times, and SQR starting at stage 4 requires 1.07×1.07\times. Their AP values are 44.444.4, 44.244.2, 43.843.8, and 42.942.9, respectively, compared with 42.542.5 for the baseline. Starting at stage 2 therefore retains nearly the performance of starting at stage 1 while reducing training cost.

  7. Knowl 7 — SQR's gain comes from late-stage emphasis and query access, not merely more losses

    empirical result

    Controlled Adamixer experiments varied how many supervised query pathways were assigned to each of six stages. Uniform supervision with three query groups used (3,3,3,3,3,3)(3,3,3,3,3,3) pathways per stage, 18 total, and achieved 43.443.4 AP. Allocating 18 pathways with emphasis on early stages, (4,4,4,3,2,1)(4,4,4,3,2,1), achieved 43.043.0 AP, while emphasizing late stages, (1,2,3,4,4,4)(1,2,3,4,4,4), achieved 43.743.7 AP. SQR allocations achieved 44.244.2 AP with (1,1,2,3,5,8)(1,1,2,3,5,8) and 44.444.4 AP with (1,2,3,5,8,13)(1,2,3,5,8,13), using 20 and 32 total supervision pathways, respectively. Uniform six-group supervision used 36 pathways but achieved only 43.643.6 AP.

    These comparisons show that the placement of supervision on later stages and the availability of intermediate queries both matter; SQR's improvement is not explained by simply increasing the number of matched queries. In a separate test, reweighting the six ordinary stage losses by (1,2,3,5,8,13)(1,2,3,5,8,13) without recollecting intermediate queries decreased AP by 0.60.6 relative to the baseline. This indicates that access to intermediate query states is necessary in addition to unequal loss weighting.

  8. Knowl 8 — SQR reduces both final-stage true-positive fading and false-positive exacerbation

    empirical result

    Applying SQR to Adamixer reduces the frequency with which an earlier stage is better than the final stage. At an IoU threshold of 0.500.50, the true-positive fading rate decreases from 28.6%28.6\% for the baseline to 23.3%23.3\% with SQR, while the false-positive exacerbation rate decreases from 50.8%50.8\% to 47.3%47.3\%. At an IoU threshold of 0.750.75, the corresponding reductions are from 26.7%26.7\% to 21.1%21.1\% for true-positive fading and from 51.2%51.2\% to 47.0%47.0\% for false-positive exacerbation. These changes support the claim that SQR trains later stages to preserve good intermediate predictions and avoid amplifying bad ones.

  9. Knowl 9 — SQR consistently improves several query-based detectors and backbones

    data/table

    The paper evaluates SQR on COCO 2017 validation with the same inference configuration for each baseline and SQR pair. Representative AP results are:

    detector and setting baseline AP SQR AP gain
    DAB-DETR-R50, 50 epochs 42.2 44.5 +2.3
    DAB-DETR-SwinB, 50 epochs 49.0 51.6 +2.6
    Deformable DETR-R50, 12 epochs 37.2 39.9 +2.7
    Deformable DETR-R50, 50 epochs 44.5 45.9 +1.4
    Adamixer-R50, 12 epochs 42.5 44.4 +1.9
    Adamixer-R50, 12 epochs, 7 stages 42.5 45.3 +2.8
    Adamixer-R50, 36 epochs, 100 queries 45.1 46.7 +1.6
    Adamixer-R50, 36 epochs, 300 queries 46.6 48.9 +2.3
    Adamixer-R101, 36 epochs, 100 queries 45.7 47.3 +1.6
    Adamixer-R101, 36 epochs, 300 queries 47.6 49.8 +2.2

    The experiments use 100 or 300 inference queries and preserve the baseline inference architecture. Across Adamixer, DAB-DETR, and Deformable DETR, SQR produces improvements from 1.41.4 to 2.82.8 AP across backbones, query counts, decoder depths, and training schedules.

  10. Knowl 10 — Training protocol and inference behavior of the evaluation

    experimental setup

    Experiments use the MS-COCO detection task, with all models trained on the train2017 split and evaluated on the COCO validation split. Unless otherwise specified, images are scaled to 800 pixels, AdamW is used, and the standard 1x schedule contains 12 epochs. The main ablation uses Adamixer with a ResNet-50 backbone; broader comparisons use 12-, 36-, and 50-epoch schedules, 100 or 300 queries, and multi-scale training for the 36- and 50-epoch settings with the shorter image side sampled from 480 to 800 pixels. Training for the broad comparison uses eight Nvidia A100 GPUs.

    SQR is applied only during training. The inference pathway remains the ordinary sequential decoder, so the method does not add inference-time query branches or change the number of inference queries. The unchanged inference pipeline is reflected in the paper's speed-versus-AP comparison: SQR-trained models occupy essentially the same inference-speed configurations as their corresponding baselines while attaining higher AP.

  11. Knowl 11 — SQR adds training computation, but longer training does not recover its benefit

    limitation

    SQR increases training computation because multiple recollected query pathways require attention, assignment, and loss computation. For Adamixer-R50, the reported training-time increase ranges from 0.350.35 to 2.852.85 hours, corresponding to approximately 0.07%0.07\% to 57%57\% depending on the configuration, implementation, and recollection starting stage. The relative impact is smaller with heavier backbones such as ResNet-101 and Swin Transformer, for which the paper reports roughly a 10%10\% increase.

    The authors also compare equal training time by extending the baseline: Adamixer receives 200%200\% more epochs and Deformable DETR receives 50%50\% more epochs. Both baselines saturate early and show no significant improvement over the extended schedules, so simply training longer does not compensate for SQR. A stochastic-depth alternative applied to the vanilla decoder reaches 40.740.7 mAP, which is not competitive with SQR. Thus, SQR's main limitation is extra training cost, while its reported benefit is not reproduced by longer baseline training or random variable-depth training.

Coverage note — The paper's detailed descriptions of individual detector architectures, related work, and implementation-specific stochastic-depth probabilities were omitted because they are background or auxiliary comparisons rather than load-bearing parts of the SQR contribution.

References

  1. 1.Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. ArXiv, abs/1607.06450, 2016. 3
  2. 2.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6154–6162, 2018. 3
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. ArXiv, abs/2005.12872, 2020. 1, 3, 8
  4. 4.Fangyi Chen, Chenchen Zhu, Zhiqiang Shen, Han Zhang, and M. Savvides. Ncms: Towards accurate anchor free object detection through l2 norm calibration and multi-feature selection. Comput. Vis. Image Underst., 200:103050, 2020. 1
  5. 5.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019. 6
  6. 6.Qiang Chen, Xiaokang Chen, Gang Zeng, and Jingdong Wang. Group detr: Fast training convergence with decoupled one-to-many label assignment. ArXiv, abs/2207.13085, 2022. 6
  7. 7.Chengjian Feng, Yujie Zhong, Yu Gao, Matthew R Scott, and Weilin Huang. Tood: Task-aligned one-stage object detection. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 3490–3499. IEEE Computer Society, 2021. 1, 3
  8. 8.Peng Gao, Minghang Zheng, Xiaogang Wang, Jifeng Dai, and Hongsheng Li. Fast convergence of detr with spatially modulated co-attention. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 3601–3610, 2021. 8
  9. 9.Ziteng Gao, Limin Wang, Bing Han, and Sheng Guo. Adamixer: A fast-converging query-based object detector. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5354–5363, 2022. 3, 6, 8
  10. 10.Ross B. Girshick. Fast r-cnn. 2015 IEEE International Conference on Computer Vision (ICCV), pages 1440–1448, 2015. 1
  11. 11.Ross B. Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. 2014 IEEE Conference on Computer Vision and Pattern Recognition, pages 580–587, 2014. 1
  12. 12.Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. 6
  13. 13.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger. Deep networks with stochastic depth. In ECCV, 2016. 7
  14. 14.Ding Jia, Yuhui Yuan, Hao He, Xiao pei Wu, Haojun Yu, Weihong Lin, Lei huan Sun, Chao Zhang, and Hanhua Hu. Detrs with hybrid matching. ArXiv, abs/2207.13080, 2022. 6
  15. 15.Kang jik Kim and Hee Seok Lee. Probabilistic anchor assignment with iou prediction for object detection. In ECCV, 2020. 3
  16. 16.Tao Kong, Fuchun Sun, Huaping Liu, Yuning Jiang, and Jianbo Shi. Foveabox: Beyond anchor-based object detector. ArXiv, abs/1904.03797, 2019. 1
  17. 17.Feng Li, Hao Zhang, Shi guang Liu, Jian Guo, Lionel M. Ni, and Lei Zhang. Dn-detr: Accelerate detr training by introducing query denoising. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13609–13617, 2022. 1, 3, 8
  18. 18.Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. ´ 2017 IEEE International Conference on Computer Vision (ICCV), pages 2999–3007, 2017. 1, 3
  19. 19.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and ´ C. Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 6
  20. 20.Shilong Liu, Feng Li, Hao Zhang, Xiao Bin Yang, Xianbiao Qi, Hang Su, Jun Zhu, and Lei Zhang. Dab-detr: Dynamic anchor boxes are better queries for detr. ArXiv, abs/2201.12329, 2022. 1, 3, 8
  21. 21.W. Liu, Dragomir Anguelov, D. Erhan, Christian Szegedy, Scott E. Reed, Cheng-Yang Fu, and Alexander C. Berg. Ssd: Single shot multibox detector. In ECCV, 2016. 1
  22. 22.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9992–10002, 2021. 7
  23. 23.Depu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng, Houqiang Li, Yuhui Yuan, Lei Sun, and Jingdong Wang. Conditional detr for fast training convergence. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 3631–3640, 2021. 1, 3, 8
  24. 24.Jeffrey Ouyang-Zhang, Jang Hyun Cho, Xingyi Zhou, and Philipp Krahenb ¨ uhl. Nms strikes back, 2022. ¨ 3
  25. 25.Joseph Redmon, Santosh Kumar Divvala, Ross B. Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 779–788, 2016. 1
  26. 26.Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39:1137–1149, 2015. 1, 3
  27. 27.Tianhe Ren, Shilong Liu, Hao Zhang, Feng Li, Xingyu Liao, and Lei Zhang. detrex. https://github.com/IDEA-Research/detrex, 2022. 6
  28. 28.Pei Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, and Ping Luo. Sparse r-cnn: End-to-end object detection with learnable proposals. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14449–14458, 2021. 1
  29. 29.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9626–9635, 2019. 1, 3
  30. 30.Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. ArXiv, abs/1706.03762, 2017. 1
  31. 31.Yingming Wang, X. Zhang, Tong Yang, and Jian Sun. Anchor detr: Query design for transformer-based detector. In AAAI, 2022. 1, 3, 8
  32. 32.Haibao Yu, Qi Han, Jianbo Li, Jianping Shi, Guangliang Cheng, and Bin Fan. Search what you want: Barrier panelty nas for mixed precision quantization. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision – ECCV 2020, pages 1–16, Cham, 2020. Springer International Publishing. 1
  33. 33.Haibao Yu, Yizhen Luo, Mao Shu, Yiyi Huo, Zebang Yang, Yifeng Shi, Zhenglong Guo, Hanyu Li, Xing Hu, Jirui Yuan, and Zaiqing Nie. Dair-v2x: A large-scale dataset for vehicle-infrastructure cooperative 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 21361–21370, June 2022. 1
  34. 34.Gongjie Zhang, Zhipeng Luo, Yingchen Yu, Kaiwen Cui, and Shijian Lu. Accelerating detr convergence via semantic-aligned matching. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 939–948, 2022. 8
  35. 35.Haoyang Zhang, Ying Wang, Feras Dayoub, and Niko Sunderhauf. Varifocalnet: An iou-aware dense object detector. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8510–8519, 2021. 1
  36. 36.Shifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei, and Stan Z. Li. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9756–9765, 2020. 1
  37. 37.Chenchen Zhu, Fangyi Chen, Zhiqiang Shen, and Marios Savvides. Soft anchor-point object detection. In ECCV, 2020. 1, 3
  38. 38.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. ArXiv, abs/2010.04159, 2021. 1, 3, 8

Citation

MLA
Chen, F., et al. “Enhanced Training of Query-Based Object Detection via Selective Query Recollection”. arXiv, 2022, http://arxiv.org/abs/2212.07593v3.
APA
Chen, F., Zhang, H., Hu, K., Huang, Y.-. kai ., Zhu, C., & Savvides, M. (2022). Enhanced Training of Query-Based Object Detection via Selective Query Recollection. arXiv. http://arxiv.org/abs/2212.07593v3
Chicago
Chen, F., H. Zhang, K. Hu, Y.-. kai . Huang, C. Zhu, and M. Savvides. 2022. “Enhanced Training of Query-Based Object Detection via Selective Query Recollection”. arXiv. http://arxiv.org/abs/2212.07593v3.
Harvard
Chen, F. et al. (2022) “Enhanced Training of Query-Based Object Detection via Selective Query Recollection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.07593v3.
Vancouver
1. Chen F, Zhang H, Hu K, Huang Y-kai, Zhu C, Savvides M (2022) Enhanced Training of Query-Based Object Detection via Selective Query Recollection. arXiv

BibTeX

@article{chen2022enhanced,
  title = {Enhanced Training of Query-Based Object Detection via Selective Query Recollection},
  author = {Chen, Fangyi and Zhang, Han and Hu, Kai and Huang, Yu-kai and Zhu, Chenchen and Savvides, Marios},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.07593v3},
  eprint = {2212.07593}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE