QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection

Chenhongyi YangZehao HuangNaiyan Wang

article2022CVPR498 citations

Proposes a cascaded sparse query mechanism that dramatically accelerates high-resolution feature pyramid detectors by predicting coarse object locations on low-resolution maps to guide sparse computation only where small objects exist.

Listen

Visual object detection has become foundational for critical technologies like autonomous driving and aerial surveillance. However, detecting small objects accurately remains a major hurdle. Standard deep learning models lose critical fine details through image down-sampling, leading to poor small-object localization. While incorporating high-resolution feature maps solves the accuracy problem, it increases computational cost quadratically. Processing these high-resolution layers wastes substantial computing power on empty background regions, drastically slowing down processing speed and preventing real-time deployment on hardware-constrained systems.

The article demonstrates an efficient computer vision framework called QueryDet, which uses a Cascade Sparse Query mechanism to accelerate high-resolution small object detection. The objective is to evaluate whether coarse-to-fine selective computation can deliver the accuracy gains of high-resolution processing while eliminating redundant background operations.

The authors evaluate their approach through extensive empirical experiments on standard benchmark datasets, including the broad Microsoft COCO dataset and the small-object-heavy VisDrone dataset. They integrate the query mechanism across multiple standard detection architectures—such as anchor-based (RetinaNet), anchor-free (FCOS), and two-stage detectors (Faster R-CNN)—as well as lightweight mobile backbones. The method operates by first predicting rough, low-resolution locations of small objects and then using those sparse positions to selectively compute high-resolution features only where needed via sparse convolutions.

The experimental findings show significant performance and efficiency gains. First, QueryDet reduces computational floating-point operations in the highest-resolution layers by approximately 99%, focusing almost entirely on relevant target areas. Second, on the COCO benchmark, adding high-resolution features with sparse querying accelerates inference speed to about 3.0 times faster than dense high-resolution computation (running at roughly 14.9 frames per second versus 4.85 frames per second), while improving small-object detection accuracy by 2.0 points. Third, on the VisDrone benchmark, the approach delivers a 2.3-fold speed increase while achieving new state-of-the-art accuracy. Fourth, when paired with lightweight mobile backbones like MobileNetV2, the method achieves an average 4.1-fold speed acceleration, proving its compatibility with edge hardware.

These results demonstrate that organizations deploying computer vision do not have to choose between detection precision and operational latency. Eliminating spatial redundancy directly translates to lower hardware and energy costs, higher frame rates, and safer response times in mission-critical applications such as autonomous driving. It also allows developers to integrate higher-resolution inputs into legacy systems without requiring costly infrastructure upgrades.

Organizations developing edge-vision systems should consider adopting sparse query mechanisms to optimize their detection pipelines. When deploying the system, engineering teams can adjust a single query threshold to balance accuracy against processing speed according to operational requirements. Before broad deployment, teams should conduct real-world pilot tests to calibrate threshold sensitivity against false-positive queries caused by large foreground objects. The authors suggest extending this sparse querying paradigm to three-dimensional point cloud data from LiDAR sensors, where spatial sparsity is even greater and computation costs are higher.

Cover for QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection

Abstract

While general object detection with deep learning has achieved great success in the past few years, the performance and efficiency of detecting small objects are far from satisfactory. The most common and effective way to promote small object detection is to use high-resolution images or feature maps. However, both approaches induce costly computation since the computational cost grows squarely as the size of images and features increases. To get the best of two worlds, we propose QueryDet that uses a novel query mechanism to accelerate the inference speed of feature-pyramid based object detectors. The pipeline composes two steps: it first predicts the coarse locations of small objects on low-resolution features and then computes the accurate detection results using high-resolution features sparsely guided by those coarse positions. In this way, we can not only harvest the benefit of high-resolution feature maps but also avoid useless computation for the background area. On the popular COCO dataset, the proposed method improves the detection mAP by 1.0 and mAP-small by 2.0, and the high-resolution inference speed is improved to 3.0× on average. On VisDrone dataset, which contains more small objects, we create a new state-of-the-art while gaining a 2.3× high-resolution acceleration on average. Code is available at https://github.com/ChenhongyiYang/QueryDet-PyTorch.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methods
  • 3.1. Revisiting RetinaNet
  • 3.2. Accelerating Inference by Sparse Query
  • 3.3. Training
  • 3.4. Relationships with Related Work
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Effectiveness of Our Approach
  • 4.3. Ablation Studies
  • 4.4. Discussions
  • 4.5. Visualization and Failure Cases
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — QueryDet’s coarse-to-fine sparse detection architecture

    model/method

    QueryDet accelerates feature-pyramid object detectors by replacing dense computation on high-resolution feature maps with a coarse-to-fine sparse process. An FPN produces maps P7,…,P2P_7,\ldots,P_2, where P2P_2 and P3P_3 are the highest-resolution maps used for detecting small objects. A query head first predicts where small objects may occur on a coarser feature map. These predicted positions are used as query keys to select corresponding regions from the next finer feature map, whose selected features are the query values.

    The classification, bounding-box regression, and query heads then use sparse convolutions only around the selected positions. The process is cascaded toward finer pyramid levels, so the high-resolution detection heads operate only on locations likely to contain small objects. QueryDet therefore retains the accuracy benefit of high-resolution features while avoiding dense background computation. The method is designed for FPN-based one-stage detectors and can also be inserted into the region proposal network of a two-stage detector.

  2. Knowl 2 — Cascade Sparse Query inference procedure

    algorithm

    Cascade Sparse Query (CSQ) performs inference by recursively propagating sparse query positions from coarse to fine FPN levels. In the paper’s standard RetinaNet configuration, querying starts at P4P_4, the query threshold is σ=0.15\sigma=0.15, and processing continues through P3P_3 and P2P_2.

    Input: FPN feature maps P7, ..., P2; query start level P4; threshold sigma = 0.15
    Output: Detection predictions from all pyramid levels
    Run the ordinary dense detection heads on the coarse levels P7, P6, P5, and P4.
    At P4, run the query head and retain positions whose query score is larger than sigma.
    For each retained query position (x, y) at level l, map it to the four positions
      (2x + i, 2y + j) at the next finer level l - 1, for i in {0, 1} and j in {0, 1}.
    Union and deduplicate all mapped positions to form the sparse key set at level l - 1.
    Gather the corresponding features from P(l - 1) to construct a sparse value feature map.
    Activate the local context around the key positions; the experiments use a 5 x 5 context patch.
    Apply sparse versions of the classification, regression, and query heads to the sparse value map.
    Store the classification and regression predictions at the active positions.
    Use the sparse query predictions to create the key set for the next finer level.
    Repeat until P2 has been processed.
    Return the accumulated detections.

    For a query at coordinates (xl,yl)(x_l,y_l) on PlP_l, the key positions on Pl−1P_{l-1} are

    {kl−1}={(2xl+i,2yl+j):i,j∈{0,1}}.\{k_{l-1}\}=\{(2x_l+i,2y_l+j): i,j\in\{0,1\}\}.

    Unlike a proposal-based two-stage detector, the coarse query prediction produces objectness locations rather than bounding boxes, so QueryDet does not require RoIAlign or RoIPooling. Its computational cost depends on the number of active query and context positions rather than the full spatial area of the high-resolution maps.

  3. Knowl 3 — Supervised query-map construction

    equation

    For each FPN level PlP_l, QueryDet defines a ground-truth query map from the centers of small ground-truth objects. Let (x,y)(x,y) denote a spatial position on PlP_l, let (xlo,ylo)(x_l^o,y_l^o) be the center of small object oo expressed in the coordinate system of PlP_l, and let sls_l be the level-specific small-object threshold. The paper sets sls_l to the minimum anchor scale at level PlP_l for anchor-based detectors and to the minimum regression range for anchor-free detectors.

    The distance from a feature position to the nearest small-object center is

    Dl[x][y]=min⁡o(x−xlo)2+(y−ylo)2.D_l[x][y]=\min_o\sqrt{(x-x_l^o)^2+(y-y_l^o)^2}.

    The binary target for the query head is

    Vl∗[x][y]={1,Dl[x][y]<sl,0,Dl[x][y]≥sl.V_l^*[x][y]= \begin{cases} 1, & D_l[x][y]<s_l,\\ 0, & D_l[x][y]\geq s_l. \end{cases}

    Thus, the query head is trained to activate a neighborhood around each small-object center rather than only a single center pixel. At inference time, query locations whose predicted probability exceeds σ\sigma are propagated to the next finer pyramid level.

  4. Knowl 4 — Layer-balanced training objective

    equation

    QueryDet retains the original RetinaNet classification and bounding-box regression training and adds a focal-loss term for the query head. At pyramid level PlP_l, let UlU_l, RlR_l, and VlV_l denote the classification, regression, and query outputs; let Ul∗U_l^*, Rl∗R_l^*, and Vl∗V_l^* denote their corresponding targets. The level loss is

    Ll(Ul,Rl,Vl)=LFL(Ul,Ul∗)+Lr(Rl,Rl∗)+LFL(Vl,Vl∗),\mathcal{L}_l(U_l,R_l,V_l)=\mathcal{L}_{FL}(U_l,U_l^*)+\mathcal{L}_{r}(R_l,R_l^*)+\mathcal{L}_{FL}(V_l,V_l^*),

    where LFL\mathcal{L}_{FL} is focal loss and Lr\mathcal{L}_{r} is the smooth-L1L_1 bounding-box regression loss. The total training loss is a weighted sum over FPN levels:

    Lall=∑lβlLl.\mathcal{L}_{all}=\sum_l \beta_l\mathcal{L}_l.

    The weights βl\beta_l compensate for the large increase in training samples caused by adding the high-resolution P2P_2 level. For COCO, the weights grow linearly from 11 on P2P_2 to 33 on P7P_7; for VisDrone, they grow linearly from 11 to 2.62.6. This prevents the numerous small-object samples on P2P_2 from dominating optimization.

  5. Knowl 5 — High-resolution computation is concentrated in sparse pyramid levels

    empirical result

    With a ResNet-50 backbone, the dense RetinaNet detection head spends nearly half of its computation on the high-resolution P3P_3 feature map: P3P_3 accounts for about 43%43\% of total FLOPs, whereas the low-resolution levels P4P_4 through P7P_7 together account for only about 15%15\%. Adding P2P_2 for improved small-object detection makes P2P_2 and P3P_3 jointly account for about 74%74\% of total computation.

    QueryDet applies sparse computation to these expensive levels. The paper reports that CSQ reduces computation on the high-resolution P2P_2 and P3P_3 heads by about 99%99\%, leaving roughly 1%1\% of the corresponding dense computation while preserving nearly the same detection accuracy. This is the main source of the speedup from high-resolution QueryDet inference.

  6. Knowl 6 — COCO and VisDrone accuracy–speed results

    data/table

    The main experiments compare RetinaNet with high-resolution QueryDet on COCO mini-val and VisDrone validation data. QueryDet without CSQ uses the high-resolution feature map densely; QueryDet with CSQ uses the same model with cascaded sparse inference. AP denotes average precision, AP50AP_{50} and AP75AP_{75} use IoU thresholds of 0.500.50 and 0.750.75, APSAP_S, APMAP_M, and APLAP_L correspond to small, medium, and large objects, and FPS is inference speed.

    On COCO, adding high-resolution features improves small-object AP substantially, while CSQ recovers and exceeds the baseline speed with only a small AP decrease:

    Could not parse LaTeX table

    Relative to the RetinaNet baseline, the standard QueryDet configuration improves overall AP from 37.4637.46 to 38.5338.53 and small-object AP from 22.6422.64 to 24.6424.64 before sparse inference. CSQ then increases speed from 4.854.85 FPS to 14.8814.88 FPS while reducing AP by only 0.170.17.

    VisDrone contains a much larger proportion of small objects, and QueryDet produces a larger accuracy gain while maintaining nearly the baseline speed:

    Could not parse LaTeX table

    On VisDrone, QueryDet improves AP by 2.142.14 points over RetinaNet, improves AP50AP_{50} by 3.243.24 points, and raises speed from 1.161.16 FPS to 2.752.75 FPS when CSQ is enabled.

  7. Knowl 7 — Ablation of high-resolution features, loss rebalance, query supervision, and CSQ

    data/table

    The COCO mini-val ablation isolates the components of QueryDet. HR means that the high-resolution P2P_2 feature is used, RB means layer-wise loss rebalancing, QH means the additional query head, and CSQ means cascaded sparse inference.

    Could not parse LaTeX table

    Adding P2P_2 without rebalancing lowers AP from 37.4637.46 to 36.1036.10, because the enlarged training set becomes dominated by small-object samples. Rebalancing raises the high-resolution model to 38.1138.11 AP. Adding the query head increases AP to 38.5338.53 and APSAP_S to 24.6424.64, demonstrating the value of explicit objectness supervision. Finally, CSQ increases speed from 4.854.85 to 14.8814.88 FPS with only a 0.170.17-point AP decrease.

  8. Knowl 8 — Empirical choices governing sparse-query efficiency

    empirical result

    The paper evaluates the starting pyramid level, query implementation, and context size. Starting CSQ at P4P_4 gives the best speed–accuracy balance on COCO mini-val: starting at P6P_6, P5P_5, P4P_4, and P3P_3 gives respectively (37.91,13.42)(37.91,13.42), (38.22,13.92)(38.22,13.92), (38.36,14.88)(38.36,14.88), and (38.45,11.51)(38.45,11.51) for (AP,FPS)(AP,\mathrm{FPS}). Starting at very coarse levels loses accuracy because small objects are difficult to distinguish there, while starting at P3P_3 leaves too little dense computation to amortize sparse-map construction.

    CSQ is faster than two alternatives with similar accuracy. Dense high-resolution processing without a query method gives 38.5338.53 AP at 4.864.86 FPS; Crop Query gives 38.3138.31 AP at 10.4910.49 FPS; Complete Convolution Query gives 38.3238.32 AP at 8.738.73 FPS; and CSQ gives 38.3638.36 AP at 14.8814.88 FPS.

    The context patch around each queried position controls another accuracy–speed trade-off. A 1×11\times1 patch gives 38.2538.25 AP at 14.0914.09 FPS, a 5×55\times5 patch gives 38.3638.36 AP at 14.0014.00 FPS, and an 11×1111\times11 patch gives 38.3838.38 AP at 13.1113.11 FPS. The paper therefore uses 5×55\times5 context as a practical compromise: smaller context harms detection, whereas larger context produces only marginal AP gains and reduces speed. Increasing the query threshold σ\sigma reduces active positions and increases speed but can lower recall; even a threshold of 0.050.05 already produces a substantial speedup.

  9. Knowl 9 — Generalization beyond the RetinaNet–ResNet-50 configuration

    empirical result

    QueryDet generalizes to lightweight backbones, anchor-free detectors, and two-stage detectors. On COCO mini-val with MobileNet V2, RetinaNet runs at 17.7517.75 FPS and 26.7226.72 AP; dense high-resolution QueryDet reaches 29.1629.16 AP at 5.315.31 FPS; and CSQ reaches 28.9428.94 AP at 21.6621.66 FPS. With ShuffleNet V2, the corresponding results are 23.0423.04 AP at 17.4517.45 FPS for RetinaNet, 26.0726.07 AP at 5.265.26 FPS for dense QueryDet, and 25.8525.85 AP at 20.0220.02 FPS with CSQ. The reported high-resolution speedups average 4.1×4.1\times with MobileNet V2 and 3.8×3.8\times with ShuffleNet V2.

    Applied to the anchor-free FCOS detector, QueryDet raises AP from 38.3738.37 to 40.0540.05 with dense high-resolution features and achieves 39.4939.49 AP with CSQ. The corresponding speeds are 17.0617.06, 7.927.92, and 14.4014.40 FPS, respectively, giving a reported 1.8×1.8\times speedup over dense high-resolution inference.

    Applied to Faster R-CNN’s FPN-based region proposal network, the implementation uses proposal levels P2P_2–P6P_6, starts querying at P4P_4, and adds separate objectness, regression, and query branches. The dense version reaches 38.4738.47 AP, 22.9822.98 small-object AP, and 17.5717.57 FPS; CSQ reaches 38.2038.20 AP, 22.2322.23 small-object AP, and 19.0319.03 FPS. In this setting, CSQ can reduce both dense high-resolution RPN computation and the number of proposals passed to the second stage.

  10. Knowl 10 — Observed failure modes and practical limitation

    limitation

    QueryDet depends on the query head producing sufficiently complete and selective coarse locations. One observed failure occurs when the query head correctly activates a small object but the sparse detection head still fails to localize the object accurately. A second occurs when large-object locations are falsely activated by the query head; these locations do not necessarily create false detections, but they trigger unnecessary sparse computation and reduce the achievable speedup.

    The method also exhibits an inherent threshold trade-off: a higher query threshold activates fewer positions and is faster, but can omit small objects and reduce recall. Conversely, low thresholds preserve recall but activate more background. QueryDet is therefore most effective when small objects are spatially sparse and the query head has high recall with few false activations.

Coverage note — No substantial contributed material was omitted; background, related work, references, and the paper’s proposed future extension to 3D detection were excluded.

References

  1. 1.Zhaowei Cai, Quanfu Fan, Rogerio S Feris, and Nuno Vasconcelos. A unified multi-scale deep convolutional neural network for fast object detection. In ECCV. Springer, 2016. 2
  2. 2.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In CVPR, 2018. 2
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-End Object Detection with Transformers. ECCV, 2020. 1
  4. 4.Chenyi Chen, Ming-Yu Liu, Oncel Tuzel, and Jianxiong Xiao. R-CNN for small object detection. In ACCV. Springer, 2016. 2
  5. 5.Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker. Learning efficient object detection models with knowledge distillation. In NeurIPS, 2017. 2
  6. 6.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4), 2017. 2
  7. 7.Kaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian. Centernet: Keypoint triplets for object detection. In ICCV, 2019. 2
  8. 8.Michael Figurnov, Maxwell D Collins, Yukun Zhu, Li Zhang, Jonathan Huang, Dmitry Vetrov, and Ruslan Salakhutdinov. Spatially adaptive computation time for residual networks. In CVPR, 2017. 3
  9. 9.Mikhail Figurnov, Aizhan Ibraimova, Dmitry P Vetrov, and Pushmeet Kohli. Perforatedcnns: Acceleration through elimination of redundant convolutions. In NeurIPS, 2016. 2
  10. 10.Cheng-Yang Fu, Wei Liu, Ananth Ranga, Ambrish Tyagi, and Alexander C Berg. DSSD: Deconvolutional single shot detector. arXiv preprint arXiv:1701.06659, 2017. 2
  11. 11.Ross Girshick. Fast r-cnn. In ICCV, 2015. 2, 4, 5
  12. 12.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014. 1, 2
  13. 13.Benjamin Graham and Laurens van der Maaten. Submanifold sparse convolutional networks. arXiv preprint arXiv:1706.01307, 2017. 2, 4
  14. 14.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In ICCV, 2017. 2, 5
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 1
  16. 16.Yihui He, Xiangyu Zhang, and Jian Sun. Channel pruning for accelerating very deep neural networks. In ICCV, 2017. 2
  17. 17.Lichao Huang, Yi Yang, Yafeng Deng, and Yinan Yu. DenseBox: Unifying landmark localization with end to end object detection. arXiv preprint arXiv:1509.04874, 2015. 2
  18. 18.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. ICLR, 2017. 3
  19. 19.Alexander Kirillov, Yuxin Wu, Kaiming He, and Ross Girshick. Pointrend: Image segmentation as rendering. In CVPR, 2020. 5
  20. 20.Mate Kisantal, Zbigniew Wojna, Jakub Murawski, Jacek Naruniec, and Kyunghyun Cho. Augmentation for small object detection. arXiv preprint arXiv:1902.07296, 2019. 2
  21. 21.Tao Kong, Fuchun Sun, Huaping Liu, Yuning Jiang, Lei Li, and Jianbo Shi. FoveaBox: Beyound anchor-based object detection. IEEE Transactions on Image Processing, 29, 2020. 2
  22. 22.Tao Kong, Anbang Yao, Yurong Chen, and Fuchun Sun. Hypernet: Towards accurate region proposal generation and joint object detection. In CVPR, 2016. 2
  23. 23.Hei Law and Jia Deng. CornerNet: Detecting objects as paired keypoints. In ECCV, 2018. 2
  24. 24.Jianan Li, Xiaodan Liang, Yunchao Wei, Tingfa Xu, Jiashi Feng, and Shuicheng Yan. Perceptual generative adversarial networks for small object detection. In CVPR, 2017. 2
  25. 25.Yanghao Li, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang. Scale-aware trident networks for object detection. In ICCV, 2019. 1, 2
  26. 26.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, 2017. 1, 2
  27. 27.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ICCV, 2017. 1, 2, 3, 4, 5
  28. 28.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV. Springer, 2014. 1, 2, 5
  29. 29.Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In ECCV. Springer, 2016. 1, 2
  30. 30.Ziming Liu, Guangyu Gao, Lin Sun, and Zhiyuan Fang. HRDNet: High-resolution Detection Network for Small Objects. arXiv preprint arXiv:2006.07607, 2020. 5
  31. 31.Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design. In ECCV, 2018. 8
  32. 32.Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. Mixed precision training. ICLR, 2018. 5
  33. 33.Mahyar Najibi, Bharat Singh, and Larry S Davis. Autofocus: Efficient multi-scale inference. In ICCV, 2019. 3, 7
  34. 34.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019. 5
  35. 35.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In CVPR, 2016. 2
  36. 36.Joseph Redmon and Ali Farhadi. YOLO9000: better, faster, stronger. In CVPR, 2017. 2
  37. 37.Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018. 1, 2
  38. 38.Mengye Ren, Andrei Pokrovsky, Bin Yang, and Raquel Urtasun. SBNet: Sparse Blocks Network for Fast Inference. In CVPR, June 2018. 3
  39. 39.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In NeurIPS, 2015. 1, 2, 8
  40. 40.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted Residuals and Linear Bottlenecks. In CVPR, 2018. 7
  41. 41.Abhinav Shrivastava, Rahul Sukthankar, Jitendra Malik, and Abhinav Gupta. Beyond skip connections: Top-down modulation for object detection. arXiv preprint arXiv:1612.06851, 2016. 2
  42. 42.Bharat Singh and Larry S Davis. An analysis of scale invariance in object detection snip. In CVPR, 2018. 2
  43. 43.Bharat Singh, Mahyar Najibi, and Larry S Davis. Sniper: Efficient multi-scale training. In NeurIPS, 2018. 2
  44. 44.Mingxing Tan, Ruoming Pang, and Quoc V Le. EfficientDet: Scalable and efficient object detection. In CVPR, 2020. 2
  45. 45.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. In ICCV, 2019. 2
  46. 46.Burak Uzkent, Christopher Yeh, and Stefano Ermon. Efficient object detection in large images using deep reinforcement learning. In WACV, 2020. 3
  47. 47.Thomas Verelst and Tinne Tuytelaars. Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference. In CVPR, June 2020. 2
  48. 48.Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, et al. Deep high-resolution representation learning for visual recognition. IEEE transactions on pattern analysis and machine intelligence, 2020. 2
  49. 49.Xinjiang Wang, Shilong Zhang, Zhuoran Yu, Litong Feng, and Wayne Zhang. Scale-equalizing Pyramid Convolution for Object Detection. In CVPR, 2020. 1
  50. 50.Yulin Wang, Kangchen Lv, Rui Huang, Shiji Song, Le Yang, and Gao Huang. Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image Classification. NeurIPS, 2020. 3
  51. 51.Yi Wei, Xinyu Pan, Hongwei Qin, Wanli Ouyang, and Junjie Yan. Quantization mimic: Towards very tiny cnn for object detection. In ECCV, 2018. 2
  52. 52.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019. 5
  53. 53.Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, and Kaiming He. Aggregated Residual Transformations for Deep Neural Networks. In CVPR, 2017. 1
  54. 54.Zhenda Xie, Zheng Zhang, Xizhou Zhu, Gao Huang, and Stephen Lin. Spatially Adaptive Inference with Stochastic Feature Sampling and Interpolation. ECCV, 2020. 2
  55. 55.Yan Yan, Yuxing Mao, and Bo Li. Second: Sparsely embedded convolutional detection. Sensors, 18(10), 2018. 2
  56. 56.Ze Yang, Shaohui Liu, Han Hu, Liwei Wang, and Stephen Lin. Reppoints: Point set representation for object detection. In ICCV, 2019. 2
  57. 57.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122, 2015. 2
  58. 58.Shifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei, and Stan Z Li. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection. In CVPR, 2020. 2
  59. 59.Pengfei Zhu, Longyin Wen, Dawei Du, Xiao Bian, Haibin Ling, Qinghua Hu, Qinqin Nie, Hao Cheng, Chenfeng Liu, Xiaoyu Liu, et al. Visdrone-det2018: The vision meets drone object detection in image challenge results. In ECCV, 2018. 2, 5
  60. 60.Barret Zoph, Ekin D Cubuk, Golnaz Ghiasi, Tsung-Yi Lin, Jonathon Shlens, and Quoc V Le. Learning data augmentation strategies for object detection. arXiv preprint arXiv:1906.11172, 2019. 2

Citation

MLA
Yang, C., et al. “QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection”. arXiv, 2021, http://arxiv.org/abs/2103.09136v2.
APA
Yang, C., Huang, Z., & Wang, N. (2021). QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection. arXiv. http://arxiv.org/abs/2103.09136v2
Chicago
Yang, C., Z. Huang, and N. Wang. 2021. “QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection”. arXiv. http://arxiv.org/abs/2103.09136v2.
Harvard
Yang, C., Huang, Z. and Wang, N. (2021) “QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2103.09136v2.
Vancouver
1. Yang C, Huang Z, Wang N (2021) QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection. arXiv

BibTeX

@article{yang2021querydet,
  title = {QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection},
  author = {Yang, Chenhongyi and Huang, Zehao and Wang, Naiyan},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2103.09136v2},
  eprint = {2103.09136}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE