QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection
Chenhongyi YangZehao HuangNaiyan Wang
Proposes a cascaded sparse query mechanism that dramatically accelerates high-resolution feature pyramid detectors by predicting coarse object locations on low-resolution maps to guide sparse computation only where small objects exist.
Visual object detection has become foundational for critical technologies like autonomous driving and aerial surveillance. However, detecting small objects accurately remains a major hurdle. Standard deep learning models lose critical fine details through image down-sampling, leading to poor small-object localization. While incorporating high-resolution feature maps solves the accuracy problem, it increases computational cost quadratically. Processing these high-resolution layers wastes substantial computing power on empty background regions, drastically slowing down processing speed and preventing real-time deployment on hardware-constrained systems.
The article demonstrates an efficient computer vision framework called QueryDet, which uses a Cascade Sparse Query mechanism to accelerate high-resolution small object detection. The objective is to evaluate whether coarse-to-fine selective computation can deliver the accuracy gains of high-resolution processing while eliminating redundant background operations.
The authors evaluate their approach through extensive empirical experiments on standard benchmark datasets, including the broad Microsoft COCO dataset and the small-object-heavy VisDrone dataset. They integrate the query mechanism across multiple standard detection architectures—such as anchor-based (RetinaNet), anchor-free (FCOS), and two-stage detectors (Faster R-CNN)—as well as lightweight mobile backbones. The method operates by first predicting rough, low-resolution locations of small objects and then using those sparse positions to selectively compute high-resolution features only where needed via sparse convolutions.
The experimental findings show significant performance and efficiency gains. First, QueryDet reduces computational floating-point operations in the highest-resolution layers by approximately 99%, focusing almost entirely on relevant target areas. Second, on the COCO benchmark, adding high-resolution features with sparse querying accelerates inference speed to about 3.0 times faster than dense high-resolution computation (running at roughly 14.9 frames per second versus 4.85 frames per second), while improving small-object detection accuracy by 2.0 points. Third, on the VisDrone benchmark, the approach delivers a 2.3-fold speed increase while achieving new state-of-the-art accuracy. Fourth, when paired with lightweight mobile backbones like MobileNetV2, the method achieves an average 4.1-fold speed acceleration, proving its compatibility with edge hardware.
These results demonstrate that organizations deploying computer vision do not have to choose between detection precision and operational latency. Eliminating spatial redundancy directly translates to lower hardware and energy costs, higher frame rates, and safer response times in mission-critical applications such as autonomous driving. It also allows developers to integrate higher-resolution inputs into legacy systems without requiring costly infrastructure upgrades.
Organizations developing edge-vision systems should consider adopting sparse query mechanisms to optimize their detection pipelines. When deploying the system, engineering teams can adjust a single query threshold to balance accuracy against processing speed according to operational requirements. Before broad deployment, teams should conduct real-world pilot tests to calibrate threshold sensitivity against false-positive queries caused by large foreground objects. The authors suggest extending this sparse querying paradigm to three-dimensional point cloud data from LiDAR sensors, where spatial sparsity is even greater and computation costs are higher.
- Paper: Focal Loss for Dense Object Detection, Tsung-Yi Lin et al. (2017). Introduces RetinaNet and Feature Pyramid Networks with Focal Loss, establishing the primary single-stage multi-scale detection architecture that QueryDet accelerates via sparse queries.
- Paper: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Shaoqing Ren et al. (2015). Provides the foundational two-stage detection framework (Faster R-CNN) upon which QueryDet integrates its coarse-to-fine query mechanism to test modular acceleration.
- Paper: Cascade R-CNN: Delving Into High Quality Object Detection, Zhaowei Cai et al. (2017). Pioneers the cascaded multi-stage architecture for progressively refining object hypotheses, providing the core design principle underlying QueryDet's cascaded query strategy.
- Paper: Sparse R-CNN: End-to-End Object Detection with Learnable Proposals, Peize Sun et al. (2020). Demonstrates the paradigm of sparse proposal interactions for object detection, motivating QueryDet's concept of query-based computation over sparse spatial locations.
- Paper: Single-Shot Refinement Neural Network for Object Detection, Shifeng Zhang et al. (2017). Presents coarse-to-fine refinement and background filtering across stages to speed up detection, directly preceding QueryDet's selective feature computation approach.
- Paper: TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios, Xingkui Zhu et al. (2021). Focuses on the problem of small object detection in aerial and drone imagery like VisDrone, establishing the domain challenges and baselines targeted by QueryDet.
- Paper: Fast Feature Pyramids for Object Detection, Piotr Dollar et al. (2014). Analyzes efficient scale-space feature extraction to avoid dense pyramid computation, motivating hierarchical coarse-to-fine computation for speedup.
- Paper: Mask Transfiner for High-Quality Instance Segmentation, Lei Ke et al. (2022). Extends the principle of sparse, selective multi-scale computation to high-resolution instance segmentation by focusing processing strictly on boundary-incoherent regions.
- Paper: Efficient Multi-Scale Attention Module with Cross-Spatial Learning, Daliang Ouyang et al. (2023). Builds upon efficient feature extraction concepts by introducing cross-spatial multi-scale attention to boost small object detection on benchmarks like VisDrone.
- Paper: DETRs Beat YOLOs on Real-time Object Detection, Yian Zhao et al. (2024). Applies real-time query-based object detection principles to transformer architectures, pushing forward end-to-end acceleration across multi-scale features.
- Paper: A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS, Juan R. Terven et al. (2023). Surveys the progression of modern real-time detector architectures and optimization strategies that succeeded and contextualize QueryDet.
