Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR
Feng LiAiling ZengShilong LiuHao ZhangHongyang LiLei ZhangLionel M. Ni
Proposes an interleaved multi-scale encoder and key-aware deformable attention for DETR architectures, cutting detection head computational cost by 60% while retaining 99% of original detection performance.
Modern computer vision systems rely heavily on Transformer-based object detection frameworks (such as DETR) to achieve high accuracy. However, deploying these models in real-world, resource-constrained environments remains difficult due to high computational demands. The primary computational bottleneck comes from the feature processing encoder, where high-resolution, low-level visual data accounts for more than 75% of all processed tokens. While these low-level features are essential for detecting small objects, processing them through dense multi-scale attention layers incurs steep computation and memory costs.
The article introduces and evaluates Lite DETR, an efficient framework designed to reduce the computational burden of Transformer encoders without compromising detection accuracy. The authors develop an interleaved update scheme that separates multi-scale features into high-level and low-level streams, updating the high-level features frequently while refreshing low-level features at a lower frequency. To preserve accuracy during these asynchronous updates, they introduce Key-aware Deformable Attention, which samples both key and value representations to generate more reliable attention weights across different image scales. The approach was evaluated on the standard Microsoft COCO benchmark across multiple established detection architectures (including Deformable DETR, DINO, and H-DETR) using standard ResNet-50 and Swin-Tiny backbones.
The evaluation demonstrated three core findings. First, Lite DETR reduces the computational operations of the detection encoder by 62% to 78% (and overall detection head computation by about 60%) while retaining 99% of original detection performance. Second, the framework maintains strong performance on small objects, avoiding the 10% performance degradation typically seen when low-level scales are omitted. Third, the plug-and-play architecture readily generalized across different base models; for instance, applying it to DINO with a Swin-Tiny backbone achieved 53.9 Average Precision at 159 GFLOPs, outperforming several state-of-the-art alternative detectors with comparable computational loads.
These findings indicate that organizations can deploy high-performing Transformer vision models on lower-cost or constrained hardware, lowering cloud computing overhead and expanding edge-deployment feasibility. System architects should consider integrating this interleaved multi-scale design into existing vision pipelines when seeking efficiency gains without losing detection accuracy. A notable limitation acknowledged in the article is that the study focuses on theoretical computational savings (measured in GFLOPs) rather than hardware-level runtime and latency optimizations. Decision-makers can have high confidence in the theoretical efficiency and accuracy trade-offs, but should validate hardware-specific latency and throughput in pilot environments before widespread deployment.
- Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). It introduces multi-scale deformable attention for DETR, which forms the foundational baseline architecture and attention mechanism that Lite DETR explicitly redesigns.
- Paper: End-to-End Object Detection with Transformers, Nicolas Carion et al. (2020). It establishes the foundational end-to-end Detection Transformer (DETR) paradigm upon which all subsequent efficient encoder modifications rely.
- Paper: DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR, Shilong Liu et al. (2022). It introduces dynamic anchor boxes as queries for DETR architectures, providing key context for the modern query-based DETR variants evaluated in Lite DETR.
- Paper: PVT v2: Improved baselines with Pyramid Vision Transformer, Wenhai Wang et al. (2021). It demonstrates efficient hierarchical multi-scale feature representations in vision transformers that motivate separating multi-scale token streams.
- Paper: QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object Detection, Chenhongyi Yang et al. (2022). It details how sparse query computation can mitigate the heavy computational burden of processing high-resolution feature maps in object detectors.
- Paper: DETRs Beat YOLOs on Real-time Object Detection, Yian Zhao et al. (2024). It pushes real-time DETR efficiency further by redesigning the multi-scale encoder into a hybrid convolution-attention architecture to outperform YOLO detectors.
- Paper: MS-DETR: Efficient DETR Training with Mixed Supervision, Chuyang Zhao et al. (2024). It enhances training efficiency and convergence across DETR models, complementing Lite DETR's architectural encoder optimizations with mixed supervision schemes.
- Paper: YOLOv12: Attention-Centric Real-Time Object Detectors, Yunjie Tian et al. (2025). It continues the pursuit of attention-driven real-time detection by integrating efficient area-attention mechanisms into real-time vision pipelines.
- Paper: Zero-TPrune: Zero-Shot Token Pruning Through Leveraging of the Attention Graph in Pre-Trained Transformers, Hongjie Wang et al. (2024). It explores training-free token pruning via attention graph modeling, offering an alternative avenue for reducing visual token computation in vision transformers.
