Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR

Feng LiAiling ZengShilong LiuHao ZhangHongyang LiLei ZhangLionel M. Ni

article2023CVPR115 citations

Proposes an interleaved multi-scale encoder and key-aware deformable attention for DETR architectures, cutting detection head computational cost by 60% while retaining 99% of original detection performance.

Listen

Modern computer vision systems rely heavily on Transformer-based object detection frameworks (such as DETR) to achieve high accuracy. However, deploying these models in real-world, resource-constrained environments remains difficult due to high computational demands. The primary computational bottleneck comes from the feature processing encoder, where high-resolution, low-level visual data accounts for more than 75% of all processed tokens. While these low-level features are essential for detecting small objects, processing them through dense multi-scale attention layers incurs steep computation and memory costs.

The article introduces and evaluates Lite DETR, an efficient framework designed to reduce the computational burden of Transformer encoders without compromising detection accuracy. The authors develop an interleaved update scheme that separates multi-scale features into high-level and low-level streams, updating the high-level features frequently while refreshing low-level features at a lower frequency. To preserve accuracy during these asynchronous updates, they introduce Key-aware Deformable Attention, which samples both key and value representations to generate more reliable attention weights across different image scales. The approach was evaluated on the standard Microsoft COCO benchmark across multiple established detection architectures (including Deformable DETR, DINO, and H-DETR) using standard ResNet-50 and Swin-Tiny backbones.

The evaluation demonstrated three core findings. First, Lite DETR reduces the computational operations of the detection encoder by 62% to 78% (and overall detection head computation by about 60%) while retaining 99% of original detection performance. Second, the framework maintains strong performance on small objects, avoiding the 10% performance degradation typically seen when low-level scales are omitted. Third, the plug-and-play architecture readily generalized across different base models; for instance, applying it to DINO with a Swin-Tiny backbone achieved 53.9 Average Precision at 159 GFLOPs, outperforming several state-of-the-art alternative detectors with comparable computational loads.

These findings indicate that organizations can deploy high-performing Transformer vision models on lower-cost or constrained hardware, lowering cloud computing overhead and expanding edge-deployment feasibility. System architects should consider integrating this interleaved multi-scale design into existing vision pipelines when seeking efficiency gains without losing detection accuracy. A notable limitation acknowledged in the article is that the study focuses on theoretical computational savings (measured in GFLOPs) rather than hardware-level runtime and latency optimizations. Decision-makers can have high confidence in the theoretical efficiency and accuracy trade-offs, but should validate hardware-specific latency and throughput in pilot environments before widespread deployment.

Cover for Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR

Abstract

Recent DEtection TRansformer-based (DETR) models have obtained remarkable performance. Its success cannot be achieved without the re-introduction of multi-scale feature fusion in the encoder. However, the excessively increased tokens in multi-scale features, especially for about 75% of low-level features, are quite computationally inefficient, which hinders real applications of DETR models. In this paper, we present Lite DETR, a simple yet efficient end-to-end object detection framework that can effectively reduce the GFLOPs of the detection head by 60% while keeping 99% of the original performance. Specifically, we design an efficient encoder block to update high-level features (corresponding to small-resolution feature maps) and low-level features (corresponding to large-resolution feature maps) in an interleaved way. In addition, to better fuse cross-scale features, we develop a key-aware deformable attention to predict more reliable attention weights. Comprehensive experiments validate the effectiveness and efficiency of the proposed Lite DETR, and the efficient encoder strategy can generalize well across existing DETR-based models. The code will be available in https://github.com/IDEA-Research/Lite-DETR.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Motivation and Analysis
  • 3.2. Model Overview
  • 3.3. Interleaved Update
  • 3.4. Iterative High-level Feature Cross-Scale Fusion
  • 3.6. Key-aware Deformable Attention
  • 3.7. Discussion with Sparse DETR and other Efficient Variants
  • 4. Experiments
  • 4.1. Setup
  • 4.2. Efficiency Improvements on Deformable DETR
  • 4.3. Efficiency Improvements on Other DETRbased Models
  • 4.4. Visualization of KDA
  • 4.5. Ablation Studies
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Lite DETR Interleaved Multi-Scale Encoder Framework

    model/method

    Lite DETR is an efficient multi-scale encoder architecture for Transformer-based object detectors (such as Deformable DETR, DINO, and H-DETR). In standard multi-scale DETR encoders, the multi-scale feature pyramid SS introduces quadratic or large linear token counts dominated by the lowest-level (highest-resolution) feature scale. Lite DETR addresses this by partitioning the multi-scale feature pyramid SS into:

    1. High-level features FH∈RNH×dmodelF_H \in \mathbb{R}^{N_H \times d_{\text{model}}}, corresponding to low-resolution feature maps (e.g., scales S1,S2,S3S_1, S_2, S_3 with downsampling factors of 1/64,1/32,1/161/64, 1/32, 1/16 relative to the input image).
    2. Low-level features FL∈RNL×dmodelF_L \in \mathbb{R}^{N_L \times d_{\text{model}}}, corresponding to high-resolution feature maps (e.g., scale S4S_4 at 1/81/8 resolution), where NL≫NHN_L \gg N_H (NHN_H accounts for only ≈5%∼25%\approx 5\% \sim 25\% of all tokens).

    The encoder is constructed by stacking BB efficient encoder blocks. Within each block, high-level features FHF_H and low-level features FLF_L are updated at different frequencies in an interleaved manner: high-level cross-scale fusion is executed for AA iterations, while low-level cross-scale fusion is executed once at the end of the block. A Lite DETR configuration is parameterized as HL−(A+1)×BHL-(A + 1) \times B, where HH is the number of high-level scales in FHF_H, LL is the number of low-level scales in FLF_L, AA is the number of iterative high-level updates per block, and +1+1 denotes the single low-level update per block stacked across BB blocks.

  2. Knowl 2 — Key-Aware Deformable Attention

    equation

    Standard multi-scale deformable attention directly predicts attention weights and sampling offsets from query tokens via linear projections without explicit key comparisons. When features across scales are updated asynchronously in an interleaved encoder, predicting attention weights purely from query features degrades performance. Key-Aware Deformable Attention (KDA) addresses this by sampling both key and value representations at the predicted offset locations and computing attention weights via scaled dot-product attention.

    Let Q∈RNQ×dmodelQ \in \mathbb{R}^{N_Q \times d_{\text{model}}} be the query tokens, SS be the multi-scale feature pyramid, p∈RNQ×2p \in \mathbb{R}^{N_Q \times 2} be the reference points of the queries, and MM be the number of attention heads. For each head and each feature level across LL scales, KK sampling points are predicted. The total number of sampling locations per query is Nv=M×L×KN_v = M \times L \times K.

    The sampling offsets Δp∈RNQ×Nv×2\Delta p \in \mathbb{R}^{N_Q \times N_v \times 2} are predicted via projection matrix Wp∈Rdmodel×(Nv×2)W^p \in \mathbb{R}^{d_{\text{model}} \times (N_v \times 2)}: Δp=QWp\Delta p = Q W^p

    Values VV and keys KK are sampled from the feature pyramid SS at locations p+Δpp + \Delta p using bilinear interpolation Samp(⋅)\text{Samp}(\cdot) and linearly projected via WV,WK∈Rdmodel×dmodelW^V, W^K \in \mathbb{R}^{d_{\text{model}} \times d_{\text{model}}}: V=Samp(S,p+Δp)WVV = \text{Samp}(S, p + \Delta p) W^V K=Samp(S,p+Δp)WKK = \text{Samp}(S, p + \Delta p) W^K

    The attention output is computed by standard scaled dot-product attention between QQ and the sampled keys KK: KDA(Q,K,V)=Softmax(QKTdk)V\text{KDA}(Q, K, V) = \text{Softmax}\left(\frac{Q K^T}{\sqrt{d_k}}\right) V

    where dk=dmodel/Md_k = d_{\text{model}} / M is the key dimension per head. The computational complexity of KDA remains linear with respect to the number of query tokens, identical to standard deformable attention.

  3. Knowl 3 — Iterative High-Level and Low-Level Feature Cross-Scale Fusion Mechanisms

    model/method

    In each efficient encoder block of Lite DETR, feature fusion proceeds through two asynchronous mechanisms:

    1. Iterative High-Level Feature Cross-Scale Fusion

    High-level features FH∈RNH×dmodelF_H \in \mathbb{R}^{N_H \times d_{\text{model}}} act as queries to extract spatial details and contextual information from all pyramid levels:

    Q=FH,K=V=Concat(FH,FL)FH′=KDA(Q,K,V)Output=Concat(FH′,FL)\begin{aligned} Q &= F_H, \quad K = V = \text{Concat}(F_H, F_L) \\ F'_H &= \text{KDA}(Q, K, V) \\ \text{Output} &= \text{Concat}(F'_H, F_L) \end{aligned}

    This step is executed iteratively for AA layers within a single block. In each step, the updated FH′F'_H replaces FHF_H in both the query QQ and the multi-scale key-value feature set for the subsequent layer.

    2. Efficient Low-Level Feature Cross-Scale Fusion

    At the end of each block (executed once, +1+1), the original low-level tokens FL∈RNL×dmodelF_L \in \mathbb{R}^{N_L \times d_{\text{model}}} query the updated high-level features FH′F'_H and unupdated low-level features FLF_L:

    Q=FL,K=V=Concat(FH′,FL)FL′=KDA(Q,K,V)S′=Concat(FL′,FH′)\begin{aligned} Q &= F_L, \quad K = V = \text{Concat}(F'_H, F_L) \\ F'_L &= \text{KDA}(Q, K, V) \\ S' &= \text{Concat}(F'_L, F'_H) \end{aligned}

    To minimize computation over the large number of low-level tokens NLN_L, the feed-forward network (FFN) following this attention layer is lightweight, using a hidden dimension reduced by a factor of λ=8\lambda = 8 (1/81/8 of the standard FFN hidden size).

  4. Knowl 4 — Token Redundancy and Degradation in Multi-Scale DETR Feature Pyramids

    data/table

    In a 4-scale feature pyramid (S1S_1 at 1/641/64, S2S_2 at 1/321/32, S3S_3 at 1/161/16, and S4S_4 at 1/81/8 of input image resolution), the number of feature tokens increases quadratically towards lower-level (larger-resolution) maps. Low-level tokens account for the vast majority of computational cost while primarily providing local detail for small objects.

    Feature Scale (SS) S1(1/64)S_1 (1/64) S2(1/32)S_2 (1/32) S3(1/16)S_3 (1/16) S4(1/8)S_4 (1/8)
    Token Ratio 1.17%1.17\% 4.71%4.71\% 18.8%18.8\% 75.3%75.3\%

    Directly dropping the lowest-level scale S4S_4 (evaluated on DINO with ResNet-50 on COCO val2017 with 12 epochs) demonstrates that while GFLOPs and memory are drastically reduced, small object detection drops substantially while large object detection remains virtually unaffected:

    Model Total GFLOPs Backbone Encoder Decoder Train Mem AP APS\text{AP}_S APL\text{AP}_L
    DINO-4scale (100%100\%) 235 70 137 28 32G 50.7 33.5 64.7
    DINO-3scale (25%25\%) 122 70 31 21 13G 48.2 30.1 63.9

    Dropping S4S_4 reduces GFLOPs by 48%48\% but incurs a 2.52.5 AP (4.9%4.9\% relative) overall drop and a 3.43.4 APS\text{AP}_S (10.2%10.2\% relative) small-object drop.

  5. Knowl 5 — Lite DETR Performance on DINO and H-DETR Detectors

    data/table

    Plugging the Lite DETR interleaved encoder into state-of-the-art DETR architectures (DINO and H-DETR) evaluated on COCO val2017 (36-epoch schedule) reduces encoder GFLOPs by 62%∼78%62\% \sim 78\% and detection head GFLOPs by ≈60%\approx 60\%, while maintaining approximately 99%99\% of original detection precision across ResNet-50 and Swin-Tiny backbones.

    Model AP AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L GFLOPs Encoder GFLOPs Params
    Swin-T backbone
    DINO 54.1 72.0 59.3 38.3 57.3 68.6 243 137 47M
    Lite-DINO H2L2-(2+1)x3 (5%) 53.1 71.4 57.9 36.6 56.0 68.8 138 30 (↓78%\downarrow 78\%) 47M
    Lite-DINO H3L1-(6+1)x1 (25%) 53.3 71.7 58.2 36.3 56.6 68.7 149 41 (↓70%\downarrow 70\%) 47M
    Lite-DINO H3L1-(2+1)x3 (25%) 53.9 72.0 58.8 37.9 57.0 69.1 159 53 (↓62%\downarrow 62\%) 47M
    H-DETR 53.2 71.5 58.2 35.9 56.4 68.2 234 137 47M
    Lite-H-DETR H2L2-(2+1)x3 (5%) 52.3 70.7 57.2 35.9 55.2 67.7 131 30 47M
    Lite-H-DETR H3L1-(6+1)x1 (25%) 52.7 71.5 58.3 35.6 56.0 68.0 142 41 47M
    Lite-H-DETR H3L1-(2+1)x3 (25%) 53.0 71.3 58.2 36.3 56.3 68.1 152 53 47M
    ResNet-50 backbone
    DINO 50.7 68.6 55.4 33.5 54.0 64.8 235 137 47M
    Lite-DINO H2L2-(2+1)x3 49.9 68.2 54.6 32.3 52.9 64.7 130 30 47M
    Lite-DINO H3L1-(6+1)x1 50.2 68.6 54.3 33.0 53.4 66.0 141 41 47M
    Lite-DINO H3L1-(2+1)x3 50.4 68.5 54.6 33.5 53.6 65.5 151 53 47M
    H-DETR 50.0 68.3 54.4 32.9 52.7 65.3 226 137 47M
    Lite-H-DETR H3L1-(2+1)x3 49.5 67.6 53.9 32.0 52.8 64.0 142 53 47M

    Lite-DINO with Swin-T achieves 53.953.9 AP at 159159 GFLOPs, outperforming real-time CNN detectors such as YOLOv7-X (52.952.9 AP at 190190 GFLOPs).

  6. Knowl 6 — Lite-Deformable DETR Performance Comparison Against Efficient DETR Variants

    data/table

    Lite DETR applied to Deformable DETR on COCO val2017 (50-epoch schedule, ResNet-50 backbone) achieves matching detection precision with significantly lower encoder GFLOPs compared to Deformable DETR and token-dropping methods like Sparse DETR and Efficient DETR.

    Model AP AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L GFLOPs Encoder GFLOPs Params
    Deformable DETR 46.8 66.0 50.6 29.8 49.7 62.0 177 90 40M
    Lite-Deformable DETR H2L2-(2+1)x3 (5%) 45.8 65.1 49.3 27.7 49.1 61.1 108 23 (↓74%\downarrow 74\%) 41M
    Lite-Deformable DETR H3L1-(6+1)x1 (25%) 45.9 65.6 49.2 27.9 49.0 61.6 115 30 (↓66%\downarrow 66\%) 41M
    Lite-Deformable DETR H3L1-(3+1)x2 (25%) 46.2 65.5 49.8 28.2 49.2 61.5 119 35 (↓61%\downarrow 61\%) 41M
    Lite-Deformable DETR H3L1-(2+1)x3 (25%) 46.7 66.1 50.6 29.1 49.7 62.2 123 39 (↓57%\downarrow 57\%) 41M
    Efficient DETR 44.2 62.2 48.0 28.4 47.5 56.6 159 79 32M
    Sparse DETR-rho-0.1 45.3 65.8 49.3 28.4 48.3 60.1 111 24 41M
    Sparse DETR-rho-0.2 45.6 65.8 49.6 28.5 48.6 60.4 119 32 41M
    Sparse DETR-rho-0.3 46.0 65.9 49.7 29.1 49.1 60.6 127 40 41M
    Sparse DETR-rho-0.5 46.3 66.0 50.1 29.0 49.5 60.8 141 54 41M

    Lite-Deformable DETR H3L1-(2+1)x3 achieves 46.746.7 AP (virtually identical to the full baseline's 46.846.8 AP) while requiring only 3939 encoder GFLOPs versus 9090 GFLOPs (57%57\% reduction). It outperforms Sparse DETR-rho-0.3 (46.046.0 AP, 4040 encoder GFLOPs) without requiring auxiliary detection loss in intermediate encoder layers.

  7. Knowl 7 — Ablation of Lite DETR Encoder Components

    data/table

    An ablation study evaluating the individual contributions of iterative high-level cross-scale fusion (HL), low-level cross-scale fusion (LL), and Key-Aware Deformable Attention (KDA) on COCO val2017 using DINO with ResNet-50 (trained for 36 epochs):

    Scale Config HL LL KDA AP APS\text{AP}_S Total GFLOPs Encoder GFLOPs
    DINO-4scale – – – 50.7 33.5 235 137
    3scale – – – 48.2 30.1 122 31
    3scale – – ✓ 49.0 (+0.8) 31.5 125 34
    3scale ✓ – – 49.0 (+0.8) 31.1 128 37
    3scale ✓ ✓ – 49.8 (+0.8) 33.0 147 49
    3scale ✓ ✓ ✓ 50.4 (+0.6) 33.5 151 53
    2scale – – – 45.2 24.1 113 14
    2scale ✓ ✓ – 49.2 (+4.0) 31.8 126 26
    2scale ✓ ✓ ✓ 49.9 (+0.7) 32.3 130 30

    Each component progressively recovers small-object performance: combining HL, LL, and KDA brings the 3-scale model from 48.248.2 AP to 50.450.4 AP, fully matching the APS\text{AP}_S of the 4-scale baseline (33.533.5) with only 5353 encoder GFLOPs compared to 137137 encoder GFLOPs.

  8. Knowl 8 — Ablation of Module Stacking and Update Frequency Hyperparameters

    data/table

    An ablation study on the number of high-level feature scales HH, encoder blocks BB, and iterative high-level cross-scale fusion layers AA within the configuration HL−(A+1)×BHL-(A+1)\times B, evaluated on Deformable DETR with ResNet-50 trained for 50 epochs on COCO val2017:

    Model Configuration AP APS\text{AP}_S Encoder GFLOPs
    Deformable DETR-4scale 46.8 29.8 90
    Deformable DETR-2scale 40.3 20.4 9
    Lite-Deformable DETR H2L2-(2+1)x3 45.8 (+5.5) 27.7 (+7.3) 23
    Deformable DETR-3scale 44.0 26.6 16
    Lite-Deformable DETR H2L2-(6+1)x1 45.9 27.9 28
    Lite-Deformable DETR H3L1-(3+1)x2 46.2 28.2 32
    Lite-Deformable DETR H3L1-(2+1)x3 46.7 (+2.7) 29.1 (+2.5) 36
    Lite-Deformable DETR H3L1-(2+1)x4 46.6 29.6 50

    Increasing the number of high-level scales from 2 to 3 and increasing block repetitions BB up to 3 improves detection accuracy to near parity with the 4-scale baseline (46.746.7 vs 46.846.8 AP). Increasing block repetitions further to (2+1)×4(2+1) \times 4 increases encoder GFLOPs to 5050 without improving overall AP (46.646.6 AP).

  9. Knowl 9 — Incompatibility of Auxiliary Encoder Detection Loss with DINO

    empirical result

    Unlike Sparse DETR, which introduces auxiliary multi-layer detection losses inside intermediate encoder layers to encourage salient token selection, Lite DETR maintains a purely feature-extraction role for the encoder. When auxiliary encoder detection loss is added to DINO with a ResNet-50 backbone, object detection accuracy drops by 1.41.4 AP on COCO val2017, demonstrating that auxiliary encoder detection losses do not generalize well to advanced DETR architectures.

  10. Knowl 10 — Limitations of Lite DETR

    limitation

    Lite DETR focuses specifically on theoretical computational complexity reduction (GFLOPs and number of query tokens). The framework does not include customized low-level hardware or run-time kernel optimizations for GPU memory access patterns or wall-clock latency.

Coverage note — None was omitted; all key architectural components (interleaved encoder, KDA, high-level/low-level fusion), formulas, benchmark comparisons (DINO, H-DETR, Deformable DETR, Sparse DETR), ablations, and stated limitations are included.

References

  1. 1.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020. 1, 2, 6
  2. 2.Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, et al. Hybrid task cascade for instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4974–4983, 2019. 2
  3. 3.Peixian Chen, Mengdan Zhang, Yunhang Shen, Kekai Sheng, Yuting Gao, Xing Sun, Ke Li, and Chunhua Shen. Efficient decoder-free object detection with transformers. arXiv preprint arXiv:2206.06829, 2022. 7
  4. 4.Qiang Chen, Xiaokang Chen, Gang Zeng, and Jingdong Wang. Group DETR: Fast Training Convergence with Decoupled One-to-Many Label Assignment. arXiv preprint arXiv:2207.13085, 2022. 3
  5. 5.Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention Mask Transformer for Universal Image Segmentation. 2022. 1
  6. 6.Bowen Cheng, Alexander G. Schwing, and Alexander Kirillov. Per-Pixel Classification is Not All You Need for Semantic Segmentation. 2021. 1
  7. 7.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 6
  8. 8.Ziteng Gao, Limin Wang, Bing Han, and Sheng Guo. AdaMixer: A Fast-Converging Query-Based Object Detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5364–5373, 2022. 7
  9. 9.Golnaz Ghiasi, Tsung-Yi Lin, and Quoc V Le. Nas-fpn: Learning scalable feature pyramid architecture for object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7036–7045, 2019. 2, 3
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. 1, 4, 6
  11. 11.Ding Jia, Yuhui Yuan, Haodi He, Xiaopei Wu, Haojun Yu, Weihong Lin, Lei Sun, Chao Zhang, and Han Hu. DETRs with Hybrid Matching. arXiv preprint arXiv:2207.13080, 2022. 1, 2, 3, 6, 7
  12. 12.Glenn Jocher, K Nishimura, T Mineeva, and R Vilariño. YoloV5, url=https://github.com/ultralytics/yolov5, 2021. 1, 7
  13. 13.Feng Li, Hao Zhang, Shilong Liu, Jian Guo, Lionel M Ni, and Lei Zhang. DN-DETR: Accelerate DETR Training by Introducing Query DeNoising. arXiv preprint arXiv:2203.01305, 2022. 1, 2, 3
  14. 14.Feng Li, Hao Zhang, Shilong Liu, Lei Zhang, Lionel M Ni, Heung-Yeung Shum, et al. Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation. arXiv preprint arXiv:2206.02777, 2022. 1
  15. 15.Junyu Lin, Xiaofeng Mao, Yuefeng Chen, Lei Xu, Yuan He, and Hui Xue. Dˆ 2ETR: Decoder-Only DETR with Computationally Efficient Cross-Scale Attention. arXiv preprint arXiv:2203.00860, 2022. 7
  16. 16.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017. 2, 3
  17. 17.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 6
  18. 18.Shilong Liu, Feng Li, Hao Zhang, Xiao Yang, Xianbiao Qi, Hang Su, Jun Zhu, and Lei Zhang. DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR. arXiv preprint arXiv:2201.12329, 2022. 1, 2, 3, 6
  19. 19.Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. Path aggregation network for instance segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8759–8768, 2018. 2, 3
  20. 20.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 1, 6
  21. 21.Depu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng, Houqiang Li, Yuhui Yuan, Lei Sun, and Jingdong Wang. Conditional DETR for Fast Training Convergence. arXiv preprint arXiv:2108.06152, 2021. 1, 2, 3, 6
  22. 22.Siyuan Qiao, Liang-Chieh Chen, and Alan Yuille. Detectors: Detecting objects with recursive feature pyramid and switchable atrous convolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10213–10224, 2021. 2
  23. 23.Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018. 1
  24. 24.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28, 2015. 1, 2
  25. 25.Byungseok Roh, JaeWoong Shin, Wuhyun Shin, and Saehoon Kim. Sparse DETR: Efficient End-to-End Object Detection with Learnable Sparsity. arXiv preprint arXiv:2111.14330, 2021. 3, 5, 6
  26. 26.Dahu Shi, Xing Wei, Liangqi Li, Ye Ren, and Wenming Tan. End-to-End Multi-Person Pose Estimation With Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11069–11078, 2022. 1
  27. 27.Hwanjun Song, Deqing Sun, Sanghyuk Chun, Varun Jampani, Dongyoon Han, Byeongho Heo, Wonjae Kim, and Ming-Hsuan Yang. An Extendable, Efficient and Effective Transformer-based Object Detector. arXiv preprint arXiv:2204.07962, 2022. 3, 7
  28. 28.Lucas Stoffl, Maxime Vidal, and Alexander Mathis. End-to-end trainable multi-instance pose estimation with transformers. arXiv preprint arXiv:2103.12115, 2021. 1
  29. 29.Mingxing Tan, Ruoming Pang, and Quoc V Le. Efficientdet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10781–10790, 2020. 1, 2, 3, 7
  30. 30.Chien-Yao Wang, Alexey Bochkovskiy, and Hong-Yuan Mark Liao. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv preprint arXiv:2207.02696, 2022. 1, 7
  31. 31.Tao Wang, Li Yuan, Yunpeng Chen, Jiashi Feng, and Shuicheng Yan. Pnp-detr: Towards efficient visual analysis with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4661–4670, 2021. 7
  32. 32.Yingming Wang, Xiangyu Zhang, Tong Yang, and Jian Sun. Anchor detr: Query design for transformer-based detector. arXiv preprint arXiv:2109.07107, 2021. 6
  33. 33.Zhuyu Yao, Jiangbo Ai, Boxun Li, and Chi Zhang. Efficient DETR: Improving End-to-End Object Detector with Dense Prior. arXiv preprint arXiv:2104.01318, 2021. 2, 3, 6
  34. 34.Chi Zhang, Lijuan Liu, Xiaoxue Zang, Frederick Liu, Hao Zhang, Xinying Song, and Jindong Chen. DETR++: Taming Your Multi-Scale Detection Transformer. arXiv preprint arXiv:2206.02977, 2022. 3
  35. 35.Gongjie Zhang, Zhipeng Luo, Yingchen Yu, Zichen Tian, Jingyi Zhang, and Shijian Lu. Towards Efficient Use of Multi-Scale Features in Transformer-Based Object Detectors. arXiv preprint arXiv:2208.11356, 2022. 3, 7
  36. 36.Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M Ni, and Heung-Yeung Shum. DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. arXiv preprint arXiv:2203.03605, 2022. 1, 2, 3, 6, 7, 8
  37. 37.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In ICLR 2021: The Ninth International Conference on Learning Representations, 2021. 1, 2, 3, 5, 6, 8

Citation

MLA
Li, F., et al. “Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR”. arXiv, 2023, http://arxiv.org/abs/2303.07335v1.
APA
Li, F., Zeng, A., Liu, S., Zhang, H., Li, H., Zhang, L., & Ni, L. M. (2023). Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR. arXiv. http://arxiv.org/abs/2303.07335v1
Chicago
Li, F., A. Zeng, S. Liu, et al. 2023. “Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR”. arXiv. http://arxiv.org/abs/2303.07335v1.
Harvard
Li, F. et al. (2023) “Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.07335v1.
Vancouver
1. Li F, Zeng A, Liu S, Zhang H, Li H, Zhang L, Ni LM (2023) Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR. arXiv

BibTeX

@article{li2023lite,
  title = {Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR},
  author = {Li, Feng and Zeng, Ailing and Liu, Shilong and Zhang, Hao and Li, Hongyang and Zhang, Lei and Ni, Lionel M.},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.07335v1},
  eprint = {2303.07335}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE