DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

Shilong LiuFeng LiHao ZhangXiao YangXianbiao QiHang SuJun ZhuLei Zhang

article2022ICLR1,195 citations

Proposes using dynamic anchor boxes as queries in detection transformers to accelerate training convergence and improve object detection accuracy on COCO through explicit spatial priors and scale-aware positional attention updated layer by layer.

Listen

Modern vision systems rely heavily on object detection to identify and localize items within images for applications such as autonomous driving and medical imaging. While transformer-based detectors eliminate complex, hand-crafted components by predicting objects directly, they suffer from exceptionally slow training convergence, often requiring 500 training cycles to reach competitive performance. This inefficiency inflates computational costs, slows development cycles, and complicates deployment.

The article demonstrates that this training bottleneck stems from how the model queries visual features. To solve this, the authors introduce DAB-DETR, a framework that directly formulates queries as dynamic four-dimensional anchor box coordinates (position, width, and height) and refines them layer by layer throughout the network. The approach was systematically evaluated using standard computer vision benchmarks on the standard COCO dataset across multiple backbone network configurations, training for only 50 epochs on modern computing clusters.

The analysis produced several key findings. First, shifting to explicit four-dimensional anchor boxes resolved the multi-mode ambiguity of prior designs, allowing the model to achieve superior accuracy in just 50 epochs—a tenfold reduction in training time compared to standard 500-epoch baselines. Second, the architecture reached top-tier accuracy, achieving up to 45.7% Average Precision with a standard ResNet-50 backbone, outperforming competing architectures under identical settings. Third, incorporating width and height to modulate spatial attention maps and tuning the temperature parameter significantly improved the model's ability to handle objects of varying shapes and sizes. Finally, the dynamic anchor box formulation proved adaptable; applying it to other detector variants required fewer than 10 lines of code while delivering immediate performance gains.

These findings indicate that transformer-based object detection can be trained significantly faster without sacrificing accuracy or incurring substantial computational overhead during inference. For organizations deploying vision models, this reduces cloud compute expenses, accelerates iteration cycles, and improves interpretability by making internal feature pooling operate like a cascading refinement process.

Teams developing vision-based detection systems should adopt dynamic anchor box query formulations to streamline training pipelines. For immediate implementation, integrating scale modulation and proper coordinate temperature tuning into existing transformer decoders provides clear accuracy gains at minimal engineering cost. Practitioners should conduct pilot tests on domain-specific datasets to confirm performance before full-scale deployment.

Confidence in these findings is high due to rigorous benchmarking and ablation studies across various backbone architectures. However, decision-makers should note that the system still faces performance challenges when detecting extremely dense, very small, or exceptionally large objects. Addressing these edge cases will require integrating multi-scale feature representations in future iterations.

Cover for DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

Abstract

We present in this paper a novel query formulation using dynamic anchor boxes for DETR (DEtection TRansformer) and offer a deeper understanding of the role of queries in DETR. This new formulation directly uses box coordinates as queries in Transformer decoders and dynamically updates them layer-by-layer. Using box coordinates not only helps using explicit positional priors to improve the query-to-feature similarity and eliminate the slow training convergence issue in DETR, but also allows us to modulate the positional attention map using the box width and height information. Such a design makes it clear that queries in DETR can be implemented as performing soft ROI pooling layer-by-layer in a cascade manner. As a result, it leads to the best performance on MS-COCO benchmark among the DETR-like detection models under the same setting, e.g., AP 45.7% using ResNet50-DC5 as backbone trained in 50 epochs. We also conducted extensive experiments to confirm our analysis and verify the effectiveness of our methods. Code is available at \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Why a Positional Prior could Speedup Training?
  • 4 DAB-DETR
  • 4.1 Overview
  • 4.2 Learning Anchor Boxes Directly
  • 4.3 Anchor Update
  • 4.4 Width & Height-Modulated Gaussian Kernel
  • 4.5 Temperature Tuning
  • 5 Experiments
  • 5.1 Main Results
  • 5.2 Ablations
  • 6 Conclusion
  • References
  • A Training Details
  • B Comparison of DETR-like models
  • C DAB-Deformable-DETR
  • D Anchors Visualization
  • E Results with Different Temperatures
  • F Results with less Decoder Layers
  • G Fixed x,yx,y For Better Performance
  • H Comparison of Box Update
  • I Analysis of Failure Cases
  • J Comparison of Runtime
  • K Comparison of Model Convergence

Knowls

  1. Knowl 1 — Dynamic Anchor Box Query Formulation in DAB-DETR

    model/method

    DAB-DETR formulates decoder queries directly as 4D bounding box coordinates Aq=(xq,yq,wq,hq)∈[0,1]4A_q = (x_q, y_q, w_q, h_q) \in [0, 1]^4, where (xq,yq)(x_q, y_q) denotes the center position and (wq,hq)(w_q, h_q) denotes the width and height of the qq-th anchor box. Unlike vanilla DETR which uses unconstrained learnable vectors as queries, DAB-DETR decouples each decoder query into two explicit components:

    1. A positional query Pq∈RDP_q \in \mathbb{R}^D generated directly from the 4D anchor coordinates AqA_q via sinusoidal positional encoding and a multi-layer perceptron (MLP).
    2. A content query (decoder embedding) Cq∈RDC_q \in \mathbb{R}^D initialized to zero (or learned pattern embeddings) and updated through self-attention and cross-attention layers.

    Using explicit 4D anchor boxes provides interpretable spatial priors for the cross-attention module: the center coordinates (xq,yq)(x_q, y_q) localize the feature pooling region around the target object, while the dimensions (wq,hq)(w_q, h_q) modulate the spatial extent of the positional attention map. This formulation enables layer-by-layer dynamic updates of coordinates across stacked decoder layers, transforming cross-attention into a continuous, soft region-of-interest (ROI) feature pooling mechanism.

  2. Knowl 2 — Positional Query Generation and Self-Attention Formulation in DAB-DETR

    equation

    In DAB-DETR, given a 4D anchor box Aq=(xq,yq,wq,hq)∈[0,1]4A_q = (x_q, y_q, w_q, h_q) \in [0, 1]^4, the positional query Pq∈RDP_q \in \mathbb{R}^D is generated by:

    Pq=MLP(PE(Aq))P_q = \text{MLP}(\text{PE}(A_q))

    where PE\text{PE} applies sinusoidal positional encoding to each coordinate independently and concatenates them:

    PE(Aq)=Cat(PE(xq),PE(yq),PE(wq),PE(hq))\text{PE}(A_q) = \text{Cat}(\text{PE}(x_q), \text{PE}(y_q), \text{PE}(w_q), \text{PE}(h_q))

    Here, PE:R→RD/2\text{PE}: \mathbb{R} \to \mathbb{R}^{D/2}, producing a concatenated vector in R2D\mathbb{R}^{2D}. The MLP:R2D→RD\text{MLP}: \mathbb{R}^{2D} \to \mathbb{R}^D consists of two linear layers with ReLU\text{ReLU} activations, where dimensionality reduction from 2D2D to DD occurs in the first linear layer, and parameters are shared across all decoder layers.

    In the decoder self-attention module, the queries QqQ_q, keys KqK_q, and values VqV_q for the qq-th query are formed by combining content queries Cq∈RDC_q \in \mathbb{R}^D and positional queries Pq∈RDP_q \in \mathbb{R}^D:

    Qq=Cq+Pq,Kq=Cq+Pq,Vq=CqQ_q = C_q + P_q, \quad K_q = C_q + P_q, \quad V_q = C_q

  3. Knowl 3 — Width- and Height-Modulated Cross-Attention Formulation

    equation

    In DAB-DETR, cross-attention decouples content and positional components by concatenating content features and 2D positional embeddings into queries QqQ_q and keys Kx,yK_{x,y}:

    Qq=Cat(Cq,PE(xq,yq)⋅MLP(csq)(Cq))Q_q = \text{Cat}\left(C_q, \text{PE}(x_q, y_q) \cdot \text{MLP}^{(\text{csq})}(C_q)\right)

    Kx,y=Cat(Fx,y,PE(x,y)),Vx,y=Fx,yK_{x,y} = \text{Cat}(F_{x,y}, \text{PE}(x, y)), \quad V_{x,y} = F_{x,y}

    where Fx,y∈RDF_{x,y} \in \mathbb{R}^D is the image feature at spatial position (x,y)(x, y), PE\text{PE} is sinusoidal positional encoding, ⋅\cdot denotes element-wise multiplication, and MLP(csq):RD→RD\text{MLP}^{(\text{csq})}: \mathbb{R}^D \to \mathbb{R}^D generates a content-dependent scale vector.

    To account for objects of varying scales, the positional query-to-key attention similarity (computed prior to the softmax operation) between query anchor (xq,yq,wq,hq)(x_q, y_q, w_q, h_q) and image coordinate (x,y)(x, y) is modulated by the relative width and height:

    ModulateAttn((x,y),(xq,yq))=1D(PE(x)⋅PE(xq)wq,refwq+PE(y)⋅PE(yq)hq,refhq)\text{ModulateAttn}((x, y), (x_q, y_q)) = \frac{1}{\sqrt{D}} \left( \text{PE}(x) \cdot \text{PE}(x_q) \frac{w_{q,\text{ref}}}{w_q} + \text{PE}(y) \cdot \text{PE}(y_q) \frac{h_{q,\text{ref}}}{h_q} \right)

    where wqw_q and hqh_q are the width and height of anchor AqA_q, and wq,ref,hq,ref=σ(MLP(Cq))∈(0,1)w_{q,\text{ref}}, h_{q,\text{ref}} = \sigma(\text{MLP}(C_q)) \in (0, 1) are reference dimensions predicted dynamically from content features CqC_q via a sigmoid activation σ\sigma.

  4. Knowl 4 — Sinusoidal Positional Encoding Temperature Tuning for Vision Coordinates

    equation

    In DAB-DETR, float coordinate values x∈[0,1]x \in [0, 1] are mapped to sinusoidal positional embeddings of dimension D/2D/2 using a temperature parameter TT:

    PE(x)2i=sin⁡(xT2i/D),PE(x)2i+1=cos⁡(xT2i/D)\text{PE}(x)_{2i} = \sin\left(\frac{x}{T^{2i/D}}\right), \quad \text{PE}(x)_{2i+1} = \cos\left(\frac{x}{T^{2i/D}}\right)

    where i∈{0,1,…,D/4−1}i \in \{0, 1, \dots, D/4 - 1\} indexes the embedding channels. While standard Natural Language Processing models set T=10000T = 10000 for integer token indices, normalized vision coordinates x∈[0,1]x \in [0, 1] with T=10000T = 10000 produce overly flattened, diffuse spatial attention maps. Setting T=20T = 20 yields a localized Gaussian-like positional prior appropriate for bounding box coordinates on image feature maps.

  5. Knowl 5 — Cascaded Layer-by-Layer Dynamic Anchor Box Update

    model/method

    DAB-DETR refines bounding box predictions across stacked Transformer decoder layers in a cascade manner. In decoder layer l∈{1,…,6}l \in \{1, \dots, 6\}, the output features of the feed-forward network are passed to a prediction head (an MLP shared across all layers) that outputs relative coordinate offsets (Δxq(l),Δyq(l),Δwq(l),Δhq(l))(\Delta x_q^{(l)}, \Delta y_q^{(l)}, \Delta w_q^{(l)}, \Delta h_q^{(l)}).

    The anchor box for query qq is updated directly in coordinate space:

    Aq(l)=Aq(l−1)+(Δxq(l),Δyq(l),Δwq(l),Δhq(l))A_q^{(l)} = A_q^{(l-1)} + \left(\Delta x_q^{(l)}, \Delta y_q^{(l)}, \Delta w_q^{(l)}, \Delta h_q^{(l)}\right)

    where Aq(0)A_q^{(0)} represents the initial learnable (or fixed) anchor box parameter. The updated anchor Aq(l)A_q^{(l)} is then fed into decoder layer l+1l+1 to regenerate the positional query Pq(l+1)P_q^{(l+1)} and modulate the subsequent cross-attention map, ensuring that both self-attention and cross-attention operate on progressively refined spatial positions.

  6. Knowl 6 — Diagnosis of DETR Convergence via Query Positional Attention Modes

    empirical result

    The slow convergence of vanilla DETR is primarily caused by the multi-modal and overly diffuse attention distributions of its unconstrained learnable positional queries in cross-attention decoder layers. When visualizing attention maps between learned queries and image positional embeddings in vanilla DETR, individual queries frequently exhibit multiple concentration peaks or nearly uniform spatial weights, failing to constrain feature pooling to a local region of interest.

    Reusing well-trained queries from a converged DETR and keeping them fixed during training only improves convergence speed in the first 25 epochs and does not resolve the overall convergence rate, indicating that query optimization difficulty is not the primary bottleneck. In contrast, replacing unconstrained queries with dynamic anchor boxes (DAB) enforces unimodal, localized spatial priors, significantly speeding up training convergence and improving validation detection loss.

  7. Knowl 7 — Benchmark Detection Performance of DAB-DETR on MS COCO 2017

    data/table

    DAB-DETR outperforms existing DETR variants and classical detectors on the MS COCO 2017 validation set under equivalent 50-epoch training schedules. Models with superscript ∗* utilize 3 pattern embeddings per anchor position.

    Model Backbone Epochs AP AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    DETR R50 500 42.0 62.4 44.2 20.5 45.8 61.1
    Faster RCNN-FPN R50 108 42.0 62.1 45.5 26.6 45.5 53.4
    Anchor DETR∗^* R50 50 42.1 63.1 44.9 22.3 46.2 60.0
    Conditional DETR R50 50 40.9 61.8 43.3 20.8 44.6 59.2
    DAB-DETR R50 50 42.2 63.1 44.7 21.5 45.7 60.3
    DAB-DETR∗^* R50 50 42.6 63.2 45.6 21.8 46.2 61.1
    DETR-DC5 R50 500 43.3 63.1 45.9 22.5 47.3 61.1
    Deformable DETR (multi-scale) R50 50 43.8 62.6 47.7 26.4 47.1 58.0
    SMCA (multi-scale) R50 50 43.7 63.6 47.2 24.2 47.0 60.4
    Conditional DETR-DC5 R50 50 43.8 64.4 46.7 24.0 47.6 60.7
    Anchor DETR-DC5∗^* R50 50 44.2 64.7 47.5 24.7 48.2 60.6
    DAB-DETR-DC5 R50 50 44.5 65.1 47.7 25.3 48.2 62.3
    DAB-DETR-DC5∗^* R50 50 45.7 66.2 49.0 26.1 49.4 63.1
    DETR R101 500 43.5 63.8 46.4 21.9 48.0 61.8
    Conditional DETR R101 50 42.8 63.7 46.0 21.7 46.6 60.9
    DAB-DETR R101 50 43.5 63.9 46.6 23.6 47.3 61.5
    DAB-DETR∗^* R101 50 44.1 64.7 47.2 24.1 48.2 62.9
    DETR-DC5 R101 500 44.9 64.7 47.7 23.7 49.5 62.3
    Conditional DETR-DC5 R101 50 45.0 65.5 48.4 26.1 48.9 62.8
    DAB-DETR-DC5 R101 50 45.8 65.9 49.3 27.0 49.8 63.8
    DAB-DETR-DC5∗^* R101 50 46.6 67.0 50.2 28.1 50.5 64.1
  8. Knowl 8 — Ablation Analysis of DAB-DETR Components

    data/table

    Ablation results on the MS COCO validation set using a ResNet-50-DC5 backbone demonstrate the specific performance contribution of each design component in DAB-DETR:

    Anchor Representation Anchor Update whwh-Modulated Attention Temperature Tuning AP (%)
    4D Box Yes Yes Yes 45.7
    4D Box No Yes Yes 44.0
    4D Box Yes No Yes 45.0
    2D Point Yes No Yes 44.0
    4D Box Yes Yes No 44.4

    Key takeaways from the ablation:

    1. Dynamic anchor updating provides a +1.7%+1.7\% AP improvement over fixed initial anchor representations.
    2. Formulating queries as 4D anchor boxes yields a +1.0%+1.0\% AP improvement over 2D anchor point queries (45.0% vs. 44.0% AP without modulation).
    3. Width and height modulation of the cross-attention map adds +0.7%+0.7\% AP (from 45.0% to 45.7% AP).
    4. Tuning the positional encoding temperature to T=20T=20 adds +1.3%+1.3\% AP over the default NLP temperature setting (from 44.4% to 45.7% AP).
  9. Knowl 9 — DAB-Deformable-DETR Formulation and Performance

    empirical result

    Integrating dynamic anchor boxes into Deformable DETR (termed DAB-Deformable-DETR) updates positional queries across decoder layers and passes refined anchor positions to both self-attention and deformable cross-attention modules, requiring under 10 lines of code modification.

    On MS COCO 2017 val with a standard ResNet-50 multi-scale backbone:

    • Original Deformable DETR with iterative bounding box refinement achieves 46.3%46.3\% AP (65.3% AP5065.3\% \text{ AP}_{50}, 50.2% AP7550.2\% \text{ AP}_{75}, 28.6% APS28.6\% \text{ AP}_S, 49.3% APM49.3\% \text{ AP}_M, 62.1% APL62.1\% \text{ AP}_L).
    • DAB-Deformable-DETR improves performance to 46.8%46.8\% AP (66.0% AP5066.0\% \text{ AP}_{50}, 50.4% AP7550.4\% \text{ AP}_{75}, 29.1% APS29.1\% \text{ AP}_S, 49.8% APM49.8\% \text{ AP}_M, 62.3% APL62.3\% \text{ AP}_L).

    During training, DAB-Deformable-DETR exhibits a higher sum of losses across all decoder layers but achieves a lower loss on the final decoder layer compared to Deformable DETR, leading to better final inference accuracy.

  10. Knowl 10 — Impact of Positional Encoding Temperature on Detection Across Object Scales

    empirical result

    The temperature TT in the sinusoidal positional encoding function controls the spatial dispersion of the Gaussian-like positional prior, directly impacting detection performance across different object scales:

    Temperature (TT) AP AP50\text{AP}_{50} AP75\text{AP}_{75} APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    2 39.6 60.7 41.9 19.3 43.3 58.0
    5 40.0 61.1 42.1 19.5 43.4 58.9
    10 40.0 61.1 42.3 19.7 43.5 59.3
    20 40.1 61.1 42.8 19.8 43.7 58.6
    50 39.8 61.0 42.2 19.7 43.2 58.8
    100 39.8 60.8 42.1 19.3 43.3 58.4
    10000 39.5 60.7 41.7 18.9 42.6 58.9

    A small temperature (e.g., T=2T = 2) produces highly concentrated attention maps that favor small and medium objects (APS=19.3\text{AP}_S = 19.3, APM=43.3\text{AP}_M = 43.3). A large temperature (e.g., T=10000T = 10000) produces flatter attention maps that favor large objects (APL=58.9\text{AP}_L = 58.9). Setting T=20T = 20 achieves the best overall trade-off (40.1%40.1\% AP).

  11. Knowl 11 — Effect of Freezing Initial Anchor Center Coordinates

    empirical result

    Fixing the randomly initialized anchor center coordinates (xq,yq)(x_q, y_q) in the first decoder layer (preventing gradient updates to (x,y)(x, y) in layer 1, while allowing them to be dynamically updated in subsequent layers 22 to 66) consistently improves detection AP across multiple backbones:

    • DAB-DETR-R50∗^*: increases from 42.6%42.6\% to 42.9%42.9\% AP (+0.3%+0.3\% AP).
    • DAB-DETR-DC5-R50: increases from 44.5%44.5\% to 44.7%44.7\% AP (+0.2%+0.2\% AP).
    • DAB-DETR-DC5-R50∗^*: increases from 45.7%45.7\% to 45.8%45.8\% AP (+0.1%+0.1\% AP).
    • DAB-DETR-R101∗^*: increases from 44.1%44.1\% to 44.8%44.8\% AP (+0.7%+0.7\% AP).
    • DAB-DETR-DC5-R101∗^*: increases from 46.6%46.6\% to 46.7%46.7\% AP (+0.1%+0.1\% AP).

    Preventing the initial anchor centers from overfitting to training data distributions provides a regularizing effect that enhances generalization.

  12. Knowl 12 — Limitations of DAB-DETR on Dense and Extreme Scale Objects

    limitation

    When evaluated under its standard single-scale feature map configuration, DAB-DETR exhibits degraded detection accuracy in three specific visual scenarios:

    1. Dense and heavily crowded object configurations, where spatial anchor overlapping can lead to ambiguous bipartite matching.
    2. Extremely small objects, where single-scale feature resolution is insufficient for fine-grained feature extraction.
    3. Extremely large objects spanning most of the image frame, where localized positional priors may under-sample full contextual cues.

    Addressing these cases requires incorporating multi-scale feature hierarchies (such as Feature Pyramid Networks or multi-scale deformable attention mechanisms) into the dynamic anchor box architecture.

Coverage note — None was omitted; all contributed models, mathematical formulations (query projection, attention modulation, temperature calibration), empirical diagnoses, COCO benchmarks, component ablations, and extensions (DAB-Deformable-DETR) are fully captured.

References

  1. 1.Alexey Bochkovskiy, Chien-Yao Wang, and Hong-Yuan Mark Liao. Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934, 2020.
  2. 2.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, pp. 213–229. Springer, 2020.
  3. 3.Xiyang Dai, Yinpeng Chen, Jianwei Yang, Pengchuan Zhang, Lu Yuan, and Lei Zhang. Dynamic detr: End-to-end object detection with dynamic attention. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 2988–2997, October 2021.
  4. 4.Peng Gao, Minghang Zheng, Xiaogang Wang, Jifeng Dai, and Hongsheng Li. Fast convergence of detr with spatially modulated co-attention. arXiv preprint arXiv:2101.07448, 2021.
  5. 5.Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, and Jian Sun. Yolox: Exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430, 2021.
  6. 6.Ross Girshick. Fast r-cnn. In 2015 IEEE International Conference on Computer Vision (ICCV), pp. 1440–1448, 2015.
  7. 7.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In 2015 IEEE International Conference on Computer Vision (ICCV), pp. 1026–1034, 2015.
  8. 8.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2016.
  9. 9.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pp. 740–755. Springer, 2014.
  10. 10.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(2):318–327, 2020.
  11. 11.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2018.
  12. 12.Depu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng, Houqiang Li, Yuhui Yuan, Lei Sun, and Jingdong Wang. Conditional detr for fast training convergence. arXiv preprint arXiv:2108.06152, 2021.
  13. 13.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 779–788, 2016.
  14. 14.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(6):1137–1149, 2017.
  15. 15.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 658–666, 2019.
  16. 16.Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, and Ping Luo. Sparse r-cnn: End-to-end object detection with learnable proposals. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14454–14463, 2021.
  17. 17.Zhiqing Sun, Shengcao Cao, Yiming Yang, and Kris Kitani. Rethinking transformer-based set prediction for object detection. arXiv preprint arXiv:2011.10881, 2020.
  18. 18.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9627–9636, 2019.
  19. 19.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.
  20. 20.Yingming Wang, Xiangyu Zhang, Tong Yang, and Jian Sun. Anchor detr: Query design for transformer-based detector. arXiv preprint arXiv:2109.07107, 2021.
  21. 21.Zhuyu Yao, Jiangbo Ai, Boxun Li, and Chi Zhang. Efficient detr: Improving end-to-end object detector with dense prior. arXiv preprint arXiv:2104.01318, 2021.
  22. 22.Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl. Objects as points. arXiv preprint arXiv:1904.07850, 2019.
  23. 23.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In ICLR 2021: The Ninth International Conference on Learning Representations, 2021.

Citation

MLA
Liu, S., et al. “DAB-DETR: Dynamic Anchor Boxes Are Better Queries for DETR”. arXiv, 2022, http://arxiv.org/abs/2201.12329v4.
APA
Liu, S., Li, F., Zhang, H., Yang, X., Qi, X., Su, H., Zhu, J., & Zhang, L. (2022). DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR. arXiv. http://arxiv.org/abs/2201.12329v4
Chicago
Liu, S., F. Li, H. Zhang, et al. 2022. “DAB-DETR: Dynamic Anchor Boxes Are Better Queries for DETR”. arXiv. http://arxiv.org/abs/2201.12329v4.
Harvard
Liu, S. et al. (2022) “DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2201.12329v4.
Vancouver
1. Liu S, Li F, Zhang H, Yang X, Qi X, Su H, Zhu J, Zhang L (2022) DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR. arXiv

BibTeX

@article{liu2022dab,
  title = {DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR},
  author = {Liu, Shilong and Li, Feng and Zhang, Hao and Yang, Xiao and Qi, Xianbiao and Su, Hang and Zhu, Jun and Zhang, Lei},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2201.12329v4},
  eprint = {2201.12329}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors