DETRs with Hybrid Matching

Ding JiaYuhui YuanHaodi HeXiaopei WuHaojun YuWeihong LinLei SunChao ZhangHan Hu

article2023CVPR339 citations

Proposes a hybrid matching strategy for DETR architectures that trains an auxiliary one-to-many matching branch alongside the standard one-to-one branch, accelerating training and boosting detection accuracy across diverse vision tasks without adding inference overhead.

Listen

Modern computer vision models increasingly rely on transformer-based architectures that directly predict target objects in an end-to-end manner. To eliminate the need for hand-crafted post-processing steps like non-maximum suppression (a technique that removes duplicate detections), these models strictly assign only one internal query to each real-world target during training. However, this one-to-one matching strategy leaves the vast majority of queries without positive target supervision, which severely restricts training efficiency, slows convergence, and limits final recognition accuracy.

The article introduces and evaluates a hybrid matching framework, designated as H-DETR, designed to accelerate and improve the training of detection transformers. The main objective is to demonstrate that introducing auxiliary training signals allows transformer models to learn better spatial features while maintaining their end-to-end deployment speed and eliminating duplicate filtering.

To achieve this, the researchers evaluated a hybrid approach across multiple standard benchmark datasets and visual tasks, including 2D object detection, panoptic segmentation, 3D object detection, multi-person pose estimation, and multi-object tracking. The approach adds an auxiliary one-to-many matching branch during training—assigning multiple queries to repeated ground-truth targets—to provide richer learning signals. Crucially, this auxiliary branch is discarded during deployment, ensuring that inference relies solely on the standard one-to-one branch without additional computational latency or architectural complexity.

The findings show consistent performance gains across all evaluated vision domains without slowing evaluation speeds. In 2D object detection on the COCO benchmark, the hybrid method improved the baseline model accuracy by +1.7% with a standard backbone and achieved 59.4% accuracy with a large backbone, outperforming leading existing transformer detectors. In 3D multi-view detection on the nuScenes dataset, overall detection scores rose by +1.7%, while human pose estimation and multi-object tracking saw baseline accuracy gains of +1.6%. Ablation analyses revealed that these gains stem primarily from significantly reduced localization errors and fewer missed targets, driven largely by better optimization of the shared visual feature encoder.

These results demonstrate that the limitation of detection transformers is not their core architecture, but rather an artificial data-supervision bottleneck during training. By resolving this training deficiency, teams can achieve higher model accuracy across varied computer vision tasks without increasing runtime inference cost, cloud serving latency, or engineering complexity in production pipelines.

Organizations developing or deploying transformer-based vision systems should integrate hybrid branch matching into their model training pipelines as a drop-in enhancement. For immediate implementation, teams should adopt the hybrid branch structure over alternate layer- or epoch-based variations, setting the auxiliary query repetition factor to at least four to six times the ground truth to ensure high-quality supervision signals. Future work should focus on implementing auxiliary matching calculations directly on graphics hardware to reduce remaining training-time overhead.

While the evaluation shows high confidence and consistent gains across varied model sizes and visual tasks, the primary limitation is a modest increase in training time (roughly 6% to 23%) and higher training-stage memory usage due to extra auxiliary queries. However, because memory-saving attention mechanisms can mitigate training memory and the auxiliary queries are entirely omitted during inference, the reported operational performance improvements remain robust.

  • Paper: DETRs Beat YOLOs on Real-time Object Detection, Yian Zhao et al. (2024). This paper advances real-time end-to-end DETR architectures, building upon the principles of efficient matching and query optimization in Transformer-based detectors.
  • Paper: YOLOv10: Real-Time End-to-End Object Detection, Ao Wang et al. (2024). This work extends dual label-assignment schemes combining one-to-many and one-to-one matching branches to create NMS-free real-time detectors.
Cover for DETRs with Hybrid Matching

Abstract

One-to-one set matching is a key design for DETR to establish its end-to-end capability, so that object detection does not require a hand-crafted NMS (non-maximum suppression) to remove duplicate detections. This end-to-end signature is important for the versatility of DETR, and it has been generalized to broader vision tasks. However, we note that there are few queries assigned as positive samples and the one-to-one set matching significantly reduces the training efficacy of positive samples. We propose a simple yet effective method based on a hybrid matching scheme that combines the original one-to-one matching branch with an auxiliary one-to-many matching branch during training. Our hybrid strategy has been shown to significantly improve accuracy. In inference, only the original one-to-one match branch is used, thus maintaining the end-to-end merit and the same inference efficiency of DETR. The method is named H-DETR, and it shows that a wide range of representative DETR methods can be consistently improved across a wide range of visual tasks, including Deformable-DETR, PETRv2, PETR, and TransTrack, among others. Code is available at: https://github.com/HDETR.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Our Approach
  • 3.1. Preliminary
  • 3.2. Hybrid Matching
  • 3.2.1 Hybrid Branch Scheme
  • 3.2.2 More Variants of Hybrid Matching
  • 4. Experiment
  • 4.1. Improving DETR-based Approaches
  • 4.2. Ablation Study
  • 4.3. Comparison with State-of-the-art
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Hybrid one-to-one and one-to-many matching

    model/method

    H-DETR addresses the low training efficiency of DETR-style one-to-one assignment. In conventional DETR, Hungarian matching assigns at most one query to each ground-truth instance, so most decoder queries receive only negative classification supervision and few queries receive localization supervision.

    H-DETR adds an auxiliary one-to-many matching branch during training. The original query group uses one-to-one matching and preserves DETR’s end-to-end prediction property; the auxiliary query group matches against repeated copies of each ground-truth instance, allowing several queries to receive localization supervision for the same object. The two branches are trained jointly, with masked self-attention preventing interactions between their query groups so that they can be processed in parallel.

    At inference time, H-DETR discards the auxiliary branch and retains only the original one-to-one branch. Consequently, inference does not require non-maximum suppression (NMS), and the inference architecture and query count of the selected DETR baseline are retained.

  2. Knowl 2 — Hybrid-branch training objective

    equation

    Let GG be the set of ground-truth instances in one training example, let LL be the number of transformer decoder layers, and let PlP^l and P~l\widetilde{P}^{l} denote the predictions from decoder layer ll of the one-to-one and one-to-many branches, respectively. The one-to-one branch uses the ordinary ground-truth set GG:

    Lone2one=∑l=1LLHungarian(Pl,G).\mathcal{L}_{\mathrm{one2one}}=\sum_{l=1}^{L}\mathcal{L}_{\mathrm{Hungarian}}(P^l,G).

    For the one-to-many branch, every ground-truth instance is repeated KK times to form G~={G1,…,GK}\widetilde{G}=\{G^1,\ldots,G^K\}, where G1=⋯=GK=GG^1=\cdots=G^K=G. Hungarian matching between the auxiliary predictions and this augmented target set gives:

    Lone2many=∑l=1LLHungarian(P~l,G~).\mathcal{L}_{\mathrm{one2many}}=\sum_{l=1}^{L}\mathcal{L}_{\mathrm{Hungarian}}(\widetilde{P}^{l},\widetilde{G}).

    The total training loss is

    Ltotal=Lone2one+λLone2many,\mathcal{L}_{\mathrm{total}}=\mathcal{L}_{\mathrm{one2one}}+\lambda\mathcal{L}_{\mathrm{one2many}},

    where λ≥0\lambda\geq 0 weights the auxiliary loss. Each Hungarian loss follows DETR and combines classification, L1L_1 box-regression, and generalized-IoU losses. Matching is performed independently at every decoder layer.

  3. Knowl 3 — Hybrid epoch and hybrid layer variants

    model/method

    H-DETR also defines two single-query-group variants of hybrid matching.

    In the hybrid-epoch scheme, one query group is trained with one-to-many matching for the first fraction ρ\rho of training epochs. The ground-truth set is repeated KeK_e times during this stage. During the remaining fraction 1−ρ1-\rho of epochs, the same query group is trained with ordinary one-to-one matching. The query group is used directly at inference without NMS.

    In the hybrid-layer scheme, one-to-many matching supervises the outputs of the first L1L_1 decoder layers, while one-to-one matching supervises the outputs of the remaining L2L_2 layers, with L=L1+L2L=L_1+L_2. Both losses are applied throughout training.

    The reported settings were: hybrid branch with 300 one-to-one queries, 1500 auxiliary queries, and K=6K=6; hybrid epoch with 1800 queries, Ke=10K_e=10, and ρ=2/3\rho=2/3; and hybrid layer with 1800 queries, K=10K=10, L1=4L_1=4, and L2=2L_2=2. The hybrid branch was selected as the default because it provided the best training-time and inference-speed trade-off.

  4. Knowl 4 — Consistent gains on COCO and LVIS object detection

    data/table

    The following results compare Deformable-DETR with its H-DETR version using the same backbone and training schedule. AP is average precision; APSAP_S, APMAP_M, and APLAP_L are AP for small, medium, and large objects. H-DETR improves every listed configuration.

    Could not parse LaTeX table

    The COCO AP gains range from +0.8+0.8 to +1.7+1.7 points across backbones and schedules, while the LVIS gains are +1.3+1.3 points with ResNet-50 and +0.9+0.9 points with Swin-L.

  5. Knowl 5 — Generalization to segmentation, pose estimation, 3D detection, and tracking

    data/table

    The hybrid matching strategy transfers to DETR-based tasks whose queries represent masks, human poses, 3D objects, or tracks. The reported representative comparisons are:

    Could not parse LaTeX table

    For the nuScenes comparison, mAP also increases from 41.0741.07 to 42.5942.59. For MOT17 validation, the false-negative count decreases from 1568015680 to 1365713657, while IDF1 increases from 68.168.1 to 68.368.3. On COCO pose estimation, the Swin-L result increases from 73.373.3 AP to 74.974.9 AP, and the ResNet-101 result increases from 69.969.9 AP to 71.071.0 AP. These results support the paper’s claim that the matching modification is not specific to 2D bounding-box detection.

  6. Knowl 6 — Hybrid branch is the most favorable matching variant

    data/table

    Ablations on a two-stage Deformable-DETR baseline compare the three hybrid schemes under matched query and positive-sample budgets. AP is reported after 12, 24, and 36 training epochs; inference GFLOPs and inference FPS are measured on the same V100 GPU.

    Could not parse LaTeX table

    All three hybrid schemes outperform their corresponding conventional baselines. The hybrid branch reaches the same 6.7 FPS as the 300-query baseline and has the lowest average training time among the hybrid variants, whereas the epoch and layer schemes evaluate 1800 queries and therefore have lower inference speed.

  7. Knowl 7 — Effect of auxiliary replication, query count, and parameter sharing

    data/table

    On COCO 2017 validation with a 12-epoch two-stage Deformable-DETR baseline, the default hybrid branch uses K=6K=6 repeated ground-truth copies and T=1500T=1500 auxiliary queries. The number of auxiliary queries is set to T=300KT=300K in the replication experiment.

    Could not parse LaTeX table

    With K=6K=6, changing the auxiliary query count TT gives AP values 47.847.8, 48.348.3, 48.448.4, 48.448.4, 48.748.7, and 48.648.6 for T=300T=300, 600600, 900900, 12001200, 15001500, and 18001800, respectively. Thus, the reported default K=6K=6, T=1500T=1500 is near the best tested setting; gains become consistent when KK exceeds 3.

    Parameter-sharing ablations give AP values of 48.748.7 when the encoder, decoder, box head, and classification head are all shared; 48.648.6 when only the classification head is independent; 48.548.5 when both prediction heads are independent; 48.348.3 when the decoder and both heads are independent; and 47.347.3 when encoder, decoder, and heads are all independent. The results indicate that retaining a shared transformer encoder is particularly important for the reported performance.

  8. Knowl 8 — Hybrid matching improves localization and false-negative errors

    empirical result

    A component ablation on COCO 2017 validation under 12 training epochs starts from a two-stage Deformable-DETR with AP 43.343.3. Increasing the feed-forward-network dimension gives AP 43.743.7; additionally setting transformer dropout to zero gives 44.344.3; adding mixed query selection gives 46.346.3; adding the look-forward-twice refinement gives 47.047.0; and finally adding hybrid matching gives 48.748.7.

    The paper’s localization-error analysis reports the following changes from Deformable-DETR to H-Deformable-DETR: AP increases from 47.047.0 to 48.748.7, overall oLRP decreases from 62.662.6 to 61.261.2, oLRP for localization errors decreases from 13.813.8 to 13.313.3, and oLRP for false negatives decreases from 41.041.0 to 39.439.4. Since lower oLRP is better, the results attribute the improvement mainly to more accurate localization of matched detections and fewer unmatched ground-truth instances.

  9. Knowl 9 — System-level COCO result with a Swin-L backbone

    empirical result

    With single-scale evaluation, a Swin-L backbone, 1333×8001333\times800 input resolution, and 36 training epochs, H-Deformable-DETR obtains 59.459.4 AP on COCO validation. Its detailed scores are 77.877.8 AP50_{50}, 65.465.4 AP75_{75}, 43.143.1 AP for small objects, 63.163.1 AP for medium objects, and 74.274.2 AP for large objects.

    Under the same reported framework, backbone, input size, and schedule, the comparison systems Group-DETR and DINO-DETR obtain 58.458.4 and 58.558.5 AP, respectively. The H-Deformable-DETR result is the strongest among the listed single-scale systems, although it includes additional enhancement techniques used in the corresponding system-level comparison.

  10. Knowl 10 — Inference efficiency is preserved, but training and memory costs increase

    limitation

    The principal deployment advantage of H-DETR is preserved because only the one-to-one branch is evaluated. On COCO with the reported Deformable-DETR setup, the one-to-one branch gives AP values 48.748.7, 49.949.9, and 50.050.0 at 12, 24, and 36 epochs without NMS, at 6.7 FPS and 80 minutes of average training time. Using only one-to-many predictions without NMS is ineffective, giving AP values 13.513.5, 13.113.1, and 12.912.9; applying NMS restores AP to 48.648.6, 49.849.8, and 49.949.9 but reduces inference speed to 5.4 FPS. Training only with one-to-many matching and using NMS reaches 49.449.4, 50.250.2, and 48.848.8 AP, but requires 95 minutes and 5.3 FPS.

    The auxiliary branch is not free during training. With ResNet-50, increasing from 300 baseline queries to the default 300 one-to-one plus 1500 auxiliary queries increases computation from 268.19 to 282.39 GFLOPs, training time from 65 to 80 minutes, and GPU memory from 5480 MB to 7728 MB. With Swin-L, the corresponding values increase from 912.29 to 926.48 GFLOPs, 202 to 215 minutes, and 8955 MB to 11203 MB. The paper attributes most memory growth to naive self- and cross-attention and most training-time growth to CPU Hungarian matching and cost/loss computation; optimized attention implementations and GPU matching are suggested as ways to reduce these overheads.

Coverage note — The supplementary-only full panoptic-segmentation table and additional loss-curve, precision-recall, and query-selection ablations were omitted because the main paper provides only summary statements for them and they are secondary to the core hybrid-matching method and its principal evaluations.

References

  1. 1.Xuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang, Yilun Chen, Hongbo Fu, and Chiew-Lan Tai. Transfusion: Robust lidar-camera fusion for 3d object detection with transformers. In CVPR, 2022. 1, 2
  2. 2.Guillem Braso, Nikita Kister, and Laura Leal-Taix ´ e. The center of attention: Center-keypoint grouping via attention for multi-person pose estimation. In ICCV, 2021. 2
  3. 3.Xipeng Cao, Peng Yuan, Bailan Feng, Kun Niu, and Yao Zhao. Cf-detr: Coarse-to-fine transformers for end-to-end object detection. In AAAI, 2022. 1, 2
  4. 4.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 1, 2, 3
  5. 5.Qiang Chen, Xiaokang Chen, Gang Zeng, and Jingdong Wang. Group detr: Fast training convergence with decoupled one-to-many label assignment. arXiv:2207.13085, 2022. 8
  6. 6.Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. Transformer tracking. In CVPR, 2021. 1
  7. 7.Bowen Cheng, Anwesa Choudhuri, Ishan Misra, Alexander Kirillov, Rohit Girdhar, and Alexander G Schwing. Mask2former for video instance segmentation. arXiv:2112.10764, 2021. 1, 2
  8. 8.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. arXiv:2112.01527, 2021. 1, 2, 3
  9. 9.Xiyang Dai, Yinpeng Chen, Jianwei Yang, Pengchuan Zhang, Lu Yuan, and Lei Zhang. Dynamic detr: End-to-end object detection with dynamic attention. In ICCV, 2021. 2
  10. 10.Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Re. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In NeurIPS, 2022. 7
  11. 11.Bin Dong, Fangao Zeng, Tiancai Wang, Xiangyu Zhang, and Yichen Wei. Solq: Segmenting objects by learning queries. NeurIPS, 2021. 1
  12. 12.Yuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu. Instances as queries. In ICCV, 2021. 1
  13. 13.Peng Gao, Minghang Zheng, Xiaogang Wang, Jifeng Dai, and Hongsheng Li. Fast convergence of detr with spatially modulated co-attention. In ICCV, 2021. 1, 2
  14. 14.Ziteng Gao, Limin Wang, Bing Han, and Sheng Guo. Adamixer: A fast-converging query-based object detector. In CVPR, 2022. 1
  15. 15.Ross Girshick. Fast r-cnn. In ICCV, 2015. 2
  16. 16.Brent A Griffin and Jason J Corso. Depth from camera motion and object detection. In CVPR, 2021. 1
  17. 17.Junjie Huang and Guan Huang. Bevdet4d: Exploit temporal cues in multi-camera 3d object detection. arXiv:2203.17054, 2022. 2
  18. 18.Junjie Huang, Guan Huang, Zheng Zhu, and Dalong Du. Bevdet: High-performance multi-camera 3d object detection in bird-eye-view. arXiv:2112.11790, 2021. 2
  19. 19.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In ICCV, 2021. 1
  20. 20.Youngwan Lee and Jongyoul Park. Centermask: Real-time anchor-free instance segmentation. In CVPR, 2020. 2
  21. 21.Feng Li, Hao Zhang, Shilong Liu, Jian Guo, Lionel M Ni, and Lei Zhang. Dn-detr: Accelerate detr training by introducing query denoising. arXiv:2203.01305, 2022. 1, 2, 3
  22. 22.Feng Li, Hao Zhang, Shilong Liu, Lei Zhang, Lionel M Ni, Heung-Yeung Shum, et al. Mask dino: Towards a unified transformer-based framework for object detection and segmentation. arXiv:2206.02777, 2022. 1, 2
  23. 23.Ke Li, Shijie Wang, Xiang Zhang, Yifan Xu, Weijian Xu, and Zhuowen Tu. Pose recognition with cascade transformers. In CVPR, 2021. 1, 2
  24. 24.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Improved multiscale vision transformers for classification and detection. arXiv:2112.01526, 2021. 8
  25. 25.Yanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang, Wankou Yang, Shu-Tao Xia, and Erjin Zhou. Tokenpose: Learning keypoint tokens for human pose estimation. In ICCV, 2021. 2
  26. 26.Zhenyu Li, Zehui Chen, Xianming Liu, and Junjun Jiang. Depthformer: Exploiting long-range correlation and local information for accurate monocular depth estimation. arXiv:2203.14211, 2022. 1
  27. 27.Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, and Jifeng Dai. Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. arXiv:2203.17270, 2022. 1, 2
  28. 28.Zhiqi Li, Wenhai Wang, Enze Xie, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, Ping Luo, and Tong Lu. Panoptic segformer: Delving deeper into panoptic segmentation with transformers. In CVPR, 2022. 1, 2, 3
  29. 29.Zhenyu Li, Xuyang Wang, Xianming Liu, and Junjun Jiang. Binsformer: Revisiting adaptive bins for monocular depth estimation. arXiv:2204.00987, 2022. 3
  30. 30.Tingting Liang, Xiaojie Chu, Yudong Liu, Yongtao Wang, Zhi Tang, Wei Chu, Jingdong Chen, and Haibin Ling. Cbnet: A composite backbone network architecture for object detection. TIP, 2022. 8
  31. 31.Tingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia, Zhiwei Lin, Yongtao Wang, Tao Tang, Bing Wang, and Zhi Tang. Bevfusion: A simple and robust lidar-camera fusion framework. arXiv:2205.13790, 2022. 2
  32. 32.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ICCV, 2017. 2
  33. 33.Shilong Liu, Feng Li, Hao Zhang, Xiao Yang, Xianbiao Qi, Hang Su, Jun Zhu, and Lei Zhang. Dab-detr: Dynamic anchor boxes are better queries for detr. arXiv:2201.12329, 2022. 1, 2
  34. 34.Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In ECCV, 2016. 2
  35. 35.Yingfei Liu, Tiancai Wang, Xiangyu Zhang, and Jian Sun. Petr: Position embedding transformation for multi-view 3d object detection. arXiv:2203.05625, 2022. 2
  36. 36.Yingfei Liu, Junjie Yan, Fan Jia, Shuailin Li, Qi Gao, Tiancai Wang, Xiangyu Zhang, and Jian Sun. Petrv2: A unified framework for 3d perception from multi-camera images. arXiv:2206.01256, 2022. 1, 2, 5, 6
  37. 37.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In CVPR, 2022. 5
  38. 38.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 8
  39. 39.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In CVPR, 2022. 8
  40. 40.Zhijian Liu, Haotian Tang, Alexander Amini, Xinyu Yang, Huizi Mao, Daniela Rus, and Song Han. Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation. arXiv:2205.13542, 2022. 2
  41. 41.Ze Liu, Zheng Zhang, Yue Cao, Han Hu, and Xin Tong. Group-free 3d object detection via transformers. In ICCV, 2021. 2
  42. 42.Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, and Michalis Raptis. Towards end-to-end unified scene text detection and layout analysis. In CVPR, 2022. 1
  43. 43.Qian Lou, Yen-Chang Hsu, Burak Uzkent, Ting Hua, Yilin Shen, and Hongxia Jin. Lite-mdetr: A lightweight multi-modal detector. In CVPR, 2022. 1
  44. 44.Weian Mao, Yongtao Ge, Chunhua Shen, Zhi Tian, Xinlong Wang, and Zhibin Wang. Tfpose: Direct human pose estimation with transformers. arXiv:2103.15320, 2021. 2
  45. 45.Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer. Trackformer: Multi-object tracking with transformers. In CVPR, 2022. 1, 2, 3
  46. 46.Depu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng, Houqiang Li, Yuhui Yuan, Lei Sun, and Jingdong Wang. Conditional detr for fast training convergence. In ICCV, 2021. 1, 2
  47. 47.Ishan Misra, Rohit Girdhar, and Armand Joulin. An end-to-end transformer model for 3d object detection. In ICCV, 2021. 1, 2
  48. 48.Alexander Neubeck and Luc Van Gool. Efficient non-maximum suppression. In ICPR, 2006. 2
  49. 49.Kemal Oksuz, Baris Can Cam, Sinan Kalkan, and Emre Akbas. One metric to measure them all: Localisation recall precision (lrp) for evaluating visual detection tasks. TPAMI, 2021. 6
  50. 50.Zobeir Raisi, Mohamed A Naiel, Georges Younes, Steven Wardell, and John S Zelek. Transformer-based text detection in the wild. In CVPR, 2021. 1
  51. 51.Zobeir Raisi, Georges Younes, and John Zelek. Arbitrary shape text detection using transformers. arXiv:2202.11221, 2022. 1
  52. 52.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. NeurIPS, 2015. 2
  53. 53.Byungseok Roh, JaeWoong Shin, Wuhyun Shin, and Saehoon Kim. Sparse detr: Efficient end-to-end object detection with learnable sparsity. arXiv:2111.14330, 2021. 1
  54. 54.Dahu Shi, Xing Wei, Liangqi Li, Ye Ren, and Wenming Tan. End-to-end multi-person pose estimation with transformers. In CVPR, 2022. 1, 2, 3, 5, 6
  55. 55.Lucas Stoffl, Maxime Vidal, and Alexander Mathis. End-to-end trainable multi-instance pose estimation with transformers. arXiv:2103.12115, 2021. 1, 2
  56. 56.Peize Sun, Jinkun Cao, Yi Jiang, Rufeng Zhang, Enze Xie, Zehuan Yuan, Changhu Wang, and Ping Luo. Transtrack: Multiple object tracking with transformer. arXiv:2012.15460, 2020. 1, 2, 3, 5, 6
  57. 57.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. In ICCV, 2019. 2
  58. 58.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 2
  59. 59.Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Max-deeplab: End-to-end panoptic segmentation with mask transformers. In CVPR, 2021. 1
  60. 60.Jianfeng Wang, Lin Song, Zeming Li, Hongbin Sun, Jian Sun, and Nanning Zheng. End-to-end object detection with fully convolutional network. In CVPR, 2021. 1, 2
  61. 61.Wen Wang, Jing Zhang, Yang Cao, Yongliang Shen, and Dacheng Tao. Towards data-efficient detection transformers. arXiv:2203.09507, 2022. 2
  62. 62.Yue Wang, Vitor Campagnolo Guizilini, Tianyuan Zhang, Yilun Wang, Hang Zhao, and Justin Solomon. Detr3d: 3d object detection from multi-view images via 3d-to-2d queries. In CoRL, 2022. 1, 2
  63. 63.Yingming Wang, Xiangyu Zhang, Tong Yang, and Jian Sun. Anchor detr: Query design for transformer-based detector. arXiv:2109.07107, 2021. 1, 2
  64. 64.Jiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan, and Ping Luo. Language as queries for referring video object segmentation. arXiv:2201.00487, 2022. 1
  65. 65.Junfeng Wu, Yi Jiang, Wenqing Zhang, Xiang Bai, and Song Bai. Seqformer: a frustratingly simple model for video instance segmentation. arXiv:2112.08275, 2021. 1
  66. 66.Yifan Xu, Weijian Xu, David Cheung, and Zhuowen Tu. Line segment detection using transformers without edges. In CVPR, 2021. 1
  67. 67.Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, and Huchuan Lu. Learning spatio-temporal transformer for visual tracking. In ICCV, 2021. 2
  68. 68.Chenglin Yang, Siyuan Qiao, Qihang Yu, Xiaoding Yuan, Yukun Zhu, Alan Yuille, Hartwig Adam, and Liang-Chieh Chen. Moat: Alternating mobile convolution and attention brings strong vision models. arXiv:2210.01820, 2022. 8
  69. 69.Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip HS Torr. Lavt: Language-aware vision transformer for referring image segmentation. In CVPR, 2022. 1
  70. 70.Qihang Yu, Huiyu Wang, Dahun Kim, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Cmt-deeplab: Clustering mask transformers for panoptic segmentation. In CVPR, 2022. 1
  71. 71.Qihang Yu, Huiyu Wang, Dahun Kim, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Cmt-deeplab: Clustering mask transformers for panoptic segmentation. In CVPR, 2022. 2
  72. 72.Qihang Yu, Huiyu Wang, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. k-means mask transformer. In ECCV, 2022. 2
  73. 73.Xiaodong Yu, Dahu Shi, Xing Wei, Ye Ren, Tingqun Ye, and Wenming Tan. Soit: Segmenting objects with instance-aware transformers. arXiv:2112.11037, 2021. 1
  74. 74.Yuhui Yuan, Xilin Chen, and Jingdong Wang. Object-contextual representations for semantic segmentation. In ECCV, 2020. 1
  75. 75.Fangao Zeng, Bin Dong, Tiancai Wang, Xiangyu Zhang, and Yichen Wei. Motr: End-to-end multiple-object tracking with transformer. arXiv:2105.03247, 2021. 2
  76. 76.Gongjie Zhang, Zhipeng Luo, Yingchen Yu, Kaiwen Cui, and Shijian Lu. Accelerating DETR convergence via semantic-aligned matching. In CVPR, 2022. 1
  77. 77.Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M Ni, and Heung-Yeung Shum. Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv:2203.03605, 2022. 1, 2, 3, 7, 8
  78. 78.Jianfeng Zhang, Yujun Cai, Shuicheng Yan, Jiashi Feng, et al. Direct multi-view multi-person 3d pose estimation. NeurIPS, 2021. 2
  79. 79.Shifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei, and Stan Z Li. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection. In CVPR, 2020. 2
  80. 80.Xiang Zhang, Yongwen Su, Subarna Tripathi, and Zhuowen Tu. Text spotting transformers. In CVPR, 2022. 1
  81. 81.Moju Zhao, Kei Okada, and Masayuki Inaba. Trtr: Visual tracking with transformer. arXiv:2105.03817, 2021. 2
  82. 82.Benjin Zhu, Jianfeng Wang, Zhengkai Jiang, Fuhang Zong, Songtao Liu, Zeming Li, and Jian Sun. Autoassign: Differentiable label assignment for dense object detection. arXiv:2007.03496, 2020. 2
  83. 83.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv:2010.04159, 2020. 1, 2, 3, 5

Citation

MLA
Jia, D., et al. “DETRs with Hybrid Matching”. arXiv, 2022, http://arxiv.org/abs/2207.13080v3.
APA
Jia, D., Yuan, Y., He, H., Wu, X., Yu, H., Lin, W., Sun, L., Zhang, C., & Hu, H. (2022). DETRs with Hybrid Matching. arXiv. http://arxiv.org/abs/2207.13080v3
Chicago
Jia, D., Y. Yuan, H. He, et al. 2022. “DETRs with Hybrid Matching”. arXiv. http://arxiv.org/abs/2207.13080v3.
Harvard
Jia, D. et al. (2022) “DETRs with Hybrid Matching”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2207.13080v3.
Vancouver
1. Jia D, Yuan Y, He H, Wu X, Yu H, Lin W, Sun L, Zhang C, Hu H (2022) DETRs with Hybrid Matching. arXiv

BibTeX

@article{jia2022detrs,
  title = {DETRs with Hybrid Matching},
  author = {Jia, Ding and Yuan, Yuhui and He, Haodi and Wu, Xiaopei and Yu, Haojun and Lin, Weihong and Sun, Lei and Zhang, Chao and Hu, Han},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2207.13080v3},
  eprint = {2207.13080}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE