Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning

Na DongYongqiang ZhangMingli DingGim Hee Lee

article2023AAAI52 citations

Presents a DETR-based framework for incremental few-shot object detection that prevents catastrophic forgetting and overfitting by combining self-supervised fine-tuning on pseudo-labeled proposals with selective knowledge distillation on class-specific layers.

Listen

Modern computer vision systems rely heavily on deep neural networks to detect objects in images. However, when these systems must learn new object categories from only a few labeled examples without retaining access to original training data, they face two major failure modes: quickly forgetting previously learned categories (catastrophic forgetting) and failing to generalize due to scarce new data (overfitting). This limitation restricts their deployment in dynamic, real-world settings where re-annotating massive datasets or retraining large systems from scratch is cost-prohibitive.

The article introduces and evaluates Incremental-DETR, a framework designed to enable the Detection Transformer architecture to continually learn novel object classes from limited samples while retaining performance on previously learned base classes. The approach demonstrates that transformer-based object detectors can successfully perform incremental few-shot detection without requiring access to original base training datasets.

To achieve this, the article establishes a two-stage training strategy based on empirical findings that separate the transformer detector into class-agnostic components (the feature extractor, transformer body, and bounding box regression head) and class-specific components (the feature projection layer and classification head). In the first stage, the full network trains on abundant base data, followed by a self-supervised tuning step that uses automated, unsupervised region proposals to expose the class-specific components to general object characteristics. In the second stage, all class-agnostic modules are frozen, and only the class-specific layers are updated using limited new examples. To prevent forgetting, knowledge distillation transfers classification and feature representations from the earlier base model while masking out regions belonging to the new classes. The framework was evaluated across standard computer vision benchmarks (MS COCO and PASCAL VOC) under both incremental and few-shot conditions using 1, 5, and 10 examples per class.

The evaluation produced four key findings. First, in standard incremental learning involving the addition of 40 new classes, the framework achieved an average precision of 37.3%, significantly outperforming previous benchmark baselines (which scored 21.3% and 23.8%) while maintaining performance on original base classes within one percentage point of dedicated base models. Second, under few-shot conditions on the same dataset, the approach consistently outperformed prior state-of-the-art methods, attaining an overall average precision of 24.9% and 24.1% for 5-shot and 10-shot settings, compared to approximately 19–20% for competing techniques. Third, in cross-dataset evaluations testing domain adaptability from COCO to VOC, the framework achieved 16.6% and 24.6% average precision in 5-shot and 10-shot tasks, whereas the prior leading baseline remained below 3%. Fourth, ablation experiments confirmed that freezing class-agnostic components is essential; unfreezing these layers during few-shot tuning caused total accuracy to drop sharply from 24.9% down to 4.1% due to catastrophic forgetting.

These findings demonstrate that organizations can reliably adapt advanced transformer-based vision models to new objects with minimal data annotation costs and without incurring the expense of storing historical training sets or executing full-scale model retraining. By decoupling model components and constraining updates to lightweight projection and classification layers, organizations can mitigate operational risks, shorten deployment cycles, and maintain reliable baseline performance across evolving visual environments.

Organizations seeking to deploy scalable visual detection should adopt modular, two-stage fine-tuning architectures and leverage automated proposal generation for self-supervision rather than collecting extensive manual annotations for every new category. When evaluating trade-offs, decision-makers should note that while performance improves dramatically with 5 to 10 training examples, single-example (1-shot) detection remains fundamentally constrained across all evaluated methods. Practitioners should ensure at least 5 to 10 annotated instances per target class in initial pilot implementations before operational rollout.

Confidence in these findings is supported by rigorous ablation studies and consistent performance gains across established academic benchmarks. However, stakeholders should recognize that the evaluations focused on controlled image sets using single-scale feature backbones. Further validation in unstructured, complex production environments is recommended to confirm real-world robustness under varied environmental conditions.

arXiv: 2205.04042
  • Paper: Few-Shot Object Detection with Foundation Models, Guangxing Han et al. (2024). Extends few-shot object detection with transformer proposal architectures by integrating frozen self-supervised vision backbones and in-context language models.
Cover for Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning

Abstract

Incremental few-shot object detection aims at detecting novel classes without forgetting knowledge of the base classes with only a few labeled training data from the novel classes. Most related prior works are on incremental object detection that rely on the availability of abundant training samples per novel class that substantially limits the scalability to real-world setting where novel data can be scarce. In this paper, we propose the Incremental-DETR that does incremental few-shot object detection via fine-tuning and self-supervised learning on the DETR object detector. To alleviate severe over-fitting with few novel class data, we first fine-tune the class-specific components of DETR with self-supervision from additional object proposals generated using Selective Search as pseudo labels. We further introduce an incremental few-shot fine-tuning strategy with knowledge distillation on the class-specific components of DETR to encourage the network in detecting novel classes without forgetting the base classes. Extensive experiments conducted on standard incremental object detection and incremental few-shot object detection settings show that our approach significantly outperforms state-of-the-art methods by a large margin. Our source code is available at https://github.com/dongnana777/Incremental-DETR.

Table of Contents

  • Introduction
  • Related Works
  • Problem Definition
  • Our Methodology Base Model Training
  • Incremental Few-Shot Fine-Tuning
  • Experiments
  • Experimental Setup
  • Implementation Details
  • Incremental Object Detection
  • Incremental Few-Shot Object Detection
  • Ablation Studies
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Class-Agnostic and Class-Specific Decomposition of DETR

    model/method

    In the DETR (specifically Deformable DETR) architecture for object detection, network components are decoupled into class-agnostic and class-specific modules to enable incremental few-shot learning:

    1. Class-Agnostic Modules: The convolutional neural network (CNN) backbone (e.g., ResNet-50), the encoder-decoder transformer, and the 3-layer Feed-Forward Network (FFN) bounding box regression head. These components model generic visual representations, spatial geometry, and cross-object relations that transfer across categories.
    2. Class-Specific Modules: The linear channel projection layer located between the CNN backbone and transformer encoder, and the linear projection classification head.

    During incremental few-shot adaptation, keeping the class-agnostic components frozen prevents catastrophic forgetting of learned spatial priors, while updating only the class-specific components enables efficient adaptation to novel categories under severe data scarcity.

  2. Knowl 2 — Self-Supervised Base Model Fine-Tuning via Selective Search Proposals

    model/method

    To enhance the generalizability and transferability of the class-specific projection layer before introducing novel classes, the base training stage of Incremental-DETR is split into two consecutive phases:

    1. Base Model Pre-Training: The entire DETR model is trained from scratch on abundant base class annotations Dbase={(x,y)}\mathcal{D}_{\text{base}} = \{(x, y)\} using the standard Hungarian bipartite matching loss.
    2. Self-Supervised Fine-Tuning: The class-agnostic components (CNN backbone, transformer, regression head) are frozen. The class-specific projection layer and classification head are fine-tuned using multi-task learning. In addition to the base class ground-truth annotations, unsupervised class-agnostic object proposals generated by Selective Search on the base images serve as pseudo ground truths. To filter noisy proposals, only the top OO ranked proposals with zero intersection-over-union (IoU) with base ground-truth bounding boxes are retained as pseudo boxes b′b', and assigned a unified pseudo-class label c′=B+1c' = B+1 (where BB is the number of base classes). This self-supervision trains the projection layer to represent unseen potential object regions.
  3. Knowl 3 — Masked Feature Knowledge Distillation

    equation

    During incremental few-shot fine-tuning on novel class data Dnovel\mathcal{D}_{\text{novel}}, a masked feature distillation loss constrains the projection layer output of the novel model to retain the base representation space on background and base-class image regions without suppressing novel feature learning:

    Lfeatkd=12Nnovel∑i=1w∑j=1h∑k=1c(1−maskijnovel)∥fijknovel−fijkbase∥22\mathcal{L}_{\text{feat}}^{\text{kd}} = \frac{1}{2 N^{\text{novel}}} \sum_{i=1}^w \sum_{j=1}^h \sum_{k=1}^c \left(1 - \text{mask}_{ij}^{\text{novel}}\right) \left\| f_{ijk}^{\text{novel}} - f_{ijk}^{\text{base}} \right\|_2^2

    where fbase,fnovel∈Rw×h×cf^{\text{base}}, f^{\text{novel}} \in \mathbb{R}^{w \times h \times c} denote the feature representations output by the projection layers of the frozen base model and the trainable novel model, respectively, with spatial width ww, height hh, and channel depth cc. The binary spatial mask masknovel∈{0,1}w×h\text{mask}^{\text{novel}} \in \{0, 1\}^{w \times h} is set to 11 for spatial locations falling inside any ground-truth bounding box of novel class objects and 00 otherwise. The normalization factor is the number of unmasked spatial positions:

    Nnovel=∑i=1w∑j=1h(1−maskijnovel)N^{\text{novel}} = \sum_{i=1}^w \sum_{j=1}^h \left(1 - \text{mask}_{ij}^{\text{novel}}\right)

  4. Knowl 4 — Bipartite Classification Knowledge Distillation via Pseudo Base Predictions

    equation

    To prevent catastrophic forgetting of base categories when training the classification head on novel samples, base model predictions on novel images are extracted as pseudo ground-truth targets for distillation. For an input novel image, an output query prediction from the frozen base model is treated as a valid pseudo base target if its predicted class probability exceeds 0.50.5 and its predicted bounding box does not overlap with any novel ground-truth bounding box.

    Bipartite Hungarian matching pairs these pseudo base targets with prediction queries from the novel model. The classification distillation loss is defined as the Kullback-Leibler (KL) divergence between the predicted class distributions of matched query pairs:

    Lclskd=Lkl_div(log⁡(qnovel),qbase)\mathcal{L}_{\text{cls}}^{\text{kd}} = \mathcal{L}_{\text{kl\_div}}\left(\log\left(q^{\text{novel}}\right), q^{\text{base}}\right)

    where qbaseq^{\text{base}} and qnovelq^{\text{novel}} denote the predicted category probability distributions over base classes produced by the base model and novel model classification heads, respectively.

  5. Knowl 5 — Incremental-DETR Novel Fine-Tuning Optimization Objective

    model/method

    In the incremental few-shot fine-tuning stage of Incremental-DETR, the novel model parameters are initialized from the self-supervised fine-tuned base model. The class-agnostic CNN backbone, encoder-decoder transformer, and regression FFN remain frozen. Only the projection layer and classification head are updated on the NN-way KK-shot novel training dataset Dnovel\mathcal{D}_{\text{novel}} by minimizing the total novel loss:

    Ltotalnovel=Lhg(y,y^)+λfeatLfeatkd+λclsLclskd\mathcal{L}_{\text{total}}^{\text{novel}} = \mathcal{L}_{\text{hg}}(y, \hat{y}) + \lambda_{\text{feat}} \mathcal{L}_{\text{feat}}^{\text{kd}} + \lambda_{\text{cls}} \mathcal{L}_{\text{cls}}^{\text{kd}}

    where:

    • Lhg(y,y^)\mathcal{L}_{\text{hg}}(y, \hat{y}) is the standard DETR Hungarian bipartite matching loss comprising sigmoid focal classification loss Lcls\mathcal{L}_{\text{cls}} and bounding box regression loss Lbox\mathcal{L}_{\text{box}} (a weighted sum of ℓ1\ell_1 loss and Generalized IoU loss) computed against novel ground-truth annotations yy.
    • Lfeatkd\mathcal{L}_{\text{feat}}^{\text{kd}} is the masked feature distillation loss applied to the projection layer.
    • Lclskd\mathcal{L}_{\text{cls}}^{\text{kd}} is the classification distillation loss applied to bipartite-matched pseudo base predictions.
    • λfeat\lambda_{\text{feat}} and λcls\lambda_{\text{cls}} are balancing hyperparameters set to 0.10.1 and 22, respectively.
  6. Knowl 6 — Incremental Few-Shot Object Detection Problem Formulation

    definition

    Incremental Few-Shot Object Detection (iFSOD) is a continual learning setting where an object detector must learn to detect novel object classes Cnovel={CB+1,…,CB+N}\mathcal{C}_{\text{novel}} = \{C_{B+1}, \dots, C_{B+N}\} from a training set Dnovel\mathcal{D}_{\text{novel}} containing only KK annotated instances per class (NN-way KK-shot), while retaining high detection performance on a disjoint set of base classes Cbase={C1,…,CB}\mathcal{C}_{\text{base}} = \{C_1, \dots, C_B\} (where Cbase∩Cnovel=∅\mathcal{C}_{\text{base}} \cap \mathcal{C}_{\text{novel}} = \emptyset). In contrast to standard few-shot object detection (FSOD), the abundant base training data Dbase\mathcal{D}_{\text{base}} is completely inaccessible during novel class learning, requiring models to resist catastrophic forgetting without replay of historical base samples.

  7. Knowl 7 — Base Model Multi-Task Hungarian Loss Formulation

    equation

    During the self-supervised base fine-tuning stage of Incremental-DETR, the base model is optimized using a joint multi-task loss combining ground-truth base supervision and pseudo-proposal self-supervision:

    Ltotalbase=Lhg(y,y^)+λ′Lhg(y′,y^)\mathcal{L}_{\text{total}}^{\text{base}} = \mathcal{L}_{\text{hg}}(y, \hat{y}) + \lambda' \mathcal{L}_{\text{hg}}(y', \hat{y})

    where λ′\lambda' is a balancing hyperparameter set to 11. The standard Hungarian loss on base ground truth y={(ci,bi)}i=1My = \{(c_i, b_i)\}_{i=1}^M and model predictions y^={(c^i,b^i)}i=1M\hat{y} = \{(\hat{c}_i, \hat{b}_i)\}_{i=1}^M is:

    Lhg(y,y^)=∑i=1M[Lcls(ci,c^σ^(i))+1{ci≠∅}Lbox(bi,b^σ^(i))]\mathcal{L}_{\text{hg}}(y, \hat{y}) = \sum_{i=1}^M \left[ \mathcal{L}_{\text{cls}}(c_i, \hat{c}_{\hat{\sigma}(i)}) + \mathbf{1}_{\{c_i \neq \emptyset\}} \mathcal{L}_{\text{box}}(b_i, \hat{b}_{\hat{\sigma}(i)}) \right]

    with optimal bipartite permutation σ^=arg⁡min⁡σ∑i=1MLmatch(yi,y^σ(i))\hat{\sigma} = \arg\min_\sigma \sum_{i=1}^M \mathcal{L}_{\text{match}}(y_i, \hat{y}_{\sigma(i)}), where Lmatch(yi,y^σ(i))=1{ci≠∅}Lcls(ci,c^σ(i))+1{ci≠∅}Lbox(bi,b^σ(i))\mathcal{L}_{\text{match}}(y_i, \hat{y}_{\sigma(i)}) = \mathbf{1}_{\{c_i \neq \emptyset\}} \mathcal{L}_{\text{cls}}(c_i, \hat{c}_{\sigma(i)}) + \mathbf{1}_{\{c_i \neq \emptyset\}} \mathcal{L}_{\text{box}}(b_i, \hat{b}_{\sigma(i)}).

    The pseudo-proposal Hungarian loss Lhg(y′,y^)\mathcal{L}_{\text{hg}}(y', \hat{y}) uses the top PP Selective Search proposals y′={(ci′,bi′)}i=1Py' = \{(c'_i, b'_i)\}_{i=1}^P with assigned binary label ci′=B+1c'_i = B+1, matching against predictions via optimal assignment σ^′\hat{\sigma}':

    Lhg(y′,y^)=∑i=1P[Lcls(ci′,c^σ^′(i))+1{ci′≠∅}Lbox(bi′,b^σ^′(i))]\mathcal{L}_{\text{hg}}(y', \hat{y}) = \sum_{i=1}^P \left[ \mathcal{L}_{\text{cls}}(c'_i, \hat{c}_{\hat{\sigma}'(i)}) + \mathbf{1}_{\{c'_i \neq \emptyset\}} \mathcal{L}_{\text{box}}(b'_i, \hat{b}_{\hat{\sigma}'(i)}) \right]

  8. Knowl 8 — Performance on COCO Same-Dataset Incremental Few-Shot Object Detection

    data/table

    Incremental few-shot object detection performance evaluated on the MS COCO 2017 validation set, partitioned into 60 base classes and 20 novel classes (shared with PASCAL VOC), across 1-shot, 5-shot, and 10-shot novel sample regimes. Average Precision (AP%) and AP at 0.50.5 IoU (AP50%) are reported for base, novel, and all combined classes:

    Shot Method Base Novel All
    AP% AP50% AP% AP50% AP% AP50%
    Base iMTFA 38.2 58.0 - - - -
    Deformable DETR 36.9 56.7 - - - -
    1 ONCE 17.9 - 0.7 - 13.6 -
    iMTFA 27.8 40.1 3.2 5.9 21.7 31.6
    Incremental-DETR (Ours) 29.4 47.1 1.9 2.7 22.5 36.0
    5 ONCE 17.9 - 1.0 - 13.7 -
    iMTFA 24.1 33.7 6.1 11.2 19.6 28.1
    Incremental-DETR (Ours) 30.5 48.4 8.3 13.3 24.9 39.6
    10 ONCE 17.9 - 1.2 - 13.7 -
    iMTFA 23.4 32.4 7.0 12.7 19.3 27.5
    Incremental-DETR (Ours) 27.3 44.0 14.4 22.4 24.1 38.6

    Incremental-DETR outperforms both ONCE and iMTFA across all shot settings in overall AP and AP50, achieving 14.4%14.4\% novel AP at 10-shot compared to 7.0%7.0\% for iMTFA and 1.2%1.2\% for ONCE, while maintaining superior retention of base class performance (27.3%27.3\% vs. 23.4%23.4\% AP).

  9. Knowl 9 — Cross-Dataset Incremental Few-Shot Performance from COCO to PASCAL VOC

    data/table

    Cross-dataset incremental few-shot object detection evaluation where the model is pre-trained on the 60 base classes of MS COCO and incrementally adapted on the 20 novel classes of PASCAL VOC 2007 (evaluated on the VOC 2007 test set):

    Shot Method Novel AP% Novel AP50%
    1 ONCE - -
    Incremental-DETR (Ours) 4.1 6.6
    5 ONCE 2.4 -
    Incremental-DETR (Ours) 16.6 26.3
    10 ONCE 2.6 -
    Incremental-DETR (Ours) 24.6 38.4

    Under cross-domain transfer, Incremental-DETR outperforms ONCE by large margins, reaching 16.6%16.6\% AP at 5-shot (compared to 2.4%2.4\%) and 24.6%24.6\% AP at 10-shot (compared to 2.6%2.6\%).

  10. Knowl 10 — Full-Data Incremental Object Detection on MS COCO

    data/table

    Performance of Incremental-DETR in standard (abundant data) incremental object detection on the MS COCO validation set for adding a single novel class (40 base classes + 1 novel class) and adding a group of novel classes (40 base classes + 40 novel classes):

    Setting Method Base Novel All
    AP% AP50% AP% AP50% AP% AP50%
    40+1 1-41 Joint Baseline 45.6 68.6 31.3 55.7 45.3 68.3
    1-40 Base Baseline 44.8 67.8 - - - -
    41 Fine-Tuned 15.6 23.9 30.7 50.3 16.0 24.6
    41 From Scratch - - 16.7 32.4 - -
    Incremental-DETR (Ours) 41.9 64.1 29.2 50.3 41.6 63.8
    40+40 1-80 Joint Baseline 46.8 69.4 36.3 54.7 41.4 61.8
    41-80 Fine-Tuned 0.0 0.0 33.0 49.6 16.5 24.8
    41-80 From Scratch - - 35.0 52.6 - -
    Shmelkov et al. (2017) - - - - 21.3 37.4
    Kj et al. (2021) - - - - 23.8 40.5
    Incremental-DETR (Ours) 44.0 66.1 30.6 47.1 37.3 56.6

    In the 40+1 setting, Incremental-DETR retains base AP at 41.9%41.9\% (close to the joint baseline of 44.8%44.8\%) while achieving 29.2%29.2\% novel AP. In the 40+40 setting, Incremental-DETR achieves 37.3%37.3\% overall AP, outperforming previous incremental methods by 13.513.5 to 16.016.0 percentage points.

  11. Knowl 11 — Ablation Analysis of Incremental-DETR Components and Freezing Strategies

    data/table

    Ablation results on the MS COCO validation set in a 5-shot per novel class setting assessing the individual and combined contributions of the two-stage fine-tuning strategy, self-supervised learning losses (Lbox,Lcls\mathcal{L}_{\text{box}}, \mathcal{L}_{\text{cls}}), masked feature distillation (Lfeatkd\mathcal{L}_{\text{feat}}^{\text{kd}}), classification distillation (Lclskd\mathcal{L}_{\text{cls}}^{\text{kd}}), and the freezing of class-agnostic components:

    Config Two-Stage Self-Supervised Distillation Base Novel All
    Lbox\mathcal{L}_{\text{box}} Lcls\mathcal{L}_{\text{cls}} Lfeatkd\mathcal{L}_{\text{feat}}^{\text{kd}} Lclskd\mathcal{L}_{\text{cls}}^{\text{kd}} AP% AP50% AP% AP50% AP% AP50%
    1 0.1 0.2 1.4 2.5 0.4 0.8
    2 ✓ 19.7 32.5 5.2 8.2 16.1 26.4
    3 ✓ ✓ ✓ 16.3 27.1 8.0 12.9 14.2 23.5
    4 ✓ ✓ ✓ ✓ 23.8 38.0 8.3 13.2 19.9 31.8
    5 ✓ ✓ ✓ 26.1 43.0 8.0 12.7 21.6 35.4
    6 ✓ ✓ ✓ 30.7 49.0 5.1 8.5 24.3 38.9
    7 ✓ ✓ ✓ ✓ 30.3 48.2 7.5 12.2 24.6 39.2
    8 (Full) ✓ ✓ ✓ ✓ ✓ 30.5 48.4 8.3 13.3 24.9 39.6

    Additionally, analyzing structural variations shows:

    • One-step base training (combining pre-training and self-supervision from scratch) degrades overall AP from 24.9%24.9\% to 23.0%23.0\% (base AP drops from 30.5%30.5\% to 28.3%28.3\%, novel AP drops from 8.3%8.3\% to 7.1%7.1\%).
    • Unfreezing class-agnostic components (CNN backbone, transformer, regression head) during novel fine-tuning causes severe performance degradation due to catastrophic forgetting and overfitting, collapsing base AP to 4.3%4.3\%, novel AP to 3.5%3.5\%, and overall AP to 4.1%4.1\%.

Coverage note — None was omitted. All key models, formulations, optimization losses, benchmark results, and ablation studies from the paper are fully covered.

References

  1. 1.Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020. End-to-end object detection with transformers. In European Conference on Computer Vision, 213–229. Springer.
  2. 2.Everingham, M.; Van Gool, L.; Williams, C. K.; Winn, J.; and Zisserman, A. 2010. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2): 303–338.
  3. 3.Ganea, D. A.; Boom, B.; and Poppe, R. 2021. Incremental Few-Shot Instance Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1185–1194.
  4. 4.Girshick, R. 2015. Fast r-cnn. In Proceedings of the IEEE international conference on computer vision, 1440–1448.
  5. 5.Girshick, R.; Donahue, J.; Darrell, T.; and Malik, J. 2014a. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 580–587.
  6. 6.Girshick, R.; Donahue, J.; Darrell, T.; and Malik, J. 2014b. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 580–587.
  7. 7.Hinton, G.; Vinyals, O.; Dean, J.; et al. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2(7).
  8. 8.Kj, J.; Rajasegaran, J.; Khan, S.; Khan, F. S.; and Balasubramanian, V. N. 2021. Incremental Object Detection via Meta-Learning. IEEE Transactions on Pattern Analysis and Machine Intelligence.
  9. 9.Kuhn, H. W. 1955. The Hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2): 83–97.
  10. 10.Lin, T. Y.; Dollar, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017a. Feature Pyramid Networks for Object Detection. In IEEE Conference on Computer Vision and Pattern Recognition.
  11. 11.Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; and Dollár, P. 2017b. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision, 2980–2988.
  12. 12.Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, 740–755. Springer.
  13. 13.Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C. Y.; and Berg, A. C. 2016. SSD: Single Shot MultiBox Detector. In European Conference on Computer Vision.
  14. 14.McCloskey, M.; and Cohen, N. J. 1989. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of learning and motivation, volume 24, 109–165. Elsevier.
  15. 15.Perez-Rua, J.-M.; Zhu, X.; Hospedales, T. M.; and Xiang, T. 2020. Incremental Few-Shot Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13846–13855.
  16. 16.Ratcliff, R. 1990. Connectionist models of recognition memory: constraints imposed by learning and forgetting functions. Psychological review, 97(2): 285.
  17. 17.Redmon, J.; Divvala, S.; Girshick, R.; and Farhadi, A. 2016. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, 779–788.
  18. 18.Ren, S.; He, K.; Girshick, R.; and Jian, S. 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In International Conference on Neural Information Processing Systems.
  19. 19.Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I.; and Savarese, S. 2019. Generalized intersection over union: A metric and a loss for bounding box regression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 658–666.
  20. 20.Shmelkov, K.; Schmid, C.; and Alahari, K. 2017. Incremental learning of object detectors without catastrophic forgetting. In Proceedings of the IEEE International Conference on Computer Vision, 3400–3409.
  21. 21.Sun, B.; Li, B.; Cai, S.; Yuan, Y.; and Zhang, C. 2021. FSCE: Few-shot object detection via contrastive proposal encoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7352–7362.
  22. 22.Uijlings, J. R.; Van De Sande, K. E.; Gevers, T.; and Smeulders, A. W. 2013. Selective search for object recognition. International journal of computer vision, 104(2): 154–171.
  23. 23.Wang, X.; Huang, T. E.; Darrell, T.; Gonzalez, J. E.; and Yu, F. 2020. Frustratingly simple few-shot object detection. arXiv preprint arXiv:2003.06957.
  24. 24.Wang, Y.-X.; Ramanan, D.; and Hebert, M. 2019. Meta-learning to detect rare objects. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 9925–9934.
  25. 25.Wu, J.; Liu, S.; Huang, D.; and Wang, Y. 2020. Multi-scale positive sample refinement for few-shot object detection. In European Conference on Computer Vision, 456–472. Springer.
  26. 26.Wu, X.; Sahoo, D.; and Hoi, S. 2020. Meta-RCNN: Meta learning for few-shot object detection. In Proceedings of the 28th ACM International Conference on Multimedia, 1679–1687.
  27. 27.Zhang, G.; Luo, Z.; Cui, K.; and Lu, S. 2021. Meta-detr: Few-shot object detection via unified image-level meta-learning. arXiv preprint arXiv:2103.11731, 2.
  28. 28.Zhou, X.; Wang, D.; and Krähenbühl, P. 2019. Objects as points. arXiv preprint arXiv:1904.07850.
  29. 29.Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2020. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159.

Citation

MLA
Dong, N., et al. “Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning”. arXiv, 2022, http://arxiv.org/abs/2205.04042v3.
APA
Dong, N., Zhang, Y., Ding, M., & Lee, G. H. (2022). Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning. arXiv. http://arxiv.org/abs/2205.04042v3
Chicago
Dong, N., Y. Zhang, M. Ding, and G. H. Lee. 2022. “Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning”. arXiv. http://arxiv.org/abs/2205.04042v3.
Harvard
Dong, N. et al. (2022) “Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.04042v3.
Vancouver
1. Dong N, Zhang Y, Ding M, Lee GH (2022) Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning. arXiv

BibTeX

@article{dong2022incremental,
  title = {Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning},
  author = {Dong, Na and Zhang, Yongqiang and Ding, Mingli and Lee, Gim Hee},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.04042v3},
  eprint = {2205.04042}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF