Incremental-DETR: Incremental Few-Shot Object Detection via Self-Supervised Learning
Na DongYongqiang ZhangMingli DingGim Hee Lee
Presents a DETR-based framework for incremental few-shot object detection that prevents catastrophic forgetting and overfitting by combining self-supervised fine-tuning on pseudo-labeled proposals with selective knowledge distillation on class-specific layers.
Modern computer vision systems rely heavily on deep neural networks to detect objects in images. However, when these systems must learn new object categories from only a few labeled examples without retaining access to original training data, they face two major failure modes: quickly forgetting previously learned categories (catastrophic forgetting) and failing to generalize due to scarce new data (overfitting). This limitation restricts their deployment in dynamic, real-world settings where re-annotating massive datasets or retraining large systems from scratch is cost-prohibitive.
The article introduces and evaluates Incremental-DETR, a framework designed to enable the Detection Transformer architecture to continually learn novel object classes from limited samples while retaining performance on previously learned base classes. The approach demonstrates that transformer-based object detectors can successfully perform incremental few-shot detection without requiring access to original base training datasets.
To achieve this, the article establishes a two-stage training strategy based on empirical findings that separate the transformer detector into class-agnostic components (the feature extractor, transformer body, and bounding box regression head) and class-specific components (the feature projection layer and classification head). In the first stage, the full network trains on abundant base data, followed by a self-supervised tuning step that uses automated, unsupervised region proposals to expose the class-specific components to general object characteristics. In the second stage, all class-agnostic modules are frozen, and only the class-specific layers are updated using limited new examples. To prevent forgetting, knowledge distillation transfers classification and feature representations from the earlier base model while masking out regions belonging to the new classes. The framework was evaluated across standard computer vision benchmarks (MS COCO and PASCAL VOC) under both incremental and few-shot conditions using 1, 5, and 10 examples per class.
The evaluation produced four key findings. First, in standard incremental learning involving the addition of 40 new classes, the framework achieved an average precision of 37.3%, significantly outperforming previous benchmark baselines (which scored 21.3% and 23.8%) while maintaining performance on original base classes within one percentage point of dedicated base models. Second, under few-shot conditions on the same dataset, the approach consistently outperformed prior state-of-the-art methods, attaining an overall average precision of 24.9% and 24.1% for 5-shot and 10-shot settings, compared to approximately 19–20% for competing techniques. Third, in cross-dataset evaluations testing domain adaptability from COCO to VOC, the framework achieved 16.6% and 24.6% average precision in 5-shot and 10-shot tasks, whereas the prior leading baseline remained below 3%. Fourth, ablation experiments confirmed that freezing class-agnostic components is essential; unfreezing these layers during few-shot tuning caused total accuracy to drop sharply from 24.9% down to 4.1% due to catastrophic forgetting.
These findings demonstrate that organizations can reliably adapt advanced transformer-based vision models to new objects with minimal data annotation costs and without incurring the expense of storing historical training sets or executing full-scale model retraining. By decoupling model components and constraining updates to lightweight projection and classification layers, organizations can mitigate operational risks, shorten deployment cycles, and maintain reliable baseline performance across evolving visual environments.
Organizations seeking to deploy scalable visual detection should adopt modular, two-stage fine-tuning architectures and leverage automated proposal generation for self-supervision rather than collecting extensive manual annotations for every new category. When evaluating trade-offs, decision-makers should note that while performance improves dramatically with 5 to 10 training examples, single-example (1-shot) detection remains fundamentally constrained across all evaluated methods. Practitioners should ensure at least 5 to 10 annotated instances per target class in initial pilot implementations before operational rollout.
Confidence in these findings is supported by rigorous ablation studies and consistent performance gains across established academic benchmarks. However, stakeholders should recognize that the evaluations focused on controlled image sets using single-scale feature backbones. Further validation in unstructured, complex production environments is recommended to confirm real-world robustness under varied environmental conditions.
- Paper: End-to-End Object Detection with Transformers, Nicolas Carion et al. (2020). Introduces the base DETR transformer-based set prediction architecture whose class-specific components and bipartite matching scheme are directly adapted by Incremental-DETR.
- Paper: OW-DETR: Open-world Detection Transformer, Akshita Gupta et al. (2022). Establishes how to formulate incremental and open-world object detection within transformer detectors using pseudo-labeling and class separation.
- Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). Presents Deformable DETR, an essential evolution of the DETR framework that resolves training convergence and scale issues in transformer-based detection.
- Paper: Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment, Guangxing Han et al. (2022). Provides fundamental background on mitigating catastrophic forgetting and proposal misalignment in few-shot object detection settings.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). Establishes foundational knowledge distillation and exemplar strategies for class-incremental representation learning.
- Paper: Few-Shot Object Detection with Foundation Models, Guangxing Han et al. (2024). Extends few-shot object detection with transformer proposal architectures by integrating frozen self-supervised vision backbones and in-context language models.
