Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment

Guangxing HanShiyuan HuangJiawei MaYicheng HeShih-Fu Chang

article2022AAAI243 citations

Proposes a meta-learning framework that enhances few-shot object detection by pairing a prototype-matching proposal generator with attentive spatial feature alignment to resolve misaligned bounding boxes and poor proposal recall for rare classes.

Listen

Standard deep learning models for visual object detection rely heavily on thousands of manually labeled examples. Collecting these annotations is expensive, time-consuming, and often impossible for rare or emerging object categories. When presented with only a handful of training examples—known as few-shot object detection—conventional models frequently overfit and fail. A central driver of this failure is that early-stage object proposal generators miss novel items, and the subsequent classification stages struggle because coarse candidate boxes misalign spatially with reference examples.

The article demonstrates a novel framework called Meta Faster R-CNN that significantly improves few-shot detection accuracy without degrading performance on pre-existing base categories. The objective is to establish an adaptable detection architecture that learns to match query images against reference examples using a two-stage, coarse-to-fine matching pipeline.

The authors developed an architecture that decouples detection into two parallel paths sharing a single visual backbone: a standard path for well-represented base categories and a dedicated few-shot path for novel categories. The few-shot pipeline introduces two core enhancements. First, a lightweight prototype matching network generates category-specific candidate regions with high recall. Second, a fine-grained classifier applies attentive spatial alignment to map semantic correspondences between candidate regions and reference examples while masking out distracting background noise. The authors evaluated the system across standard benchmark datasets, including MS COCO and PASCAL VOC, under varying data constraints ranging from one to thirty training examples per category.

The evaluation produced four primary findings. First, the proposed candidate proposal generator substantially increased target recall, achieving an average recall of 33.8% at 100 proposals compared to 22.9% for standard region proposal baselines. Second, the spatial alignment and foreground attention mechanisms increased final detection precision across benchmarks, outperforming competing fine-tuning techniques on the 10-shot MS COCO benchmark with a 12.7 average precision compared to 9.6–10.7 in prior methods. Third, the meta-learning architecture proved exceptionally effective in extreme low-data regimes; without any post-deployment fine-tuning, the 1-shot model achieved competitive or superior precision compared to fine-tuned alternatives. Fourth, separating the base and novel detection pathways preserved high performance on base categories (achieving a 36.9 average precision) while maintaining fast inference speeds of approximately 0.21 seconds per image.

These results demonstrate that organizations can deploy computer vision systems that rapidly enroll new target items without expensive re-annotation or extensive model re-training. By addressing spatial misalignment and proposal quality directly, the framework reduces operational computing costs and avoids the performance drops common when adding new classes to legacy detectors. In practical deployment, systems can rely entirely on pure meta-inference for instant 1-shot onboarding, or apply targeted fine-tuning when ten or more examples are available to maximize overall accuracy.

Decision-makers adopting this framework should maintain a two-branch architecture to preserve base category accuracy and execute full fine-tuning only when sufficient examples exist. In high-throughput settings requiring rapid class enrollment, deploying the meta-testing pipeline without fine-tuning provides the most efficient operational workflow. Future technical exploration should focus on incorporating broader image context and external semantic knowledge to improve performance on small objects and fine-grained visual distinctions.

Confidence in these findings is supported by consistent performance gains across multiple standard vision benchmarks and controlled ablation studies. However, practitioners should exercise caution in settings with heavy background clutter or extreme scale variations, as the authors note persistent limitations in detecting very small objects and distinguishing visually similar categories in complex environments.

  • Paper: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Shaoqing Ren et al. (2015). It establishes the foundational two-stage detection and Region Proposal Network framework that Meta Faster R-CNN directly adapts for few-shot learning.
  • Paper: Prototypical Networks for Few-shot Learning, Jake Snell et al. (2017). It introduces the metric-learning prototype matching principle used by the source to replace standard linear classifiers in the region proposal network.
  • Paper: Fast R-CNN, Ross B. Girshick (2015). It formalizes the RoI-based feature pooling and multi-task loss formulation that underlies the region classification head refined in this work.
  • Paper: Feature Pyramid Networks for Object Detection, Tsung-Yi Lin et al. (2017). It provides the multi-scale feature pyramid representations commonly integrated with Faster R-CNN architectures to generate proposals across varying object scales.
Cover for Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment

Abstract

Few-shot object detection (FSOD) aims to detect objects using only a few examples. How to adapt state-of-the-art object detectors to the few-shot domain remains challenging. Object proposal is a key ingredient in modern object detectors. However, the quality of proposals generated for few-shot classes using existing methods is far worse than that of many-shot classes, e.g., missing boxes for few-shot classes due to misclassification or inaccurate spatial locations with respect to true objects. To address the noisy proposal problem, we propose a novel meta-learning based FSOD model by jointly optimizing the few-shot proposal generation and fine-grained few-shot proposal classification. To improve proposal generation for few-shot classes, we propose to learn a lightweight metric-learning based prototype matching network, instead of the conventional simple linear object/nonobject classifier, e.g., used in RPN. Our non-linear classifier with the feature fusion network could improve the discriminative prototype matching and the proposal recall for few-shot classes. To improve the fine-grained few-shot proposal classification, we propose a novel attentive feature alignment method to address the spatial misalignment between the noisy proposals and few-shot classes, thus improving the performance of few-shot object detection. Meanwhile we learn a separate Faster R-CNN detection head for many-shot base classes and show strong performance of maintaining base-classes knowledge. Our model achieves state-of-the-art performance on multiple FSOD benchmarks over most of the shots and metrics.

Table of Contents

  • Introduction
  • Related Work
  • Our Approach
  • Task Definition
  • The Model Architecture
  • The Training Framework
  • Experimental Results
  • Datasets
  • Ablation Study
  • Comparison with State-of-the-arts
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Two-branch coarse-to-fine architecture for few-shot detection

    model/method

    Meta Faster R-CNN addresses few-shot object detection with disjoint base classes Cbase\mathcal{C}_{\mathrm{base}} and novel classes Cnovel\mathcal{C}_{\mathrm{novel}}. The model uses a shared feature backbone but decouples detection into two branches. The base-class branch is a conventional Faster R-CNN: a category-agnostic RPN generates proposals, and an R-CNN head predicts one of the base classes or background together with bounding-box regression. The novel-class branch first uses a category-specific Meta-RPN to generate proposals conditioned on few-shot class prototypes, then uses a Meta-Classifier for fine-grained proposal classification and refinement.

    For each novel class, the detector receives KK annotated support examples, typically with K∈{1,5,10}K\in\{1,5,10\}, and a query image. Support objects are cropped with surrounding context, resized, and processed by the same backbone as the query image. After meta-training, a novel class can be enrolled by computing its prototype from the support examples, without further training during meta-testing. The separate base-class head preserves the efficiency and accuracy of conventional Faster R-CNN while the novel-class branch supports incremental class addition.

  2. Knowl 2 — Meta-RPN with nonlinear prototype matching

    model/method

    The Meta-RPN generates category-specific proposals for each novel class instead of applying the base-class RPN's linear object/nonobject classifier. Let the query feature map be fq=F(Iq)∈RH×W×Cf_q=F(I_q)\in\mathbb{R}^{H\times W\times C}, where HH, WW, and CC are its spatial height, spatial width, and channel dimension. For novel class cc, the support feature-map prototype is the average of the KK support features, and its spatially pooled prototype is

    fpoolc=1HW∑h=1H∑w=1Wfc(h,w)∈RC.f_{\mathrm{pool}}^c=\frac{1}{HW}\sum_{h=1}^{H}\sum_{w=1}^{W}f^c(h,w)\in\mathbb{R}^{C}.

    The pooled prototype is broadcast over the query feature map. At every anchor location, the Meta-RPN forms a fused representation

    fqc=[ΦMult(fq⊙fpoolc),  ΦSub(fq−fpoolc),  ΦCat[fq,fpoolc]],f_q^c=\left[\Phi_{\mathrm{Mult}}(f_q\odot f_{\mathrm{pool}}^c),\;\Phi_{\mathrm{Sub}}(f_q-f_{\mathrm{pool}}^c),\;\Phi_{\mathrm{Cat}}[f_q,f_{\mathrm{pool}}^c]\right],

    where ⊙\odot is element-wise multiplication, [⋅][\cdot] denotes channel concatenation, and each Φ\Phi is a convolution followed by ReLU. Multiplication emphasizes shared features, subtraction measures feature differences, and concatenation provides a learnable joint representation. A binary classification layer and a bounding-box regression layer then predict proposals for class cc. The implementation first applies a 3×33\times3 convolution and ReLU to query features, and the default inference setting retains 100 proposals per novel class.

  3. Knowl 3 — Proposal-to-prototype spatial alignment

    model/method

    The Meta-Classifier performs fine-grained matching between a noisy proposal and a novel-class prototype using high-resolution feature maps. A proposal feature map fpf_p and a support prototype feature map f^c\hat f^c both have spatial size H′×W′=7×7H'\times W'=7\times7 and channel dimension C′=2048C'=2048. For proposal grid location ii and prototype grid location jj, their affinity is

    A(i,j)=fp(i)Tf^c(j),A(i,j)=f_p(i)^{\mathsf T}\hat f^c(j),

    where A∈RH′W′×H′W′A\in\mathbb{R}^{H'W'\times H'W'} and each feature vector lies in RC′\mathbb{R}^{C'}. The affinity is normalized over all prototype locations for every proposal location:

    A′(i,j)=exp⁡(A(i,j))∑kexp⁡(A(i,k)).A'(i,j)=\frac{\exp(A(i,j))}{\sum_k\exp(A(i,k))}.

    The prototype is then spatially aligned to the proposal by aggregating prototype features according to these correspondence weights:

    fˉc(i)=∑jA′(i,j)f^c(j).\bar f^c(i)=\sum_j A'(i,j)\hat f^c(j).

    Thus, semantic regions in the support prototype are compared with the proposal at corresponding, rather than fixed, spatial positions. The alignment direction is deliberately asymmetric: transforming the prototype into the proposal's spatial organization preserves the proposal's original structure for subsequent bounding-box regression.

  4. Knowl 4 — Foreground attention and nonlinear fine-grained classification

    model/method

    Because a proposal can include substantial background, the Meta-Classifier derives a foreground attention mask from the same affinity matrix used for alignment. For proposal location ii, the mask is

    M(i)=σ(∑jA(i,j))=11+exp⁡(−∑jA(i,j)),M(i)=\sigma\left(\sum_jA(i,j)\right)=\frac{1}{1+\exp\left(-\sum_jA(i,j)\right)},

    where σ\sigma is the sigmoid function and M∈RH′×W′M\in\mathbb{R}^{H'\times W'}. Locations with strong correspondence to some prototype region receive larger mask values. The mask is applied to both the proposal and aligned prototype features:

    fattc=M⊙fˉc,fattp=M⊙fp.f^c_{\mathrm{att}}=M\odot\bar f^c,\qquad f^p_{\mathrm{att}}=M\odot f_p.

    To stabilize training, learnable scalars γ1\gamma_1 and γ2\gamma_2, both initialized to zero, gate the attended features through residual connections:

    f~c=γ1fattc+f^c,f~p=γ2fattp+fp.\tilde f^c=\gamma_1f^c_{\mathrm{att}}+\hat f^c,\qquad \tilde f^p=\gamma_2f^p_{\mathrm{att}}+f_p.

    The final similarity representation uses three nonlinear fusion subnetworks:

    f=[ΨMult(f~c⊙f~p),  ΨSub(f~c−f~p),  ΨCat[f~c,f~p]].f=\left[\Psi_{\mathrm{Mult}}(\tilde f^c\odot\tilde f^p),\;\Psi_{\mathrm{Sub}}(\tilde f^c-\tilde f^p),\;\Psi_{\mathrm{Cat}}[\tilde f^c,\tilde f^p]\right].

    Each Ψ\Psi contains three convolutional layers with ReLU activations. A binary classifier and bounding-box regressor operate on ff to produce the final novel-class detection and refinement.

  5. Knowl 5 — Three-stage meta-training, base training, and fine-tuning procedure

    model/method

    Meta Faster R-CNN is trained in three stages. First, meta-training uses only data from base classes. Each episode samples several base classes, kk support examples per sampled class, and query images with ground-truth boxes, thereby simulating novel-class detection. Binary cross-entropy and smooth-L1L_1 losses train the Meta-RPN and Meta-Classifier, while positive and negative matching pairs are sampled at a 1:31{:}3 ratio to prevent the many negative pairs from dominating the loss. At meta-testing, no novel-class optimization is performed; the support images are only used to compute class prototypes.

    Second, after meta-training, the shared backbone is fixed and a separate base-class RPN and R-CNN head are trained using the abundant base-class annotations. Third, fine-tuning uses a small balanced set of original images containing both base and novel classes. Exactly kk novel-class instances per class are included, and these images are used as query images to update both the Meta-RPN and Meta-Classifier. Fine-tuning generally helps when more novel examples are available, but with extremely few examples such as two shots it can overfit and provide little improvement.

  6. Knowl 6 — Separate base-class head preserves detection accuracy and speed

    empirical result

    On MSCOCO with a ResNet-101 backbone, the separate conventional Faster R-CNN head substantially outperforms using the few-shot Meta-Classifier for base-class detection and retains the speed of the base detector. Accuracy is reported in AP points and speed in seconds per image.

    Could not parse LaTeX table

    The separate-head model is more accurate than the TFA comparison while running at the same 0.21 seconds per image. Applying the class-conditioned few-shot detector directly to base classes is both slower because classes are detected independently and less accurate because averaged few-shot prototypes are inferior to a learned softmax classifier when abundant base data are available.

  7. Knowl 7 — Meta-RPN substantially improves novel-class proposal recall

    empirical result

    On MSCOCO under 10-shot evaluation, the Meta-RPN produces higher proposal average recall (AR) than a base-class RPN and Attention-RPN at every proposal budget. The number of proposals is counted per novel class.

    Could not parse LaTeX table

    The results show that conditioning proposals on novel-class prototypes is more effective than reusing a base-class object/nonobject classifier. The nonlinear Meta-RPN remains advantageous after fine-tuning, and its gain is especially relevant when only a small number of proposals can be retained.

  8. Knowl 8 — MSCOCO few-shot detection results

    data/table

    On MSCOCO, using ResNet-101 and the same few-shot instances as prior methods, Meta Faster R-CNN improves over the attention-based meta-learning baseline in every reported meta-testing setting. The table reports AP, AP50, and AP75 for novel classes; values are AP points.

    Could not parse LaTeX table

    Meta-training alone is particularly effective in the 1- and 2-shot regimes, where fine-tuning has little data and is vulnerable to overfitting. Fine-tuning produces larger gains at 10 and 30 shots, reaching 16.6 AP and 31.8 AP50 at 30 shots.

  9. Knowl 9 — PASCAL VOC few-shot detection results

    data/table

    On PASCAL VOC, Meta Faster R-CNN uses ResNet-101 and is evaluated on three standard novel-class splits. The metric is AP50 for novel classes, and each split is tested at 1, 2, 3, 5, and 10 shots. The meta-testing model requires no novel-class training after meta-training, whereas the fine-tuned model uses the corresponding novel-class images as query training data.

    Could not parse LaTeX table

    Meta Faster R-CNN improves over the meta-testing baseline across all listed split-and-shot settings. Fine-tuning yields especially large improvements on Novel Set 1, while the meta-testing model remains competitive in the extremely low-shot settings.

  10. Knowl 10 — Complementarity of multiplication, subtraction, and concatenation fusion

    empirical result

    Ablations on MSCOCO with 10-shot meta-testing show that the three fusion subnetworks in Meta-RPN and Meta-Classifier are complementary. The entries report AP, AP50, and AP75. Each row activates the indicated element-wise or concatenation fusion operation.

    Could not parse LaTeX table

    Using concatenation alone is weaker, particularly for the Meta-Classifier, because it must learn a more complex fusion function. Combining it with multiplication and subtraction produces the best result for both proposal generation and proposal classification.

Coverage note — The paper's qualitative failure-case visualization—missed small persons and boat misclassification—and its associated suggestions for multiscale detection and image-context modeling were omitted because they are lower-significance diagnostic analyses rather than load-bearing components of the proposed method.

References

  1. 1.Everingham, M.; Van Gool, L.; Williams, C. K.; Winn, J.; and Zisserman, A. 2010. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2): 303–338.
  2. 2.Fan, Q.; Zhuo, W.; Tang, C.-K.; and Tai, Y.-W. 2020. Few-shot object detection with attention-RPN and multi-relation detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4013–4022.
  3. 3.Finn, C.; Abbeel, P.; and Levine, S. 2017. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. In International Conference on Machine Learning, 1126–1135.
  4. 4.Gidaris, S.; and Komodakis, N. 2018. Dynamic few-shot visual learning without forgetting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 4367–4375.
  5. 5.Girshick, R. 2015. Fast R-CNN. In Proceedings of the IEEE international conference on computer vision, 1440–1448.
  6. 6.Han, G.; He, Y.; Huang, S.; Ma, J.; and Chang, S.-F. 2021. Query Adaptive Few-Shot Object Detection With Heterogeneous Graph Convolutional Networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 3263–3272.
  7. 7.Han, G.; Ma, J.; Huang, S.; Chen, L.; and Chang, S.-F. 2022a. Few-Shot Object Detection with Fully Cross-Transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  8. 8.Han, G.; Ma, J.; Huang, S.; Chen, L.; Chellappa, R.; and Chang, S.-F. 2022b. Multimodal Few-Shot Object Detection with Meta-Learning Based Cross-Modal Prompting.
  9. 9.Han, G.; Zhang, X.; and Li, C. 2017a. Revisiting Faster R-CNN: A Deeper Look at Region Proposal Network. In International Conference on Neural Information Processing, 14–24.
  10. 10.Han, G.; Zhang, X.; and Li, C. 2017b. Single shot object detection with top-down refinement. In 2017 IEEE International Conference on Image Processing (ICIP), 3360–3364.
  11. 11.Han, G.; Zhang, X.; and Li, C. 2018. Semi-supervised DFF: Decoupling detection and feature flow for video object detectors. In Proceedings of the 26th ACM international conference on Multimedia, 1811–1819.
  12. 12.He, K.; Gkioxari, G.; Dollar, P.; and Girshick, R. 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, 2961–2969.
  13. 13.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 770–778.
  14. 14.Hosang, J.; Benenson, R.; Dollar, P.; and Schiele, B. 2015. What makes for effective detection proposals? IEEE transactions on pattern analysis and machine intelligence, 38(4): 814–830.
  15. 15.Hsieh, T.-I.; Lo, Y.-C.; Chen, H.-T.; and Liu, T.-L. 2019. One-shot object detection with co-attention and co-excitation. In Advances in Neural Information Processing Systems, 2725–2734.
  16. 16.Huang, S.; Ma, J.; Han, G.; and Chang, S.-F. 2022. Task-Adaptive Negative Class Envision for Few-Shot Open-Set Recognition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  17. 17.Kang, B.; Liu, Z.; Wang, X.; Yu, F.; Feng, J.; and Darrell, T. 2019. Few-shot object detection via feature reweighting. In Proceedings of the IEEE International Conference on Computer Vision, 8420–8429.
  18. 18.Karlinsky, L.; Shtok, J.; Harary, S.; Schwartz, E.; Aides, A.; Feris, R.; Giryes, R.; and Bronstein, A. M. 2019. Repmet: Representative-based metric learning for classification and few-shot object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5197–5206.
  19. 19.Koch, G.; Zemel, R.; and Salakhutdinov, R. 2015. Siamese neural networks for one-shot image recognition. In ICML deep learning workshop, volume 2. Lille.
  20. 20.Krizhevsky, A.; Sutskever, I.; and Hinton, G. E. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In Advances in neural information processing systems, 1097–1105.
  21. 21.Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollar, P.; and Zitnick, C. L. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, 740–755. Springer.
  22. 22.Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; and Berg, A. C. 2016. Ssd: Single shot multibox detector. In European conference on computer vision, 21–37. Springer.
  23. 23.Ma, J.; Xie, H.; Han, G.; Chang, S.-F.; Galstyan, A.; and Abd-Almageed, W. 2021. Partner-Assisted Learning for Few-Shot Image Classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10573–10582.
  24. 24.Osokin, A.; Sumin, D.; and Lomakin, V. 2020. OS2D: One-Stage One-Shot Object Detection by Matching Anchor Features. In European Conference on Computer Vision.
  25. 25.Perez-Rua, J.-M.; Zhu, X.; Hospedales, T. M.; and Xiang, T. 2020. Incremental Few-Shot Object Detection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13846–13855.
  26. 26.Redmon, J.; Divvala, S.; Girshick, R.; and Farhadi, A. 2016. You only look once: Unified, real-time object detection. In IEEE conference on computer vision and pattern recognition, 779–788.
  27. 27.Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems, 91–99.
  28. 28.Snell, J.; Swersky, K.; and Zemel, R. 2017. Prototypical networks for few-shot learning. In Advances in neural information processing systems, 4077–4087.
  29. 29.Sun, B.; Li, B.; Cai, S.; Yuan, Y.; and Zhang, C. 2021. FSCE: Few-Shot Object Detection via Contrastive Proposal Encoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7352–7362.
  30. 30.Sung, F.; Yang, Y.; Zhang, L.; Xiang, T.; Torr, P. H.; and Hospedales, T. M. 2018. Learning to compare: Relation network for few-shot learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1199–1208.
  31. 31.Tian, Z.; Shen, C.; Chen, H.; and He, T. 2019. Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE international conference on computer vision, 9627–9636.
  32. 32.Vinyals, O.; Blundell, C.; Lillicrap, T.; Wierstra, D.; et al. 2016. Matching networks for one shot learning. In Advances in neural information processing systems, 3630–3638.
  33. 33.Wang, X.; Huang, T. E.; Darrell, T.; Gonzalez, J. E.; and Yu, F. 2020. Frustratingly Simple Few-Shot Object Detection. In International Conference on Machine Learning (ICML).
  34. 34.Wang, Y.-X.; Ramanan, D.; and Hebert, M. 2019. Meta-learning to detect rare objects. In Proceedings of the IEEE International Conference on Computer Vision, 9925–9934.
  35. 35.Wu, J.; Liu, S.; Huang, D.; and Wang, Y. 2020. Multi-Scale Positive Sample Refinement for Few-Shot Object Detection. In European Conference on Computer Vision, 456–472. Springer.
  36. 36.Wu, X.; Sahoo, D.; and Hoi, S. 2020. Meta-rcnn: Meta learning for few-shot object detection. In Proceedings of the 28th ACM International Conference on Multimedia, 1679–1687.
  37. 37.Xiao, Y.; and Marlet, R. 2020. Few-Shot Object Detection and Viewpoint Estimation for Objects in the Wild. In European Conference on Computer Vision.
  38. 38.Yan, X.; Chen, Z.; Xu, A.; Wang, X.; Liang, X.; and Lin, L. 2019. Meta r-cnn: Towards general solver for instance-level low-shot learning. In Proceedings of the IEEE International Conference on Computer Vision, 9577–9586.
  39. 39.Ypsilantis, N.-A.; Garcia, N.; Han, G.; Ibrahimi, S.; Van Noord, N.; and Tolias, G. 2021. The Met Dataset: Instance-level Recognition for Artworks. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  40. 40.Zhang, W.; and Wang, Y.-X. 2021. Hallucination Improves Few-Shot Object Detection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13008–13017.
  41. 41.Zhu, C.; Chen, F.; Ahmed, U.; Shen, Z.; and Savvides, M. 2021. Semantic Relation Reasoning for Shot-Stable Few-Shot Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8782–8791.

Citation

MLA
Han, G., et al. “Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 1, 2022, pp. 780–89, https://doi.org/10.1609/AAAI.V36I1.19959.
APA
Han, G., Huang, S., Ma, J., He, Y., & Chang, S.-F. (2022). Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment. Proceedings of the AAAI Conference on Artificial Intelligence, 36(1), 780–789. https://doi.org/10.1609/AAAI.V36I1.19959
Chicago
Han, G., S. Huang, J. Ma, Y. He, and S.-F. Chang. 2022. “Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (1): 780–89. https://doi.org/10.1609/AAAI.V36I1.19959.
Harvard
Han, G. et al. (2022) “Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(1), pp. 780–789. Available at: https://doi.org/10.1609/AAAI.V36I1.19959.
Vancouver
1. Han G, Huang S, Ma J, He Y, Chang S-F (2022) Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment. Proceedings of the AAAI Conference on Artificial Intelligence 36:780–789

BibTeX

@article{Han_2022, title={Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I1.19959}, DOI={10.1609/aaai.v36i1.19959}, number={1}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Han, Guangxing and Huang, Shiyuan and Ma, Jiawei and He, Yicheng and Chang, Shih-Fu}, year={2022}, month=June, pages={780–789} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF