Task Discrepancy Maximization for Fine-grained Few-Shot Classification

Su Been LeeWonJun MoonJae-Pil Heo

article2022CVPR83 citations

Proposes Task Discrepancy Maximization, a plug-and-play module that combines support and query attention mechanisms to reweight feature channels for capturing subtle, class-discriminative regions in fine-grained few-shot classification.

Listen

Modern computer vision models achieve remarkable accuracy but traditionally require massive amounts of labeled training data. In specialized, fine-grained tasks—such as distinguishing visually similar bird species, aircraft models, or animal breeds—collecting vast labeled datasets is prohibitively expensive and time-consuming. While few-shot learning aims to recognize novel categories from only a handful of labeled examples, existing feature alignment techniques struggle in fine-grained settings because they treat all visual features broadly rather than isolating the subtle, distinct details necessary to differentiate look-alike classes.

The article introduces and evaluates Task Discrepancy Maximization (TDM), a modular plug-in designed to enhance fine-grained few-shot image classification. The primary objective is to demonstrate that dynamically weighting feature channels based on their class-specific distinctiveness allows vision models to focus on critical discriminative regions and significantly boost classification performance.

To achieve this, the authors designed a lightweight mechanism comprising two complementary components: a Support Attention Module (SAM) and a Query Attention Module (QAM). SAM assesses labeled support images to highlight channels that exhibit distinct properties for a specific class while suppressing shared, generic features. QAM analyzes the unlabeled query image to emphasize object-relevant regions, counteracting potential bias from small labeled sample sizes. Combining these two modules yields adaptive, task-specific channel weights. The framework was evaluated across seven fine-grained benchmark datasets (including CUB-200-2011, Aircraft, meta-iNat, Stanford Cars, Stanford Dogs, and Oxford Pets) integrated into four established metric-based few-shot architectures (ProtoNet, DSN, CTX, and FRN) using both standard 4-layer convolutional and 12-layer residual network backbones.

The experimental findings show that TDM consistently improves classification accuracy across all baseline models and datasets, setting new state-of-the-art benchmarks in nearly every evaluated setting. On benchmark bird classification (CUB), TDM increased baseline 1-shot accuracy by over 7 percentage points on simple backbones and reached up to 84.36% on raw images with standard ResNet architectures. In aircraft identification, TDM boosted baseline models by up to 7 percentage points. Furthermore, on challenging datasets with severe domain shifts between training and test distributions (such as tiered meta-iNat), TDM demonstrated strong generalization and resistance to overfitting, confirming that both the support and query attention submodules provide complementary benefits.

These results indicate that automated, fine-grained visual recognition can be effectively deployed even when training data is extremely scarce, directly reducing the cost and operational risk associated with manual data labeling. Rather than requiring complex end-to-end model redesigns, the findings show that existing metric-based computer vision pipelines can achieve significant performance gains simply by appending lightweight channel-weighting modules.

Organizations seeking to deploy few-shot vision systems in specialized domains should integrate task-adaptive channel weighting into their existing metric-based architectures. Technical teams should also test spatial pooling options during implementation, as average pooling proved more robust to visual noise than maximum pooling. Researchers and developers should next explore extending the core dynamic channel-weighting mechanism to other computer vision domains, such as object detection and fine-grained retrieval.

Confidence in these findings is high given the extensive validation across seven benchmarks, two neural network backbones, and 10,000 evaluation episodes with tight confidence intervals. However, a primary limitation noted in the article is that TDM is specifically tailored to localize subtle, fine-grained details; its performance advantages may be more limited in coarse-grained classification tasks where broad global features are sufficient to distinguish widely different object categories.

  • Paper: TADAM: Task dependent adaptive metric for improved few-shot learning, Boris N. Oreshkin et al. (2018). TADAM establishes the foundational framework for dynamic, task-dependent feature conditioning and metric scaling in few-shot learning that informs task-adaptive representations.
  • Paper: Matching Networks for One Shot Learning, Oriol Vinyals et al. (2016). Matching Networks introduces episodic support-set to query-set attention mechanisms that form the standard baseline architecture and formulation for few-shot classification.
  • Paper: A Closer Look at Few-shot Classification, Wei-Yu Chen et al. (2019). This benchmark study rigorously characterizes few-shot classification architectures and evaluation protocols, providing essential context for fine-grained few-shot benchmarks like CUB.
  • Paper: CBAM: Convolutional Block Attention Module, Sanghyun Woo et al. (2018). CBAM establishes the standard channel-attention and spatial-attention formulations used to refine discriminative intermediate feature maps in convolutional neural networks.
  • Paper: Maximum Classifier Discrepancy for Unsupervised Domain Adaptation, Kuniaki Saito et al. (2017). This work introduces the principle of classifier and task discrepancy maximization to align representations and emphasize decision boundaries.
  • Paper: Part-Based R-CNNs for Fine-Grained Category Detection, Ning Zhang et al. (2014). This foundational paper outlines the central challenge of fine-grained recognition by emphasizing the necessity of localizing subtle, discriminative object parts.
Cover for Task Discrepancy Maximization for Fine-grained Few-Shot Classification

Abstract

Recognizing discriminative details such as eyes and beaks is important for distinguishing fine-grained classes since they have similar overall appearances. In this regard, we introduce Task Discrepancy Maximization (TDM), a simple module for fine-grained few-shot classification. Our objective is to localize the class-wise discriminative regions by highlighting channels encoding distinct information of the class. Specifically, TDM learns task-specific channel weights based on two novel components: Support Attention Module (SAM) and Query Attention Module (QAM). SAM produces a support weight to represent channel-wise discriminative power for each class. Still, since the SAM is basically only based on the labeled support sets, it can be vulnerable to bias toward such support set. Therefore, we propose QAM which complements SAM by yielding a query weight that grants more weight to object-relevant channels for a given query image. By combining these two weights, a class-wise task-specific channel weight is defined. The weights are then applied to produce task-adaptive feature maps more focusing on the discriminative details. Our experiments validate the effectiveness of TDM and its complementary benefits with prior methods in fine-grained few-shot classification.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Few-Shot Classification
  • 2.2. Feature Alignment
  • 3. Method
  • 3.1. Problem Formulation
  • 3.2. Channel-wise Representativeness Scores
  • 3.3. Support Attention Module (SAM)
  • 3.4. Query Attention Module (QAM)
  • 3.5. Task Discrepancy Maximization (TDM)
  • 3.6. Discussion
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Implementation Details
  • 4.3. Comparison to Existing Methods
  • 5. Ablation Study
  • 5.1. Effect of Submodules
  • 5.2. Choice of Pooling Function
  • 5.3. Metric Compatibility
  • 6. Limitation
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — Task Discrepancy Maximization

    model/method

    Task Discrepancy Maximization (TDM) is a plug-in feature-alignment module for fine-grained few-shot classification. In an NN-way KK-shot episode, a feature extractor produces CC-channel feature maps for the support and query images. TDM assigns a channel-weight vector to each episode class by combining two complementary signals: a support-derived weight that emphasizes class-discriminative channels and a query-derived weight that emphasizes channels containing object-relevant information.

    The resulting class-specific weights rescale feature-map channels before metric-based classification. This allows TDM to suppress channels encoding features shared by several fine-grained classes while highlighting details whose relevance depends on the classes present in the current episode. TDM is designed to operate with existing metric-based methods, including ProtoNet, DSN, CTX, and FRN.

  2. Knowl 2 — Channel-wise representativeness scores

    equation

    For an episode, let xi,jSx^S_{i,j} be the jj-th support image from class ii, let xQx^Q be a query image, and let gθg_\theta be the feature extractor. The corresponding feature maps are Fi,jS=gθ(xi,jS)F^S_{i,j}=g_\theta(x^S_{i,j}) and FQ=gθ(xQ)F^Q=g_\theta(x^Q), where every feature map lies in RC×H×W\mathbb{R}^{C\times H\times W}. The class prototype is FiP=1K∑j=1KFi,jSF_i^P=\frac{1}{K}\sum_{j=1}^{K}F^S_{i,j}. If fi,cP∈RH×Wf^P_{i,c}\in\mathbb{R}^{H\times W} is channel cc of prototype FiPF_i^P, its mean spatial feature is MiP=1C∑c=1Cfi,cPM_i^P=\frac{1}{C}\sum_{c=1}^{C}f^P_{i,c}.

    For class ii and channel cc, TDM defines an intra-class representativeness score and an inter-class representativeness score as

    Ri,cintra=1HW∥fi,cP−MiP∥22,Ri,cinter=1HWmin⁡1≤j≤Nj≠i∥fi,cP−MjP∥22.R^{\mathrm{intra}}_{i,c}=\frac{1}{H W}\left\|f^P_{i,c}-M_i^P\right\|_2^2, \qquad R^{\mathrm{inter}}_{i,c}=\frac{1}{H W}\min_{\substack{1\leq j\leq N\\j\neq i}}\left\|f^P_{i,c}-M_j^P\right\|_2^2.

    The intra score measures how closely a channel’s activated spatial pattern matches the salient spatial pattern of its own class; a smaller value indicates stronger within-class representativeness. The inter score measures the channel’s distance from the salient patterns of the other N−1N-1 episode classes; a larger value indicates stronger class distinctiveness. TDM uses both scores to estimate channel importance.

  3. Knowl 3 — Support Attention Module

    model/method

    The Support Attention Module (SAM) converts the channel-wise intra and inter scores of each class into a support weight. For class ii, let Riintra,Riinter∈RCR_i^{\mathrm{intra}},R_i^{\mathrm{inter}}\in\mathbb{R}^{C} collect the scores over all CC channels. Two separate fully connected blocks, bintrab^{\mathrm{intra}} and binterb^{\mathrm{inter}}, produce

    wiintra=bintra(Riintra),wiinter=binter(Riinter).w_i^{\mathrm{intra}}=b^{\mathrm{intra}}(R_i^{\mathrm{intra}}),\qquad w_i^{\mathrm{inter}}=b^{\mathrm{inter}}(R_i^{\mathrm{inter}}).

    The class-specific support weight is

    wiS=αwiintra+(1−α)wiinter,α∈[0,1].w_i^S=\alpha w_i^{\mathrm{intra}}+(1-\alpha)w_i^{\mathrm{inter}},\qquad \alpha\in[0,1].

    The two fully connected blocks have the same architecture: for a batch of BB score vectors of length CC, a fully connected layer maps B×CB\times C to B×2CB\times2C, followed by batch normalization and ReLU; a second fully connected layer maps B×2CB\times2C to B×CB\times C, followed by 1+tanh⁡1+\tanh. In SAM, BB is the number of episode classes. The resulting wiSw_i^S assigns high values to channels that are representative of class ii and distinct from the other classes in the episode.

  4. Knowl 4 — Query Attention Module

    model/method

    The Query Attention Module (QAM) supplies information unavailable from the few labeled support images. For a query feature map FQF^Q, let fcQ∈RH×Wf_c^Q\in\mathbb{R}^{H\times W} denote channel cc, and define the query’s mean spatial feature by MQ=1C∑c=1CfcQM^Q=\frac{1}{C}\sum_{c=1}^{C}f_c^Q. Because the query label is unknown, QAM uses only within-image channel relationships and computes

    RQintra(c)=1HW∥fcQ−MQ∥22,R_Q^{\mathrm{intra}}(c)=\frac{1}{H W}\left\|f_c^Q-M^Q\right\|_2^2,

    for every channel cc. A fully connected block bQb^Q with the same C→2C→CC\rightarrow2C\rightarrow C architecture, batch normalization, ReLU, and final 1+tanh⁡1+\tanh activation produces the query weight

    wQ=bQ(RQintra).w^Q=b^Q(R_Q^{\mathrm{intra}}).

    QAM emphasizes channels whose spatial activations resemble the query’s mean spatial object representation. It is complementary to SAM: object-relevant channels need not be class-discriminative, while class-discriminative channels identified from a small support set may be unreliable for a particular query.

  5. Knowl 5 — Task-adaptive feature transformation and inference

    model/method

    For class ii, TDM combines the support weight wiS∈RCw_i^S\in\mathbb{R}^{C} and query weight wQ∈RCw^Q\in\mathbb{R}^{C} using

    wiT=βwiS+(1−β)wQ,β∈[0,1].w_i^T=\beta w_i^S+(1-\beta)w^Q,\qquad \beta\in[0,1].

    If F∈RC×H×WF\in\mathbb{R}^{C\times H\times W} has channel maps f1,…,fCf_1,\ldots,f_C, the task-adaptive feature map is obtained by channel-wise scaling:

    A=wiT⊙F=[wi,1Tf1,wi,2Tf2,…,wi,CTfC].A=w_i^T\odot F=[w^T_{i,1}f_1,w^T_{i,2}f_2,\ldots,w^T_{i,C}f_C].

    Support images from class ii are transformed with wiTw_i^T. When evaluating a query as class ii, the same wiTw_i^T is applied to that query, producing Ai,jS=wiT⊙Fi,jSA^S_{i,j}=w_i^T\odot F^S_{i,j} and AiQ=wiT⊙FQA_i^Q=w_i^T\odot F^Q. For ProtoNet, the transformed support prototype AiSA_i^S is the mean of the transformed support maps, and classification uses

    pθ(y=i∣x)=exp⁡[−d(AiS,AiQ)]∑j=1Nexp⁡[−d(AjS,AjQ)],p_\theta(y=i\mid x)=\frac{\exp[-d(A_i^S,A_i^Q)]}{\sum_{j=1}^{N}\exp[-d(A_j^S,A_j^Q)]},

    where dd is the feature distance and yy is the class label. The experiments use Euclidean distance by default and also show compatibility with cosine distance.

  6. Knowl 6 — Experimental protocol and benchmark splits

    experimental setup

    TDM was evaluated with Conv-4 and ResNet-12 backbones on seven fine-grained few-shot benchmarks: CUB-200-2011, Aircraft, meta-iNat, tiered meta-iNat, Stanford Cars, Stanford Dogs, and Oxford Pets. Images are resized to 84×8484\times84; Conv-4 produces 64×5×564\times5\times5 feature maps and ResNet-12 produces 640×5×5640\times5\times5 feature maps. CUB was tested both with human-annotated bounding-box crops and in raw-image form, while Aircraft images were bounding-box preprocessed. Standard random crop, horizontal flip, and color jitter were used for augmentation.

    The balancing parameters were fixed to α=0.5\alpha=0.5 and β=0.5\beta=0.5. To reduce overfitting, uniform random noise in [−0.2,0.2][-0.2,0.2] was added to each TDM task weight. For each NN-way KK-shot evaluation, the authors sampled 10,000 episodes with 16 query images per class and reported mean accuracy with 95% confidence intervals. The class partitions were disjoint:

    Could not parse LaTeX table
  7. Knowl 7 — CUB-200-2011 performance

    data/table

    On CUB-200-2011 with bounding-box-cropped images, TDM consistently improves the reproduced metric-based baselines under both Conv-4 and ResNet-12. The values are classification accuracies in percent; the two columns for each backbone are 1-shot and 5-shot results. Confidence intervals for the authors’ implementations were all below 0.23 percentage points.

    Could not parse LaTeX table

    For raw CUB images using ResNet-12, the reproduced baselines and their TDM versions were ProtoNet: 78.58±0.22→79.11±0.2278.58\pm0.22\rightarrow79.11\pm0.22 (1-shot) and 89.83±0.12→90.83±0.1189.83\pm0.12\rightarrow90.83\pm0.11 (5-shot); DSN: 80.47±0.20→80.58±0.2080.47\pm0.20\rightarrow80.58\pm0.20 and 89.92±0.12→89.95±0.1289.92\pm0.12\rightarrow89.95\pm0.12; CTX: 80.95±0.21→83.45±0.1980.95\pm0.21\rightarrow83.45\pm0.19 and 91.54±0.11→92.49±0.1191.54\pm0.11\rightarrow92.49\pm0.11; and FRN: 83.54±0.19→84.36±0.1983.54\pm0.19\rightarrow84.36\pm0.19 and 92.96±0.10→93.37±0.1092.96\pm0.10\rightarrow93.37\pm0.10. The results show that TDM remains beneficial without relying on cropped inputs.

  8. Knowl 8 — Aircraft and meta-iNat performance

    data/table

    TDM improves nearly every reproduced baseline on Aircraft, meta-iNat, and tiered meta-iNat. Results are classification accuracies in percent; the Aircraft columns compare Conv-4 and ResNet-12, while the meta-iNat columns use Conv-4. The meta-iNat evaluation uses standard 5-way episodes, including for tiered meta-iNat, whose test super-categories are disjoint from the training super-categories.

    Could not parse LaTeX table
    Could not parse LaTeX table

    The only reported decrease is FRN on tiered meta-iNat in the 5-shot setting, where accuracy changes from 63.4563.45 to 62.9162.91. On the additional Stanford Cars, Stanford Dogs, and Oxford Pets experiments with Conv-4, TDM improved every evaluated baseline with non-overlapping confidence intervals; the paper reports average gains of 4.44 percentage points in 1-shot and 3.27 percentage points in 5-shot classification.

  9. Knowl 9 — Ablation of TDM components and design choices

    empirical result

    A ProtoNet with a Conv-4 backbone was evaluated on cropped CUB and Aircraft to isolate the effects of TDM’s components. Both SAM and QAM improve the baseline, and their combination is stronger than either component alone. On cropped CUB 1-shot, SAM alone raises accuracy from 62.9062.90 to 68.5368.53, QAM alone raises it to 65.1165.11, and the combination reaches 69.9469.94. Average pooling is better than max pooling in the reported CUB 1-shot, CUB 5-shot, Aircraft 1-shot, and Aircraft 5-shot settings. TDM also improves ProtoNet when cosine distance replaces the default Euclidean distance.

    Could not parse LaTeX table
    Could not parse LaTeX table

    With cosine distance, CUB cropped accuracy changes from 68.6968.69 to 70.4770.47 in 1-shot and from 82.8982.89 to 84.3484.34 in 5-shot; Aircraft accuracy changes from 48.3648.36 to 49.2149.21 in 1-shot and from 63.4563.45 to 66.2666.26 in 5-shot. The authors therefore use average pooling and conclude that TDM is not tied to a particular commonly used distance metric.

  10. Knowl 10 — Scope limitation

    limitation

    TDM is designed to suppress broadly shared object information and focus on fine-grained discriminative details. Consequently, its benefit may be limited on coarse-grained classification tasks, where representing the whole object rather than selectively emphasizing small distinguishing regions may be more appropriate.

Coverage note — The complete prior-method comparison rows and the individual bar-chart values for Stanford Cars, Stanford Dogs, and Oxford Pets were omitted; the central TDM results, aggregate gains, component ablations, and limitations are retained.

References

  1. 1.Arman Afrasiyabi, Jean-François Lalonde, and Christian Gagné. Associative alignment for few-shot image classification. In European Conference on Computer Vision, pages 18–35. Springer, 2020. 6
  2. 2.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations, 2015. 3
  3. 3.Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, and Jia-Bin Huang. A closer look at few-shot classification. In International Conference on Learning Representations, 2019. 1, 6, 7
  4. 4.Zhengyu Chen, Jixie Ge, Heshen Zhan, Siteng Huang, and Donglin Wang. Pareto self-supervised training for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13663–13672, 2021. 7
  5. 5.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 1
  6. 6.Yao Ding, Yanzhao Zhou, Yi Zhu, Qixiang Ye, and Jianbin Jiao. Selective sparse sampling for fine-grained image recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6599–6608, 2019. 2, 5
  7. 7.Carl Doersch, Ankush Gupta, and Andrew Zisserman. Crosstransformers: spatially-aware few-shot transfer. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. 2, 3, 5, 6, 7, 8
  8. 8.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning, pages 1126–1135. PMLR, 2017. 1, 2, 6
  9. 9.Weifeng Ge, Xiangru Lin, and Yizhou Yu. Weakly supervised complementary parts models for fine-grained image classification from the bottom up. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3034–3043, 2019. 2, 5
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 1
  11. 11.Jie Hong, Pengfei Fang, Weihao Li, Tong Zhang, Christian Simon, Mehrtash Harandi, and Lars Petersson. Reinforced attention for few-shot learning and beyond. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 913–923, 2021. 7
  12. 12.Ruibing Hou, Hong Chang, Bingpeng Ma, Shiguang Shan, and Xilin Chen. Cross attention network for few-shot classification. Advances in Neural Information Processing Systems, 32, 2019. 2
  13. 13.Zeyi Huang, Haohan Wang, Eric P Xing, and Dong Huang. Self-challenging improves cross-domain generalization. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16, pages 124–140. Springer, 2020. 5
  14. 14.Dahyun Kang, Heeseung Kwon, Juhong Min, and Minsu Cho. Relational embedding for few-shot classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8822–8833, 2021. 2, 3, 6
  15. 15.Aditya Khosla, Nityananda Jayadevaprakash, Bangpeng Yao, and Fei-Fei Li. Novel dataset for fine-grained image categorization: Stanford dogs. In Proc. CVPR Workshop on Fine-Grained Visual Categorization (FGVC), volume 2. Citeseer, 2011. 6
  16. 16.Valentin Khrulkov, Leyla Mirvakhabova, Evgeniya Ustinova, Ivan Oseledets, and Victor Lempitsky. Hyperbolic image embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6418–6428, 2020. 1, 2
  17. 17.Jongmin Kim, Taesup Kim, Sungwoong Kim, and Chang D Yoo. Edge-labeling graph neural network for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11–20, 2019. 7
  18. 18.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, pages 554–561, 2013. 6
  19. 19.Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, and Stefano Soatto. Meta-learning with differentiable convex optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10657–10665, 2019. 2
  20. 20.Aoxue Li, Weiran Huang, Xu Lan, Jiashi Feng, Zhenguo Li, and Liwei Wang. Boosting few-shot learning with adaptive margin loss. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12576–12584, 2020. 2
  21. 21.Bohao Li, Boyu Yang, Chang Liu, Feng Liu, Rongrong Ji, and Qixiang Ye. Beyond max-margin: Class margin equilibrium for few-shot object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7363–7372, 2021. 5
  22. 22.Congcong Li, Dawei Du, Libo Zhang, Longyin Wen, Tiejian Luo, Yanjun Wu, and Pengfei Zhu. Spatial attention pyramid network for unsupervised domain adaptation. In European Conference on Computer Vision, pages 481–497. Springer, 2020. 2
  23. 23.Junjie Li, Zilei Wang, and Xiaoming Hu. Learning intact features by erasing-inpainting for few-shot classification. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 8401–8409, 2021. 5
  24. 24.Wenbin Li, Lei Wang, Jinglin Xu, Jing Huo, Yang Gao, and Jiebo Luo. Revisiting local descriptor based image-to-class measure for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7260–7268, 2019. 6
  25. 25.Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. A structured self-attentive sentence embedding. In International Conference on Learning Representations, 2017. 3
  26. 26.Bin Liu, Yue Cao, Yutong Lin, Qi Li, Zheng Zhang, Mingsheng Long, and Han Hu. Negative margin matters: Understanding margin in few-shot classification. In European Conference on Computer Vision, pages 438–455. Springer, 2020. 6
  27. 27.Chuanbin Liu, Hongtao Xie, Zheng-Jun Zha, Lingfeng Ma, Lingyun Yu, and Yongdong Zhang. Filtration and distillation: Enhancing region attention for fine-grained visual categorization. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11555–11562, 2020. 2, 5
  28. 28.S. Maji, J. Kannala, E. Rahtu, M. Blaschko, and A. Vedaldi. Fine-grained visual classification of aircraft. Technical report, 2013. 6
  29. 29.Puneet Mangla, Nupur Kumari, Abhishek Sinha, Mayank Singh, Balaji Krishnamurthy, and Vineeth N Balasubramanian. Charting the right manifold: Manifold mixup for few-shot learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2218–2227, 2020. 6
  30. 30.Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pages 3498–3505. IEEE, 2012. 6
  31. 31.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. In International Conference on Learning Representations, 2017. 2
  32. 32.Ryne Roady, Tyler L Hayes, Ronald Kemker, Ayesha Gonzales, and Christopher Kanan. Are open set classification methods effective on large-scale datasets? Plos one, 15(9):e0238302, 2020. 1, 3
  33. 33.Christian Simon, Piotr Koniusz, Richard Nock, and Mehrtash Harandi. Adaptive subspaces for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4136–4145, 2020. 2, 3, 5, 6, 7, 8
  34. 34.Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical networks for few-shot learning. Advances in neural information processing systems, 30, 2017. 1, 2, 4, 5, 6, 7, 8
  35. 35.Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip HS Torr, and Timothy M Hospedales. Learning to compare: Relation network for few-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1199–1208, 2018. 1, 2, 6
  36. 36.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8769–8778, 2018. 6
  37. 37.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017. 3
  38. 38.Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. Advances in neural information processing systems, 29:3630–3638, 2016. 1, 2, 6
  39. 39.C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The Caltech-UCSD Birds-200-2011 Dataset. Technical report, 2011. 6
  40. 40.Yan Wang, Wei-Lun Chao, Kilian Q Weinberger, and Laurens van der Maaten. Simpleshot: Revisiting nearest-neighbor classification for few-shot learning. arXiv preprint arXiv:1911.04623, 2019. 7
  41. 41.Davis Wertheimer and Bharath Hariharan. Few-shot learning with localization in realistic settings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6558–6567, 2019. 6
  42. 42.Davis Wertheimer, Luming Tang, and Bharath Hariharan. Few-shot classification with feature map reconstruction networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8012–8021, 2021. 2, 3, 5, 6, 7, 8
  43. 43.Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. Cbam: Convolutional block attention module. In Proceedings of the European conference on computer vision (ECCV), pages 3–19, 2018. 2
  44. 44.Chengming Xu, Yanwei Fu, Chen Liu, Chengjie Wang, Jilin Li, Feiyue Huang, Li Zhang, and Xiangyang Xue. Learning dynamic alignment via meta-filter for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5182–5191, 2021. 2, 3
  45. 45.Han-Jia Ye, Hexiang Hu, De-Chuan Zhan, and Fei Sha. Few-shot learning via embedding adaptation with set-to-set functions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8808–8817, 2020. 2, 3, 6, 7
  46. 46.Baoquan Zhang, Xutao Li, Yunming Ye, Zhichao Huang, and Lisai Zhang. Prototype completion with primitive knowledge for few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3754–3762, 2021. 1, 2, 3
  47. 47.Chi Zhang, Yujun Cai, Guosheng Lin, and Chunhua Shen. Deepemd: Few-shot image classification with differentiable earth mover’s distance and structured classifiers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12203–12213, 2020. 2, 6, 7
  48. 48.Hongguang Zhang, Piotr Koniusz, Songlei Jian, Hongdong Li, and Philip HS Torr. Rethinking class relations: Absolute-relative supervised and unsupervised few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9432–9441, 2021. 7
  49. 49.Jiabao Zhao, Yifan Yang, Xin Lin, Jing Yang, and Liang He. Looking wider for better adaptive representation in few-shot learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 10981–10989, 2021. 7
  50. 50.Heliang Zheng, Jianlong Fu, Zheng-Jun Zha, and Jiebo Luo. Looking for the devil in the details: Learning trilinear attention sampling network for fine-grained image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5012–5021, 2019. 2, 5

Citation

MLA
Lee, S., et al. “Task Discrepancy Maximization for Fine-grained Few-Shot Classification”. IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR), 2022, 2022, http://arxiv.org/abs/2207.01376v1.
APA
Lee, S., Moon, W., & Heo, J.-P. (2022). Task Discrepancy Maximization for Fine-grained Few-Shot Classification. IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR), 2022. http://arxiv.org/abs/2207.01376v1
Chicago
Lee, S., W. Moon, and J.-P. Heo. 2022. “Task Discrepancy Maximization for Fine-grained Few-Shot Classification”. IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR), 2022. http://arxiv.org/abs/2207.01376v1.
Harvard
Lee, S., Moon, W. and Heo, J.-P. (2022) “Task Discrepancy Maximization for Fine-grained Few-Shot Classification”, IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR), 2022 [Preprint]. Available at: http://arxiv.org/abs/2207.01376v1.
Vancouver
1. Lee S, Moon W, Heo J-P (2022) Task Discrepancy Maximization for Fine-grained Few-Shot Classification. IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR), 2022

BibTeX

@article{lee2022task,
  title = {Task Discrepancy Maximization for Fine-grained Few-Shot Classification},
  author = {Lee, SuBeen and Moon, WonJun and Heo, Jae-Pil},
  year = {2022},
  journal = {IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR), 2022},
  url = {http://arxiv.org/abs/2207.01376v1},
  eprint = {2207.01376}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE