Revisiting Prototypical Network for Cross Domain Few-Shot Learning

Fei ZhouPeng WangLei ZhangWei WeiYanning Zhang

article2023CVPR97 citations

Proposes a local-global distillation framework that mitigates simplicity bias in Prototypical Networks by enforcing semantic consistency between full images and local crops, achieving state-of-the-art cross-domain few-shot classification without target-domain fine-tuning.

Listen

Deploying image classification models to new operational environments using very few labeled training samples remains a major challenge in computer vision. Standard metric-based methods, such as Prototypical Networks, identify classes by comparing query images against category prototypes derived from support examples. However, their recognition accuracy deteriorates significantly when deployed across different domains because deep neural networks suffer from simplicity bias. Instead of learning generalizable, object-level semantic representations, models latch onto basic shortcut cues like simple colors or shapes that separate training classes within a known domain but fail to transfer to novel contexts.

The article proposes and evaluates the Local-global Distillation Prototypical Network (LDP-net) to overcome this limitation. The primary objective is to demonstrate that enforcing prediction consistency between whole images and their localized sub-regions during training produces robust, transferable semantic representations that generalize across varied target domains without requiring target-domain model fine-tuning.

The researchers developed a two-branch framework. The global branch processes full images, while the local branch evaluates random image crops. The system transfers knowledge between branches via distillation across three levels: self-image distillation (aligning a full image with its own local crops), cross-image distillation (aligning local crops with another image from the same class to reduce intra-class variation), and cross-episode distillation (updating the local branch via exponential moving average to preserve cumulative knowledge). The model was trained solely on a single source dataset (mini-ImageNet) and evaluated across eight target domains spanning fine-grained natural categories (CUB, Cars, Places, Plantae) and distinct specialized disciplines (CropDisease, EuroSAT satellite imagery, ISIC dermatology, and ChestX radiography).

The evaluation produced several key findings. First, LDP-net consistently outperformed the baseline Prototypical Network, yielding accuracy gains of 3 to 10 percentage points across evaluated benchmarks. Second, when tested across eight domains without target-domain adaptation, LDP-net achieved new state-of-the-art results, recording average accuracies of 46.34% in 1-shot tasks and 62.60% in 5-shot tasks—surpassing prior leading approaches by 1.49% and 1.02%, respectively. Third, when utilizing full task data via iterative semi-supervised classifier updates, LDP-net reached 50.85% (1-shot) and 64.10% (5-shot) average accuracy, even outperforming competing models that underwent full target-domain parameter fine-tuning. Fourth, visual analyses confirmed that the model broadened its attention across entire objects rather than fixating on localized shortcut patches, resulting in a smoother, more robust optimization landscape.

These findings indicate that enforcing multi-scale semantic consistency successfully eliminates shortcut bias, allowing models to learn durable features that generalize across domains. For practical applications, this provides substantial operational value: organizations can deploy a single pre-trained feature extractor into diverse downstream environments without costly, compute-heavy fine-tuning or specialized target adaptation workflows.

Based on these results, decision-makers should consider adopting local-global distillation strategies when building vision systems intended for cross-domain few-shot deployment. Organizations can implement the pre-trained feature extractor out-of-the-box for low-latency, resource-constrained tasks, or pair it with iterative label expansion when task data is available for additional accuracy gains. However, leadership should note that absolute classification accuracy remains low in highly dissimilar domains featuring extreme visual shift, such as chest radiography (22.21% to 26.88%) and dermatological imaging (33.44% to 48.44%). Before deploying systems to safety-critical or specialized medical applications, further research is required to explore expanded multi-domain source training and more data-efficient target adaptation strategies.

  • Paper: Prototypical Networks for Few-shot Learning, Jake Snell et al. (2017). It introduces Prototypical Networks, the foundational metric-based few-shot classification framework that this paper directly analyzes and builds upon to solve cross-domain degradation.
  • Paper: A Closer Look at Few-shot Classification, Wei-Yu Chen et al. (2019). It systematically establishes the cross-domain few-shot classification benchmark and reveals significant generalization drops in Prototypical Networks under domain shifts.
  • Paper: Relational Knowledge Distillation, Wonpyo Park et al. (2019). It explores relational knowledge distillation across metric embeddings, providing relevant foundations for distillation-based representation alignment across tasks.
  • Paper: Learning to Generalize: Meta-Learning for Domain Generalization, Da Li et al. (2017). It formalizes episodic meta-learning for domain generalization, establishing concepts for training models to generalize to unseen target domains.
  • Paper: TADAM: Task dependent adaptive metric for improved few-shot learning, Boris N. Oreshkin et al. (2018). It demonstrates how adaptive metric conditioning and auxiliary feature regularization enhance the generalization capabilities of prototype-based metric learners.

No sufficiently relevant recommendations were found.

Cover for Revisiting Prototypical Network for Cross Domain Few-Shot Learning

Abstract

Prototypical Network is a popular few-shot solver that aims at establishing a feature metric generalizable to novel few-shot classification (FSC) tasks using deep neural networks. However, its performance drops dramatically when generalizing to the FSC tasks in new domains. In this study, we revisit this problem and argue that the devil lies in the simplicity bias pitfall in neural networks. In specific, the network tends to focus on some biased shortcut features (e.g., color, shape, etc.) that are exclusively sufficient to distinguish very few classes in the meta-training tasks within a pre-defined domain, but fail to generalize across domains as some desirable semantic features. To mitigate this problem, we propose a Local-global Distillation Prototypical Network (LDP-net). Different from the standard Prototypical Network, we establish a two-branch network to classify the query image and its random local crops, respectively. Then, knowledge distillation is conducted among these two branches to enforce their class affiliation consistency. The rationale behind is that since such global-local semantic relationship is expected to hold regardless of data domains, the local-global distillation is beneficial to exploit some cross-domain transferable semantic features for feature metric establishment. Moreover, such local-global semantic consistency is further enforced among different images of the same class to reduce the intra-class semantic variation of the resultant feature. In addition, we propose to update the local branch as Exponential Moving Average (EMA) over training episodes, which makes it possible to better distill cross-episode knowledge and further enhance the generalization performance. Experiments on eight cross-domain FSC benchmarks empirically clarify our argument and show the state-of-the-art results of LDP-net. Code is available in https://github.com/NWPUZhoufei/LDP-Net

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Preliminary knowledge about the ProtoNet
  • 3.2. The proposed LDP-net
  • 4. Experimental Analysis
  • 4.1. Experimental details
  • 4.2. Experimental results
  • 4.2.1 Comparison with the ProtoNet baseline
  • 4.2.2 Comparison with state-of-the-art methods
  • 4.2.3 Ablation study
  • 5. Conclusions
  • 6. Acknowledgment
  • References

Knowls

  1. Knowl 1 — Local-Global Distillation Prototypical Network (LDP-net) Architecture

    model/method

    The Local-Global Distillation Prototypical Network (LDP-net) addresses the simplicity bias pitfall of standard Prototypical Networks (ProtoNet) in cross-domain few-shot classification (CD-FSC). Standard ProtoNet often latches onto simple shortcut features (such as color or shape) that distinguish classes within the source domain but fail to generalize to novel target domains.

    LDP-net adopts a two-branch architecture:

    1. Global Branch (fhetasf_{ heta_s}): Operates on the full raw query image XQiX_{Qi} and extracts global feature embeddings, maintaining identical structure and prototype-matching classification behavior as standard ProtoNet.
    2. Local Branch (fhetatf_{ heta_t}): Operates on RR local crops Xqi,rX_{qi,r} (r∈[1,R]r \in [1, R]) generated via multi-crop data augmentation from the query image and resized to a lower resolution.

    Knowledge distillation is performed between the global and local branches to enforce class affiliation prediction consistency. Because global-local semantic relationships are inherently domain-invariant, transferring fine-grained local semantic features into the global representation reduces reliance on shortcut patterns. Furthermore, the local branch parameters θt\theta_t are updated as an Exponential Moving Average (EMA) of the global branch parameters θs\theta_s across meta-training episodes to accumulate cross-episode knowledge without requiring extra gradient updates. At test time, the local branch is discarded, and only the frozen global feature extractor fhetasf_{ heta_s} is deployed.

  2. Knowl 2 — Self-Image and Cross-Image Local-Global Distillation Formulations

    equation

    Let an NN-way KK-shot episode T\mathcal{T} consist of a support set TS={XSn,k}n=1,k=1N,K\mathcal{T}_S = \{X_{S_{n,k}}\}_{n=1, k=1}^{N, K} and a query set TQ={XQi}i=1I\mathcal{T}_Q = \{X_{Qi}\}_{i=1}^I, where NN is the number of classes, KK is the number of support samples per class, and II is the total number of query images. For each class n∈[1,N]n \in [1, N], the class prototype CnC_n is computed using the global feature extractor fhetasf_{ heta_s}:

    Cn=1K∑k=1Kfhetas(XSn,k)C_n = \frac{1}{K} \sum_{k=1}^K f_{ heta_s}(X_{S_{n,k}})

    The global classification prediction vector PQi=[PQi1,…,PQiN]∈RNP_{Qi} = [P_{Qi}^1, \dots, P_{Qi}^N] \in \mathbb{R}^N for a query image XQiX_{Qi} is obtained by metric matching between fhetas(XQi)f_{ heta_s}(X_{Qi}) and prototypes {Cn}n=1N\{C_n\}_{n=1}^N.

    For RR local crops {Xqi,r}r=1R\{X_{qi,r}\}_{r=1}^R of query image XQiX_{Qi}, local features fhetat(Xqi,r)f_{ heta_t}(X_{qi,r}) produce local prediction vectors Pqi,r∈RNP_{qi,r} \in \mathbb{R}^N by matching against the same class prototypes.

    1. Self-Image Local-Global Distillation Loss: Enforces consistency between the global prediction and local crop predictions of the same query image:

    LXQiself=1R∑r=1RH(Pqi,r,PQi)\mathcal{L}_{X_{Qi}}^{\text{self}} = \frac{1}{R} \sum_{r=1}^R H(P_{qi,r}, P_{Qi})

    where H(p,q)=−∑n=1Nqnlog⁡pnH(p, q) = -\sum_{n=1}^N q^n \log p^n denotes the cross-entropy function.

    2. Cross-Image Local-Global Distillation Loss: To reduce intra-class semantic variation while preventing model collapse (which occurs if all query samples of the same class are constrained simultaneously), one query sample XQjX_{Qj} (j≠ij \neq i) from the same category is randomly selected to provide local crop predictions Pqj,rP_{qj,r}:

    LXQicross=1R∑r=1RH(Pqj,r,PQi),j≠i\mathcal{L}_{X_{Qi}}^{\text{cross}} = \frac{1}{R} \sum_{r=1}^R H(P_{qj,r}, P_{Qi}), \quad j \neq i

  3. Knowl 3 — Cross-Episode Knowledge Distillation via Exponential Moving Average

    equation

    In meta-learning, each sampled episode draws a distinct combination of classes from the source dataset. Directly updating local and global branches simultaneously via episode gradients introduces optimization instability and fails to accumulate visual knowledge smoothly across episodes.

    To enable cross-episode knowledge transfer and stabilize distillation, the parameters θt\theta_t of the local branch network fhetatf_{ heta_t} are updated after each episode as an Exponential Moving Average (EMA) of the global network parameters θs\theta_s:

    θt←mθt+(1−m)θs\theta_t \leftarrow m \theta_t + (1 - m) \theta_s

    where m∈[0,1)m \in [0, 1) is the momentum hyperparameter (set to m=0.998m = 0.998 in implementation). This allows the local network to act as a slowly evolving, stable teacher that distills accumulated cross-episode representations into the global branch.

  4. Knowl 4 — Total Meta-Training Objective of LDP-net

    equation

    For a given training episode T\mathcal{T} with II query samples, the overall loss function LTS\mathcal{L}_{TS} used to optimize the global feature extractor parameters θs\theta_s is defined as:

    LTS=∑i=1ILXQice+λ1∑i=1ILXQiself+λ2∑i=1ILXQicross\mathcal{L}_{TS} = \sum_{i=1}^I \mathcal{L}_{X_{Qi}}^{\text{ce}} + \lambda_1 \sum_{i=1}^I \mathcal{L}_{X_{Qi}}^{\text{self}} + \lambda_2 \sum_{i=1}^I \mathcal{L}_{X_{Qi}}^{\text{cross}}

    where:

    • LXQice=H(y^Qi,yQi)\mathcal{L}_{X_{Qi}}^{\text{ce}} = H(\hat{y}_{Qi}, y_{Qi}) is the standard supervised cross-entropy loss between the predicted label distribution y^Qi\hat{y}_{Qi} (derived from global prediction PQiP_{Qi}) and the ground truth class label yQiy_{Qi}.
    • LXQiself\mathcal{L}_{X_{Qi}}^{\text{self}} is the self-image local-global distillation loss.
    • LXQicross\mathcal{L}_{X_{Qi}}^{\text{cross}} is the cross-image local-global distillation loss.
    • λ1\lambda_1 and λ2\lambda_2 are non-negative weighting coefficients (empirically set to λ1=1.0\lambda_1 = 1.0 and λ2=0.15\lambda_2 = 0.15).
  5. Knowl 5 — LDP-net Meta-Training Algorithm

    algorithm

    The meta-training pipeline trains the global feature extractor fθsf_{\theta_s} using episodic mini-batches from the source domain Ds\mathcal{D}_s alongside local-global distillation against an EMA-updated local extractor fθtf_{\theta_t}.

    Input: Source domain dataset Ds\mathcal{D}_s, global feature extractor fθsf_{\theta_s} with parameters θs\theta_s, local feature extractor fθtf_{\theta_t} with parameters θt\theta_t, loss weights λ1,λ2\lambda_1, \lambda_2, momentum coefficient mm, crop count RR.
    Output: Optimized global feature extractor fθsf_{\theta_s}.
    while not converged do
        Sample a few-shot episode T=(TS,TQ)\mathcal{T} = (\mathcal{T}_S, \mathcal{T}_Q) from Ds\mathcal{D}_s
        Compute class prototypes Cn=1K∑k=1Kfθs(XSn,k)C_n = \frac{1}{K} \sum_{k=1}^K f_{\theta_s}(X_{S_{n,k}}) for all classes n∈[1,N]n \in [1, N]
        
        for each query image XQi∈TQX_{Qi} \in \mathcal{T}_Q do
            Generate RR local crops {Xqi,r}r=1R\{X_{qi,r}\}_{r=1}^R using multi-crop augmentation
            Compute global prediction PQi=matching(fθs(XQi),{Cn}n=1N)P_{Qi} = \text{matching}(f_{\theta_s}(X_{Qi}), \{C_n\}_{n=1}^N)
            Compute local predictions Pqi,r=matching(fθt(Xqi,r),{Cn}n=1N)P_{qi,r} = \text{matching}(f_{\theta_t}(X_{qi,r}), \{C_n\}_{n=1}^N) for r∈[1,R]r \in [1, R]
            Randomly sample XQj∈TQX_{Qj} \in \mathcal{T}_Q from the same class as XQiX_{Qi} with j≠ij \neq i
            Compute local predictions Pqj,r=matching(fθt(Xqj,r),{Cn}n=1N)P_{qj,r} = \text{matching}(f_{\theta_t}(X_{qj,r}), \{C_n\}_{n=1}^N) for r∈[1,R]r \in [1, R]
            Compute LXQiself=1R∑r=1RH(Pqi,r,PQi)\mathcal{L}_{X_{Qi}}^{\text{self}} = \frac{1}{R} \sum_{r=1}^R H(P_{qi,r}, P_{Qi})
            Compute LXQicross=1R∑r=1RH(Pqj,r,PQi)\mathcal{L}_{X_{Qi}}^{\text{cross}} = \frac{1}{R} \sum_{r=1}^R H(P_{qj,r}, P_{Qi})
            Compute LXQice=H(PQi,yQi)\mathcal{L}_{X_{Qi}}^{\text{ce}} = H(P_{Qi}, y_{Qi})
        end for
        
        Compute total loss LTS=∑i=1I(LXQice+λ1LXQiself+λ2LXQicross)\mathcal{L}_{TS} = \sum_{i=1}^I (\mathcal{L}_{X_{Qi}}^{\text{ce}} + \lambda_1 \mathcal{L}_{X_{Qi}}^{\text{self}} + \lambda_2 \mathcal{L}_{X_{Qi}}^{\text{cross}})
        Update θs\theta_s via gradient descent on LTS\mathcal{L}_{TS}
        Update θt←mθt+(1−m)θs\theta_t \leftarrow m \theta_t + (1 - m) \theta_s
    end while
    return fθsf_{\theta_s}
  6. Knowl 6 — Cross-Domain Evaluation Protocol and Unsupervised Query Augmentation (LDP-net†)

    model/method

    Cross-domain few-shot evaluation evaluates the model trained on source domain Ds\mathcal{D}_s (mini-ImageNet) directly on target domain datasets Dt\mathcal{D}_t without test-time fine-tuning of the feature extractor backbone.

    1. Standard Evaluation (LDP-net): The local branch is removed. For each sampled target episode, the frozen global network fθsf_{\theta_s} extracts feature vectors for support and query samples. A Logistic Regression classifier is trained exclusively on the support feature representations and used to classify query features.
    2. Full-Data Evaluation Protocol (LDP-net†\text{LDP-net}^\dagger): When evaluating under the setting where unlabeled target query data can be utilized transductively without fine-tuning backbone weights, an iterative pseudo-labeling procedure is employed:
      • Train the initial classifier on the support set features.
      • Predict labels and class probabilities for query samples.
      • Select query samples with high prediction confidence and add them to the support set with their predicted pseudo-labels.
      • Retrain the classifier on the augmented support set.
      • Repeat this augmentation cycle for several iterations, and use the final classifier to produce the final query predictions.
  7. Knowl 7 — Single-Source CD-FSC Experimental Setup

    experimental setup

    The single-source cross-domain few-shot classification benchmark is configured as follows:

    • Source Domain: The 64 training classes of mini-ImageNet.
    • Target Domains (8 datasets): Natural images with domain shifts: CUB, Cars, Places, Plantae; Cross-domain datasets with extreme domain shifts (medical, agricultural, remote sensing): ChestX, ISIC, EuroSAT, CropDisease. All images are resized to 224×224224 \times 224 pixels.
    • Backbone: ResNet-10, pre-trained on the source domain via standard supervised classification prior to episodic meta-training.
    • Training Details: Meta-trained for 100 epochs using the Adam optimizer with a learning rate of 0.0010.001. Each epoch consists of 100 randomly sampled episodes. Default training episodes use N=5N=5 classes, K=5K=5 support shots, and 15 query samples per class.
    • Hyperparameters: Loss weights λ1=1.0\lambda_1 = 1.0, λ2=0.15\lambda_2 = 0.15; EMA momentum m=0.998m = 0.998; number of local crops R=6R = 6. The model checkpoint from the final epoch (epoch 100) is evaluated directly without validation set selection.
    • Testing Protocol: 600 randomly sampled 5-way 1-shot or 5-way 5-shot episodes (each with 15 query images per class) per target dataset, reporting mean top-1 accuracy (%) and 95% confidence intervals.
  8. Knowl 8 — Performance Comparison on Cross-Domain FSC Benchmarks

    data/table

    LDP-net and its transductive variant LDP-net†\text{LDP-net}^\dagger outperform previous state-of-the-art CD-FSC methods across 8 target domains in 5-way 1-shot and 5-way 5-shot settings under various protocols (Case 1: no fine-tuning, no query data used; Case 2: no fine-tuning, query data used via †\dagger; Case 3: target fine-tuning with query data via ∗†*\dagger).

    Method Ft CUB Cars Places Plantae Chest ISIC EuroSAT CropDisease Ave.
    5-way 1-shot
    RelationNet 41.27 30.09 48.16 31.23 21.95 30.53 49.08 53.58 38.24
    GNN 44.40 31.72 52.42 33.60 21.94 30.14 54.61 59.19 41.00
    GNN+AFA 46.86 34.25 54.04 36.76 22.92 33.21 63.12 67.61 44.85
    LDP-net (ours) 49.82 35.51 53.82 39.84 23.01 33.97 65.11 69.64 46.34
    TPN+AFA†^\dagger 50.85 38.43 60.29 40.27 21.69 34.25 66.17 72.44 48.05
    RDC†^\dagger 47.77 38.74 58.82 41.88 22.66 32.29 67.58 80.88 48.83
    LDP-net†^\dagger (ours) 55.94 37.44 62.21 41.04 22.21 33.44 73.25 81.24 50.85
    RDC∗†^{*\dagger} 50.09 39.04 61.17 41.30 22.32 36.28 70.51 85.79 50.81
    5-way 5-shot
    GNN+AFA 68.25 49.28 76.21 54.26 25.02 46.01 85.58 88.06 61.58
    LDP-net (ours) 70.39 52.84 72.90 58.49 26.67 48.06 82.01 89.40 62.60
    TPN+AFA†^\dagger 65.86 47.89 72.81 55.67 23.47 46.29 80.12 85.69 59.73
    RDC†^\dagger 63.39 52.75 72.83 55.30 25.10 42.10 79.12 88.03 59.83
    LDP-net†^\dagger (ours) 73.34 53.06 75.47 59.64 26.88 48.44 84.05 91.89 64.10
    TPN+ATA∗†^{*\dagger} 70.14 55.23 73.87 59.02 24.74 49.83 85.47 93.56 63.98
    RDC∗†^{*\dagger} 67.23 53.49 74.91 57.47 25.07 49.91 84.29 93.30 63.21

    Without fine-tuning on the target domain, LDP-net outperforms GNN+AFA by 1.49% on 1-shot average accuracy and by 1.02% on 5-shot average accuracy. In the transductive full-data setting, LDP-net†\text{LDP-net}^\dagger achieves 50.85% (1-shot) and 64.10% (5-shot) average accuracy, surpassing previous best methods including those that fine-tune backbone parameters on target tasks (RDC∗†RDC^{*\dagger} at 50.81% and 63.21%).

  9. Knowl 9 — Ablation Study of LDP-net Distillation Components

    data/table

    An ablation study isolating the self-image, cross-image, and cross-episode distillation modules demonstrates that each component provides incremental generalization gains over standard ProtoNet and pre-trained ProtoNet++ on both natural (CUB, Cars) and extreme (EuroSAT, ISIC) domain shift benchmarks using Euclidean distance classification.

    Distillation Component CUB Cars EuroSAT ISIC
    Self-img Cross-img Cross-ep 1-shot 5-shot 1-shot 5-shot 1-shot 5-shot 1-shot 5-shot
    ProtoNet 41.77 58.98 29.79 41.16 57.50 74.44 30.65 40.42
    ProtoNet++ 40.34 61.94 31.63 46.56 59.11 81.44 31.73 44.01
    43.99 63.80 32.87 49.58 62.31 80.22 32.88 45.00
    44.04 64.01 33.33 49.69 63.03 80.07 33.28 45.35
    47.70 68.94 34.65 51.61 63.70 80.26 33.51 46.42

    Key observations:

    • Self-image distillation alone outperforms ProtoNet++ across datasets (e.g., +3.65% 1-shot on CUB, +3.20% 1-shot on EuroSAT).
    • Cross-image distillation adds consistency across instances of the same category, yielding consistent further gains.
    • Cross-episode EMA distillation provides substantial performance boosts (e.g., +3.66% 1-shot and +4.93% 5-shot on CUB over the two-loss baseline without EMA), confirming the effectiveness of stable cross-episode knowledge accumulation.
  10. Knowl 10 — Qualitative Feature Activation and Loss Landscape Generalization

    empirical result

    Qualitative feature analysis confirms that LDP-net overcomes the simplicity bias of baseline prototypical networks:

    1. Class Activation Mapping (CAM): Visualizations show that ProtoNet++ focuses narrowly on small, isolated shortcut regions (e.g., single patches of color or contour), whereas LDP-net activates broad and comprehensive semantic regions of target objects.
    2. Loss Landscape Analysis: When evaluating models on the CUB target domain under 2,000 random parameter perturbation directions from the source domain solution, LDP-net produces a smoother loss landscape with significantly wider optimal contour basins than ProtoNet++. Wider, flatter loss minima indicate enhanced cross-domain generalization and robustness to domain shift.
  11. Knowl 11 — Limitations Under Extreme Domain Discrepancies

    limitation

    Although LDP-net achieves state-of-the-art cross-domain transfer performance across eight benchmarks, its classification accuracy remains comparatively low on target domains characterized by extreme distributional shifts from natural images, specifically ChestX (23.01% 1-shot, 26.67% 5-shot) and ISIC (33.97% 1-shot, 48.06% 5-shot). Mitigating such massive domain gaps may require expanding the diversity of source pre-training datasets or developing more expressive target-domain adaptation mechanisms.

Coverage note — None was omitted. All contributed methodology, loss formulations, training algorithms, evaluation settings, empirical benchmark comparisons, ablation studies, qualitative analyses, and stated limitations are fully represented.

References

  1. 1.Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems, pages 4077–4087, 2017.
  2. 2.Yifei Huang, Lijin Yang, and Yoichi Sato. Compound prototype matching for few-shot action recognition. In European Conference on Computer Vision, pages 351–368. Springer, 2022.
  3. 3.Tao Zhang and Wu Huang. Kernel relative-prototype spectral filtering for few-shot learning. In European Conference on Computer Vision, pages 541–557. Springer, 2022.
  4. 4.Carl Doersch, Ankush Gupta, and Andrew Zisserman. Crosstransformers: spatially-aware few-shot transfer. Advances in Neural Information Processing Systems, 33:21981–21993, 2020.
  5. 5.Pan Li, Shaogang Gong, Chengjie Wang, and Yanwei Fu. Ranking distance calibration for cross-domain few-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9099–9108, 2022.
  6. 6.Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. In Advances in neural information processing systems, pages 3630–3638, 2016.
  7. 7.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The caltech-ucsd birds-200-2011 dataset. 2011.
  8. 8.Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli. The pitfalls of simplicity bias in neural networks. Advances in Neural Information Processing Systems, 33:9573–9585, 2020.
  9. 9.Damien Teney, Ehsan Abbasnejad, Simon Lucey, and Anton van den Hengel. Evading the simplicity bias: Training a diverse set of models discovers solutions with superior ood generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16761–16772, 2022.
  10. 10.Kaifeng Lyu, Zhiyuan Li, Runzhe Wang, and Sanjeev Arora. Gradient descent on two-layer nets: Margin maximization and simplicity bias. Advances in Neural Information Processing Systems, 34:12978–12991, 2021.
  11. 11.Adrien Bardes, Jean Ponce, and Yann LeCun. Vicregl: Self-supervised learning of local visual features. In NeurIPS, 2022.
  12. 12.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 1126–1135. JMLR. org, 2017.
  13. 13.Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, and Stefano Soatto. Meta-learning with differentiable convex optimization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 10657–10665, 2019.
  14. 14.Andrei A Rusu, Dushyant Rao, Jakub Sygnowski, Oriol Vinyals, Razvan Pascanu, Simon Osindero, and Raia Hadsell. Meta-learning with latent embedding optimization. In International Conference on Learning Representations, 2019.
  15. 15.Andrey Zhmoginov, Mark Sandler, and Maksym Vladymyrov. Hypertransformer: Model generation for supervised and semi-supervised few-shot learning. In International Conference on Machine Learning, pages 27075–27098. PMLR, 2022.
  16. 16.Sungyong Baik, Myungsub Choi, Janghoon Choi, Heewon Kim, and Kyoung Mu Lee. Meta-learning with adaptive hyperparameters. Advances in Neural Information Processing Systems, 33:20755–20765, 2020.
  17. 17.Yonglong Tian, Yue Wang, Dilip Krishnan, Joshua B Tenenbaum, and Phillip Isola. Rethinking few-shot image classification: a good embedding is all you need? In European Conference on Computer Vision, pages 266–282. Springer, 2020.
  18. 18.Puneet Mangla, Nupur Kumari, Abhishek Sinha, Mayank Singh, Balaji Krishnamurthy, and Vineeth N Balasubramanian. Charting the right manifold: Manifold mixup for few-shot learning. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 2218–2227, 2020.
  19. 19.Chen Liu, Yanwei Fu, Chengming Xu, Siqian Yang, Jilin Li, Chengjie Wang, and Li Zhang. Learning a few-shot embedding model with contrastive learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 8635–8643, 2021.
  20. 20.Dominic Zhao, Johannes von Oswald, Seijin Kobayashi, Joao Sacramento, and Benjamin F Grewe. Meta-learning via hypernetworks. In 4th Workshop on Meta-Learning at NeurIPS 2020 (MetaLearn 2020), 2020.
  21. 21.Bin Liu, Yue Cao, Yutong Lin, Qi Li, Zheng Zhang, Mingsheng Long, and Han Hu. Negative margin matters: Understanding margin in few-shot classification. In European conference on computer vision, pages 438–455. Springer, 2020.
  22. 22.Haoqing Wang and Zhi-Hong Deng. Cross-domain few-shot classification via adversarial task augmentation. In Zhi-Hua Zhou, editor, Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, pages 1075–1081. International Joint Conferences on Artificial Intelligence Organization, 8 2021. Main Track.
  23. 23.Yunhui Guo, Noel C Codella, Leonid Karlinsky, James V Codella, John R Smith, Kate Saenko, Tajana Rosing, and Rogerio Feris. A broader study of cross-domain few-shot learning. In European conference on computer vision, pages 124–141. Springer, 2020.
  24. 24.Hanwen Liang, Qiong Zhang, Peng Dai, and Juwei Lu. Boosting the generalization capability in cross-domain few-shot learning via noise-enhanced supervised autoencoder. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9424–9434, 2021.
  25. 25.Debasmit Das, Sungrack Yun, and Fatih Porikli. Confess: A framework for single source cross-domain few-shot learning. In International Conference on Learning Representations, 2021.
  26. 26.Wei-Hong Li, Xialei Liu, and Hakan Bilen. Universal representation learning from multiple domains for few-shot classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9526–9535, 2021.
  27. 27.Yuqian Fu, Yanwei Fu, and Yu-Gang Jiang. Meta-fdmixup: Cross-domain few-shot learning guided by labeled target data. In Proceedings of the 29th ACM International Conference on Multimedia, pages 5326–5334, 2021.
  28. 28.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  29. 29.Hung-Yu Tseng, Hsin-Ying Lee, Jia-Bin Huang, and Ming-Hsuan Yang. Cross-domain few-shot classification via learned feature-wise transformation. In International Conference on Learning Representations, 2019.
  30. 30.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. Advances in Neural Information Processing Systems, 33:9912–9924, 2020.
  31. 31.Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, and Luc V Gool. Revisiting contrastive methods for unsupervised learning of visual representations. Advances in Neural Information Processing Systems, 34:16238–16250, 2021.
  32. 32.Tri Huynh, Simon Kornblith, Matthew R Walter, Michael Maire, and Maryam Khademi. Boosting contrastive self-supervised learning with false negative cancellation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2785–2795, 2022.
  33. 33.Akash Srivastava, Lazar Valkov, Chris Russell, Michael U Gutmann, and Charles Sutton. Veegan: Reducing mode collapse in gans using implicit variational learning. Advances in neural information processing systems, 30, 2017.
  34. 34.Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip HS Torr, and Timothy M Hospedales. Learning to compare: Relation network for few-shot learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1199–1208, 2018.
  35. 35.Victor Garcia and Joan Bruna. Few-shot learning with graph neural networks. In 6th International Conference on Learning Representations, ICLR 2018, 2018.
  36. 36.Yanxu Hu and Andy J Ma. Adversarial feature augmentation for cross-domain few-shot classification. In European Conference on Computer Vision, pages 20–37. Springer, 2022.
  37. 37.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2921–2929, 2016.
  38. 38.Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. Visualizing the loss landscape of neural nets. Advances in neural information processing systems, 31, 2018.

Citation

MLA
Zhou, F., et al. “Revisiting Prototypical Network for Cross Domain Few-Shot Learning”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 20061–70, https://doi.org/10.1109/CVPR52729.2023.01921.
APA
Zhou, F., Wang, P., Zhang, L., Wei, W., & Zhang, Y. (2023). Revisiting Prototypical Network for Cross Domain Few-Shot Learning. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 20061–20070. https://doi.org/10.1109/CVPR52729.2023.01921
Chicago
Zhou, F., P. Wang, L. Zhang, W. Wei, and Y. Zhang. 2023. “Revisiting Prototypical Network for Cross Domain Few-Shot Learning”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 20061–70. https://doi.org/10.1109/CVPR52729.2023.01921.
Harvard
Zhou, F. et al. (2023) “Revisiting Prototypical Network for Cross Domain Few-Shot Learning”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 20061–20070. Available at: https://doi.org/10.1109/CVPR52729.2023.01921.
Vancouver
1. Zhou F, Wang P, Zhang L, Wei W, Zhang Y (2023) Revisiting Prototypical Network for Cross Domain Few-Shot Learning. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 20061–20070

BibTeX

@inproceedings{Zhou_2023, title={Revisiting Prototypical Network for Cross Domain Few-Shot Learning}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01921}, DOI={10.1109/cvpr52729.2023.01921}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhou, Fei and Wang, Peng and Zhang, Lei and Wei, Wei and Zhang, Yanning}, year={2023}, month=June, pages={20061–20070} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE