Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly

Yongqin XianChristoph H. LampertBernt SchieleZeynep Akata

article2017TPAMI1,899 citations

Establishes a unified evaluation benchmark and standardized data splits to resolve widespread test-set overlap in zero-shot learning, while introducing the Animals with Attributes 2 (AWA2) dataset and systematically comparing leading methods under both standard and generalized settings.

Listen

Deploying machine learning models in dynamic operational environments requires recognizing new, previously unseen object categories without requiring expensive, time-consuming data collection and annotation. Zero-shot learning addresses this need by transferring knowledge from familiar categories to unfamiliar ones using shared auxiliary descriptions, such as semantic attributes. However, recent rapid growth in proposed methods has occurred without standardized evaluation benchmarks, resulting in inconsistent experimental setups, parameter tuning on test data, and data contamination that overstates real-world model performance.

The article establishes a rigorous, unified benchmarking framework to systematically evaluate contemporary zero-shot learning methods and quantify genuine algorithmic progress. It investigates how different model architectures perform under both standard zero-shot conditions (where test queries come exclusively from unseen classes) and generalized zero-shot conditions (where test queries may belong to either seen or unseen classes).

To conduct this evaluation, the authors re-evaluated thirteen representative zero-shot learning algorithms across five benchmark image datasets—including scene, bird, and general object collections—and evaluated ten methods on a large-scale twenty-one-thousand-category dataset. The authors corrected methodological flaws by designing new dataset splits to prevent feature-extractor pre-training data from overlapping with evaluation classes. They also introduced a new fifty-class animal dataset with thirty-seven thousand publicly licensed images to ensure open reproducibility, and evaluated performance using class-balanced top-one accuracy and the harmonic mean of seen and unseen class accuracy.

The investigation produced four central findings. First, existing benchmark evaluations significantly overstated model performance due to training class contamination; under corrected splits, performance dropped across several datasets, falling by roughly fifteen to twenty percentage points on coarse-grained benchmarks. Second, bilinear compatibility learning frameworks (such as Attribute Label Embedding and Deep Visual Semantic Embedding) and generative models consistently outperformed two-stage independent attribute classifiers across standard zero-shot benchmarks. Third, in the realistic generalized zero-shot setting, overall performance fell drastically across all evaluated methods because models exhibited a strong prediction bias toward familiar training classes, which act as distractors. Finally, while transductive approaches that incorporate unlabeled test images during training improved standard zero-shot recognition, they failed to yield consistent gains under generalized zero-shot conditions.

These results demonstrate that reported zero-shot performance in early literature did not accurately reflect real-world viability, creating operational risk if deployed without recalibration. Systems deployed in open environments will frequently encounter a mix of familiar and novel inputs, where default models will heavily misclassify novel items as familiar categories. Incorporating simple novelty detection mechanisms mitigated this bias and improved balanced generalized accuracy.

Organizations developing or deploying zero-shot vision systems should adopt corrected, non-overlapping dataset splits and evaluate systems using the harmonic mean of seen and unseen performance rather than standard zero-shot accuracy alone. Machine learning teams should prioritize compatibility learning architectures or generative models over independent attribute classifiers, while integrating explicit novelty detection mechanisms to manage open-set classification trade-offs before deploying models to production.

The findings provide high confidence regarding the relative ranking of evaluated algorithms under controlled benchmark conditions. However, performance remains heavily constrained on large, fine-grained, or highly imbalanced class distributions, where top-one accuracy across broad vocabularies fell below one percent for all methods. Decision-makers should treat current zero-shot systems as assistive rather than fully autonomous tools when scaling to massive or highly rare category distributions.

Cover for Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly

Abstract

Due to the importance of zero-shot learning, i.e. classifying images where there is a lack of labeled training data, the number of proposed approaches has recently increased steadily. We argue that it is time to take a step back and to analyze the status quo of the area. The purpose of this paper is three-fold. First, given the fact that there is no agreed upon zero-shot learning benchmark, we first define a new benchmark by unifying both the evaluation protocols and data splits of publicly available datasets used for this task. This is an important contribution as published results are often not comparable and sometimes even flawed due to, e.g. pre-training on zero-shot test classes. Moreover, we propose a new zero-shot learning dataset, the Animals with Attributes 2 (AWA2) dataset which we make publicly available both in terms of image features and the images themselves. Second, we compare and analyze a significant number of the state-of-the-art methods in depth, both in the classic zero-shot setting but also in the more realistic generalized zero-shot setting. Finally, we discuss in detail the limitations of the current status of the area which can be taken as a basis for advancing it.

Table of Contents

  • I Introduction
  • II Related Work
  • III Evaluated Methods
  • III-A Learning Linear Compatibility
  • III-B Learning Nonlinear Compatibility
  • III-C Learning Intermediate Attribute Classifiers
  • III-D Hybrid Models
  • III-E Transductive Zero-Shot Learning Setting
  • IV Datasets
  • IV-A Attribute Datasets
  • IV-B Large-Scale ImageNet
  • V Evaluation Protocol
  • V-A Image and Class Embedding
  • V-B Dataset Splits
  • V-C Evaluation Criteria
  • VI Experiments
  • VI-A Zero-Shot Learning Experiments
  • VI-B Generalized Zero-Shot Learning Results
  • VI-C Transductive (Generalized) Zero-Shot Learning
  • VII Conclusion
  • References

Knowls

  1. Knowl 1 — Proposed Split Protocol for Zero-Shot Learning Benchmarks

    experimental setup

    Standard dataset splits (SS) historically used in zero-shot learning (ZSL) evaluate models on test classes that overlap with the 1,000 classes of ImageNet used to pre-train deep neural network feature extractors (e.g., ResNet-101). Specifically, 7 out of 12 test classes in aPY, 6 out of 10 in AWA1, 1 out of 50 in CUB, and 6 out of 72 in SUN appear in ImageNet 1K, violating the core zero-shot premise that test classes are unseen during feature extractor training.

    The Proposed Split (PS) protocol remedies this data contamination by ensuring that no zero-shot test class belongs to ImageNet 1K, while reserving disjoint sets of training, validation, and test classes for parameter tuning and evaluation:

    • SUN: 14,340 images across 717 fine-grained scene classes with 102 attributes. Training split contains 580 training and 65 validation classes; test split contains 72 classes (1,440 test images).
    • CUB: 11,788 images across 200 fine-grained bird species with 312 attributes. Training split contains 100 training and 50 validation classes; test split contains 50 classes (2,967 test images in PS).
    • AWA1: 30,475 images across 50 coarse animal classes with 85 attributes. Training split contains 27 training and 13 validation classes; test split contains 10 classes (5,685 test images in PS).
    • AWA2: 37,322 images across 50 coarse animal classes with 85 attributes. Training split contains 27 training and 13 validation classes; test split contains 10 classes (7,913 test images in PS).
    • aPY: 15,339 images across 32 coarse object classes with 64 attributes. Training split contains 15 training and 5 validation classes; test split contains 12 classes (7,924 test images in PS).

    Hyperparameter optimization must be conducted exclusively on the disjoint validation classes rather than on the test classes.

  2. Knowl 2 — Per-Class Accuracy and Generalized Zero-Shot Harmonic Mean Metric

    equation

    To account for class imbalance where densely populated classes might dominate evaluation, classification accuracy is computed as the average per-class top-1 accuracy for a set of evaluation classes Y\mathcal{Y}:

    \text{acc}_{\mathcal{Y}} = \frac{1}{|\mathcal{Y}|} \sum_{c=1}^{|\mathcal{Y}|} \frac{\text{# correct predictions in class } c}{\text{# total samples in class } c}

    In Generalized Zero-Shot Learning (GZSL), test images can belong to either seen training classes Ytr\mathcal{Y}_{tr} or unseen test classes Yts\mathcal{Y}_{ts}, and the search space at prediction time spans Ytr∪Yts\mathcal{Y}_{tr} \cup \mathcal{Y}_{ts}. Because models often exhibit extreme bias toward seen classes at the expense of unseen classes, the overall GZSL performance is measured by the harmonic mean HH of seen class accuracy accYtr\text{acc}_{\mathcal{Y}_{tr}} and unseen class accuracy accYts\text{acc}_{\mathcal{Y}_{ts}}:

    H=2⋅accYtr⋅accYtsaccYtr+accYtsH = \frac{2 \cdot \text{acc}_{\mathcal{Y}_{tr}} \cdot \text{acc}_{\mathcal{Y}_{ts}}}{\text{acc}_{\mathcal{Y}_{tr}} + \text{acc}_{\mathcal{Y}_{ts}}}

    The harmonic mean is used rather than the arithmetic mean because an arithmetic mean can remain high if seen class accuracy is very large even when unseen class accuracy is near zero, whereas HH is high only if performance is strong on both seen and unseen classes.

  3. Knowl 3 — Animals with Attributes 2 (AWA2) Dataset and Cross-Dataset Compatibility with AWA1

    data/table

    The Animals with Attributes 2 (AWA2) dataset consists of 37,322 images with public copyright licenses, providing a publicly redistributable replacement for the original Animals with Attributes (AWA1) dataset (30,475 images whose raw images could not be distributed due to licensing constraints). AWA2 matches AWA1 with the exact same 50 animal categories and 85 attribute annotations.

    To verify dataset compatibility, models are evaluated across both datasets under the Proposed Splits (PS). In the cross-dataset setup, a model is trained on the training set of dataset AA and evaluated on the test set of dataset BB (denoted A:BA:B):

    Method AWA1:AWA1 AWA1:AWA2 AWA2:AWA2 AWA2:AWA1
    DAP 44.1% 44.2% 46.1% 46.2%
    IAP 35.9% 36.1% 35.9% 35.3%
    CONSE 45.6% 46.5% 44.5% 43.7%
    CMT 39.5% 40.7% 37.9% 37.7%
    SSE 60.1% 61.6% 61.0% 59.8%
    LATEM 55.1% 55.4% 55.8% 53.5%
    ALE 59.9% 59.9% 62.5% 60.9%
    DEVISE 54.2% 55.2% 59.7% 57.7%
    SJE 65.6% 65.5% 61.9% 62.0%
    ESZSL 58.2% 58.5% 58.6% 59.9%
    SYNC 54.0% 53.7% 46.6% 46.9%
    SAE 53.0% 52.4% 54.1% 53.1%

    For 8 of the 12 methods, the performance difference between AWA1 and AWA2 is within 2%. A paired t-test comparing the 24 paired configurations where the evaluation set is held constant while the training set alternates (i.e. AWA1:AWA2 vs AWA2:AWA2, and AWA1:AWA1 vs AWA2:AWA1) yields p=0.007p = 0.007, indicating that variation stems from training set sample differences rather than underlying domain mismatch, confirming AWA2 as a drop-in replacement for AWA1.

  4. Knowl 4 — Unified Formulations for Linear Compatibility Zero-Shot Learning

    model/method

    Linear compatibility models score an image xx with visual feature embedding θ(x)∈Rd\theta(x) \in \mathbb{R}^d and a candidate class yy with semantic attribute embedding ϕ(y)∈Ra\phi(y) \in \mathbb{R}^a using a bilinear function parameterized by matrix W∈Rd×aW \in \mathbb{R}^{d \times a}:

    F(x,y;W)=θ(x)TWϕ(y)F(x, y; W) = \theta(x)^T W \phi(y)

    Classification assigns f(x;W)=arg⁡max⁡y∈YF(x,y;W)f(x; W) = \arg\max_{y \in \mathcal{Y}} F(x, y; W). On training data S={(xn,yn)}n=1N\mathcal{S} = \{(x_n, y_n)\}_{n=1}^N with yn∈Ytry_n \in \mathcal{Y}_{tr}, specific linear compatibility frameworks learn WW via different objective functions:

    1. DEVISE: Optimizes an unregularized ranking margin objective with Stochastic Gradient Descent (SGD): ∑y∈Ytr[Δ(yn,y)+F(xn,y;W)−F(xn,yn;W)]+\sum_{y \in \mathcal{Y}_{tr}} [\Delta(y_n, y) + F(x_n, y; W) - F(x_n, y_n; W)]_+ where Δ(yn,y)=1\Delta(y_n, y) = 1 if yn≠yy_n \neq y and 00 otherwise, and [z]+=max⁡(0,z)[z]_+ = \max(0, z).

    2. ALE: Optimizes a weighted approximate ranking objective emphasizing top rank violations: ∑y∈YtrlrΔ(xn,yn)rΔ(xn,yn)[Δ(yn,y)+F(xn,y;W)−F(xn,yn;W)]+\sum_{y \in \mathcal{Y}_{tr}} \frac{l_{r_\Delta(x_n, y_n)}}{r_\Delta(x_n, y_n)} [\Delta(y_n, y) + F(x_n, y; W) - F(x_n, y_n; W)]_+ where rΔ(xn,yn)=∑y∈Ytr1(F(xn,y;W)+Δ(yn,y)≥F(xn,yn;W))r_\Delta(x_n, y_n) = \sum_{y \in \mathcal{Y}_{tr}} \mathbf{1}(F(x_n, y; W) + \Delta(y_n, y) \ge F(x_n, y_n; W)) and lk=∑i=1k1il_k = \sum_{i=1}^k \frac{1}{i}.

    3. SJE: Uses the structured SVM loss focusing exclusively on the maximum violating class: [max⁡y∈Ytr(Δ(yn,y)+F(xn,y;W))−F(xn,yn;W)]+\left[ \max_{y \in \mathcal{Y}_{tr}} (\Delta(y_n, y) + F(x_n, y; W)) - F(x_n, y_n; W) \right]_+

    4. ESZSL: Optimizes a square loss with explicit Frobenius norm regularization: 1N∑n=1N∑y∈Ytr(F(xn,y;W)−1(y=yn))2+γ∥Wϕ(y)∥22+λ∥θ(x)TW∥22+β∥W∥F2\frac{1}{N} \sum_{n=1}^N \sum_{y \in \mathcal{Y}_{tr}} \left(F(x_n, y; W) - \mathbf{1}(y=y_n)\right)^2 + \gamma \|W \phi(y)\|_2^2 + \lambda \|\theta(x)^T W\|_2^2 + \beta \|W\|_F^2 which has a closed-form convex solution.

    5. SAE: Uses a linear autoencoder formulation enforcing reconstruction of visual embeddings: min⁡W∥θ(x)−WTϕ(y)∥2+λ∥Wθ(x)−ϕ(y)∥2\min_W \|\theta(x) - W^T \phi(y)\|^2 + \lambda \|W \theta(x) - \phi(y)\|^2

  5. Knowl 5 — Nonlinear, Hybrid, and Generative Zero-Shot Learning Formulations

    model/method

    Beyond linear bilinear mappings, zero-shot models learn non-linear associations, convex combinations of seen classifiers, or generative distributions:

    • LATEM (Nonlinear Piecewise Compatibility): Selects the maximum score across KK latent linear matrices {Wi}i=1K\{W_i\}_{i=1}^K: F(x,y;{Wi})=max⁡1≤i≤Kθ(x)TWiϕ(y)F(x, y; \{W_i\}) = \max_{1 \le i \le K} \theta(x)^T W_i \phi(y)

    • CMT (Nonlinear Neural Mapping): Maps visual features into semantic embedding space using a two-layer neural network with weights (W1,W2)(W_1, W_2): ∑y∈Ytr∑x∈Xy∥ϕ(y)−W1tanh⁡(W2θ(x))∥2\sum_{y \in \mathcal{Y}_{tr}} \sum_{x \in \mathcal{X}_y} \|\phi(y) - W_1 \tanh(W_2 \theta(x))\|^2

    • CONSE (Convex Combination of Seen Embeddings): Predicts the semantic representation of an unseen instance by taking the convex combination of semantic vectors s(⋅)s(\cdot) of the top TT most likely seen classes predicted by a softmax classifier ptr(y∣x)p_{tr}(y \mid x): 1Z∑i=1Tptr(f(x,i)∣x)⋅s(f(x,i)),Z=∑i=1Tptr(f(x,i)∣x)\frac{1}{Z} \sum_{i=1}^T p_{tr}(f(x, i) \mid x) \cdot s(f(x, i)), \quad Z = \sum_{i=1}^T p_{tr}(f(x, i) \mid x)

    • SYNC (Synthesized Classifiers): Synthesizes unseen class classifiers wcw_c as linear combinations of phantom class basis classifiers {vr}r=1R\{v_r\}_{r=1}^R via bipartite graph weights scrs_{cr}: min⁡wc∥wc−∑r=1Rscrvr∥22\min_{w_c} \left\| w_c - \sum_{r=1}^R s_{cr} v_r \right\|_2^2

    • GFZSL (Generative Gaussian Framework): Models each class conditional distribution as a multivariate Gaussian N(μy,diag(σy))\mathcal{N}(\mu_y, \text{diag}(\sigma_y)). Regressors fμf_\mu and fσf_\sigma map class semantic embeddings ϕ(y)\phi(y) to parameters μy=fμ(ϕ(y))\mu_y = f_\mu(\phi(y)) and σy=fσ(ϕ(y))\sigma_y = f_\sigma(\phi(y)). At test time, prediction is arg⁡max⁡yp(x∣σy,μy)\arg\max_y p(x \mid \sigma_y, \mu_y).

  6. Knowl 6 — Graph-Based Label Propagation Algorithm for Transductive Zero-Shot Learning

    algorithm

    In the transductive zero-shot learning setting, unlabeled visual features from unseen classes are accessible during training. Any compatibility model (such as ALE or DSRL) can be extended to the transductive setting by initializing class prediction scores and applying graph Laplacian label propagation.

    Input: Unlabeled test feature vectors Xts={x1,x2,…,xN}⊂RdX_{ts} = \{x_1, x_2, \dots, x_N\} \subset \mathbb{R}^d, learned linear mapping W∈Rd×aW \in \mathbb{R}^{d \times a}, unseen class attribute embeddings {ϕ(y)}y∈Yts⊂Ra\{\phi(y)\}_{y \in \mathcal{Y}_{ts}} \subset \mathbb{R}^a, nearest-neighbor size kk, RBF parameter σ>0\sigma > 0, regularization parameter α∈[0,1]\alpha \in [0, 1].
    Output: Predicted class labels {y^i}i=1N\{\hat{y}_i\}_{i=1}^N with y^i∈Yts\hat{y}_i \in \mathcal{Y}_{ts}.
    Initialize score matrix S0∈RN×∣Yts∣S_0 \in \mathbb{R}^{N \times |\mathcal{Y}_{ts}|}:
    for i=1i = 1 to NN do
        for each class y∈Ytsy \in \mathcal{Y}_{ts} do
            S0(i,y)←xiTWϕ(y)S_0(i, y) \leftarrow x_i^T W \phi(y)
        end for
    end for
    Construct affinity matrix M∈RN×NM \in \mathbb{R}^{N \times N} over mutual kk-nearest neighbors:
    for i=1i = 1 to NN do
        for j=1j = 1 to NN do
            if i∈KNN(j)i \in \text{KNN}(j) or j∈KNN(i)j \in \text{KNN}(i) then
                Mij←exp⁡(−∥xi−xj∥222σ2)M_{ij} \leftarrow \exp\left( -\frac{\|x_i - x_j\|_2^2}{2\sigma^2} \right)
            else
                Mij←0M_{ij} \leftarrow 0
            end if
        end for
    end for
    Compute diagonal degree matrix Q∈RN×NQ \in \mathbb{R}^{N \times N} where Qii←∑j=1NMijQ_{ii} \leftarrow \sum_{j=1}^N M_{ij}
    Compute normalized Laplacian L←Q−1/2MQ−1/2L \leftarrow Q^{-1/2} M Q^{-1/2}
    Compute smoothed score matrix via closed-form label propagation:
    S←(I−αL)−1S0S \leftarrow (I - \alpha L)^{-1} S_0
    for i=1i = 1 to NN do
        y^i←arg⁡max⁡y∈YtsS(i,y)\hat{y}_i \leftarrow \arg\max_{y \in \mathcal{Y}_{ts}} S(i, y)
    end for
    return {y^i}i=1N\{\hat{y}_i\}_{i=1}^N
  7. Knowl 7 — Zero-Shot Learning Benchmark Performance on Attribute Datasets (SS vs PS)

    data/table

    Zero-shot learning Top-1 per-class accuracy (%) across SUN, CUB, AWA1, AWA2, and aPY datasets using 2048-dimensional ResNet-101 image features, comparing Standard Splits (SS) and Proposed Splits (PS):

    SUN CUB AWA1 AWA2 aPY
    Method SS PS SS PS SS PS SS PS SS PS
    DAP 38.9 39.9 37.5 40.0 57.1 44.1 58.7 46.1 35.2 33.8
    IAP 17.4 19.4 27.1 24.0 48.1 35.9 46.9 35.9 22.4 36.6
    CONSE 44.2 38.0 36.7 33.6 63.6 46.3 67.9 44.6 25.9 26.4
    CMT 41.9 40.1 37.3 34.6 58.9 39.5 66.3 37.9 26.9 28.0
    SSE 54.5 51.5 43.7 43.9 68.8 60.1 67.5 61.0 31.1 35.0
    LATEM 56.9 55.3 49.4 49.6 74.8 55.1 68.7 55.8 34.5 36.8
    ALE 59.1 58.1 53.2 54.9 78.6 59.9 80.3 62.5 30.9 39.7
    DEVISE 57.5 56.5 53.2 52.0 72.9 54.2 68.6 59.7 35.4 37.0
    SJE 57.1 52.7 55.3 53.9 76.7 65.6 69.5 61.9 32.0 31.7
    ESZSL 57.3 54.5 55.1 51.9 74.7 58.2 75.6 58.6 34.4 38.3
    SYNC 59.1 56.2 54.1 56.0 72.2 51.8 71.2 49.3 39.7 23.9
    SAE 42.4 40.3 33.4 33.3 80.6 53.0 80.7 54.1 8.3 8.3
    GFZSL 62.9 60.6 53.0 49.3 80.5 68.2 79.3 63.8 51.3 38.4

    When overlapping ImageNet 1K classes are eliminated under PS, accuracies drop significantly on AWA1, AWA2, and aPY (e.g. SAE drops by 27.6% on AWA1 and 26.6% on AWA2). In contrast, on fine-grained datasets (SUN, CUB) where test classes did not heavily intersect ImageNet 1K, accuracies between SS and PS remain close. On PS, bilinear compatibility methods (ALE, SJE, DEVISE) and generative GFZSL consistently outperform two-stage intermediate classifier approaches (DAP, IAP).

  8. Knowl 8 — Generalized Zero-Shot Learning (GZSL) Performance on Proposed Splits

    data/table

    Generalized Zero-Shot Learning top-1 per-class accuracy (%) on unseen test classes (tsts), seen training classes (trtr), and their harmonic mean (HH) across datasets under the Proposed Splits (PS) using ResNet-101 features:

    SUN CUB AWA1 AWA2 aPY
    Method ts tr H ts tr H ts tr H ts tr H ts tr H
    DAP 4.2 25.1 7.2 1.7 67.9 3.3 0.0 88.7 0.0 0.0 84.7 0.0 4.8 78.3 9.0
    IAP 1.0 37.8 1.8 0.2 72.8 0.4 2.1 78.2 4.1 0.9 87.6 1.8 5.7 65.6 10.4
    CONSE 6.8 35.9 11.4 2.0 70.6 3.9 0.4 89.6 0.8 0.5 90.6 1.0 0.0 91.2 0.0
    CMT 8.1 21.8 11.8 7.2 49.8 12.6 0.9 87.6 1.8 0.5 90.0 1.0 1.4 85.2 2.8
    CMT* 8.7 28.0 13.3 4.7 60.1 8.7 8.4 86.9 15.3 8.7 89.0 15.9 10.9 74.2 19.0
    SSE 2.1 36.4 4.0 8.5 46.9 14.4 7.0 80.5 12.9 8.1 82.5 14.8 0.3 78.4 0.6
    LATEM 14.7 28.8 19.5 15.2 57.3 24.0 7.3 71.7 13.3 11.5 77.3 20.0 1.3 71.4 2.6
    ALE 21.8 33.1 26.3 23.7 62.8 34.4 16.8 76.1 27.5 14.0 81.8 23.9 4.6 73.7 8.7
    DEVISE 16.9 27.4 20.9 23.8 53.0 32.8 13.4 68.7 22.4 17.1 74.7 27.8 3.5 78.4 6.7
    SJE 14.4 29.7 19.4 23.5 59.2 33.6 11.3 74.6 19.6 8.0 73.9 14.4 1.3 71.4 2.6
    ESZSL 11.0 27.9 15.8 14.7 56.5 23.3 6.6 75.6 12.1 5.9 77.8 11.0 2.4 70.1 4.6
    SYNC 7.9 43.3 13.4 11.5 70.9 19.8 9.0 88.9 16.3 9.7 89.7 17.5 7.4 66.3 13.3
    SAE 8.8 18.0 11.8 7.8 54.0 13.6 1.8 77.1 3.5 1.1 82.2 2.2 0.4 80.9 0.9
    GFZSL 0.0 39.6 0.0 0.0 45.7 0.0 1.8 80.3 3.5 2.5 80.1 4.8 0.0 83.3 0.0

    In GZSL, seen classes act as distractors, causing a severe drop in unseen class accuracy across all methods compared to pure ZSL. Two-stage and hybrid models (DAP, IAP, CONSE) achieve high seen-class accuracy (>70%>70\%) but near-zero unseen-class accuracy (<5%<5\%), resulting in poor harmonic mean scores. Linear compatibility methods (ALE, DEVISE, SJE) attain the highest harmonic means across datasets (ALE achieves top HH on SUN with 26.3%, CUB with 34.4%, and AWA1 with 27.5%; DEVISE achieves top HH on AWA2 with 27.8%). CMT* with novelty detection leads on aPY (H=19.0%H = 19.0\%).

  9. Knowl 9 — Large-Scale Zero-Shot Learning Performance on ImageNet Splits

    data/table

    Zero-shot Top-1 per-class accuracy (%) on ImageNet using ResNet-101 features and Word2Vec class embeddings. Models are trained on the standard 1,000 ImageNet classes and evaluated on:

    • Semantic hierarchy splits: classes 2 hops (1,509 classes) and 3 hops (7,678 classes) away in WordNet.
    • Most populated splits: top 500, 1K, and 5K most populated unseen classes.
    • Least populated splits: bottom 500, 1K, and 5K least populated unseen classes.
    • All: all remaining ~20,000 unseen classes of ImageNet.
    Hierarchy Most Populated Least Populated All
    Method 2H 3H 500 1K 5K 500 1K 5K 20K
    CONSE 7.63% 2.18% 12.33% 8.31% 3.22% 3.53% 2.69% 1.05% 0.95%
    CMT 2.88% 0.67% 5.10% 3.04% 1.04% 1.87% 1.08% 0.33% 0.29%
    LATEM 5.45% 1.32% 10.81% 6.63% 1.90% 4.53% 2.74% 0.76% 0.50%
    ALE 5.38% 1.32% 10.40% 6.77% 2.00% 4.27% 2.85% 0.79% 0.50%
    DEVISE 5.25% 1.29% 10.36% 6.68% 1.94% 4.23% 2.86% 0.78% 0.49%
    SJE 5.31% 1.33% 9.88% 6.53% 1.99% 4.93% 2.93% 0.78% 0.52%
    ESZSL 6.35% 1.51% 11.91% 7.69% 2.34% 4.50% 3.23% 0.94% 0.62%
    SYNC 9.26% 2.29% 15.83% 10.75% 3.42% 5.83% 3.52% 1.26% 0.96%
    SAE 4.89% 1.26% 9.96% 6.57% 2.09% 2.50% 2.17% 0.72% 0.56%
    GFZSL 1.45% – 2.01% 1.35% – 1.40% 1.11% 0.13% –

    SYNC outperforms all other methods across all evaluated splits on ImageNet (e.g. 9.26% on 2H, 15.83% on M500, 0.96% on All 20K), followed by ESZSL. Performance on populated classes is consistently higher than on sparse, fine-grained, least populated classes (e.g. 15.83% for SYNC on M500 vs 5.83% on L500). Generative GFZSL drops sharply on ImageNet, showing that generative density models struggle when transitioning from explicit attributes to noisy Word2Vec embeddings. On the full 20K test set, top-1 accuracy for all methods remains below 1.0%.

  10. Knowl 10 — Divergent Impact of Transductive Learning in ZSL versus GZSL

    empirical result

    Transductive zero-shot learning leverages unlabeled images from test classes to refine representations (evaluated across GFZSL-tran, DSRL, and ALE-tran under the Proposed Splits):

    • In pure Zero-Shot Learning (ZSL), transductive learning provides consistent accuracy gains over inductive models across datasets. On AWA2, GFZSL-tran achieves 78.6% top-1 accuracy compared to 63.8% for inductive GFZSL. On aPY, ALE-tran reaches 45.5% compared to 37.1% for inductive ALE. GFZSL-tran yields the highest transductive ZSL performance on SUN, AWA1, and AWA2, while ALE-tran leads on CUB and aPY.
    • In Generalized Zero-Shot Learning (GZSL), transductive learning fails to produce consistent improvements. For ALE-tran, harmonic mean performance does not improve over inductive ALE on any of the evaluated datasets. While GFZSL-tran improves GZSL harmonic mean on AWA1 and AWA2, GFZSL models collapse to near-zero harmonic means on SUN, CUB, and aPY in both inductive and transductive modes.

    Thus, access to unlabeled unseen class data aids discrimination among unseen classes when the search space is restricted to YtsY_{ts}, but does not solve the distractor bias toward seen classes when the search space includes Ytr∪YtsY_{tr} \cup Y_{ts}.

Coverage note — None was omitted.

References

  1. 1.C. Lampert, H. Nickisch, and S. Harmeling, “Attribute-based classification for zero-shot visual object categorization,” in TPAMI, 2013. 1, 2, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13
  2. 2.H. Larochelle, D. Erhan, and Y. Bengio, “Zero-data learning of new tasks,” in AAAI, 2008. 1
  3. 3.M. Rohrbach, M. Stark, and B.Schiele, “Evaluating knowledge transfer and zero-shot learning in a large-scale setting,” in CVPR, 2011. 1, 3, 6
  4. 4.X. Yu and Y. Aloimonos, “Attribute-based transfer learning for object categorization with zero or one training example,” in ECCV, 2010. 1
  5. 5.X. Xu, Y. Yang, D. Zhang, H. T. Shen, and J. Song, “Matrix trifactorization with manifold regularizations for zero-shot learning,” in CVPR, 2017. 1
  6. 6.Z. Ding, M. Shao, and Y. Fu, “Low-rank embedded ensemble semantic dictionary for zero-shot learning,” in CVPR, 2017. 1
  7. 7.A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. A. Ranzato, and T. Mikolov, “Devise: A deep visual-semantic embedding model,” in NIPS, 2013, pp. 2121–2129. 1, 2, 3, 8, 9, 10, 11, 12, 13
  8. 8.Z. Akata, F. Perronnin, Z. Harchaoui, and C. Schmid, “Label embedding for attribute-based classification,” in CVPR, 2013. 1, 6
  9. 9.Z. Akata, S. Reed, D. Walter, H. Lee, and B. Schiele, “Evaluation of output embeddings for fine-grained image classification,” in CVPR, 2015. 1, 2, 3, 4, 6, 7, 8, 9, 10, 11, 12, 13
  10. 10.B. Romera-Paredes and P. H. Torr, “An embarrassingly simple approach to zero-shot learning,” ICML, 2015. 1, 2, 3, 4, 7, 8, 9, 10, 11, 12, 13
  11. 11.Y. Xian, Z. Akata, G. Sharma, Q. Nguyen, M. Hein, and B. Schiele, “Latent embeddings for zero-shot classification,” in CVPR, 2016. 1, 2, 3, 4, 7, 8, 9, 10, 11, 12, 13
  12. 12.R. Socher, M. Ganjoo, C. D. Manning, and A. Ng, “Zero-shot learning through cross-modal transfer,” in NIPS, 2013. 1, 2, 3, 4, 8, 9, 10, 11, 12, 13
  13. 13.Z. Zhang and V. Saligrama, “Zero-shot learning via semantic similarity embedding,” in ICCV, 2015. 1, 2, 4, 7, 8, 9, 10, 11, 12, 13
  14. 14.S. Changpinyo, W.-L. Chao, B. Gong, and F. Sha, “Synthesized classifiers for zero-shot learning,” in CVPR, 2016. 1, 2, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13
  15. 15.M. Norouzi, T. Mikolov, S. Bengio, Y. Singer, J. Shlens, A. Frome, G. Corrado, and J. Dean, “Zero-shot learning by convex combination of semantic embeddings,” in ICLR, 2014. 1, 2, 4, 8, 9, 10, 11, 12, 13
  16. 16.G. Patterson and J. Hays, “Sun attribute database: Discovering, annotating, and recognizing scene attributes,” in CVPR, 2012. 1, 5, 7
  17. 17.P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona, “Caltech-UCSD Birds 200,” Caltech, Tech. Rep. CNS-TR-2010-001, 2010. 1, 5, 7
  18. 18.A. Farhadi, I. Endres, D. Hoiem, and D. Forsyth, “Describing objects by their attributes,” CVPR, 2009. 1, 2, 5, 6, 7
  19. 19.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in CVPR, 2009. 1, 5, 6
  20. 20.D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” IJCV, vol. 60, no. 2, pp. 91–110, 2004. 1
  21. 21.J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell, “Decaf: A deep convolutional activation feature for generic visual recognition.” in ICML, 2014. 1
  22. 22.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014. 1
  23. 23.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016. 1, 6
  24. 24.Z. Al-Halah, M. Tapaswi, and R. Stiefelhagen, “Recovering the missing link: Predicting class-attribute associations for unsupervised zero-shot learning,” in CVPR, 2016. 2
  25. 25.D. Jayaraman and K. Grauman, “Zero-shot recognition with unreliable attributes,” in NIPS, 2014. 2
  26. 26.P. Kankuekul, A. Kawewong, S. Tangruamsub, and O. Hasegawa, “Online incremental attribute-based zero-shot learning,” in CVPR, 2012. 2
  27. 27.T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in NIPS, 2013. 2, 3, 6
  28. 28.Y. Fu, T. M. Hospedales, T. Xiang, and S. Gong, “Transductive multiview zero-shot learning,” TPAMI, 2015. 2, 3
  29. 29.M. Palatucci, D. Pomerleau, G. E. Hinton, and T. M. Mitchell, “Zero-shot learning with semantic output codes,” in NIPS, 2009. 2
  30. 30.Z. Akata, F. Perronnin, Z. Harchaoui, and C. Schmid, “Label-embedding for image classification,” TPAMI, 2016. 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13
  31. 31.R. Qiao, L. Liu, C. Shen, and A. van den Hengel, “Less is more: Zeroshot learning from online textual documents with noise suppression,” in CVPR, 2016. 2, 3
  32. 32.M. Bucher, S. Herbin, and F. Jurie, “Improving semantic embedding consistency by metric learning for zero-shot classiffication,” in ECCV, 2016, pp. 730–746. 2
  33. 33.E. Kodirov, T. Xiang, and S. Gong, “Semantic autoencoder for zero-shot learning,” in CVPR, 2017. 2, 3, 4, 7, 8, 9, 10, 11, 12, 13
  34. 34.J. Lei Ba, K. Swersky, S. Fidler et al., “Predicting deep zero-shot convolutional neural networks using textual descriptions,” in ICCV, 2015. 2, 3
  35. 35.L. Zhang, T. Xiang, and S. Gong, “Learning a deep embedding model for zero-shot learning,” in CVPR, 2017. 2
  36. 36.S. Changpinyo, W.-L. Chao, and F. Sha, “Predicting visual exemplars of unseen classes for zero-shot learning,” ICCV, pp. 3496–3505, 2017. 2
  37. 37.Z. Zhang and V. Saligrama, “Zero-shot learning via joint semantic similarity embedding,” in CVPR, 2016. 2
  38. 38.Z. Fu, T. Xiang, E. Kodirov, and S. Gong, “Zero-shot object recognition by semantic manifold distance,” in CVPR, June 2015. 2
  39. 39.Z. Akata, M. Malinowski, M. Fritz, and B. Schiele, “Multi-cue zero-shot learning with strong supervision,” in CVPR, 2016. 2
  40. 40.Y. Long, L. Liu, L. Shao, F. Shen, G. Ding, and J. Han, “From zero-shot learning to conventional supervised classification: Unseen visual data synthesis,” in CVPR, 2017. 2
  41. 41.V. K. Verm and P. Rai, “A simple exponential family framework for zeroshot learning,” in ECML, 2017, pp. 792–808. 2, 5, 7, 8, 9, 10, 11, 12, 13
  42. 42.Y. Li and D. Wang, “Zero-shot learning with generative latent prototype model,” arXiv preprint arXiv:1705.09474, 2017. 2
  43. 43.T. Mukherjee and T. Hospedales, “Gaussian visual-linguistic embedding for zero-shot recognition,” in EMNLP, 2016. 2
  44. 44.M. Rohrbach, S. Ebert, and B. Schiele, “Transfer learning in a transductive setting,” in NIPS, 2013. 3
  45. 45.E. Kodirov, T. Xiang, Z. Fu, and S. Gong, “Unsupervised domain adaptation for zero-shot learning,” in ICCV, 2015. 3
  46. 46.X. Li, Y. Guo, and D. Schuurmans, “Semi-supervised zero-shot classification with label representation learning,” in ICCV, 2015. 3
  47. 47.X. Li and Y. Guo, “Max-margin zero-shot learning for multi-class classification.” in AISTATS, 2015. 3
  48. 48.Y. Fu and L. Sigal, “Semi-supervised vocabulary-informed learning,” in CVPR, 2016. 3
  49. 49.Z. Zhang and V. Saligrama, “Zero-shot recognition via structured prediction,” in CVPR, 2016. 3
  50. 50.T. Mensink, E. Gavves, and C. G. Snoek, “Costa: Co-occurrence statistics for zero-shot classification,” in CVPR, 2014. 3
  51. 51.M. Rohrbach, M. Stark, G. Szarvas, I. Gurevych, and B. Schiele, “What helps where–and why? semantic relatedness for knowledge transfer,” in CVPR, 2010. 3
  52. 52.S. Reed, Z. Akata, H. Lee, and B. Schiele, “Learning deep representations of fine-grained visual descriptions,” in CVPR, 2016. 3
  53. 53.M. Elhoseiny, B. Saleh, and A. Elgammal, “Write a classifier: Zero-shot learning using purely textual descriptions,” in ICCV, 2013. 3
  54. 54.S. Antol, C. L. Zitnick, and D. Parikh, “Zero-shot learning via visual abstraction,” in European Conference on Computer Vision. Springer, 2014, pp. 401–416. 3
  55. 55.T. Mensink, J. Verbeek, F. Perronnin, and G. Csurka, “Metric learning for large scale image classification: Generalizing to new classes at near-zero cost,” Computer Vision–ECCV 2012, pp. 488–501, 2012. 3
  56. 56.J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in EMNLP, 2014, pp. 1532–1543. 3
  57. 57.G. A. Miller, “Wordnet: a lexical database for english,” CACM, vol. 38, pp. 39–41, 1995. [Online]. Available: http://doi.acm.org/10.1145/219717.219748 3, 6
  58. 58.N. Karessli, Z. Akata, B. Schiele, and A. Bulling, “Gaze embeddings for zero-shot image classification,” in IEEE Computer Vision and Pattern Recognition (CVPR), 2017. 3
  59. 59.W. J. Scheirer, A. Rocha, A. Sapkota, and T. E. Boult, “Towards open set recognition,” TPAMI, vol. 36, 2013. 3
  60. 60.L. Jain, W. Scheirer, and T. Boult, “Multi-class open set recognition using probability of inclusion,” in ECCV, 2014. 3
  61. 61.H. Zhang, X. Shang, W. Yang, H. Xu, H. Luan, and T.-S. Chua, “Online collaborative learning for open-vocabulary visual classifiers,” in CVPR, 2016. 3
  62. 62.A. Bendale and T. E. Boult, “Towards open set deep networks,” in CVPR, 2016. 3
  63. 63.W.-L. Chao, S. Changpinyo, B. Gong, and F. Sha, “An empirical study and analysis of generalized zero-shot learning for object recognition in the wild,” in ECCV, 2016. 3
  64. 64.T. Joachims, “Optimizing search engines using clickthrough data,” in KDD. ACM, 2002. 3
  65. 65.N. Usunier, D. Buffoni, and P. Gallinari, “Ranking with ordered weighted pairwise classification,” in ICML, 2009. 3
  66. 66.J. Weston, S. Bengio, and N. Usunier, “Wsabie: Scaling up to large vocabulary image annotation,” in IJCAI, 2011. 3
  67. 67.I. Tsochantaridis, T. Joachims, T. Hofmann, and Y. Altun, “Large margin methods for structured and interdependent output variables,” JMLR, 2005. 4
  68. 68.R. H. Bartels and G. Stewart, “Solution of the matrix equation ax+ xb= c [f4],” Commun. ACM, vol. 15, no. 9, pp. 820–826, 1972. 4
  69. 69.O. Chapelle, B. Scholkopf, and A. Zien, “Semi-supervised learning,” IEEE Transactions on Neural Networks, vol. 20, no. 3, pp. 542–542, 2009. 5
  70. 70.D. Zhou, O. Bousquet, T. N. Lal, J. Weston, and B. Scholkopf, “Learning ¨with local and global consistency,” in NIPS, 2004. 5
  71. 71.M. Ye and Y. Guo, “Zero-shot classification with discriminative semantic representation learning,” in CVPR, 2017. 5, 7, 8, 12, 13
  72. 72.Y. Fujiwara and G. Irie, “Efficient label propagation,” in ICML, 2014, pp. 784–792. 5
  73. 73.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR, 2015. 6
  74. 74.B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, “Learning deep features for scene recognition using places database,” in NIPS, 2014. 8
  75. 75.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” CVPR, 2015. 8
  76. 76.S. Garcia and F. Herrera, “An extension on“statistical comparisons of classifiers over multiple data sets”for all pairwise comparisons,” JLMR, vol. 9, pp. 2677–2694, 2008. 9

Citation

MLA
Xian, Y., et al. “Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 9, 2019, pp. 2251–65, https://doi.org/10.1109/TPAMI.2018.2857768.
APA
Xian, Y., Lampert, C. H., Schiele, B., & Akata, Z. (2019). Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(9), 2251–2265. https://doi.org/10.1109/TPAMI.2018.2857768
Chicago
Xian, Y., C. H. Lampert, B. Schiele, and Z. Akata. 2019. “Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly”. IEEE Transactions on Pattern Analysis and Machine Intelligence 41 (9): 2251–65. https://doi.org/10.1109/TPAMI.2018.2857768.
Harvard
Xian, Y. et al. (2019) “Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(9), pp. 2251–2265. Available at: https://doi.org/10.1109/TPAMI.2018.2857768.
Vancouver
1. Xian Y, Lampert CH, Schiele B, Akata Z (2019) Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly. IEEE Transactions on Pattern Analysis and Machine Intelligence 41:2251–2265

BibTeX

@article{Xian_2019, title={Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly}, volume={41}, ISSN={1939-3539}, url={http://dx.doi.org/10.1109/TPAMI.2018.2857768}, DOI={10.1109/tpami.2018.2857768}, number={9}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Xian, Yongqin and Lampert, Christoph H. and Schiele, Bernt and Akata, Zeynep}, year={2019}, month=Sept, pages={2251–2265} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF