An embarrassingly simple approach to zero-shot learning

Bernardino Romera-ParedesPhilip H. S. Torr

article2015ICML1,354 citations

Proposes a two-layer linear framework for zero-shot learning that is implementable in a single line of code, establishes theoretical generalization bounds through domain adaptation, and surpasses complex state-of-the-art methods across standard benchmarks by up to 17%.

Listen

Modern automated classification systems struggle when encountering new categories after the initial training phase, a frequent issue in practical applications where the number of categories continuously expands or collecting examples for every specific class is impractical. Zero-shot learning addresses this by identifying unseen categories based solely on high-level attribute descriptions that link them to previously seen concepts. However, existing methods often separate attribute training from the final category prediction step, rely on unrealistic independence assumptions, require complex optimization procedures, or fail to handle the unreliability of predicted attributes effectively.

The main objective of the article is to formulate, theoretically analyze, and empirically evaluate an extremely simple, computationally efficient linear model for zero-shot learning that directly links features, attributes, and categories in a unified optimization problem. The article also aims to provide mathematical guarantees on generalisation error by formalizing zero-shot learning as a domain adaptation problem.

To achieve this, the article establishes a two-layer linear framework where the first layer maps raw data features into an intermediate semantic attribute space, while the second layer utilizes provided class attribute descriptions that are interchangeable during evaluation. The authors pair a standard multiclass loss function with a tailored regularization strategy that controls representation scales and bounds variance across domains. This mathematical formulation results in an exact, closed-form solution that requires minimal code and requires no iterative solvers. The authors evaluated this method on synthetic benchmarks and three standard real-world image datasets covering animals, scenes, and diverse objects.

The experimental findings show that this streamlined method consistently outperforms state-of-the-art benchmarks across the evaluated datasets. On the scene recognition benchmark, the model achieved an accuracy of 65.75%, delivering an improvement ratio of approximately 17% over prior leading approaches. On the animal categorization benchmark, the method reached 49.30% accuracy, representing an improvement of over 14% compared to existing techniques, while executing training in roughly four seconds compared to more than eleven hours for a prior related model. On object recognition with limited training classes, a variant utilizing individual instance signatures achieved 27.27% accuracy, outperforming the prior state-of-the-art result of 26.02%. Synthetic tests further demonstrated that the model remains robust against uninformative or corrupted attributes and scales effectively as the number of training classes grows.

These findings demonstrate that highly complex, multi-stage architectures are not required to achieve superior accuracy in zero-shot classification. By integrating attribute learning and class prediction into a single objective with principled regularization, organizations can significantly cut computational and engineering costs while achieving faster deployment cycles. The dramatic reduction in training time—from several hours to mere seconds—enables rapid model iteration and adaptation across large-scale classification pipelines without demanding specialized infrastructure.

Organizations handling dynamically evolving category sets should adopt this simple linear model as a high-performing baseline for zero-shot recognition tasks. When deploying to datasets where attribute signatures are defined at the individual instance level rather than the class level, teams should select the instance-signature variant to avoid performance drops when class diversity is low. Future work should focus on extending this framework to incorporate deep non-linear layers and testing compatibility with learned word embeddings rather than manually curated visual attributes.

The primary limitation of the method is its sensitivity to the ratio of classes to attributes in its standard form, which can degrade performance if the number of training classes is significantly smaller than the attribute count. Additionally, theoretical error bounds assume that the source and target feature distributions remain reasonably aligned and that unseen classes share measurable attribute overlap with training classes. The reported results provide strong confidence for practical adoption in visual recognition domains where high-quality attribute descriptions are available.

  • Paper: DeViSE: A Deep Visual-Semantic Embedding Model, Andrea Frome et al. (2013). DeViSE establishes visual-semantic embeddings mapping images into semantic label spaces, providing the conceptual foundation for ESZSL's bilinear feature-attribute-class framework.
  • Paper: Describing Objects by their Attributes, Ali Farhadi et al. (2009). This seminal work introduces attribute-based visual recognition and standard attribute datasets (e.g., Animals with Attributes) that serve as the testbeds for zero-shot learning methods.
  • Paper: CNN Features Off-the-Shelf: An Astounding Baseline for Recognition, Ali Sharif Razavian et al. (2014). This paper demonstrates that off-the-shelf deep CNN activations serve as powerful fixed representations, which ESZSL directly adopts as input image features.
  • Paper: Adapting Visual Category Models to New Domains, Kate Saenko et al. (2010). Saenko et al. formulate domain adaptation using learned regularized metric transforms, motivating ESZSL's theoretical characterization of zero-shot learning as domain adaptation with generalization error bounds.
Cover for An embarrassingly simple approach to zero-shot learning

Abstract

Zero-shot learning consists in learning how to recognise new concepts by just having a description of them. Many sophisticated approaches have been proposed to address the challenges this problem comprises. In this paper we describe a zero-shot learning approach that can be implemented in just one line of code, yet it is able to outperform state of the art approaches on standard datasets. The approach is based on a more general framework which models the relationships between features, attributes, and classes as a two linear layers network, where the weights of the top layer are not learned but are given by the environment. We further provide a learning bound on the generalisation error of this kind of approaches, by casting them as domain adaptation methods. In experiments carried out on three standard real datasets, we found that our approach is able to perform significantly better than the state of art on all of them, obtaining a ratio of improvement up to 17%.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Embarrassingly simple ZSL
  • 3.1. Regularisation and loss function choices
  • 4. Risk bounds
  • 4.1. Simple ZSL as a domain adaptation problem
  • 4.2. Risk bounds for domain adaptation
  • 5. Experiments
  • 5.1. Synthetic experiments
  • 5.2. Real data experiments
  • 6. Discussion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Embarrassingly Simple Zero-Shot Learning Optimization and Closed-Form Solution

    model/method

    In zero-shot learning (ZSL), let X∈Rd×mX \in \mathbb{R}^{d \times m} denote the matrix of mm training instances of dimension dd, Y∈{−1,1}m×zY \in \{-1, 1\}^{m \times z} denote the binary ground-truth class indicator matrix over zz training classes, and S∈[0,1]a×zS \in [0, 1]^{a \times z} denote the matrix containing the aa-dimensional continuous or binary attribute signatures for each of the zz training classes.

    Rather than learning a direct linear classifier matrix W∈Rd×zW \in \mathbb{R}^{d \times z} via standard empirical risk minimization min⁡WL(X⊤W,Y)+Ω(W)\min_{W} L(X^\top W, Y) + \Omega(W), ESZSL decomposes the classifier as W=VS⊤W = V S^\top, where V∈Rd×aV \in \mathbb{R}^{d \times a} maps feature vectors to the intermediate attribute space. The optimization problem is:

    min⁡V∈Rd×aL(X⊤VS,Y)+Ω(V;S,X)\min_{V \in \mathbb{R}^{d \times a}} L(X^\top V S, Y) + \Omega(V; S, X)

    where L(P,Y)=∥P−Y∥Fro2L(P, Y) = \|P - Y\|_{\text{Fro}}^2 is the squared Frobenius norm loss. The regularizer Ω(V;S,X)\Omega(V; S, X) is designed to control both the norm of attribute signatures projected into the feature space (VSV S) and the variance of data instances projected into the attribute space (X⊤VX^\top V), combined with weight decay:

    Ω(V;S,X)=γ∥VS∥Fro2+λ∥X⊤V∥Fro2+β∥V∥Fro2\Omega(V; S, X) = \gamma \|V S\|_{\text{Fro}}^2 + \lambda \|X^\top V\|_{\text{Fro}}^2 + \beta \|V\|_{\text{Fro}}^2

    where γ,λ,β>0\gamma, \lambda, \beta > 0 are regularization hyperparameters. By setting β=γλ\beta = \gamma \lambda, the objective function is strictly convex and yields the exact closed-form solution:

    V=(XX⊤+γId)−1XYS⊤(SS⊤+λIa)−1V = (X X^\top + \gamma I_d)^{-1} X Y S^\top (S S^\top + \lambda I_a)^{-1}

    where Id∈Rd×dI_d \in \mathbb{R}^{d \times d} and Ia∈Ra×aI_a \in \mathbb{R}^{a \times a} are identity matrices.

  2. Knowl 2 — Zero-Shot Inference via Linear Attribute Compatibility

    model/method

    Given a learned feature-to-attribute mapping matrix V∈Rd×aV \in \mathbb{R}^{d \times a} trained on source classes, zero-shot classification at inference time predicts the class of a new test instance x∈Rdx \in \mathbb{R}^d among a set of z′z' unseen test classes. Each test class i∈{1,…,z′}i \in \{1, \dots, z'\} is defined by its semantic attribute signature Si′∈[0,1]aS'_i \in [0, 1]^a, arranged into a test attribute matrix S′∈[0,1]a×z′S' \in [0, 1]^{a \times z'}.

    The linear model parameter for unseen classes is constructed as W′=VS′∈Rd×z′W' = V S' \in \mathbb{R}^{d \times z'}. Classification of instance xx is performed by selecting the class signature that maximizes the linear compatibility score:

    y^=argmax⁡i∈{1,…,z′}x⊤VSi′\hat{y} = \operatorname{argmax}_{i \in \{1, \dots, z'\}} x^\top V S'_i

    where Si′S'_i is the ii-th column of S′S'. This inference operation requires only a single matrix-vector multiplication x⊤(VS′)x^\top (V S').

  3. Knowl 3 — Zero-Shot Generalization Error Bound via Domain Adaptation

    theoretical result

    Zero-shot learning can be cast as a domain adaptation problem where the source task classifies instances using training attribute signatures S∈[0,1]a×zS \in [0, 1]^{a \times z} and the target task classifies instances using test attribute signatures S′∈[0,1]a×z′S' \in [0, 1]^{a \times z'}. By mapping instance-class pairs (xi,st)(x_i, s_t) to x~t,i=vec⁡(xist⊤)∈Rda\tilde{x}_{t,i} = \operatorname{vec}(x_i s_t^\top) \in \mathbb{R}^{da} and model parameters to v=vec⁡(V)∈Rdav = \operatorname{vec}(V) \in \mathbb{R}^{da}, source instances are drawn from a distribution D\mathcal{D} and target instances from D′\mathcal{D}'.

    Let H\mathcal{H} be the hypothesis space of linear classifiers on Rda\mathbb{R}^{da} with VC-dimension dˉ=da+1\bar{d} = da + 1, and let mˉ=mz\bar{m} = mz be the effective number of training instance-class pairs. For any hypothesis h∈Hh \in \mathcal{H}, with probability at least 1−δ1 - \delta over the choice of sample sets U∼DmˉU \sim \mathcal{D}^{\bar{m}} and U′∼(D′)mˉU' \sim (\mathcal{D}')^{\bar{m}}, the target domain error ϵ′(h)\epsilon'(h) is bounded by:

    ϵ′(h)≤ϵ(h)+42dˉmˉ(log⁡2mˉdˉ+log⁡4δ)+λ∗+12d^HΔH(U,U′)\epsilon'(h) \le \epsilon(h) + 4 \sqrt{\frac{2\bar{d}}{\bar{m}} \left(\log \frac{2\bar{m}}{\bar{d}} + \log \frac{4}{\delta}\right)} + \lambda^* + \frac{1}{2} \hat{d}_{\mathcal{H}\Delta\mathcal{H}}(U, U')

    where ϵ(h)\epsilon(h) is the source domain error, λ∗=inf⁡h∈H[ϵ(h)+ϵ′(h)]\lambda^* = \inf_{h \in \mathcal{H}} [\epsilon(h) + \epsilon'(h)] measures the combined error of the ideal joint hypothesis (with λ∗=0\lambda^* = 0 if the ground truth labeling is realizable in H\mathcal{H}), and d^HΔH(U,U′)\hat{d}_{\mathcal{H}\Delta\mathcal{H}}(U, U') is the empirical A\mathcal{A}-distance over the symmetric difference space HΔH={h(x)⊕h′(x):h,h′∈H}\mathcal{H}\Delta\mathcal{H} = \{h(x) \oplus h'(x) : h, h' \in \mathcal{H}\}.

    Boundary behavior of this bound includes:

    1. Identical Signatures (S=S′S = S'): The domain distance dHΔH(D,D′)=0d_{\mathcal{H}\Delta\mathcal{H}}(\mathcal{D}, \mathcal{D}') = 0, reducing the bound to the standard Vapnik-Chervonenkis generalization bound.
    2. Orthogonal Signatures: If every training class signature is orthogonal to every test class signature (∥⟨si,sj′⟩=0∥\|\langle s_i, s'_j \rangle = 0\| for all i,ji, j), the right-hand side term exceeds 1, rendering the bound vacuous and reflecting the impossibility of cross-class transfer without shared attribute structure.
  4. Knowl 4 — ESZSL-AS Formulation for Instance-Level Attribute Datasets

    model/method

    When datasets provide attribute annotations per individual instance rather than solely per class, the ESZSL All Signatures (ESZSL-AS) variant treats each training instance's attribute signature as a distinct class in its own right.

    Let X∈Rd×mX \in \mathbb{R}^{d \times m} be the matrix of mm training instances and S∈Ra×mS \in \mathbb{R}^{a \times m} be the matrix containing the aa-dimensional attribute signature of each instance. In this setting, the ground-truth label matrix is replaced by the identity matrix Im∈Rm×mI_m \in \mathbb{R}^{m \times m} (effectively omitting YY). The closed-form solution for the feature-to-attribute projection matrix V∈Rd×aV \in \mathbb{R}^{d \times a} becomes:

    V=(XX⊤+γId)−1XS⊤(SS⊤+λIa)−1V = (X X^\top + \gamma I_d)^{-1} X S^\top (S S^\top + \lambda I_a)^{-1}

    where γ>0\gamma > 0 and λ>0\lambda > 0 are regularization parameters, Id∈Rd×dI_d \in \mathbb{R}^{d \times d} is the identity matrix on feature space, and Ia∈Ra×aI_a \in \mathbb{R}^{a \times a} is the identity matrix on attribute space. At test time, zero-shot prediction for an instance x∈Rdx \in \mathbb{R}^d across z′z' unseen classes with signatures S′∈[0,1]a×z′S' \in [0, 1]^{a \times z'} continues to use class-level signatures via argmax⁡ix⊤VSi′\operatorname{argmax}_i x^\top V S'_i.

    ESZSL-AS is particularly advantageous when the number of training classes zz is small compared to the attribute dimensionality aa, as it effectively increases the number of training classes to mm.

  5. Knowl 5 — Kernelized Bilinear Zero-Shot Learning Formulation

    theoretical result

    By the Representer Theorem, if the regularizer on the linear zero-shot parameter matrix V∈Rd×aV \in \mathbb{R}^{d \times a} can be expressed in terms of inner products as Ω(V)=Ψ(V⊤V)\Omega(V) = \Psi(V^\top V), the bilinear zero-shot optimization problem can be formulated in dual kernel form.

    Let K∈Rm×mK \in \mathbb{R}^{m \times m} be the Gram matrix with entries Ki,j=⟨ϕ(xi),ϕ(xj)⟩K_{i,j} = \langle \phi(x_i), \phi(x_j) \rangle computed via a feature map ϕ(⋅)\phi(\cdot) for mm training instances. The optimization problem over coefficient matrix A∈Rm×aA \in \mathbb{R}^{m \times a} is:

    min⁡A∈Rm×aL(KAS,Y)+Ψ(S⊤A⊤KAS)\min_{A \in \mathbb{R}^{m \times a}} L(K A S, Y) + \Psi(S^\top A^\top K A S)

    where S∈[0,1]a×zS \in [0, 1]^{a \times z} contains the class attribute signatures, Y∈{−1,1}m×zY \in \{-1, 1\}^{m \times z} contains the ground-truth labels for zz training classes, and LL is a convex loss function. The resulting problem is convex and solvable globally in closed form or via standard convex optimization routines using only inner products between data instances.

  6. Knowl 6 — Multi-Class Accuracy Benchmark Comparison on Real Zero-Shot Datasets

    data/table

    Zero-shot classification performance of ESZSL and ESZSL-AS was evaluated against Direct Attribute Prediction (DAP) and Zero-Shot Recognition with Unreliable Attributes (ZSRwUA) across three benchmark datasets using combined χ2\chi^2-kernels over standard visual features (such as SIFT and PHOG). Results report multiclass accuracy (%) as mean ±\pm standard deviation over 20 random trials.

    Method AwA aPY SUN
    DAP 40.50 18.12 52.50
    ZSRwUA 43.01 ±\pm 0.07 26.02 ±\pm 0.05 56.18 ±\pm 0.27
    ESZSL 49.30 ±\pm 0.21 15.11 ±\pm 2.24 65.75 ±\pm 0.51
    ESZSL-AS — 27.27 ±\pm 1.62 61.53 ±\pm 1.03

    The datasets have the following properties:

    • Animals with Attributes (AwA): 30,475 instances, 85 attributes, 40 training classes, 10 test classes (class-level attribute annotations; ESZSL-AS not applicable). ESZSL outperforms ZSRwUA by over 14.6% relative improvement.
    • aPascal/aYahoo (aPY): 15,339 instances, 65 attributes, 20 training classes, 12 test classes (instance-level attribute annotations). Standard ESZSL degrades due to the low number of training classes relative to attributes (z=20<a=65z=20 < a=65), while ESZSL-AS achieves top accuracy with a 4.8% improvement over ZSRwUA.
    • SUN scene attributes: 14,340 instances, 102 attributes, 707 training classes, 10 test classes (instance-level attribute annotations). ESZSL achieves a 17.0% relative improvement over ZSRwUA (65.75%65.75\% vs 56.18%56.18\%) as the large number of training classes (z=707≫a=102z=707 \gg a=102) favors standard class-level pooling.
  7. Knowl 7 — Regularizer Ablation and Computational Efficiency Compared to Latent Embeddings

    empirical result

    On the Animals with Attributes (AwA) dataset using DeCAF deep convolutional features, ESZSL was compared against the Structured Joint Embedding approach of Akata et al. (2013) across varying training set sizes, and evaluated in an ablation study assessing the impact of its regularization terms:

    Training Instances Akata et al. (2013) ESZSL
    500 32.30% 33.09%
    1000 38.57% 42.44%
    2000 40.21% 44.82%

    Key empirical findings:

    1. Accuracy and Scalability: ESZSL consistently outperforms the approach of Akata et al. at all sample sizes. On 2,000 training instances, training the model of Akata et al. required over 11 hours, whereas ESZSL computed the closed-form solution in 4.12 seconds.
    2. Regularizer Component Ablation: Using all training instances with DeCAF features, ESZSL achieved 50.37% classification accuracy. Removing the cross-space regularization terms by setting γ=λ=0\gamma = \lambda = 0 in Ω(V;S,X)=γ∥VS∥Fro2+λ∥X⊤V∥Fro2+β∥V∥Fro2\Omega(V; S, X) = \gamma \|V S\|_{\text{Fro}}^2 + \lambda \|X^\top V\|_{\text{Fro}}^2 + \beta \|V\|_{\text{Fro}}^2 (leaving only standard Frobenius norm weight decay on VV) resulted in a drop to 45.02% accuracy. This demonstrates that simultaneously penalizing the projection norm in feature space and variance in attribute space is critical to generalization.
  8. Knowl 8 — Sample Efficiency and Robustness to Corrupted Attributes in Synthetic Experiments

    empirical result

    Controlled synthetic data experiments comparing ESZSL with Direct Attribute Prediction (DAP) across simulated feature dimension d=10d = 10, attribute dimension a=100a = 100, and 50 instances per class demonstrate two key properties:

    1. Training Class Sample Efficiency: When varying the number of training classes zz from 50 to 500 (evaluated on 100 unseen test classes), ESZSL significantly outperforms DAP across all settings. ESZSL trained on only 100 classes achieves higher multiclass accuracy (~63%) than DAP trained on 500 classes (~58%). Furthermore, ESZSL performance plateaus near z=200z = 200, demonstrating high data efficiency in learning attribute-feature correlations.
    2. Robustness to Non-Informative Attributes: When ψ∈[5,45]\psi \in [5, 45] out of 100 attributes are corrupted by replacing their values with non-informative random Bernoulli noise, ESZSL maintains superior performance over DAP across all noise levels. Because ESZSL optimizes multiclass classification directly while learning the weighting of each attribute implicitly, it suppresses uninformative attributes more effectively than two-stage attribute predictors.

Coverage note — Omitted only supplementary appendix proofs and secondary experimental plots from the appendix.

References

  1. 1.Akata, Zeynep, Perronnin, Florent, Harchaoui, Zaid, and Schmid, Cordelia. Label-embedding for attribute-based classification. In Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on, pp. 819–826. IEEE, 2013.
  2. 2.Argyriou, Andreas, Evgeniou, Theodoros, and Pontil, Massimiliano. Convex multi-task feature learning. Machine Learning, 73(3):243–272, 2008.
  3. 3.Argyriou, Andreas, Micchelli, Charles A, and Pontil, Massimiliano. When is there a representer theorem? vector versus matrix regularizers. The Journal of Machine Learning Research, 10:2507–2529, 2009.
  4. 4.Ben-David, Shai, Blitzer, John, Crammer, Koby, Pereira, Fernando, et al. Analysis of representations for domain adaptation. Advances in neural information processing systems, 19:137, 2007.
  5. 5.Blitzer, John, Crammer, Koby, Kulesza, Alex, Pereira, Fernando, and Wortman, Jennifer. Learning bounds for domain adaptation. In Advances in neural information processing systems, pp. 129–136, 2008.
  6. 6.Bosch, Anna, Zisserman, Andrew, and Munoz, Xavier. Representing shape with a spatial pyramid kernel. In Proceedings of the 6th ACM international conference on Image and video retrieval, pp. 401–408. ACM, 2007.
  7. 7.Croonenborghs, Tom, Driessens, Kurt, and Bruynooghe, Maurice. Learning relational options for inductive transfer in relational reinforcement learning. In Inductive Logic Programming, pp. 88–97. Springer, 2008.
  8. 8.Daume III, Hal. Frustratingly easy domain adaptation. arXiv preprint arXiv:0907.1815, 2009.
  9. 9.Dietterich, Thomas G. and Bakiri, Ghulum. Solving multiclass learning problems via error-correcting output codes. arXiv preprint cs/9501101, 1995.
  10. 10.Farhadi, Ali, Endres, Ian, Hoiem, Derek, and Forsyth, David. Describing objects by their attributes. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pp. 1778–1785. IEEE, 2009.
  11. 11.Ferrari, Vittorio and Zisserman, Andrew. Learning visual attributes. In Advances in Neural Information Processing Systems, pp. 433–440, 2007.
  12. 12.Frome, Andrea, Corrado, Greg S, Shlens, Jon, Bengio, Samy, Dean, Jeff, Mikolov, Tomas, et al. Devise: A deep visual-semantic embedding model. In Advances in Neural Information Processing Systems, pp. 2121–2129, 2013.
  13. 13.Fu, Yanwei, Hospedales, Timothy M, Xiang, Tao, and Gong, Shaogang. Learning multimodal latent attributes. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 36(2):303–316, 2014.
  14. 14.Hariharan, Bharath, Vishwanathan, SVN, and Varma, Manik. Efficient max-margin multi-label classification with applications to zero-shot learning. Machine learning, 88(1-2):127–155, 2012.
  15. 15.Hwang, Sung Ju, Sha, Fei, and Grauman, Kristen. Sharing features between objects and their attributes. In Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on, pp. 1761–1768. IEEE, 2011.
  16. 16.Jayaraman, Dinesh and Grauman, Kristen. Zero-shot recognition with unreliable attributes. pp. 3464–3472, 2014.
  17. 17.Jayaraman, Dinesh, Sha, Fei, and Grauman, Kristen. Decorrelating semantic visual attributes by resisting the urge to share. In Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on, pp. 1629–1636. IEEE, 2014.
  18. 18.Jiang, Jing and Zhai, ChengXiang. Instance weighting for domain adaptation in nlp. In ACL, volume 7, pp. 264–271, 2007.
  19. 19.Kifer, Daniel, Ben-David, Shai, and Gehrke, Johannes. Detecting change in data streams. In Proceedings of the Thirtieth international conference on Very large data bases-Volume 30, pp. 180–191. VLDB Endowment, 2004.
  20. 20.Lampert, Christoph H, Nickisch, Hannes, and Harmeling, Stefan. Learning to detect unseen object classes by between-class attribute transfer. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pp. 951–958. IEEE, 2009.
  21. 21.Lampert, Christoph H, Nickisch, Hannes, and Harmeling, Stefan. Attribute-based classification for zero-shot visual object categorization. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 36(3):453–465, 2014.
  22. 22.Lawrence, Neil D and Platt, John C. Learning to learn with the informative vector machine. In Proceedings of the twenty-first international conference on Machine learning, pp. 65. ACM, 2004.
  23. 23.Liu, Jingen, Kuipers, Benjamin, and Savarese, Silvio. Recognizing human actions by attributes. In Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on, pp. 3337–3344. IEEE, 2011.
  24. 24.Lowe, David G. Distinctive image features from scale-invariant keypoints. International journal of computer vision, 60(2):91–110, 2004.
  25. 25.Mahajan, Dhruv, Sellamanickam, Sundararajan, and Nair, Vinod. A joint learning framework for attribute models and object descriptions. In Computer Vision (ICCV), 2011 IEEE International Conference on, pp. 1227–1234. IEEE, 2011.
  26. 26.Palatucci, Mark, Hinton, Geoffrey, Pomerleau, Dean, and Mitchell, Tom M. Zero-Shot Learning with Semantic Output Codes. Neural Information Processing Systems, pp. 1–9, 2009.
  27. 27.Pan, Sinno Jialin and Yang, Qiang. A survey on transfer learning. Knowledge and Data Engineering, IEEE Transactions on, 22(10):1345–1359, 2010.
  28. 28.Patterson, Genevieve and Hays, James. Sun attribute database: Discovering, annotating, and recognizing scene attributes. In Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, pp. 2751–2758. IEEE, 2012.
  29. 29.Raykar, Vikas C, Krishnapuram, Balaji, Bi, Jinbo, Dundar, Murat, and Rao, R Bharat. Bayesian multiple instance learning: automatic feature selection and inductive transfer. In Proceedings of the 25th international conference on Machine learning, pp. 808–815. ACM, 2008.
  30. 30.Romera-Paredes, Bernardino, Aung, Hane, Bianchi-Berthouze, Nadia, and Pontil, Massimiliano. Multilinear multitask learning. In Proceedings of the 30th International Conference on Machine Learning, pp. 1444–1452, 2013.
  31. 31.Ruckert, Ulrich and Kramer, Stefan. Kernel-based inductive transfer. In Machine Learning and Knowledge Discovery in Databases, pp. 220–233. Springer, 2008.
  32. 32.Suzuki, Masahiro, Sato, Haruhiko, Oyama, Satoshi, and Kurihara, Masahito. Transfer learning based on the observation probability of each attribute. In Systems, Man and Cybernetics (SMC), 2014 IEEE International Conference on, pp. 3627–3631. IEEE, 2014.
  33. 33.Wang, Yang and Mori, Greg. A discriminative latent model of object classes and attributes. In Computer Vision–ECCV 2010, pp. 155–168. Springer, 2010.

Citation

MLA
Romera-Paredes, B., and P. H. S. Torr. “An Embarrassingly Simple Approach to Zero-Shot Learning”. Advances in Computer Vision and Pattern Recognition, Springer International Publishing, 2017, pp. 11–30, https://doi.org/10.1007/978-3-319-50077-5_2.
APA
Romera-Paredes, B., & Torr, P. H. S. (2017). An Embarrassingly Simple Approach to Zero-Shot Learning. In Advances in Computer Vision and Pattern Recognition (pp. 11–30). Springer International Publishing. https://doi.org/10.1007/978-3-319-50077-5_2
Chicago
Romera-Paredes, B., and P. H. S. Torr. 2017. “An Embarrassingly Simple Approach to Zero-Shot Learning”. In Advances in Computer Vision and Pattern Recognition. Springer International Publishing. https://doi.org/10.1007/978-3-319-50077-5_2.
Harvard
Romera-Paredes, B. and Torr, P.H.S. (2017) “An Embarrassingly Simple Approach to Zero-Shot Learning”, Advances in Computer Vision and Pattern Recognition. Springer International Publishing, pp. 11–30. Available at: https://doi.org/10.1007/978-3-319-50077-5_2.
Vancouver
1. Romera-Paredes B, Torr PHS (2017) An Embarrassingly Simple Approach to Zero-Shot Learning. In: Advances in Computer Vision and Pattern Recognition. Springer International Publishing, pp 11–30

BibTeX

@inbook{Romera_Paredes_2017, title={An Embarrassingly Simple Approach to Zero-Shot Learning}, ISBN={9783319500775}, ISSN={2191-6594}, url={http://dx.doi.org/10.1007/978-3-319-50077-5_2}, DOI={10.1007/978-3-319-50077-5_2}, booktitle={Visual Attributes}, publisher={Springer International Publishing}, author={Romera-Paredes, Bernardino and Torr, Philip H. S.}, year={2017}, pages={11–30} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors