keyword
Fisher kernel
A Fisher kernel is a function used in machine learning that measures the similarity between two data points by leveraging a generative probabilistic model. It bridges generative modeling and discriminative classification by calculating the gradient of the log-likelihood of an input sample with respect to the parameters of a fitted probability distribution, producing a gradient representation known as the Fisher score. The kernel measures similarity by computing the inner product of these score vectors, normalized by the inverse Fisher information matrix of the underlying model. This mathematical formulation enables complex or variable-sized data structures, such as sequences and image descriptor collections, to be mapped into fixed-dimensional vector spaces that are well-suited for linear classifiers and support vector machines.
2 items

Self-taught learning: transfer learning from unlabeled data
Rajat Raina, Alexis Battle, Honglak Lee, Benjamin Packer, Andrew Y. Ng
Why you should read this
Proposes a machine learning framework that applies sparse coding to easily accessible, uncurated, and unlabeled data from entirely different classes to build higher-level feature representations that improve supervised classification performance across image, audio, and text tasks.
We present a new machine learning framework called “self-taught learning” for using unlabeled data in supervised classification tasks. We do not assume that the unlabeled data follows the same class labels or generative distribution as the labeled data. Thus, we would like to use a large number of unlabeled images (or audio samples, or text documents) randomly downloaded from the Internet to improve performance on a given image (or audio, or text) classification task. Such unlabeled data is significantly easier to obtain than in typical semi-supervised or transfer learning settings, making self-taught learning widely applicable to many practical learning problems. We describe an approach to self-taught learning that uses sparse coding to construct higher-level features using the unlabeled data. These features form a succinct input representation and significantly improve classification performance. When using an SVM for classification, we further show how a Fisher kernel can be learned for this representation.
Added
2026-09-18

Improving the Fisher Kernel for Large-Scale Image Classification
Florent Perronnin, Jorge Sánchez, Thomas Mensink
Why you should read this
Proposes key improvements to the Fisher vector framework—including power normalization and L2 normalization—that allow fast linear classifiers to match or exceed complex non-linear methods and achieve state-of-the-art image classification accuracy at scale.
The Fisher kernel (FK) is a generic framework which combines the benefits of generative and discriminative approaches. In the context of image classification the FK was shown to extend the popular bag-of-visual-words (BOV) by going beyond count statistics. However, in practice, this enriched representation has not yet shown its superiority over the BOV. In the first part we show that with several well-motivated modifications over the original framework we can boost the accuracy of the FK. On PASCAL VOC 2007 we increase the Average Precision (AP) from 47.9% to 58.3%. Similarly, we demonstrate state-of-the-art accuracy on CalTech 256. A major advantage is that these results are obtained using only SIFT descriptors and costless linear classifiers. Equipped with this representation, we can now explore image classification on a larger scale. In the second part, as an application, we compare two abundant resources of labeled images to learn classifiers: ImageNet and Flickr groups. In an evaluation involving hundreds of thousands of training images we show that classifiers learned on Flickr groups perform surprisingly well (although they were not intended for this purpose) and that they can complement classifiers learned on more carefully annotated datasets.
Added
2026-09-14
