keyword
Pascal Visual Object Classes Challenge
The Pascal Visual Object Classes Challenge was an influential annual benchmark competition and standardized evaluation framework in computer vision that ran from 2005 through 2012. Organized to advance the development and evaluation of machine perception algorithms, the challenge provided publicly available datasets of natural images paired with ground-truth annotations and standardized evaluation software. Participating algorithms were evaluated across multiple fundamental visual recognition tasks, primarily image classification, object detection, pixel-level semantic segmentation, and action classification across defined categories such as vehicles, animals, and household objects. By introducing uniform testing protocols and popularizing standard performance metrics like mean average precision, the challenge established a critical foundation for objectively comparing models and assessing progress in automated visual understanding.
4 items

Single-Shot Refinement Neural Network for Object Detection
Shifeng Zhang, Longyin Wen, Xiao Bian, Zhen Lei, Stan Z. Li
Why you should read this
Proposes RefineDet, an object detector that combines the high accuracy of two-stage methods with the fast inference of single-stage models by using anchor refinement and feature transfer modules to filter false positives and optimize bounding boxes before final classification.
For object detection, the two-stage approach (e.g., Faster R-CNN) has been achieving the highest accuracy, whereas the one-stage approach (e.g., SSD) has the advantage of high efficiency. To inherit the merits of both while overcoming their disadvantages, in this paper, we propose a novel single-shot based detector, called RefineDet, that achieves better accuracy than two-stage methods and maintains comparable efficiency of one-stage methods. RefineDet consists of two inter-connected modules, namely, the anchor refinement module and the object detection module. Specifically, the former aims to (1) filter out negative anchors to reduce search space for the classifier, and (2) coarsely adjust the locations and sizes of anchors to provide better initialization for the subsequent regressor. The latter module takes the refined anchors as the input from the former to further improve the regression and predict multi-class label. Meanwhile, we design a transfer connection block to transfer the features in the anchor refinement module to predict locations, sizes and class labels of objects in the object detection module. The multi-task loss function enables us to train the whole network in an end-to-end way. Extensive experiments on PASCAL VOC 2007, PASCAL VOC 2012, and MS COCO demonstrate that RefineDet achieves state-of-the-art detection accuracy with high efficiency. Code is available at this https URL
Added
2026-09-25

Evaluating Color Descriptors for Object and Scene Recognition
K. V. D. Sande, T. Gevers, Cees G. M. Snoek
Why you should read this
Establishes a systematic taxonomy and evaluation framework for color descriptors under varying photometric conditions, demonstrating that methods like OpponentSIFT substantially improve object and scene recognition accuracy over standard intensity-based SIFT.
Image category recognition is important to access visual information on the level of objects and scene types. So far, intensity-based descriptors have been widely used for feature extraction at salient points. To increase illumination invariance and discriminative power, color descriptors have been pro- posed. Because many different descriptors exist, a structured overview is required of color invariant descriptors in the context of image category recognition. Therefore, this paper studies the invariance properties and the distinctiveness of color descriptors¹ in a structured way. The analytical invariance properties of color descriptors are explored, using a taxonomy based on invariance properties with respect to photometric transformations, and tested experimentally using a dataset with known illumination conditions. In addition, the distinctiveness of color descriptors is assessed experimentally using two benchmarks, one from the image domain and one from the video domain. From the theoretical and experimental results, it can be derived that invariance to light intensity changes and light color changes affects category recognition. The results reveal further that, for light intensity shifts, the usefulness of invariance is category-specific. Overall, when choosing a single descriptor and no prior knowledge about the dataset and object and scene categories is available, the OpponentSIFT is recommended. Furthermore, a combined set of color descriptors outperforms intensity- based SIFT and improves category recognition by 8% on the PASCAL VOC 2007 and by 7% on the Mediamill Challenge.
Added
2026-09-16

Improving the Fisher Kernel for Large-Scale Image Classification
Florent Perronnin, Jorge Sánchez, Thomas Mensink
Why you should read this
Proposes key improvements to the Fisher vector framework—including power normalization and L2 normalization—that allow fast linear classifiers to match or exceed complex non-linear methods and achieve state-of-the-art image classification accuracy at scale.
The Fisher kernel (FK) is a generic framework which combines the benefits of generative and discriminative approaches. In the context of image classification the FK was shown to extend the popular bag-of-visual-words (BOV) by going beyond count statistics. However, in practice, this enriched representation has not yet shown its superiority over the BOV. In the first part we show that with several well-motivated modifications over the original framework we can boost the accuracy of the FK. On PASCAL VOC 2007 we increase the Average Precision (AP) from 47.9% to 58.3%. Similarly, we demonstrate state-of-the-art accuracy on CalTech 256. A major advantage is that these results are obtained using only SIFT descriptors and costless linear classifiers. Equipped with this representation, we can now explore image classification on a larger scale. In the second part, as an application, we compare two abundant resources of labeled images to learn classifiers: ImageNet and Flickr groups. In an evaluation involving hundreds of thousands of training images we show that classifiers learned on Flickr groups perform surprisingly well (although they were not intended for this purpose) and that they can complement classifiers learned on more carefully annotated datasets.
Added
2026-09-14

The Pascal Visual Object Classes Challenge: A Retrospective
Mark Everingham, S. Eslami, Luc Van Gool, Christopher K. I. Williams, John Winn, Andrew Zisserman
Why you should read this
Introduces novel statistical evaluation techniques and distills critical lessons from five years of the PASCAL VOC challenge to help researchers diagnose object recognition errors and design better vision benchmarks.
Abstract The PASCAL Visual Object Classes (VOC) challenge consists of two components: (i) a publicly available dataset of images together with ground truth annotation and standardised evaluation software; and (ii) an annual competition and workshop. There are five challenges: classification, detection, segmentation, action classification, and person layout. In this paper we provide a review of the challenge from 2008–2012. The paper is intended for two audiences: algorithm designers, researchers who want to see what the state of the art is, as measured by performance on the VOC datasets, along with the limitations and weak points of the current generation of algorithms; and, challenge designers, who want to see what we as organisers have learnt from the process and our recommendations for the organisation of future challenges. To analyse the performance of submitted algorithms on the VOC datasets we introduce a number of novel evaluation methods: a bootstrapping method for determining whether differences in the performance of two algorithms are significant or not; a normalised average precision so that performance can be compared across classes with different proportions of positive instances; a clustering method for visualising the performance across multiple algorithms so that the hard and easy images can be identified; and the use of a joint classifier over the submitted algorithms in order to measure their complementarity and combined performance. We also analyse the community’s progress through time using the methods of Hoiem et al (2012) to identify the types of occurring errors. We conclude the paper with an appraisal of the aspects of the challenge that worked well, and those that could be improved in future challenges.
Added
2026-09-09
