Evaluating Color Descriptors for Object and Scene Recognition

K. V. D. SandeT. GeversCees G. M. Snoek

article2010TPAMI2,124 citations
Listen

Automated object and scene recognition in digital images and video is essential for modern visual search, categorization, and retrieval systems. However, real-world variations in illuminationsuch as changes in light intensity, diffuse lighting shifts, and light color fluctuationsfrequently degrade the accuracy of standard visual descriptors. While intensity-based methods like the standard Scale-Invariant Feature Transform (SIFT) are widely used, incorporating color information can increase distinctiveness and robustness. Because many distinct color representations exist without a unifying framework, system architects have lacked clear guidance on how these methods relate to one another and which descriptors to select for specific visual categorization tasks.

The article establishes a systematic taxonomy of color descriptors based on their analytical photometric invariance properties and evaluates their discriminative power across standard object and scene recognition benchmarks. The goal is to provide evidence-based recommendations on selecting and combining descriptors to maximize recognition performance under real-world lighting variations.

To evaluate these techniques, the authors established a physical reflectance framework that models five primary lighting transformations: light intensity changes, light intensity shifts, combined intensity changes and shifts, light color changes, and full color changes with shifts. They analyzed several families of descriptorsincluding color histograms, generalized color moments, and color extensions of SIFTwithin a bag-of-words machine learning pipeline. The theoretical invariance of each descriptor was experimentally verified on the Amsterdam Library of Object Images (ALOI), comprising over 48,000 images of 1,000 objects under controlled illumination, viewpoint, and compression changes. The distinctiveness and overall recognition performance were then benchmarked on two large, real-world collections: the PASCAL VOC 2007 image dataset (nearly 10,000 photographs across 20 object categories) and the Mediamill Challenge video benchmark (over 43,000 keyframes across 39 concepts).

The analysis produced several key findings regarding descriptor design and performance. First, derivative-based color SIFT descriptors consistently and significantly outperform color histograms and color moments in distinctiveness, as histograms and moments lack local spatial context and degrade rapidly under moderate image compression. Second, invariant properties must be matched to operational conditions: while invariance to light intensity scaling is universally beneficial, invariance to light intensity shifts is category-specific. For categories with severe diffuse or mixed lighting (such as outdoor scenes, buildings, and vehicles), shift-invariant descriptors like OpponentSIFT deliver superior results; conversely, for categories where diffuse lighting shifts are minimal, excess invariance reduces discriminative power, making scale-invariant descriptors like C-SIFT preferable. Third, when choosing a single general-purpose descriptor without prior knowledge of the target dataset, OpponentSIFT achieved the highest overall reliability. Finally, combining multiple complementary color descriptors achieved state-of-the-art results, improving mean average precision over standard intensity-based SIFT by 8% on PASCAL VOC 2007 (reaching 0.605) and by 7% on the Mediamill Challenge (reaching 0.510).

These findings demonstrate that adopting color-invariant descriptors yields substantial performance gains in visual categorization without requiring fundamental redesigns of existing classification pipelines. However, system designers should avoid assuming that maximal invariance is always optimal. Over-engineering invariance to light color changes or shifts when they do not occur in practice discards useful discriminative data, diminishing category recognition accuracy.

For engineering and data science teams implementing visual categorization pipelines, the article supports practical recommendations. When constrained to deploying a single feature representation without prior knowledge of dataset lighting, OpponentSIFT should be the default choice, followed by C-SIFT. For applications where maximizing accuracy is paramount, teams should fuse multiple descriptors with varying invariance levels (including SIFT, OpponentSIFT, C-SIFT, rgSIFT, and RGB-SIFT) alongside spatial pyramid pooling and dense sampling. Further performance gains can be pursued by developing category-specific feature selection strategies rather than applying a uniform descriptor set across all visual classes.

The findings are supported by rigorous benchmark testing and statistical bootstrapping across thousands of diverse images and video frames. Readers should note that extreme lighting conditions causing severe color clipping (where pixel values saturate to maximum or minimum limits) degrade the performance of all evaluated descriptors. Additionally, because standard recording equipment frequently performs automatic white balancing, light color changes occur less often in standard benchmarks than intensity variations. Within these operating boundaries, the conclusions provide high confidence for guiding descriptor selection in production computer vision systems.

Sande et al (2010).pdf
  • Paper: Distinctive Image Features from Scale-Invariant Keypoints, David G. Lowe (2004). This seminal paper introduces the standard intensity-based SIFT descriptor and scale-space keypoint framework upon which the source's color-invariant SIFT variants are systematically built and evaluated.
  • Paper: A performance evaluation of local descriptors, Krystian Mikolajczyk et al. (2005). This benchmark established the foundational methodology for evaluating local descriptor distinctiveness and invariance under image transformations, providing the empirical precedent for the source's evaluation protocol.
  • Paper: Visual categorization with bags of keypoints, Gabriella Csurka et al. (2004). This paper presents the bag-of-keypoints categorization architecture that serves as the baseline visual recognition pipeline employed throughout the source.
  • Paper: Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories, Svetlana Lazebnik et al. (2006). This work introduces spatial pyramid matching to preserve spatial layout within visual vocabularies, which the source adopts to evaluate dense descriptor sampling and category recognition.
  • Paper: Color indexing, Michael J. Swain et al. (1991). This classic study introduced color histograms and opponent-color indexing for object recognition, defining the baseline color representations that the source analyzes and contrasts against spatial derivative descriptors.
  • Paper: A reflectance model for computer graphics, Robert L. Cook et al. (1981). This foundational paper establishes the physical reflectance and illumination models essential for formalizing the photometric invariance classes analyzed in the source.
  • Paper: A Comparison of Affine Region Detectors, K. Mikolajczyk et al. (2005). This comparative benchmark evaluates region detectors across photometric and geometric variations, underpinning the interest-region and invariance assumptions tested in the source.
  • Paper: SURF: Speeded Up Robust Features, Herbert Bay et al. (2006). This work introduces SURF and fast local feature computation, representing key contemporary local descriptor baselines evaluated alongside SIFT variants.

Table of Contents

  • I. INTRODUCTION
  • II. REFLECTANCE MODEL
  • A. Diagonal Model
  • B. Photometric Analysis
  • III. COLOR DESCRIPTORS AND INVARIANT PROPERTIES
  • A. Histograms
  • B. Color Moments and Moment Invariants
  • C. Color SIFT Descriptors
  • D. Conclusion
  • IV. EXPERIMENTAL SETUP
  • A. Feature Extraction Pipelines
  • B. Classification
  • C. Experiment 1: Illumination Changes
  • D. Experiment 2: Image Benchmark
  • E. Experiment 3: Video Benchmark
  • F. Evaluation Criteria
  • V. RESULTS
  • A. Experiment 1: Illumination Changes
  • B. Experiment 2: Image Benchmark
  • C. Experiment 3: Video Benchmark
  • D. Comparison with state-of-the-art
  • E. Discussion
  • VI. CONCLUSION
  • ACKNOWLEDGMENTS
  • REFERENCES

Knowls

  1. Knowl 1 — Taxonomy of Photometric Variations via the Diagonal-Offset Model

    model/method

    Photometric changes between an image taken under an unknown illuminant fu=(Ru,Gu,Bu)T\mathbf{f}^u = (R^u, G^u, B^u)^T and under a canonical reference illuminant fc=(Rc,Gc,Bc)T\mathbf{f}^c = (R^c, G^c, B^c)^T are represented using the diagonal-offset model:

    (RcGcBc)=(a000b000c)(RuGuBu)+(o1o2o3)\begin{pmatrix} R^c \\ G^c \\ B^c \end{pmatrix} = \begin{pmatrix} a & 0 & 0 \\ 0 & b & 0 \\ 0 & 0 & c \end{pmatrix} \begin{pmatrix} R^u \\ G^u \\ B^u \end{pmatrix} + \begin{pmatrix} o_1 \\ o_2 \\ o_3 \end{pmatrix}

    where diagonal elements a,b,c>0a, b, c > 0 account for illuminant spectral changes and scaling, while the offset (o1,o2,o3)T(o_1, o_2, o_3)^T accounts for diffuse light, ambient scattering, sensor infrared sensitivity, and interreflections. Five physical variation types are categorized based on constraints on this model:

    1. Light Intensity Change (a=b=ca = b = c, o1=o2=o3=0o_1 = o_2 = o_3 = 0): Channels scale uniformly by a factor aa. This models shadows, shading, and changes in light source brightness. Invariance corresponds to scale-invariance with respect to intensity.
    2. Light Intensity Shift (a=b=c=1a = b = c = 1, o1=o2=o3o_1 = o_2 = o_3): Channels shift by an equal additive offset o1o_1. This models white diffuse lighting, specular highlights under white light, interreflections, and uniform sensor bias. Invariance corresponds to shift-invariance with respect to intensity.
    3. Light Intensity Change and Shift (a=b=ca = b = c, o1=o2=o3o_1 = o_2 = o_3): Simultaneous uniform scaling and uniform shift of all color channels.
    4. Light Color Change (abca \neq b \neq c, o1=o2=o3=0o_1 = o_2 = o_3 = 0): Color channels scale independently under the full diagonal (von Kries) model, accounting for changes in illuminant color and spectral scattering.
    5. Light Color Change and Shift (abca \neq b \neq c, o1o2o3o_1 \neq o_2 \neq o_3): Independent channel scaling and arbitrary independent channel offsets under the full diagonal-offset model.
  2. Knowl 2 — Analytical Invariance Taxonomy of Color Descriptors

    theoretical result

    Under the analytical condition that no color clipping occurs (pixel values remain inside the valid recording range [0,255][0, 255]), image descriptors exhibit distinct invariance properties with respect to the five categories of the diagonal-offset illumination model:

    Descriptor Light Intensity Change Light Intensity Shift Light Intensity Change Shift Light Color Change Light Color Change Shift
    RGB Histogram - - - - -
    O1,O2O_1, O_2 - + - - -
    O3O_3 (Intensity) - - - - -
    Hue Histogram + + + - -
    Saturation - - - - -
    r,gr, g Chromaticity + - - - -
    Transformed Color + + + + +
    Color Moments - + - - -
    Moment Invariants + + + + +
    SIFT (I\nabla I) + + + - -
    HSV-SIFT - - - - -
    HueSIFT + + + - -
    OpponentSIFT + + + - -
    C-SIFT + - - - -
    rgSIFT + - - - -
    Transformed Color SIFT + + + + +
    RGB-SIFT + + + + +

    In the table, '+' denotes analytical invariance and '-' denotes lack of invariance.

  3. Knowl 3 — Opponent Color Space and OpponentSIFT Formulation

    model/method

    The opponent color space transforms (R,G,B)(R, G, B) color channels into three orthogonal components:

    (O1O2O3)=(RG2R+G2B6R+G+B3)\begin{pmatrix} O_1 \\ O_2 \\ O_3 \end{pmatrix} = \begin{pmatrix} \frac{R - G}{\sqrt{2}} \\[4pt] \frac{R + G - 2B}{\sqrt{6}} \\[4pt] \frac{R + G + B}{\sqrt{3}} \end{pmatrix}

    Channel O3O_3 represents light intensity, whereas O1O_1 and O2O_2 represent chromatic opponent differences. For identical channel offsets (o1=o2=o3o_1 = o_2 = o_3), subtraction cancels out the shift in O1O_1 and O2O_2:

    (Ru+o1)(Gu+o1)2=RuGu2=O1\frac{(R^u + o_1) - (G^u + o_1)}{\sqrt{2}} = \frac{R^u - G^u}{\sqrt{2}} = O_1

    OpponentSIFT computes standard 128-dimensional SIFT descriptors independently on each of the three opponent channels O1,O2,O3O_1, O_2, O_3, producing a 384-dimensional descriptor vector. Because the spatial derivative operations inside SIFT cancel additive offsets and SIFT vector length normalization cancels multiplicative scaling factors, OpponentSIFT achieves scale-invariance and shift-invariance with respect to light intensity changes and shifts.

  4. Knowl 4 — Equivalence of RGB-SIFT and Transformed Color SIFT

    theoretical result

    The transformed color model standardizes each color channel independently over a local region or image:

    R=RμRσR,G=GμGσG,B=BμBσBR' = \frac{R - \mu_R}{\sigma_R}, \quad G' = \frac{G - \mu_G}{\sigma_G}, \quad B' = \frac{B - \mu_B}{\sigma_B}

    where μC\mu_C and σC\sigma_C denote the mean and standard deviation of channel C{R,G,B}C \in \{R, G, B\}, yielding zero-mean, unit-variance distributions invariant to independent channel scalings and arbitrary additive offsets.

    Transformed Color SIFT applies SIFT to each normalized channel (R,G,B)(R', G', B'). RGB-SIFT applies standard SIFT to the raw color channels (R,G,B)(R, G, B) independently.

    Because SIFT computes gradients C\nabla C, any additive offset μC\mu_C is inherently eliminated ((CμC)=C)(\nabla(C - \mu_C) = \nabla C). Furthermore, because SIFT normalizes its output descriptor vector to unit Euclidean length, division by σC\sigma_C is implicitly performed during vector normalization. Consequently, the resulting descriptor values of RGB-SIFT and Transformed Color SIFT are identical. Both achieve invariance to light intensity changes, light intensity shifts, and light color changes and shifts under the full diagonal-offset model.

  5. Knowl 5 — Formulations of C-SIFT, rgSIFT, and HueSIFT Descriptors

    model/method

    Color extensions of the SIFT descriptor incorporate color information while targeting specific photometric invariances:

    1. C-SIFT: Uses the normalized opponent color invariant channels C1=O1/O3C_1 = O_1 / O_3 and C2=O2/O3C_2 = O_2 / O_3. Dividing by the intensity channel O3=(R+G+B)/3O_3 = (R + G + B)/\sqrt{3} cancels out the multiplicative scaling factor aa in the diagonal model, conferring scale-invariance to light intensity changes. However, non-zero offsets do not cancel when taking derivatives of these channel ratios; thus C-SIFT is not shift-invariant.
    2. rgSIFT: Computes SIFT descriptors on the normalized chromaticity channels r=R/(R+G+B)r = R / (R + G + B) and g=G/(R+G+B)g = G / (R + G + B) (bb is redundant since r+g+b=1r + g + b = 1). Normalization removes uniform intensity scaling, making rgSIFT scale-invariant to light intensity changes, but it lacks shift-invariance.
    3. HueSIFT: Concatenates a 128-dimensional intensity SIFT descriptor with a saturation-weighted hue histogram. Because hue HH is unstable near the neutral grey axis where certainty is inversely proportional to saturation SS, each sample in the hue histogram is weighted by its saturation value. HueSIFT is scale- and shift-invariant to light intensity changes and shifts.
    4. HSV-SIFT: Computes SIFT across H,S,VH, S, V channels separately (3×128=3843 \times 128 = 384 dimensions). Because saturation SS and intensity VV lack scale/shift invariance and hue instability at low saturation is unweighted, the combined HSV-SIFT descriptor possesses no analytical photometric invariance.
  6. Knowl 6 — Bag-of-Words and Multi-Feature Spatial Pyramid Classification Pipeline

    experimental setup

    The image and video category recognition architecture uses a bag-of-words representation with support vector machines:

    1. Feature Sampling: Salient point extraction using the scale-invariant Harris-Laplace point detector (Harris corner detector followed by Laplacian-of-Gaussians scale selection) and/or regular dense pixel grid sampling.
    2. Codebook Construction: kk-means clustering is performed on 200,000 local descriptors sampled randomly from training data to construct a visual codebook of k=4,000k = 4{,}000 visual words.
    3. Quantization and Normalization: Local descriptors are mapped to the nearest visual word by Euclidean distance. Histograms of visual words are normalized to sum to 1.
    4. Spatial Pyramid Matching: Feature vectors are extracted across multi-level spatial subdivisions: the full image (1×11 \times 1, weight w=1w = 1), a 2×22 \times 2 grid (4 quadrants, weight w=1/4w = 1/4 each), and a 1×31 \times 3 horizontal strip layout (3 bars, weight w=1/3w = 1/3 each).
    5. SVM Kernel Classification: For multi-feature combinations of mm feature vectors {F(1),,F(m)}\{\vec{F}^{(1)}, \dots, \vec{F}^{(m)}\}, a weighted χ2\chi^2 kernel is employed:

    k({F(1),,F(m)},{F(1),,F(m)})=exp(1j=1mwjj=1mwjDjdistχ2(F(j),F(j)))k\left(\{\vec{F}^{(1)}, \dots, \vec{F}^{(m)}\}, \{\vec{F}'^{(1)}, \dots, \vec{F}'^{(m)}\}\right) = \exp\left( -\frac{1}{\sum_{j=1}^m w_j} \sum_{j=1}^m \frac{w_j}{D_j} \operatorname{dist}_{\chi^2}(\vec{F}^{(j)}, \vec{F}'^{(j)}) \right)

    where distχ2(F,F)=12i=1n(FiFi)2Fi+Fi\operatorname{dist}_{\chi^2}(\vec{F}, \vec{F}') = \frac{1}{2} \sum_{i=1}^n \frac{(F_i - F'_i)^2}{F_i + F'_i} (with 0/0=00/0 = 0) and DjD_j is the mean χ2\chi^2 distance between training examples for feature jj. Binary SVM cost parameters are tuned by 3-fold cross-validation over 242^{-4} to 242^4, using class weights (#pos+#neg)/#pos(\#\text{pos} + \#\text{neg})/\#\text{pos} and (#pos+#neg)/#neg(\#\text{pos} + \#\text{neg})/\#\text{neg}.

  7. Knowl 7 — Empirical Verification of Invariance on the ALOI Dataset

    empirical result

    Testing on 1,000 objects from the Amsterdam Library of Object Images (ALOI; >48,000 images) under single-example nearest-neighbor classification confirms theoretical invariance properties:

    • Light Intensity Scaling (a[0.33,3.0]a \in [0.33, 3.0]): Non-invariant descriptors (RGB histogram, opponent histogram, color moments) drop from ~100% correct identification to <40% at a=0.33a = 0.33 or a=3.0a = 3.0. Scale-invariant descriptors (SIFT, OpponentSIFT, C-SIFT, rgSIFT, RGB-SIFT, Moment Invariants) maintain >90% accuracy until extreme scaling where color clipping (>50% clipped pixels at range boundaries [0,255][0, 255]) degrades all descriptors.
    • Light Intensity Shifts (o1[20,20]o_1 \in [-20, 20]): Non-shift-invariant descriptors (RGB histogram, rg-histogram, C-SIFT, rgSIFT) drop to 40-70% at shifts of ±20\pm 20. OpponentSIFT and RGB-SIFT retain >95% accuracy. Color moments and moment invariants tolerate only small shifts before degrading.
    • Light Color Changes (Color temperature 3075 K to 2175 K): Histograms lacking color invariance degrade to <20%. In practice, SIFT, OpponentSIFT, C-SIFT, and rgSIFT remain robust (>80-95%), whereas HSV-SIFT and HueSIFT degrade steadily to ~60%.
    • JPEG Compression (Quality 90% down to 30%): Hue histograms, rg-histograms, transformed color histograms, and moment invariants degrade rapidly (dropping to 30-70%), whereas SIFT and all color SIFT variants maintain near 100% accuracy, demonstrating superior robustness to compression artifacts.
  8. Knowl 8 — Category Recognition Performance on Image and Video Benchmarks

    empirical result

    Evaluation on PASCAL VOC Challenge 2007 (image benchmark, 20 categories, 9,963 images, 11-point interpolated mean average precision [MAP]) and Mediamill Challenge / TRECVID 2005 (video benchmark, 39 LSCOM-Lite categories, 43,907 keyframes, non-interpolated MAP) reveals domain-specific descriptor behaviors:

    1. SIFT vs. Histogram/Moment Descriptors: SIFT and color SIFT descriptors consistently outperform color histograms and moment invariants across all benchmarks due to greater spatial distinctiveness.
    2. Image Benchmark (PASCAL VOC 2007): C-SIFT achieves the highest single-descriptor MAP (~0.44), followed closely by rgSIFT (~0.43) and OpponentSIFT (~0.42). Scale invariance to light intensity changes is critical for categories like bird, boat, horse, motorbike, person, potted plant, and sheep. RGB-SIFT exhibits reduced performance compared to C-SIFT because invariance to light color changes reduces discriminative power when illuminant shifts are negligible.
    3. Video Benchmark (Mediamill Challenge): OpponentSIFT achieves the top single-descriptor MAP (~0.405), significantly outperforming C-SIFT (~0.39) and RGB-SIFT (~0.39). In broadcast video with frequent indoor-to-outdoor transitions, weather changes, and diffuse lighting (e.g., building, meeting, mountain, sky, studio, vegetation), invariance to both intensity changes and shifts is critical.
  9. Knowl 9 — Performance of Color Descriptor Combinations Across Benchmarks

    data/table

    Fusing complementary color descriptors with intensity SIFT yields substantial performance gains over intensity SIFT alone, achieving state-of-the-art results on PASCAL VOC 2007 and the Mediamill Challenge:

    Dataset Author Descriptors Spatial Pyramid MAP
    PASCAL VOC 2007 Marszałek et al. SIFT, HueSIFT, other 1x1+2x2+1x3 0.575
    PASCAL VOC 2007 Marszałek et al. SIFT, HueSIFT, other (+ feature selection) 1x1+2x2+1x3 0.594
    PASCAL VOC 2007 This paper SIFT 1x1+2x2+1x3 0.558
    PASCAL VOC 2007 This paper C-SIFT 1x1+2x2+1x3 0.566
    PASCAL VOC 2007 This paper SIFT, OpponentSIFT, rgSIFT, C-SIFT, RGB-SIFT 1x1+2x2+1x3 0.605
    Mediamill Challenge Snoek et al. Weibull baseline 1x1 0.250
    Mediamill Challenge This paper SIFT 1x1+2x2+1x3 0.476
    Mediamill Challenge This paper OpponentSIFT 1x1+2x2+1x3 0.494
    Mediamill Challenge This paper SIFT, OpponentSIFT, rgSIFT, C-SIFT, RGB-SIFT 1x1+2x2+1x3 0.510

    On PASCAL VOC 2007, fusing five descriptors ({SIFT, OpponentSIFT, rgSIFT, C-SIFT, RGB-SIFT}) using Harris-Laplace and dense sampling yields a MAP of 0.605 (an 8% relative increase over SIFT at 0.558). On the Mediamill Challenge, the combined set yields 0.510 MAP (a 7% relative improvement over SIFT at 0.476, and a 104% improvement over the 0.250 baseline). The same combined descriptor set submitted to PASCAL VOC 2008 achieved 0.549 MAP and to NIST TRECVID 2008 achieved 0.194 inferred MAP.

  10. Knowl 10 — Guidelines for Single Color Descriptor Selection

    model/method

    When selecting a single color descriptor, the optimal choice depends on the dataset domain and prior knowledge regarding lighting conditions:

    Rank PASCAL VOC 2007 Mediamill Challenge Unknown Data
    1 C-SIFT OpponentSIFT OpponentSIFT
    2 OpponentSIFT RGB-SIFT C-SIFT
    3 RGB-SIFT C-SIFT RGB-SIFT
    4 SIFT SIFT SIFT
    • Default Recommendation for Unknown Data: OpponentSIFT is recommended when no prior knowledge of the dataset or lighting conditions is available. It incorporates scale- and shift-invariance to light intensity changes while preserving discriminative color information.
    • Controlled/Natural Photography (e.g., PASCAL VOC): C-SIFT performs best because scale invariance to intensity is essential, while illuminant color variation is rare due to white balancing.
    • Uncontrolled Broadcast Video (e.g., Mediamill Challenge): OpponentSIFT performs best because diffuse lighting shifts and indoor/outdoor transitions require both scale and shift intensity invariance.

Coverage note — None was omitted; all key theoretical definitions, invariant taxonomies, descriptor formulations, experimental setups, and empirical evaluation results across datasets have been fully represented.

References

  1. 1.R. Datta, D. Joshi, J. Li, and J. Z. Wang, “Image retrieval: Ideas, influences, and trends of the new age,” ACM Computing Surveys, vol. 40, no. 2, pp. 1–60, 2008.
  2. 2.R. Fergus, F.-F. Li, P. Perona, and A. Zisserman, “Learning object categories from Google’s image search,” in IEEE International Conference on Computer Vision, Beijing, China, 2005, pp. 1816–1823.
  3. 3.S. Lazebnik, C. Schmid, and J. Ponce, “Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories.” in IEEE Conference on Computer Vision and Pattern Recognition, vol. 2, New York, USA, 2006, pp. 2169–2178.
  4. 4.J. Vogel and B. Schiele, “Semantic modeling of natural scenes for content-based image retrieval,” International Journal of Computer Vision, vol. 72, no. 2, pp. 133–157, 2007.
  5. 5.J. Zhang, M. Marszałek, S. Lazebnik, and C. Schmid, “Local features and kernels for classification of texture and object categories: A comprehensive study,” International Journal of Computer Vision, vol. 73, no. 2, pp. 213–238, 2007.
  6. 6.S.-F. Chang, D. Ellis, W. Jiang, K. Lee, A. Yanagawa, A. C. Loui, and J. Luo, “Large-Scale Multimodal Semantic Concept Detection for Consumer Video,” in ACM International Workshop on Multimedia Information Retrieval, Augsburg, Germany, 2007, pp. 255–264.
  7. 7.A. F. Smeaton, P. Over, and W. Kraaij, “Evaluation campaigns and TRECVid,” in ACM International Workshop on Multimedia Information Retrieval, Santa Barbara, USA, 2006, pp. 321–330.
  8. 8.Y.-G. Jiang, C.-W. Ngo, and J. Yang, “Towards optimal bag-of-features for object categorization and semantic video retrieval,” in ACM International Conference on Image and Video Retrieval, Amsterdam, The Netherlands, 2007, pp. 494–501.
  9. 9.D. G. Lowe, “Distinctive image features from scale-invariant keypoints.” International Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, 2004.
  10. 10.K. Mikolajczyk, T. Tuytelaars, C. Schmid, A. Zisserman, J. Matas, F. Schaffalitzky, T. Kadir, and L. Van Gool, “A comparison of affine region detectors,” International Journal of Computer Vision, vol. 65, no. 1-2, pp. 43–72, 2005.
  11. 11.T. Tuytelaars and K. Mikolajczyk, “Local invariant feature detectors: A survey,” Foundations and Trends in Computer Graphics and Vision, vol. 3, no. 3, pp. 177–280, 2008.
  12. 12.A. E. Abdel-Hakim and A. A. Farag, “CSIFT: A SIFT descriptor with color invariant characteristics,” in IEEE Conference on Computer Vision and Pattern Recognition, New York, USA, 2006, pp. 1978–1983.
  13. 13.J. M. Geusebroek, R. van den Boomgaard, A. W. M. Smeulders, and H. Geerts, “Color invariance,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 23, no. 12, pp. 1338–1350, 2001.
  14. 14.J. van de Weijer, T. Gevers, and A. Bagdanov, “Boosting color saliency in image feature detection,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 28, no. 1, pp. 150–156, 2006.
  15. 15.G. J. Burghouts and J. M. Geusebroek, “Performance evaluation of local color invariants,” Computer Vision and Image Understanding, vol. 113, pp. 48–62, 2009.
  16. 16.A. Bosch, A. Zisserman, and X. Muoz, “Scene classification using a hybrid generative/discriminative approach,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, no. 04, pp. 712–727, 2008.
  17. 17.G. D. Finlayson, M. S. Drew, and B. V. Funt, “Spectral sharpening: sensor transformations for improved color constancy,” Journal of the Optical Society of America A, vol. 11, no. 5, p. 1553, 1994.
  18. 18.J. von Kries, “Influence of adaptation on the effects produced by luminous stimuli,” In MacAdam, D.L. (Ed.), Sources of Color Vision. MIT Press, Cambridge, MS., 1970.
  19. 19.K. E. A. van de Sande, T. Gevers, and C. G. M. Snoek, “Evaluation of color descriptors for object and scene recognition,” in IEEE Conference on Computer Vision and Pattern Recognition, Anchorage, Alaska, USA, June 2008.
  20. 20.J. M. Geusebroek, G. J. Burghouts, and A. W. M. Smeulders, “The Amsterdam library of object images,” International Journal of Computer Vision, vol. 61, no. 1, pp. 103–112, 2005.
  21. 21.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results.” [Online]. Available: http://www.pascal-network.org/challenges/VOC/voc2007/
  22. 22.C. G. M. Snoek, M. Worring, J. C. van Gemert, J.-M. Geusebroek, and A. W. M. Smeulders, “The challenge problem for automated detection of 101 semantic concepts in multimedia,” in ACM International Conference on Multimedia, Santa Barbara, USA, 2006, pp. 421–430.
  23. 23.M. Shafer, “Using color to seperate reflection components,” Color Research and Applications, vol. 10, no. 4, pp. 210–218, 1985.
  24. 24.G. D. Finlayson, S. D. Hordley, and R. Xu, “Convex programming colour constancy with a diagonal-offset model,” in IEEE International Conference on Image Processing, 2005, pp. 948–951.
  25. 25.T. Gevers, J. van de Weijer, and H. Stokman, Color image processing: methods and applications: color feature detection: an overview. CRC press, 2006, ch. 9, pp. 203–226.
  26. 26.F. Mindru, T. Tuytelaars, L. Van Gool, and T. Moons, “Moment invariants for recognition under changing viewpoint and illumination,” Computer Vision and Image Understanding, vol. 94, no. 1-3, pp. 3–27, 2004.
  27. 27.J. Matas, O. Chum, M. Urban, and T. Pajdla, “Robust wide-baseline stereo from maximally stable extremal regions,” Image and Vision Computing, vol. 22, no. 10, pp. 761 – 767, 2004.
  28. 28.P.-E. Forssén, “Maximally stable colour regions for recognition and matching,” in IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, USA, June 2007.
  29. 29.J. Sivic and A. Zisserman, “Video Google: A text retrieval approach to object matching in videos,” in IEEE International Conference on Computer Vision, Nice, France, 2003, pp. 1470–1477.
  30. 30.T. K. Leung and J. Malik, “Representing and recognizing the visual appearance of materials using three-dimensional textons,” International Journal of Computer Vision, vol. 43, no. 1, pp. 29–44, 2001.
  31. 31.R. Fergus, P. Perona, and A. Zisserman, “Object class recognition by unsupervised scale-invariant learning,” in IEEE Conference on Computer Vision and Pattern Recognition, vol. 2, 2003, pp. 264–271.
  32. 32.F. Jurie and B. Triggs, “Creating efficient codebooks for visual recognition.” in IEEE International Conference on Computer Vision, Beijing, China, 2005, pp. 604–610.
  33. 33.B. Leibe and B. Schiele, “Interleaved object categorization and segmentation,” in British Machine Vision Conference, Norwich, UK, 2003, pp. 759–768.
  34. 34.C.-C. Chang and C.-J. Lin, LIBSVM: a library for support vector machines, 2001, software available at http://www.csie.ntu.edu.tw/∼cjlin/libsvm.
  35. 35.M. Naphade, J. R. Smith, J. Tesic, S.-F. Chang, W. Hsu, L. Kennedy, A. Hauptmann, and J. Curtis, “Large-scale concept ontology for multimedia,” IEEE Multimedia, vol. 13, no. 3, pp. 86–91, 2006.
  36. 36.C. M. Bishop, Pattern Recognition and Machine Learning. Springer, August 2006.
  37. 37.B. Efron, “Bootstrap methods: Another look at the jackknife,” Annals of Statistics, vol. 7, pp. 1–26, 1979.
  38. 38.M. Marszałek, C. Schmid, H. Harzallah, and J. van de Weijer, “Learning object representations for visual object class recognition,” 2007, Visual Recognition Challenge workshop, in conjunction with IEEE International Conference on Computer Vision, Rio de Janeiro, Brazil. [Online]. Available: http://lear.inrialpes.fr/pubs/2007/MSHV07
  39. 39.J. C. van Gemert, J.-M. Geusebroek, C. J. Veenman, C. G. M. Snoek, and A. W. M. Smeulders, “Robust scene categorization by learning image statistics in context,” in CVPR Workshop on Semantic Learning Applications in Multimedia (SLAM), 2006.
  40. 40.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “The PASCAL Visual Object Classes Challenge 2008 (VOC2008) Results.” [Online]. Available: http://www.pascal-network.org/challenges/VOC/voc2008/
  41. 41.M. A. Tahir, K. E. A. van de Sande, J. R. R. Uijlings, and et al. , “University of Amsterdam and University of Surrey at PASCAL VOC 2008,” 2008, PASCAL Visual Object Classes Challenge Workshop, in conjunction with IEEE European Conference on Computer Vision, Marseille, France. [Online]. Available: http://staff.science.uva.nl/∼ksande/pub/vandesande-pascalvoc2008.pdf
  42. 42.J. C. van Gemert, C. J. Veenman, A. W. M. Smeulders, and J.-M. Geusebroek, “Visual word ambiguity,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2009.
  43. 43.C. G. M. Snoek, K. E. A. van de Sande, O. de Rooij, B. Huurnink, J. C. van Gemert, J. R. R. Uijlings, and et al. , “The MediaMill TRECVID 2008 semantic video search engine,” in Proceedings of the 6th TRECVID Workshop, Gaithersburg, USA, November 2008.

Citation

MLA
van de Sande, K. E. A., et al. “Evaluating Color Descriptors for Object and Scene Recognition”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 32, no. 9, 2010, pp. 1582–96, https://doi.org/10.1109/TPAMI.2009.154.
APA
van de Sande, K. E. A., Gevers, T., & Snoek, C. G. M. (2010). Evaluating Color Descriptors for Object and Scene Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(9), 1582–1596. https://doi.org/10.1109/TPAMI.2009.154
Chicago
van de Sande, K. E. A., T. Gevers, and C. G. M. Snoek. 2010. “Evaluating Color Descriptors for Object and Scene Recognition”. IEEE Transactions on Pattern Analysis and Machine Intelligence 32 (9): 1582–96. https://doi.org/10.1109/TPAMI.2009.154.
Harvard
van de Sande, K.E.A., Gevers, T. and Snoek, C.G.M. (2010) “Evaluating Color Descriptors for Object and Scene Recognition”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(9), pp. 1582–1596. Available at: https://doi.org/10.1109/TPAMI.2009.154.
Vancouver
1. van de Sande KEA, Gevers T, Snoek CGM (2010) Evaluating Color Descriptors for Object and Scene Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 32:1582–1596

BibTeX

@article{van_de_Sande_2010, title={Evaluating Color Descriptors for Object and Scene Recognition}, volume={32}, ISSN={0162-8828}, url={http://dx.doi.org/10.1109/TPAMI.2009.154}, DOI={10.1109/tpami.2009.154}, number={9}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={van de Sande, Koen E A and Gevers, T and Snoek, Cees G M}, year={2010}, month=Sept, pages={1582–1596} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF