Food-101 - Mining Discriminative Components with Random Forests

Lukas BossardMatthieu GuillauminLuc Van Gool

article2014ECCV3,528 citations

Introduces the Food-101 benchmark dataset alongside a Random Forest framework that mines discriminative superpixel components across all classes simultaneously to achieve efficient and accurate food recognition.

Listen

The paper presents a method for automatically recognizing images of food dishes, a task complicated by high visual variability in ingredients, lighting, viewpoint, and preparation, with no consistent global layout to exploit. Food photography is widespread in social media and mobile health applications, yet existing systems depend on manual labeling by experts or crowdsourcing, limiting scalability. The work introduces both a new weakly supervised mining technique and a large public benchmark to advance practical solutions.

The authors developed a Random Forest framework that jointly mines discriminative local image regions, called components, across all categories by clustering superpixels rather than exhaustively sampling sliding windows. They created the Food-101 dataset of 101,000 real-world images spanning 101 categories, with noisy training labels drawn from a photo-sharing site and cleaned test images. Classification proceeds by scoring a small number of superpixels with linear SVMs trained on the mined components, followed by spatial pooling and a final multi-class SVM.

On Food-101 the approach reached 50.76 percent average accuracy, exceeding Improved Fisher Vector classification by 11.88 percentage points and a recent mid-level discriminative superpixel method by 8.13 points. It proved robust to parameter choices such as tree depth and number of components per class, and delivered competitive results on the MIT-Indoor scene dataset while evaluating far fewer regions than sliding-window alternatives. A convolutional neural network still led by roughly 5.6 points, at the cost of six days of training on a high-end GPU.

These gains matter because they enable more accurate, efficient recognition that can support automatic photo organization and calorie tracking without expert intervention. The superpixel restriction sharply reduces test-time computation compared with thousands of sliding windows, making deployment on mobile devices more feasible. The dataset itself provides a challenging, realistic benchmark that highlights limitations of global descriptors on food imagery.

The method is generic enough to apply to other fine-grained classification tasks. Further gains would require either stronger local features, larger training sets, or hybrid models that combine the efficiency of mined components with the representational power of deep networks. Results on noisy web data indicate the approach tolerates label noise, yet performance on the cleanest subsets and on visually similar classes remains well below human levels, so additional labeled data or refined component selection would be needed before reliable large-scale use.

Cover for Food-101 - Mining Discriminative Components with Random Forests

Abstract

Abstract. In this paper we address the problem of automatically recognizing pictured dishes. To this end, we introduce a novel method to mine discriminative parts using Random Forests (RF), which allows us to mine for parts simultaneously for all classes and to share knowledge among them. To improve efficiency of mining and classification, we only consider patches that are aligned with image superpixels, which we call components. To measure the performance of our RF component mining for food recognition, we introduce a novel and challenging dataset of 101 food categories, with 101'000 images. With an average accuracy of 50.76%, our model outperforms alternative classification methods except for CNN, including SVM classification on Improved Fisher Vectors and existing discriminative part-mining algorithms by 11.88% and 8.13%, respectively. On the challenging MIT-Indoor dataset, our method compares nicely to other s-o-a component-based classification methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Dataset: Food-101
  • 4 Random Forest Component Mining
  • 4.1 Candidate Component Generation
  • 4.2 Mining Components
  • 4.3 Training Component Models
  • 4.4 Recognition from Mined Components
  • 5 Experimental Evaluation
  • 5.1 Implementation Details
  • 5.2 Influence of Parameters for Component Mining
  • 5.3 Comparison on Food-101
  • 5.4 Results on MIT-Indoor
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Food-101 Benchmark Dataset

    definition

    The Food-101 dataset is an image classification benchmark consisting of 101,000 real-world food images distributed equally across 101 categories (e.g., apple pie, bibimbap, chocolate mousse, dumplings, edamame, macarons, mussels, pad thai, risotto, sashimi).

    • Composition: Each class contains exactly 1,000 images, partitioned into 750 training images and 250 test images.
    • Data Collection and Annotation: Images were downloaded from foodspotting.com. The 250 test images per category are manually inspected and cleaned. In contrast, the 750 training images per category were intentionally left uncleaned and weakly labeled, preserving real-world label noise, intense colors, and variations in lighting, exposure, and perspective.
    • Image Preprocessing: All images are rescaled to have a maximum side length of 512 pixels; images smaller than 512 pixels were excluded from the dataset.
  2. Knowl 2 — Random Forest Discriminative Component Mining (RFDC) Framework

    model/method

    Random Forest Discriminant Components (RFDC) is a weakly-supervised part-mining and classification framework designed to discover discriminative visual regions (components) aligned with image superpixels across all object classes simultaneously.

    The end-to-end framework operates in four stages:

    1. Superpixel Segmentation and Feature Extraction: Images are segmented into graph-based superpixels (parameterized by smoothing σ=0.1\sigma = 0.1, scale parameter k=300k = 300, and a minimum superpixel area of 1% of the image, producing approximately 30 superpixels per image). For each superpixel, dense Speeded Up Robust Features (SURF) and L∗a∗b∗L^*a^*b^* color values are extracted from the superpixel region and its bounding box context. Both feature channels are encoded using Improved Fisher Vectors (IFV) with a 64-mode Gaussian Mixture Model (GMM), whitened with PCA, and concatenated into an 8,576-dimensional descriptor per superpixel.
    2. Candidate Component Generation: A multi-class Random Forest is trained on weakly-labeled superpixels (each inheriting its image's category label). Tree nodes recursively split the sample space by training linear Support Vector Machines (SVMs) on random binary class partitions to maximize information gain.
    3. Component Selection and Mining: The leaf nodes of the forest define candidate clusters. For each class yy, leaves are ranked using a distinctiveness score based on validation set statistics. A diversity filter based on non-maxima suppression eliminates near-duplicate clusters that share more than 50% of superpixels with higher-scoring leaves, yielding the top NN visual components per class.
    4. Classification: The Random Forest is discarded after mining. For each selected component, a dedicated linear binary SVM is trained using hard-negative mining. Superpixels of a query image are scored by all K×NK \times N component SVMs, aggregated into spatial representations via a 3-level spatial pyramid with loose spatial assignment, and classified by a multi-class structured-output SVM.
  3. Knowl 3 — Candidate Component Tree Node Splitting Criterion

    equation

    In the RFDC framework, each decision tree TtT_t in a Random Forest F={Tt}\mathcal{F} = \{T_t\} is grown on a random bootstrap sample of superpixels S={si=(xi,yi)}S = \{s_i = (\mathbf{x}_i, y_i)\}, where xi∈Rd\mathbf{x}_i \in \mathbb{R}^d is the feature vector of superpixel sis_i and yi∈{1,…,K}y_i \in \{1, \dots, K\} is the category label of the image containing sis_i.

    At each internal node, candidate decision functions ϕ(x):Rd→{0,1}\phi(\mathbf{x}): \mathbb{R}^d \to \{0, 1\} are generated by randomly assigning the present classes to two pseudo-classes and training a linear binary Support Vector Machine: ϕ(x)=1[wTx+b>0]\phi(\mathbf{x}) = \mathbf{1}[\mathbf{w}^T \mathbf{x} + b > 0] where w∈Rd\mathbf{w} \in \mathbb{R}^d is the weight vector, b∈Rb \in \mathbb{R} is the bias term, and 1[⋅]\mathbf{1}[\cdot] denotes the indicator function. The function routes sample x\mathbf{x} to the right child subset SrS_r if wTx+b>0\mathbf{w}^T \mathbf{x} + b > 0, and to the left child subset SlS_l otherwise.

    From a set of 100 randomly sampled binary SVM partitions, the node selects the decision function that maximizes the information gain criterion: I(S,ϕ)=H(S)−(∣Sl∣∣S∣H(Sl)+∣Sr∣∣S∣H(Sr))I(S, \phi) = H(S) - \left( \frac{|S_l|}{|S|} H(S_l) + \frac{|S_r|}{|S|} H(S_r) \right) where H(S)H(S) is the Shannon class entropy of sample set SS: H(S)=−∑k=1Kp(y=k∣S)log⁡2p(y=k∣S)H(S) = -\sum_{k=1}^K p(y=k \mid S) \log_2 p(y=k \mid S) Node splitting stops when a maximum tree depth is reached, when a node contains fewer than 25 samples, or when all samples in the node belong to a single class.

  4. Knowl 4 — Leaf Distinctiveness Metric and Diversity Filtering for Component Mining

    equation

    Let L=⋃tLt\mathcal{L} = \bigcup_t \mathcal{L}_t be the set of all leaf nodes across all trees in the trained Random Forest F\mathcal{F}, where each leaf l∈Ll \in \mathcal{L} carries an empirical class posterior distribution p(y∣l)p(y \mid l) computed from the training samples that reached ll.

    For a validation set of superpixels, let δl,s=1\delta_{l,s} = 1 if sample ss reaches leaf ll, and δl,s=0\delta_{l,s} = 0 otherwise. The class confidence score of sample ss for class yy is computed across the forest by averaging leaf posteriors: p(y∣s)=1∣F∣∑l∈Lδl,sp(y∣l)p(y \mid s) = \frac{1}{|\mathcal{F}|} \sum_{l \in \mathcal{L}} \delta_{l,s} p(y \mid l) where ∣F∣|\mathcal{F}| is the total number of trees in the forest, ensuring ∑l∈Lδl,s=∣F∣\sum_{l \in \mathcal{L}} \delta_{l,s} = |\mathcal{F}|.

    The distinctiveness of a leaf ll with respect to category yy is defined as: distinctiveness(l∣y)=∑sδl,sp(y∣s)\text{distinctiveness}(l \mid y) = \sum_{s} \delta_{l,s} p(y \mid s) where the summation runs over all validation samples ss.

    Diversity Filtering via Non-Maxima Suppression: Leaves are ranked in descending order of distinctiveness(l∣y)\text{distinctiveness}(l \mid y). Starting from the highest-ranked leaf, candidate leaves are evaluated sequentially: any leaf whose assigned superpixel set overlaps by more than 50% with the superpixels of any already accepted higher-scoring leaf is discarded. The top NN surviving leaves are retained as the mined visual components for category yy.

  5. Knowl 5 — Image Classification via Superpixel Component Scoring and Spatial Pooling

    model/method

    After identifying NN component leaves for each of the KK classes, dedicated component detectors are trained and used for image-level classification:

    1. Component Model Training: For each selected leaf representing class yy, a linear binary SVM is trained. The positive training set consists of the most confident samples of class yy within that leaf cluster (e.g., top 100 to 500 samples), and negatives are gathered from a large negative repository using iterative hard-negative mining. After training these SVMs, the Random Forest is completely discarded.
    2. Superpixel Scoring: For a query image segmented into superpixels {si}\{s_i\}, each superpixel is scored by all K×NK \times N component SVMs, yielding a confidence score vector c(si)∈RK×N\mathbf{c}(s_i) \in \mathbb{R}^{K \times N}.
    3. Spatial Pyramid Pooling: A 3-level spatial pyramid partitions the image into spatial regions. Loose spatial assignment is applied: each superpixel fully contributes its score vector c(si)\mathbf{c}(s_i) to every spatial region that it intersects. Scores within each spatial region are averaged across its assigned superpixels.
    4. Final Category Classification: The pooled score vectors from all pyramid cells are concatenated and passed to a structured-output multi-class SVM, trained using the cutting-plane algorithm, to predict the final image class.
  6. Knowl 6 — Classification Performance Comparison on the Food-101 Benchmark

    data/table

    Classification accuracy comparison on the Food-101 benchmark dataset (101 classes, 750 training images and 250 test images per class). Evaluated methods include global image classification pipelines, direct Random Forest classifiers, local part-mining methods using 20 components per class, and deep convolutional neural networks.

    Method Avg. Accuracy [%]
    Global Approaches
    Bag-of-Words (BoW) 28.51
    Improved Fisher Vectors (IFV) 38.88
    Convolutional Neural Network (CNN) 56.40
    Local / Part-Based Approaches
    Random Forest (RF direct) 32.72
    Randomized Clustering Forests (RCF) 28.46
    Mid-Level Discriminative Superpixels (MLDS) 42.63
    Random Forest Discriminant Components (RFDC) 50.76

    Key Comparisons and Observations:

    • RFDC achieves 50.76% average accuracy, outperforming the global Improved Fisher Vector (IFV) baseline by 11.88% and the sliding-window-adapted part-mining method MLDS by 8.13%.
    • Using the Random Forest directly for classification (RF) or as a feature generator (RCF) yields substantially lower accuracy (32.72% and 28.46%), proving the value of the component mining and dedicated SVM training stages.
    • Deep CNN sets the highest accuracy at 56.40% (+5.64% over RFDC), at the cost of extensive training time (6 days on an NVIDIA Tesla K20X GPU).
  7. Knowl 7 — Classification Performance on the MIT-Indoor Scene Benchmark

    data/table

    Classification results on the MIT-Indoor scene categorization dataset (67 scene categories), evaluated under both the standard restricted training protocol (~80 training and 20 testing images per class) and the full training set protocol. RFDC evaluates 50 mined components per class on approximately 30 superpixels per image.

    Method Avg. Accuracy [%]
    Part-Based Baselines
    HOG Patches 38.10
    Bag of Parts (BoP) 46.10
    MI-SVM 46.40
    MMDL 50.15
    D-Parts 51.40
    Discriminative Mode Seeking (DMS) 64.03
    Part-Based (RFDC)
    RFDC (restricted train set) 54.40
    RFDC (full train set) 58.36
    Global or Combined
    IFV 60.77
    IFV + BoP 63.10
    IFV + DMS 66.87

    Key Computational and Empirical Findings:

    • RFDC achieves 54.40% accuracy on the restricted set and 58.36% on the full train set, outperforming part-based baselines including HOG Patches, BoP, MI-SVM, MMDL, and D-Parts.
    • Efficiency: RFDC evaluates only 50 components on ~30 superpixels per test image, requiring ~0.8 seconds per test image on 8 CPU cores (70% encoding, 25% component SVM evaluation), compared to thousands of sliding windows in competing part-based approaches.
    • Full pipeline training uses approximately 250 CPU hours (55% tree training, 15% component model training, 20% final structured SVM classifier training).
  8. Knowl 8 — Superpixel Feature Representation and Encoding Ablation for RFDC

    data/table

    Ablation study comparing the classification accuracy of RFDC on the Food-101 dataset under different visual feature descriptors (HOG, dense SURF, L∗a∗b∗L^*a^*b^* color) and encoding schemes (Bag-of-Words [BoW] versus Improved Fisher Vectors [IFV]) with varying codebook/dictionary sizes (@K@K).

    Encoding Features Avg. Accuracy [%]
    HOG 8.85
    BoW SURF @ 1024 33.47
    BoW SURF @ 1024 + Color @ 256 38.83
    IFV Color @ 64 14.24
    IFV SURF @ 64 44.79
    IFV SURF @ 64 + Color @ 64 49.40

    Key Insights:

    • HOG features perform poorly on food recognition (8.85%), because food categories lack rigid spatial contours and rely primarily on color and texture distributions.
    • Dense SURF with IFV encoding outperforms BoW by more than 11% (44.79% vs. 33.47%).
    • Combining texture (SURF) and color (L∗a∗b∗L^*a^*b^*) channels provides consistent accuracy gains across both BoW (+5.36%) and IFV (+4.61%), achieving the highest accuracy at 49.40%.
  9. Knowl 9 — Hyperparameter Sensitivity of Random Forest Component Mining

    empirical result

    Analysis of the hyperparameters governing RFDC component mining on Food-101 demonstrates consistent robustness across parameter variations:

    • Number of Trees: Varying the forest size from 10 to 30 trees results in minimal accuracy variation (from ~48.8% to ~49.4%), showing that large ensembles are not required to mine effective components.
    • Tree Depth: Accuracy rises sharply up to depth 4, peaks between depth 5 and 7 (~49.4%), and gradually declines at deeper levels (depth ≥9\ge 9) as leaf sample counts become too small to maintain generalizable statistics.
    • Positive Training Samples per Component: Classification accuracy improves with the number of top confident leaf samples used to train each component SVM, reaching a performance plateau around 200 positive samples (plateauing near 50.0%). Training on 200 samples per model significantly reduces training time with negligible performance loss compared to larger sample pools.
    • Number of Components per Class (NN): Accuracy increases with component count (N=1N=1 achieves ~43.8%, N=10N=10 achieves ~46.9%, N=20N=20 achieves ~49.4%, N=30N=30 achieves ~50.4%, N=50N=50 achieves ~50.8%). Beyond N=20N=20, classification gains saturate while feature dimensionality scales linearly from 42,420 dimensions (N=20N=20) to 106,050 dimensions (N=50N=50), substantially increasing memory and computation overhead.

Coverage note — Qualitative component heat maps (Figure 8) and individual per-class accuracy breakdowns (Figure 6) were omitted as their conclusions are fully reflected in the aggregate benchmark tables and empirical analysis.

References

  1. 1.Arandjelović, R., Zisserman, A.: Three things everyone should know to improve object retrieval. In: CVPR (2012)
  2. 2.Bay, H., Tuytelaars, T., Van Gool, L.: SURF: Speeded Up Robust Features. In: ICCV (2006)
  3. 3.Bosch, A., Zisserman, A., Munoz, X.: Image Classification using Random Forests and Ferns. In: ICCV (2007)
  4. 4.Breiman, L.: Random forests. Machine Learning (2001)
  5. 5.Chen, M., Dhingra, K., Wu, W., Yang, L., Sukthankar, R., Yang, J.: PFID: Pittsburgh fast-food image dataset. In: ICIP (2009)
  6. 6.Chen, M.y., Yang, Y.h., Ho, C.j., Wang, S.h., Liu, S.m., Chang, E., Yeh, C.h., Ouhyoung, M.: Automatic Chinese food identification and quantity estimation. In: SIGGRAPH Asia 2012 Technical Briefs (2012)
  7. 7.Doersch, C., Gupta, A., Efros, A.A.: Mid-level visual element discovery as discriminative mode seeking. In: NIPS (2013)
  8. 8.Endres, I., Shih, K., Jiaa, J., Hoiem, D.: Learning Collections of Part Models for Object Recognition. In: CVPR (2013)
  9. 9.Felzenszwalb, P.F., Girshick, R., McAllester, D., Ramanan, D.: Object detection with discriminatively trained part based models. PAMI (2010)
  10. 10.Felzenszwalb, P.F., Huttenlocher, D.P.: Efficient Graph-Based Image Segmentation. IJCV (2004)
  11. 11.Gall, J., Yao, A., Razavi, N., Van Gool, L., Lempitsky, V.: Hough forests for object detection, tracking, and action recognition. PAMI (2011)
  12. 12.Girshick, R., Donahue, J., Darrell, T., Malik, J.: Rich feature hierarchies for accurate object detection and semantic segmentation. In: CVPR (2014)
  13. 13.Hariharan, B., Malik, J., Ramanan, D.: Discriminative decorrelation for clustering and classification. In: ECCV (2012)
  14. 14.Ho, T.K.: Random decision forests. In: ICDAR (1995)
  15. 15.Hoashi, H., Joutou, T., Yanai, K.: Image Recognition of 85 Food Categories by Feature Fusion. In: ISM (2010)
  16. 16.Jia, Y.: Caffe: An open source convolutional architecture for fast feature embedding. http://caffe.berkeleyvision.org/ (2013)
  17. 17.Joachims, T., Finley, T., Yu, C.N.J.: Cutting-plane training of structural SVMs. Machine Learning (2009)
  18. 18.Joutou, T., Yanai, K.: A food image recognition system with Multiple Kernel Learning. In: ICIP (2009)
  19. 19.Juneja, M., Vedaldi, A., Jawahar, C., Zisserman, A.: Blocks That Shout: Distinctive Parts for Scene Classification. In: CVPR (2013)
  20. 20.Kawano, Y., Yanai, K.: Real-Time Mobile Food Recognition System. In: IEEE Conference on Computer Vision and Pattern Recognition Workshops (2013)
  21. 21.King, D.E.: Dlib-ml: A machine learning toolkit. JMLR (2009)
  22. 22.Kontschieder, P., Rota Bul`o, S., Bischof, H., Pelillo, M.: Structured class-labels in random forests for semantic image labelling. In: ICCV (2011)
  23. 23.Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: NIPS (2012)
  24. 24.Lazebnik, S., Schmid, C., Ponce, J.: Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. In: CVPR (2006)
  25. 25.Li, Q., Wu, J., Tu, Z.: Harvesting mid-level visual concepts from large-scale internet images. In: CVPR (2013)
  26. 26.Malisiewicz, T., Gupta, A., Efros, A.A.: Ensemble of exemplar-svms for object detection and beyond. In: ICCV (2011)
  27. 27.Martin, C., Correa, J., Han, H., Allen, H., Rood, J., Champagne, C., Gunturk, B., Bray, G.: Validity of the remote food photography method (RFPM) for estimating energy and nutrient intake in near real-time. Obesity (2011)
  28. 28.Matsuda, Y., Hoashi, H., Yanai, K.: Multiple-Food Recognition Considering Co-occurrence Employing Manifold Ranking. In: ICPR (2012)
  29. 29.Moosmann, F., Nowak, E., Jurie, F.: Randomized clustering forests for image classification. PAMI (2008)
  30. 30.Noronha, J., Hysen, E., Zhang, H., Gajos, K.Z.: Platemate: crowdsourcing nutritional analysis from food photographs. In: ACM Symposium on UI Software and Technology (2011)
  31. 31.Quattoni, A., Torralba, A.: Recognizing indoor scenes. In: CVPR (2009)
  32. 32.Sánchez, J., Perronnin, F., Mensink, T., Verbeek, J.: Image Classification with the Fisher Vector: Theory and Practice. IJCV (2013)
  33. 33.Shotton, J., Johnson, M., Cipolla, R.: Semantic texton forests for image categorization and segmentation. In: CVPR (2008)
  34. 34.Singh, S., Gupta, A., Efros, A.A.: Unsupervised discovery of mid-level discriminative patches. In: ECCV (2012)
  35. 35.Sun, J., Ponce, J.: Learning discriminative part detectors for image classification and cosegmentation. In: ICCV (2013)
  36. 36.Uijlings, J.R.R., van de Sande, K.E.A., Gevers, T., Smeulders, A.W.M.: Selective search for object recognition. IJCV (2013)
  37. 37.Vedaldi, A., Fulkerson, B.: VLFeat: An open and portable library of computer vision algorithms. http://www.vlfeat.org/ (2008)
  38. 38.Wang, X., Wang, B., Bai, X., Liu, W., Tu, Z.: Max-margin multiple-instance dictionary learning. In: NIPS (2013)
  39. 39.Yang, S.L., Chen, M., Pomerleau, D., Sukthankar, R.: Food recognition using statistics of pairwise local features. In: CVPR (2010)
  40. 40.Yao, B., Khosla, A., Fei-Fei, L.: Combining randomization and discrimination for fine-grained image categorization. In: CVPR (2011)

Citation

MLA
Bossard, L., et al. “Food-101 – Mining Discriminative Components with Random Forests”. Lecture Notes in Computer Science, Springer International Publishing, 2014, pp. 446–61, https://doi.org/10.1007/978-3-319-10599-4_29.
APA
Bossard, L., Guillaumin, M., & Van Gool, L. (2014). Food-101 – Mining Discriminative Components with Random Forests. In Lecture Notes in Computer Science (pp. 446–461). Springer International Publishing. https://doi.org/10.1007/978-3-319-10599-4_29
Chicago
Bossard, L., M. Guillaumin, and L. Van Gool. 2014. “Food-101 – Mining Discriminative Components with Random Forests”. In Lecture Notes in Computer Science. Springer International Publishing. https://doi.org/10.1007/978-3-319-10599-4_29.
Harvard
Bossard, L., Guillaumin, M. and Van Gool, L. (2014) “Food-101 – Mining Discriminative Components with Random Forests”, Lecture Notes in Computer Science. Springer International Publishing, pp. 446–461. Available at: https://doi.org/10.1007/978-3-319-10599-4_29.
Vancouver
1. Bossard L, Guillaumin M, Van Gool L (2014) Food-101 – Mining Discriminative Components with Random Forests. In: Lecture Notes in Computer Science. Springer International Publishing, pp 446–461

BibTeX

@inbook{Bossard_2014, title={Food-101 – Mining Discriminative Components with Random Forests}, ISBN={9783319105994}, ISSN={1611-3349}, url={http://dx.doi.org/10.1007/978-3-319-10599-4_29}, DOI={10.1007/978-3-319-10599-4_29}, booktitle={Computer Vision – ECCV 2014}, publisher={Springer International Publishing}, author={Bossard, Lukas and Guillaumin, Matthieu and Van Gool, Luc}, year={2014}, pages={446–461} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF