Food-101 - Mining Discriminative Components with Random Forests
Lukas BossardMatthieu GuillauminLuc Van Gool
Introduces the Food-101 benchmark dataset alongside a Random Forest framework that mines discriminative superpixel components across all classes simultaneously to achieve efficient and accurate food recognition.
The paper presents a method for automatically recognizing images of food dishes, a task complicated by high visual variability in ingredients, lighting, viewpoint, and preparation, with no consistent global layout to exploit. Food photography is widespread in social media and mobile health applications, yet existing systems depend on manual labeling by experts or crowdsourcing, limiting scalability. The work introduces both a new weakly supervised mining technique and a large public benchmark to advance practical solutions.
The authors developed a Random Forest framework that jointly mines discriminative local image regions, called components, across all categories by clustering superpixels rather than exhaustively sampling sliding windows. They created the Food-101 dataset of 101,000 real-world images spanning 101 categories, with noisy training labels drawn from a photo-sharing site and cleaned test images. Classification proceeds by scoring a small number of superpixels with linear SVMs trained on the mined components, followed by spatial pooling and a final multi-class SVM.
On Food-101 the approach reached 50.76 percent average accuracy, exceeding Improved Fisher Vector classification by 11.88 percentage points and a recent mid-level discriminative superpixel method by 8.13 points. It proved robust to parameter choices such as tree depth and number of components per class, and delivered competitive results on the MIT-Indoor scene dataset while evaluating far fewer regions than sliding-window alternatives. A convolutional neural network still led by roughly 5.6 points, at the cost of six days of training on a high-end GPU.
These gains matter because they enable more accurate, efficient recognition that can support automatic photo organization and calorie tracking without expert intervention. The superpixel restriction sharply reduces test-time computation compared with thousands of sliding windows, making deployment on mobile devices more feasible. The dataset itself provides a challenging, realistic benchmark that highlights limitations of global descriptors on food imagery.
The method is generic enough to apply to other fine-grained classification tasks. Further gains would require either stronger local features, larger training sets, or hybrid models that combine the efficiency of mined components with the representational power of deep networks. Results on noisy web data indicate the approach tolerates label noise, yet performance on the cleanest subsets and on visually similar classes remains well below human levels, so additional labeled data or refined component selection would be needed before reliable large-scale use.
- Paper: Object Detection with Discriminatively Trained Part-Based Models, Pedro F. Felzenszwalb et al. (2010). This paper establishes the foundational part-based object detection framework and latent SVM mining techniques that the source paper adapts and extends for food recognition.
- Paper: Visual categorization with bags of keypoints, Gabriella Csurka et al. (2004). Reading this paper provides the bag-of-features and vocabulary quantization background assumed by traditional component-based classification approaches.
- Paper: Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories, Svetlana Lazebnik et al. (2006). This spatial pyramid matching technique establishes core image partitioning and histogram representation methods that inform part-based visual recognition pipelines.
- Paper: Efficient Graph-Based Image Segmentation, PEDRO F. FELZENSZWALB et al. (2004). Understanding this foundational graph-based image segmentation algorithm helps explain how image superpixels and components are generated for efficient part mining.
- Paper: SLIC Superpixels Compared to State-of-the-Art Superpixel Methods, Radhakrishna Achanta et al. (2012). This superpixel benchmark provides the boundary-adherence methodology necessary for understanding superpixel-aligned component extraction in food image analysis.
- Paper: Rich feature hierarchies for accurate object detection and semantic segmentation, Ross Girshick et al. (2014). This paper demonstrates how convolutional neural networks superseded traditional part-mining and handcrafted feature hierarchies for accurate object detection.
- Paper: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Shaoqing Ren et al. (2015). This work directly continues the evolution of object detection by introducing region proposal networks to replace slower component mining strategies.
- Paper: CNN Features Off-the-Shelf: An Astounding Baseline for Recognition, Ali Sharif Razavian et al. (2014). This study extends image classification baselines by evaluating off-the-shelf CNN features across diverse benchmarks, including the Food-101 dataset introduced by the source.
- Paper: Fast R-CNN, Ross B. Girshick (2015). This paper follows up on region-based detection architectures by streamlining multi-stage pipelines into a single fast shared-convolution framework.
