Part-Based R-CNNs for Fine-Grained Category Detection

Ning ZhangJeff DonahueRoss B. GirshickTrevor Darrell

article2014ECCV1,323 citations

Develops a part-based R-CNN framework that localizes semantic parts and enforces geometric constraints across region proposals, achieving state-of-the-art fine-grained visual recognition without relying on test-time bounding box annotations.

Listen

Visual fine-grained categorization—distinguishing between closely related subcategories such as specific animal species, product variants, or plant types—presents a major challenge for automated systems because visual differences are often subtle and dependent on object pose. Traditional solutions rely heavily on isolating specific semantic parts (such as the head or body) to capture these nuances, but they suffer from a major operational bottleneck: they require manual bounding box annotations around the target object during testing because automated part localization has historically been unreliable.

The main objective of the article is to develop and evaluate an automated, end-to-end visual recognition model that jointly localizes whole objects and their semantic parts to achieve high fine-grained classification accuracy without requiring any manual bounding box input at test time.

The approach extends the Region-Based Convolutional Neural Network (R-CNN) framework by training deep learning detectors for both whole objects and their specific parts using candidate image regions generated via selective search. To ensure that localized parts remain physically plausible, the system applies learned geometric constraints that rescore part candidates based on spatial priors, including an appearance-based nearest-neighbor model. The overall model extracts feature descriptors from the localized regions using a fine-tuned convolutional neural network and classifies categories using a linear support vector machine. The evaluation was conducted on the standard Caltech-UCSD Birds benchmark, which comprises 11,788 images across 200 bird species.

The analysis yielded several key findings regarding accuracy and localization capabilities. In the realistic setting where object bounding boxes are unknown at test time, the proposed system achieved a 73.89% classification accuracy when fine-tuned, significantly outperforming prior baseline methods that scored around 44.94%. Even without fine-tuning, the fully automated model achieved 66.0% accuracy, matching or exceeding prior state-of-the-art methods that required manual bounding boxes. On part localization, the model achieved a 65% improvement over strong baseline models for detecting bird heads in an unassisted setting. In addition, incorporating geometric part constraints directly improved whole-object localization, lifting object-only recognition accuracy by over 11 percentage points compared to standard single-object detectors.

These findings demonstrate that automated, pose-normalized part detection can replace costly manual annotations in practical visual recognition workflows. Eliminating the requirement for human-provided bounding boxes substantially reduces operational overhead, latency, and integration friction for fine-grained computer vision applications. Furthermore, the results show that combining deep feature representations with geometric priors allows systems to remain robust against pose variations, background clutter, and partial occlusions.

Based on these results, organizations deploying fine-grained visual recognition should adopt joint object-and-part localization architectures rather than whole-image classifiers. For immediate technical enhancements, practitioners should explore dense window sampling or alternative region proposal methods, as the candidate generation step currently limits recall for smaller parts (dropping below 40% at higher spatial precision thresholds). Future initiatives should also evaluate weakly supervised approaches that automatically discover parts without requiring detailed manual part annotations during the training phase.

Confidence in these findings is high for structured benchmark conditions, but several practical limitations remain. The model relies on supervised part annotations during training, which can be expensive to collect across diverse domains. In addition, performance depends on candidate region generation quality and hyperparameter tuning for geometric priors. System performance should be validated carefully in operational environments with heavy occlusions or novel camera viewpoints before full-scale deployment.

Cover for Part-Based R-CNNs for Fine-Grained Category Detection

Abstract

Semantic part localization can facilitate fine-grained categorization by explicitly isolating subtle appearance differences associated with specific object parts. Methods for pose-normalized representations have been proposed, but generally presume bounding box annotations at test time due to the difficulty of object detection. We propose a model for fine-grained categorization that overcomes these limitations by leveraging deep convolutional features computed on bottom-up region proposals. Our method learns whole-object and part detectors, enforces learned geometric constraints between them, and predicts a fine-grained category from a pose-normalized representation. Experiments on the Caltech-UCSD bird dataset confirm that our method outperforms state-of-the-art fine-grained categorization methods in an end-to-end evaluation without requiring a bounding box at test time.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 2.1 Part-based models for detection and pose localization
  • 2.2 Fine-grained categorization
  • 2.3 Convolutional networks
  • 3 Part-based R-CNNs
  • 3.1 Geometric constraints
  • 3.2 Fine-grained categorization
  • 4 Evaluation
  • 4.1 Fine-grained categorization
  • 4.2 Part localization
  • 4.3 Component Analysis
  • 5 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Part-Based R-CNN Architecture

    model/method

    Part-Based R-CNN is an object detection and fine-grained categorization framework that eliminates the requirement of test-time ground truth bounding boxes by extending the Region-based Convolutional Neural Network (R-CNN) paradigm to semantic parts.

    The framework consists of four primary stages:

    1. Region Proposal Generation: Candidate bounding box proposals are generated for an input image using a bottom-up segmentation method (Selective Search).
    2. Whole-Object and Part Detection: Deep convolutional neural network (CNN) feature descriptors are extracted for each proposal. Independent linear Support Vector Machine (SVM) classifiers are trained for the whole object (root) and for each predefined semantic part.
    3. Joint Geometric Rescoring and Configuration Selection: Candidate object and part proposals are jointly selected by maximizing an objective function combining individual detector confidence scores and learned geometric layout constraints relative to the object root.
    4. Pose-Normalized Representation and Classification: Deep features are extracted from the localized whole object and parts. These features are concatenated into a single pose-normalized feature vector, which is classified by a final one-versus-all multiclass SVM to predict the fine-grained category.
  2. Knowl 2 — Joint Object and Part Localization Optimization

    equation

    Let X={x0,x1,…,xn}X = \{x_0, x_1, \dots, x_n\} denote the candidate bounding box locations of the whole object p0p_0 and nn semantic parts {pi}i=1n\{p_i\}_{i=1}^n in a test image. For each detector i∈{0,1,…,n}i \in \{0, 1, \dots, n\}, the detector probability score on a candidate region xx is defined as: di(x)=σ(wi⊤ϕ(x))=11+exp⁡(−wi⊤ϕ(x))d_i(x) = \sigma(w_i^\top \phi(x)) = \frac{1}{1 + \exp(-w_i^\top \phi(x))} where ϕ(x)∈RD\phi(x) \in \mathbb{R}^D is the CNN feature descriptor extracted at region xx, wi∈RDw_i \in \mathbb{R}^D is the learned linear SVM weight vector for category or part ii, and σ(⋅)\sigma(\cdot) is the standard logistic sigmoid function.

    The optimal joint object and part configuration X∗X^* is computed by solving: X∗=arg⁡max⁡XΔ(X)∏i=0ndi(xi)X^* = \arg\max_{X} \Delta(X) \prod_{i=0}^n d_i(x_i) where Δ(X)\Delta(X) is a geometric configuration scoring function that enforces spatial consistency between the object root box x0x_0 and part boxes x1,…,xnx_1, \dots, x_n.

    The basic bounding box inclusion constraint Δbox(X)\Delta_{\text{box}}(X) is defined as: Δbox(X)=∏i=1ncx0(xi)\Delta_{\text{box}}(X) = \prod_{i=1}^n c_{x_0}(x_i) where cx(y)={1if region y falls outside region x by at most ϵ pixels0otherwisec_{x}(y) = \begin{cases} 1 & \text{if region } y \text{ falls outside region } x \text{ by at most } \epsilon \text{ pixels} \\ 0 & \text{otherwise} \end{cases} with tolerance parameter ϵ=10\epsilon = 10 pixels.

  3. Knowl 3 — Geometric Constraint Scoring Models for Part Localization

    model/method

    To constrain part locations relative to the whole-object root bounding box and filter out false positive part detections caused by detector imperfection or occlusion, the geometric scoring function Δgeometric(X)\Delta_{\text{geometric}}(X) incorporates spatial prior distributions: Δgeometric(X)=Δbox(X)(∏i=1nδi(xi))α\Delta_{\text{geometric}}(X) = \Delta_{\text{box}}(X) \left( \prod_{i=1}^n \delta_i(x_i) \right)^\alpha where α≥0\alpha \ge 0 is a scalar hyperparameter that balances the detector confidence and geometric prior likelihood, and δi(xi)\delta_i(x_i) is a spatial density function for part pip_i given training data.

    Two definitions for δi(xi)\delta_i(x_i) are utilized:

    1. Parametric Mixture of Gaussians Prior (δiMG\delta^{\text{MG}}_i): A mixture of Gaussians with Ng=4N_g = 4 components is fit to the normalized 2D spatial displacements of part pip_i relative to the object root box across all training instances.
    2. Non-Parametric Appearance-Conditioned Prior (δiNP\delta^{\text{NP}}_i): Let x~0=arg⁡max⁡x0d0(x0)\tilde{x}_0 = \arg\max_{x_0} d_0(x_0) be the top-scoring root detection. The system finds the K=20K = 20 nearest neighbors in the training dataset to x~0\tilde{x}_0 based on cosine distance over deep convolutional features (specifically pool5 layer activations). A unimodal Gaussian distribution is fit to the relative positions of part pip_i in these KK retrieved nearest neighbors, and δiNP(xi)\delta^{\text{NP}}_i(x_i) is evaluated as the likelihood of xix_i under this neighbor-conditioned Gaussian.
  4. Knowl 4 — Part-Specific CNN Fine-Tuning and Feature Representation

    model/method

    To extract discriminative features tailored to fine-grained categorization, the feature representation combines domain-specific convolutional network fine-tuning and pose normalization.

    • Network Architecture: An AlexNet-equivalent network pretrained on ImageNet (1000 categories) is used as the base model.
    • Fine-Tuning Procedure: Separate CNN models are independently fine-tuned for the whole object and for each semantic part. The original 1000-way fc8 classification layer is replaced with a 200-way fc8 layer initialized from a Gaussian distribution with mean μ=0\mu = 0 and standard deviation σ=0.01\sigma = 0.01. Ground truth crops of the whole object and semantic parts are warped to 227×227227 \times 227 pixels with 16 pixels of image context added around each border. The global fine-tuning learning rate is initialized to one-tenth of the initial ImageNet learning rate and decreased by a factor of 10 during training, while the new fc8 classification layer uses 10 times the global learning rate.
    • Feature Concatenation: Features ϕ(xi)\phi(x_i) are extracted from the fc6 layer (4096 dimensions) of the CNN fine-tuned for component ii. The joint representation of an image is the concatenated vector [ϕ(x0),ϕ(x1),…,ϕ(xn)][\phi(x_0), \phi(x_1), \dots, \phi(x_n)]. If all candidate proposals for a specific part ii fall below the part detector's threshold, ϕ(xi)\phi(x_i) is set to a zero vector 0\mathbf{0}.
  5. Knowl 5 — CUB-200-2011 Experimental Setup and Training Protocol

    experimental setup

    Experiments are conducted on the Caltech-UCSD Birds 200-2011 (CUB200-2011) dataset, which contains 11,788 images across 200 bird species with approximately 30 training samples per species. The dataset includes annotations for whole-object bounding boxes and 15 keypoints (beak, back, breast, belly, forehead, crown, left eye, left leg, left wing, right eye, right leg, right wing, tail, nape, throat). The 15 keypoints are grouped into n=2n=2 semantic parts: head and body.

    • Proposals: Candidate windows are generated using Selective Search.
    • Part Detector SVM Training: Region proposals with Intersection over Union (IoU) overlap ≥0.7\ge 0.7 with a ground truth object or part box are positive training examples; proposals with IoU overlap ≤0.3\le 0.3 with all ground truth regions are negative training examples.
    • Evaluation Protocols:
      1. Bounding Box Given: The ground-truth whole-object bounding box is provided at test time.
      2. Bounding Box Unknown: Fully automatic evaluation where neither object nor part bounding boxes are provided at test time.
  6. Knowl 6 — Fine-Grained Classification Performance on CUB-200-2011

    data/table

    Classification accuracy on CUB200-2011 under both test-time bounding box settings is shown below. Models with -ft utilize feature extraction from part-specific fine-tuned CNN models. Oracle denotes evaluating classification using ground-truth bounding box and part annotations at test time.

    Method Accuracy
    Bounding Box Given
    DPD 50.98%
    DPD+DeCAF feature 64.96%
    POOF 56.78%
    Symbiotic Segmentation 59.40%
    Alignment 62.70%
    Oracle 72.83%
    Oracle-ft 82.02%
    Ours (Δbox\Delta_{\text{box}}) 67.55%
    Ours (Δgeometric\Delta_{\text{geometric}} with δMG\delta^{\text{MG}}) 67.98%
    Ours (Δgeometric\Delta_{\text{geometric}} with δNP\delta^{\text{NP}}) 68.07%
    Ours-ft (Δbox\Delta_{\text{box}}) 75.34%
    Ours-ft (Δgeometric\Delta_{\text{geometric}} with δMG\delta^{\text{MG}}) 76.37%
    Ours-ft (Δgeometric\Delta_{\text{geometric}} with δNP\delta^{\text{NP}}) 76.34%
    Bounding Box Unknown
    DPD+DeCAF with no bounding box 44.94%
    Ours (Δnull\Delta_{\text{null}}) 64.57%
    Ours (Δbox\Delta_{\text{box}}) 65.22%
    Ours (Δgeometric\Delta_{\text{geometric}} with δMG\delta^{\text{MG}}) 65.98%
    Ours (Δgeometric\Delta_{\text{geometric}} with δNP\delta^{\text{NP}}) 65.96%
    Ours-ft (Δbox\Delta_{\text{box}}) 72.73%
    Ours-ft (Δgeometric\Delta_{\text{geometric}} with δMG\delta^{\text{MG}}) 72.95%
    Ours-ft (Δgeometric\Delta_{\text{geometric}} with δNP\delta^{\text{NP}}) 73.89%

    When bounding boxes are unknown at test time, Part-Based R-CNN with fine-tuning (Ours-ft with δNP\delta^{\text{NP}}) achieves 73.89% accuracy, which is 28.95 percentage points higher than DPD+DeCAF without bounding boxes (44.94%) and outperforms previous state-of-the-art methods that required ground-truth bounding boxes at test time (e.g., DPD+DeCAF with bounding box at 64.96%).

  7. Knowl 7 — Part Localization Accuracy Under Percentage of Correctly Localized Parts Metric

    data/table

    Part localization is evaluated using the Percentage of Correctly Localized Parts (PCP) metric, where a predicted part window is counted as correct if its Intersection over Union (IoU) overlap with the ground-truth part box is ≥0.5\ge 0.5.

    Method Head PCP Body PCP
    Bounding Box Given
    Strong DPM 43.49% 75.15%
    Ours (Δbox\Delta_{\text{box}}) 61.40% 65.42%
    Ours (Δgeometric\Delta_{\text{geometric}} with δMG\delta^{\text{MG}}) 66.03% 76.62%
    Ours (Δgeometric\Delta_{\text{geometric}} with δNP\delta^{\text{NP}}) 68.19% 79.82%
    Bounding Box Unknown
    Strong DPM 37.44% 47.08%
    Ours (Δnull\Delta_{\text{null}}) 60.50% 64.43%
    Ours (Δbox\Delta_{\text{box}}) 60.56% 65.31%
    Ours (Δgeometric\Delta_{\text{geometric}} with δMG\delta^{\text{MG}}) 61.94% 70.16%
    Ours (Δgeometric\Delta_{\text{geometric}} with δNP\delta^{\text{NP}}) 61.42% 70.68%

    Part-Based R-CNN substantially outperforms Strongly Supervised Deformable Part Models (Strong DPM). In the bounding-box-unknown setting, head localization improves from 37.44% (Strong DPM) to 61.94% (Ours with δMG\delta^{\text{MG}}), and body localization improves from 47.08% to 70.68% (Ours with δNP\delta^{\text{NP}}). Introducing non-parametric geometric constraints δNP\delta^{\text{NP}} provides an absolute gain of 14.40% on body localization over Δbox\Delta_{\text{box}} alone when the bounding box is given.

  8. Knowl 8 — Effect of Part Localization on Fine-Grained Object Classification

    data/table

    To isolate the contribution of semantic part representations versus whole-object representation, linear SVM classifiers were trained using only the whole-object CNN feature descriptors (no part features included) extracted from predicted or ground truth bounding boxes on CUB200-2011.

    Method Accuracy
    Oracle (ground truth bounding box) 57.94%
    Oracle-ft 68.29%
    Strong DPM 38.02%
    R-CNN 51.05%
    Ours (Δbox\Delta_{\text{box}}) 50.17%
    Ours (Δgeometric\Delta_{\text{geometric}} with δMG\delta^{\text{MG}}) 51.83%
    Ours (Δgeometric\Delta_{\text{geometric}} with δNP\delta^{\text{NP}}) 52.38%
    Ours-ft (Δbox\Delta_{\text{box}}) 62.13%
    Ours-ft (Δgeometric\Delta_{\text{geometric}} with δMG\delta^{\text{MG}}) 62.06%
    Ours-ft (Δgeometric\Delta_{\text{geometric}} with δNP\delta^{\text{NP}}) 62.75%

    Comparing object-only classification to the combined object-and-part classification reveals:

    1. Including localized part features increases classification accuracy by 11.14 percentage points for fine-tuned models in the bounding-box-unknown setting (62.75% object-only vs. 73.89% with parts).
    2. For oracle ground truth boxes, adding parts increases fine-tuned accuracy from 68.29% to 82.02% (a 13.73 percentage point gain).
    3. Applying joint geometric constraints over parts during detection improves root object window localization, raising accuracy from 51.05% (vanilla R-CNN) to 52.38% (Ours with δNP\delta^{\text{NP}}) without fine-tuning.
  9. Knowl 9 — Recall Performance of Bottom-Up Proposals on Semantic Parts

    data/table

    The recall of bottom-up Selective Search region proposals on CUB200-2011 was measured across different Intersection over Union (IoU) overlap thresholds for the whole object bounding box and semantic parts (head and body).

    Region Type IoU ≥0.50\ge 0.50 IoU ≥0.60\ge 0.60 IoU ≥0.70\ge 0.70
    Bounding box 96.70% 97.68% 89.50%
    Head 93.34% 73.87% 37.57%
    Body 96.70% 85.97% 54.68%

    At an overlap threshold of 0.50, Selective Search achieves high recall on bird parts (93.34% on heads and 96.70% on bodies, averaging ~95%), confirming that bottom-up region proposals provide a viable basis for part localization. However, at a stricter overlap threshold of 0.70, head recall drops to 37.57%, indicating that bottom-up region proposal accuracy can become a limiting bottleneck for fine-grained part localization at high overlap requirements.

  10. Knowl 10 — Hyperparameter Sensitivity to Geometric Prior Weight and Nearest Neighbors

    empirical result

    Five-fold cross-validation on the CUB200-2011 training set was used to evaluate the sensitivity of fine-grained classification accuracy to the geometric prior exponent α\alpha and the nearest neighbor count KK in δNP\delta^{\text{NP}}:

    • Geometric Weight (α\alpha): When α=0\alpha = 0, the model reduces to Δbox\Delta_{\text{box}} with no geometric layout prior. Accuracy peaks in the range α∈[0.05,0.10]\alpha \in [0.05, 0.10] (reaching ~64% cross-validation accuracy) and declines monotonically as α\alpha increases above 0.10 (dropping below 55% at α=0.40\alpha = 0.40). A small α\alpha is necessary because the probability density value of the continuous Gaussian distribution is substantially larger in magnitude than the sigmoid detector outputs, requiring fractional exponent attenuation to prevent the spatial prior from overpowering detector scores.
    • Nearest Neighbors (KK): Performance increases sharply from K=1K = 1 (~54%) to K=10K = 10 (~62%) and plateaus for K≥20K \ge 20 (reaching ~63.5% around K=80K = 80). Classification accuracy is robust and insensitive to the choice of KK as long as K≥10K \ge 10.

Coverage note — None was omitted; all key contributions including framework architecture, joint geometric optimization formulation, prior definitions, fine-tuning protocols, and experimental results on classification, localization, ablations, proposal recall, and hyperparameter sensitivity are covered.

References

  1. 1.Angelova, A., Zhu, S.: Efficient object detection and segmentation for fine-grained recognition. In: CVPR (2013)
  2. 2.Angelova, A., Zhu, S., Lin, Y.: Image segmentation for large-scale subcategory flower recognition. In: WACV (2013)
  3. 3.Azizpour, H., Laptev, I.: Object detection using strongly-supervised deformable part models. In: Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., Schmid, C. (eds.) ECCV 2012, Part I. LNCS, vol. 7572, pp. 836–849. Springer, Heidelberg (2012)
  4. 4.Belhumeur, P.N., Jacobs, D., Kriegman, D., Kumar, N.: Localizing parts of faces using a consensus of exemplars. In: CVPR (2011)
  5. 5.Belhumeur, P.N., et al.: Searching the world’s herbaria: A system for visual identification of plant species. In: Forsyth, D., Torr, P., Zisserman, A. (eds.) ECCV 2008, Part IV. LNCS, vol. 5305, pp. 116–129. Springer, Heidelberg (2008)
  6. 6.Belhumeur, P.N., Jacobs, D.W., Kriegman, D.J., Kumar, N.: Localizing parts of faces using a consensus of exemplars. In: CVPR (2011)
  7. 7.Berg, T., Belhumeur, P.N.: POOF: Part-based one-vs.-one features for fine-grained categorization, face verification, and attribute estimation. In: CVPR (2013)
  8. 8.Bourdev, L., Malik, J.: Poselets: Body part detectors trained using 3D human pose annotations. In: ICCV (2009), http://www.eecs.berkeley.edu/~lbourdev/poselets
  9. 9.Branson, S., Wah, C., Schroff, F., Babenko, B., Welinder, P., Perona, P., Belongie, S.: Visual recognition with humans in the loop. In: Daniilidis, K., Maragos, P., Paragios, N. (eds.) ECCV 2010, Part IV. LNCS, vol. 6314, pp. 438–451. Springer, Heidelberg (2010)
  10. 10.Chai, Y., Lempitsky, V., Zisserman, A.: Symbiotic segmentation and part localization for fine-grained categorization. In: ICCV (2013)
  11. 11.Chai, Y., Rahtu, E., Lempitsky, V., Van Gool, L., Zisserman, A.: TriCoS: A tri-level class-discriminative co-segmentation method for image classification. In: Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., Schmid, C. (eds.) ECCV 2012, Part I. LNCS, vol. 7572, pp. 794–807. Springer, Heidelberg (2012)
  12. 12.Dalal, N., Triggs, B.: Histograms of oriented gradients for human detection. In: CVPR (2005)
  13. 13.Deng, J., Krause, J., Fei-Fei, L.: Fine-grained crowdsourcing for fine-grained recognition. In: CVPR (2013)
  14. 14.Donahue, J., Jia, Y., Vinyals, O., Hoffman, J., Zhang, N., Tzeng, E., Darrell, T.: DeCAF: A deep convolutional activation feature for generic visual recognition. In: ICML (2014)
  15. 15.Duan, K., Parkh, D., Crandall, D., Grauman, K.: Discovering localized attributes for fine-grained recognition. In: CVPR (2012)
  16. 16.Farrell, R., Oza, O., Zhang, N., Morariu, V.I., Darrell, T., Davis, L.S.: Birdlets: Subordinate categorization using volumetric primitives and pose-normalized appearance. In: ICCV (2011)
  17. 17.Felzenszwalb, P.F., Girshick, R.B., McAllester, D., Ramanan, D.: Object detection with discriminatively trained part based models. IEEE Transactions on Pattern Analysis and Machine Intelligence (2010)
  18. 18.Felzenszwalb, P.F., Huttenlocher, D.: Efficient matching of pictorial structure. In: CVPR (2000)
  19. 19.Fischler, M.A., Elschlager, R.A.: The representation and matching of pictorial structures. IEEE Transactions on Computers (January 1973), http://dx.doi.org/10.1109/T-C.1973.223602
  20. 20.Gavves, E., Fernando, B., Snoek, C., Smeulders, A., Tuytelaars, T.: Fine-grained categorization by alignments. In: ICCV (2013)
  21. 21.Girshick, R., Donahue, J., Darrell, T., Malik, J.: Rich feature hierarchies for accurate object detection and semantic segmentation. In: CVPR (2014)
  22. 22.ILSVRC: ImageNet Large-scale Visual Recognition Challenge (2010-2012), http://www.image-net.org/challenges/LSVRC/2011/
  23. 23.Jarrett, K., Kavukcuoglu, K., Ranzato, M., LeCun, Y.: What is the best multi-stage architecture for object recognition? In: ICCV (2009)
  24. 24.Jia, Y.: Caffe: An open source convolutional architecture for fast feature embedding (2013), http://caffe.berkeleyvision.org/
  25. 25.Khosla, A., Jayadevaprakash, N., Yao, B., Fei-Fei, L.: Novel dataset for fine-grained image categorization. In: FGVC Workshop, CVPR (2011)
  26. 26.Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: NIPS (2012)
  27. 27.LeCun, Y., Boser, B., Denker, J., Henderson, D., Howard, R.E., Hubbard, W., Jackel, L.D.: Backpropagation applied to hand-written zip code recognition. Neural Computation (1989)
  28. 28.Lecun, Y., Bottou, L., Bengio, Y., Haffner, P.: Gradient-based learning applied to document recognition. Proceedings of the IEEE, 2278–2324 (1998)
  29. 29.Liu, J., Belhumeur, P.N.: Bird part localization using exemplar-based models with enforced pose and subcategory consistency. In: ICCV (2013)
  30. 30.Liu, J., Kanazawa, A., Jacobs, D., Belhumeur, P.: Dog breed classification using part localization. In: Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., Schmid, C. (eds.) ECCV 2012, Part I. LNCS, vol. 7572, pp. 172–185. Springer, Heidelberg (2012)
  31. 31.Maji, S., Kannala, J., Rahtu, E., Blaschko, M., Vedaldi, A.: Fine-grained visual classification of aircraft. Tech. rep. (2013)
  32. 32.Martinez-Munoz, G., Larios, N., Mortensen, E., Zhang, W., Yamamuro, A., Paasch, R., Payet, N., Lytle, D., Shapiro, L., Todorovic, S., Moldenke, A., Dietterich, T.: Dictionary-free categorization of very similar objects via stacked evidence trees. In: CVPR (2009)
  33. 33.Nilsback, M.E., Zisserman, A.: A visual vocabulary for flower classification. In: CVPR (2006)
  34. 34.Nilsback, M.E., Zisserman, A.: Automated flower classification over a large number of classes. In: ICVGIP (2008)
  35. 35.Parkhi, O.M., Vedaldi, A., Jawahar, C.V., Zisserman, A.: The truth about cats and dogs. In: ICCV (2011)
  36. 36.Parkhi, O.M., Vedaldi, A., Zisserman, A., Jawahar, C.V.: Cats and dogs. In: CVPR (2012)
  37. 37.Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R., LeCun, Y.: OverFeat: Integrated recognition, localization and detection using convolutional networks. CoRR abs/1312.6229 (2013)
  38. 38.Sfar, A.R., Boujemaa, N., Geman, D.: Vantage feature frames for fine-grained categorization. In: CVPR (2013)
  39. 39.Stark, M., Krause, J., Pepik, B., Meger, D., Little, J.J., Schiele, B., Koller, D.: Fine-grained categorization for 3D scene understanding. In: BMVC (2012)
  40. 40.Uijlings, J., van de Sande, K., Gevers, T., Smeulders, A.: Selective search for object recognition. IJCV (2013)
  41. 41.Welinder, P., Branson, S., Mita, T., Wah, C., Schroff, F., Belongie, S., Perona, P.: Caltech-UCSD Birds 200. Tech. Rep. CNS-TR-2010-001, California Institute of Technology (2010)
  42. 42.Xie, L., Tian, Q., Hong, R., Yan, S., Zhang, B.: Hierarchical part matching for fine-grained visual categorization. In: ICCV (2013)
  43. 43.Yang, S., Bo, L., Wang, J., Shapiro, L.: Unsupervised template learning for fine-grained object recognition. In: NIPS (2012)
  44. 44.Yao, B., Bradski, G., Fei-Fei, L.: A codebook-free and annotation-free approach for fine-grained image categorization. In: CVPR (2012)
  45. 45.Yao, B., Khosla, A., Fei-Fei, L.: Combining randomization and discrimination for fine-grained image categorization. In: CVPR (2011)
  46. 46.Zhang, N., Farrell, R., Darrell, T.: Pose pooling kernels for sub-category recognition. In: CVPR (2012)
  47. 47.Zhang, N., Farrell, R., Iandola, F., Darrell, T.: Deformable part descriptors for fine-grained recognition and attribute prediction. In: ICCV (2013)
  48. 48.Zhang, N., Paluri, M., Ranzato, M., Darrell, T., Bourdev, L.: PANDA: Pose aligned networks for deep attribute modeling. In: CVPR (2014)

Citation

MLA
Zhang, N., et al. “Part-based R-CNNs for Fine-grained Category Detection”. arXiv, 2014, http://arxiv.org/abs/1407.3867v1.
APA
Zhang, N., Donahue, J., Girshick, R., & Darrell, T. (2014). Part-based R-CNNs for Fine-grained Category Detection. arXiv. http://arxiv.org/abs/1407.3867v1
Chicago
Zhang, N., J. Donahue, R. Girshick, and T. Darrell. 2014. “Part-based R-CNNs for Fine-grained Category Detection”. arXiv. http://arxiv.org/abs/1407.3867v1.
Harvard
Zhang, N. et al. (2014) “Part-based R-CNNs for Fine-grained Category Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1407.3867v1.
Vancouver
1. Zhang N, Donahue J, Girshick R, Darrell T (2014) Part-based R-CNNs for Fine-grained Category Detection. arXiv

BibTeX

@article{zhang2014part,
  title = {Part-based R-CNNs for Fine-grained Category Detection},
  author = {Zhang, Ning and Donahue, Jeff and Girshick, Ross and Darrell, Trevor},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1407.3867v1},
  eprint = {1407.3867}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF