Fine-Grained Visual Classification of Aircraft

Subhransu MajiEsa RahtuJuho KannalaMatthew BlaschkoAndrea Vedaldi

article2013arXiv3,082 citations

Introduces the FGVC-Aircraft benchmark to evaluate fine-grained visual recognition on rigid man-made objects, challenging models to distinguish subtle structural differences across 100 aircraft variants organized in a three-level hierarchy.

Listen

The FGVC-Aircraft dataset was created to support research on fine-grained visual classification, a task that requires distinguishing visually similar object categories. Aircraft provide a useful alternative to common subjects such as birds or pets because they are rigid and exhibit structured variations in size, design history, purpose, and branding that are measurable from exterior images.

The work assembled 10,000 images covering 100 model variants, grouped into 70 families and 30 manufacturers. Images were sourced from aircraft enthusiast collections, filtered for diversity across photographers, airports, and time periods, then annotated with bounding boxes through crowdsourcing. Three classification tasks were defined at increasing levels of granularity, with performance measured by class-normalized accuracy on balanced train-validation-test splits.

A bag-of-visual-words baseline using dense SIFT features and a nonlinear SVM reached 48.7 percent accuracy on the 100-way variant task. Accuracy rose to 58.5 percent at the family level and 71.3 percent at the manufacturer level. Distinctive models such as the Eurofighter Typhoon exceeded 90 percent accuracy, while many Airbus and Boeing variants showed substantial confusion because differences are often limited to length or minor structural details.

These results indicate that current standard methods capture coarse manufacturer distinctions reasonably well but struggle with the subtle cues needed for variant-level recognition. The dataset therefore supplies a controlled benchmark that isolates the contribution of fine appearance differences from object deformation.

The authors plan to enlarge the collection as additional photographers grant permission and to apply the same construction process to other object categories. Researchers should treat the current baseline as a starting point rather than a performance ceiling; further gains will likely require methods that exploit part-level geometry or hierarchical label structure. The main limitations are the modest total size, the reliance on a restricted set of photographers even after diversity filtering, and the age range of the images, which may affect generalization to modern high-resolution photography.

arXiv: 1306.5151
  • Paper: Visual categorization with bags of keypoints, Gabriella Csurka et al. (2004). Its bag-of-keypoints framework establishes the visual-word and SIFT-based classification approach that the aircraft paper adapts for its baseline.
Cover for Fine-Grained Visual Classification of Aircraft

Abstract

This paper introduces FGVC-Aircraft, a new dataset containing 10,000 images of aircraft spanning 100 aircraft models, organised in a three-level hierarchy. At the finer level, differences between models are often subtle but always visually measurable, making visual recognition challenging but possible. A benchmark is obtained by defining corresponding classification tasks and evaluation protocols, and baseline results are presented. The construction of this dataset was made possible by the work of aircraft enthusiasts, a strategy that can extend to the study of number of other object classes. Compared to the domains usually considered in fine-grained visual classification (FGVC), for example animals, aircraft are rigid and hence less deformable. They, however, present other interesting modes of variation, including purpose, size, designation, structure, historical style, and branding.

Table of Contents

  • 1. Introduction
  • 2. The dataset: content, tasks, and evaluation
  • 3. Dataset construction
  • 3.1. Initial data collection
  • 3.2. Diversity maximisation
  • 3.3. Bounding boxes
  • 3.4. Hierarchy
  • 4. Baselines
  • 5. Summary
  • 5.1. Acknowledgments
  • References

Knowls

  1. Knowl 1 — FGVC-Aircraft Dataset Structure and Label Hierarchy

    definition

    FGVC-Aircraft is a benchmark dataset designed for fine-grained visual classification of aircraft. It contains 10,00010{,}000 high-resolution images (typically 1–2 megapixels) across 100 distinct aircraft model variants, balanced with exactly 100 images per variant. Each image contains an aircraft exterior where the dominant plane is annotated with a 2D bounding box.

    The aircraft categories are organized into a three-level evaluation hierarchy:

    • Variant (100 classes): The finest visually detectable distinction level, formed by grouping visually indistinguishable production models (e.g., the Boeing 737-700 variant consolidates models such as 737-7H4, 737-76N, and 737-7K2).
    • Family (70 classes): Groups of variants that share common design foundations and differ only in subtle structural attributes such as fuselage length or engine options (e.g., the Boeing 737 family contains variants 737-200 through 737-900).
    • Manufacturer (30 classes): Groups of aircraft families built by the same parent company (e.g., Boeing, Airbus, Cessna, Embraer).

    The dataset is divided into training, validation, and test subsets, each containing either 33 or 34 images per variant. Bounding box annotations are permitted for training classifiers but are withheld during test evaluation. Raw source images contain a 20-pixel-high photographer copyright banner along the bottom edge, which must be cropped out prior to visual feature extraction.

  2. Knowl 2 — Class-Normalised Average Accuracy Metric and Benchmark Tasks

    equation

    FGVC-Aircraft defines three classification tasks according to the label hierarchy: variant recognition (M=100M = 100), family recognition (M=70M = 70), and manufacturer recognition (M=30M = 30).

    Performance is evaluated using class-normalised average classification accuracy, calculated as the mean of the diagonal entries of the row-normalised confusion matrix. Let yi∈{1,…,M}y_i \in \{1, \dots, M\} denote the true label for image i∈{1,…,N}i \in \{1, \dots, N\}, and let y^i\hat{y}_i denote the predicted label. The element CpqC_{pq} of the confusion matrix represents the proportion of images of true class pp predicted as class qq:

    Cpq=∣{i∈{1,…,N}:y^i=q∧yi=p}∣∣{i∈{1,…,N}:yi=p}∣C_{pq} = \frac{|\{i \in \{1, \dots, N\} : \hat{y}_i = q \wedge y_i = p\}|}{|\{i \in \{1, \dots, N\} : y_i = p\}|}

    The overall class-normalised average accuracy is given by:

    Accuracy=1M∑p=1MCpp\text{Accuracy} = \frac{1}{M} \sum_{p=1}^M C_{pp}

    This normalization ensures that performance across imbalanced or balanced classes is weighted equally per category.

  3. Knowl 3 — Metadata-Driven Greedy Diversity Maximisation Algorithm

    algorithm

    When constructing visual datasets from aircraft spotter repositories, images from a small number of photographers often exhibit strong temporal, airport, and operator correlations. To decorrelate the dataset without accessing image pixels, candidate images are greedily filtered using an a priori similarity metric derived exclusively from four metadata fields: photographer identity, capture timestamp, airline operator, and airport location.

    Input: Candidate pool of images U\mathcal{U} for an aircraft variant (∣U∣≥120|\mathcal{U}| \ge 120), target size K=100K = 100, pairwise metadata similarity function sim(u,v)∈R\text{sim}(u, v) \in \mathbb{R}
    Output: Selected subset S⊂U\mathcal{S} \subset \mathcal{U} with ∣S∣=K|\mathcal{S}| = K
    S←∅\mathcal{S} \leftarrow \emptyset
    Select an initial image u∗∈Uu^* \in \mathcal{U}
    S←S∪{u∗}\mathcal{S} \leftarrow \mathcal{S} \cup \{u^*\}
    U←U∖{u∗}\mathcal{U} \leftarrow \mathcal{U} \setminus \{u^*\}
    while ∣S∣<K|\mathcal{S}| < K do
        u∗←arg⁡min⁡u∈Umax⁡s∈Ssim(u,s)u^* \leftarrow \arg\min_{u \in \mathcal{U}} \max_{s \in \mathcal{S}} \text{sim}(u, s)
        S←S∪{u∗}\mathcal{S} \leftarrow \mathcal{S} \cup \{u^*\}
        U←U∖{u∗}\mathcal{U} \leftarrow \mathcal{U} \setminus \{u^*\}
    end while
    Randomly partition S\mathcal{S} into training (33 or 34 images), validation (33 or 34 images), and test (33 or 34 images) splits
    return S\mathcal{S}

    This greedy dispersion breaks photographic burst sequences and prevents classifiers from exploiting background or livery artifacts.

  4. Knowl 4 — Crowdsourced Bounding Box Annotation and Consensus Protocol

    experimental setup

    Bounding boxes for the dominant aircraft in FGVC-Aircraft were obtained via Amazon Mechanical Turk using a multi-annotator agreement pipeline:

    1. Task Batching & Exclusion: Approximately 110 images per variant were submitted to crowd workers in batches of 10 images at a rate of 0.03 USD per batch (total cost of 110 USD, completed in under 48 hours). Annotators were instructed to skip images lacking an aircraft exterior (e.g., cockpit views), filtering out invalid images.
    2. Triplicate Annotation: Three separate bounding box annotations were collected for every image.
    3. Intersection-over-Union (IoU) Thresholding: The set of annotations was accepted only if at least two workers achieved an IoU overlap score exceeding 85%85\% (0.850.85). Annotations failing this consistency check were discarded.
    4. Box Averaging: Qualifying boxes were averaged to produce the final ground truth bounding box. Images failing quality checks were removed until exactly 100 validated images per variant remained.
  5. Knowl 5 — Baseline Bag-of-Visual-Words Classifier for Aircraft Recognition

    model/method

    The standard baseline classifier for FGVC-Aircraft utilizes a bag-of-visual-words representation with spatial pyramids and non-linear Support Vector Machines (SVMs):

    • Feature Extraction: Dense SIFT descriptors are extracted at multiple scales over the entire image area without cropping to the ground truth bounding box.
    • Dictionary & Quantization: A visual codebook of 600 words is constructed using kk-means clustering.
    • Spatial Pyramid Pooling: Visual word occurrences are pooled over spatial grids of 1×11 \times 1 and 2×22 \times 2 sub-regions.
    • Classification: A multi-class non-linear SVM with a χ2\chi^2 kernel is trained in a one-vs-rest configuration.
    • Hierarchical Classification Strategy: For evaluating coarser hierarchy levels (family and manufacturer), the model is trained exclusively on the 100-way variant classification task, and variant posterior predictions are mapped and summed to their parent family or manufacturer classes. Training SVMs directly on family or manufacturer labels yields significantly lower accuracy than this bottom-up label merging approach.
  6. Knowl 6 — Hierarchical Classification Performance of the Baseline Model

    empirical result

    Applying the multi-scale dense SIFT spatial pyramid χ2\chi^2-SVM baseline to FGVC-Aircraft across the three hierarchy levels yields the following class-normalised average accuracies:

    • Variant Recognition (100 classes): 48.69%48.69\%
    • Family Recognition (70 classes): 58.48%58.48\% (via bottom-up variant label merging)
    • Manufacturer Recognition (30 classes): 71.30%71.30\% (via bottom-up variant label merging)

    Performance varies sharply by airframe uniqueness. Visually distinct aircraft with unique wing or fuselage geometries achieve high recognition rates (e.g., DR-400 at 94.1%94.1\%, Eurofighter Typhoon at 94.1%94.1\%, F-16A/B at 90.9%90.9\%, and Cessna 172 at 88.2%88.2\%). Conversely, commercial airliner variants within the same family (e.g., Boeing 737, Boeing 747, Airbus A320/A330/A340 families) exhibit extensive intra-family confusion, with accuracies falling as low as 6.1%6.1\% for the Boeing 737-300 and 11.8%11.8\% for the Airbus A321. At the manufacturer level, the highest confusion occurs between Boeing and Airbus due to similar structural configurations across commercial passenger airliners.

  7. Knowl 7 — Per-Variant Classification Accuracy of Baseline Classifier on FGVC-Aircraft

    data/table

    The table below lists the test set class-normalised classification accuracy for each of the 100 aircraft variants evaluated with the multi-scale dense SIFT spatial pyramid χ2\chi^2-SVM baseline, sorted from highest to lowest accuracy.

    Model Variant Accuracy Model Variant Accuracy Model Variant Accuracy
    DR-400 94.1% DHC-8-100 57.6% ERJ 135 35.3%
    Eurofighter Typhoon 94.1% Embraer Legacy 600 57.6% 747-100 33.3%
    F-16A/B 90.9% F/A-18 57.6% 747-300 33.3%
    Cessna 172 88.2% 757-300 54.5% 767-200 33.3%
    SR-20 88.2% 767-400 54.5% 777-200 33.3%
    BAE-125 84.8% A340-500 54.5% BAE 146-200 33.3%
    DH-82 84.8% Cessna 208 54.5% DC-10 33.3%
    Tornado 84.8% Challenger 600 54.5% DC-8 33.3%
    C-130 81.8% E-170 54.5% MD-87 33.3%
    Hawk T1 81.8% Gulfstream V 54.5% 737-500 32.4%
    Model B200 81.8% ATR-42 51.5% 727-200 30.3%
    DHC-1 78.8% CRJ-900 51.5% A300B4 30.3%
    Il-76 76.5% EMB-120 51.5% A330-300 30.3%
    An-12 75.8% DC-3 50.0% E-190 29.4%
    Falcon 900 75.8% DHC-6 50.0% BAE 146-300 26.5%
    PA-28 75.8% Tu-134 48.5% 737-700 24.2%
    Spitfire 70.6% Gulfstream IV 47.1% A340-300 24.2%
    DC-6 69.7% Tu-154 47.1% MD-80 23.5%
    E-195 69.7% 737-900 45.5% A310 21.2%
    Cessna 560 67.6% Fokker 100 42.4% A319 21.2%
    Fokker 50 67.6% L-1011 42.4% A330-200 21.2%
    Cessna 525 66.7% Boeing 717 41.2% C-47 21.2%
    Global Express 66.7% CRJ-200 41.2% 747-200 20.6%
    Saab 2000 66.7% DHC-8-300 39.4% 737-200 17.6%
    Yak-42 66.7% ERJ 145 39.4% 737-800 17.6%
    A318 64.7% ATR-72 38.2% 757-200 17.6%
    Falcon 2000 64.7% 707-320 36.4% A320 15.2%
    Metroliner 64.7% 747-400 36.4% 767-300 14.7%
    Beechcraft 1900 63.6% CRJ-700 36.4% DC-9-30 14.7%
    Dornier 328 63.6% MD-11 36.4% 737-400 12.1%
    Fokker 70 63.6% MD-90 36.4% A321 11.8%
    Saab 340 63.6% 777-300 35.3% 737-300 06.1%
    737-600 57.6% A340-200 35.3%
    A380 57.6% A340-600 35.3% Average 48.69%

Coverage note — No substantial contributed material was omitted.

References

  1. 1.K. Chatfield, V. Lempitsky, A. Vedaldi, and A. Zisserman. The devil is in the details: an evaluation of recent feature encoding methods. In Proc. BMVC, 2011. 5
  2. 2.Aditya Khosla, Nityananda Jayadevaprakash, Bangpeng Yao, and Li Fei-Fei. Novel dataset for fine-grained image categorization. In CVPR Workshop on Fine-Grained Visual Categorization, 2011. 1
  3. 3.J. Liu, A. Kanazawa, D. Jacobs, and P. Belhumeur. Dog breed classification using part localization. In Proc. ECCV, 2012.
  4. 4.O. Parkhi, A. Vedaldi, C. V. Jawahar, and A. Zisserman. Cats vs dogs. In Proc. CVPR, 2012. 1
  5. 5.C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The caltech-ucsd birds-200-2011 dataset. Technical report, California Institute of Technology, 2011. 1

Citation

MLA
Maji, S., et al. “Fine-Grained Visual Classification of Aircraft”. arXiv, 2013, http://arxiv.org/abs/1306.5151v1.
APA
Maji, S., Rahtu, E., Kannala, J., Blaschko, M., & Vedaldi, A. (2013). Fine-Grained Visual Classification of Aircraft. arXiv. http://arxiv.org/abs/1306.5151v1
Chicago
Maji, S., E. Rahtu, J. Kannala, M. Blaschko, and A. Vedaldi. 2013. “Fine-Grained Visual Classification of Aircraft”. arXiv. http://arxiv.org/abs/1306.5151v1.
Harvard
Maji, S. et al. (2013) “Fine-Grained Visual Classification of Aircraft”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1306.5151v1.
Vancouver
1. Maji S, Rahtu E, Kannala J, Blaschko M, Vedaldi A (2013) Fine-Grained Visual Classification of Aircraft. arXiv

BibTeX

@article{maji2013fine,
  title = {Fine-Grained Visual Classification of Aircraft},
  author = {Maji, Subhransu and Rahtu, Esa and Kannala, Juho and Blaschko, Matthew and Vedaldi, Andrea},
  year = {2013},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1306.5151v1},
  eprint = {1306.5151}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors