Fine-Grained Visual Classification of Aircraft
Subhransu MajiEsa RahtuJuho KannalaMatthew BlaschkoAndrea Vedaldi
Introduces the FGVC-Aircraft benchmark to evaluate fine-grained visual recognition on rigid man-made objects, challenging models to distinguish subtle structural differences across 100 aircraft variants organized in a three-level hierarchy.
The FGVC-Aircraft dataset was created to support research on fine-grained visual classification, a task that requires distinguishing visually similar object categories. Aircraft provide a useful alternative to common subjects such as birds or pets because they are rigid and exhibit structured variations in size, design history, purpose, and branding that are measurable from exterior images.
The work assembled 10,000 images covering 100 model variants, grouped into 70 families and 30 manufacturers. Images were sourced from aircraft enthusiast collections, filtered for diversity across photographers, airports, and time periods, then annotated with bounding boxes through crowdsourcing. Three classification tasks were defined at increasing levels of granularity, with performance measured by class-normalized accuracy on balanced train-validation-test splits.
A bag-of-visual-words baseline using dense SIFT features and a nonlinear SVM reached 48.7 percent accuracy on the 100-way variant task. Accuracy rose to 58.5 percent at the family level and 71.3 percent at the manufacturer level. Distinctive models such as the Eurofighter Typhoon exceeded 90 percent accuracy, while many Airbus and Boeing variants showed substantial confusion because differences are often limited to length or minor structural details.
These results indicate that current standard methods capture coarse manufacturer distinctions reasonably well but struggle with the subtle cues needed for variant-level recognition. The dataset therefore supplies a controlled benchmark that isolates the contribution of fine appearance differences from object deformation.
The authors plan to enlarge the collection as additional photographers grant permission and to apply the same construction process to other object categories. Researchers should treat the current baseline as a starting point rather than a performance ceiling; further gains will likely require methods that exploit part-level geometry or hierarchical label structure. The main limitations are the modest total size, the reliance on a restricted set of photographers even after diversity filtering, and the age range of the images, which may affect generalization to modern high-resolution photography.
- Paper: Visual categorization with bags of keypoints, Gabriella Csurka et al. (2004). Its bag-of-keypoints framework establishes the visual-word and SIFT-based classification approach that the aircraft paper adapts for its baseline.
- Paper: Part-Based R-CNNs for Fine-Grained Category Detection, Ning Zhang et al. (2014). It advances fine-grained recognition with automatic part localization, addressing the aircraft benchmark’s need for methods that capture subtle structural cues.
- Paper: Task Discrepancy Maximization for Fine-grained Few-Shot Classification, Su Been Lee et al. (2022). It applies a task-adaptive few-shot method to the Aircraft benchmark, extending aircraft recognition to settings with very few labeled examples.
