PACO: Parts and Attributes of Common Objects

Vignesh RamanathanAnmol KaliaVladan PetrovicYi WenBaixue ZhengBaishan GuoRui WangAaron MarquezRama KovvuriAbhishek Kadian

article2023CVPR194 citations

Introduces a large-scale visual dataset and benchmark spanning 75 object categories with over 640,000 part masks and detailed attribute annotations across images and videos to advance joint part segmentation and zero-shot instance detection.

Listen

Modern computer vision models increasingly need to provide detailed, fine-grained visual descriptions rather than simply identifying broad object category labels. Applications such as open-vocabulary retrieval and natural language interaction require systems to localize specific components of an object and recognize their detailed characteristics. However, existing public datasets either focus on narrow domains like fashion and birds or lack joint annotations that capture object masks, part masks, and attributes together across a broad set of common objects.

The article introduces PACO (Parts and Attributes of Common Objects), a large-scale benchmark designed to advance research in joint object detection, part segmentation, and attribute recognition across common objects. The objective is to establish an open-source evaluation suite and baseline performance benchmarks across image and video domains for fine-grained object understanding.

To construct PACO, the authors combined image data from the LVIS benchmark with video frames from the Ego4D dataset. The dataset spans 75 common object categories, 456 object-specific part categories, and 55 curated attributes covering colors, patterns, materials, and reflectance properties. In total, the authors collected 641,400 part masks across 260,300 object instances in approximately 81,500 images, making it roughly ten times larger in annotated object instances with parts than previous common object benchmarks. To establish standard reference points, the authors extended standard detection models—specifically Mask R-CNN and Vision Transformer architectures—and evaluated them across three benchmark tasks: part mask segmentation, instance-level attribute prediction, and zero-shot instance detection using multi-attribute queries.

The experimental findings show that segmenting parts and predicting fine-grained attributes are substantially more challenging than standard object detection. On part segmentation, the best Vision Transformer model achieved a mask Average Precision (AP) of 17.7 for object parts compared to 43.4 for whole objects, primarily because object parts are significantly smaller and visually distinct across different parent categories. On instance-level attribute detection, the highest-performing model scored an overall box AP of 19.5 on object attributes and 13.8 on part attributes; performance bounds analysis revealed that the theoretical upper bound with perfect attribute classification exceeds 72 AP, exposing a massive capability gap in current models. In zero-shot instance retrieval, queries with a single attribute yielded higher recall than complex two-attribute queries because prediction errors compounded as query complexity grew. Furthermore, off-the-shelf open-vocabulary detectors struggled significantly on fine-grained queries, registering low single-digit recall scores compared to task-adapted baselines.

These results indicate that current computer vision systems are fundamentally limited in their ability to resolve fine visual details and individual component properties. Improving performance in these areas is crucial for commercial systems that require precise localization, such as inventory management, robotics, and advanced search platforms, where misidentifying part attributes introduces significant operational error. The findings also demonstrate that open-vocabulary foundation models cannot yet reliably substitute for specialized, fine-grained visual detectors without targeted architectural enhancements.

Organizations developing fine-grained perception systems should adopt joint object-part-attribute training objectives and utilize larger vision transformer backbones, which consistently outperformed standard convolutional networks across all benchmark tasks. System architects should also pursue new modeling techniques that reduce compound attribute prediction errors in complex language queries and adapt open-vocabulary frameworks to handle dense attribute reasoning.

While the dataset provides rigorous quality control—including high agreement with expert gold annotations—limitations remain. The vocabulary is constrained to 75 common categories and 55 attributes, and the data distribution exhibits a long-tail imbalance where certain rare parts and attributes have limited examples. Nevertheless, the benchmark provides a highly credible and standardized foundation for evaluating progress in detailed visual understanding.

  • Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Segment Anything scales promptable mask prediction to a foundation model level, evaluating zero-shot part and object delineation across fine-grained datasets like PACO.
Cover for PACO: Parts and Attributes of Common Objects

Abstract

Object models are gradually progressing from predicting just category labels to providing detailed descriptions of object instances. This motivates the need for large datasets which go beyond traditional object masks and provide richer annotations such as part masks and attributes. Hence, we introduce PACO: Parts and Attributes of Common Objects. It spans 75 object categories, 456 object-part categories and 55 attributes across image (LVIS) and video (Ego4D) datasets. We provide 641K part masks annotated across 260K object boxes, with roughly half of them exhaustively annotated with attributes as well. We design evaluation metrics and provide benchmark results for three tasks on the dataset: part mask segmentation, object and part attribute prediction and zero-shot instance detection. Dataset, models, and code are open-sourced at https://github.com/facebookresearch/paco.

Table of Contents

  • 1. Introduction
  • 1.1. Related work
  • 2. Dataset construction
  • 2.1. Image sources
  • 2.2. Object vocabulary selection
  • 2.3. Parts vocabulary selection
  • 2.4. Attribute vocabulary selection
  • 2.5. Annotation pipeline
  • 2.5.1 Object annotation
  • 2.5.2 Part mask annotation
  • 2.5.3 Attributes annotation
  • 2.5.4 Instance annotation
  • 2.5.5 Managing annotation quality
  • 3. Dataset statistics
  • 4. Tasks and evaluation benchmark
  • 4.1. Dataset splits
  • 4.2. Federated dataset for object categories
  • 4.3. Part segmentation
  • 4.4. Instance-level attributes prediction
  • 4.5. Zero-shot instance detection
  • 5. Benchmarking experiments
  • 5.1. Part segmentation
  • 5.2. Instance-level attributes prediction
  • 5.3. Zero-shot instance detection
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — PACO Dataset Specification and Structure

    definition

    PACO (Parts and Attributes of Common Objects) is a benchmark dataset designed for the joint tasks of object detection, part mask segmentation, and instance-level attribute recognition across both image and video domains.

    The dataset is characterized by the following specifications:

    • Object Categories: 75 common object categories selected from the intersection of frequent LVIS categories and Ego4D video narrations having at least 20 instances in Ego4D.
    • Object-Part Categories: 456 object-specific part categories derived from 200 object-agnostic part classes mined from web diagrams and curated to ensure visibility and distinguishability across instances.
    • Attributes: 55 visual attributes grouped into four categories: 29 colors, 10 patterns and markings, 13 materials, and 3 reflectance levels.
    • Image Sources and Annotation Counts:
      • PACO-LVIS (static images): 57.6K images with object masks, 52.7K images with part masks, 274K object masks, 502K part masks, 74.4K objects with attributes, and 186K parts with attributes.
      • PACO-EGO4D (egocentric video frames): 23.9K frames with object masks, 24K frames with part masks, 58.4K object masks, 139.3K part masks, 49.6K objects with attributes, and 110.6K parts with attributes.
      • Joint Totals: 641.4K part masks across 260.3K object instances in 76.7K images/frames. Attribute annotations cover 124K object instances and 296.6K part instances, exhaustively annotated across all 55 attributes per annotated box (averaging 23.4 attribute annotations per image) and providing explicit negative attribute labels.
    • Dataset Splits:
      • PACO-LVIS: 45,790 train, 2,410 validation, and 9,443 test images (the test split is a strict subset of the LVIS-v1 validation set).
      • PACO-EGO4D: 15,667 train, 825 validation, and 9,892 test frames, with mutually disjoint object instance IDs between train and test sets (16,908 unique object instances annotated in total).
  2. Knowl 2 — Federated Evaluation Protocol for Object-Part and Attribute Detection

    model/method

    Because PACO is a federated dataset where not every image is exhaustively annotated for all categories and attributes, evaluation metrics adapt the Average Precision (AP) protocol across Intersection over Union (IoU) thresholds to handle missing labels.

    1. Part Segmentation (APopartAP^{opart})

    For an object-part category (o,p)(o, p) combining object oo and part pp:

    • Negative images of object oo: All predicted masks for (o,p)(o, p) are marked as false positives.
    • Non-exhaustive positive images: Predicted masks overlapping an annotated ground-truth (o,p)(o, p) mask above the IoU threshold are true positives. All other predicted masks are ignored.
    • Exhaustive positive images: Predicted masks matching ground-truth (o,p)(o, p) masks above the IoU threshold are true positives; unmatched predictions are false positives.

    2. Instance-Level Category-Attribute Detection (APattobjAP^{obj}_{att} and APattopartAP^{opart}_{att})

    For a category c∈{o,(o,p)}c \in \{o, (o, p)\} and an attribute aa, AP is computed only if the evaluation split contains ≥1\ge 1 positive instance of (c,a)(c, a) and ≥40\ge 40 negative instances of category cc for attribute aa:

    • Negative images of object oo: All predicted masks for (c,a)(c, a) are false positives.
    • Positive images (both exhaustive and non-exhaustive):
      • Predictions overlapping ground-truth category cc annotated positively for aa are true positives.
      • Predictions overlapping ground-truth category cc annotated negatively for aa are false positives.
      • Predictions overlapping ground-truth category cc with un-annotated attribute labels are ignored.
      • Predictions not overlapping any ground-truth mask of category cc are ignored in non-exhaustive images and marked as false positives in exhaustive images.

    To reduce metric variance from rare category-attribute combinations, AP values for (c,a)(c, a) are averaged across all valid categories cc for attribute aa to yield APaAP_a. These are subsequently averaged across attribute types: color (APcolAP_{col}), pattern & markings (APpatAP_{pat}), material (APmatAP_{mat}), reflectance (APrefAP_{ref}), and overall (APattAP_{att}).

  3. Knowl 3 — Zero-Shot Instance Detection Task and Query Hierarchy

    definition

    Zero-shot instance detection is the task of retrieving and localizing the bounding box of a specific instance of an object from an image collection using a compositional language query, where no sample images of that instance have been previously seen.

    Query Hierarchy

    Queries are structured by compositionality into level-kk (LkLk) queries containing kk attributes of the object and/or its parts:

    • Level 1 (L1L1): 1 attribute (e.g., "blue mug" or "mug with a blue handle").
    • Level 2 (L2L2): 2 attributes (e.g., "blue striped mug").
    • Level 3 (L3L3): 3 attributes (e.g., "blue striped mug with white handle").

    Evaluation Setup and Metrics

    • Target and Distractors: Each query is evaluated against 1 positive target image and up to 100 distractor images. Distractor images contain hard negatives: other instances of the same object class that differ by at least one attribute from the query. On average, >40%>40\% of distractors per query are hard negatives.
    • Query Statistics: PACO-LVIS contains 931 L1L1, 2,348 L2L2, and 2,000 L3L3 queries. PACO-EGO4D contains 793 L1L1, 1,437 L2L2, and 2,115 L3L3 queries.
    • Metric: Average Recall at top-kk (AR@kAR@k for k∈{1,5}k \in \{1, 5\}), calculated across bounding box IoU thresholds from 0.50.5 to 0.950.95.
  4. Knowl 4 — Joint Object, Part, and Attribute Detector Architecture and Scoring

    model/method

    Joint detection of objects, parts, and attributes extends standard two-stage object detectors (Mask R-CNN and ViT-det) with multi-task prediction heads:

    1. Object and Part Detection Head: The standard classification head is configured to predict 531 mutually distinct classes (75 object categories plus 456 object-part categories) along with bounding boxes and segmentation masks. Models are trained using federated cross-entropy loss and Large Scale Jittering (LSJ) data augmentation.
    2. Attribute Head: An auxiliary multi-task head attached to the shared detector backbone uses the same Region of Interest (RoI)-pooled features as the detection head to predict object and part attributes. It outputs independent multi-class probability distributions for each of the four attribute types (color, pattern, material, reflectance) optimized via separate cross-entropy loss terms.
    3. Inference Ranking:
      • For a single category-attribute pair (c,a)(c, a) where cc is an object or object-part and aa is an attribute, candidate bounding boxes are ranked by the product of their category score ScS_c and attribute score SaS_a: S(c,a)=Sc×SaS(c, a) = S_c \times S_a
      • For zero-shot instance queries composed of multiple object, part, and attribute terms, candidate bounding boxes are ranked using the geometric mean of the constituent detection and attribute probabilities predicted by the model.
  5. Knowl 5 — Object and Object-Part Segmentation Benchmark Results on PACO-LVIS

    data/table

    Instance segmentation performance for object categories (APobjAP^{obj}) and object-part categories (APopartAP^{opart}) on the PACO-LVIS test split across ResNet and Vision Transformer (ViT-det) backbones with Feature Pyramid Networks (FPN) and Cascade FPN:

    Model Mask APobjAP^{obj} Mask APopartAP^{opart} Box APobjAP^{obj} Box APopartAP^{opart}
    ResNet-50 FPN 31.5±0.331.5 \pm 0.3 12.3±0.112.3 \pm 0.1 34.6±0.334.6 \pm 0.3 16.0±0.116.0 \pm 0.1
    ResNet-50 Cascade FPN 32.6±1.332.6 \pm 1.3 12.5±0.712.5 \pm 0.7 37.4±1.637.4 \pm 1.6 16.3±1.116.3 \pm 1.1
    ResNet-101 FPN 31.5±0.631.5 \pm 0.6 12.3±0.312.3 \pm 0.3 34.8±0.834.8 \pm 0.8 16.1±0.316.1 \pm 0.3
    ResNet-101 Cascade FPN 35.1±0.135.1 \pm 0.1 13.7±0.113.7 \pm 0.1 40.2±0.140.2 \pm 0.1 17.9±0.217.9 \pm 0.2
    ViT-B FPN 33.6±0.333.6 \pm 0.3 13.5±0.113.5 \pm 0.1 38.7±0.438.7 \pm 0.4 17.5±0.017.5 \pm 0.0
    ViT-B Cascade FPN 33.6±0.333.6 \pm 0.3 13.5±0.113.5 \pm 0.1 38.7±0.438.7 \pm 0.4 17.5±0.017.5 \pm 0.0
    ViT-L FPN 42.8±0.342.8 \pm 0.3 17.3±0.117.3 \pm 0.1 47.3±0.247.3 \pm 0.2 22.0±0.122.0 \pm 0.1
    ViT-L Cascade FPN 43.4±0.343.4 \pm 0.3 17.7±0.017.7 \pm 0.0 49.7±0.249.7 \pm 0.2 22.9±0.022.9 \pm 0.0

    Object-part segmentation (APopartAP^{opart}) is significantly lower than object segmentation (APobjAP^{obj}) across all models (e.g., 17.7%17.7\% vs. 43.4%43.4\% mask AP for ViT-L Cascade FPN). This degradation is primarily driven by the smaller physical size distribution of parts relative to full objects. Scaling the backbone architecture to ViT-L provides consistent improvements across both tasks.

  6. Knowl 6 — Instance-Level Attribute Prediction Results and Oracle Bounds on PACO-LVIS

    data/table

    Box Average Precision (Box AP) performance for object attributes (APobjAP^{obj}) and object-part attributes (APopartAP^{opart}) on the PACO-LVIS test split, categorized by overall attribute score (attatt), color (colcol), pattern/marking (patpat), material (matmat), and reflectance (refref):

    Model APattobjAP^{obj}_{att} APcolobjAP^{obj}_{col} APpatobjAP^{obj}_{pat} APmatobjAP^{obj}_{mat} APrefobjAP^{obj}_{ref} APattopartAP^{opart}_{att} APcolopartAP^{opart}_{col} APpatopartAP^{opart}_{pat} APmatopartAP^{opart}_{mat} APrefopartAP^{opart}_{ref}
    R50 FPN 13.5±0.313.5 \pm 0.3 10.8±0.110.8 \pm 0.1 14.1±0.614.1 \pm 0.6 9.9±0.49.9 \pm 0.4 19.1±0.719.1 \pm 0.7 9.7±0.29.7 \pm 0.2 10.7±0.210.7 \pm 0.2 10.6±0.510.6 \pm 0.5 6.9±0.06.9 \pm 0.0 10.7±0.210.7 \pm 0.2
    R50 Cascade 15.0±1.015.0 \pm 1.0 12.4±0.712.4 \pm 0.7 16.1±0.716.1 \pm 0.7 11.0±0.911.0 \pm 0.9 20.6±1.620.6 \pm 1.6 10.5±0.710.5 \pm 0.7 11.6±0.811.6 \pm 0.8 11.6±0.811.6 \pm 0.8 7.6±0.77.6 \pm 0.7 11.2±0.711.2 \pm 0.7
    R101 FPN 13.5±0.313.5 \pm 0.3 11.0±0.211.0 \pm 0.2 13.9±0.313.9 \pm 0.3 9.9±0.49.9 \pm 0.4 19.1±0.619.1 \pm 0.6 9.9±0.19.9 \pm 0.1 11.0±0.411.0 \pm 0.4 10.8±0.410.8 \pm 0.4 7.1±0.27.1 \pm 0.2 10.9±0.310.9 \pm 0.3
    R101 Cascade 16.0±0.116.0 \pm 0.1 13.4±0.213.4 \pm 0.2 16.7±0.216.7 \pm 0.2 12.3±0.112.3 \pm 0.1 21.5±0.421.5 \pm 0.4 11.5±0.211.5 \pm 0.2 12.6±0.112.6 \pm 0.1 12.5±0.312.5 \pm 0.3 8.5±0.38.5 \pm 0.3 12.6±0.312.6 \pm 0.3
    ViT-B FPN 15.0±0.215.0 \pm 0.2 11.9±0.111.9 \pm 0.1 14.9±0.514.9 \pm 0.5 12.8±0.412.8 \pm 0.4 20.4±0.820.4 \pm 0.8 10.9±0.210.9 \pm 0.2 11.3±0.311.3 \pm 0.3 11.4±0.611.4 \pm 0.6 9.0±0.19.0 \pm 0.1 11.8±0.311.8 \pm 0.3
    ViT-B Cascade 15.7±0.215.7 \pm 0.2 12.6±0.112.6 \pm 0.1 16.0±0.516.0 \pm 0.5 13.2±0.413.2 \pm 0.4 20.9±0.520.9 \pm 0.5 11.0±0.211.0 \pm 0.2 11.6±0.211.6 \pm 0.2 11.7±0.411.7 \pm 0.4 9.0±0.29.0 \pm 0.2 11.5±0.311.5 \pm 0.3
    ViT-L FPN 18.8±0.318.8 \pm 0.3 14.9±0.214.9 \pm 0.2 18.9±1.018.9 \pm 1.0 16.0±0.716.0 \pm 0.7 25.4±0.725.4 \pm 0.7 13.5±0.213.5 \pm 0.2 14.0±0.214.0 \pm 0.2 14.0±0.414.0 \pm 0.4 11.7±0.411.7 \pm 0.4 14.3±0.614.3 \pm 0.6
    ViT-L Cascade 19.5±0.319.5 \pm 0.3 15.6±0.315.6 \pm 0.3 19.1±0.519.1 \pm 0.5 16.3±0.316.3 \pm 0.3 27.0±0.427.0 \pm 0.4 13.8±0.113.8 \pm 0.1 14.4±0.314.4 \pm 0.3 15.1±0.015.1 \pm 0.0 11.5±0.211.5 \pm 0.2 14.5±0.414.5 \pm 0.4

    Oracle Sensitivity Bounds on Attribute Prediction

    To isolate attribute classification accuracy from object detection quality, APattobjAP^{obj}_{att} bounds are computed by freezing the model's box detections and replacing predicted attribute scores with (1) Lower Bound (LB): uniform attribute scores (ignoring attribute predictions), and (2) Upper Bound (UB): ground-truth attribute oracle scores:

    Model LB (No Attribute) Original UB (Perfect Attribute)
    ResNet-50 FPN 8.6±0.38.6 \pm 0.3 13.5±0.313.5 \pm 0.3 61.4±0.361.4 \pm 0.3
    ResNet-101 FPN 8.6±0.38.6 \pm 0.3 13.5±0.313.5 \pm 0.3 63.0±0.363.0 \pm 0.3
    ViT-B FPN 9.0±0.19.0 \pm 0.1 15.0±0.215.0 \pm 0.2 60.5±0.160.5 \pm 0.1
    ViT-L FPN 10.6±0.210.6 \pm 0.2 18.8±0.318.8 \pm 0.3 72.6±0.372.6 \pm 0.3

    The substantial gap between baseline model performance (13.5%−18.8%13.5\% - 18.8\%) and the upper bound (60.5%−72.6%60.5\% - 72.6\%) demonstrates that attribute classification errors constitute the primary bottleneck in joint instance-attribute prediction.

  7. Knowl 7 — Zero-Shot Instance Detection Performance Across Query Complexity Levels

    data/table

    Evaluation of FPN models on the PACO-LVIS zero-shot instance detection test set across query complexity levels (L1L1, L2L2, L3L3, and all queries combined) using Average Recall at top-1 (AR@1AR@1) and top-5 (AR@5AR@5):

    L1 queries L2 queries L3 queries All queries
    Model AR@1AR@1 AR@5AR@5 AR@1AR@1 AR@5AR@5 AR@1AR@1 AR@5AR@5 AR@1AR@1 AR@5AR@5
    ResNet-50 FPN 22.5±0.722.5 \pm 0.7 39.2±0.539.2 \pm 0.5 20.1±0.420.1 \pm 0.4 38.5±0.138.5 \pm 0.1 22.3±0.922.3 \pm 0.9 44.5±1.144.5 \pm 1.1 21.4±0.621.4 \pm 0.6 40.9±0.340.9 \pm 0.3
    ResNet-101 FPN 23.1±0.723.1 \pm 0.7 40.5±1.440.5 \pm 1.4 20.0±0.620.0 \pm 0.6 39.3±1.039.3 \pm 1.0 23.1±0.723.1 \pm 0.7 45.2±0.645.2 \pm 0.6 21.7±0.621.7 \pm 0.6 41.8±0.841.8 \pm 0.8
    ViT-B FPN 26.8±0.226.8 \pm 0.2 45.8±0.245.8 \pm 0.2 22.7±0.522.7 \pm 0.5 40.0±0.740.0 \pm 0.7 24.1±0.524.1 \pm 0.5 42.5±1.542.5 \pm 1.5 23.9±0.423.9 \pm 0.4 42.0±0.942.0 \pm 0.9
    ViT-L FPN 35.3±0.735.3 \pm 0.7 57.3±0.657.3 \pm 0.6 29.7±0.629.7 \pm 0.6 50.1±0.250.1 \pm 0.2 31.1±0.831.1 \pm 0.8 52.3±0.952.3 \pm 0.9 31.2±0.431.2 \pm 0.4 52.2±0.552.2 \pm 0.5

    Across all model architectures, performance exhibits a consistent non-monotonic trend: L1>L3>L2L1 > L3 > L2. This trend reflects the trade-off between two countervailing effects:

    1. Information gain: Highly compositional queries (L3L3) provide more descriptive constraints, making target instance retrieval easier than moderately compositional queries (L2L2).
    2. Error compounding: Evaluating multiple attribute predictions simultaneously increases the likelihood of classification failure, causing single-attribute queries (L1L1) to outperform multi-attribute queries (L3L3).
  8. Knowl 8 — Zero-Shot Instance Detection Comparison: Open-Vocabulary Detectors vs PACO Baselines

    data/table

    Zero-shot instance detection evaluation on PACO-LVIS comparing off-the-shelf open-vocabulary detectors (MDETR and Detic) against models trained directly on PACO for Level 1 queries (L1L1), broken down by queries containing only object attributes (L1objL1_{obj}) and only part attributes (L1partL1_{part}):

    Model L1objL1_{obj} L1partL1_{part} All L1L1
    MDETR (ResNet-101) 4.1±0.64.1 \pm 0.6 5.3±0.65.3 \pm 0.6 4.9±0.34.9 \pm 0.3
    PACO ResNet-101 FPN 20.3±0.920.3 \pm 0.9 24.4±1.024.4 \pm 1.0 23.1±0.723.1 \pm 0.7
    Detic (Swin-B) 5.2±0.75.2 \pm 0.7 6.2±0.36.2 \pm 0.3 5.9±0.25.9 \pm 0.2
    PACO ViT-B FPN 22.6±0.822.6 \pm 0.8 28.9±0.628.9 \pm 0.6 26.8±0.226.8 \pm 0.2

    All numbers report AR@1AR@1. Open-vocabulary detectors achieve low recall (4.9%4.9\% for MDETR and 5.9%5.9\% for Detic on All L1L1) compared to PACO-trained baselines (23.1%23.1\% and 26.8%26.8\%) due to architectural and training limitations:

    • Detic is primarily trained on noun vocabularies and has limited fine-grained attribute grounding capability.
    • MDETR is trained for referring expressions on positive images and exhibits poor discrimination against hard negative distractor images.
  9. Knowl 9 — Zero-Shot vs Few-Shot Instance Detection Discrepancy on PACO-EGO4D

    empirical result

    On the PACO-EGO4D benchmark, where multiple video frames depict the same physical object instance, few-shot instance detection was compared against zero-shot attribute-based instance detection across 1,992 queries that have at least 6 bounding box instances of the target object.

    • Few-Shot Model Setup: A two-stage baseline employing a pre-trained ResNet-50 FPN object detector followed by nearest-neighbor ranking of RoI-pooled features against kk visual exemplar bounding boxes (k∈{1,2,3,4,5}k \in \{1, 2, 3, 4, 5\}).
    • Performance Gap:
      • Zero-shot FPN detectors trained on the joint PACO dataset achieve between 15%15\% and 22%22\% AR@1AR@1 (with ResNet-101 FPN achieving ≈22%\approx 22\%, ViT-B FPN ≈17%\approx 17\%, and ResNet-50 FPN ≈15%\approx 15\%).
      • The 1-shot visual baseline (k=1k = 1) achieves over 42%42\% AR@1AR@1, creating an absolute margin of >20>20 points over the strongest zero-shot model.
      • Scaling visual exemplars from k=1k = 1 to k=5k = 5 further elevates few-shot retrieval performance toward 50%50\%.

    This indicates that current attribute and part language descriptions in zero-shot models are far less effective at capturing instance-level identity than even a single visual bounding box exemplar.

Coverage note — No substantial contributed material was omitted; all key definitions, dataset statistics, federated metric formulations, baseline architectures, and experimental results from the main paper have been fully captured.

References

  1. 1.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6154–6162, 2018. 7
  2. 2.Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu, Sanja Fidler, Raquel Urtasun, and Alan Yuille. Detect what you can: Detecting and representing objects using holistic models and body parts. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1971–1978, 2014. 1, 2
  3. 3.Zhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K. Wong, and Qi Wu. Cops-ref: A new dataset and task on compositional referring expression comprehension. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 1, 3, 6
  4. 4.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 2
  5. 5.Li Deng. The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE signal processing magazine, 29(6):141–142, 2012. 2
  6. 6.N. Dinesh Reddy, Minh Vo, and Srinivasa G. Narasimhan. Carfusion: Combining point tracking and part detection for dynamic 3d reconstruction of vehicles. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 2
  7. 7.M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The pascal visual object classes challenge: A retrospective. International Journal of Computer Vision, 111(1):98–136, Jan. 2015. 2
  8. 8.Chengjian Feng, Yujie Zhong, Zequn Jie, Xiangxiang Chu, Haibing Ren, Xiaolin Wei, Weidi Xie, and Lin Ma. Promptdet: Expand your detector vocabulary with uncurated images. arXiv preprint arXiv:2203.16513, 2022. 1
  9. 9.Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, Tsung-Yi Lin, Ekin D Cubuk, Quoc V Le, and Barret Zoph. Simple copy-paste is a strong data augmentation method for instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2918–2928, 2021. 7
  10. 10.Ke Gong, Xiaodan Liang, Dongyu Zhang, Xiaohui Shen, and Liang Lin. Look into person: Self-supervised structuresensitive learning and a new benchmark for human parsing. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 932–940, 2017. 2
  11. 11.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18995–19012, 2022. 2, 3
  12. 12.Sheng Guo, Weilin Huang, Xiao Zhang, Prasanna Srikhanta, Yin Cui, Yuan Li, Hartwig Adam, Matthew R Scott, and Serge Belongie. The imaterialist fashion attribute dataset. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 2
  13. 13.Agrim Gupta, Piotr Dollar, and Ross Girshick. LVIS: A dataset for large vocabulary instance segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019. 2, 3, 5, 6
  14. 14.Akshita Gupta, Sanath Narayan, KJ Joseph, Salman Khan, Fahad Shahbaz Khan, and Mubarak Shah. Ow-detr: Open-world detection transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9235–9244, 2022. 1
  15. 15.Ju He, Shuo Yang, Shaokang Yang, Adam Kortylewski, Xiaoding Yuan, Jie-Neng Chen, Shuai Liu, Cheng Yang, and Alan Yuille. Partimagenet: A large, high-quality dataset of parts. arXiv preprint arXiv:2112.00933, 2021. 1, 2
  16. 16.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 2
  17. 17.Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6700–6709, 2019. 1, 3
  18. 18.Menglin Jia, Mengyun Shi, Mikhail Sirotenko, Yin Cui, Claire Cardie, Bharath Hariharan, Hartwig Adam, and Serge Belongie. Fashionpedia: Ontology, segmentation, and an attribute localization dataset. In European conference on computer vision, pages 316–332. Springer, 2020. 1, 2, 3
  19. 19.Licheng Jiao, Ruohan Zhang, Fang Liu, Shuyuan Yang, Biao Hou, Lingling Li, and Xu Tang. New generation deep learning for video object detection: A survey. IEEE Transactions on Neural Networks and Learning Systems, 2021. 2
  20. 20.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetrmodulated detection for end-to-end multi-modal understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1780–1790, 2021. 1, 8
  21. 21.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 787–798, 2014. 1, 3, 6
  22. 22.Ronak Kosti, Jose M Alvarez, Adria Recasens, and Agata Lapedriza. Emotion recognition in context. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1667–1675, 2017. 2
  23. 23.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123(1):32–73, 2017. 1
  24. 24.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al. The open images dataset v4. International Journal of Computer Vision, 128(7):1956–1981, 2020. 2
  25. 25.Dangwei Li, Zhang Zhang, Xiaotang Chen, Haibin Ling, and Kaiqi Huang. A richly annotated dataset for pedestrian attribute recognition. arXiv preprint arXiv:1603.07054, 2016. 1
  26. 26.Jianshu Li, Jian Zhao, Yunchao Wei, Congyan Lang, Yidong Li, Terence Sim, Shuicheng Yan, and Jiashi Feng. Multiplehuman parsing in the wild. arXiv preprint arXiv:1705.07206, 2017. 2
  27. 27.Yining Li, Chen Huang, Chen Change Loy, and Xiaoou Tang. Human attribute recognition by deep hierarchical contexts. In European conference on computer vision, pages 684–700. Springer, 2016. 2
  28. 28.Yanghao Li, Hanzi Mao, Ross B. Girshick, and Kaiming He. Exploring plain vision transformer backbones for object detection. ArXiv, abs/2203.16527, 2022. 2, 7
  29. 29.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017. 7
  30. 30.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 2
  31. 31.Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1096–1104, 2016. 2
  32. 32.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 11–20, 2016. 1, 3, 6
  33. 33.Panagiotis Meletis, Xiaoxiao Wen, Chenyang Lu, Daan de Geus, and Gijs Dubbelman. Cityscapes-panoptic-parts and pascal-panoptic-parts datasets for scene understanding. arXiv preprint arXiv:2004.07944, 2020. 2
  34. 34.Genevieve Patterson and James Hays. Coco attributes: Attributes for people, animals, and objects. European Conference on Computer Vision, 2016. 1
  35. 35.Khoi Pham, Kushal Kafle, Zhe Lin, Zhihong Ding, Scott Cohen, Quan Tran, and Abhinav Shrivastava. Learning to predict visual attributes in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13018–13028, 2021. 1, 2
  36. 36.Bernardino Romera-Paredes and Philip Torr. An embarrassingly simple approach to zero-shot learning. In International conference on machine learning, pages 2152–2161. PMLR, 2015. 3
  37. 37.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In Proceedings of the IEEE/CVF international conference on computer vision, pages 8430–8439, 2019. 2
  38. 38.Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. Flava: A foundational language and vision alignment model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15638–15650, 2022. 7
  39. 39.Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang, Wojciech Galuba, and Tal Hassner. Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8802–8812, 2021. 2
  40. 40.Xibin Song, Peng Wang, Dingfu Zhou, Rui Zhu, Chenye Guan, Yuchao Dai, Hao Su, Hongdong Li, and Ruigang Yang. Apollocar3d: A large 3d car instance understanding benchmark for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5452–5462, 2019. 2
  41. 41.Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, and Serge Belongie. Coco-text: Dataset and benchmark for text detection and recognition in natural images. arXiv preprint arXiv:1601.07140, 2016. 2, 4, 5, 6, 7
  42. 42.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The caltech-ucsd birds-200-2011 dataset. Technical report, 2011. 1, 2
  43. 43.Chenyun Wu, Zhe Lin, Scott Cohen, Trung Bui, and Subhransu Maji. Phrasecut: Language-based image segmentation in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10216–10225, 2020. 3
  44. 44.Yongqin Xian, Bernt Schiele, and Zeynep Akata. Zero-shot learning-the good, the bad and the ugly. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4582–4591, 2017. 3
  45. 45.Ke Yan, Youbao Tang, Yifan Peng, Veit Sandfort, Mohammadhadi Bagheri, Zhiyong Lu, and Ronald M Summers. Mulan: multitask universal lesion analysis network for joint lesion detection, tagging, and segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 194–202. Springer, 2019. 2
  46. 46.Lu Yang, Qing Song, Zhihui Wang, and Ming Jiang. Parsing r-cnn for instance-level human analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 364–373, 2019. 2
  47. 47.Shuai Zheng, Fan Yang, M Hadi Kiapour, and Robinson Piramuthu. Modanet: A large-scale street fashion dataset with polygon annotations. In Proceedings of the 26th ACM international conference on Multimedia, pages 1670–1678, 2018. 1, 2
  48. 48.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE transactions on pattern analysis and machine intelligence, 40(6):1452–1464, 2017. 2
  49. 49.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 633–641, 2017. 2
  50. 50.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ade20k dataset. International Journal of Computer Vision, 127(3):302–321, 2019. 1, 2
  51. 51.Xingyi Zhou, Rohit Girdhar, Armand Joulin, Phillip Krahenb ü uhl, and Ishan Misra. Detecting twenty-thousand ¨ classes using image-level supervision. arXiv preprint arXiv:2201.02605, 2022. 1, 8
  52. 52.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020. 7

Citation

MLA
Ramanathan, V., et al. “PACO: Parts and Attributes of Common Objects”. arXiv, 2023, http://arxiv.org/abs/2301.01795v1.
APA
Ramanathan, V., Kalia, A., Petrovic, V., Wen, Y., Zheng, B., Guo, B., Wang, R., Marquez, A., Kovvuri, R., Kadian, A., Mousavi, A., Song, Y., Dubey, A., & Mahajan, D. (2023). PACO: Parts and Attributes of Common Objects. arXiv. http://arxiv.org/abs/2301.01795v1
Chicago
Ramanathan, V., A. Kalia, V. Petrovic, et al. 2023. “PACO: Parts and Attributes of Common Objects”. arXiv. http://arxiv.org/abs/2301.01795v1.
Harvard
Ramanathan, V. et al. (2023) “PACO: Parts and Attributes of Common Objects”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.01795v1.
Vancouver
1. Ramanathan V, Kalia A, Petrovic V, et al (2023) PACO: Parts and Attributes of Common Objects. arXiv

BibTeX

@article{ramanathan2023paco,
  title = {PACO: Parts and Attributes of Common Objects},
  author = {Ramanathan, Vignesh and Kalia, Anmol and Petrovic, Vladan and Wen, Yi and Zheng, Baixue and Guo, Baishan and Wang, Rui and Marquez, Aaron and Kovvuri, Rama and Kadian, Abhishek and Mousavi, Amir and Song, Yiwen and Dubey, Abhimanyu and Mahajan, Dhruv},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.01795v1},
  eprint = {2301.01795}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/publicdomain/zero/1.0/