PACO: Parts and Attributes of Common Objects
Vignesh RamanathanAnmol KaliaVladan PetrovicYi WenBaixue ZhengBaishan GuoRui WangAaron MarquezRama KovvuriAbhishek Kadian
Introduces a large-scale visual dataset and benchmark spanning 75 object categories with over 640,000 part masks and detailed attribute annotations across images and videos to advance joint part segmentation and zero-shot instance detection.
Modern computer vision models increasingly need to provide detailed, fine-grained visual descriptions rather than simply identifying broad object category labels. Applications such as open-vocabulary retrieval and natural language interaction require systems to localize specific components of an object and recognize their detailed characteristics. However, existing public datasets either focus on narrow domains like fashion and birds or lack joint annotations that capture object masks, part masks, and attributes together across a broad set of common objects.
The article introduces PACO (Parts and Attributes of Common Objects), a large-scale benchmark designed to advance research in joint object detection, part segmentation, and attribute recognition across common objects. The objective is to establish an open-source evaluation suite and baseline performance benchmarks across image and video domains for fine-grained object understanding.
To construct PACO, the authors combined image data from the LVIS benchmark with video frames from the Ego4D dataset. The dataset spans 75 common object categories, 456 object-specific part categories, and 55 curated attributes covering colors, patterns, materials, and reflectance properties. In total, the authors collected 641,400 part masks across 260,300 object instances in approximately 81,500 images, making it roughly ten times larger in annotated object instances with parts than previous common object benchmarks. To establish standard reference points, the authors extended standard detection models—specifically Mask R-CNN and Vision Transformer architectures—and evaluated them across three benchmark tasks: part mask segmentation, instance-level attribute prediction, and zero-shot instance detection using multi-attribute queries.
The experimental findings show that segmenting parts and predicting fine-grained attributes are substantially more challenging than standard object detection. On part segmentation, the best Vision Transformer model achieved a mask Average Precision (AP) of 17.7 for object parts compared to 43.4 for whole objects, primarily because object parts are significantly smaller and visually distinct across different parent categories. On instance-level attribute detection, the highest-performing model scored an overall box AP of 19.5 on object attributes and 13.8 on part attributes; performance bounds analysis revealed that the theoretical upper bound with perfect attribute classification exceeds 72 AP, exposing a massive capability gap in current models. In zero-shot instance retrieval, queries with a single attribute yielded higher recall than complex two-attribute queries because prediction errors compounded as query complexity grew. Furthermore, off-the-shelf open-vocabulary detectors struggled significantly on fine-grained queries, registering low single-digit recall scores compared to task-adapted baselines.
These results indicate that current computer vision systems are fundamentally limited in their ability to resolve fine visual details and individual component properties. Improving performance in these areas is crucial for commercial systems that require precise localization, such as inventory management, robotics, and advanced search platforms, where misidentifying part attributes introduces significant operational error. The findings also demonstrate that open-vocabulary foundation models cannot yet reliably substitute for specialized, fine-grained visual detectors without targeted architectural enhancements.
Organizations developing fine-grained perception systems should adopt joint object-part-attribute training objectives and utilize larger vision transformer backbones, which consistently outperformed standard convolutional networks across all benchmark tasks. System architects should also pursue new modeling techniques that reduce compound attribute prediction errors in complex language queries and adapt open-vocabulary frameworks to handle dense attribute reasoning.
While the dataset provides rigorous quality control—including high agreement with expert gold annotations—limitations remain. The vocabulary is constrained to 75 common categories and 55 attributes, and the data distribution exhibits a long-tail imbalance where certain rare parts and attributes have limited examples. Nevertheless, the benchmark provides a highly credible and standardized foundation for evaluating progress in detailed visual understanding.
- Paper: LVIS: A Dataset for Large Vocabulary Instance Segmentation, Agrim Gupta et al. (2019). LVIS provides the core large-vocabulary image benchmark and instance segmentation framework upon which PACO directly builds its part- and attribute-level annotations.
- Paper: Ego4D: Around the World in 3,000 Hours of Egocentric Video, Kristen Grauman et al. (2021). Ego4D contributes the foundational egocentric video data that PACO utilizes to annotate object parts and attributes in video domains.
- Paper: Describing Objects by their Attributes, Ali Farhadi et al. (2009). This seminal paper introduced the paradigm of recognizing objects via visual attributes and part descriptors that PACO scales to dense modern benchmarks.
- Paper: Scene Parsing through ADE20K Dataset, Bolei Zhou et al. (2017). ADE20K pioneered hierarchical scene parsing with multi-level part segmentations, establishing foundational methodologies for evaluating part masks alongside objects.
- Paper: Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations, Ranjay Krishna et al. (2016). Visual Genome established the standard for dense crowdsourced object and attribute annotations, motivating PACO's structured part-attribute ontology.
- Paper: Mask R-CNN, Kaiming He et al. (2017). Mask R-CNN provides the standard instance segmentation baseline architecture and evaluation metrics extended by PACO for part mask segmentation.
- Paper: Masked-attention Mask Transformer for Universal Image Segmentation, Bowen Cheng et al. (2022). Mask2Former represents the modern universal image segmentation architecture evaluated as a primary baseline on PACO's part and instance segmentation benchmarks.
- Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). ViLD establishes vision-language knowledge distillation methods for zero-shot object detection that inform PACO's zero-shot instance detection evaluation.
- Paper: Object Detection with Discriminatively Trained Part-Based Models, Pedro F. Felzenszwalb et al. (2010). This paper establishes the foundational computer vision principle of representing objects via deformable configurations of semantic parts.
- Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Segment Anything scales promptable mask prediction to a foundation model level, evaluating zero-shot part and object delineation across fine-grained datasets like PACO.
