Simple Multi-dataset Detection
Xingyi ZhouVladlen KoltunPhilipp Krähenbühl
Presents a multi-dataset object detection framework that automatically integrates disparate label spaces via 0-1 integer programming without manual taxonomy reconciliation, enabling a single detector to match dataset-specific baselines and generalize directly to unseen domains.
Building general-purpose computer vision systems requires models that can detect a wide range of objects across varied environments. However, current object detection models remain largely confined to individual datasets, which have limited vocabularies, differing category definitions, and distinct training protocols. Past attempts to combine multiple datasets relied heavily on time-consuming manual taxonomy reconciliation, struggled with imbalanced data distributions, or experienced notable performance degradation when evaluated across domains.
The article demonstrates an automated method for training a unified object detector across disparate large-scale datasets. The primary objective is to train a single high-performing model that automatically reconciles inconsistent label spaces into a shared taxonomy using visual data alone, without requiring manual human mapping.
To achieve this, the authors first trained a shared backbone network with dataset-specific classification heads. This partitioned detector used tailored sampling strategies and loss functions, such as hierarchy-aware objectives, for each dataset. Next, the authors formulated an integer linear optimization problem to automatically merge category labels based on the visual prediction correlations of the partitioned model. They evaluated the framework at scale on COCO, Objects365, and OpenImages—encompassing 945 original classes—and tested transfer performance on seven distinct, unseen benchmarks.
The findings show that the automated taxonomy optimization successfully condensed 945 disjoint classes into a cohesive 701-concept vocabulary, outperforming human-expert and language-based baselines across all training sets. When given sufficient training iterations, the unified multi-dataset model matched or surpassed the accuracy of dataset-specific models on their native domains. In zero-shot cross-dataset evaluations on benchmarks like Pascal VOC, ScanNet, and Cityscapes, the unified detector averaged a 47.3 mean Average Precision (mAP50), outperforming both single-dataset baselines and multi-model ensembles. Finally, scaling up to a large ResNeSt200 backbone established top-tier performance on COCO (52.9 mAP) and Objects365 (33.7 mAP), outperforming the previous competition-winning Objects365 model by 2 mAP points.
These results demonstrate that multi-dataset training no longer requires manual label unification or sacrificial trade-offs in per-domain accuracy. By eliminating the need to know the target domain at test time, the resulting unified detector reduces deployment complexity, mitigates duplicate classification errors, and improves generalization in real-world out-of-domain environments.
Organizations developing broad computer vision applications should adopt automated visual taxonomy reconciliation to scale up perception systems across legacy and newly labeled datasets. Teams should avoid simple data concatenation in favor of balanced dataset sampling and loss-specific supervision. For immediate implementation, development can leverage the authors' open-source codebase.
Confidence in these findings is supported by extensive empirical validation across multiple standard benchmarks and repeated experimental runs. However, key limitations remain: the current taxonomy formulation optimizes strictly using visual cues rather than integrating textual semantics, and it treats hierarchical labels as distinct classes rather than formal parent-child relationships. Future work should address hierarchical reasoning to further enhance semantic consistency.
- Paper: Unbiased look at dataset bias, Antonio Torralba et al. (2011). This seminal paper quantifies cross-dataset bias and the resulting drop in generalization performance, which directly motivates the multi-dataset unified detection paradigm.
- Paper: The Open Images Dataset V4, Alina Kuznetsova et al. (2018). This paper presents the large-scale, hierarchical Open Images benchmark that serves as one of the core multi-class training datasets unified and evaluated in the source work.
- Paper: Unified Perceptual Parsing for Scene Understanding, Tete Xiao et al. (2018). This work establishes multi-dataset learning across heterogeneous annotations using task-specific heads and selective sampling strategies that directly inform the partitioned detector baseline.
- Paper: Focal Loss for Dense Object Detection, Tsung-Yi Lin et al. (2017). This paper introduces Focal Loss and one-stage dense detection principles necessary for understanding loss weighting and class imbalance management in large-scale detector heads.
- Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). This study introduces open-vocabulary distillation across disparate category spaces like LVIS and Objects365, contextualizing the challenges of scaling vocabularies beyond single-dataset boundaries.
- Paper: Uni3D: A Unified Baseline for Multi-Dataset 3D Object Detection, Bo Zhang et al. (2023). This paper extends the concept of multi-dataset unification and taxonomy reconciliation from 2D vision models into 3D LiDAR object detection.
- Paper: Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt Training, Xiaoyang Wu et al. (2024). This work builds on multi-dataset representation learning by using language guidance and prompt training to overcome negative transfer across conflicting 3D point cloud label spaces.
- Paper: Hierarchical Open-vocabulary Universal Image Segmentation, Xudong Wang et al. (2023). This work advances multi-dataset open-vocabulary parsing by explicitly addressing the structural and hierarchical label ambiguities identified as a key limitation of visual-only taxonomy merging.
- Paper: PACO: Parts and Attributes of Common Objects, Vignesh Ramanathan et al. (2023). This research provides a fine-grained part and attribute benchmark that directly tests multi-granularity detection beyond standard bounding-box object taxonomies.
