Simple Multi-dataset Detection

Xingyi ZhouVladlen KoltunPhilipp Krähenbühl

article2022CVPR165 citations

Presents a multi-dataset object detection framework that automatically integrates disparate label spaces via 0-1 integer programming without manual taxonomy reconciliation, enabling a single detector to match dataset-specific baselines and generalize directly to unseen domains.

Listen

Building general-purpose computer vision systems requires models that can detect a wide range of objects across varied environments. However, current object detection models remain largely confined to individual datasets, which have limited vocabularies, differing category definitions, and distinct training protocols. Past attempts to combine multiple datasets relied heavily on time-consuming manual taxonomy reconciliation, struggled with imbalanced data distributions, or experienced notable performance degradation when evaluated across domains.

The article demonstrates an automated method for training a unified object detector across disparate large-scale datasets. The primary objective is to train a single high-performing model that automatically reconciles inconsistent label spaces into a shared taxonomy using visual data alone, without requiring manual human mapping.

To achieve this, the authors first trained a shared backbone network with dataset-specific classification heads. This partitioned detector used tailored sampling strategies and loss functions, such as hierarchy-aware objectives, for each dataset. Next, the authors formulated an integer linear optimization problem to automatically merge category labels based on the visual prediction correlations of the partitioned model. They evaluated the framework at scale on COCO, Objects365, and OpenImages—encompassing 945 original classes—and tested transfer performance on seven distinct, unseen benchmarks.

The findings show that the automated taxonomy optimization successfully condensed 945 disjoint classes into a cohesive 701-concept vocabulary, outperforming human-expert and language-based baselines across all training sets. When given sufficient training iterations, the unified multi-dataset model matched or surpassed the accuracy of dataset-specific models on their native domains. In zero-shot cross-dataset evaluations on benchmarks like Pascal VOC, ScanNet, and Cityscapes, the unified detector averaged a 47.3 mean Average Precision (mAP50), outperforming both single-dataset baselines and multi-model ensembles. Finally, scaling up to a large ResNeSt200 backbone established top-tier performance on COCO (52.9 mAP) and Objects365 (33.7 mAP), outperforming the previous competition-winning Objects365 model by 2 mAP points.

These results demonstrate that multi-dataset training no longer requires manual label unification or sacrificial trade-offs in per-domain accuracy. By eliminating the need to know the target domain at test time, the resulting unified detector reduces deployment complexity, mitigates duplicate classification errors, and improves generalization in real-world out-of-domain environments.

Organizations developing broad computer vision applications should adopt automated visual taxonomy reconciliation to scale up perception systems across legacy and newly labeled datasets. Teams should avoid simple data concatenation in favor of balanced dataset sampling and loss-specific supervision. For immediate implementation, development can leverage the authors' open-source codebase.

Confidence in these findings is supported by extensive empirical validation across multiple standard benchmarks and repeated experimental runs. However, key limitations remain: the current taxonomy formulation optimizes strictly using visual cues rather than integrating textual semantics, and it treats hierarchical labels as distinct classes rather than formal parent-child relationships. Future work should address hierarchical reasoning to further enhance semantic consistency.

  • Paper: Unbiased look at dataset bias, Antonio Torralba et al. (2011). This seminal paper quantifies cross-dataset bias and the resulting drop in generalization performance, which directly motivates the multi-dataset unified detection paradigm.
  • Paper: The Open Images Dataset V4, Alina Kuznetsova et al. (2018). This paper presents the large-scale, hierarchical Open Images benchmark that serves as one of the core multi-class training datasets unified and evaluated in the source work.
  • Paper: Unified Perceptual Parsing for Scene Understanding, Tete Xiao et al. (2018). This work establishes multi-dataset learning across heterogeneous annotations using task-specific heads and selective sampling strategies that directly inform the partitioned detector baseline.
  • Paper: Focal Loss for Dense Object Detection, Tsung-Yi Lin et al. (2017). This paper introduces Focal Loss and one-stage dense detection principles necessary for understanding loss weighting and class imbalance management in large-scale detector heads.
  • Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). This study introduces open-vocabulary distillation across disparate category spaces like LVIS and Objects365, contextualizing the challenges of scaling vocabularies beyond single-dataset boundaries.
Cover for Simple Multi-dataset Detection

Abstract

How do we build a general and broad object detection system? We use all labels of all concepts ever annotated. These labels span diverse datasets with potentially inconsistent taxonomies. In this paper, we present a simple method for training a unified detector on multiple large-scale datasets. We use dataset-specific training protocols and losses, but share a common detection architecture with dataset-specific outputs. We show how to automatically integrate these dataset-specific outputs into a common semantic taxonomy. In contrast to prior work, our approach does not require manual taxonomy reconciliation. Experiments show our learned taxonomy outperforms a expert-designed taxonomy in all datasets. Our multi-dataset detector performs as well as dataset-specific models on each training domain, and can generalize to new unseen dataset without fine-tuning on them. Code is available at https://github.com/xingyizhou/UniDet.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 4. Training a multi-dataset detector
  • 4.1. Learning a unified label space
  • 4.2. Loss functions
  • 5. Experiments
  • 5.1. Multi-dataset detection
  • 5.2. Unified multi-dataset detection
  • 5.3. Cross-dataset evaluation
  • 5.4. Scale up to large models
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Partitioned Multi-Dataset Object Detector Architecture and Training Protocol

    model/method

    To train a multi-dataset object detector across KK distinct datasets D1,…,DK\mathcal{D}_1, \dots, \mathcal{D}_K with disparate label spaces L1,…,LK\mathcal{L}_1, \dots, \mathcal{L}_K and differing dataset-specific loss functions ℓ1,…,ℓK\ell_1, \dots, \ell_K, the architecture shares a single common backbone network (e.g., Cascade R-CNN) and region proposal network (RPN) while assigning a separate classification head to each dataset Dk\mathcal{D}_k. For cascade architectures, the classification layers across all cascade stages are split per dataset.

    The partitioned detector is trained across all datasets simultaneously by optimizing the expected loss over all datasets:

    min⁡ΘEDk[E(I^,B^)∼Dk[ℓk(Mk(I^;Θ),B^)]]\min_{\Theta} \mathbb{E}_{\mathcal{D}_k} \left[ \mathbb{E}_{(\hat{I}, \hat{B}) \sim \mathcal{D}_k} [\ell_k(M_k(\hat{I}; \Theta), \hat{B})] \right]

    where Θ\Theta represents the shared network parameters, I^\hat{I} is an image sampled from Dk\mathcal{D}_k, B^\hat{B} denotes its ground-truth bounding boxes, and Mk(⋅)M_k(\cdot) is the dataset-specific output head.

    The training protocol integrates three components:

    1. Uniform dataset sampling: Each dataset Dk\mathcal{D}_k is sampled with equal probability ({1/K}\{1/K\}) per iteration, ensuring small datasets are not dominated by large ones.
    2. Intra-dataset class-aware sampling: Long-tailed datasets (such as Objects365 and OpenImages) use class-aware image sampling within the dataset.
    3. Dataset-specific hierarchy-aware losses: For datasets with hierarchical annotations (such as OpenImages), the detector uses a hierarchy-aware sigmoid cross-entropy loss that treats all parent classes of an annotated class as positive targets while ignoring descendant class losses, avoiding the mutual exclusivity assumption of standard cross-entropy.
  2. Knowl 2 — 0-1 Integer Linear Program Formulation for Automatic Label Space Unification

    model/method

    Given KK datasets with label spaces L1,…,LK\mathcal{L}_1, \dots, \mathcal{L}_K, automatic taxonomy reconciliation seeks a unified label space L\mathcal{L} and a collection of Boolean projection matrices Tk∈{0,1}∣Lk∣×∣L∣T_k \in \{0, 1\}^{|\mathcal{L}_k| \times |\mathcal{L}|} mapping the unified label space to each dataset-specific label space.

    The mapping adheres to direct 1-to-1 matching constraints:

    1. Each dataset-specific class matches exactly one unified class: Tk1=1T_k \mathbf{1} = \mathbf{1}.
    2. Each unified class contains at most one class from each dataset: Tk⊤1≤1T_k^\top \mathbf{1} \le \mathbf{1}.

    Given partitioned detector output vectors dik∈R∣Lk∣d_i^k \in \mathbb{R}^{|\mathcal{L}_k|} for a bounding box bib_i, the joint detection score vector di∈R∣L∣d_i \in \mathbb{R}^{|\mathcal{L}|} is computed via elementwise averaging over participating datasets:

    di=∑kTk⊤dik∑kTk⊤1d_i = \frac{\sum_k T_k^\top d_i^k}{\sum_k T_k^\top \mathbf{1}}

    Dataset-specific predictions are reconstructed via d~ik=Tkdi\tilde{d}_i^k = T_k d_i.

    Let T=T1×⋯×TK\mathcal{T} = \mathcal{T}_1 \times \dots \times \mathcal{T}_K be the Cartesian product of possible column assignments, where t∈Tt \in \mathcal{T} represents a class combination across datasets. Introducing binary decision variables xt∈{0,1}x_t \in \{0, 1\} indicating whether combination tt is included in the unified label space L\mathcal{L}, the taxonomy optimization is formulated as a 0-1 Integer Linear Program (ILP):

    min⁡x∑t∈Txt(ct+λ)subject to∑t∈T∣t(c)=1xt=1∀c∈⋃kLk\min_{x} \sum_{t \in \mathcal{T}} x_t (c_t + \lambda) \quad \text{subject to} \quad \sum_{t \in \mathcal{T} \mid t(c) = 1} x_t = 1 \quad \forall c \in \bigcup_k \mathcal{L}_k

    where λ>0\lambda > 0 is a cardinality regularization penalty penalizing large unified label spaces ∣L∣=∑txt|\mathcal{L}| = \sum_t x_t, and ctc_t is the precomputed merge cost for combining the classes in tuple tt. For K=2K = 2, this formulation is equivalent to minimum-weight bipartite matching; for K≥3K \ge 3, it corresponds to weighted multidimensional graph matching.

  3. Knowl 3 — Merge Cost Objectives for Label Space Optimization

    equation

    The merge cost ctc_t for a potential class combination t∈Tt \in \mathcal{T} across datasets in the label unification integer linear program is defined as the expected class-wise loss incurred when replacing partitioned detector scores DckD_c^k with reprojected unified scores D~ck\tilde{D}_c^k:

    ct=EDk[∑c∈Lk∣t(c)=1Lc(Dck,D~ck)]c_t = \mathbb{E}_{\mathcal{D}_k} \left[ \sum_{c \in \mathcal{L}_k \mid t(c) = 1} \mathcal{L}_c(D_c^k, \tilde{D}_c^k) \right]

    Two specific loss functions Lc\mathcal{L}_c are utilized:

    1. Output Score Distortion: An unsupervised metric measuring the Euclidean distance between partitioned and reprojected detection score distributions:

    Lcdist(Dck,D~ck)=(Dck−D~ck)2\mathcal{L}_c^{\text{dist}}(D_c^k, \tilde{D}_c^k) = \left( D_c^k - \tilde{D}_c^k \right)^2

    1. Validation Average Precision Degradation: A supervised metric measuring the drop in class-level Average Precision on the validation set of dataset Dk\mathcal{D}_k:

    LcAP(Dck,D~ck)=1∣Lk∣(APc(Dck)−APc(D~ck))\mathcal{L}_c^{\text{AP}}(D_c^k, \tilde{D}_c^k) = \frac{1}{|\mathcal{L}_k|} \left( \text{AP}_c(D_c^k) - \text{AP}_c(\tilde{D}_c^k) \right)

  4. Knowl 4 — Effectiveness of Multi-Dataset Training Strategies

    empirical result

    Evaluating multi-dataset training components on COCO (80 classes), Objects365 (365 classes), and OpenImages (500 classes) demonstrates that naive dataset pooling degrades accuracy on smaller datasets, while uniform dataset sampling, class-aware intra-dataset sampling, and hierarchy-aware losses are necessary to achieve optimal performance across all domains.

    Method COCO Objects365 OpenImages Mean
    Simple merge 34.2 14.6 50.8 33.2
    w/ uniform dataset sampling 41.1 16.5 46.0 34.5
    w/ class-aware sampling 35.3 18.5 61.8 38.5
    w/ dataset + class-aware sampling 41.8 20.3 60.0 40.6
    Partitioned detector (ours) 41.8 20.6 62.7 41.7

    All models use a ResNet-50 Cascade R-CNN trained with a 2×2\times schedule (180k iterations). Simple merging heavily biases towards OpenImages due to its 18×18\times larger image volume. Combining uniform dataset-level sampling with intra-dataset class-aware sampling improves the mean mAP from 33.2 to 40.6. Incorporating the dataset-specific hierarchy-aware sigmoid cross-entropy loss for OpenImages provides an additional +2.7+2.7 mAP improvement on OpenImages (reaching 62.7 mAP) without degrading COCO or Objects365 performance.

  5. Knowl 5 — Convergence and Training Schedule Scaling for Partitioned vs Single-Dataset Models

    empirical result

    When comparing a partitioned multi-dataset detector to individual dataset-specific models across training schedules (2×=180k2\times = 180\text{k}, 6×=540k6\times = 540\text{k}, and 8×=720k8\times = 720\text{k} iterations), the partitioned detector requires longer schedules to match single-dataset models because each dataset receives only 1/K1/K of the gradient updates per epoch under uniform dataset sampling.

    2×2\times 6×6\times 8×8\times
    Model COCO O365 OImg COCO O365 OImg COCO O365 OImg
    Partitioned detector 41.8 20.6 62.7 44.6 23.6 64.8 45.5 24.6 66.0
    COCO only 41.5 – – 42.5 – – 42.5 – –
    Objects365 only – 23.8 – – 25.0 – – 24.9 –
    OpenImages only – – 64.6 – – 65.4 – – 65.7

    At a 2×2\times schedule, single-dataset models outperform the partitioned model on Objects365 and OpenImages. At a 6×6\times schedule, the partitioned model matches or exceeds single-dataset models. At an 8×8\times schedule, the partitioned model converges, outperforming the single-dataset model on COCO (45.545.5 vs 42.542.5 mAP) and matching performance on Objects365 (24.624.6 vs 24.924.9 mAP) and OpenImages (66.066.0 vs 65.765.7 mAP).

  6. Knowl 6 — Taxonomy Quality: Automatically Learned vs Expert Human and Language Baselines

    empirical result

    Unifying 945 disjoint classes from COCO, Objects365, and OpenImages using the 0-1 ILP formulation produces a unified label space of ∣L∣=701|\mathcal{L}| = 701 classes that outperforms both human expert manual unification and GloVe word-embedding-based unification when retraining a ResNet-50 Cascade R-CNN (2×2\times schedule, mean ±\pm std over 3 runs).

    Unification Method ∣L∣|\mathcal{L}| COCO Objects365 OpenImages Mean mAP
    GloVe embedding 696 41.6 ±\pm 0.00 20.3 ±\pm 0.12 62.4 ±\pm 0.06 41.4 ±\pm 0.05
    Learned, distortion 682 41.6 ±\pm 0.15 20.7 ±\pm 0.06 62.6 ±\pm 0.06 41.7 ±\pm 0.09
    Learned, AP (ours) 701 41.9 ±\pm 0.10 20.8 ±\pm 0.10 63.0 ±\pm 0.21 41.9 ±\pm 0.02
    Expert human 659 41.5 ±\pm 0.06 20.7 ±\pm 0.06 62.6 ±\pm 0.06 41.6 ±\pm 0.04

    The visual AP-learned label space achieves a mean mAP of 41.941.9, exceeding the human expert taxonomy (41.641.6 mAP) by 0.30.3 mAP. Visual unification correctly merges semantically equivalent classes with differing vocabulary names (e.g., OpenImages "cattle" and COCO/Objects365 "cow") while separating visually distinct concepts sharing identical names (e.g., distinguishing COCO ovens with cooktops, OpenImages ovens with control panels, and Objects365 oven frontal doors).

  7. Knowl 7 — Performance of Unified Retrained Detector vs Partitioned and Oracle Ensembles

    empirical result

    A comparison on the training domains (COCO, Objects365, OpenImages) using ResNet-50 Cascade R-CNN under an 8×8\times training schedule shows that retraining on the automatically unified taxonomy recovers accuracy lost during naive weight merging and matches domain-aware partitioned models without requiring test-time dataset identification.

    Model COCO Objects365 OpenImages Mean mAP
    Unified (naive merge) 44.4 23.6 65.3 44.4
    Unified (retrained) 45.4 24.4 66.0 45.3
    Partitioned (oracle) 45.5 24.6 66.0 45.4
    Ensemble (oracle) 42.5 24.9 65.7 44.4

    While naive offline linear merging of partitioned weights yields 44.444.4 mean mAP, retraining under the unified label space achieves 45.345.3 mean mAP. This matches the partitioned detector oracle (45.445.4 mAP) and outperforms an ensemble of three single-dataset models (44.444.4 mAP), while operating as a single unified head requiring no test domain routing.

  8. Knowl 8 — Zero-Shot Cross-Dataset Transfer to Unseen Target Domains

    empirical result

    Evaluating object detectors on 7 unseen test datasets (Pascal VOC, VIPER, Cityscapes, ScanNet, WildDash, CrowdHuman, KITTI) using GloVe nearest-neighbor label matching demonstrates that multi-dataset detectors generalize substantially better than single-dataset models or their 4-model ensemble.

    Training Data VOC VIPER Cityscapes ScanNet WildDash CrowdHuman KITTI Mean
    COCO 80.0 13.9 39.6 17.4 25.9 73.9 30.5 40.2
    Objects365 71.9 20.7 43.4 24.9 27.6 71.8 32.2 41.8
    OpenImages 64.4 10.4 29.8 24.2 20.3 66.7 21.8 33.9
    Mapillary 11.4 15.2 44.7 0.0 23.4 49.3 37.8 26.0
    Ensemble (4 models) 79.7 16.8 46.0 30.1 32.1 73.9 34.3 44.7
    Partitioned 83.1 20.9 48.4 32.2 34.4 70.0 38.9 46.8
    Unified (retrained) 82.9 21.3 52.6 29.8 34.7 70.7 39.9 47.3
    Dataset-specific (oracle) 80.3 31.8 54.6 44.7 – 80.0 – –

    All models use ResNet-50 Cascade R-CNN. The retrained unified detector achieves a mean mAP50 of 47.347.3, outperforming the best single-dataset model (Objects365 at 41.841.8) by +5.5+5.5 mAP50 and the 4-model ensemble (44.744.7) by +2.6+2.6 mAP50. On Pascal VOC, the multi-dataset unified detector (82.982.9 mAP50) exceeds the in-domain trained VOC oracle (80.380.3 mAP50) without training on any VOC images.

  9. Knowl 9 — Scaling Unified Object Detection with ResNeSt-200

    empirical result

    Scaling the unified detector architecture to a ResNeSt-200 backbone trained jointly on COCO, OpenImages, Objects365, and Mapillary using an 8×8\times schedule matches or exceeds specialized state-of-the-art single-dataset models across major benchmarks.

    Model COCO (test-dev) OpenImages (test) Mapillary (test) Objects365 (val)
    Unified ResNeSt-200 (ours) 52.9 60.6 / 56.8 25.3 33.7
    ResNeSt-200 (single) 50.9 – – –
    TSD (SENet154-DCN) – 60.5 / – – –
    CACascade RCNN – – – 31.6

    On COCO test-challenge, the unified model achieves 52.952.9 mAP, outperforming the single-dataset ResNeSt-200 baseline (50.950.9 mAP) by +2.0+2.0 mAP. On OpenImages 2019 Challenge test sets (public/private), the unified model achieves 60.6/56.860.6 / 56.8 mAP, matching the 2019 challenge winner TSD (60.560.5 mAP). On Objects365, the single model obtains 33.733.7 mAP, surpassing the 2019 Objects365 challenge winner CACascade R-CNN (31.631.6 mAP) by +2.1+2.1 mAP.

  10. Knowl 10 — Limitations of Vision-Only Flat Taxonomy Learning

    limitation

    The taxonomy learning formulation has two primary limitations:

    1. Sole Reliance on Visual Cues: The 0-1 ILP optimization relies exclusively on visual detector firing correlations or validation AP, without incorporating linguistic embeddings as auxiliary cues during optimization.
    2. Flat Direct-Mapping Assumption: The formulation enforces a 1-to-1 direct mapping constraint (Tk⊤1≤1T_k^\top \mathbf{1} \le \mathbf{1}) and does not model label hierarchies across datasets. Consequently, hierarchical parent-child relationships across datasets (e.g., COCO "person" and OpenImages "boy") are treated as completely independent disjoint classes in the unified output space rather than structured hierarchical concepts.

Coverage note — None was omitted; all key algorithmic contributions (partitioned detector training, ILP taxonomy optimization formulation, loss functions), major ablation tables (sampling/loss ablations, schedule scaling, taxonomy comparisons, unified vs partitioned), generalization results (cross-dataset zero-shot evaluation, ResNeSt-200 scaling), and limitations are included.

References

  1. 1.Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran. Zero-shot object detection. In ECCV, 2018. 2
  2. 2.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: High quality object detection and instance segmentation. TPAMI, 2019. 5
  3. 3.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016. 5
  4. 4.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017. 5
  5. 5.Xiyang Dai, Yinpeng Chen, Bin Xiao, Dongdong Chen, Mengchen Liu, Lu Yuan, and Lei Zhang. Dynamic head: Unifying object detection heads with attentions. In CVPR, 2021. 1
  6. 6.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The PASCAL visual object classes (VOC) challenge. IJCV, 2010. 2, 5
  7. 7.Ali Farhadi, Ian Endres, Derek Hoiem, and David Forsyth. Describing objects by their attributes. In CVPR, 2009. 2
  8. 8.Yanwei Fu, Tao Xiang, Yu-Gang Jiang, Xiangyang Xue, Leonid Sigal, and Shaogang Gong. Recent advances in zero-shot recognition: Toward data-efficient understanding of visual content. IEEE Signal Processing Magazine, 2018. 2
  9. 9.Yuan Gao, Hui Shen, Donghong Zhong, Jian Wang, Zeyu Liu, Ti Bai, Xiang Long, and Shilei Wen. A solution for densely annotated large scale object detection task. 2019. 3, 8
  10. 10.Zheng Ge, Songtao Liu, Feng Wang, Zeming Li, and Jian Sun. Yolox: Exceeding yolo series in 2021. arXiv:2107.08430, 2021. 1
  11. 11.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR, 2012. 5
  12. 12.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In CVPR, 2019. 2
  13. 13.Irtiza Hasan, Shengcai Liao, Jinpeng Li, Saad Ullah Akram, and Ling Shao. Generalizable pedestrian detection: The elephant in the room. In CVPR, 2021. 2
  14. 14.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In CVPR, 2017. 1
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 5
  16. 16.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In CVPR, 2018. 8
  17. 17.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 2017. 2
  18. 18.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. IJCV, 2020. 1, 3, 5
  19. 19.John Lambert, Zhuang Liu, Ozan Sener, James Hays, and Vladlen Koltun. MSeg: A composite dataset for multi-domain semantic segmentation. In CVPR, 2020. 1, 2, 3, 5
  20. 20.Xiang Li, Wenhai Wang, Xiaolin Hu, Jun Li, Jinhui Tang, and Jian Yang. Generalized focal loss v2: Learning reliable localization quality estimation for dense object detection. CVPR, 2021. 1
  21. 21.Zhihui Li, Lina Yao, Xiaoqin Zhang, Xianzhi Wang, Salil Kanhere, and Huaxiang Zhang. Zero-shot object detection with textual descriptions. In AAAI, 2019. 2
  22. 22.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 1, 2, 5
  23. 23.Jeffrey T Linderoth and Ted K Ralphs. Noncommercial software for mixed-integer linear programming. Integer programming: theory and practice, 3(253-303):144–189, 2005. 4
  24. 24.Gerhard Neuhold, Tobias Ollmann, Samuel Rota Bulo`, and Peter Kontschieder. The mapillary vistas dataset for semantic understanding of street scenes. In ICCV, 2017. 1, 5, 7
  25. 25.Mohammad Norouzi, Tomas Mikolov, Samy Bengio, Yoram Singer, Jonathon Shlens, Andrea Frome, Greg S Corrado, and Jeffrey Dean. Zero-shot learning by convex combination of semantic embeddings. In ICLR, 2014. 2
  26. 26.Junran Peng, Xingyuan Bu, Ming Sun, Zhaoxiang Zhang, Tieniu Tan, and Junjie Yan. Large-scale object detection in the wild from imbalanced multi-labels. In CVPR, 2020. 3, 5
  27. 27.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In EMNLP, 2014. 7
  28. 28.Shafin Rahman, Salman Khan, and Nick Barnes. Transductive learning for zero-shot object detection. In ICCV, 2019. 2
  29. 29.Rene´ Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. TPAMI, 2020. 2
  30. 30.Joseph Redmon and Ali Farhadi. Yolo9000: better, faster, stronger. In CVPR, 2017. 2
  31. 31.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In NIPS, 2015. 1
  32. 32.Stephan R Richter, Zeeshan Hayder, and Vladlen Koltun. Playing for benchmarks. In ICCV, 2017. 5
  33. 33.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In ICCV, 2019. 1, 3, 5
  34. 34.Shuai Shao, Zijian Zhao, Boxun Li, Tete Xiao, Gang Yu, Xiangyu Zhang, and Jian Sun. Crowdhuman: A benchmark for detecting human in a crowd. arXiv:1805.00123, 2018. 5
  35. 35.Li Shen, Zhouchen Lin, and Qingming Huang. Relay backpropagation for effective learning of deep convolutional neural networks. In ECCV, 2016. 3, 5
  36. 36.Guanglu Song, Yu Liu, and Xiaogang Wang. Revisiting the sibling head in object detector. In CVPR, 2020. 8
  37. 37.Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In CVPR, 2015. 2
  38. 38.Jingru Tan, Changbao Wang, Buyu Li, Quanquan Li, Wanli Ouyang, Changqing Yin, and Junjie Yan. Equalization loss for long-tailed object recognition. In CVPR, 2020. 3
  39. 39.Chien-Yao Wang, Alexey Bochkovskiy, and Hong-Yuan Mark Liao. Scaled-yolov4: Scaling cross stage partial network. CVPR, 2021. 1
  40. 40.Xudong Wang, Zhaowei Cai, Dashan Gao, and Nuno Vasconcelos. Towards universal object detection by domain attention. In CVPR, 2019. 2, 5
  41. 41.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019. 2, 5
  42. 42.Hang Xu, Linpu Fang, Xiaodan Liang, Wenxiong Kang, and Zhenguo Li. Universal-rcnn: Universal object detector via transferable graph r-cnn. In AAAI, 2020. 2
  43. 43.Gengshan Yang, Joshua Manela, Michael Happold, and Deva Ramanan. Hierarchical deep stereo matching on high-resolution images. In CVPR, 2019. 2
  44. 44.Oliver Zendel, Katrin Honauer, Markus Murschitz, Daniel Steininger, and Gustavo Fernandez Dominguez. Wilddashcreating hazard-aware benchmarks. In ECCV, 2018. 5
  45. 45.Haoyang Zhang, Ying Wang, Feras Dayoub, and Niko Su¨nderhauf. Varifocalnet: An iou-aware dense object detector. CVPR, 2021. 1
  46. 46.Hang Zhang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Zhi Zhang, Haibin Lin, Yue Sun, Tong He, Jonas Mueller, R Manmatha, et al. Resnest: Split-attention networks. arXiv: 2004.08955, 2020. 6, 8
  47. 47.Xiangyun Zhao, Samuel Schulter, Gaurav Sharma, Yi-Hsuan Tsai, Manmohan Chandraker, and Ying Wu. Object detection with a unified label space from multiple datasets. In ECCV, 2020. 1, 2, 3
  48. 48.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In CVPR, 2017. 2

Citation

MLA
Zhou, X., et al. “Simple Multi-dataset Detection”. arXiv, 2021, http://arxiv.org/abs/2102.13086v2.
APA
Zhou, X., Koltun, V., & Krähenbühl, P. (2021). Simple multi-dataset detection. arXiv. http://arxiv.org/abs/2102.13086v2
Chicago
Zhou, X., V. Koltun, and P. Krähenbühl. 2021. “Simple Multi-dataset Detection”. arXiv. http://arxiv.org/abs/2102.13086v2.
Harvard
Zhou, X., Koltun, V. and Krähenbühl, P. (2021) “Simple multi-dataset detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2102.13086v2.
Vancouver
1. Zhou X, Koltun V, Krähenbühl P (2021) Simple multi-dataset detection. arXiv

BibTeX

@article{zhou2021simple,
  title = {Simple multi-dataset detection},
  author = {Zhou, Xingyi and Koltun, Vladlen and Krähenbühl, Philipp},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2102.13086v2},
  eprint = {2102.13086}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE