HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining

Shixiang TangCheng ChenQingsong XieMeilin ChenYizhou WangYuanzheng CiLei BaiFeng ZhuHaiyang YangLi Yi

article2023CVPR59 citations

Presents HumanBench, a 19-dataset benchmark spanning six human-centric vision tasks, alongside PATH, a hierarchical projector-assisted pretraining framework that achieves state-of-the-art transfer performance across diverse granularities of human perception.

Listen

Human-centric perception tasks—including person re-identification, pose estimation, human parsing, pedestrian attribute recognition, and pedestrian detection—are critical for industries such as autonomous driving, surveillance, and virtual environments. Historically, organizations have developed specialized models for each distinct task, resulting in high engineering costs, duplicated data labeling efforts, and inefficient multi-model deployment pipelines. While general vision foundation models exist, they are primarily trained on natural scenes or language associations rather than the multi-scale, fine-grained structural features inherent to human bodies.

The article evaluates whether a unified, multi-dataset supervised pretraining framework can create a single, highly generalizable human-centric visual representation. It demonstrates a scalable methodology to overcome task conflicts and annotation discrepancies across disparate human vision tasks, establishing a standardized evaluation benchmark to measure performance.

To achieve this, the article introduces HumanBench, a large-scale benchmark comprising 11 million images across 37 datasets covering five core human-centric tasks. To resolve the negative task conflicts common in multi-task learning, the article proposes a pretraining framework called Projector Assisted Hierarchical Pretraining (PATH). This method employs a hierarchical weight-sharing architecture: a single transformer backbone is shared across all tasks, lightweight attention and gating projectors are shared strictly among datasets of the same task, and independent output heads handle dataset-specific distributions. During downstream evaluation, the projectors and heads are discarded, assessing the pure backbone's transferable features through full finetuning, head-only finetuning, and partial layer finetuning.

The findings show that PATH establishes state-of-the-art results across 17 downstream datasets and matches top existing methods on 2 others. First, when adapted to 12 seen pretraining datasets, PATH outperforms prior specialized models in major categories, improving human parsing accuracy by up to 2.5% mean intersection-over-union and person re-identification by 4.9% mean average precision. Second, PATH demonstrates robust out-of-dataset generalization across 5 unseen datasets, improving recognition accuracy by 4.4% and reducing pedestrian detection miss rates. Third, when evaluated on crowd counting—a completely unseen task outside the pretraining scope—PATH reduces mean squared error by up to 10.4% compared to standard self-supervised baselines. Finally, scaling the backbone from a base model to a larger architecture yields substantial further accuracy gains, while controlled ablation experiments confirm that PATH surpasses general-purpose foundation models like CLIP and MAE by 1.8% to 6.2% on human-centric benchmarks.

These results demonstrate that a single, unified human-centric backbone can replace fragmented, task-specific computer vision models without sacrificing accuracy. For decision-makers, this consolidation translates directly into reduced deployment footprints, lower annotation and training overhead, and faster development cycles for new human-centric applications. Notably, the success of partial and head-only finetuning proves that organizations can deploy high-performing solutions on small target datasets without expensive full-model retuning.

Engineering and research teams should consider adopting unified pretraining backbones for human-centric vision pipelines and using HumanBench as a standard evaluation protocol. For edge deployments or applications with scarce training data, partial finetuning of the final layers is recommended over full retraining to balance development speed, compute costs, and peak accuracy. Further work should focus on testing unified multi-task architectures during inference, exploring unified single-head output decoders, and assessing performance on streaming video.

Confidence in these findings is high given the broad evaluation across 19 benchmark datasets and multiple task domains. However, stakeholders should note that the pretraining phase requires substantial compute to process the 11 million images, and certain highly specialized, compute-intensive re-identification methods may still achieve marginal advantages in narrow, single-domain surveillance deployments.

Cover for HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining

Abstract

Human-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the meta-verse. It is desirable to have a general pretrain model for versatile human-centric downstream tasks. This paper forges ahead along this path from the aspects of both benchmark and pretraining methods. Specifically, we propose a HumanBench based on existing datasets to comprehensively evaluate on the common ground the generalization abilities of different pretraining methods on 19 datasets from 6 diverse downstream tasks, including person ReID, pose estimation, human parsing, pedestrian attribute recognition, pedestrian detection, and crowd counting. To learn both coarse-grained and fine-grained knowledge in human bodies, we further propose a Projector AssisTed Hierarchical pretraining method (PATH) to learn diverse knowledge at different granularity levels. Comprehensive evaluations on HumanBench show that our PATH achieves new state-of-the-art results on 17 downstream datasets and on-par results on the other 2 datasets. The code will be publicly at https://github.com/OpenGVLab/HumanBench.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. HumanBench
  • 3.1. Pretraining Datasets
  • 3.2. Evaluation Scenarios and Protocols
  • 4. Methodology
  • 4.1. Hierarchical Weight Sharing
  • 4.2. Design of Task-specific Projector
  • 4.3. Dataset-specific Head and Objective Functions
  • 4.4. Technical Details
  • 5. Experiment
  • 5.1. Experimental Setup
  • 5.2. Experimental Results
  • 5.3. Ablation Study
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Projector AssisTed Hierarchical Pretraining (PATH) Framework

    model/method

    Projector AssisTed Hierarchical Pretraining (PATH) is a multi-task, multi-dataset supervised pretraining architecture designed to learn general human-centric visual representations from heterogeneous datasets and annotations without suffering from task conflicts.

    PATH organizes model parameters into a three-level hierarchical weight-sharing structure:

    1. Shared Backbone F\mathcal{F}: A standard Vision Transformer (e.g., ViT-Base or ViT-Large) shared across all datasets and all tasks to extract global, general-purpose human-centric features.
    2. Task-Specific Projectors Pt\mathcal{P}^t (t∈{1,…,T}t \in \{1, \dots, T\}): A set of TT lightweight projector modules, where each projector Pt\mathcal{P}^t is shared exclusively among all datasets belonging to the tt-th human-centric task (e.g., pose estimation, human parsing, person ReID). The projector selectively attends to task-relevant features from the backbone feature maps.
    3. Dataset-Specific Heads Hjt\mathcal{H}_j^t (j∈{1,…,Nt}j \in \{1, \dots, N_t\}): A total of N=∑t=1TNtN = \sum_{t=1}^T N_t prediction heads, where each head Hjt\mathcal{H}_j^t is dedicated to a single dataset jj within task tt without parameter sharing, absorbing dataset-specific label spaces, class mappings, and domain shifts.

    Given an image xx from dataset Djt\mathcal{D}_j^t, intermediate and final features F\mathbf{F} from the backbone are routed through the task projector p=Pt(F)p = \mathcal{P}^t(\mathbf{F}) to produce activation Zi=Hjt(p)Z_{i} = \mathcal{H}_j^t(p). The overall training objective across NN datasets is:

    L=∑i=1NλiLi(Zi,Yi)\mathcal{L} = \sum_{i=1}^N \lambda_i \mathcal{L}_i(Z_i, Y_i)

    where Li\mathcal{L}_i is the dataset-specific loss with weighting hyperparameter λi\lambda_i, and YiY_i is the corresponding label.

    During downstream evaluation or adaptation, the task-specific projectors Pt\mathcal{P}^t and dataset-specific heads Hjt\mathcal{H}_j^t are discarded; only the backbone F\mathcal{F} is transferred as the general feature extractor.

  2. Knowl 2 — Task-Specific Projector Design and Layer-Gating Aggregation

    model/method

    In the Projector AssisTed Hierarchical Pretraining (PATH) framework, each task-specific projector Pt\mathcal{P}^t extracts task-relevant spatial and channel features from intermediate layers of the shared Vision Transformer backbone and aggregates them across network depth.

    Given intermediate feature maps flf_l extracted at the ll-th transformer block (l∈{1,…,L}l \in \{1, \dots, L\} for an LL-layer backbone) for an image from the tt-th task, the projector first applies channel attention via a Squeeze-and-Excitation block Et\mathcal{E}^t followed by spatial attention via a standard self-attention block At\mathcal{A}^t:

    zl=At(Et(fl))z_l = \mathcal{A}^t(\mathcal{E}^t(f_l))

    To aggregate attended representations from all LL backbone layers, a recursive gating mechanism combines layer outputs sequentially:

    p1=z1p_1 = z_1

    pl=μlzl+(1−μl)pl−1,for l=2,…,Lp_l = \mu_l z_l + (1 - \mu_l) p_{l-1}, \quad \text{for } l = 2, \dots, L

    where the layer gate μl∈(0,1)\mu_l \in (0, 1) is parameterized with a learnable scalar αl\alpha_l (initialized to zero) and a fixed temperature Ttemp=0.1T_{\text{temp}} = 0.1:

    μl=σ(αlTtemp)\mu_l = \sigma\left(\frac{\alpha_l}{T_{\text{temp}}}\right)

    with σ(⋅)\sigma(\cdot) being the standard sigmoid function. The final aggregated feature map p=pLp = p_L is then provided to the dataset-specific head.

  3. Knowl 3 — HumanBench Pretraining Dataset Suite and Curation Protocols

    definition

    HumanBench provides a standardized pretraining dataset collection comprising 37 publicly available human-centric visual datasets with a total of 11,019,187 RGB images spanning 5 core tasks:

    • Person Re-Identification (ReID): 7 datasets, 5,446,419 images.
    • Human Pose Estimation: 11 datasets, 3,739,291 images.
    • Human Parsing: 7 datasets, 1,419,910 images.
    • Pedestrian Attribute Recognition: 6 datasets, 242,880 images.
    • Pedestrian Detection: 6 datasets, 170,687 images.

    The pretraining collection employs specific curation procedures:

    1. Noisy Label Filtering: In the weakly labeled ReID dataset LUPerson-NL, only person identities containing between 15 and 200 images are retained (yielding 151,595 identities and 5,178,420 images), removing long-tail noisy instances.
    2. De-duplication via Difference Hash (dHash): To prevent test set contamination and data leakage, dHash is computed for all pretraining images against downstream evaluation datasets; any pretraining image sharing a hash code with an evaluation image is discarded.
    3. Temporal Redundancy Subsampling: In video-derived datasets (such as AIST++ and PennAction), only 1 frame every 8 consecutive frames is sampled.
  4. Knowl 4 — HumanBench Evaluation Benchmark: Tasks, Datasets, and Protocols

    experimental setup

    HumanBench evaluates the generalizability of human-centric pretrained models across 19 benchmark datasets structured into 3 distinct evaluation scenarios and 3 evaluation protocols:

    Evaluation Scenarios:

    1. In-Dataset Evaluation (12 datasets): Tests models on validation/test subsets of datasets seen during pretraining (similar data distribution and task):
      • ReID: Market1501, MSMT, CUHK03
      • Pose Estimation: COCO, Human3.6M, AIC
      • Human Parsing: Human3.6M, LIP, CIHP
      • Pedestrian Attribute Recognition: PA-100K, RAPv2
      • Pedestrian Detection: CrowdHuman
    2. Out-of-Dataset Evaluation (5 datasets): Tests models on unseen datasets belonging to tasks seen during pretraining (assessing cross-domain generalization):
      • ReID: SenseReID
      • Pose: MPII
      • Parsing: ATR
      • Attribute: PETA
      • Detection: Caltech
    3. Unseen-Task Evaluation (2 datasets): Tests models on a task completely excluded from pretraining:
      • Crowd Counting: ShanghaiTech (ShTech) PartA, ShanghaiTech (ShTech) PartB

    Evaluation Protocols:

    • Full Finetuning (FT): All backbone layers and downstream task heads are updated jointly on 100% downstream training data.
    • Head Finetuning (Head FT): The pretrained backbone is frozen, and only a lightweight task head is trained (linear probing).
    • Partial Finetuning (Partial FT): The backbone is frozen except for the final two transformer blocks, which are finetuned alongside the downstream task head.
  5. Knowl 5 — Implementation Adaptations: Positional Embedding Interpolation and LayerNorm Decoders

    model/method

    To train a shared Vision Transformer backbone across multi-task, multi-dataset human-centric images with variable resolutions and domain gaps, two adaptations are applied:

    1. Shared Positional Embedding via 2D Interpolation: Human-centric datasets feature diverse input resolutions (from tight bounding box crops in ReID to large scene images in detection). Positional embeddings are parameterized on a canonical 224×224224 \times 224 grid across all tasks and dynamically bicubically interpolated during pretraining to match the exact patch grid of each dataset's input size.
    2. Layer Normalization in Pose and Parsing Decoders: Pose estimation and human parsing decoders traditionally employ Batch Normalization (BatchNorm). In distributed multi-task pretraining where batches originate from distinct datasets with severe domain differences, BatchNorm running statistics become unstable. Replacing BatchNorm with Layer Normalization (LayerNorm) within decoder heads stabilizes feature representation learning.
  6. Knowl 6 — Generalization Performance of PATH across HumanBench Tasks

    data/table

    The Projector AssisTed Hierarchical pretraining method (PATH) was evaluated across 19 human-centric datasets in HumanBench under Full Finetuning (FT), Head Finetuning (Head FT), and Partial Finetuning (Partial FT), using ViT-Base (ViT-B) and ViT-Large (ViT-L) backbones against domain state-of-the-art (SoTA), Masked Autoencoders (MAE), and CLIP.

    Human Parsing (mIoU / pACC) Person ReID (mAP / Top-1) Pedestrian Detection
    Method H3.6M LIP CIHP ATR Market MSMT CUHK03 SenseReID CrowdHuman Caltech (MR−2↓\text{MR}^{-2} \downarrow)
    SoTA 62.5 60.3 65.6 97.4 86.8 61.0 76.4 34.6 92.1 46.6
    SoTA †\dagger - - - - 93.0 71.8 77.7 - 92.5 28.8
    ViT-Base Backbone
    MAE 62.0 57.2 62.9 97.4 79.2 51.5 65.8 43.8 89.6 48.1
    CLIP 58.2 53.4 61.7 97.0 78.6 53.6 66.9 42.5 82.1 -
    PATH (w/o FT) 63.9 56.3 63.9 - 88.6 66.3 77.2 - 89.1 -
    PATH (FT) 65.0 61.4 66.8 97.5 89.5 69.1 82.6 47.7 90.6 30.1
    PATH (Head FT) 64.1 59.9 63.3 97.1 - - - - 90.0 31.1
    PATH (Partial FT) 63.7 60.0 63.1 97.2 88.7 66.1 79.5 48.2 90.9 28.3
    ViT-Large Backbone
    PATH (w/o FT) 65.0 62.9 67.1 - 91.6 72.7 83.7 - 89.4 -
    PATH (Partial FT) 66.2 62.6 67.5 97.4 91.8 74.7 86.0 60.0 90.8 28.7
    Pose Estimation (AP / MR−2↓\text{MR}^{-2} \downarrow / PCKh) Pedestrian Attribute (mA) Counting (MSE ↓\downarrow)
    Method COCO H3.6M AIC MPII PA-100K RAPv2 PETA ShTech A ShTech B
    SoTA 75.8 7.4 - 92.3 83.5 81.0 87.1 94.3 11.0
    SoTA †\dagger 77.1 - 32.0 93.3 - - - - -
    ViT-Base Backbone
    MAE 75.8 8.2 31.8 90.1 82.3 80.8 84.6 102.1 15.5
    CLIP 74.4 9.9 31.1 88.1 76.1 77.0 81.2 117.9 16.3
    PATH (w/o FT) 75.0 6.9 31.1 - - - - - -
    PATH (FT) 76.3 6.2 35.0 93.3 85.0 81.2 88.0 91.7 10.8
    PATH (Head FT) 75.2 6.1 31.6 92.7 77.4 72.4 79.0 - -
    PATH (Partial FT) 76.0 6.1 33.3 93.0 86.9 83.1 89.8 - 14.0
    ViT-Large Backbone
    PATH (w/o FT) 74.7 7.1 25.6 - - - - - -
    PATH (Partial FT) 77.1 5.8 36.3 93.7 90.8 87.4 90.7 - -

    Metric specifics: ATR uses pixel accuracy (pACC); Caltech and Human3.6M pose use log-average miss rate (extMR−2 ext{MR}^{-2}); crowd counting uses Mean Squared Error (MSE); other parsing tasks report mean Intersection over Union (mIoU); ReID uses mAP (SenseReID uses Top-1); detection uses AP; pose uses AP/PCKh; attributes use mean accuracy (mA). †\dagger denotes prior SoTA methods using specialized techniques such as part-based models or multi-task tuning.

    Key takeaways: With ViT-Base, PATH outperforms prior SoTA on 15 of 19 datasets under full finetuning, and partial finetuning matches or exceeds full finetuning on smaller datasets (e.g., Caltech, PETA, PA-100K). ViT-Large scales performance further, reaching new SoTA results on 17 datasets.

  7. Knowl 7 — Ablation of Projector Sharing Strategies and Positional Embedding

    data/table

    Ablation experiments conducted on a 1.28M-image subset of HumanBench (matching ImageNet-1K scale) analyze the impact of different projector weight-sharing strategies and shared positional embeddings across 4 in-dataset evaluations (PA-100K, LIP, Market1501, MSMT) and 3 out-of-dataset evaluations (Caltech, PETA, MPII).

    Design (a) (b) (c) (d)
    Shared Pos. Embedding ✓
    Projector Share Type All-Shared (A) Dataset-Specific (S) Task-Shared (T) Task-Shared (T)
    Detection: Caltech (1−MR−21 - \text{MR}^{-2}) 60.6 57.5 60.2 60.9
    Attribute: PA-100K (mA) 84.4 84.6 84.6 84.4
    Attribute: PETA (mA) 87.6 87.2 87.9 87.5
    Pose: MPII ([email protected]) 92.0 92.4 92.5 92.4
    Parsing: LIP (mIoU) 59.7 59.7 60.5 61.0
    ReID: Market1501 (mAP) 86.2 86.5 87.1 87.6
    ReID: MSMT (mAP) 65.8 66.0 66.1 66.8
    On average 76.6 76.3 76.9 77.2

    Key observations:

    1. Projector Granularity: Sharing the projector at the task level (configuration (c), average 76.9%) outperforms both sharing a single projector across all tasks (configuration (a), average 76.6%) and using separate unshared projectors per dataset (configuration (b), average 76.3%). Task-shared projectors map general backbone features to unified task subspaces while insulating the backbone from cross-task interference.
    2. Shared Positional Embeddings: Sharing interpolated positional embeddings across tasks (configuration (d)) provides an additional +0.3% average gain over unshared task-specific positional embeddings (configuration (c)), ensuring consistent spatial coordinate representations across task domains.
  8. Knowl 8 — Comparison of Pretraining Paradigms on HumanBench Subset

    data/table

    To evaluate the domain effectiveness of human-centric data versus general natural images, and supervised multi-task pretraining versus self-supervised pretraining, models were trained on a 1.28M-image subset of HumanBench and evaluated on downstream tasks.

    Pretraining Data ImageNet-1K (1.28M) HumanBench Subset (1.28M)
    Pretraining Method MAE (800 ep) MAE (800 ep) MoCoV3 (800 ep) PATH (Ours)
    Detection: Caltech (1−MR−21 - \text{MR}^{-2}) 51.9 58.2 57.5 60.9
    Attribute: PA-100K (mA) 82.3 83.5 82.9 84.4
    Attribute: PETA (mA) 84.6 85.3 84.3 87.5
    Pose: MPII ([email protected]) 90.1 91.3 90.4 92.4
    Parsing: LIP (mIoU) 57.2 60.1 58.6 61.0
    ReID: Market1501 (mAP) 79.2 84.6 86.8 87.6
    ReID: MSMT (mAP) 51.5 64.5 67.2 66.8
    On average 71.0 75.4 75.4 77.2

    Key observations:

    1. Data Domain Domain Advantage: Pretraining MAE on the HumanBench subset achieves an average performance of 75.4%, outperforming MAE pretrained on ImageNet-1K (71.0%) by +4.4%, demonstrating that domain-specific human-centric imagery is superior to generic natural images for human perception tasks.
    2. Supervised Multi-Task Superiority: Pretrained on the exact same HumanBench subset, PATH achieves 77.2% average accuracy, outperforming self-supervised baselines (MAE and MoCoV3, both 75.4%) by +1.8%, showing that supervised multi-granularity annotations provide complementary task knowledge.

Coverage note — Specific architectural implementations of standard downstream task heads (VitPose, TransReID, Segformer, Anchor DETR, Label2Label, DR.VIC) and the complete listing of all 37 individual source datasets are omitted as they adopt existing methods detailed in supplementary material.

References

  1. 1.Mykhaylo Andriluka, Umar Iqbal, Eldar Insafutdinov, Leonid Pishchulin, Anton Milan, Juergen Gall, and Bernt Schiele. Posetrack: A benchmark for human pose estimation and tracking. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5167–5176, 2018. 15
  2. 2.Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele. 2d human pose estimation: New benchmark and state of the art analysis. In Proceedings of the IEEE Conference on computer Vision and Pattern Recognition, pages 3686–3693, 2014. 4
  3. 3.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. 6
  4. 4.Markus Braun, Sebastian Krebs, Fabian Flohr, and Dariu M Gavrila. The eurocity persons dataset: A novel benchmark for object detection. arXiv preprint arXiv:1805.07193, 2018. 15
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 1
  6. 6.Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987, 2019. 14
  7. 7.Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich. Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In International conference on machine learning, pages 794–803. PMLR, 2018. 3
  8. 8.Xuangeng Chu, Anlin Zheng, Xiangyu Zhang, and Jian Sun. Detection in crowded scenes: One proposal, multiple predictions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12214–12223, 2020. 1
  9. 9.Charles Darwin and Phillip Prodger. The expression of the emotions in man and animals. Oxford University Press, USA, 1998. 3
  10. 10.Yubin Deng, Ping Luo, Chen Change Loy, and Xiaoou Tang. Pedestrian attribute recognition at far distance. In Proceedings of the 22nd ACM international conference on Multimedia, pages 789–792, 2014. 4
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 1
  12. 12.Piotr Dollar, Christian Wojek, Bernt Schiele, and Pietro Perona. Pedestrian detection: An evaluation of the state of the art. IEEE transactions on pattern analysis and machine intelligence, 34(4):743–761, 2011. 4
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 1, 4
  14. 14.Hao-Shu Fang, Jiefeng Li, Hongyang Tang, Chao Xu, Haoyi Zhu, Yuliang Xiu, Yong-Lu Li, and Cewu Lu. Alphapose: Whole-body regional multi-person pose estimation and tracking in real-time. arXiv preprint arXiv:2211.03375, 2022. 15
  15. 15.Dengpan Fu, Dongdong Chen, Jianmin Bao, Hao Yang, Lu Yuan, Lei Zhang, Houqiang Li, and Dong Chen. Unsupervised pre-training for person re-identification. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14750–14759, 2021. 15
  16. 16.Yixiao Ge, Dapeng Chen, and Hongsheng Li. Mutual mean-teaching: Pseudo label refinery for unsupervised domain adaptation on person re-identification. arXiv preprint arXiv:2001.01526, 2020. 1
  17. 17.Yixiao Ge, Feng Zhu, Dapeng Chen, Rui Zhao, et al. Self-paced contrastive learning with hybrid memory for domain adaptive object re-id. Advances in Neural Information Processing Systems, 33:11309–11321, 2020. 1
  18. 18.Ke Gong, Xiaodan Liang, Yicheng Li, Yimin Chen, Ming Yang, and Liang Lin. Instance-level human parsing via part grouping network. In Proceedings of the European conference on computer vision (ECCV), pages 770–785, 2018. 4, 15, 21
  19. 19.Ke Gong, Xiaodan Liang, Dongyu Zhang, Xiaohui Shen, and Liang Lin. Look into person: Self-supervised structure-sensitive learning and a new benchmark for human parsing. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 932–940, 2017. 4, 15
  20. 20.Michelle Guo, Albert Haque, De-An Huang, Serena Yeung, and Li Fei-Fei. Dynamic task prioritization for multitask learning. In Proceedings of the European conference on computer vision (ECCV), pages 270–287, 2018. 3
  21. 21.Tao Han, Lei Bai, Junyu Gao, Qi Wang, and Wanli Ouyang. Dr. vic: Decomposition and reasoning for video individual counting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3083–3092, 2022. 6
  22. 22.Irtiza Hasan, Shengcai Liao, Jinpeng Li, Saad Ullah Akram, and Ling Shao. Generalizable pedestrian detection: The elephant in the room. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11328–11337, 2021. 7
  23. 23.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022. 4, 6
  24. 24.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729–9738, 2020. 4
  25. 25.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 3
  26. 26.Shuting He, Hao Luo, Pichao Wang, Fan Wang, Hao Li, and Wei Jiang. Transreid: Transformer-based object re-identification. In Proceedings of the IEEE/CVF international conference on computer vision, pages 15013–15022, 2021. 6, 7
  27. 27.Yinan He, Gengshi Huang, Siyu Chen, Jianing Teng, Kun Wang, Zhenfei Yin, Lu Sheng, Ziwei Liu, Yu Qiao, and Jing Shao. X-learner: Learning cross sources and tasks for universal visual representation. In European Conference on Computer Vision, pages 509–528. Springer, 2022. 3
  28. 28.Alexander Hermans, Lucas Beyer, and Bastian Leibe. In defense of the triplet loss for person re-identification. arXiv preprint arXiv:1703.07737, 2017. 19
  29. 29.Fangzhou Hong, Liang Pan, Zhongang Cai, and Ziwei Liu. Versatile multi-modal pre-training for human-centric perception. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16156–16166, 2022. 1, 3
  30. 30.Fangzhou Hong, Liang Pan, Zhongang Cai, and Ziwei Liu. Versatile multi-modal pre-training for human-centric perception. arXiv preprint arXiv:2203.13815, 2022. 7
  31. 31.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7132–7141, 2018. 5
  32. 32.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pages 448–456. PMLR, 2015. 6, 19
  33. 33.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7):1325–1339, jul 2014. 4, 15
  34. 34.Jian Jia, Naiyu Gao, Fei He, Xiaotang Chen, and Kaiqi Huang. Learning disentangled attribute representations for robust pedestrian attribute recognition. 2022. 7
  35. 35.Xin Jin, Cuiling Lan, Wenjun Zeng, Guoqiang Wei, and Zhibo Chen. Semantics-aligned representation learning for person re-identification. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11173–11180, 2020. 7
  36. 36.Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7482–7491, 2018. 3
  37. 37.Iasonas Kokkinos. Ubernet: Training a universal convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6129–6138, 2017. 3
  38. 38.Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668, 2020. 1
  39. 39.Dangwei Li, Zhang Zhang, Xiaotang Chen, and Kaiqi Huang. A richly annotated pedestrian dataset for person retrieval in real surveillance scenarios. IEEE transactions on image processing, 28(4):1575–1590, 2018. 4, 15
  40. 40.Jianshu Li, Jian Zhao, Yunchao Wei, Congyan Lang, Yidong Li, Terence Sim, Shuicheng Yan, and Jiashi Feng. Multiple-human parsing in the wild. arXiv preprint arXiv:1705.07206, 2017. 15
  41. 41.Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa. Ai choreographer: Music conditioned 3d dance generation with aist++. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13401–13412, 2021. 3, 15
  42. 42.Suichan Li, Dapeng Chen, Bin Liu, Nenghai Yu, and Rui Zhao. Memory-based neighbourhood embedding for visual recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6102–6111, 2019. 7
  43. 43.Tianjiao Li, Jun Liu, Wei Zhang, Yun Ni, Wenqian Wang, and Zhiheng Li. Uav-human: A large benchmark for human behavior understanding with unmanned aerial vehicles. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16266–16275, 2021. 15
  44. 44.Wanhua Li, Zhexuan Cao, Jianjiang Feng, Jie Zhou, and Jiwen Lu. Label2label: A language modeling framework for multi-attribute learning. In European Conference on Computer Vision, pages 562–579. Springer, 2022. 6, 20, 21
  45. 45.Wei Li, Rui Zhao, Tong Xiao, and Xiaogang Wang. Deepreid: Deep filter pairing neural network for person re-identification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 152–159, 2014. 4, 15
  46. 46.Yining Li, Chen Huang, Chen Change Loy, and Xiaoou Tang. Human attribute recognition by deep hierarchical contexts. In European conference on computer vision, pages 684–700. Springer, 2016. 15
  47. 47.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Mvitv2: Improved multiscale vision transformers for classification and detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4804–4814, 2022. 1
  48. 48.Yanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang, Wankou Yang, Shu-Tao Xia, and Erjin Zhou. Tokenpose: Learning keypoint tokens for human pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11313–11322, 2021. 1
  49. 49.Xiaodan Liang, Ke Gong, Xiaohui Shen, and Liang Lin. Look into person: Joint body parsing & pose estimation network and a new benchmark. IEEE transactions on pattern analysis and machine intelligence, 41(4):871–885, 2018. 1, 19
  50. 50.Xiaodan Liang, Chunyan Xu, Xiaohui Shen, Jianchao Yang, Si Liu, Jinhui Tang, Liang Lin, and Shuicheng Yan. Human parsing with contextualized convolutional neural network. In Proceedings of the IEEE international conference on computer vision, pages 1386–1394, 2015. 4
  51. 51.Matthieu Lin, Chuming Li, Xingyuan Bu, Ming Sun, Chen Lin, Junjie Yan, Wanli Ouyang, and Zhidong Deng. Detr for crowd pedestrian detection. arXiv preprint arXiv:2012.06785, 2020. 1
  52. 52.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 4, 14, 15
  53. 53.Bo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone, and Qiang Liu. Conflict-averse gradient descent for multi-task learning. Advances in Neural Information Processing Systems, 34:18878–18890, 2021. 2
  54. 54.Kunliang Liu, Ouk Choi, Jianming Wang, and Wonjun Hwang. Cdgnet: Class distribution guided network for human parsing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4473–4482, 2022. 7
  55. 55.Shikun Liu, Edward Johns, and Andrew J Davison. End-to-end multi-task learning with attention. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1871–1880, 2019. 3
  56. 56.Xihui Liu, Haiyu Zhao, Maoqing Tian, Lu Sheng, Jing Shao, Shuai Yi, Junjie Yan, and Xiaogang Wang. Hydraplus-net: Attentive deep features for pedestrian analysis. In Proceedings of the IEEE international conference on computer vision, pages 350–359, 2017. 4, 15
  57. 57.Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1096–1104, 2016. 14, 15
  58. 58.Chen Change Loy, Dahua Lin, Wanli Ouyang, Yuanjun Xiong, Shuo Yang, Qingqiu Huang, Dongzhan Zhou, Wei Xia, Quanquan Li, Ping Luo, et al. Wider face and pedestrian challenge 2018: Methods and results. arXiv preprint arXiv:1902.06854, 2019. 15
  59. 59.Yongxi Lu, Abhishek Kumar, Shuangfei Zhai, Yu Cheng, Tara Javidi, and Rogerio Feris. Fully-adaptive feature sharing in multi-task networks with applications in person attribute classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5334–5343, 2017. 3
  60. 60.Hao Luo, Youzhi Gu, Xingyu Liao, Shenqi Lai, and Wei Jiang. Bag of tricks and a strong baseline for deep person re-identification. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 0–0, 2019. 1, 19, 21
  61. 61.Liqian Ma, Xu Jia, Qianru Sun, Bernt Schiele, Tinne Tuytelaars, and Luc Van Gool. Pose guided person image generation. Advances in neural information processing systems, 30, 2017. 1
  62. 62.Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3d human pose estimation in the wild using improved cnn supervision. In 2017 international conference on 3D vision (3DV), pages 506–516. IEEE, 2017. 15
  63. 63.Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, and Martial Hebert. Cross-stitch networks for multi-task learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3994–4003, 2016. 3
  64. 64.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021. 6
  65. 65.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018. 1
  66. 66.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. 1
  67. 67.Mert Bulent Sariyildiz, Yannis Kalantidis, Diane Larlus, and Karteek Alahari. Concept generalization in visual representation learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9629–9639, 2021. 2
  68. 68.Ozan Sener and Vladlen Koltun. Multi-task learning as multi-objective optimization. Advances in neural information processing systems, 31, 2018. 3
  69. 69.Jing Shao, Siyu Chen, Yangguang Li, Kun Wang, Zhenfei Yin, Yinan He, Jianing Teng, Qinghong Sun, Mengya Gao, Jihao Liu, et al. Intern: A new learning paradigm towards general vision. arXiv preprint arXiv:2111.08687, 2021. 3
  70. 70.Shuai Shao, Zijian Zhao, Boxun Li, Tete Xiao, Gang Yu, Xiangyu Zhang, and Jian Sun. Crowdhuman: A benchmark for detecting human in a crowd. arXiv preprint arXiv:1805.00123, 2018. 4, 15
  71. 71.Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR, 2018. 6
  72. 72.Weibo Shu, Jia Wan, Kay Chen Tan, Sam Kwong, and Antoni B Chan. Crowd counting in the frequency domain. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19618–19627, 2022. 7
  73. 73.Xiujun Shu, Xiao Wang, Xianghao Zang, Shiliang Zhang, Yuanqi Chen, Ge Li, and Qi Tian. Large-scale spatio-temporal person re-identification: Algorithms and benchmark. IEEE Transactions on Circuits and Systems for Video Technology, 2021. 15
  74. 74.Patrick Sudowe, Hannah Spitzer, and Bastian Leibe. Person Attribute Recognition with a Jointly-trained Holistic CNN Model. In ICCV’15 ChaLearn Looking at People Workshop, 2015. 15
  75. 75.Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. Deep high-resolution representation learning for human pose estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5693–5703, 2019. 7
  76. 76.Yifan Sun, Liang Zheng, Yi Yang, Qi Tian, and Shengjin Wang. Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline). In Proceedings of the European conference on computer vision (ECCV), pages 480–496, 2018. 6
  77. 77.Deepa Suryawanshi et al. Image Recognition: Detection of nearly duplicate images. PhD thesis, California State University Channel Islands, 2018. 3
  78. 78.Marvin Teichmann, Michael Weber, Marius Zoellner, Roberto Cipolla, and Raquel Urtasun. Multinet: Real-time joint semantic reasoning for autonomous driving. In 2018 IEEE intelligent vehicles symposium (IV), pages 1013–1020. IEEE, 2018. 3
  79. 79.Hanyue Tu, Chunyu Wang, and Wenjun Zeng. Voxelpose: Towards multi-camera 3d human pose estimation in wild environment. In European Conference on Computer Vision, pages 197–212. Springer, 2020. 1
  80. 80.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 1, 5, 19
  81. 81.Edward Vendrow, Duy Tho Le, and Hamid Rezatofighi. Jrdb-pose: A large-scale dataset for multi-person pose estimation and tracking. arXiv preprint arXiv:2210.11940, 2022. 15
  82. 82.Timo Von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. In Proceedings of the European Conference on Computer Vision (ECCV), pages 601–617, 2018. 15
  83. 83.Guanshuo Wang, Yufeng Yuan, Xiong Chen, Jiwei Li, and Xi Zhou. Learning discriminative features with multiple granularities for person re-identification. In Proceedings of the 26th ACM international conference on Multimedia, pages 274–282, 2018. 6
  84. 84.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning, pages 23318–23340. PMLR, 2022. 3
  85. 85.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022. 1, 3
  86. 86.Yizhou Wang, Shixiang Tang, Feng Zhu, Lei Bai, Rui Zhao, Donglian Qi, and Wanli Ouyang. Revisiting the transferability of supervised pretraining: an mlp perspective. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9183–9193, 2022. 2, 3, 4
  87. 87.Yingming Wang, Xiangyu Zhang, Tong Yang, and Jian Sun. Anchor detr: Query design for transformer-based detector. In Proceedings of the AAAI conference on artificial intelligence, volume 36, pages 2567–2575, 2022. 1, 6, 20
  88. 88.Longhui Wei, Shiliang Zhang, Wen Gao, and Qi Tian. Person transfer gan to bridge domain gap for person re-identification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 79–88, 2018. 4, 15
  89. 89.Jiahong Wu, He Zheng, Bo Zhao, Yixin Li, Baoming Yan, Rui Liang, Wenjia Wang, Shipei Zhou, Guosen Lin, Yanwei Fu, et al. Large-scale datasets for going deeper in image understanding. In 2019 IEEE International Conference on Multimedia and Expo (ICME), pages 1480–1485. IEEE, 2019. 4, 15
  90. 90.Bin Xiao, Haiping Wu, and Yichen Wei. Simple baselines for human pose estimation and tracking. In Proceedings of the European conference on computer vision (ECCV), pages 466–481, 2018. 1
  91. 91.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. Advances in Neural Information Processing Systems, 34:12077–12090, 2021. 6
  92. 92.Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. Vitpose: Simple vision transformer baselines for human pose estimation. arXiv preprint arXiv:2204.12484, 2022. 1, 6, 7, 19, 21
  93. 93.Kota Yamaguchi, M Hadi Kiapour, and Tamara L Berg. Paper doll parsing: Retrieving similar styles to parse clothing items. In Proceedings of the IEEE international conference on computer vision, pages 3519–3526, 2013. 15
  94. 94.Qize Yang, Ancong Wu, and Wei-Shi Zheng. Person re-identification by contour sketch under moderate clothing change. IEEE transactions on pattern analysis and machine intelligence, 43(6):2029–2046, 2019. 15
  95. 95.Sen Yang, Zhibin Quan, Mu Nie, and Wankou Yang. Transpose: Keypoint localization via transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11802–11812, 2021. 7
  96. 96.Shijie Yu, Feng Zhu, Dapeng Chen, Rui Zhao, Haobin Chen, Shixiang Tang, Jinguo Zhu, and Yu Qiao. Multiple domain experts collaborative learning: Multi-source domain generalization for person re-identification. arXiv preprint arXiv:2105.12355, 2021. 1
  97. 97.Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. Gradient surgery for multi-task learning. Advances in Neural Information Processing Systems, 33:5824–5836, 2020. 2
  98. 98.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021. 3
  99. 99.Y Yuan, F Rao, H Lang, W Lin, C Zhang, X Chen, and J Wang. Hrformer: High-resolution transformer for dense prediction. arxiv 2021. arXiv preprint arXiv:2110.09408. 19
  100. 100.Shanshan Zhang, Rodrigo Benenson, and Bernt Schiele. Citypersons: A diverse dataset for pedestrian detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3213–3221, 2017. 15
  101. 101.Shifeng Zhang, Yiliang Xie, Jun Wan, Hansheng Xia, Stan Z Li, and Guodong Guo. Widerperson: A diverse dataset for dense pedestrian detection in the wild. IEEE Transactions on Multimedia, 22(2):380–393, 2019. 15
  102. 102.Weiyu Zhang, Menglong Zhu, and Konstantinos G Derpanis. From actemes to action: A strongly-supervised representation for detailed action understanding. In Proceedings of the IEEE international conference on computer vision, pages 2248–2255, 2013. 3, 15
  103. 103.Yuanhan Zhang, Zhenfei Yin, Jing Shao, and Ziwei Liu. Benchmarking omni-vision representation through the lens of visual realms. In European Conference on Computer Vision, pages 594–611. Springer, 2022. 14
  104. 104.Yingying Zhang, Desen Zhou, Siqin Chen, Shenghua Gao, and Yi Ma. Single-image crowd counting via multi-column convolutional neural network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 589–597, 2016. 4
  105. 105.Zhilu Zhang and Mert Sabuncu. Generalized cross entropy loss for training deep neural networks with noisy labels. Advances in neural information processing systems, 31, 2018. 19
  106. 106.Haiyu Zhao, Maoqing Tian, Shuyang Sun, Jing Shao, Junjie Yan, Shuai Yi, Xiaogang Wang, and Xiaoou Tang. Spindle net: Person re-identification with human body region guided feature decomposition and fusion. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1077–1085, 2017. 4, 7
  107. 107.Nanxuan Zhao, Zhirong Wu, Rynson WH Lau, and Stephen Lin. What makes instance discrimination good for transfer learning? arXiv preprint arXiv:2006.06606, 2020. 2
  108. 108.Anlin Zheng, Yuang Zhang, Xiangyu Zhang, Xiaojuan Qi, and Jian Sun. Progressive end-to-end object detection in crowded scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 857–866, 2022. 7, 21
  109. 109.Liang Zheng, Liyue Shen, Lu Tian, Shengjin Wang, Jingdong Wang, and Qi Tian. Scalable person re-identification: A benchmark. In Proceedings of the IEEE international conference on computer vision, pages 1116–1124, 2015. 4, 15
  110. 110.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6881–6890, 2021. 19
  111. 111.Shuai Zheng, Fan Yang, M Hadi Kiapour, and Robinson Piramuthu. Modanet: A large-scale street fashion dataset with polygon annotations. In Proceedings of the 26th ACM international conference on Multimedia, pages 1670–1678, 2018. 15
  112. 112.Yi Zheng, Shixiang Tang, Guolong Teng, Yixiao Ge, Kaijian Liu, Jing Qin, Donglian Qi, and Dapeng Chen. Online pseudo label generation by hierarchical cluster dynamics for adaptive person re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8371–8381, 2021. 1
  113. 113.Zhedong Zheng, Xiaodong Yang, Zhiding Yu, Liang Zheng, Yi Yang, and Jan Kautz. Joint discriminative and generative learning for person re-identification. In proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2138–2147, 2019. 15
  114. 114.Qixian Zhou, Xiaodan Liang, Ke Gong, and Liang Lin. Adaptive temporal encoding network for video instance-level human parsing. In Proceedings of the 26th ACM international conference on Multimedia, pages 1527–1535, 2018. 15
  115. 115.Jinguo Zhu, Xizhou Zhu, Wenhai Wang, Xiaohua Wang, Hongsheng Li, Xiaogang Wang, and Jifeng Dai. Uni-perceiver-moe: Learning sparse generalist models with conditional moes. arXiv preprint arXiv:2206.04674, 2022. 1, 3
  116. 116.Kuan Zhu, Haiyun Guo, Tianyi Yan, Yousong Zhu, Jinqiao Wang, and Ming Tang. Pass: Part-aware self-supervised pre-training for person re-identification. In European Conference on Computer Vision, pages 198–214. Springer, 2022. 6, 7

Citation

MLA
Tang, S., et al. “HumanBench: Towards General Human-centric Perception with Projector Assisted Pretraining”. arXiv, 2023, http://arxiv.org/abs/2303.05675v1.
APA
Tang, S., Chen, C., Xie, Q., Chen, M., Wang, Y., Ci, Y., Bai, L., Zhu, F., Yang, H., Yi, L., Zhao, R., & Ouyang, W. (2023). HumanBench: Towards General Human-centric Perception with Projector Assisted Pretraining. arXiv. http://arxiv.org/abs/2303.05675v1
Chicago
Tang, S., C. Chen, Q. Xie, et al. 2023. “HumanBench: Towards General Human-centric Perception with Projector Assisted Pretraining”. arXiv. http://arxiv.org/abs/2303.05675v1.
Harvard
Tang, S. et al. (2023) “HumanBench: Towards General Human-centric Perception with Projector Assisted Pretraining”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.05675v1.
Vancouver
1. Tang S, Chen C, Xie Q, et al (2023) HumanBench: Towards General Human-centric Perception with Projector Assisted Pretraining. arXiv

BibTeX

@article{tang2023humanbench,
  title = {HumanBench: Towards General Human-centric Perception with Projector Assisted Pretraining},
  author = {Tang, Shixiang and Chen, Cheng and Xie, Qingsong and Chen, Meilin and Wang, Yizhou and Ci, Yuanzheng and Bai, Lei and Zhu, Feng and Yang, Haiyang and Yi, Li and Zhao, Rui and Ouyang, Wanli},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.05675v1},
  eprint = {2303.05675}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE