HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining
Shixiang TangCheng ChenQingsong XieMeilin ChenYizhou WangYuanzheng CiLei BaiFeng ZhuHaiyang YangLi Yi
Presents HumanBench, a 19-dataset benchmark spanning six human-centric vision tasks, alongside PATH, a hierarchical projector-assisted pretraining framework that achieves state-of-the-art transfer performance across diverse granularities of human perception.
Human-centric perception tasks—including person re-identification, pose estimation, human parsing, pedestrian attribute recognition, and pedestrian detection—are critical for industries such as autonomous driving, surveillance, and virtual environments. Historically, organizations have developed specialized models for each distinct task, resulting in high engineering costs, duplicated data labeling efforts, and inefficient multi-model deployment pipelines. While general vision foundation models exist, they are primarily trained on natural scenes or language associations rather than the multi-scale, fine-grained structural features inherent to human bodies.
The article evaluates whether a unified, multi-dataset supervised pretraining framework can create a single, highly generalizable human-centric visual representation. It demonstrates a scalable methodology to overcome task conflicts and annotation discrepancies across disparate human vision tasks, establishing a standardized evaluation benchmark to measure performance.
To achieve this, the article introduces HumanBench, a large-scale benchmark comprising 11 million images across 37 datasets covering five core human-centric tasks. To resolve the negative task conflicts common in multi-task learning, the article proposes a pretraining framework called Projector Assisted Hierarchical Pretraining (PATH). This method employs a hierarchical weight-sharing architecture: a single transformer backbone is shared across all tasks, lightweight attention and gating projectors are shared strictly among datasets of the same task, and independent output heads handle dataset-specific distributions. During downstream evaluation, the projectors and heads are discarded, assessing the pure backbone's transferable features through full finetuning, head-only finetuning, and partial layer finetuning.
The findings show that PATH establishes state-of-the-art results across 17 downstream datasets and matches top existing methods on 2 others. First, when adapted to 12 seen pretraining datasets, PATH outperforms prior specialized models in major categories, improving human parsing accuracy by up to 2.5% mean intersection-over-union and person re-identification by 4.9% mean average precision. Second, PATH demonstrates robust out-of-dataset generalization across 5 unseen datasets, improving recognition accuracy by 4.4% and reducing pedestrian detection miss rates. Third, when evaluated on crowd counting—a completely unseen task outside the pretraining scope—PATH reduces mean squared error by up to 10.4% compared to standard self-supervised baselines. Finally, scaling the backbone from a base model to a larger architecture yields substantial further accuracy gains, while controlled ablation experiments confirm that PATH surpasses general-purpose foundation models like CLIP and MAE by 1.8% to 6.2% on human-centric benchmarks.
These results demonstrate that a single, unified human-centric backbone can replace fragmented, task-specific computer vision models without sacrificing accuracy. For decision-makers, this consolidation translates directly into reduced deployment footprints, lower annotation and training overhead, and faster development cycles for new human-centric applications. Notably, the success of partial and head-only finetuning proves that organizations can deploy high-performing solutions on small target datasets without expensive full-model retuning.
Engineering and research teams should consider adopting unified pretraining backbones for human-centric vision pipelines and using HumanBench as a standard evaluation protocol. For edge deployments or applications with scarce training data, partial finetuning of the final layers is recommended over full retraining to balance development speed, compute costs, and peak accuracy. Further work should focus on testing unified multi-task architectures during inference, exploring unified single-head output decoders, and assessing performance on streaming video.
Confidence in these findings is high given the broad evaluation across 19 benchmark datasets and multiple task domains. However, stakeholders should note that the pretraining phase requires substantial compute to process the 11 million images, and certain highly specialized, compute-intensive re-identification methods may still achieve marginal advantages in narrow, single-domain surveillance deployments.
- Paper: Simple Multi-dataset Detection, Xingyi Zhou et al. (2022). Understanding multi-dataset training strategies with dataset-specific heads provides key foundational context for the architectural choices explored in HumanBench's unified pretraining framework.
- Paper: Unified Perceptual Parsing for Scene Understanding, Tete Xiao et al. (2018). Examining unified perceptual parsing across heterogeneous visual annotations establishes the multi-task learning paradigms that HumanBench builds upon to unify human-centric perception.
- Paper: Taskonomy: Disentangling Task Transfer Learning, Amir Zamir et al. (2018). Studying task transferability and cross-task relationships clarifies the negative transfer challenges that HumanBench specifically aims to mitigate with hierarchical projectors.
- Paper: DensePose: Dense Human Pose Estimation in the Wild, Rıza Alp Güler et al. (2018). Familiarity with dense human pose estimation benchmarks provides essential background on the fine-grained, body-centric visual tasks integrated into the HumanBench suite.
- Paper: Bag of Tricks and a Strong Baseline for Deep Person Re-Identification, Hao Luo et al. (2019). Reviewing standard baseline and optimization practices in person re-identification illuminates one of the core downstream perception tasks benchmarked in HumanBench.
- Paper: Deep Learning for Person Re-Identification: A Survey and Outlook, Mang Ye et al. (2020). Surveying the landscape and architectural requirements of person re-identification highlights the fragmented task-specific designs that HumanBench's unified backbone seeks to replace.
- Paper: Learning Multiple Dense Prediction Tasks from Partially Annotated Data, Wei-Hong Li et al. (2022). Reading how shared encoders handle partially annotated, multi-task dense prediction setups provides critical context for multi-dataset human-centric pretraining.
- Paper: Big Transfer (BiT): General Visual Representation Learning, Alexander Kolesnikov et al. (2019). Exploring scalable supervised pretraining and transfer rules establishes the foundational transfer learning methodology that HumanBench adapts to specialized human perception.
- Paper: VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene Understanding, Yi Xin et al. (2024). Learn how parameter-efficient multi-task adapter architectures build upon unified pre-trained representations to efficiently transfer to simultaneous dense visual tasks.
- Paper: Towards Large-Scale 3D Representation Learning with Multi-Dataset Point Prompt Training, Xiaoyang Wu et al. (2024). Discover how the principles of mitigating negative transfer during multi-dataset pretraining are extended to large-scale 3D representation learning via prompt training.
