Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer

Wang ZengSheng JinWentao LiuChen QianPing LuoWanli OuyangXiaogang Wang

article2022CVPR191 citations

Proposes Token Clustering Transformer (TCFormer), a vision architecture that dynamically clusters and merges image tokens into flexible shapes and sizes to allocate finer resolution to critical human regions while efficiently compressing background details for tasks like pose estimation and 3D mesh reconstruction.

Listen

Vision transformers are critical tools for human-centric visual analysis applications, such as augmented reality, action recognition, and human-computer interaction. However, standard vision transformers divide images into rigid, uniform grids, assigning equal computational focus to all image patches regardless of content. This uniform allocation is inefficient and sub-optimal for human analysis, where intricate foreground details—such as faces and hands—require high precision, while expansive background regions require minimal processing.

The article introduces and evaluates the Token Clustering Transformer (TCFormer), a novel framework designed to dynamically generate flexible vision tokens—computational image units—based on semantic content. The primary objective is to demonstrate that grouping tokens via progressive clustering allows models to preserve crucial fine details in target regions while reducing computational focus on background areas.

To achieve this, the approach employs a multi-stage architecture featuring two key components: a Clustering Token Merge block and a Multi-stage Token Aggregation head. The merge block uses a density-based nearest-neighbor clustering algorithm to group semantically similar tokens into flexible, non-grid shapes and sizes, guided by an importance score. The aggregation head progressively combines features across all stages while preserving fine-grained details. The authors evaluated the system across several benchmark datasets, including COCO-WholeBody for 2D whole-body pose estimation, 3DPW and Human3.6M for 3D human mesh reconstruction, WFLW for facial landmark alignment, and ImageNet-1K for general image classification.

The findings confirm that TCFormer consistently outperforms existing state-of-the-art models across all tested domains. On 2D whole-body pose estimation, TCFormer achieved 57.2% Average Precision, notably outperforming established baselines on difficult detail-heavy parts like hands by 6.2% Average Precision. On 3D mesh reconstruction, it reduced mean joint position error to 49.3 mm on 3DPW, achieving competitive or superior accuracy compared to specialized methods. For 2D face alignment, it delivered a lower normalized error (4.28%) than models utilizing extra boundary information. In general image classification, TCFormer attained an 82.4% Top-1 accuracy on ImageNet-1K, proving its broad utility beyond human-specific tasks.

These results demonstrate that dynamic, clustering-based token allocation substantially enhances feature representation without requiring prohibitive memory or computational overhead. By dedicating finer tokens to complex body areas and coarse tokens to backgrounds, systems can achieve higher operational accuracy in detail-sensitive vision tasks, directly benefiting real-time human tracking and interactive visual technologies.

Organizations developing vision systems should consider adopting dynamic token clustering mechanisms over rigid grid transformers for fine-grained localization tasks. Future technical roadmaps should explore applying this architecture to broader visual domains, including general object detection and semantic segmentation.

The primary limitation of the current method is that its clustering algorithm exhibits quadratic computational complexity relative to the number of input tokens, which can constrain processing speeds on very high-resolution images. While the authors suggest mitigating this by partitioning tokens into subsets prior to clustering, decision-makers should evaluate this scaling trade-off when deploying the model on high-resolution, real-time pipelines.

Cover for Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer

Abstract

Vision transformers have achieved great successes in many computer vision tasks. Most methods generate vision tokens by splitting an image into a regular and fixed grid and treating each cell as a token. However, not all regions are equally important in human-centric vision tasks, e.g., the human body needs a fine representation with many tokens, while the image background can be modeled by a few tokens. To address this problem, we propose a novel Vision Transformer, called Token Clustering Transformer (TCFormer), which merges tokens by progressive clustering, where the tokens can be merged from different locations with flexible shapes and sizes. The tokens in TCFormer can not only focus on important areas but also adjust the token shapes to fit the semantic concept and adopt a fine resolution for regions containing critical details, which is beneficial to capturing detailed information. Extensive experiments show that TCFormer consistently outperforms its counterparts on different challenging human-centric tasks and datasets, including whole-body pose estimation on COCO-WholeBody and 3D human mesh reconstruction on 3DPW. Code is available at https://github.com/zengwang430521/TCFormer.git.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Transformers in Human-Centric Vision Tasks
  • 2.2. Dynamic Token Generation
  • 2.3. Clustering for Feature Aggregation
  • 3. Method
  • 3.1. Overview Architecture
  • 3.2. Transformer Block
  • 3.3. Clustering-based Token Merge (CTM) Block
  • 3.4. Multi-stage Token Aggregation (MTA) Head
  • 3.5. Implementation Details
  • 4. Experiments
  • 4.1. 2D Whole-body Pose Estimation
  • 4.2. 3D Human Mesh Reconstruction
  • 4.3. 2D Face Keypoint Localization
  • 4.4. Image Classification
  • 4.5. Ablation Study
  • 4.6. Qualitative Evaluation
  • 5. Analysis
  • 6. Conclusions and Limitations
  • References

Knowls

  1. Knowl 1 — Token Clustering Transformer Architecture

    model/method

    The Token Clustering Transformer (TCFormer) is a hierarchical vision transformer designed for human-centric visual analysis tasks (such as 2D whole-body pose estimation, 3D human mesh recovery, and face alignment). Standard vision transformers divide an image into fixed, uniform grid tokens regardless of image content. TCFormer dynamically adjusts token location, shape, and spatial scale across 4 hierarchical stages through progressive semantic clustering, allocating dense, fine-grained tokens to semantically critical regions (e.g., human faces, hands, and body joints) and sparse, coarse tokens to background regions.

    The overall architecture consists of:

    1. Initial Tokenization: A convolutional embedding layer transforms the input RGB image into a regular feature map where each pixel is initialized as an individual token.
    2. Hierarchical Stages: 4 sequential stages containing stacked transformer blocks. In each block, spatial reduction layers reduce key/value resolution to bound computational complexity, and depth-wise convolutional layers capture local spatial context and implicit positional information without hand-crafted explicit positional embeddings.
    3. Clustering-based Token Merge (CTM) Blocks: Inserted between consecutive stages to group semantically similar tokens via density peaks clustering, reducing the token count by a factor of 4 between adjacent stages.
    4. Multi-stage Token Aggregation (MTA) Head: Aggregates multi-scale token representations from all stages top-down at the token level, preserving fine-grained details before projecting to the final output representation.
  2. Knowl 2 — Clustering-based Token Merge via Density Peaks Clustering

    algorithm

    The Clustering-based Token Merge (CTM) module reduces the token sequence between adjacent transformer stages by clustering tokens according to their feature representations using a kk-nearest-neighbor density peaks clustering algorithm (DPC-KNN).

    Input: Token feature set X={x1,x2,…,xN}X = \{x_1, x_2, \dots, x_N\} with xi∈Rdx_i \in \mathbb{R}^d, number of target clusters M=N/4M = N/4, neighborhood size kk
    Output: Merged cluster token features Y={y1,y2,…,yM}Y = \{y_1, y_2, \dots, y_M\}, cluster assignments C={C1,C2,…,CM}C = \{C_1, C_2, \dots, C_M\}
    for each token xi∈Xx_i \in X do
        Find kk-nearest neighbors KNN(xi)\mathrm{KNN}(x_i) based on Euclidean distance in feature space
        Compute local density ρi=exp⁡(−1k∑xj∈KNN(xi)∥xi−xj∥22)\rho_i = \exp\left(-\frac{1}{k} \sum_{x_j \in \mathrm{KNN}(x_i)} \|x_i - x_j\|_2^2\right)
    end for
    for each token xi∈Xx_i \in X do
        if there exists jj such that ρj>ρi\rho_j > \rho_i then
            Compute distance indicator δi=min⁡j:ρj>ρi∥xi−xj∥2\delta_i = \min_{j: \rho_j > \rho_i} \|x_i - x_j\|_2
        else
            Compute distance indicator δi=max⁡j∥xi−xj∥2\delta_i = \max_j \|x_i - x_j\|_2
        end if
        Compute cluster center priority score Si=ρi×δiS_i = \rho_i \times \delta_i
    end for
    Select the MM tokens with the highest priority score SiS_i as cluster centers {μ1,μ2,…,μM}\{\mu_1, \mu_2, \dots, \mu_M\}
    Assign each token xi∈Xx_i \in X to the cluster CmC_m of its nearest cluster center μm\mu_m in Euclidean feature space

    The spatial extent (token region) of each merged cluster token ymy_m is defined as the union of the spatial token regions belonging to all constituent tokens in cluster CmC_m.

  3. Knowl 3 — Importance-Guided Token Feature Merging and Attention Weighting

    model/method

    Within the Clustering-based Token Merge (CTM) module, tokens grouped into the same cluster CiC_i are merged into a single representative token feature yiy_i using an importance-weighted average rather than uniform averaging. An importance score pj∈Rp_j \in \mathbb{R} is estimated for each token xjx_j from its feature vector via a lightweight projection:

    yi=∑j∈Ciepjxj∑j∈Ciepjy_i = \frac{\sum_{j \in C_i} e^{p_j} x_j}{\sum_{j \in C_i} e^{p_j}}

    where CiC_i is the index set of tokens assigned to the ii-th cluster, xjx_j is the feature vector of token jj, and pjp_j is the predicted importance score of token jj.

    Following feature merging, the merged cluster tokens serve as queries Q∈RM×dkQ \in \mathbb{R}^{M \times d_k} and interact with the original pre-merge tokens serving as keys K∈RN×dkK \in \mathbb{R}^{N \times d_k} and values V∈RN×dkV \in \mathbb{R}^{N \times d_k} via a cross-attention transformer layer. The token importance score vector P∈R1×NP \in \mathbb{R}^{1 \times N} is directly incorporated into the attention matrix to bias attention toward high-importance tokens:

    Attention⁡(Q,K,V)=softmax⁡(QKTdk+P)V\operatorname{Attention}(Q, K, V) = \operatorname{softmax}\left(\frac{Q K^T}{\sqrt{d_k}} + P\right) V

    where dkd_k is the channel dimension of the queries. This allows the newly generated cluster tokens to selectively aggregate critical detailed features from the original fine tokens.

  4. Knowl 4 — Multi-stage Token Aggregation Head

    model/method

    The Multi-stage Token Aggregation (MTA) head aggregates hierarchical features across all transformer stages without converting intermediate tokens into low-resolution convolutional feature maps, preventing the information loss that occurs when averaging fine tokens over coarse pixel grids.

    MTA operates top-down starting from the coarsest stage (Stage 4) to the finest stage (Stage 1):

    1. Token Upsampling: Using the cluster assignment relationships recorded during the forward pass of each Clustering-based Token Merge (CTM) module, each merged cluster token feature is copied and broadcast back to all original tokens that formed that cluster.
    2. Feature Fusion: The upsampled token features from the deeper stage are element-wise added to the corresponding token features of the preceding shallower stage.
    3. Transformer Refinement: The fused token sequence is processed through a transformer block to harmonize features across scales.
    4. Progressive Recursion: The upsample-fuse-refine sequence repeats recursively until the tokens reach Stage 1 resolution, where each token corresponds directly to a single pixel on the initial high-resolution grid.
    5. Final Output: The Stage 1 tokens are reshaped into a standard 2D feature map and processed by a convolutional prediction head to produce keypoint heatmaps.
  5. Knowl 5 — Bidirectional Mapping Between Irregular Vision Tokens and Grid Feature Maps

    model/method

    Because convolutional operations and spatial reduction modules require regular 2D feature grid representations, TCFormer employs area-weighted bidirectional mappings between irregular semantic tokens and standard rectangular grid feature maps:

    1. Tokens to Feature Map: Each grid pixel is treated as a rectangular geometric unit. To construct the feature for a given pixel, all vision tokens whose spatial coverage areas overlap with that pixel's grid cell are identified. The pixel feature is computed as the weighted average of these token features, with weights proportional to the intersection area between the token's spatial region and the pixel's grid cell.
    2. Feature Map to Tokens: When transforming a regular 2D feature map back into vision tokens, each token's feature vector is computed by taking the spatial average of the feature map values across all grid pixels contained within that token's defined spatial region.
  6. Knowl 6 — Whole-Body Human Pose Estimation Benchmark on COCO-WholeBody

    data/table

    TCFormer was evaluated on the COCO-WholeBody V1.0 test dataset containing 133 whole-body keypoints (17 body, 6 feet, 68 face, 42 hands). Evaluation metrics are Object Keypoint Similarity (OKS)-based Average Precision (AP) and Average Recall (AR) across individual body parts and the whole body.

    Method Resolution body foot face hand whole-body
    AP AR AP AR AP AR AP AR AP AR
    SN* 480 × 480 0.427 0.583 0.099 0.369 0.649 0.697 0.408 0.580 0.327 0.456
    OpenPose 480 × 480 0.563 0.612 0.532 0.645 0.765 0.840 0.386 0.433 0.442 0.523
    PAF* 480 × 480 0.381 0.526 0.053 0.278 0.655 0.701 0.359 0.528 0.295 0.405
    AE + HRNet-w48 512 × 512 0.592 0.686 0.443 0.595 0.619 0.674 0.347 0.438 0.422 0.532
    HigherHRNet-w48 512 × 512 0.630 0.706 0.440 0.573 0.730 0.777 0.389 0.477 0.487 0.574
    ZoomNet 384 × 288 0.743 0.802 0.798 0.869 0.623 0.701 0.401 0.498 0.541 0.658
    SBL-Res50 256 × 192 0.652 0.739 0.614 0.746 0.608 0.716 0.460 0.584 0.520 0.633
    SBL-Res101 256 × 192 0.670 0.754 0.640 0.767 0.611 0.723 0.463 0.589 0.533 0.647
    SBL-Res152 256 × 192 0.682 0.764 0.662 0.788 0.624 0.728 0.482 0.606 0.548 0.661
    HRNet-w32 256 × 192 0.700 0.746 0.567 0.645 0.637 0.688 0.473 0.546 0.553 0.626
    TCFormer w/o CTM 256 × 192 0.667 0.749 0.562 0.695 0.617 0.621 0.479 0.590 0.535 0.639
    TCFormer w/o MTA Head 256 × 192 0.679 0.761 0.658 0.780 0.634 0.732 0.499 0.619 0.553 0.662
    TCFormer 256 × 192 0.691 0.770 0.698 0.813 0.649 0.746 0.535 0.650 0.572 0.678

    TCFormer achieves 0.572 whole-body AP, outperforming HRNet-w32 by 1.9% AP at the same input resolution (256×192256 \times 192). On hand keypoints—which occupy small spatial areas and depend heavily on high-resolution details—TCFormer achieves 0.535 AP, exceeding HRNet-w32 by 6.2% AP and SBL-Res152 by 5.3% AP.

  7. Knowl 7 — Ablation of Token Clustering Merge and Multi-stage Token Aggregation Head

    empirical result

    Ablation studies on COCO-WholeBody pose estimation at 256×192256 \times 192 resolution isolate the individual contributions of the Clustering-based Token Merge (CTM) module and the Multi-stage Token Aggregation (MTA) head:

    • Impact of CTM: Replacing the CTM block with a standard strided convolutional downsampling layer causes the overall whole-body performance to drop by 3.7% AP (from 57.2% to 53.5%) and 3.9% AR (from 67.8% to 63.9%). The performance drop is most pronounced on parts requiring high detail: foot AP drops by 13.6% (from 69.8% to 56.2%) and hand AP drops by 5.6% (from 53.5% to 47.9%), whereas main body AP decreases by only 2.4% (from 69.1% to 66.7%).
    • Impact of MTA Head: Replacing the token-level MTA head with a standard deconvolutional head leads to an overall drop of 1.9% AP (from 57.2% to 55.3%) and 1.6% AR (from 67.8% to 66.2%). Performance drops on foot AP by 4.0% and hand AP by 3.6%, demonstrating that token-level aggregation prevents the loss of fine details caused by intermediate rasterization into coarse feature maps.
  8. Knowl 8 — 3D Human Mesh Reconstruction on 3DPW and Human3.6M

    data/table

    TCFormer combined with a standard HMR regression head was evaluated on 3D human pose and mesh estimation using the 3DPW (in-the-wild) and Human3.6M (indoor) benchmarks. Metrics reported are Mean Per Joint Position Error (MPJPE in mm) and Procrustes-Aligned MPJPE (PA-MPJPE in mm).

    Method 3DPW Human3.6M
    PA-MPJPE ↓\downarrow MPJPE ↓\downarrow PA-MPJPE ↓\downarrow MPJPE ↓\downarrow
    HMR* 76.7 130.0 56.8 88.0
    SPIN* 59.2 96.9 41.1 62.5
    Zanfir et al. 57.1 90.0 - -
    EFT 52.2 - 43.8 -
    DSR 51.7 85.7 41.4 62.0
    TCFormer (Ours) 49.3 80.6 42.8 62.9

    On 3DPW, TCFormer achieves 49.3 mm PA-MPJPE and 80.6 mm MPJPE, outperforming prior model-based regression frameworks (including EFT at 52.2 mm and DSR with dense supervision at 51.7 mm PA-MPJPE) at an input resolution of 224×224224 \times 224.

  9. Knowl 9 — 2D Facial Landmark Localization on WFLW

    data/table

    A lightweight variant (TCFormer-Light) was evaluated on the WFLW dataset (98 facial landmarks; 7,500 train, 2,500 test images) across the full test set and 6 specific subsets: Large Pose (LP), Expression (Expr.), Illumination (Illu.), Make-up (Mu.), Occlusion (Occu.), and Blur. The evaluation metric is Normalized Mean Error (NME in %), normalized by inter-ocular distance.

    Method Backbone Test ↓\downarrow LP ↓\downarrow Expr. ↓\downarrow Illu. ↓\downarrow Mu. ↓\downarrow Occu. ↓\downarrow Blur ↓\downarrow
    ESR - 11.13 25.88 11.47 10.49 11.05 13.75 12.20
    SDM - 10.29 24.10 11.45 9.32 9.38 13.03 11.28
    CFSS - 9.07 21.36 10.09 8.30 8.74 11.76 9.96
    DVLN VGG-16 6.08 11.54 6.78 5.73 5.98 7.33 6.88
    HRNetV2 HRNetV2-W18 4.60 7.94 4.85 4.55 4.29 5.44 5.42
    TCFormer (Ours) TCFormer-Light 4.28 7.27 4.56 4.18 4.27 5.18 4.87
    Models trained with extra information
    LAB (w/ B) Hourglass 5.27 10.24 5.51 5.23 5.15 6.79 6.32
    PDB (w/ DA) ResNet-50 5.11 8.75 5.36 4.93 5.41 6.37 5.81

    TCFormer-Light achieves an overall NME of 4.28% on the WFLW test set, improving over HRNetV2-W18 (4.60%) and outperforming methods that use extra boundary supervision (LAB, 5.27%) or stronger data augmentations (PDB, 5.11%).

  10. Knowl 10 — Computational Complexity Limitation of DPC-KNN Token Clustering

    limitation

    The primary computational bottleneck in TCFormer stems from the Density Peaks Clustering algorithm based on kk-nearest neighbors (DPC-KNN) used inside the Clustering-based Token Merge (CTM) block. The pairwise distance computation in DPC-KNN scales quadratically (O(N2)O(N^2)) with respect to the total number of input tokens NN. While GPU implementation keeps clustering and merging to 9.4% of the forward pass time under standard human-centric input resolutions (256×192256 \times 192 or 224×224224 \times 224), the quadratic scaling limits inference throughput when scaling to very large input image resolutions. This can be alleviated by partitioning tokens spatially into independent sub-regions and executing clustering part-wise.

Coverage note — ImageNet-1K classification benchmark results (Table 4) were omitted as ImageNet was solely included as an auxiliary verification of general vision feature representation, whereas the paper's core contributions and evaluations center on human-centric visual analysis tasks.

References

  1. 1.Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele. 2d human pose estimation: New benchmark and state of the art analysis. In IEEE Conf. Comput. Vis. Pattern Recog., pages 3686–3693, 2014.
  2. 2.Xavier P Burgos-Artizzu, Pietro Perona, and Piotr Dollar. Robust face landmark estimation under occlusion. In Int. Conf. Comput. Vis., pages 1513–1520, 2013.
  3. 3.Xudong Cao, Yichen Wei, Fang Wen, and Jian Sun. Face alignment by explicit shape regression. Int. J. Comput. Vis., 107(2):177–190, 2014.
  4. 4.Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh. Openpose: realtime multi-person 2d pose estimation using part affinity fields. IEEE Trans. Pattern Anal. Mach. Intell., 2018.
  5. 5.Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. Realtime multi-person 2d pose estimation using part affinity fields. In IEEE Conf. Comput. Vis. Pattern Recog., 2017.
  6. 6.Bowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi, Thomas S Huang, and Lei Zhang. Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 5386–5395, 2020.
  7. 7.Hongsuk Choi, Gyeongsik Moon, and Kyoung Mu Lee. Pose2mesh: Graph convolutional network for 3d human pose and mesh recovery from a 2d human pose. In Eur. Conf. Comput. Vis., pages 769–787, 2020.
  8. 8.Xiao Chu, Wei Yang, Wanli Ouyang, Cheng Ma, Alan L Yuille, and Xiaogang Wang. Multi-context attention for human pose estimation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 1831–1840, 2017.
  9. 9.MMPose Contributors. Openmmlab pose estimation tool-box and benchmark. https://github.com/open-mmlab/mmpose, 2020.
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. Int. Conf. Learn. Represent., 2021.
  11. 11.Mingjing Du, Shifei Ding, and Hongjie Jia. Study on density peaks clustering based on k-nearest neighbors and principal component analysis. Knowledge-Based Systems, 99:135–145, 2016.
  12. 12.Haodong Duan, Kwan-Yee Lin, Sheng Jin, Wentao Liu, Chen Qian, and Wanli Ouyang. Trb: a novel triplet representation for understanding 2d human body. In Int. Conf. Comput. Vis., pages 9479–9488, 2019.
  13. 13.Sai Kumar Dwivedi, Nikos Athanasiou, Muhammed Kocabas, and Michael J Black. Learning to regress bodies from images using differentiable semantic rendering. In Int. Conf. Comput. Vis., 2021.
  14. 14.Zhen-Hua Feng, Josef Kittler, Muhammad Awais, Patrik Huber, and Xiao-Jun Wu. Wing loss for robust facial landmark localisation with convolutional neural networks. In IEEE Conf. Comput. Vis. Pattern Recog., 2018.
  15. 15.Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Handsformer: Keypoint transformer for monocular 3d pose estimation ofhands and object in interaction. arXiv preprint arXiv:2104.14639, 2021.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conf. Comput. Vis. Pattern Recog., 2016.
  17. 17.Gines Hidalgo, Yaadhav Raaj, Haroon Idrees, Donglai Xiang, Hanbyul Joo, Tomas Simon, and Yaser Sheikh. Singlenetwork whole-body pose estimation. In IEEE Conf. Comput. Vis. Pattern Recog., 2019.
  18. 18.Kun Hu, Zhiyong Wang, Wei Wang, Kaylena A Ehgoetz Martens, Liang Wang, Tieniu Tan, Simon JG Lewis, and David Dagan Feng. Graph sequence recurrent neural network for vision-based freezing of gait detection. IEEE Trans. Image Process., 29:1890–1901, 2019.
  19. 19.Lin Huang, Jianchao Tan, Jingjing Meng, Ji Liu, and Junsong Yuan. Hot-net: Non-autoregressive transformer for 3d hand-object pose estimation. In ACM Int. Conf. Multimedia, pages 3136–3145, 2020.
  20. 20.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE Trans. Pattern Anal. Mach. Intell., 36(7):1325–1339, 2013.
  21. 21.Sheng Jin, Wentao Liu, Wanli Ouyang, and Chen Qian. Multi-person articulated tracking with spatial and temporal embeddings. In IEEE Conf. Comput. Vis. Pattern Recog., pages 5664–5673, 2019.
  22. 22.Sheng Jin, Wentao Liu, Enze Xie, Wenhai Wang, Chen Qian, Wanli Ouyang, and Ping Luo. Differentiable hierarchical graph grouping for multi-person pose estimation. In Eur. Conf. Comput. Vis., pages 718–734. Springer, 2020.
  23. 23.Sheng Jin, Lumin Xu, Jin Xu, Can Wang, Wentao Liu, Chen Qian, Wanli Ouyang, and Ping Luo. Whole-body human pose estimation in the wild. In Eur. Conf. Comput. Vis., pages 196–214, 2020.
  24. 24.Sam Johnson and Mark Everingham. Clustered pose and nonlinear appearance models for human pose estimation. In Brit. Mach. Vis. Conf., 2010. doi:10.5244/C.24.12.
  25. 25.Sam Johnson and Mark Everingham. Learning effective human pose estimation from inaccurate annotation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 1465–1472, 2011.
  26. 26.Hanbyul Joo, Natalia Neverova, and Andrea Vedaldi. Exemplar fine-tuning for 3d human model fitting towards in-the-wild 3d human pose estimation. In International Conference on 3D Vision (3DV), pages 42–52, 2021.
  27. 27.Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In IEEE Conf. Comput. Vis. Pattern Recog., 2018.
  28. 28.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. Int. Conf. Learn. Represent., 2015.
  29. 29.Nikos Kolotouros, Georgios Pavlakos, Michael J Black, and Kostas Daniilidis. Learning to reconstruct 3d human pose and shape via model-fitting in the loop. Int. Conf. Comput. Vis., 2019.
  30. 30.Nikos Kolotouros, Georgios Pavlakos, and Kostas Daniilidis. Convolutional mesh regression for single-image human shape reconstruction. In IEEE Conf. Comput. Vis. Pattern Recog., pages 4501–4510, 2019.
  31. 31.Jiefeng Li, Siyuan Bian, Ailing Zeng, Can Wang, Bo Pang, Wentao Liu, and Cewu Lu. Human pose regression with residual log-likelihood estimation. In Int. Conf. Comput. Vis., pages 11025–11034, 2021.
  32. 32.Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool. Localvit: Bringing locality to vision transformers. arXiv preprint arXiv:2104.05707, 2021.
  33. 33.Yanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang, Wankou Yang, Shu-Tao Xia, and Erjin Zhou. Tokenpose: Learning keypoint tokens for human pose estimation. Int. Conf. Comput. Vis., 2021.
  34. 34.Kevin Lin, Lijuan Wang, and Zicheng Liu. End-to-end human pose and mesh reconstruction with transformers. In IEEE Conf. Comput. Vis. Pattern Recog., pages 1954–1963, 2021.
  35. 35.Kevin Lin, Lijuan Wang, and Zicheng Liu. Mesh graphormer. Int. Conf. Comput. Vis., 2021.
  36. 36.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Eur. Conf. Comput. Vis., 2014.
  37. 37.Wentao Liu, Jie Chen, Cheng Li, Chen Qian, Xiao Chu, and Xiaolin Hu. A cascaded inception of inception network with attention modulated feature fusion for human pose estimation. In AAAI, 2018.
  38. 38.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. Int. Conf. Comput. Vis., 2021.
  39. 39.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. Int. Conf. Learn. Represent., 2019.
  40. 40.Weian Mao, Yongtao Ge, Chunhua Shen, Zhi Tian, Xinlong Wang, and Zhibin Wang. Tfpose: Direct human pose estimation with transformers. arXiv preprint arXiv:2103.15320, 2021.
  41. 41.Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3d human pose estimation in the wild using improved cnn supervision. In International Conference on 3D vision (3DV), pages 506–516, 2017.
  42. 42.Gyeongsik Moon and Kyoung Mu Lee. I2l-meshnet: Image-to-lixel prediction network for accurate 3d human pose and mesh estimation from a single rgb image. In Eur. Conf. Comput. Vis., pages 752–768, 2020.
  43. 43.Alejandro Newell, Zhiao Huang, and Jia Deng. Associative embedding: End-to-end learning for joint detection and grouping. In Adv. Neural Inform. Process. Syst., 2017.
  44. 44.Alejandro Newell, Kaiyu Yang, and Jia Deng. Stacked hourglass networks for human pose estimation. In Eur. Conf. Comput. Vis., 2016.
  45. 45.Wanli Ouyang, Xiao Chu, and Xiaogang Wang. Multi-source deep learning for human pose estimation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 2329–2336, 2014.
  46. 46.Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Adv. Neural Inform. Process. Syst., 2017.
  47. 47.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. Dynamicvit: Efficient vision transformers with dynamic token sparsification. Adv. Neural Inform. Process. Syst., 2021.
  48. 48.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. Int. J. Comput. Vis., 115(3):211–252, 2015.
  49. 49.Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. Deep high-resolution representation learning for human pose estimation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 5693–5703, 2019.
  50. 50.Alexander Toshev and Christian Szegedy. Deeppose: Human pose estimation via deep neural networks. In IEEE Conf. Comput. Vis. Pattern Recog., pages 1653–1660, 2014.
  51. 51.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. ICML, 2021.
  52. 52.Gul Varol, Duygu Ceylan, Bryan Russell, Jimei Yang, Ersin Yumer, Ivan Laptev, and Cordelia Schmid. Bodynet: Volumetric inference of 3d human body shapes. In Eur. Conf. Comput. Vis., 2018.
  53. 53.Timo von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. In Eur. Conf. Comput. Vis., pages 601–617, 2018.
  54. 54.Jiahang Wang, Sheng Jin, Wentao Liu, Weizhong Liu, Chen Qian, and Ping Luo. When human pose estimation meets robustness: Adversarial algorithms and benchmarks. In IEEE Conf. Comput. Vis. Pattern Recog., pages 11855–11864, 2021.
  55. 55.Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, et al. Deep high-resolution representation learning for visual recognition. IEEE Trans. Pattern Anal. Mach. Intell., 2020.
  56. 56.Tao Wang, Li Yuan, Yunpeng Chen, Jiashi Feng, and Shuicheng Yan. Pnp-detr: Towards efficient visual analysis with transformers. In Int. Conf. Comput. Vis., pages 4661–4670, 2021.
  57. 57.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. Int. Conf. Comput. Vis., 2021.
  58. 58.Yulin Wang, Rui Huang, Shiji Song, Zeyi Huang, and Gao Huang. Not all images are worth 16x16 words: Dynamic vision transformers with adaptive sequence length. Adv. Neural Inform. Process. Syst., 2021.
  59. 59.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. Int. Conf. Comput. Vis., 2021.
  60. 60.Size Wu, Sheng Jin, Wentao Liu, Lei Bai, Chen Qian, Dong Liu, and Wanli Ouyang. Graph-based 3d multi-person pose estimation using multi-view images. In Int. Conf. Comput. Vis., pages 11148–11157, 2021.
  61. 61.Wayne Wu, Chen Qian, Shuo Yang, Quan Wang, Yici Cai, and Qiang Zhou. Look at boundary: A boundary-aware face alignment algorithm. In IEEE Conf. Comput. Vis. Pattern Recog., 2018.
  62. 62.Wenyan Wu and Shuo Yang. Leveraging intra and interdataset variations for robust face alignment. In IEEE Conf. Comput. Vis. Pattern Recog. Worksh., pages 150–159, 2017.
  63. 63.Bin Xiao, Haiping Wu, and Yichen Wei. Simple baselines for human pose estimation and tracking. In Eur. Conf. Comput. Vis., 2018.
  64. 64.Xuehan Xiong and Fernando De la Torre. Supervised descent method and its applications to face alignment. In IEEE Conf. Comput. Vis. Pattern Recog., pages 532–539, 2013.
  65. 65.Lumin Xu, Yingda Guan, Sheng Jin, Wentao Liu, Chen Qian, Ping Luo, Wanli Ouyang, and Xiaogang Wang. Vipnas: Efficient video pose estimation via neural architecture search. In IEEE Conf. Comput. Vis. Pattern Recog., pages 16072–16081, 2021.
  66. 66.Sen Yang, Zhibin Quan, Mu Nie, and Wankou Yang. Transpose: Towards explainable human pose estimation by transformer. Int. Conf. Comput. Vis., 2021.
  67. 67.Wei Yang, Shuang Li, Wanli Ouyang, Hongsheng Li, and Xiaogang Wang. Learning feature pyramids for human pose estimation. In Int. Conf. Comput. Vis., pages 1281–1290, 2017.
  68. 68.Rex Ying, Jiaxuan You, Christopher Morris, Xiang Ren, William L Hamilton, and Jure Leskovec. Hierarchical graph representation learning with differentiable pooling. Adv. Neural Inform. Process. Syst., 2018.
  69. 69.Kaimin Yu, Zhiyong Wang, Markus Hagenbuchner, and David Dagan Feng. Spectral embedding based facial expression recognition with multiple features. Neurocomputing, 129:136–145, 2014.
  70. 70.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. Int. Conf. Comput. Vis., 2021.
  71. 71.Yuhui Yuan, Rao Fu, Lang Huang, Weihong Lin, Chao Zhang, Xilin Chen, and Jingdong Wang. Hrformer: High-resolution transformer for dense prediction. Adv. Neural Inform. Process. Syst., 2021.
  72. 72.Xiaoyu Yue, Shuyang Sun, Zhanghui Kuang, Meng Wei, Philip HS Torr, Wayne Zhang, and Dahua Lin. Vision transformer with progressive sampling. In Int. Conf. Comput. Vis., pages 387–396, 2021.
  73. 73.Andrei Zanfir, Eduard Gabriel Bazavan, Hongyi Xu, William T Freeman, Rahul Sukthankar, and Cristian Sminchisescu. Weakly supervised 3d human pose and shape reconstruction with normalizing flows. In Eur. Conf. Comput. Vis., pages 465–481, 2020.
  74. 74.Wang Zeng, Wanli Ouyang, Ping Luo, Wentao Liu, and Xiaogang Wang. 3d human mesh regression with dense correspondence. In IEEE Conf. Comput. Vis. Pattern Recog., pages 7054–7063, 2020.
  75. 75.Ce Zheng, Sijie Zhu, Matias Mendieta, Taojiannan Yang, Chen Chen, and Zhengming Ding. 3d human pose estimation with spatial and temporal transformers. Int. Conf. Comput. Vis., 2021.
  76. 76.Shizhan Zhu, Cheng Li, Chen Change Loy, and Xiaoou Tang. Face alignment by coarse-to-fine shape searching. In IEEE Conf. Comput. Vis. Pattern Recog., pages 4998–5006, 2015.

Citation

MLA
Zeng, W., et al. “Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer”. arXiv, 2022, http://arxiv.org/abs/2204.08680v3.
APA
Zeng, W., Jin, S., Liu, W., Qian, C., Luo, P., Ouyang, W., & Wang, X. (2022). Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer. arXiv. http://arxiv.org/abs/2204.08680v3
Chicago
Zeng, W., S. Jin, W. Liu, et al. 2022. “Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer”. arXiv. http://arxiv.org/abs/2204.08680v3.
Harvard
Zeng, W. et al. (2022) “Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.08680v3.
Vancouver
1. Zeng W, Jin S, Liu W, Qian C, Luo P, Ouyang W, Wang X (2022) Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer. arXiv

BibTeX

@article{zeng2022not,
  title = {Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer},
  author = {Zeng, Wang and Jin, Sheng and Liu, Wentao and Qian, Chen and Luo, Ping and Ouyang, Wanli and Wang, Xiaogang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.08680v3},
  eprint = {2204.08680}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE