Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer
Wang ZengSheng JinWentao LiuChen QianPing LuoWanli OuyangXiaogang Wang
Proposes Token Clustering Transformer (TCFormer), a vision architecture that dynamically clusters and merges image tokens into flexible shapes and sizes to allocate finer resolution to critical human regions while efficiently compressing background details for tasks like pose estimation and 3D mesh reconstruction.
Vision transformers are critical tools for human-centric visual analysis applications, such as augmented reality, action recognition, and human-computer interaction. However, standard vision transformers divide images into rigid, uniform grids, assigning equal computational focus to all image patches regardless of content. This uniform allocation is inefficient and sub-optimal for human analysis, where intricate foreground details—such as faces and hands—require high precision, while expansive background regions require minimal processing.
The article introduces and evaluates the Token Clustering Transformer (TCFormer), a novel framework designed to dynamically generate flexible vision tokens—computational image units—based on semantic content. The primary objective is to demonstrate that grouping tokens via progressive clustering allows models to preserve crucial fine details in target regions while reducing computational focus on background areas.
To achieve this, the approach employs a multi-stage architecture featuring two key components: a Clustering Token Merge block and a Multi-stage Token Aggregation head. The merge block uses a density-based nearest-neighbor clustering algorithm to group semantically similar tokens into flexible, non-grid shapes and sizes, guided by an importance score. The aggregation head progressively combines features across all stages while preserving fine-grained details. The authors evaluated the system across several benchmark datasets, including COCO-WholeBody for 2D whole-body pose estimation, 3DPW and Human3.6M for 3D human mesh reconstruction, WFLW for facial landmark alignment, and ImageNet-1K for general image classification.
The findings confirm that TCFormer consistently outperforms existing state-of-the-art models across all tested domains. On 2D whole-body pose estimation, TCFormer achieved 57.2% Average Precision, notably outperforming established baselines on difficult detail-heavy parts like hands by 6.2% Average Precision. On 3D mesh reconstruction, it reduced mean joint position error to 49.3 mm on 3DPW, achieving competitive or superior accuracy compared to specialized methods. For 2D face alignment, it delivered a lower normalized error (4.28%) than models utilizing extra boundary information. In general image classification, TCFormer attained an 82.4% Top-1 accuracy on ImageNet-1K, proving its broad utility beyond human-specific tasks.
These results demonstrate that dynamic, clustering-based token allocation substantially enhances feature representation without requiring prohibitive memory or computational overhead. By dedicating finer tokens to complex body areas and coarse tokens to backgrounds, systems can achieve higher operational accuracy in detail-sensitive vision tasks, directly benefiting real-time human tracking and interactive visual technologies.
Organizations developing vision systems should consider adopting dynamic token clustering mechanisms over rigid grid transformers for fine-grained localization tasks. Future technical roadmaps should explore applying this architecture to broader visual domains, including general object detection and semantic segmentation.
The primary limitation of the current method is that its clustering algorithm exhibits quadratic computational complexity relative to the number of input tokens, which can constrain processing speeds on very high-resolution images. While the authors suggest mitigating this by partitioning tokens into subsets prior to clustering, decision-makers should evaluate this scaling trade-off when deploying the model on high-resolution, real-time pipelines.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). This foundational paper establishes the standard uniform-grid patch tokenization for vision transformers that TCFormer directly aims to redesign through dynamic clustering.
- Paper: Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet, Li Yuan et al. (2021). It introduces progressive token folding to reduce token redundancy and capture local structure, providing an early precedent for progressive token aggregation.
- Paper: Transformer in Transformer, Kai Han et al. (2021). It analyzes the limitations of coarse, uniform patch partitioning and proposes nested tokenization to preserve fine-grained visual details.
- Paper: PVT v2: Improved baselines with Pyramid Vision Transformer, Wenhai Wang et al. (2021). It outlines hierarchical pyramid vision transformer backbones that serve as the foundational multi-stage baseline architecture adapted by TCFormer.
- Paper: Efficient Transformers: A Survey, Yi Tay et al. (2020). It provides a systematic taxonomy of efficient transformer mechanisms, including token clustering and sequence downsampling strategies.
- Paper: Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers, Siyuan Wei et al. (2023). This work advances dynamic token reduction by matching and squeezing discarded tokens into retained host tokens to prevent information loss during aggressive compression.
- Paper: CF-ViT: A General Coarse-to-Fine Method for Vision Transformer, Mengzhao Chen et al. (2023). It builds on the principle of non-uniform spatial allocation by implementing a dynamic coarse-to-fine mechanism that selectively re-splits informative regions.
- Paper: BiFormer: Vision Transformer with Bi-Level Routing Attention, Lei Zhu et al. (2023). It extends content-aware token routing by using dynamic bipartite graph clustering to allocate attention compute adaptively across semantic regions.
- Paper: TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation, Sai Kumar Dwivedi et al. (2024). It pushes human-centric visual analysis further by formalizing human mesh recovery as a discrete tokenized prediction problem.
- Paper: VisionZip: Longer is Better but Not Necessary in Vision Language Models, Senqiao Yang et al. (2025). It scales semantic token clustering and merging concepts to vision-language models by selectively merging redundant background tokens based on dominant attention proxies.
