Built independently by an author, for readers. Read the story and support ChapterPal

keyword

tokenized pose representation

A tokenized pose representation is an approach in computer vision where human body postures are encoded as a collection of discrete tokens selected from a learned codebook instead of continuous joint coordinates or rotation parameters. By mapping continuous body configurations into discrete indices, this technique reframes continuous 3D pose estimation as a token prediction task, effectively constraining predicted postures to a distribution of anatomically valid human poses. This structured representation functions as a strong prior that reduces projection ambiguities, prevents physically implausible joint configurations, and improves the robustness of human mesh recovery when dealing with occlusions or limited visual evidence.

1 item

TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation

TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation

Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, Yao Feng, Michael J. Black

OrganizationsETH ZurichMax Planck Institute for Intelligent SystemsMeshcapade GmbH

Why you should read this

Presents a discrete token-based pose representation and an adaptive thresholding loss to prevent camera projection errors from degrading 3D human mesh recovery accuracy during 2D supervision.

We address the problem of regressing 3D human pose and shape from a single image, with a focus on 3D accuracy. The current best methods leverage large datasets of 3D pseudo-ground-truth (p-GT) and 2D keypoints, leading to robust performance. With such methods, however, we observe a paradoxical decline in 3D pose accuracy with increasing 2D accuracy. This is caused by biases in the p-GT and the use of an approximate camera projection model. We quantify the error induced by current camera models and show that fitting 2D keypoints and p-GT accurately causes incorrect 3D poses. Our analysis defines the invalid distances within which minimizing 2D and p-GT losses is detrimental. We use this to formulate a new loss, “Threshold-Adaptive Loss Scaling” (TALS), that penalizes gross 2D and p-GT errors but not smaller ones. With such a loss, there are many 3D poses that could equally explain the 2D evidence. To reduce this ambiguity we need a prior over valid human poses but such priors can introduce unwanted bias. To address this, we exploit a tokenized representation of human pose and reformulate the problem as token prediction. This restricts the estimated poses to the space of valid poses, effectively improving robustness to occlusion. Extensive experiments on the EMDB and 3DPW datasets show that our reformulated loss and tokenization allows us to train on in-the-wild data while improving 3D accuracy over the state-of-the-art. Our models and code are available for research at https://tokenhmr.is.tue.mpg.de.

Added

2026-09-26