TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation
Sai Kumar DwivediYu SunPriyanka PatelYao FengMichael J. Black
Presents a discrete token-based pose representation and an adaptive thresholding loss to prevent camera projection errors from degrading 3D human mesh recovery accuracy during 2D supervision.
Estimating accurate three-dimensional human pose and body shape from single, unconstrained photographs is essential for applications across computer vision, digital avatars, and motion analysis. Current state-of-the-art methods rely heavily on large-scale internet imagery paired with two-dimensional joint coordinates and estimated 3D annotations to achieve wide generalization. However, because precise camera calibration is unavailable in everyday imagery, standard models use approximate camera geometries. This introduces an unexpected trade-off: forcing algorithms to closely match 2D image coordinates systematically degrades 3D accuracy, frequently distorting poses by bending joints or tilting bodies to compensate for camera perspective mismatches.
The article aims to evaluate the source and magnitude of this camera-induced trade-off and demonstrate a new regression framework, named TokenHMR, that maintains high 3D accuracy while leveraging large, uncalibrated web datasets.
To address this challenge, the authors conducted quantitative evaluations using high-fidelity synthetic and real-world motion benchmarks to measure the error introduced by approximate camera models. Building on these empirical insights, the method introduces two main innovations. First, it implements Threshold-Adaptive Loss Scaling, a training objective that penalizes large 2D and 3D annotation errors while ignoring errors that fall below the expected threshold of camera distortion. Second, it discretizes continuous human joint angles into a learned vocabulary of valid body configurations via a discrete tokenized representation, converting the continuous pose regression task into a classification-based token prediction problem.
The empirical evaluations reveal several primary findings. In synthetic tests, forcing precise 2D alignment without ground-truth camera parameters drove average joint positioning errors above 146 to 300 millimeters in 3D space despite seemingly excellent 2D alignment. Incorporating unconstrained web datasets into standard baseline models degraded 3D accuracy on the EMDB benchmark by more than 17%. By replacing continuous regression with tokenized pose classification and applying the adaptive threshold loss, TokenHMR achieved a 7.6% reduction in mean per-joint position error and a 9.0% reduction in mean vertex error compared to the leading HMR2.0 baseline on the EMDB dataset. Furthermore, in challenging stress tests where images were cropped by up to 50% around the boundaries, the discrete prior limited accuracy degradation significantly better than continuous regression alternatives.
These findings demonstrate that optimizing exclusively for pixel-level 2D alignment under unknown camera settings introduces hidden risks and skews 3D spatial reconstructions. By constraining predictions to anatomically valid poses through a discrete vocabulary, organizations can use readily available, low-cost internet datasets without needing specialized multicamera capture rigs or precise lens metadata. This directly mitigates common visual artifacts, such as unnatural bent knees, and improves the reliability of single-image 3D body estimation in real-world environments with occlusions or visual truncation.
The authors recommend that teams building single-image 3D body estimation systems adopt relaxed 2D loss thresholds aligned with expected camera discrepancies rather than forcing strict 2D convergence. Organizations should also transition from continuous angle regression to discrete, tokenized pose representations to maintain anatomical validity. Future efforts should explore integrating dynamic camera estimation mechanisms and testing whether tokenized priors generalize effectively to temporal video tracking and multi-person interaction scenes.
The conclusions are supported by evaluations across standard benchmarks, including EMDB, 3DPW, BEDLAM, and AMASS. However, some limitations remain. Discretizing continuous poses incurs an inherent theoretical resolution loss of approximately 2.5 millimeters, although this remains negligible compared to overall baseline errors. Additionally, while the model significantly reduces perspective-induced distortions, single-view 3D depth recovery remains fundamentally ambiguous in the absence of exact intrinsic camera parameters.
- Paper: End-to-End Recovery of Human Shape and Pose, Angjoo Kanazawa et al. (2017). This paper establishes the foundational Human Mesh Recovery (HMR) framework that uses parametric SMPL regression and 2D-to-3D reprojection losses, which TokenHMR directly diagnoses and redesigns to avoid perspective distortions.
- Paper: Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image, Federica Bogo et al. (2016). It introduces the SMPLify optimization paradigm for fitting parametric body models to 2D keypoints, providing essential context on the 2D joint alignment objectives evaluated and modified in TokenHMR.
- Paper: AMASS: Archive of Motion Capture As Surface Shapes, Naureen Mahmood et al. (2019). This work introduces the AMASS motion capture dataset and standardized SMPL parameterizations that form the training backbone and anatomical ground truth used by TokenHMR.
- Paper: A Simple Yet Effective Baseline for 3d Human Pose Estimation, Julieta Martinez et al. (2017). It establishes fundamental baselines for decoupling 2D joint estimation from 3D spatial pose lifting, highlighting error propagation mechanisms relevant to TokenHMR.
- Paper: Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments, Catalin Ionescu et al. (2014). This benchmark paper establishes standard evaluation protocols and metrics for 3D human pose recovery that are foundational to measuring camera and mesh recovery errors.
No sufficiently relevant recommendations were found.
