TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation

Sai Kumar DwivediYu SunPriyanka PatelYao FengMichael J. Black

article2024CVPR131 citations

Presents a discrete token-based pose representation and an adaptive thresholding loss to prevent camera projection errors from degrading 3D human mesh recovery accuracy during 2D supervision.

Listen

Estimating accurate three-dimensional human pose and body shape from single, unconstrained photographs is essential for applications across computer vision, digital avatars, and motion analysis. Current state-of-the-art methods rely heavily on large-scale internet imagery paired with two-dimensional joint coordinates and estimated 3D annotations to achieve wide generalization. However, because precise camera calibration is unavailable in everyday imagery, standard models use approximate camera geometries. This introduces an unexpected trade-off: forcing algorithms to closely match 2D image coordinates systematically degrades 3D accuracy, frequently distorting poses by bending joints or tilting bodies to compensate for camera perspective mismatches.

The article aims to evaluate the source and magnitude of this camera-induced trade-off and demonstrate a new regression framework, named TokenHMR, that maintains high 3D accuracy while leveraging large, uncalibrated web datasets.

To address this challenge, the authors conducted quantitative evaluations using high-fidelity synthetic and real-world motion benchmarks to measure the error introduced by approximate camera models. Building on these empirical insights, the method introduces two main innovations. First, it implements Threshold-Adaptive Loss Scaling, a training objective that penalizes large 2D and 3D annotation errors while ignoring errors that fall below the expected threshold of camera distortion. Second, it discretizes continuous human joint angles into a learned vocabulary of valid body configurations via a discrete tokenized representation, converting the continuous pose regression task into a classification-based token prediction problem.

The empirical evaluations reveal several primary findings. In synthetic tests, forcing precise 2D alignment without ground-truth camera parameters drove average joint positioning errors above 146 to 300 millimeters in 3D space despite seemingly excellent 2D alignment. Incorporating unconstrained web datasets into standard baseline models degraded 3D accuracy on the EMDB benchmark by more than 17%. By replacing continuous regression with tokenized pose classification and applying the adaptive threshold loss, TokenHMR achieved a 7.6% reduction in mean per-joint position error and a 9.0% reduction in mean vertex error compared to the leading HMR2.0 baseline on the EMDB dataset. Furthermore, in challenging stress tests where images were cropped by up to 50% around the boundaries, the discrete prior limited accuracy degradation significantly better than continuous regression alternatives.

These findings demonstrate that optimizing exclusively for pixel-level 2D alignment under unknown camera settings introduces hidden risks and skews 3D spatial reconstructions. By constraining predictions to anatomically valid poses through a discrete vocabulary, organizations can use readily available, low-cost internet datasets without needing specialized multicamera capture rigs or precise lens metadata. This directly mitigates common visual artifacts, such as unnatural bent knees, and improves the reliability of single-image 3D body estimation in real-world environments with occlusions or visual truncation.

The authors recommend that teams building single-image 3D body estimation systems adopt relaxed 2D loss thresholds aligned with expected camera discrepancies rather than forcing strict 2D convergence. Organizations should also transition from continuous angle regression to discrete, tokenized pose representations to maintain anatomical validity. Future efforts should explore integrating dynamic camera estimation mechanisms and testing whether tokenized priors generalize effectively to temporal video tracking and multi-person interaction scenes.

The conclusions are supported by evaluations across standard benchmarks, including EMDB, 3DPW, BEDLAM, and AMASS. However, some limitations remain. Discretizing continuous poses incurs an inherent theoretical resolution loss of approximately 2.5 millimeters, although this remains negligible compared to overall baseline errors. Additionally, while the model significantly reduces perspective-induced distortions, single-view 3D depth recovery remains fundamentally ambiguous in the absence of exact intrinsic camera parameters.

arXiv: 2404.16752

No sufficiently relevant recommendations were found.

Cover for TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation

Abstract

We address the problem of regressing 3D human pose and shape from a single image, with a focus on 3D accuracy. The current best methods leverage large datasets of 3D pseudo-ground-truth (p-GT) and 2D keypoints, leading to robust performance. With such methods, however, we observe a paradoxical decline in 3D pose accuracy with increasing 2D accuracy. This is caused by biases in the p-GT and the use of an approximate camera projection model. We quantify the error induced by current camera models and show that fitting 2D keypoints and p-GT accurately causes incorrect 3D poses. Our analysis defines the invalid distances within which minimizing 2D and p-GT losses is detrimental. We use this to formulate a new loss, “Threshold-Adaptive Loss Scaling” (TALS), that penalizes gross 2D and p-GT errors but not smaller ones. With such a loss, there are many 3D poses that could equally explain the 2D evidence. To reduce this ambiguity we need a prior over valid human poses but such priors can introduce unwanted bias. To address this, we exploit a tokenized representation of human pose and reformulate the problem as token prediction. This restricts the estimated poses to the space of valid poses, effectively improving robustness to occlusion. Extensive experiments on the EMDB and 3DPW datasets show that our reformulated loss and tokenization allows us to train on in-the-wild data while improving 3D accuracy over the state-of-the-art. Our models and code are available for research at https://tokenhmr.is.tue.mpg.de.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. HPS Regression
  • 2.2. Pose Prior
  • 3. Camera/pose Bias
  • 4. TokenHMR
  • 4.1. Preliminaries
  • 4.2. Threshold-Adaptive Loss Scaling: TALS
  • 4.3. Tokenization
  • 4.4. Architecture
  • 4.5. Losses
  • 5. Experiments
  • 5.1. Implementation Details
  • 5.2. How to Alleviate the 3D Degradation Problem?
  • 5.3. How does the Token-based Prior Help?
  • 5.4. Ablation Study of Tokenizer
  • 6. Conclusion
  • 7. Acknowledgement
  • References

Knowls

  1. Knowl 1 — Camera-model mismatch creates a 2D–3D accuracy trade-off

    empirical result

    TokenHMR identifies incorrect camera assumptions as the cause of a systematic conflict between 2D alignment and 3D pose accuracy in monocular human mesh recovery. Existing regressors commonly use scaled-orthographic or perspective projection with fixed camera parameters rather than estimating the image’s focal length, rotation, and translation. For typical near-eye-height views, the legs are farther from the camera than the upper body and are therefore foreshortened; minimizing 2D reprojection error can consequently produce bent knees or an incorrectly tilted body in 3D. Pseudo-ground-truth obtained by fitting bodies to 2D observations inherits the same bias.

    On BEDLAM, which provides exact 3D bodies, 2D keypoints, and cameras, projecting ground-truth bodies with the HMR2.0 camera produced PCK scores of 0.660.66 at PCK0.5 and 0.860.86 at PCK1.0, whereas HMR2.0b predictions achieved 0.780.78 and 0.880.88, respectively. With a correct camera, both scores would be 1.01.0; thus, the lower 2D error of HMR2.0b is obtained by deviating from the true 3D pose and shape.

    The authors also optimized a modified SMPLify objective that preserves 2D alignment while maximizing 3D error:

    L=w2D∥Π(J3D,C)−J2Dg∥2−w3D∥J3D−J3Dg∥2+m,\mathcal{L}=w_{2D}\left\|\Pi(\mathbf{J}_{3D},\mathbf{C})-\mathbf{J}_{2D}^{g}\right\|_2-w_{3D}\left\|\mathbf{J}_{3D}-\mathbf{J}_{3D}^{g}\right\|_2+m,

    where J3D\mathbf{J}_{3D} and J3Dg\mathbf{J}_{3D}^{g} are predicted and ground-truth 3D joints, J2Dg\mathbf{J}_{2D}^{g} is the ground-truth 2D joint vector, Π\Pi is 3D-to-2D projection under the fixed HMR2.0 camera C\mathbf{C}, and w2D=4w_{2D}=4, w3D=40.5w_{3D}=40.5, and m=20m=20. After 100 optimization iterations, MPJPE reached 146146 mm while 2D alignment remained high; after 200 iterations, MPJPE exceeded 300300 mm. This demonstrates that good 2D alignment alone does not determine an accurate 3D pose when the camera is unknown or misspecified.

  2. Knowl 2 — Threshold-Adaptive Loss Scaling suppresses harmful small errors

    model/method

    Threshold-Adaptive Loss Scaling (TALS) retains 2D keypoint and pseudo-ground-truth supervision for large errors but weakens its influence once the error falls below the amount explainable by the incorrect camera model. The thresholds are estimated on BEDLAM by replacing predicted SMPL parameters with ground-truth parameters and then measuring the residual error under the HMR2.0 camera. The 2D threshold εJ2D\varepsilon_{J_{2D}} is the mean L1 error between these projected joints and the ground-truth 2D joints after normalizing coordinates by image width and scaling them to [−0.5,0.5][-0.5,0.5]. The pose threshold εθ\varepsilon_{\theta} is computed separately for each pose parameter from the mean geodesic distance between predicted and ground-truth joint rotations on SO(3)SO(3).

    For a pose-error magnitude eθe_{\theta} and a 2D-joint-error magnitude eJ2De_{J_{2D}}, TALS uses

    LθpGT={eθ2,eθ>εθ,αθeθ2,eθ≤εθ,LJ2D,pGT={eJ2D,eJ2D>εJ2D,αJ2DeJ2D,eJ2D≤εJ2D.\mathcal{L}_{\theta_{pGT}}= \begin{cases} e_{\theta}^{2}, & e_{\theta}>\varepsilon_{\theta},\\ \alpha_{\theta}e_{\theta}^{2}, & e_{\theta}\leq\varepsilon_{\theta}, \end{cases} \qquad \mathcal{L}_{J_{2D,pGT}}= \begin{cases} e_{J_{2D}}, & e_{J_{2D}}>\varepsilon_{J_{2D}},\\ \alpha_{J_{2D}}e_{J_{2D}}, & e_{J_{2D}}\leq\varepsilon_{J_{2D}}. \end{cases}

    Here αθ\alpha_{\theta} and αJ2D\alpha_{J_{2D}} are small positive scale factors; the authors use 0.010.01 for both. Thus, large disagreements still provide a strong corrective signal, while fitting the final camera-biased residual more accurately is discouraged.

  3. Knowl 3 — VQ-VAE converts SMPL pose into a discrete valid-pose vocabulary

    model/method

    TokenHMR learns a discrete prior over valid human poses with a Vector-Quantized Variational Autoencoder (VQ-VAE) trained on motion-capture data. The tokenizer represents the SMPL body pose as θ=[θ1,…,θ21]\boldsymbol{\theta}=[\theta_1,\ldots,\theta_{21}], where each joint representation θi∈R6\theta_i\in\mathbb{R}^{6}. An encoder EE maps the pose to MM latent vectors z=E(θ)=[z1,…,zM]\mathbf{z}=E(\boldsymbol{\theta})=[\mathbf{z}_1,\ldots,\mathbf{z}_M], with each zi∈Rdc\mathbf{z}_i\in\mathbb{R}^{d_c}. A learned codebook is CB={ck}k=1K\mathrm{CB}=\{\mathbf{c}_k\}_{k=1}^{K}, where each code ck∈Rdc\mathbf{c}_k\in\mathbb{R}^{d_c}. Each latent is quantized by nearest-neighbor lookup:

    z^i=arg⁡min⁡ck∈CB  ∥zi−ck∥2.\hat{\mathbf{z}}_i=\underset{\mathbf{c}_k\in\mathrm{CB}}{\arg\min}\;\|\mathbf{z}_i-\mathbf{c}_k\|_2.

    A decoder reconstructs continuous SMPL pose parameters from the quantized tokens. The tokenizer is trained with reconstruction, embedding, and commitment losses:

    LVQ=λRELRE+λE∥sg⁡[z]−e∥2+λC∥z−sg⁡[e]∥2,\mathcal{L}_{VQ}=\lambda_{RE}\mathcal{L}_{RE}+\lambda_E\|\operatorname{sg}[\mathbf{z}]-\mathbf{e}\|_2+\lambda_C\|\mathbf{z}-\operatorname{sg}[\mathbf{e}]\|_2,

    where sg⁡\operatorname{sg} stops gradients, e\mathbf{e} denotes the selected codebook embeddings, and λRE,λE,λC\lambda_{RE},\lambda_E,\lambda_C weight the three terms. Reconstruction combines pose and 3D-joint errors: LRE=L1(θg,θ)+L1(J3Dg,J3D)\mathcal{L}_{RE}=\mathcal{L}_1(\boldsymbol{\theta}^{g},\boldsymbol{\theta})+\mathcal{L}_1(\mathbf{J}_{3D}^{g},\mathbf{J}_{3D}), where superscript gg denotes ground truth. Exponential-moving-average codebook updates and codebook resets are used to avoid codebook collapse. The resulting codebook acts as a vocabulary that restricts predictions to poses represented by motion-capture data without explicitly favoring a frequently occurring individual pose.

  4. Knowl 4 — Differentiable logits connect image features to the frozen pose decoder

    model/method

    The image-to-pose model uses a Vision Transformer backbone and a transformer decoder to extract image features. Instead of predicting a discrete code index with a non-differentiable lookup, the network predicts logits Q∈RM×K\mathbf{Q}\in\mathbb{R}^{M\times K} for the MM pose tokens and KK codebook entries. The logits are converted into soft assignments and multiplied by the pretrained codebook:

    zˉ=softmax⁡(Q) CB∈RM×dc≈z^.\bar{\mathbf{z}}=\operatorname{softmax}(\mathbf{Q})\,\mathrm{CB}\in\mathbb{R}^{M\times d_c}\approx\hat{\mathbf{z}}.

    The approximate quantized features zˉ\bar{\mathbf{z}} are passed through the VQ-VAE decoder to produce the body pose. The decoder is frozen during TokenHMR training, so the learned motion-capture prior remains fixed while the image encoder learns to select appropriate tokens. Global orientation, hand pose, body shape, and camera are predicted with separate linear heads; only the body-pose branch is routed through the token logits and frozen decoder. In the final model, the body-pose branch uses four blocks of linear layers, each containing two MLPs with GELU activations.

  5. Knowl 5 — TokenHMR combines ordinary ground-truth supervision with TALS

    equation

    For an input image II, TokenHMR predicts SMPL pose θ\boldsymbol{\theta}, shape β∈R10\boldsymbol{\beta}\in\mathbb{R}^{10}, and perspective camera parameters T\mathbf{T}. SMPL produces a mesh with N=6890N=6890 vertices and 3D joints J3D\mathbf{J}_{3D} obtained by a learned linear joint regressor. On datasets with reliable 3D ground truth, TokenHMR uses pose, shape, 3D-joint, and 2D-reprojection losses:

    LGT=λθLθ(θ,θg)+λβLβ(β,βg)+λ3DL3D(J3D,J3Dg)+λ2DL2D(J2D,J2Dg).\mathcal{L}_{GT}=\lambda_{\theta}\mathcal{L}_{\theta}(\boldsymbol{\theta},\boldsymbol{\theta}^{g})+\lambda_{\beta}\mathcal{L}_{\beta}(\boldsymbol{\beta},\boldsymbol{\beta}^{g})+\lambda_{3D}\mathcal{L}_{3D}(\mathbf{J}_{3D},\mathbf{J}_{3D}^{g})+\lambda_{2D}\mathcal{L}_{2D}(\mathbf{J}_{2D},\mathbf{J}_{2D}^{g}).

    Here θg\boldsymbol{\theta}^{g}, βg\boldsymbol{\beta}^{g}, J3Dg\mathbf{J}_{3D}^{g}, and J2Dg\mathbf{J}_{2D}^{g} are ground-truth pose, shape, 3D joints, and 2D joints, respectively. For in-the-wild data with SMPL pseudo-ground truth, the total objective is

    LTotal=LGT+LθpGT+LJ2D,pGT,\mathcal{L}_{Total}=\mathcal{L}_{GT}+\mathcal{L}_{\theta_{pGT}}+\mathcal{L}_{J_{2D,pGT}},

    where the last two terms are the threshold-scaled pose and 2D losses. This lets TokenHMR use large in-the-wild datasets for generalization while reducing their camera and annotation bias.

  6. Knowl 6 — Training uses motion capture, standard 3D, in-the-wild, and synthetic data

    experimental setup

    The tokenizer is trained in a first stage on the standard AMASS training split and MOYO training data. Its encoder and decoder contain one ResNet block and four 1D convolutions each; λRE=50\lambda_{RE}=50, λE=1\lambda_E=1, and λC=1\lambda_C=1. It is trained for 150,000 iterations with batch size 256 and learning rate 2×10−42\times10^{-4}, with progressively increasing random joint noise beginning at 10−310^{-3} and updated every 5,000 iterations. The selected tokenizer has M=160M=160 tokens and a 2048×2562048\times256 codebook.

    TokenHMR uses a ViT-H/16 image backbone, a standard transformer decoder, and ViTPose initialization. It is trained for 100,000 iterations on four NVIDIA RTX 6000 GPUs with batch size 256 and learning rate 10−510^{-5}. Its loss weights are λθ=10−3\lambda_{\theta}=10^{-3}, λβ=5×10−4\lambda_{\beta}=5\times10^{-4}, λJ2D=10−2\lambda_{J_{2D}}=10^{-2}, and λJ3D=5×10−2\lambda_{J_{3D}}=5\times10^{-2}; both TALS scale factors are 0.010.01.

    Training combines standard datasets—Human3.6M, MPI-INF-3DHP, COCO, and MPII—denoted SD; in-the-wild 2D datasets InstaVariety, AVA, and AI Challenger with pseudo-ground truth, denoted ITW; and BEDLAM synthetic data with accurate 3D ground truth, denoted BL. Evaluation uses MVE, MPJPE, and PA-MPJPE in millimeters on EMDB and 3DPW, with lower values better. MVE is mean vertex error, MPJPE is mean per-joint position error, and PA-MPJPE is MPJPE after Procrustes alignment.

  7. Knowl 7 — TokenHMR improves 3D accuracy on EMDB and 3DPW

    data/table

    The following results compare methods trained with different data and loss/prior choices. All entries are errors in millimeters; lower is better. The strongest combined model, SD + ITW + BL TokenHMR, obtains the best or near-best results across the reported metrics and reduces EMDB MPJPE from 99.399.3 mm for the matched HMR2.0 baseline to 91.791.7 mm, a 7.6%7.6\% reduction.

    Could not parse LaTeX table

    Adding ITW data to SD reduces HMR2.0’s EMDB MPJPE from 97.897.8 to 118.5118.5 mm despite improving 2D alignment. Adding BEDLAM recovers some accuracy but does not remove the degradation. TALS improves the matched HMR2.0 model, while replacing continuous pose regression with the token prior gives a larger improvement; combining the tokenized pose representation with the full training mixture gives the best overall result.

  8. Knowl 8 — The token prior is more robust to image truncation

    data/table

    To test whether the discrete pose prior helps when image evidence is ambiguous, the authors evenly crop either 30% or 50% from the image boundaries on EMDB. TokenHMR degrades less than HMR2.0 at both crop levels, especially under the more severe 50% crop. Values are errors in millimeters, and the parenthesized values are increases relative to the uncropped condition.

    Could not parse LaTeX table

    Both models use identical backbones and training data. Under the 50% crop, TokenHMR’s MPJPE increase is 34.2834.28 mm versus 38.5938.59 mm for HMR2.0, and its PA-MPJPE increase is 23.2723.27 mm versus 27.4827.48 mm. The constrained pose vocabulary therefore improves robustness when parts of the person are absent or visually ambiguous.

  9. Knowl 9 — Tokenizer capacity and training data determine reconstruction quality

    data/table

    The tokenizer ablation evaluates mean vertex error (MVE) and MPJPE in millimeters on the AMASS test set and MOYO validation set. Most variants are trained on AMASS alone to expose out-of-distribution behavior on MOYO; the final starred variant is trained on AMASS plus MOYO and is used by TokenHMR. Increasing codebook entries helps more than increasing code dimension, while increasing the number of pose tokens improves reconstruction until a practical plateau around 160 tokens.

    Could not parse LaTeX table

    Pose noise slightly worsens in-distribution AMASS reconstruction but improves the out-of-distribution MOYO result. Training on both motion-capture datasets substantially improves MOYO reconstruction, supporting the use of the 2048×2562048\times256, 160-token tokenizer for TokenHMR.

  10. Knowl 10 — Discretization introduces a small, measurable pose-reconstruction cost

    empirical result

    The tokenized representation is not lossless: the authors report approximately a 2.52.5 mm loss in 3D accuracy attributable to discretizing pose. They characterize this cost as roughly 20 times smaller than the error of state-of-the-art human mesh regressors on real data. The tokenizer ablations also show that increasing the number of pose tokens from 80 to 160 substantially reduces reconstruction error, while increasing it further to 320 yields only a small additional improvement. Thus, TokenHMR accepts a limited reconstruction error in exchange for restricting predictions to a valid-pose vocabulary and improving robustness to camera bias and occlusion.

Coverage note — No substantial contributed material was omitted; implementation details, camera-bias analysis, TALS, tokenization, architecture, objectives, ablations, robustness tests, and benchmark results are covered.

References

  1. 1.Ijaz Akhter and Michael J. Black. Pose-conditioned joint angle limits for 3D human pose reconstruction. In Computer Vision and Pattern Recognition (CVPR), 2015.
  2. 2.Mykhaylo Andriluka, Leonid Pishchulin, Peter Gehler, and Bernt Schiele. 2d human pose estimation: New benchmark and state of the art analysis. In Computer Vision and Pattern Recognition (CVPR), 2014.
  3. 3.Michael J Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang. Bedlam: A synthetic dataset of bodies exhibiting detailed lifelike animated motion. In CVPR, pages 8726–8737, 2023.
  4. 4.Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J. Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In European Conference on Computer Vision (ECCV), 2016.
  5. 5.Hongsuk Choi, Gyeongsik Moon, JoonKyu Park, and Kyoung Mu Lee. Learning to estimate robust 3d human mesh from in-the-wild crowded scenes. In Computer Vision and Pattern Recognition (CVPR), pages 1475–1484, 2022.
  6. 6.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021.
  7. 7.Sai Kumar Dwivedi, Nikos Athanasiou, Muhammed Kocabas, and Michael J. Black. Learning to regress bodies from images using differentiable semantic rendering. In International Conference on Computer Vision (ICCV), 2021.
  8. 8.Sai Kumar Dwivedi, Cordelia Schmid, Hongwei Yi, Michael J. Black, and Dimitrios Tzionas. POCO: 3D pose and shape estimation using confidence. In International Conference on 3D Vision (3DV), 2024.
  9. 9.Zigang Geng, Chunyu Wang, Yixuan Wei, Ze Liu, Houqiang Li, and Han Hu. Human pose as compositional tokens. In Computer Vision and Pattern Recognition (CVPR), 2023.
  10. 10.G. Georgakis, Ren Li, S. Karanam, Terrence Chen, J. Kosecka, and Ziyan Wu. Hierarchical kinematic human mesh recovery. European Conference on Computer Vision (ECCV), 2020.
  11. 11.Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Malik. Humans in 4D: Reconstructing and tracking humans with transformers. In International Conference on Computer Vision (ICCV), 2023.
  12. 12.Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. Ava: A video dataset of spatio-temporally localized atomic visual actions. In Computer Vision and Pattern Recognition (CVPR), pages 6047–6056, 2018.
  13. 13.Chuan Guo, Xinxin Zuo, Sen Wang, and Li Cheng. Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts. In European Conference on Computer Vision (ECCV), 2022.
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Computer Vision and Pattern Recognition (CVPR), 2016.
  15. 15.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv: 1606.08415, 2016.
  16. 16.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments. Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2014.
  17. 17.Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen. Motiongpt: Human motion as a foreign language. arXiv preprint arXiv: 2306.14795, 2023.
  18. 18.Sam Johnson and Mark Everingham. Learning effective human pose estimation from inaccurate annotation. In Computer Vision and Pattern Recognition (CVPR), 2011.
  19. 19.Hanbyul Joo, Natalia Neverova, and Andrea Vedaldi. Exemplar fine-tuning for 3D human pose fitting towards in-the-wild 3D human pose estimation. In International Conference on 3D Vision (3DV), 2020.
  20. 20.Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In Computer Vision and Pattern Recognition (CVPR), 2018.
  21. 21.Angjoo Kanazawa, Jason Y. Zhang, Panna Felsen, and Jitendra Malik. Learning 3D human dynamics from video. In Computer Vision and Pattern Recognition (CVPR), 2019.
  22. 22.Manuel Kaufmann, Jie Song, Chen Guo, Kaiyue Shen, Tianjian Jiang, Chengcheng Tang, Juan Jose Zárate, and Otmar Hilliges. EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild. In International Conference on Computer Vision (ICCV), 2023.
  23. 23.Muhammed Kocabas, Nikos Athanasiou, and Michael J. Black. VIBE: Video inference for human body pose and shape estimation. In Computer Vision and Pattern Recognition (CVPR), 2020.
  24. 24.Muhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, and Michael J. Black. PARE: Part attention regressor for 3D human body estimation. In International Conference on Computer Vision (ICCV), 2021.
  25. 25.Muhammed Kocabas, Chun-Hao P. Huang, Joachim Tesch, Lea Muller, Otmar Hilliges, and Michael J. Black. SPEC: Seeing people in the wild with an estimated camera. In International Conference on Computer Vision (ICCV), 2021.
  26. 26.Nikos Kolotouros, Georgios Pavlakos, Michael J. Black, and Kostas Daniilidis. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. In International Conference on Computer Vision (ICCV), 2019.
  27. 27.Nikos Kolotouros, Georgios Pavlakos, and Kostas Daniilidis. Convolutional mesh regression for single-image human shape reconstruction. In Computer Vision and Pattern Recognition (CVPR), 2019.
  28. 28.Nikos Kolotouros, Georgios Pavlakos, Dinesh Jayaraman, and Kostas Daniilidis. Probabilistic modeling for human mesh recovery. In International Conference on Computer Vision (ICCV), 2021.
  29. 29.Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J. Black, and Peter Gehler. Unite the people: Closing the loop between 3D and 2D human representations. In Computer Vision and Pattern Recognition (CVPR), 2017.
  30. 30.Jiefeng Li, Chao Xu, Zhicun Chen, Siyuan Bian, Lixin Yang, and Cewu Lu. HybrIK: A hybrid analytical-neural inverse kinematics solution for 3D human pose and shape estimation. In Computer Vision and Pattern Recognition (CVPR), 2021.
  31. 31.Zhihao Li, Jianzhuang Liu, Zhensong Zhang, Songcen Xu, and Youliang Yan. CLIFF: Carrying location information in full frames into human pose and shape estimation. In European Conference on Computer Vision (ECCV), 2022.
  32. 32.Jing Lin, Ailing Zeng, Haoqian Wang, Lei Zhang, and Yu Li. One-stage 3D whole-body mesh recovery with component aware transformer. In Computer Vision and Pattern Recognition (CVPR), 2023.
  33. 33.Kevin Lin, Lijuan Wang, and Zicheng Liu. Mesh graphormer. In International Conference on Computer Vision (ICCV), 2021.
  34. 34.Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision (ECCV), 2014.
  35. 35.Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. In Transactions on Graphics (TOG), 2015.
  36. 36.Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: Archive of motion capture as surface shapes. In International Conference on Computer Vision (ICCV), 2019.
  37. 37.Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal V. Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3d human pose estimation in the wild using improved cnn supervision. In International Conference on 3D Vision (3DV), 2017.
  38. 38.Marko Mihajlovic, Shunsuke Saito, Aayush Bansal, Michael Zollhoefer, and Siyu Tang. COAP: Compositional articulated occupancy of people. In Computer Vision and Pattern Recognition (CVPR), 2022.
  39. 39.Gyeongsik Moon, Hongsuk Choi, and Kyoung Mu Lee. Neuralannot: Neural annotator for 3d human mesh training sets. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2299–2307, 2022.
  40. 40.Mohamed Omran, Christoph Lassner, Gerard Pons-Moll, Peter Gehler, and Bernt Schiele. Neural body fitting: Unifying deep learning and model based human pose and shape estimation. In International Conference on 3D Vision (3DV), 2018.
  41. 41.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Computer Vision and Pattern Recognition (CVPR), 2019.
  42. 42.A. Ramesh, Mikhail Pavlov, Gabriel Goh, S. Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. International Conference on Machine Learning (ICML), 2021.
  43. 43.Ali Razavi, Aaron Van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with vq-vae-2. Conference on Neural Information Processing Systems (NeurIPS), 2019.
  44. 44.Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul Joo. PIFuHD: Multi-level pixel-aligned implicit function for high-resolution 3D human digitization. In Computer Vision and Pattern Recognition (CVPR), 2020.
  45. 45.Istvan Sárándi, Timm Linder, Kai O. Arras, and Bastian Leibe. Metric-scale truncation-robust heatmaps for 3D human pose estimation. In IEEE Int Conf Automatic Face and Gesture Recognition (FG), 2020.
  46. 46.Akash Sengupta, Ignas Budvytis, and Roberto Cipolla. HuManiFlow: Ancestor-Conditioned Normalising Flows on SO(3) Manifolds for Human Pose and Shape Distribution Estimation. In Computer Vision and Pattern Recognition (CVPR), 2023.
  47. 47.Yu Sun, Qian Bao, Wu Liu, Yili Fu, Black Michael J., and Tao Mei. Monocular, One-stage, Regression of Multiple 3D People. In International Conference on Computer Vision (ICCV), 2021.
  48. 48.Yu Sun, Wu Liu, Qian Bao, Yili Fu, Tao Mei, and Michael J Black. Putting People in their Place: Monocular Regression of 3D People in Depth. In Computer Vision and Pattern Recognition (CVPR), 2022.
  49. 49.Garvita Tiwari, Dimitrije Antic, Jan Eric Lenssen, Nikolaos Sarafianos, Tony Tung, and Gerard Pons-Moll. Pose-ndf: Modeling human pose manifolds with neural distance fields. In European Conference on Computer Vision (ECCV), 2022.
  50. 50.Shashank Tripathi, Lea Muller, Chun-Hao P. Huang, Taheri Omid, Michael J. Black, and Dimitrios Tzionas. 3D human pose estimation via intuitive physics. In Computer Vision and Pattern Recognition (CVPR), 2023.
  51. 51.Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Conference on Neural Information Processing Systems (NeurIPS), 2017.
  52. 52.Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Conference on Neural Information Processing Systems (NeurIPS), 2017.
  53. 53.Timo von Marcard, Roberto Henschel, Michael J. Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3D human pose in the wild using IMUs and a moving camera. In European Conference on Computer Vision (ECCV), 2018.
  54. 54.Wenjia Wang, Yongtao Ge, Haiyi Mei, Zhongang Cai, Qingping Sun, Yanjun Wang, Chunhua Shen, Lei Yang, and Taku Komura. Zolly: Zoom focal length correctly for perspective-distorted human mesh reconstruction. In International Conference on Computer Vision (ICCV), 2023.
  55. 55.Jiahong Wu, He Zheng, Bo Zhao, Yixin Li, Baoming Yan, Rui Liang, Wenjia Wang, Shipei Zhou, Guosen Lin, Yanwei Fu, et al. Ai challenger: A large-scale dataset for going deeper in image understanding. arXiv, 2017.
  56. 56.Yuliang Xiu, Jinlong Yang, Xu Cao, Dimitrios Tzionas, and Michael J. Black. ECON: Explicit clothed humans optimized via normal integration. In Computer Vision and Pattern Recognition (CVPR), 2023.
  57. 57.Hongyi Xu, Eduard Gabriel Bazavan, Andrei Zanfir, William T. Freeman, Rahul Sukthankar, and Cristian Sminchisescu. GHUM & GHUML: Generative 3D human shape and articulated pose models. In Computer Vision and Pattern Recognition (CVPR), 2020.
  58. 58.Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. ViTPose: Simple vision transformer baselines for human pose estimation. In Conference on Neural Information Processing Systems (NeurIPS), 2022.
  59. 59.Hongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang, Yebin Liu, Limin Wang, and Zhenan Sun. Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop. In International Conference on Computer Vision (ICCV), pages 11446–11456, 2021.
  60. 60.Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen. T2m-gpt: Generating human motion from textual descriptions with discrete representations. In Computer Vision and Pattern Recognition (CVPR), 2023.

Citation

MLA
Dwivedi, S. K., et al. “TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation”. arXiv, 2024, http://arxiv.org/abs/2404.16752v1.
APA
Dwivedi, S. K., Sun, Y., Patel, P., Feng, Y., & Black, M. J. (2024). TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation. arXiv. http://arxiv.org/abs/2404.16752v1
Chicago
Dwivedi, S. K., Y. Sun, P. Patel, Y. Feng, and M. J. Black. 2024. “TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation”. arXiv. http://arxiv.org/abs/2404.16752v1.
Harvard
Dwivedi, S.K. et al. (2024) “TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.16752v1.
Vancouver
1. Dwivedi SK, Sun Y, Patel P, Feng Y, Black MJ (2024) TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation. arXiv

BibTeX

@article{dwivedi2024tokenhmr,
  title = {TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation},
  author = {Dwivedi, Sai Kumar and Sun, Yu and Patel, Priyanka and Feng, Yao and Black, Michael J.},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.16752v1},
  eprint = {2404.16752}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE