MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video

Jinlu ZhangZhigang TuJianyu YangYujin ChenJunsong Yuan

article2022CVPR436 citations

Proposes MixSTE, a sequence-to-sequence transformer architecture that alternates between modeling individual joint trajectories over time and spatial joint dependencies within frames to achieve state-of-the-art accuracy in video-based 3D human pose estimation.

Listen

Estimating three-dimensional human body poses from standard two-dimensional video is a fundamental computer vision capability with significant value for robotics, action recognition, and virtual human systems. However, extracting accurate 3D positions from monocular video remains challenging because different 3D poses can correspond to the identical 2D projection. While recent approaches have leveraged attention-based transformer models across video sequences, they typically fail to capture the distinct, independent motion trajectories of individual body joints over time, and they commonly predict only a single central frame per sequence, which introduces substantial computational redundancy.

The article demonstrates and evaluates MixSTE, a sequence-to-sequence neural network architecture designed to reconstruct full 3D pose sequences from 2D keypoints. MixSTE alternately stacks spatial attention blocks to capture relationships across body joints within each frame and temporal attention blocks to separately track the movement path of each individual joint across time.

To establish credibility and performance benchmarks, the authors evaluated MixSTE across three widely recognized experimental datasets: Human3.6M, MPI-INF-3DHP, and HumanEva. The model was evaluated using standard 2D pose detection inputs as well as ground-truth 2D keypoints, testing both short- and long-term sequence lengths ranging from 1 to 243 frames.

The findings confirm that MixSTE establishes new state-of-the-art performance across all benchmarked datasets. On the Human3.6M benchmark using standard 2D inputs, the model reduces mean joint position error to 40.9 mm, achieving a 7.6% error reduction over prior leading transformer methods, while yielding a 10.9% improvement under rigid alignment metrics. When provided with ground-truth 2D keypoints, MixSTE achieves a 31.0% error reduction over previous transformer baselines. Ablation experiments demonstrated that treating each joint as a separate temporal token drastically cut computational operations per frame from over 186,000 to 645 megaflops, while the full sequence-to-sequence output significantly increased inference processing speed.

These results show that separating joint trajectories and enforcing end-to-end sequence coherence substantially boosts accuracy while resolving the high latency and computational inefficiencies common in prior methods. For enterprise applications, these improvements reduce hardware infrastructure costs and enable practical deployment in real-time or resource-constrained environments where rapid video processing is critical.

Stakeholders deploying 3D human pose systems should consider adopting sequence-to-sequence alternating spatio-temporal architectures to maximize throughput and positional precision. For applications involving small datasets, fine-tuning pretrained models from larger datasets is recommended to ensure strong generalization. Future engineering and research efforts should prioritize modeling input noise distributions and pairing the architecture with robust upstream 2D detectors to mitigate performance degradation caused by missing or inaccurate 2D keypoints.

No sufficiently relevant recommendations were found.

Cover for MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video

Abstract

Recent transformer-based solutions have been introduced to estimate 3D human pose from 2D keypoint sequence by considering body joints among all frames globally to learn spatio-temporal correlation. We observe that the motions of different joints differ significantly. However, the previous methods cannot efficiently model the solid inter-frame correspondence of each joint, leading to insufficient learning of spatial-temporal correlation. We propose MixSTE (Mixed Spatio-Temporal Encoder), which has a temporal transformer block to separately model the temporal motion of each joint and a spatial transformer block to learn inter-joint spatial correlation. These two blocks are utilized alternately to obtain better spatio-temporal feature encoding. In addition, the network output is extended from the central frame to entire frames of the input video, thereby improving the coherence between the input and output sequences. Extensive experiments are conducted on three benchmarks (Human3.6M, MPI-INF-3DHP, and HumanEva). The results show that our model outperforms the state-of-the-art approach by 10.9% P-MPJPE and 7.6% MPJPE. The code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Our Approach
  • 3.1 Mixed Spatio-Temporal Encoder
  • 3.1.1 Separate Temporal Correlation Learning
  • 3.1.2 Spatial Correlation Learning
  • 3.1.3 Alternating design with Seq2seq
  • 3.2 Transformer Block in MixSTE
  • 3.3 Loss Function
  • 4 Experiment
  • 4.1 Datasets and Evaluation Protocols
  • 4.2 Implementation Details
  • 4.3 Comparison with State-of-the-art Methods
  • 4.4 Ablation Study
  • 4.5 Qualitative Results
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — MixSTE Spatio-Temporal Transformer Architecture for 3D Pose Sequence Estimation

    model/method

    The Mixed Spatio-Temporal Encoder (MixSTE) is a sequence-to-sequence (seq2seq) transformer framework designed to lift a sequence of 2D human pose keypoints CN,T∈RN×T×2C_{N,T} \in \mathbb{R}^{N \times T \times 2} (NN body joints across TT consecutive frames) directly into a complete 3D human pose sequence Out∈RN×T×3\text{Out} \in \mathbb{R}^{N \times T \times 3}.

    The input 2D coordinates are first linearly projected into a feature tensor PN,T∈RN×T×dmP_{N,T} \in \mathbb{R}^{N \times T \times d_m}, where dmd_m denotes the joint feature embedding dimension. Spatial positional embeddings Es-pos∈RN×dmE_{s\text{-}pos} \in \mathbb{R}^{N \times d_m} and temporal positional embeddings Et-pos∈RT×dmE_{t\text{-}pos} \in \mathbb{R}^{T \times d_m} are incorporated in the initial encoder layer to retain spatial skeleton layout and temporal frame order.

    The core of the encoder consists of an alternating stack of dld_l loops, where each loop comprises:

    1. A Spatial Transformer Block (STB) that executes multi-head self-attention (MSA) across the NN joint tokens within each frame, capturing inter-joint structural and kinematic dependencies.
    2. A Temporal Transformer Block (TTB) that applies MSA across the TT frame tokens independently for each joint, modeling individual joint motion trajectories.

    Throughout the dld_l loops, the joint feature dimension is held constant at dmd_m. After processing through the stacked alternating blocks, a linear regression head projects the resulting representation X∈RN×T×dmX \in \mathbb{R}^{N \times T \times d_m} from dimension dmd_m to 33, yielding the output 3D joint coordinate sequence.

  2. Knowl 2 — Separate Temporal Motion Modeling via Joint Separation

    model/method

    Prior spatio-temporal transformers for 3D pose estimation flatten all NN joints of a single frame into a joint feature token of dimension N×dmN \times d_m for temporal modeling. In contrast, MixSTE introduces joint separation, which isolates the temporal trajectory of each individual joint i∈{1,…,N}i \in \{1, \dots, N\} across all TT frames as an independent sequence of tokens pi=(pi,1,pi,2,…,pi,T)∈R1×T×dmp_i = (p_{i,1}, p_{i,2}, \dots, p_{i,T}) \in \mathbb{R}^{1 \times T \times d_m}, where pi,j∈Rdmp_{i,j} \in \mathbb{R}^{d_m} is the feature of joint ii at frame jj.

    The temporal transformer function F\mathcal{F} processes each joint trajectory in parallel:

    Xlt=Concat(F(p1,1,…,p1,T),…,F(pN,1,…,pN,T))∈RN×T×dmX_l^t = \text{Concat}\left(\mathcal{F}(p_{1,1}, \dots, p_{1,T}), \dots, \mathcal{F}(p_{N,1}, \dots, p_{N,T})\right) \in \mathbb{R}^{N \times T \times d_m}

    where XltX_l^t is the output of the ll-th Temporal Transformer Block (TTB). This formulation reduces the temporal transformer token dimension from N×dmN \times d_m to dmd_m, drastically lowering computational complexity (FLOPs) and enabling the model to scale to significantly longer sequence lengths TT while capturing joint-specific motion trajectories.

  3. Knowl 3 — Theoretical Inference Speedup of Seq2seq over Seq2frame Lifting

    equation

    In 2D-to-3D pose lifting, seq2frame methods estimate only the 3D pose of the central frame by processing an input window of tt frames, requiring sliding-window evaluation with heavy overlap to reconstruct a full video. Given a total video length TT, input receptive field t<Tt < T, and sequence padding length δ\delta, seq2frame methods require T(1+2δ)T(1+2\delta) inference passes, whereas a seq2seq method predicting all tt frames simultaneously requires (T+2δ)/t(T + 2\delta)/t inference passes.

    The theoretical inference pass speedup ratio GG of seq2seq over seq2frame is formulated as:

    G=T(1+2δ)T+2δt=T(1+2δ)T+2δ⋅t≈(1+2δ)⋅tG = \frac{T(1+2\delta)}{\frac{T+2\delta}{t}} = \frac{T(1+2\delta)}{T + 2\delta} \cdot t \approx (1+2\delta) \cdot t

    As the input sequence length tt increases, seq2seq eliminates redundant window calculations, yielding inference speeds exceeding 1,000 frames per second.

  4. Knowl 4 — Hierarchically Weighted MPJPE and Temporal Consistency Loss Formulation

    model/method

    MixSTE is trained end-to-end using a composite loss function L\mathcal{L} that balances per-joint position accuracy and temporal sequence smoothness:

    L=Lw+λtLt+λmLm\mathcal{L} = \mathcal{L}_w + \lambda_t \mathcal{L}_t + \lambda_m \mathcal{L}_m

    where Lw\mathcal{L}_w is the Weighted Mean Per-Joint Position Error (WMPJPE), Lt\mathcal{L}_t is the temporal consistency loss (TCLoss), Lm\mathcal{L}_m is the Mean Per-Joint Velocity Error (MPJVE) loss, and λt,λm\lambda_t, \lambda_m are balancing hyperparameters.

    The WMPJPE loss applies body-part-specific importance weights WiW_i based on kinematic flexibility:

    Lw=1Ns∑i=1Ns(Wi⋅1T∑j=1T∥pi,j−gti,j∥22)\mathcal{L}_w = \frac{1}{N^s} \sum_{i=1}^{N^s} \left( W_i \cdot \frac{1}{T} \sum_{j=1}^T \|p_{i,j} - gt_{i,j}\|_2^2 \right)

    where NsN^s is the number of skeleton joints, TT is the sequence length, pi,j∈R3p_{i,j} \in \mathbb{R}^3 is the predicted 3D position of joint ii at frame jj, and gti,j∈R3gt_{i,j} \in \mathbb{R}^3 is the ground-truth 3D position.

    The weighting vector WiW_i is structured hierarchically across four joint groups:

    • Torso joints: W=1.0W = 1.0
    • Head joints: W=1.5W = 1.5
    • Middle limb joints (elbows, knees): W=2.5W = 2.5
    • Terminal limb joints (wrists, feet): W=4.0W = 4.0

    Lt\mathcal{L}_t enforces first-order consistency across consecutive frames, while Lm\mathcal{L}_m directly minimizes the L2L_2 error between predicted and ground-truth inter-frame joint velocities.

  5. Knowl 5 — 3D Human Pose Estimation Performance on Human3.6M

    data/table

    The performance of MixSTE on Human3.6M was evaluated under Protocol 1 (MPJPE in mm, without rigid alignment), Protocol 2 (P-MPJPE in mm, with Procrustes alignment), and MPJVE (velocity error in mm/frame) across 15 action classes using detected 2D poses from Cascaded Pyramid Network (CPN) and High-Resolution Network (HRNet), as well as 2D ground truth (GT).

    Method 2D Input Sequence Length (TT) MPJPE (mm) ↓\downarrow P-MPJPE (mm) ↓\downarrow
    VideoPose3D (Pavllo et al., 2019) CPN 243 46.8 -
    Attention Mesh (Liu et al., 2020) CPN 243 45.1 35.6
    Anatomy-aware (Chen et al., 2021) CPN 243 44.1 -
    PoseFormer (Zheng et al., 2021) CPN 81 44.3 36.5
    MixSTE (Ours) CPN 81 42.4 33.9
    MixSTE (Ours) CPN 243 40.9 32.6
    Motion Guided (Wang et al., 2020) HRNet 96 42.6 32.7
    MixSTE (Ours) HRNet 243 39.8 30.6
    PoseFormer (Zheng et al., 2021) 2D GT 81 31.3 -
    MixSTE (Ours) 2D GT 81 25.9 -
    MixSTE (Ours) 2D GT 243 21.6 -

    Under CPN detections, MixSTE (T=243T=243) improves MPJPE by 7.6% (40.9 mm vs. 44.3 mm) and P-MPJPE by 10.9% (32.6 mm vs. 36.5 mm) compared to PoseFormer (T=81T=81). With ground-truth 2D inputs, MixSTE (T=243T=243) improves MPJPE by 31.0% over PoseFormer (21.6 mm vs. 31.3 mm). MixSTE also achieves the lowest velocity error (MPJVE of 2.3 mm/frame compared to 3.1 mm/frame for PoseFormer).

  6. Knowl 6 — Ablation of MixSTE Architectural Components and Computational Complexity

    data/table

    An ablation study conducted on Human3.6M with CPN 2D detections demonstrates the individual and cumulative impact of the seq2seq setup, alternating spatio-temporal design, joint separation, and the proposed loss formulation on reconstruction accuracy (MPJPE) and computational cost (FLOPs per frame).

    Configuration Seq2seq Alternating Design Joint Separation Custom Loss MPJPE (mm) ↓\downarrow FLOPs (M) ↓\downarrow
    Baseline ✓ 51.7 186,405
    + Alternating ✓ ✓ 45.5 186,405
    + Separation ✓ ✓ ✓ 41.7 645
    MixSTE (Full) ✓ ✓ ✓ ✓ 40.9 645

    Alternating the spatial and temporal transformer blocks improves MPJPE by 6.2 mm (from 51.7 mm to 45.5 mm). Incorporating joint separation further reduces MPJPE by 3.8 mm (to 41.7 mm) while decreasing computational complexity by a factor of 289×\times (from 186,405 M FLOPs to 645 M FLOPs per frame). Adding the weighted position and temporal consistency losses yields an overall 20.9% error reduction (from 51.7 mm to 40.9 mm).

  7. Knowl 7 — Evaluation on MPI-INF-3DHP and HumanEva Benchmarks

    data/table

    MixSTE was evaluated on MPI-INF-3DHP using ground-truth 2D inputs and on HumanEva-I using 2D detector inputs with and without pretraining (fine-tuning) on Human3.6M.

    Dataset / Method Setting / Pretrain PCK (%) ↑\uparrow AUC (%) ↑\uparrow MPJPE (mm) ↓\downarrow
    MPI-INF-3DHP
    Anatomy-aware (Chen et al., 2021) T=243T=243 87.8 53.8 79.1
    PoseFormer (Zheng et al., 2021) T=27T=27 88.6 56.4 77.1
    MixSTE (Ours) T=1T=1 94.2 63.8 57.9
    MixSTE (Ours) T=27T=27 94.4 66.5 54.9
    HumanEva-I Walk Jog Average
    VideoPose3D (Pavllo et al., 2019) T=81T=81, FT 14.0 / 12.5 20.3 / 17.5 18.2
    PoseFormer (Zheng et al., 2021) T=43T=43, FT 14.4 / 10.2 22.7 / 13.4 20.1
    MixSTE (Ours) T=43T=43 (from scratch) 20.3 / 22.4 27.3 / 34.3 28.5
    MixSTE (Ours) T=43T=43, stride=1 16.2 / 14.2 24.6 / 25.8 20.9
    MixSTE (Ours) T=43T=43, FT 12.7 / 10.9 22.6 / 17.0 16.1

    On MPI-INF-3DHP (T=27T=27), MixSTE achieves 94.4% PCK, 66.5% AUC, and 54.9 mm MPJPE, outperforming PoseFormer by 22.2 mm MPJPE. On HumanEva-I with Human3.6M pretraining (FT), MixSTE attains an average MPJPE of 16.1 mm.

  8. Knowl 8 — Ablation of Loss Objectives and Model Hyperparameters

    data/table

    Ablations on Human3.6M analyze the contribution of individual loss terms to position error (MPJPE) and motion smoothness (MPJVE), as well as the impact of encoder depth (dld_l), feature dimension (dmd_m), and sequence length (TT).

    Loss Formulation MPJPE (mm) MPJVE Depth (dld_l) Dim (dmd_m) Length (TT) MPJPE (mm)
    MPJPE Loss 41.7 5.0 4 64 27 54.3
    WMPJPE Loss 41.3 4.6 6 64 27 53.2
    WMPJPE + Motion Loss 41.3 4.3 8 64 27 51.8
    WMPJPE + TCLoss 41.2 3.6 10 64 27 51.1
    WMPJPE + MPJVE Loss 41.2 2.6 8 128 27 47.9
    WMPJPE + T-Loss (Full) 40.9 2.3 8 256 27 46.1
    8 512 27 45.1
    8 640 27 46.0
    8 512 81 42.7
    8 512 128 42.0
    8 512 243 40.9
    8 512 300 41.8

    WMPJPE improves position accuracy over uniform MPJPE (41.3 mm vs. 41.7 mm). Adding TCLoss and MPJVE loss significantly improves smoothness, cutting MPJVE from 4.6 to 2.3 mm/frame. Optimal hyperparameters are depth dl=8d_l=8, channel dimension dm=512d_m=512, and sequence length T=243T=243. Although dl=10d_l=10 achieves slightly lower error (51.1 mm vs 51.8 mm at dm=64d_m=64), dl=8d_l=8 is selected to avoid parameter inflation (33.7M vs. 42.2M parameters).

  9. Knowl 9 — Vulnerability to 2D Keypoint Detector Noise and Missing Detections

    limitation

    MixSTE's 3D pose reconstruction accuracy depends heavily on the reliability of the upstream 2D pose detector. Errors in 2D keypoint estimation—such as occlusions, missing keypoints, or spatial jitter—propagate directly into the spatio-temporal self-attention layers. On small datasets (such as HumanEva-I trained from scratch without fine-tuning), the lack of inductive bias in full transformer encoders leads to suboptimal performance without data sample stride adjustments or pretraining on larger datasets.

Coverage note — None. All key contributions, mathematical formulations, architectural mechanisms, ablation findings, experimental benchmarks, and stated limitations are fully represented.

References

  1. 1.Yujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai, Tat-Jen Cham, Junsong Yuan, and Nadia Magnenat Thalmann. Exploiting spatial-temporal relationships for 3d pose estimation via graph convolutional networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2272–2281, 2019.
  2. 2.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, pages 213–229. Springer, 2020.
  3. 3.Cristian Sminchisescu Catalin Ionescu, Fuxin Li. Latent structured models for human pose estimation. In International Conference on Computer Vision, 2011.
  4. 4.Tianlang Chen, Chen Fang, Xiaohui Shen, Yiheng Zhu, Zhili Chen, and Jiebo Luo. Anatomy-aware 3d human pose estimation with bone-based pose decomposition. IEEE Transactions on Circuits and Systems for Video Technology, 2021.
  5. 5.Yujin Chen, Zhigang Tu, Liuhao Ge, Dejun Zhang, Ruizhi Chen, and Junsong Yuan. So-handnet: Self-organizing network for 3d hand pose estimation with semi-supervised learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6961–6970, 2019.
  6. 6.Yujin Chen, Zhigang Tu, Di Kang, Linchao Bao, Ying Zhang, Xuefei Zhe, Ruizhi Chen, and Junsong Yuan. Model-based 3d hand reconstruction via self-supervised learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10451–10460, 2021.
  7. 7.Yujin Chen, Zhigang Tu, Di Kang, Ruizhi Chen, Linchao Bao, Zhengyou Zhang, and Junsong Yuan. Joint hand-object 3d reconstruction from a single image with cross-branch feature fusion. IEEE Transactions on Image Processing, 30:4008–4021, 2021.
  8. 8.Yilun Chen, Zhicheng Wang, Yuxiang Peng, Zhiqiang Zhang, Gang Yu, and Jian Sun. Cascaded pyramid network for multi-person pose estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7103–7112, 2018.
  9. 9.Hai Ci, Chunyu Wang, Xiaoxuan Ma, and Yizhou Wang. Optimizing network structure for 3d human pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019.
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020.
  11. 11.Jia Gong, Zhipeng Fan, Qiuhong Ke, Hossein Rahmani, and Jun Liu. Meta agent teaming active learning for pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  12. 12.Kehong Gong, Jianfeng Zhang, and Jiashi Feng. Poseaug: A differentiable pose augmentation framework for 3d human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8575–8584, June 2021.
  13. 13.Kaiming He, Georgia Gkioxari, Piotr Doll'ar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017.
  14. 14.Yihui He, Rui Yan, Katerina Fragkiadaki, and Shoou-I Yu. Epipolar transformers. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  15. 15.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  16. 16.Mir Rayat Imtiaz Hossain and James J. Little. Exploiting temporal information for 3d human pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV), September 2018.
  17. 17.Catalin Ionescu, Joao Carreira, and Cristian Sminchisescu. Iterated second-order label sensitive pooling for 3d human pose estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2014.
  18. 18.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7):1325–1339, July 2014.
  19. 19.Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7):1325–1339, jul 2014.
  20. 20.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization, 2014.
  21. 21.Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017.
  22. 22.Shichao Li, Lei Ke, Kevin Pratama, Yu-Wing Tai, Chi-Keung Tang, and Kwang-Ting Cheng. Cascaded deep monocular 3d human pose estimation with evolutionary training data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  23. 23.Wenhao Li, Hong Liu, Runwei Ding, Mengyuan Liu, Pichao Wang, and Wenming Yang. Exploiting temporal contexts with strided transformer for 3d human pose estimation. IEEE Transactions on Multimedia, 2022.
  24. 24.Jiahao Lin and Gim Hee Lee. Trajectory space factorization for deep video-based 3d human pose estimation. arXiv preprint arXiv:1908.08289, 2019.
  25. 25.Kevin Lin, Lijuan Wang, and Zicheng Liu. End-to-end human pose and mesh reconstruction with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1954–1963, 2021.
  26. 26.Mude Lin, Liang Lin, Xiaodan Liang, Keze Wang, and Hui Cheng. Recurrent 3d pose sequence machines. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
  27. 27.Kenkun Liu, Rongqi Ding, Zhiming Zou, Le Wang, and Wei Tang. A comprehensive study of weight sharing in graph networks for 3d human pose estimation. In European Conference on Computer Vision, pages 318–334. Springer, 2020.
  28. 28.Ruixu Liu, Ju Shen, He Wang, Chen Chen, Sen-ching Cheung, and Vijayan Asari. Attention mechanism exploits temporal contexts: Real-time 3d human pose reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5064–5073, 2020.
  29. 29.Xianzheng Ma, Hossein Rahmani, Zhipeng Fan, Bin Yang, Jun Chen, and Jun Liu. Remote: Reinforced motion transformation network for semi-supervised 2d pose estimation in videos. In Proceedings of the AAAI Conference on Artificial Intelligence, 2022.
  30. 30.Xiaoxuan Ma, Jiajun Su, Chunyu Wang, Hai Ci, and Yizhou Wang. Context modeling in 3d human pose estimation: A unified perspective. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6238–6247, 2021.
  31. 31.Julieta Martinez, Rayat Hossain, Javier Romero, and James J Little. A simple yet effective baseline for 3d human pose estimation. In Proceedings of the IEEE international conference on computer vision, pages 2640–2649, 2017.
  32. 32.Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3d human pose estimation in the wild using improved cnn supervision. In 3D Vision (3DV), 2017 Fifth International Conference on. IEEE, 2017.
  33. 33.Dushyant Mehta, Srinath Sridhar, Oleksandr Sotnychenko, Helge Rhodin, Mohammad Shafiei, Hans-Peter Seidel, Weipeng Xu, Dan Casas, and Christian Theobalt. Vnect: Real-time 3d human pose estimation with a single rgb camera. ACM Transactions on Graphics (TOG), 36(4):1–14, 2017.
  34. 34.Alejandro Newell, Kaiyu Yang, and Jia Deng. Stacked hourglass networks for human pose estimation. In European conference on computer vision, pages 483–499. Springer, 2016.
  35. 35.Georgios Pavlakos, Xiaowei Zhou, and Kostas Daniilidis. Ordinal depth supervision for 3d human pose estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018.
  36. 36.Georgios Pavlakos, Xiaowei Zhou, Konstantinos G. Derpanis, and Kostas Daniilidis. Coarse-to-fine volumetric prediction for single-image 3d human pose. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
  37. 37.Dario Pavllo, Christoph Feichtenhofer, David Grangier, and Michael Auli. 3d human pose estimation in video with temporal convolutions and semi-supervised training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7753–7762, 2019.
  38. 38.Varun Ramakrishna, Takeo Kanade, and Yaser Sheikh. Reconstructing 3D Human Pose from 2D Image Landmarks. In Andrew Fitzgibbon, Svetlana Lazebnik, Pietro Perona, Yoichi Sato, and Cordelia Schmid, editors, Computer Vision – ECCV 2012, Lecture Notes in Computer Science, pages 573–586, Berlin, Heidelberg, 2012. Springer.
  39. 39.Varun Ramakrishna, Takeo Kanade, and Yaser Sheikh. Reconstructing 3d human pose from 2d image landmarks. In European conference on computer vision, pages 573–586. Springer, 2012.
  40. 40.Leonid Sigal, Alexandru O Balan, and Michael J Black. Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. International journal of computer vision, 87(1-2):4, 2010.
  41. 41.Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. Deep high-resolution representation learning for human pose estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5693–5703, 2019.
  42. 42.Xiao Sun, Bin Xiao, Fangyin Wei, Shuang Liang, and Yichen Wei. Integral human pose regression. In Proceedings of the European Conference on Computer Vision (ECCV), pages 529–545, 2018.
  43. 43.Mikael Svenstrup, Soren Tranberg, Hans Jorgen Andersen, and Thomas Bak. Pose estimation and adaptive robot behaviour for human-robot interaction. In 2009 IEEE International Conference on Robotics and Automation, pages 3571–3576, 2009.
  44. 44.Bugra Tekin, Artem Rozantsev, Vincent Lepetit, and Pascal Fua. Direct prediction of 3d body poses from motion compensated sequences. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 991–1000, 2016.
  45. 45.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  46. 46.Jingbo Wang, Sijie Yan, Yuanjun Xiong, and Dahua Lin. Motion guided 3d pose estimation from videos. In European Conference on Computer Vision, pages 764–780. Springer, 2020.
  47. 47.Tom Wehrbein, Marco Rudolph, Bodo Rosenhahn, and Bastian Wandt. Probabilistic monocular 3d human pose estimation with normalizing flows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 11199–11208, October 2021.
  48. 48.Tianhan Xu and Wataru Takano. Graph stacked hourglass networks for 3d human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16105–16114, 2021.
  49. 49.Sen Yang, Zhibin Quan, Mu Nie, and Wankou Yang. Transpose: Towards explainable human pose estimation by transformer. arXiv preprint arXiv:2012.14214, 2020.
  50. 50.Mang Ye, He Li, Bo Du, Jianbing Shen, Ling Shao, and Steven C. H. Hoi. Collaborative refining for person re-identification with label noise. IEEE Transactions on Image Processing, 31:379–391, 2022.
  51. 51.Raymond Yeh, Yuan-Ting Hu, and Alexander Schwing. Chirality nets for human pose regression. Advances in Neural Information Processing Systems, 32:8163–8173, 2019.
  52. 52.Jae Shin Yoon, Lingjie Liu, Vladislav Golyanik, Kripasindhu Sarkar, Hyun Soo Park, and Christian Theobalt. Pose-guided human animation from a single image in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15039–15048, June 2021.
  53. 53.Ailing Zeng, Xiao Sun, Lei Yang, Nanxuan Zhao, Minhao Liu, and Qiang Xu. Learning skeletal graph neural networks for hard 3d pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 11436–11445, October 2021.
  54. 54.Can Zhang, Tianyu Yang, Junwu Weng, Meng Cao, Jue Wang, and Zou Yuexian. Unsupervised pre-training for temporal action localization tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  55. 55.Jiaxu Zhang, Gaoxiang Ye, Zhigang Tu, Yongtao Qin, Jinlu Zhang, Xiangjian Liu, and Shixu Luo. A spatial attentive and temporal dilated (satd) gcn for skeleton-based action recognition. CAAI Transactions on Intelligence Technology, 2020.
  56. 56.Long Zhao, Xi Peng, Yu Tian, Mubbasir Kapadia, and Dimitris N. Metaxas. Semantic graph convolutional networks for 3d human pose regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019.
  57. 57.Ce Zheng, Sijie Zhu, Matias Mendieta, Taojiannan Yang, Chen Chen, and Zhengming Ding. 3d human pose estimation with spatial and temporal transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 11656–11665, October 2021.
  58. 58.Xingyi Zhou, Qixing Huang, Xiao Sun, Xiangyang Xue, and Yichen Wei. Towards 3d human pose estimation in the wild: a weakly-supervised approach. In Proceedings of the IEEE International Conference on Computer Vision, pages 398–407, 2017.

Citation

MLA
Zhang, J., et al. “MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video”. arXiv, 2022, http://arxiv.org/abs/2203.00859v4.
APA
Zhang, J., Tu, Z., Yang, J., Chen, Y., & Yuan, J. (2022). MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video. arXiv. http://arxiv.org/abs/2203.00859v4
Chicago
Zhang, J., Z. Tu, J. Yang, Y. Chen, and J. Yuan. 2022. “MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video”. arXiv. http://arxiv.org/abs/2203.00859v4.
Harvard
Zhang, J. et al. (2022) “MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.00859v4.
Vancouver
1. Zhang J, Tu Z, Yang J, Chen Y, Yuan J (2022) MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video. arXiv

BibTeX

@article{zhang2022mixste,
  title = {MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video},
  author = {Zhang, Jinlu and Tu, Zhigang and Yang, Jianyu and Chen, Yujin and Yuan, Junsong},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.00859v4},
  eprint = {2203.00859}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE