MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video
Jinlu ZhangZhigang TuJianyu YangYujin ChenJunsong Yuan
Proposes MixSTE, a sequence-to-sequence transformer architecture that alternates between modeling individual joint trajectories over time and spatial joint dependencies within frames to achieve state-of-the-art accuracy in video-based 3D human pose estimation.
Estimating three-dimensional human body poses from standard two-dimensional video is a fundamental computer vision capability with significant value for robotics, action recognition, and virtual human systems. However, extracting accurate 3D positions from monocular video remains challenging because different 3D poses can correspond to the identical 2D projection. While recent approaches have leveraged attention-based transformer models across video sequences, they typically fail to capture the distinct, independent motion trajectories of individual body joints over time, and they commonly predict only a single central frame per sequence, which introduces substantial computational redundancy.
The article demonstrates and evaluates MixSTE, a sequence-to-sequence neural network architecture designed to reconstruct full 3D pose sequences from 2D keypoints. MixSTE alternately stacks spatial attention blocks to capture relationships across body joints within each frame and temporal attention blocks to separately track the movement path of each individual joint across time.
To establish credibility and performance benchmarks, the authors evaluated MixSTE across three widely recognized experimental datasets: Human3.6M, MPI-INF-3DHP, and HumanEva. The model was evaluated using standard 2D pose detection inputs as well as ground-truth 2D keypoints, testing both short- and long-term sequence lengths ranging from 1 to 243 frames.
The findings confirm that MixSTE establishes new state-of-the-art performance across all benchmarked datasets. On the Human3.6M benchmark using standard 2D inputs, the model reduces mean joint position error to 40.9 mm, achieving a 7.6% error reduction over prior leading transformer methods, while yielding a 10.9% improvement under rigid alignment metrics. When provided with ground-truth 2D keypoints, MixSTE achieves a 31.0% error reduction over previous transformer baselines. Ablation experiments demonstrated that treating each joint as a separate temporal token drastically cut computational operations per frame from over 186,000 to 645 megaflops, while the full sequence-to-sequence output significantly increased inference processing speed.
These results show that separating joint trajectories and enforcing end-to-end sequence coherence substantially boosts accuracy while resolving the high latency and computational inefficiencies common in prior methods. For enterprise applications, these improvements reduce hardware infrastructure costs and enable practical deployment in real-time or resource-constrained environments where rapid video processing is critical.
Stakeholders deploying 3D human pose systems should consider adopting sequence-to-sequence alternating spatio-temporal architectures to maximize throughput and positional precision. For applications involving small datasets, fine-tuning pretrained models from larger datasets is recommended to ensure strong generalization. Future engineering and research efforts should prioritize modeling input noise distributions and pairing the architecture with robust upstream 2D detectors to mitigate performance degradation caused by missing or inaccurate 2D keypoints.
- Paper: A Simple Yet Effective Baseline for 3d Human Pose Estimation, Julieta Martinez et al. (2017). This seminal work establishes the decoupled 2D-to-3D pose lifting paradigm and standard baseline benchmark on Human3.6M that MixSTE adopts and extends across video sequences.
- Paper: Is Space-Time Attention All You Need for Video Understanding?, Gedas Bertasius et al. (2021). It introduces factorized divided space-time attention in transformers, providing the foundational architectural principle for MixSTE's alternating spatial and temporal transformer blocks.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). It explores factorized spatial and temporal attention mechanisms in vision transformers, directly informing the spatio-temporal factorization strategy utilized by MixSTE.
- Paper: Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition, Sijie Yan et al. (2018). It formalizes the representation of human skeletal joint sequences across spatial and temporal dimensions, which MixSTE reimagines using transformer attention instead of graph convolutions.
- Paper: Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments, Catalin Ionescu et al. (2014). It presents the Human3.6M dataset and standard evaluation protocols used as the primary benchmark for 3D pose sequence estimation in MixSTE.
- Paper: HumanEva: Synchronized Video and Motion Capture Dataset and Baseline Algorithm for Evaluation of Articulated Human Motion, L. Sigal et al. (2010). It establishes the standardized multi-view video and motion capture benchmark HumanEva, which serves as one of the core evaluation datasets in MixSTE.
- Paper: Monocular 3D Human Pose Estimation in the Wild Using Improved CNN Supervision, Dushyant Mehta et al. (2016). It introduces the MPI-INF-3DHP benchmark dataset, providing another essential 3D human pose evaluation benchmark used to test MixSTE.
No sufficiently relevant recommendations were found.
