Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences
Yujie ZhouHaodong DuanAnyi RaoBing SuJiaqi Wang
Proposes a negative-sample-free self-supervised skeleton representation framework that uses centrality-based spatial masking and motion-weighted temporal masking to capture local spatio-temporal dependencies and maintain high recognition accuracy even when joints are missing.
Human action recognition based on 3D skeleton data is increasingly valuable in surveillance, robotics, and interactive technologies because it remains robust across changing backgrounds and appearances. However, building supervised models requires expensive and time-consuming manual data annotation. Existing self-supervised methods that learn without labels rely heavily on global data alterations and contrastive pairs, often ignoring fine-grained spatial and temporal relationships across body joints and video frames. Consequently, conventional models experience severe accuracy drops in real-world scenarios when cameras face visual obstructions or miss skeletal joints.
The article demonstrates and evaluates a self-supervised framework called Partial Spatio-Temporal Learning (PSTL), which trains models to recognize actions effectively even when significant portions of spatial joints or video frames are missing. The approach aims to improve recognition accuracy and maintain high reliability when deployed in occluded environments.
To achieve this, the authors developed a three-stream neural network structure that does not require negative sample pairs or large memory banks. The system compares a complete skeleton video stream against two partially masked streams: a spatial stream where structurally central joints are systematically removed from calculations, and a temporal stream where high-motion keyframes are selectively omitted. The model minimizes feature redundancy by aligning representations across these streams. The authors evaluated the framework across three standard benchmark datasets containing tens of thousands of video sequences across varied subjects, camera setups, and challenging noise conditions.
The experimental findings show clear performance advantages. First, PSTL consistently surpassed existing leading self-supervised methods across linear, semi-supervised, and fine-tuning benchmarks. Second, the framework showed exceptional robustness in partial-body evaluations simulating camera occlusions: when ten joints were missing, competing models suffered performance drops of approximately 17% to 20%, whereas PSTL experienced only a 3.5% decrease. Third, in data-constrained semi-supervised settings using only 1% or 10% of labeled samples, PSTL significantly outperformed prior approaches, demonstrating up to an 8.6% accuracy gain on noisy benchmark splits.
These findings indicate that incorporating localized spatio-temporal masking enables models to learn generalized, occlusion-resilient motion patterns without requiring costly manual annotations or memory-intensive contrastive training. Organizations deploying vision systems can reduce annotation expenditures, lower hardware overhead during pre-training, and improve operational reliability in visually degraded or crowded environments where subjects are frequently blocked from view.
Decision-makers should consider adopting negative-sample-free masking frameworks when developing vision pipelines for occluded real-world conditions or when working with constrained labeled data. Next steps include testing this masking approach across other skeleton-based architectures and running pilot deployments on live camera feeds. Because the evaluation relied primarily on benchmark datasets with fixed frame lengths and predefined joint topologies, practitioners should validate the model's performance on continuous, unsegmented video streams before full-scale deployment.
- Paper: Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition, Sijie Yan et al. (2018). Introduces Spatial Temporal Graph Convolutional Networks (ST-GCN), the foundational backbone architecture for modeling spatial and temporal skeleton dynamics used in skeleton-based representation learning.
- Paper: Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition, Lei Shi et al. (2018). Presents adaptive graph convolutional networks that dynamically model non-local joint relationships, establishing key principles for joint interaction modeling built upon in masked spatio-temporal learning.
- Paper: NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis, Amir Shahroudy et al. (2016). Introduces the NTU RGB+D benchmark and foundational cross-subject/cross-view evaluation protocols widely utilized to assess PSTL's downstream performance.
- Paper: NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding, Jun Liu et al. (2019). Establishes the expanded NTU RGB+D 120 benchmark dataset used in the source paper to validate self-supervised skeleton representations.
- Paper: Self-Supervised Visual Feature Learning With Deep Neural Networks: A Survey, Longlong Jing et al. (2019). Surveys self-supervised visual feature learning and pretext task formulations, providing the broader conceptual foundation for mask-based and contrastive representation learning.
- Paper: BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition, Yuxuan Zhou et al. (2024). Extends skeleton topology learning by explicitly preserving physical bone connectivity and decoupling multi-relational feature projections in Graph Convolutional Networks.
