Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences

Yujie ZhouHaodong DuanAnyi RaoBing SuJiaqi Wang

article2023AAAI57 citations

Proposes a negative-sample-free self-supervised skeleton representation framework that uses centrality-based spatial masking and motion-weighted temporal masking to capture local spatio-temporal dependencies and maintain high recognition accuracy even when joints are missing.

Listen

Human action recognition based on 3D skeleton data is increasingly valuable in surveillance, robotics, and interactive technologies because it remains robust across changing backgrounds and appearances. However, building supervised models requires expensive and time-consuming manual data annotation. Existing self-supervised methods that learn without labels rely heavily on global data alterations and contrastive pairs, often ignoring fine-grained spatial and temporal relationships across body joints and video frames. Consequently, conventional models experience severe accuracy drops in real-world scenarios when cameras face visual obstructions or miss skeletal joints.

The article demonstrates and evaluates a self-supervised framework called Partial Spatio-Temporal Learning (PSTL), which trains models to recognize actions effectively even when significant portions of spatial joints or video frames are missing. The approach aims to improve recognition accuracy and maintain high reliability when deployed in occluded environments.

To achieve this, the authors developed a three-stream neural network structure that does not require negative sample pairs or large memory banks. The system compares a complete skeleton video stream against two partially masked streams: a spatial stream where structurally central joints are systematically removed from calculations, and a temporal stream where high-motion keyframes are selectively omitted. The model minimizes feature redundancy by aligning representations across these streams. The authors evaluated the framework across three standard benchmark datasets containing tens of thousands of video sequences across varied subjects, camera setups, and challenging noise conditions.

The experimental findings show clear performance advantages. First, PSTL consistently surpassed existing leading self-supervised methods across linear, semi-supervised, and fine-tuning benchmarks. Second, the framework showed exceptional robustness in partial-body evaluations simulating camera occlusions: when ten joints were missing, competing models suffered performance drops of approximately 17% to 20%, whereas PSTL experienced only a 3.5% decrease. Third, in data-constrained semi-supervised settings using only 1% or 10% of labeled samples, PSTL significantly outperformed prior approaches, demonstrating up to an 8.6% accuracy gain on noisy benchmark splits.

These findings indicate that incorporating localized spatio-temporal masking enables models to learn generalized, occlusion-resilient motion patterns without requiring costly manual annotations or memory-intensive contrastive training. Organizations deploying vision systems can reduce annotation expenditures, lower hardware overhead during pre-training, and improve operational reliability in visually degraded or crowded environments where subjects are frequently blocked from view.

Decision-makers should consider adopting negative-sample-free masking frameworks when developing vision pipelines for occluded real-world conditions or when working with constrained labeled data. Next steps include testing this masking approach across other skeleton-based architectures and running pilot deployments on live camera feeds. Because the evaluation relied primarily on benchmark datasets with fixed frame lengths and predefined joint topologies, practitioners should validate the model's performance on continuous, unsegmented video streams before full-scale deployment.

Cover for Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences

Abstract

Self-supervised learning has demonstrated remarkable capability in representation learning for skeleton-based action recognition. Existing methods mainly focus on applying global data augmentation to generate different views of the skeleton sequence for contrastive learning. However, due to the rich action clues in the skeleton sequences, existing methods may only take a global perspective to learn to discriminate different skeletons without thoroughly leveraging the local relationship between different skeleton joints and video frames, which is essential for real-world applications. In this work, we propose a Partial Spatio-Temporal Learning (PSTL) framework to exploit the local relationship from a partial skeleton sequences built by a unique spatio-temporal masking strategy. Specifically, we construct a negative-sample-free triplet steam structure that is composed of an anchor stream without any masking, a spatial masking stream with Central Spatial Masking (CSM), and a temporal masking stream with Motion Attention Temporal Masking (MATM). The feature cross-correlation matrix is measured between the anchor stream and the other two masking streams, respectively. (1) Central Spatial Masking discards selected joints from the feature calculation process, where the joints with a higher degree of centrality have a higher possibility of being selected. (2) Motion Attention Temporal Masking leverages the motion of action and remove frames that move faster with a higher possibility. Our method achieves state-of-the-art performance on NTURGB+D 60, NTURGB+D 120 and PKU-MMD under various downstream tasks. Furthermore, to simulate the real-world scenarios, a practical evaluation is performed where some skeleton joints are lost in downstream tasks. In contrast to previous methods that suffer from large performance drops, our PSTL can still achieve remarkable results under this challenging setting, validating the robustness of our method. Our code is available at https://github.com/YujieOuO/PSTL.git.

Table of Contents

  • Introduction
  • Related Work
  • Method
  • Skeleton Barlow Twins
  • Central Spatial Masking
  • Motion Attention Temporal Masking
  • Loss Function
  • Experiments
  • Datasets
  • Implementation Details
  • Experiment Setting
  • Ablation Studies of PSTL
  • Comparison with State-of-the-art
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Partial Spatio-Temporal Learning (PSTL) Framework

    model/method

    Partial Spatio-Temporal Learning (PSTL) is a negative-sample-free, triplet-stream self-supervised learning framework designed for 3D skeleton-based action representation learning. Given an input skeleton sequence s∈RC×T×Vs \in \mathbb{R}^{C \times T \times V} (where CC is the 3D coordinate channel dimension, TT is the number of frames, and VV is the number of skeleton joints), the framework processes three parallel streams:

    1. Anchor Stream: An ordinary data augmentation is applied to generate view xx. A feature encoder f(⋅;θ)f(\cdot; \theta) extracts representation h=f(x;θ)∈Rchh = f(x; \theta) \in \mathbb{R}^{c_h}, which is projected by a multi-layer projector g(⋅)g(\cdot) into embedding z=g(h)∈Rczz = g(h) \in \mathbb{R}^{c_z}. This stream undergoes no masking to retain full semantic information.

    2. Spatial Masking Stream: An ordinary data augmentation is applied to generate view x′x', which is then processed by Central Spatial Masking (CSM) to produce partial skeleton sequence xc′x'_c. The shared encoder and projector yield feature hc′=f(xc′;θ)h'_c = f(x'_c; \theta) and embedding zc′=g(hc′)z'_c = g(h'_c).

    3. Temporal Masking Stream: An ordinary data augmentation is applied to generate view x^\hat{x}, which is processed by Motion Attention Temporal Masking (MATM) to yield sequence x^m\hat{x}_m. The shared encoder and projector yield feature h^m=f(x^m;θ)\hat{h}_m = f(\hat{x}_m; \theta) and embedding z^m=g(h^m)\hat{z}_m = g(\hat{h}_m).

    The framework avoids negative sample pairs and memory banks by optimizing empirical cross-correlation matrices between the anchor embedding zz and each of the masked embeddings zc′z'_c and z^m\hat{z}_m towards identity matrices, simultaneously promoting invariance and minimizing feature redundancy.

  2. Knowl 2 — Central Spatial Masking (CSM) Strategy

    model/method

    Central Spatial Masking (CSM) constructs partial skeleton inputs during pre-training by removing specific joints from the feature calculation. Rather than setting joint coordinates to zero (which introduces misleading 3D position information), CSM removes the selected joints from the sequence and discards their corresponding rows and columns in the graph adjacency matrix of the Spatial-Temporal Graph Convolutional Network (ST-GCN) backbone, ensuring that representations are derived exclusively from unmasked joints.

    Joint selection is governed by degree centrality on the predefined skeleton graph topology. Let ViV_i for i∈{1,…,n}i \in \{1, \dots, n\} denote the skeleton joints, and let did_i denote the graph degree of joint ViV_i. The probability pip_i of masking joint ViV_i is:

    pi=di∑j=1ndjp_i = \frac{d_i}{\sum_{j=1}^n d_j}

    Joints with higher connectivity (such as the upper and lower torso centers with degrees 3 and 4) have higher probabilities of being masked than peripheral joints (with degree 1), forcing the encoder to learn relationships across wider spatial spans. In standard pre-training, 9 joints are masked per sample.

  3. Knowl 3 — Motion Attention Temporal Masking (MATM) Strategy

    model/method

    Motion Attention Temporal Masking (MATM) identifies and masks key semantic frames in skeleton sequences based on motion magnitude. Given an augmented skeleton sequence x^∈RC×T×V\hat{x} \in \mathbb{R}^{C \times T \times V}, temporal displacement between adjacent frames is computed as:

    m:,t,:=x^:,t+1,:−x^:,t,:m_{:,t,:} = \hat{x}_{:,t+1,:} - \hat{x}_{:,t,:}

    for t∈{1,…,T−1}t \in \{1, \dots, T-1\}. The overall motion rate mtm_t of frame tt defines its temporal attention weight ata_t:

    at=mt2∑i=1Tmi2a_t = \frac{m_t^2}{\sum_{i=1}^T m_i^2}

    The top-KK frames with highest attention weights {at1,…,atK}\{a_{t_1}, \dots, a_{t_K}\} are identified as key frames and masked. To preserve sequence diversity, an additional KK frames are randomly selected and masked from the remaining frames in the sequence, producing the final temporally masked sequence x^m\hat{x}_m. In practice, KK is set to 10.

  4. Knowl 4 — PSTL Cross-Correlation Loss Objective

    equation

    For batch embeddings z,z′∈RB×czz, z' \in \mathbb{R}^{B \times c_z} (with batch size BB and embedding dimension czc_z), the empirical cross-correlation matrix C∈Rcz×czC \in \mathbb{R}^{c_z \times c_z} along the batch dimension bb is given by:

    Cij=∑bzb,izb,j′∑b(zb,i)2∑b(zb,j′)2C_{ij} = \frac{\sum_b z_{b,i} z'_{b,j}}{\sqrt{\sum_b (z_{b,i})^2} \sqrt{\sum_b (z'_{b,j})^2}}

    where i,j∈{1,…,cz}i, j \in \{1, \dots, c_z\} index embedding vector dimensions.

    The overall pre-training loss Lp\mathcal{L}_p of Partial Spatio-Temporal Learning (PSTL) balances the spatial masking stream loss L1\mathcal{L}_1 (between anchor embedding zz and spatial-masked embedding zc′z'_c) and the temporal masking stream loss L2\mathcal{L}_2 (between anchor embedding zz and temporal-masked embedding z^m\hat{z}_m):

    Lp=L1+L2\mathcal{L}_p = \mathcal{L}_1 + \mathcal{L}_2

    where

    L1=∑i(1−Cii′)2+λ∑i∑j≠i(Cij′)2\mathcal{L}_1 = \sum_i (1 - C'_{ii})^2 + \lambda \sum_i \sum_{j \neq i} (C'_{ij})^2

    L2=∑i(1−C^ii)2+λ∑i∑j≠i(C^ij)2\mathcal{L}_2 = \sum_i (1 - \hat{C}_{ii})^2 + \lambda \sum_i \sum_{j \neq i} (\hat{C}_{ij})^2

    Here C′C' is the cross-correlation matrix between zz and zc′z'_c, C^\hat{C} is the cross-correlation matrix between zz and z^m\hat{z}_m, and λ=2×10−4\lambda = 2 \times 10^{-4} is a trade-off parameter balancing the invariance term (diagonal elements forced to 1) and the redundancy reduction term (off-diagonal elements forced to 0).

  5. Knowl 5 — Skeleton Data Augmentation Suite

    model/method

    Prior to spatial or temporal masking, skeleton views are generated using four ordinary transformations:

    1. Shear: A linear transformation applying shear matrix SS to 3D joint coordinates:

    S=[1s12s13s211s23s31s321]S = \begin{bmatrix} 1 & s_{12} & s_{13} \\ s_{21} & 1 & s_{23} \\ s_{31} & s_{32} & 1 \end{bmatrix}

    where shear factors sij∼U[−β,β]s_{ij} \sim \mathcal{U}[-\beta, \beta] with amplitude β=1.0\beta = 1.0.

    1. Crop: The skeleton sequence of length TT is padded by γT\gamma T frames (padding ratio γ=1/6\gamma = 1/6) and randomly cropped back to the standardized sequence length T=50T = 50.

    2. Rotate: A main axis M∈{X,Y,Z}M \in \{X, Y, Z\} is randomly selected and rotated by an angle sampled uniformly from [0,π/6][0, \pi/6], while the remaining two axes are rotated by angles sampled uniformly from [0,π/180][0, \pi/180].

    3. Spatial Flip: Swaps corresponding left and right skeleton joints across the body symmetry axis with probability p=0.5p = 0.5.

  6. Knowl 6 — Partial Body Evaluation Protocol for Occlusion Robustness

    definition

    Partial Body Evaluation is an evaluation protocol designed to test representation robustness against joint occlusions and shading during downstream action recognition tasks:

    1. Joint shaded: A set number of skeleton joints (from 0 to 10 out of 25 joints) are randomly dropped during testing.

    2. Body part shaded: The 25 skeleton joints are grouped into 5 anatomical body parts (left arm, right arm, left leg, right leg, torso). During evaluation, a subset of these parts (from 0 to 4 parts) is randomly removed.

    A linear classifier is trained on representations from the frozen pre-trained encoder and tested under these missing-joint conditions without retraining the encoder.

  7. Knowl 7 — Linear Evaluation Performance under Joint and Body Part Shading

    data/table

    Top-1 linear evaluation accuracy (%) on the NTU-RGB+D 60 Cross-Subject (xsub) split under increasing levels of joint occlusion (0 to 10 masked joints) and body part occlusion (0 to 4 masked parts):

    # of masked joints 0 2 4 6 8 10
    SkeletonCLR 68.3 65.6 61.9 58.6 54.5 50.7
    SkeletonBT 68.5 65.0 61.7 58.6 55.0 50.6
    AimCLR 74.3 70.4 67.2 63.0 58.5 54.3
    PSTL (ours) 77.3 76.9 76.2 75.7 74.6 73.8
    # of masked body parts 0 1 2 3 4 –
    SkeletonCLR 68.3 63.6 57.7 47.5 33.5 –
    SkeletonBT 68.5 64.0 58.5 50.2 38.4 –
    AimCLR 74.3 69.3 63.9 53.7 37.9 –
    PSTL (ours) 77.3 73.6 68.3 60.0 47.3 –

    When 10 joints are masked, PSTL performance drops by only 3.5% (from 77.3% to 73.8%), whereas AimCLR, SkeletonBT, and SkeletonCLR drop by 20.0%, 17.9%, and 17.6% respectively. Under 4 body parts masked, PSTL maintains 47.3% accuracy, outperforming AimCLR (37.9%) by +9.4% and SkeletonCLR (33.5%) by +13.8%.

  8. Knowl 8 — Linear Evaluation Comparison on Action Recognition Benchmarks

    data/table

    Top-1 accuracy (%) for linear evaluation on NTU-RGB+D 60 (xsub, xview), NTU-RGB+D 120 (xsub, xset), and PKU-MMD (Part I, Part II). Modalities include joint (J), motion (M), bone (B), and three-stream fusion (3s=J+M+B3s = \text{J}+\text{M}+\text{B}):

    Method NTU-60 (%) NTU-120 (%) PKU-MMD (%)
    xsub xview xsub xset Part I Part II
    Single-stream:
    LongT GAN 39.1 48.1 – – 67.7 26.0
    MS2L 52.6 – – – 64.9 27.6
    AS-CAL 58.5 64.8 48.6 49.2 – –
    PC 50.7 76.3 42.7 41.7 – –
    SkeletonCLR (J) 68.3 – – – 80.9 35.2
    ISC – – – – 80.9 36.0
    AimCLR (J) 74.3 79.7 63.4 63.4 83.4 36.8
    SkeletonBT (J) 68.5 71.6 53.9 52.7 84.7 38.2
    PSTL (J) 77.3 81.8 66.2 67.7 88.4 49.3
    Three-stream (J+M+B):
    3s-SkeletonCLR 75.0 79.8 60.7 62.6 – –
    3s-Colorization 75.2 83.1 – – – –
    3s-CrosSCLR 77.8 83.4 67.9 67.1 84.9 21.2
    3s-AimCLR 78.9 83.8 68.2 68.8 87.8 38.5
    3s-SkeletonBT 69.3 72.2 56.3 56.8 84.8 44.1
    3s-PSTL 79.1 83.8 69.2 70.3 89.2 52.3

    Single-stream PSTL (J) improves over AimCLR (J) by +3.0% (xsub) and +2.1% (xview) on NTU-60, +2.8% (xsub) and +4.3% (xset) on NTU-120, and +5.0% (Part I) and +12.5% (Part II) on PKU-MMD. Three-stream PSTL reaches 79.1% on NTU-60 xsub, 69.2% / 70.3% on NTU-120, and 52.3% on PKU-MMD Part II.

  9. Knowl 9 — Fine-Tuning and Semi-Supervised Performance of PSTL

    data/table

    Top-1 accuracy (%) for full fine-tuning on NTU-RGB+D 60 and NTU-RGB+D 120 (bone stream or three-stream fusion) and semi-supervised fine-tuning on PKU-MMD (1% and 10% labeled data):

    Method (Fine-tuning) NTU-60 (%) NTU-120 (%)
    xsub xview xsub xset
    SkeletonCLR (B) 82.2 88.9 73.6 75.3
    AimCLR (B) 83.0 89.2 77.2 76.0
    PSTL (B) 84.5 92.0 78.6 78.9
    3s-ST-GCN (Supervised) 85.2 91.4 77.2 77.1
    3s-CrosSCLR 86.2 92.5 80.5 80.4
    3s-AimCLR 86.9 92.8 80.1 80.9
    3s-PSTL 87.1 93.9 81.3 82.6
    Method (Semi-supervised PKU-MMD) Part I (%) Part II (%)
    1% labeled data:
    LongT GAN 35.8 12.4
    MS2L 36.4 13.0
    ISC 37.7 –
    3s-CrosSCLR 49.7 10.2
    3s-AimCLR 57.5 15.1
    3s-PSTL 62.5 16.9
    10% labeled data:
    LongT GAN 69.5 25.7
    MS2L 70.3 26.1
    ISC 72.1 –
    3s-CrosSCLR 82.9 28.6
    3s-AimCLR 86.1 33.4
    3s-PSTL 86.9 42.0

    Single-stream bone PSTL (78.6% xsub, 78.9% xset on NTU-120) outperforms the fully supervised 3-stream ST-GCN baseline (77.2% xsub, 77.1% xset). Under semi-supervised fine-tuning with 10% labels on PKU-MMD Part II, 3s-PSTL achieves 42.0%, outperforming 3s-AimCLR (33.4%) by +8.6%.

  10. Knowl 10 — Ablation of CSM, MATM, and Random Masking Strategies

    empirical result

    Linear evaluation accuracy on NTU-RGB+D 60 under Cross-Subject (xsub) and Cross-View (xview) shows the individual and joint contributions of Central Spatial Masking (CSM) and Motion Attention Temporal Masking (MATM) relative to baseline SkeletonBT and random masking strategies:

    1. Component Ablation on SkeletonBT:

      • Baseline SkeletonBT (no masking): 68.5% (xsub), 71.6% (xview).
      • SkeletonBT + CSM: 71.8% (xsub, +3.3%), 75.4% (xview, +3.8%).
      • SkeletonBT + MATM: 73.7% (xsub, +5.2%), 78.0% (xview, +6.4%).
      • Full PSTL (SkeletonBT + CSM + MATM): 77.3% (xsub, +8.8%), 81.8% (xview, +10.2%).
    2. Masking Policy Comparison:

      • Random Spatial Masking (RSM) + Random Temporal Masking (RTM): 75.6% (xsub), 79.7% (xview).
      • RSM + MATM: 76.4% (xsub), 81.2% (xview).
      • CSM + RTM: 76.1% (xsub), 80.4% (xview).
      • CSM + MATM: 77.3% (xsub), 81.8% (xview).

    Both degree-centrality spatial masking and motion-attention temporal masking provide consistent performance gains over uniform random masking (+1.7% on xsub and +2.1% on xview over RSM + RTM).

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Cao, Z.; Simon, T.; Wei, S.-E.; and Sheikh, Y. 2017. Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  2. 2.Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning, 1597–1607. PMLR.
  3. 3.Chen, X.; and He, K. 2021. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 15750–15758.
  4. 4.Chen, Z.; Li, S.; Yang, B.; Li, Q.; and Liu, H. 2021. Multi-scale spatial temporal graph convolutional network for skeleton-based action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, 1113–1122.
  5. 5.Du, Y.; Wang, W.; and Wang, L. 2015. Hierarchical recurrent neural network for skeleton based action recognition. IEEE Conference on Computer Vision and Pattern Recognition.
  6. 6.Gidaris, S.; Singh, P.; and Komodakis, N. 2018. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728.
  7. 7.Grill, J.-B.; Strub, F.; Altché, F.; Tallec, C.; Richemond, P.; Buchatskaya, E.; Doersch, C.; Avila Pires, B.; Guo, Z.; Gheshlaghi Azar, M.; et al. 2020. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems, 33: 21271–21284.
  8. 8.Guo, T.; Liu, H.; Chen, Z.; Liu, M.; Wang, T.; and Ding, R. 2022. Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-supervised Action Recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 762–770.
  9. 9.He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2022. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16000–16009.
  10. 10.He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 9729–9738.
  11. 11.Hua, G.; Liu, H.; Li, W.; Zhang, Q.; Ding, R.; and Xu, X. 2022. Weakly-supervised 3D Human Pose Estimation with Cross-view U-shaped Graph Convolutional Network. IEEE Transactions on Multimedia.
  12. 12.Ke, Q.; Bennamoun, M.; An, S.; Sohel, A. F.; and Boussaïd, F. 2017. A New Representation of Skeleton Sequences for 3D Action Recognition. CVPR.
  13. 13.Kim, S. T.; and Reiter, A. 2017. Interpretable 3D Human Action Analysis with Temporal Convolutional Networks. IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, 1623–1631.
  14. 14.Larsson, G.; Maire, M.; and Shakhnarovich, G. 2016. Learning representations for automatic colorization. In European conference on computer vision, 577–593. Springer.
  15. 15.Li, L.; Wang, M.; Ni, B.; Wang, H.; Yang, J.; and Zhang, W. 2021. 3D Human Action Representation Learning via Cross-View Consistency Pursuit. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4741–4750.
  16. 16.Lin, L.; Song, S.; Yang, W.; and Liu, J. 2020. Ms2l: Multi-task self-supervised learning for skeleton based action recognition. In Proceedings of the 28th ACM International Conference on Multimedia, 2490–2498.
  17. 17.Liu, J.; Shahroudy, A.; Perez, M.; Wang, G.; Duan, L.-Y.; and Kot, A. C. 2019. Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding. IEEE transactions on pattern analysis and machine intelligence, 42(10): 2684–2701.
  18. 18.Liu, J.; Song, S.; Liu, C.; Li, Y.; and Hu, Y. 2020. A benchmark dataset and comparison study for multi-modal human action analytics. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 16(2): 1–24.
  19. 19.Noroozi, M.; and Favaro, P. 2016. Unsupervised learning of visual representations by solving jigsaw puzzles. In European conference on computer vision, 69–84. Springer.
  20. 20.Rao, H.; Xu, S.; Hu, X.; Cheng, J.; and Hu, B. 2021. Augmented skeleton based contrastive action learning with momentum lstm for unsupervised action recognition. Information Sciences, 569: 90–109.
  21. 21.Shahroudy, A.; Liu, J.; Ng, T.-T.; and Wang, G. 2016. Ntu rgb+ d: A large scale dataset for 3d human activity analysis. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1010–1019.
  22. 22.Shi, L.; Zhang, Y.; Cheng, J.; and Lu, H. 2019. Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 12026–12035.
  23. 23.Shotton, J.; Fitzgibbon, A.; Cook, M.; Sharp, T.; Finocchio, M.; Moore, R.; Kipman, A.; and Blake, A. 2011. Real-time human pose recognition in parts from single depth images. In CVPR 2011, 1297–1304. Ieee.
  24. 24.Song, Y.-F.; Zhang, Z.; Shan, C.; and Wang, L. 2020. Stronger, faster and more explainable: A graph convolutional baseline for skeleton-based action recognition. In proceedings of the 28th ACM international conference on multimedia, 1625–1633.
  25. 25.Su, K.; Liu, X.; and Shlizerman, E. 2020. Predict & cluster: Unsupervised skeleton based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9631–9640.
  26. 26.Thoker, F. M.; Doughty, H.; and Snoek, C. G. 2021. Skeleton-contrastive 3D action representation learning. In Proceedings of the 29th ACM International Conference on Multimedia, 1655–1663.
  27. 27.Vemulapalli, R.; Arrate, F.; and Chellappa, R. 2014. Human action recognition by representing 3d skeletons as points in a lie group. In Proceedings of the IEEE conference on computer vision and pattern recognition, 588–595.
  28. 28.Wei, C.; Xie, L.; Ren, X.; Xia, Y.; Su, C.; Liu, J.; Tian, Q.; and Yuille, A. L. 2019. Iterative reorganization with weak spatial constraints: Solving arbitrary jigsaw puzzles for unsupervised representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1910–1919.
  29. 29.Xia, L.; Chen, C.-C.; and Aggarwal, J. K. 2012. View invariant human action recognition using histograms of 3d joints. In 2012 IEEE computer society conference on computer vision and pattern recognition workshops, 20–27. IEEE.
  30. 30.Yan, S.; Xiong, Y.; and Lin, D. 2018. Spatial temporal graph convolutional networks for skeleton-based action recognition. In Thirty-second AAAI conference on artificial intelligence.
  31. 31.Zbontar, J.; Jing, L.; Misra, I.; LeCun, Y.; and Deny, S. 2021. Barlow twins: Self-supervised learning via redundancy reduction. In International Conference on Machine Learning, 12310–12320. PMLR.
  32. 32.Zhang, P.; Lan, C.; Xing, J.; Zeng, W.; Xue, J.; and Zheng, N. 2017. View adaptive recurrent neural networks for high performance human action recognition from skeleton data. In Proceedings of the IEEE international conference on computer vision, 2117–2126.
  33. 33.Zhang, R.; Isola, P.; and Efros, A. A. 2016. Colorful image colorization. In European conference on computer vision, 649–666. Springer.
  34. 34.Zhang, Z. 2012. Microsoft kinect sensor and its effect. IEEE multimedia, 19(2): 4–10.
  35. 35.Zheng, N.; Wen, J.; Liu, R.; Long, L.; Dai, J.; and Gong, Z. 2018. Unsupervised representation learning with long-term dynamics for skeleton based action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32.

Citation

MLA
Zhou, Y., et al. “Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 3, 2023, pp. 3825–33, https://doi.org/10.1609/aaai.v37i3.25495.
APA
Zhou, Y., Duan, H., Rao, A., Su, B., & Wang, J. (2023). Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences. Proceedings of the AAAI Conference on Artificial Intelligence, 37(3), 3825–3833. https://doi.org/10.1609/aaai.v37i3.25495
Chicago
Zhou, Y., H. Duan, A. Rao, B. Su, and J. Wang. 2023. “Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences”. Proceedings of the AAAI Conference on Artificial Intelligence 37 (3): 3825–33. https://doi.org/10.1609/aaai.v37i3.25495.
Harvard
Zhou, Y. et al. (2023) “Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences”, Proceedings of the AAAI Conference on Artificial Intelligence, 37(3), pp. 3825–3833. Available at: https://doi.org/10.1609/aaai.v37i3.25495.
Vancouver
1. Zhou Y, Duan H, Rao A, Su B, Wang J (2023) Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences. Proceedings of the AAAI Conference on Artificial Intelligence 37:3825–3833

BibTeX

@article{Zhou_2023, title={Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton Sequences}, volume={37}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/aaai.v37i3.25495}, DOI={10.1609/aaai.v37i3.25495}, number={3}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Zhou, Yujie and Duan, Haodong and Rao, Anyi and Su, Bing and Wang, Jiaqi}, year={2023}, month=June, pages={3825–3833} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF