Revisiting Skeleton-based Action Recognition

Haodong DuanYue ZhaoKai ChenDahua LinBo Dai

article2022CVPR761 citations

Proposes PoseConv3D, a 3D-CNN framework that replaces traditional graph-based representations with 3D heatmap volumes of 2D skeletons to achieve state-of-the-art action recognition while handling multi-person scenes and multi-modal fusion without added computational complexity.

Listen

Automated human action recognition from video is vital for modern applications ranging from sports analytics to public safety monitoring. While skeleton-based representations offer clean, privacy-preserving signals by ignoring complex backgrounds and lighting shifts, the dominant approach—graph convolutional networks—faces serious operational hurdles. Graph networks represent human joints as discrete coordinate points, which makes them fragile when upstream pose estimators introduce errors. Furthermore, their computational load multiplies linearly with each additional person in a scene, and their irregular structure makes them difficult to combine smoothly with standard video and optical data.

The article introduces and evaluates PoseConv3D, an alternative framework that replaces graph-based coordinates with three-dimensional heatmap volumes processed by three-dimensional convolutional neural networks. The study demonstrates that converting 2D joint and limb positions into continuous volumetric heatmaps delivers superior accuracy, greater noise resilience, and seamless integration with other video modalities.

To evaluate the framework, the authors conducted comprehensive experiments across six major video action benchmarks (including NTURGB+D, FineGYM, Kinetics400, and UCF101) spanning individual and group activities. The method applies standard two-dimensional human pose estimators, aggregates the resulting joint and limb confidence maps across video frames, and leverages subjects-centered cropping alongside uniform temporal sampling to maximize computational efficiency without losing global action context.

The results show substantial improvements across key operational metrics. First, PoseConv3D achieved state-of-the-art accuracy on five of six standard skeleton-based benchmarks and all eight multi-modality benchmarks tested, reaching up to 94.1% accuracy on NTU-60 and 94.3% on FineGYM. Second, the model demonstrated extreme robustness against tracking noise: when limb keypoints were randomly dropped during testing, PoseConv3D suffered less than a 1% performance loss, compared to a steep 14.3% drop in standard graph-based models. Third, it scaled exceptionally well to crowded environments; on a group volleyball benchmark with 13 people per frame, PoseConv3D reduced computational floating-point operations by over 75% compared to graph models while improving accuracy by 2.1%. Finally, early cross-modality fusion combining video frames and skeleton heatmaps consistently surpassed traditional late-stage fusion across all datasets.

These findings indicate that organizations deploying action recognition systems can achieve higher reliability at significantly lower computational expense, particularly in multi-person or real-world video settings where input data is inherently noisy. Unlike graph neural networks, which require delicate dataset-specific tuning and fail when pose estimators shift, 3D convolutional networks provide a unified, predictable architecture that easily integrates into existing video processing pipelines.

Organizations developing computer vision systems should consider transitioning from coordinate-based graph networks to 3D heatmap-based convolutional pipelines. Technical teams should adopt 2D top-down pose estimators and test multi-modal early fusion architectures when combining video feeds with skeletal data. Before wide-scale operational deployment, teams should conduct internal pilot tests to verify performance across domain-specific camera views and validate resource utilization on target edge or cloud hardware.

Cover for Revisiting Skeleton-based Action Recognition

Abstract

Human skeleton, as a compact representation of human action, has received increasing attention in recent years. Many skeleton-based action recognition methods adopt GCNs to extract features on top of human skeletons. Despite the positive results shown in these attempts, GCN-based methods are subject to limitations in robustness, interoperability, and scalability. In this work, we propose PoseConv3D, a new approach to skeleton-based action recognition. PoseConv3D relies on a 3D heatmap volume instead of a graph sequence as the base representation of human skeletons. Compared to GCN-based methods, PoseConv3D is more effective in learning spatiotemporal features, more robust against pose estimation noises, and generalizes better in cross-dataset settings. Also, PoseConv3D can handle multiple-person scenarios without additional computation costs. The hierarchical features can be easily integrated with other modalities at early fusion stages, providing a great design space to boost the performance. PoseConv3D achieves the state-of-the-art on five of six standard skeleton-based action recognition benchmarks. Once fused with other modalities, it achieves the state-of-the-art on all eight multi-modality action recognition benchmarks. Code has been made available at: https://github.com/kennymckormick/pyskl.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Framework
  • 3.1 Good Practices for Pose Extraction
  • 3.2 From 2D Poses to 3D Heatmap Volumes
  • 3.3 3D-CNN for Skeleton-based Action Recognition
  • 4 Experiments
  • 4.1 Dataset Preparation
  • 4.2 Good properties of PoseConv3D
  • 4.3 Multi-Modality Fusion with RGBPose-Conv3D
  • 4.4 Comparisons with the state-of-the-art
  • 4.5 Ablation on Heatmap Processing
  • 5 Conclusion
  • A Visualization
  • B Generating Pseudo Heatmap Volumes.
  • C Detailed Architectures of PoseConv3D
  • C.1 Different variants of PoseConv3D.
  • C.2 RGBPose-Conv3D instantiated with SlowOnly.
  • D Supplementary Experiments
  • D.1 Ablation Study on Pose Extraction
  • D.2 Multi-Modality Action Recognition Results on UCF101 and HMDB51
  • D.3 Using 3D Skeletons in PoseConv3D
  • D.4 Practice for Group Activity Recognition
  • D.5 Uniform Sampling for RGB-based recognition
  • D.6 NTU-60 Error Analysis
  • D.7 Why skeleton-based pose estimation performs poorly on Kinetics400
  • References

Knowls

  1. Knowl 1 — PoseConv3D represents pose sequences as volumes for 3D convolution

    model/method

    PoseConv3D classifies an action from a 3D heatmap volume rather than applying graph convolutions directly to joint coordinates. For each video frame, the 2D human pose is represented by heatmaps for its joints or limbs; heatmaps from successive frames are stacked along time to produce an input of shape K×T×H×WK\times T\times H\times W, where KK is the number of heatmap channels, TT is the number of sampled frames, and H,WH,W are the spatial dimensions. A 3D convolutional neural network (3D-CNN) processes the volume and predicts the action class. The framework illustration on page 4 depicts pose estimation, heatmap stacking and preprocessing, and 3D-CNN classification as this end-to-end pipeline.

  2. Knowl 2 — Joint and limb heatmaps can be constructed from pose coordinates and confidences

    equation

    When pose-estimator heatmaps are unavailable, PoseConv3D constructs a 2D heatmap for each joint from its coordinate-confidence triplet. For joint channel kk, pixel (i,j)(i,j), joint coordinate (xk,yk)(x_k,y_k) in pixel units, confidence score ckc_k, and Gaussian scale σ\sigma in pixels, the joint heatmap is

    Jkij=exp⁡ ⁣(−(i−xk)2+(j−yk)22σ2)ck.J_{kij}=\exp\!\left(-\frac{(i-x_k)^2+(j-y_k)^2}{2\sigma^2}\right)c_k.

    A limb channel can instead represent the segment between joints aka_k and bkb_k. Let D((i,j),seg⁡[ak,bk])D((i,j),\operatorname{seg}[a_k,b_k]) be the pixel distance from pixel (i,j)(i,j) to the segment joining the two joints, and let cak,cbkc_{a_k},c_{b_k} be their confidence scores. The limb heatmap is

    Lkij=exp⁡ ⁣(−D((i,j),seg⁡[ak,bk])22σ2)min⁡(cak,cbk).L_{kij}=\exp\!\left(-\frac{D((i,j),\operatorname{seg}[a_k,b_k])^2}{2\sigma^2}\right)\min(c_{a_k},c_{b_k}).

    The resulting joint or limb maps are stacked across frames along the temporal dimension. For multiple people, the maps for the same joint or limb channel are accumulated across people without increasing the heatmap dimensions.

  3. Knowl 3 — Pose extraction quality and estimator choice affect skeleton recognition

    model/method

    PoseConv3D uses 2D poses, which the paper reports are generally of better quality than available 3D poses. Its standard extraction pipeline is top-down: a human detector supplies person proposals to a 2D pose estimator. In the reported experiments, Faster R-CNN with a ResNet-50 backbone detects people and an HRNet pretrained on COCO keypoints estimates poses; top-down estimators are favored over bottom-up alternatives based on their performance on standard pose-estimation benchmarks. For datasets other than FineGYM, poses are extracted by applying this pipeline to RGB frames. FineGYM instead uses ground-truth athlete boxes in its default pose extraction. Pose estimates may be stored as coordinate-confidence triplets (x,y,c)(x,y,c) rather than full heatmaps to reduce storage substantially, at the cost of a small reported performance drop; Gaussian joint or limb maps can subsequently be reconstructed from those triplets.

  4. Knowl 4 — Subject-centered cropping and uniform temporal sampling reduce heatmap input redundancy

    model/method

    PoseConv3D reduces the spatial and temporal size of its heatmap volumes while retaining the pose sequence. Subject-centered cropping finds the smallest spatial box enclosing the poses of interest across the video, crops the frames to that box, and resizes them to the target input resolution. On FineGYM, with input size 32×56×5632\times56\times56, this preprocessing raises Mean Top-1 accuracy from 91.7% to 92.7%. For temporal reduction, uniform sampling divides the full video into nn equal-duration segments and randomly selects one frame from each segment, producing nn sampled frames. Unlike fixed-stride sampling from a short temporal window, this covers the full clip. Tests on FineGYM and NTU-60 found uniform sampling consistently better than fixed strides; on FineGYM, 1-clip uniform testing could outperform fixed-stride 10-clip testing. The improvement was especially associated with longer videos, whose full action dynamics are less likely to fit within a short sampled window.

  5. Knowl 5 — Lightweight 3D-CNN backbones retain accuracy on pose volumes

    data/table

    PoseConv3D adapts 3D-CNNs for pose volumes by removing early downsampling and using shallower or narrower networks than are typically used for RGB clips. In the reported setting, pose-volume spatial resolution is four times smaller than the RGB input resolution. The variants below were evaluated on NTU-60 cross-subject (X-Sub); accuracy is reported in percent, FLOPs in giga floating-point operations, and parameters in millions. The results show that smaller variants can reduce computation substantially with at most a 0.3 percentage-point accuracy reduction within each backbone family. SlowOnly was selected as the default backbone for its simplicity and recognition performance.

    Could not parse LaTeX table

    Here, HR denotes doubled height and width, wd denotes doubled channel width, and s denotes a shallower network.

  6. Knowl 6 — PoseConv3D is less affected than MS-G3D by missing limb keypoints

    empirical result

    Robustness was tested on FineGYM by randomly dropping one of the eight limb keypoints in each frame with probability pp and measuring Mean Top-1 accuracy. The table reports the observed accuracy for MS-G3D, MS-G3D trained with the same keypoint-drop perturbation, and Pose-SlowOnly. Without robust training, MS-G3D falls from 92.0% at p=0p=0 to 77.7% at p=1p=1, whereas Pose-SlowOnly falls from 92.4% to 91.5%. Robust training reduces MS-G3D's sensitivity to test-time dropping, but has a lower score at p=0p=0 than ordinary MS-G3D.

    Could not parse LaTeX table

    The reported result supports greater tolerance of Pose-SlowOnly to this particular input perturbation; it does not establish robustness to every kind of pose-estimation error.

  7. Knowl 7 — PoseConv3D transfers better across pose-estimator and person-box quality shifts

    empirical result

    A FineGYM cross-annotation evaluation compared models trained and tested with high-quality (HQ) or low-quality (LQ) annotations. In the estimator experiment, HRNet poses were HQ and MobileNet poses were LQ. In the person-box experiment, ground-truth boxes were HQ and tracking boxes were LQ. Each entry is Mean Top-1 accuracy in percent; the train-to-test directions are shown explicitly. PoseConv3D scores above MS-G3D in all six train/test conditions, including when training and testing annotation quality differ.

    Could not parse LaTeX table

    The comparison indicates that PoseConv3D loses less performance than the graph baseline under these particular changes in pose estimator or person-box source.

  8. Knowl 8 — A single heatmap volume handles many people without person-proportional computation

    empirical result

    On the Volleyball group-activity dataset, each video contains about 12 people (the evaluated input example has 13 people) and 20 frames. A graph model's input has shape 13×20×17×313\times20\times17\times3, so its computation grows with the number of people; the reported MS-G3D configuration uses 2.8M parameters and 7.2 GFLOPs. PoseConv3D accumulates all people into one heatmap volume of shape 17×12×56×5617\times12\times56\times56 rather than allocating a volume per person. With Pose-SlowOnly base channel width 16, this configuration uses 0.52M parameters and 1.6 GFLOPs and reaches 91.3% Top-1 accuracy on Volleyball validation, 2.1 percentage points above the graph-based approach. The fixed-size accumulated volume is the mechanism by which the pose representation avoids computation increasing in proportion to the number of people.

  9. Knowl 9 — RGBPose-Conv3D fuses RGB and pose features through asymmetric pathways

    model/method

    RGBPose-Conv3D is a two-pathway 3D-CNN for early fusion of RGB frames and pose heatmap volumes. The pose pathway is narrower, shallower, and lower in spatial resolution than the RGB pathway, reflecting the different input modalities. Bidirectional lateral connections transfer features between the pathways to enable early-stage fusion. To limit overfitting, training uses a separate cross-entropy loss for each pathway. The reported training procedure first trains RGB-only and pose-only models, initializes the corresponding pathways from those models, and then fine-tunes the two-pathway network to learn the lateral connections. At inference, the pathway prediction scores can also be fused, combining early feature fusion with late score fusion.

  10. Knowl 10 — Bidirectional early fusion improves RGB-pose recognition across modality balances

    data/table

    The RGBPose-Conv3D fusion ablation compares late score fusion alone with one-way and two-way early lateral connections. Accuracy entries are reported as 1-clip / 10-clip results in percent. Bidirectional RGB-to-pose and pose-to-RGB connections outperform either one-way direction; early-plus-late fusion reaches 94.1% at 10 clips, compared with 93.4% for late fusion alone. The early-plus-late strategy also improves results on both FineGYM, where pose is more important, and NTU-60, where RGB is more important.

    Could not parse LaTeX table
    Could not parse LaTeX table

    For the dataset comparison, paired values are 1-clip / 10-clip accuracy in percent. The gains on both datasets show that early-plus-late fusion was useful in the tested pose-dominant and RGB-dominant cases.

  11. Knowl 11 — PoseConv3D matches or exceeds graph baselines on most skeleton benchmarks

    data/table

    The skeleton-recognition comparison evaluates SlowOnly PoseConv3D with 10-clip testing and heatmap-volume input of size 48×56×5648\times56\times56. The models use high-quality 2D poses; MS-G3D++ is the graph baseline evaluated on those same poses. Accuracies are percentages. Joint-plus-limb PoseConv3D exceeds MS-G3D++ on five of the six listed benchmarks, with NTU-120 X-Sub the exception. The joint-only variant also provides a strong result across the benchmarks.

    Could not parse LaTeX table

    The comparison with MS-G3D++ controls for the pose annotations by using the same extracted 2D skeletons: the graph model consumes coordinate-confidence triplets, while PoseConv3D consumes heatmaps generated from them. The reported results therefore compare the recognition models under a shared pose source, while the published-pose MS-G3D row reflects a different pose setup.

  12. Knowl 12 — RGB-pose and late multimodal fusion set strong results on eight benchmarks

    data/table

    The paper evaluates RGBPose-Conv3D on FineGYM and four NTU-60/NTU-120 splits, and late fusion of PoseConv3D scores with existing RGB or RGB-plus-flow systems on Kinetics-400, UCF101, and HMDB51. The table gives the previous comparison result and the reported fused result, with modalities identified in parentheses. On the three late-fusion benchmarks, adding pose gives 85.5% on Kinetics-400, 98.8% on UCF101, and 85.0% on HMDB51, each above the listed previous result. RGB-pose fusion also improves over the listed prior values on the FineGYM and NTU splits.

    Could not parse LaTeX table

    Here RR, FF, and PP denote RGB, optical flow, and pose, respectively; reported values are recognition accuracies in percent. The paper uses RGBPose-Conv3D for the FineGYM and NTU results and late score fusion for Kinetics-400, UCF101, and HMDB51.

  13. Knowl 13 — 3D heatmap volumes outperform 2D heatmap aggregation baselines

    data/table

    An apples-to-apples evaluation compares 3D heatmap volumes processed by 3D-CNNs against 2D heatmap aggregation methods processed as 2D-CNN inputs. The datasets are HMDB51, UCF101, and NTU-60 X-Sub; accuracy is reported in percent, FLOPs in giga operations, and parameter counts as listed. Pose-SlowOnly has the highest accuracy in all three datasets. The lightweight Pose-X3D-s is also more accurate than PoTion and PA3D in all three, with FLOPs equal to PA3D's reported 0.60G and fewer parameters than either aggregation baseline.

    Could not parse LaTeX table

    These results support the paper's motivation for retaining the pose heatmaps as a spatiotemporal volume rather than aggregating them into a 2D representation, while also showing that the accuracy-computation tradeoff depends on the chosen 3D-CNN backbone.

Coverage note — No major contributed method or experimental finding was deliberately omitted; dataset descriptions, related work, acknowledgements, and implementation details that do not add standalone findings were excluded.

References

  1. 1.Sadjad Asghari-Esfeden, Mario Sznaier, and Octavia Camps. Dynamic motion representation for human action recognition. In WACV, pages 557–566, 2020.
  2. 2.Carlos Caetano, Jessica Sena, François Bremond, Jeferson A Dos Santos, and William Robson Schwartz. Skelemotion: A new representation of skeleton joint sequences based on motion information for 3d action recognition. In AVSS, pages 1–8. IEEE, 2019.
  3. 3.Jinmiao Cai, Nianjuan Jiang, Xiaoguang Han, Kui Jia, and Jiangbo Lu. Jolo-gcn: mining joint-centered light-weight information for skeleton-based action recognition. In WACV, pages 2735–2744, 2021.
  4. 4.Z. Cao, G. Hidalgo Martinez, T. Simon, S. Wei, and Y. A. Sheikh. Openpose: Realtime multi-person 2d pose estimation using part affinity fields. TPAMI, 2019.
  5. 5.Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. Realtime multi-person 2d pose estimation using part affinity fields. In CVPR, 2017.
  6. 6.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, pages 6299–6308, 2017.
  7. 7.Yuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li, Ying Deng, and Weiming Hu. Channel-wise topology refinement graph convolution for skeleton-based action recognition. In ICCV, pages 13359–13368, 2021.
  8. 8.Bowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi, Thomas S Huang, and Lei Zhang. Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation. In CVPR, 2020.
  9. 9.Ke Cheng, Yifan Zhang, Xiangyu He, Weihan Chen, Jian Cheng, and Hanqing Lu. Skeleton-based action recognition with shift graph convolutional network. In CVPR, pages 183–192, 2020.
  10. 10.Vasileios Choutas, Philippe Weinzaepfel, Jérôme Revaud, and Cordelia Schmid. Potion: Pose motion representation for action recognition. In CVPR, pages 7024–7033, 2018.
  11. 11.MMAction2 Contributors. Openmmlab’s next generation video understanding toolbox and benchmark. https://github.com/open-mmlab/mmaction2, 2020.
  12. 12.Srijan Das, Rui Dai, Di Yang, and Francois Bremond. Vpn++: Rethinking video-pose embeddings for understanding activities of daily living. arXiv:2105.08141, 2021.
  13. 13.Srijan Das, Saurav Sharma, Rui Dai, Francois Bremond, and Monique Thonnat. Vpn: Learning video-pose embedding for activities of daily living. In ECCV, pages 72–90. Springer, 2020.
  14. 14.Mahdi Davoodikakhki and KangKang Yin. Hierarchical action classification with network pruning. In ISVC, pages 291–305. Springer, 2020.
  15. 15.Yong Du, Wei Wang, and Liang Wang. Hierarchical recurrent neural network for skeleton based action recognition. In CVPR, pages 1110–1118, 2015.
  16. 16.Haodong Duan, Yue Zhao, Yuanjun Xiong, Wentao Liu, and Dahua Lin. Omni-sourced webly-supervised learning for video recognition. In ECCV, pages 670–688. Springer, 2020.
  17. 17.Christoph Feichtenhofer. X3d: Expanding architectures for efficient video recognition. In CVPR, pages 203–213, 2020.
  18. 18.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In ICCV, pages 6202–6211, 2019.
  19. 19.Pranay Gupta, Anirudh Thatipelli, Aditya Aggarwal, Shubh Maheshwari, Neel Trivedi, Sourav Das, and Ravi Kiran Sarvadevabhatla. Quo vadis, skeleton action recognition? IJCV, 129(7):2097–2112, 2021.
  20. 20.Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet? In CVPR, pages 6546–6555, 2018.
  21. 21.Alejandro Hernandez Ruiz, Lorenzo Porzi, Samuel Rota Bulo, and Francesc Moreno-Noguer. 3d cnns on distance matrices for human action recognition. In MM, pages 1087–1095, 2017.
  22. 22.Mostafa S Ibrahim, Srikanth Muralidharan, Zhiwei Deng, Arash Vahdat, and Greg Mori. A hierarchical deep temporal model for group activity recognition. In CVPR, 2016.
  23. 23.Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 3d convolutional neural networks for human action recognition. TPAMI, 35(1):221–231, 2012.
  24. 24.Hamid Reza Vaezi Joze, Amirreza Shaban, Michael L Iuzzolino, and Kazuhito Koishida. Mmtm: Multimodal transfer module for cnn fusion. In CVPR, pages 13289–13299, 2020.
  25. 25.Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, and Farid Boussaid. A new representation of skeleton sequences for 3d action recognition. In CVPR, 2017.
  26. 26.Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre. Hmdb: a large video database for human motion recognition. In ICCV, pages 2556–2563. IEEE, 2011.
  27. 27.Heeseung Kwon, Manjin Kim, Suha Kwak, and Minsu Cho. Learning self-similarity in space and time as generalized motion for action recognition. arXiv:2102.07092, 2021.
  28. 28.Bin Li, Xi Li, Zhongfei Zhang, and Fei Wu. Spatio-temporal graph routing for skeleton-based action recognition. In AAAI, volume 33, pages 8561–8568, 2019.
  29. 29.Chao Li, Qiaoyong Zhong, Di Xie, and Shiliang Pu. Co-occurrence feature learning from skeleton data for action recognition and detection with hierarchical aggregation. arXiv preprint arXiv:1804.06055, 2018.
  30. 30.Maosen Li, Siheng Chen, Xu Chen, Ya Zhang, Yanfeng Wang, and Qi Tian. Actional-structural graph convolutional networks for skeleton-based action recognition. In CVPR, 2019.
  31. 31.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
  32. 32.Zeyi Lin, Wei Zhang, Xiaoming Deng, Cuixia Ma, and Hongan Wang. Image-based pose representation for action recognition and hand gesture recognition. In FG, pages 532–539. IEEE, 2020.
  33. 33.Hong Liu, Juanhui Tu, and Mengyuan Liu. Two-stream 3d convolutional neural network for skeleton-based action recognition. arXiv:1705.08106, 2017.
  34. 34.Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C. Kot. Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding. TPAMI, 2019.
  35. 35.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv:2106.13230, 2021.
  36. 36.Ziyu Liu, Hongwen Zhang, Zhenghao Chen, Zhiyong Wang, and Wanli Ouyang. Disentangling and unifying graph convolutions for skeleton-based action recognition. In CVPR, 2020.
  37. 37.Diogo C Luvizon, David Picard, and Hedi Tabia. 2d/3d pose estimation and action recognition using multitask deep learning. In CVPR, pages 5137–5146, 2018.
  38. 38.Alejandro Newell, Zhiao Huang, and Jia Deng. Associative embedding: End-to-end learning for joint detection and grouping. In NeurIPS, pages 2277–2287, 2017.
  39. 39.Alejandro Newell, Kaiyu Yang, and Jia Deng. Stacked hourglass networks for human pose estimation. In ECCV, pages 483–499. Springer, 2016.
  40. 40.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv:1506.01497, 2015.
  41. 41.Kohei Sendo and Norimichi Ukita. Heatmapping of people involved in group activities. In ICMVA, 2019.
  42. 42.Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang. Ntu rgb+d: A large scale dataset for 3d human activity analysis. In CVPR, June 2016.
  43. 43.Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin. Finegym: A hierarchical video dataset for fine-grained action understanding. In CVPR, pages 2616–2625, 2020.
  44. 44.Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. Skeleton-based action recognition with directed graph neural networks. In CVPR, pages 7912–7921, 2019.
  45. 45.Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In CVPR, pages 12026–12035, 2019.
  46. 46.Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. Decoupled spatial-temporal attention network for skeleton-based action recognition. arXiv:2007.03263, 2020.
  47. 47.Karen Simonyan and Andrew Zisserman. Two-stream convolutional networks for action recognition in videos. arXiv:1406.2199, 2014.
  48. 48.Yi-Fan Song, Zhang Zhang, Caifeng Shan, and Liang Wang. Richly activated graph convolutional network for robust skeleton-based action recognition. TSCVT, 31(5):1915–1925, 2020.
  49. 49.Yi-Fan Song, Zhang Zhang, Caifeng Shan, and Liang Wang. Stronger, faster and more explainable: A graph convolutional baseline for skeleton-based action recognition. In MM, pages 1625–1633, 2020.
  50. 50.Yi-Fan Song, Zhang Zhang, Caifeng Shan, and Liang Wang. Constructing stronger and faster baselines for skeleton-based action recognition. arXiv preprint arXiv:2106.15125, 2021.
  51. 51.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv:1212.0402, 2012.
  52. 52.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. JMLR, 2014.
  53. 53.Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. Deep high-resolution representation learning for human pose estimation. In CVPR, pages 5693–5703, 2019.
  54. 54.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3d convolutional networks. In ICCV, 2015.
  55. 55.Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli. Video classification with channel-separated convolutional networks. In ICCV, pages 5552–5561, 2019.
  56. 56.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In CVPR, pages 6450–6459, 2018.
  57. 57.Raviteja Vemulapalli, Felipe Arrate, and Rama Chellappa. Human action recognition by representing 3d skeletons as points in a lie group. In CVPR, pages 588–595, 2014.
  58. 58.Jiang Wang, Zicheng Liu, Ying Wu, and Junsong Yuan. Mining actionlet ensemble for action recognition with depth cameras. In CVPR, pages 1290–1297. IEEE, 2012.
  59. 59.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks: Towards good practices for deep action recognition. In ECCV, 2016.
  60. 60.Philippe Weinzaepfel and Gregory Rogez. Mimetics: Towards understanding human actions out of context. IJCV, 129(5):1675–1690, 2021.
  61. 61.Bin Xiao, Haiping Wu, and Yichen Wei. Simple baselines for human pose estimation and tracking. In ECCV, 2018.
  62. 62.Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer. Audiovisual slowfast networks for video recognition. arXiv:2001.08740, 2020.
  63. 63.An Yan, Yali Wang, Zhifeng Li, and Yu Qiao. Pa3d: Pose-action 3d machine for video recognition. In CVPR, pages 7922–7931, 2019.
  64. 64.Sijie Yan, Yuanjun Xiong, and Dahua Lin. Spatial temporal graph convolutional networks for skeleton-based action recognition. In AAAI, volume 32, 2018.
  65. 65.Hao Yang, Dan Yan, Li Zhang, Dong Li, YunDa Sun, ShaoDi You, and Stephen J Maybank. Feedback graph convolutional network for skeleton-based action recognition. arXiv:2003.07564, 2020.
  66. 66.Dingyuan Zhu, Ziwei Zhang, Peng Cui, and Wenwu Zhu. Robust graph convolutional networks against adversarial attacks. In KDD, pages 1399–1407, 2019.

Citation

MLA
Duan, H., et al. “Revisiting Skeleton-based Action Recognition”. arXiv, 2021, http://arxiv.org/abs/2104.13586v2.
APA
Duan, H., Zhao, Y., Chen, K., Lin, D., & Dai, B. (2021). Revisiting Skeleton-based Action Recognition. arXiv. http://arxiv.org/abs/2104.13586v2
Chicago
Duan, H., Y. Zhao, K. Chen, D. Lin, and B. Dai. 2021. “Revisiting Skeleton-based Action Recognition”. arXiv. http://arxiv.org/abs/2104.13586v2.
Harvard
Duan, H. et al. (2021) “Revisiting Skeleton-based Action Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2104.13586v2.
Vancouver
1. Duan H, Zhao Y, Chen K, Lin D, Dai B (2021) Revisiting Skeleton-based Action Recognition. arXiv

BibTeX

@article{duan2021revisiting,
  title = {Revisiting Skeleton-based Action Recognition},
  author = {Duan, Haodong and Zhao, Yue and Chen, Kai and Lin, Dahua and Dai, Bo},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2104.13586v2},
  eprint = {2104.13586}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE