BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition

Yuxuan ZhouXudong YanZhi-Qi ChengYan YanQi DaiXian-Sheng Hua

article2024CVPR86 citations

Proposes BlockGCN, an architecture for skeleton-based action recognition that preserves physical bone connectivity through graph distances and persistent homology while cutting graph convolution parameters by over 40% using block-diagonal weight matrices.

Listen

Human action recognition based on 3D skeleton data is an essential technology for practical applications such as healthcare monitoring and automated video analysis, where both computational efficiency and robustness to environmental conditions are critical. Current state-of-the-art systems rely on graph neural networks to model body joints and their structural relationships. However, existing models suffer from two major deficiencies: they progressively forget physical bone connectivity during training as they optimize joint connections, and they rely on redundant, computationally heavy network architectures to capture diverse joint interactions across complex actions.

The article demonstrates that explicitly encoding physical bone structure and decoupling multi-relational feature projections solves the forgetting and efficiency issues in graph-based action recognition. To achieve this, the authors introduce BlockGCN, an architecture that combines static relative graph distances, persistent homology analysis for dynamic action-specific movement patterns, and an efficient block-diagonal graph convolution module named BlockGC.

The authors evaluated the framework across three standard benchmark datasets: NTU RGB+D, NTU RGB+D 120, and Northwestern-UCLA. The experimental setup compared BlockGCN against prevailing state-of-the-art models using standardized evaluation protocols, including both multi-modal fusion and single-modality evaluations, while conducting controlled ablation tests to isolate the contribution of each architectural component.

The evaluation revealed several key findings. First, BlockGCN achieved new benchmark performance across all datasets, outperforming the previous state-of-the-art by an average of 0.5% in accuracy under the standard four-modality setup and by 0.8% to 1.4% when using the single-joint modality. Second, the BlockGC design eliminated parameter redundancy, cutting total model parameters by over 40% compared to baseline graph convolutions and reducing overall network parameters by 35% to 38% compared to leading alternative models. Third, empirical tests showed that traditional network initializations fail to preserve skeletal structure during training, whereas the proposed static and dynamic topological encodings successfully retained physical constraints and boosted standalone accuracy by up to 1.7%.

These findings indicate that artificial intelligence systems can achieve superior recognition accuracy with significantly fewer parameters and lower computational overhead. By moving away from bloated ensemble convolutions and attention mechanisms, organizations can deploy high-performing human action recognition models on resource-constrained devices, reducing inference costs and processing latency while maintaining high reliability in safety-critical and medical environments.

Organizations developing or deploying video analytics systems should transition to decoupled, topology-preserving graph architectures like BlockGCN to optimize resource utilization without sacrificing predictive power. Further research should explore applying these topological encoding principles and block-diagonal convolution strategies to other domains that process graph-structured data.

The presented results are supported by extensive benchmark experiments on standard public datasets. However, decision-makers should note that evaluations were conducted in controlled benchmark settings on pose sequences resized to fixed frame lengths. Additional pilot testing in live, unconstrained production environments is advised to ensure generalization across irregular camera setups and noisier skeleton tracking pipelines.

No sufficiently relevant recommendations were found.

Cover for BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition

Abstract

Graph Convolutional Networks (GCNs) have long set the state-of-the-art in skeleton-based action recognition, leveraging their ability to unravel the complex dynamics of human joint topology through the graph’s adjacency matrix. However, an inherent flaw has come to light in these cutting-edge models: they tend to optimize the adjacency matrix jointly with the model weights. This process, while seemingly efficient, causes a gradual decay of bone connectivity data, resulting in a model indifferent to the very topology it sought to represent. To remedy this, we propose a two-fold strategy: (1) We introduce an innovative approach that encodes bone connectivity by harnessing the power of graph distances to describe the physical topology; we further incorporate action-specific topological representation via persistent homology analysis to depict systemic dynamics. This preserves the vital topological nuances often lost in conventional GCNs. (2) Our investigation also reveals the redundancy in existing GCNs for multi-relational modeling, which we address by proposing an efficient refinement to Graph Convolutions (GC) - the BlockGC. This significantly reduces parameters while improving performance beyond original GCNs. Our full model, BlockGCN, establishes new benchmarks in skeleton-based action recognition across all model categories. Its high accuracy and lightweight design, most notably on the large-scale NTU RGB+D 120 dataset, stand as strong validation of the efficacy of BlockGCN.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Skeleton-based Action Recognition
  • 2.2. GCNs for Skeleton-based Action Recognition
  • 3. Method
  • 3.1. Problem Formulation
  • 3.2. Topological Encoding
  • 3.2.1 Static Topological Encoding
  • 3.2.2 Dynamic Topological Encoding
  • 3.3. Efficient Multi-Relational Modeling
  • 3.4. Model Architecture
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Comparison with State-of-the-art
  • 4.3. Ablation Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Block Graph Convolution (BlockGC)

    model/method

    Block Graph Convolution (BlockGC) is an efficient graph convolution layer designed for multi-relational modeling in skeleton-based action recognition. In standard Graph Convolutional Networks (GCNs), updating node representations with a full weight projection matrix W(l)∈Rd×dW^{(l)} \in \mathbb{R}^{d \times d} incurs redundant parameters and computational cost O(∣V∣d2)O(|\mathcal{V}|d^2), where ∣V∣|\mathcal{V}| is the number of skeleton joints and dd is the feature channel dimension.

    BlockGC splits the feature channels dd across the hidden representations into KK disjoint feature groups of size d/Kd/K. It employs a block diagonal feature projection weight matrix W(l)=diag⁡(W1(l),W2(l),…,WK(l))W^{(l)} = \operatorname{diag}(W_1^{(l)}, W_2^{(l)}, \dots, W_K^{(l)}), where each sub-matrix Wk(l)∈R(d/K)×(d/K)W_k^{(l)} \in \mathbb{R}^{(d/K) \times (d/K)} operates exclusively on the kk-th feature group. Spatial aggregation and channel projection are executed in parallel across the KK groups:

    Hk(l)=σ((Ak(l)+Bk(l))(Hk(l−1)+Ck(l−1))Wk(l))for k∈{1,…,K}H_k^{(l)} = \sigma\left( (A_k^{(l)} + B_k^{(l)}) (H_k^{(l-1)} + C_k^{(l-1)}) W_k^{(l)} \right) \quad \text{for } k \in \{1, \dots, K\}

    where for layer ll and group kk:

    • Hk(l−1)∈R∣V∣×T×(d/K)H_k^{(l-1)} \in \mathbb{R}^{|\mathcal{V}| \times T \times (d/K)} is the input representation across ∣V∣|\mathcal{V}| joints and TT temporal frames,
    • Ak(l)∈R∣V∣×∣V∣A_k^{(l)} \in \mathbb{R}^{|\mathcal{V}| \times |\mathcal{V}|} is a learnable group-specific adjacency matrix,
    • Bk(l)∈R∣V∣×∣V∣B_k^{(l)} \in \mathbb{R}^{|\mathcal{V}| \times |\mathcal{V}|} is the static topological encoding matrix,
    • Ck(l−1)∈R∣V∣×(d/K)C_k^{(l-1)} \in \mathbb{R}^{|\mathcal{V}| \times (d/K)} is the group-partitioned dynamic topological encoding,
    • σ(⋅)\sigma(\cdot) is the ReLU activation function.

    The overall output H(l)∈R∣V∣×T×dH^{(l)} \in \mathbb{R}^{|\mathcal{V}| \times T \times d} is formed by concatenating the group representations [H1(l),…,HK(l)][H_1^{(l)}, \dots, H_K^{(l)}]. Decoupling the feature projections enforces semantic independence across groups while cutting computational complexity to O(∣V∣d2K)O(\frac{|\mathcal{V}|d^2}{K}) and feature-projection weight parameters to d2K\frac{d^2}{K}.

  2. Knowl 2 — Static Topological Encoding via Shortest Path Distances

    model/method

    Static Topological Encoding preserves the physical connectivity structure of human skeletons throughout GCN training, counteracting the decay of initial connectivity information caused by gradient updates on learnable adjacency matrices.

    Given the physical human skeleton graph GS=(V,E)\mathcal{G}_S = (\mathcal{V}, \mathcal{E}), where V\mathcal{V} denotes the set of joints and E\mathcal{E} denotes the physical bones, the relative graph distance between joints viv_i and vjv_j is defined by the Shortest Path Distance (SPD):

    di,j=min⁡P∈Paths⁡(GS){∣P∣∣P1=vi,P∣P∣=vj}d_{i, j} = \min_{P \in \operatorname{Paths}(\mathcal{G}_S)} \{|P| \mid P_1 = v_i, P_{|P|} = v_j\}

    where PP is a sequence of connected vertices in GS\mathcal{G}_S, ∣P∣|P| denotes path length in terms of hops, and P1,P∣P∣P_1, P_{|P|} are the start and end vertices.

    A shared 1D trainable lookup table E={eindex}E = \{e_{\text{index}}\} maps each integer path distance di,jd_{i,j} to a learnable scalar weight:

    Bij=edi,jB_{ij} = e_{d_{i,j}}

    The resulting matrix B∈R∣V∣×∣V∣B \in \mathbb{R}^{|\mathcal{V}| \times |\mathcal{V}|} is directly added to the learnable spatial adjacency matrix AA during graph aggregation. Because the mapping from joint pairs to parameter indices in EE is hard-coded by the fixed bone graph topology, the structural physical constraints cannot be forgotten during model optimization.

  3. Knowl 3 — Dynamic Topological Encoding via Persistent Homology

    model/method

    Dynamic Topological Encoding extracts multi-scale topological invariants of skeleton pose dynamics for each action sequence using persistent homology from algebraic topology.

    Given an input skeleton pose sequence, a dynamic weighted complete graph GD\mathcal{G}_D is constructed where the nodes are joints V\mathcal{V} and edge weights wij=∥xi−xj∥2w_{ij} = \|x_i - x_j\|_2 represent pairwise Euclidean 3D spatial distances. Using the Vietoris-Rips complex construction, a nested sequence of simplicial complexes (filtration) is formed:

    ∅=K0⊆K1⊆K2⊆⋯⊆Km=K\emptyset = \mathcal{K}_0 \subseteq \mathcal{K}_1 \subseteq \mathcal{K}_2 \subseteq \dots \subseteq \mathcal{K}_m = \mathcal{K}

    where Ki\mathcal{K}_i corresponds to the subcomplex formed at scale ϵi\epsilon_i.

    From the filtration, 0-dimensional persistence diagrams/barcodes {D10,D20,…,Dp0}\{\mathcal{D}_1^0, \mathcal{D}_2^0, \dots, \mathcal{D}_p^0\} are extracted, summarizing the birth and death scales (bi0,di0)(b_i^0, d_i^0) of 0-dimensional topological features (connected components). The multiset of barcodes is transformed into a fixed-size vector representation using a differentiable vectorization mapping Ψ0:{D10,…,Dp0}→R∣V∣×d′\Psi^0: \{\mathcal{D}_1^0, \dots, \mathcal{D}_p^0\} \to \mathbb{R}^{|\mathcal{V}| \times d'}. The vectorized representation is then projected to match the hidden feature dimension of layer ll via a learned linear mapping fθ:R∣V∣×d′→R∣V∣×df_\theta: \mathbb{R}^{|\mathcal{V}| \times d'} \to \mathbb{R}^{|\mathcal{V}| \times d}:

    C(l)=fθ(Ψ0(D10,D20,…,Dp0))C^{(l)} = f_\theta\left(\Psi^0\left(\mathcal{D}_1^0, \mathcal{D}_2^0, \dots, \mathcal{D}_p^0\right)\right)

    The dynamic topological feature matrix C(l)∈R∣V∣×dC^{(l)} \in \mathbb{R}^{|\mathcal{V}| \times d} is added directly to the node hidden feature matrix H(l−1)H^{(l-1)} prior to BlockGC spatial aggregation.

  4. Knowl 4 — BlockGCN Architecture for Skeleton-Based Action Recognition

    model/method

    BlockGCN is an end-to-end deep neural network for skeleton action recognition built by stacking Block Graph Convolution (BlockGC) modules and Multi-Scale Temporal Convolution (MS-TCN) modules.

    The network comprises:

    1. Input Encoding: An input skeleton sequence of TT frames and ∣V∣|\mathcal{V}| joints is augmented with a learnable absolute positional embedding added to the joint features.
    2. Stacked Spatio-Temporal Blocks: 10 sequential stages, each consisting of:
      • A BlockGC layer integrating Static Topological Encoding B∈R∣V∣×∣V∣B \in \mathbb{R}^{|\mathcal{V}| \times |\mathcal{V}|} and Dynamic Topological Encoding C∈R∣V∣×dC \in \mathbb{R}^{|\mathcal{V}| \times d} with KK parallel channel groups (default K=8K=8).
      • A Multi-Scale Temporal Convolution Network (MS-TCN) module consisting of three temporal convolution branches with different kernel sizes and dilation rates preceded by a 1×11 \times 1 temporal convolution for channel reduction. The branch outputs are concatenated along the feature dimension.
    3. Classification Head: A global average pooling (GAP) layer over both the joint dimension ∣V∣|\mathcal{V}| and temporal dimension TT, followed by a linear layer and a softmax classifier optimizing standard cross-entropy loss.
  5. Knowl 5 — Computational and Parameter Complexity Comparison for Multi-Relational Graph Convolutions

    theoretical result

    Multi-relational graph convolution methods differ in theoretical computational time complexity and parameter count. For ∣V∣|\mathcal{V}| body joints, hidden dimension dd (where d≫∣V∣d \gg |\mathcal{V}|), and KK relation groups or ensemble branches, the asymptotic scaling is:

    Multi-Relational Modeling Method Computational Complexity Parameter Count
    Vanilla Graph Convolution (Baseline) O(∣V∣d2)O(|\mathcal{V}|d^2) d2+∣V∣2d^2 + |\mathcal{V}|^2
    Ensemble of Graph Convolutions O(K∣V∣d2)O(K|\mathcal{V}|d^2) Kd2+K∣V∣2Kd^2 + K|\mathcal{V}|^2
    Ensemble of Adjacency Matrices (DecouplingGC) O(∣V∣d2)O(|\mathcal{V}|d^2) d2+K∣V∣2d^2 + K|\mathcal{V}|^2
    Proposed BlockGC O(∣V∣d2K)O\left(\frac{|\mathcal{V}|d^2}{K}\right) d2K+K∣V∣2\frac{d^2}{K} + K|\mathcal{V}|^2

    Because d2≫∣V∣2d^2 \gg |\mathcal{V}|^2, dividing the weight projection matrix into KK diagonal blocks reduces both the dominant projection parameter cost and the computational floating point operations by a factor of KK, making BlockGC strictly more efficient than baseline GC, multi-GC ensembles, and multi-adjacency ensembles.

  6. Knowl 6 — Catastrophic Forgetting of Skeleton Topology in Learnable Adjacency Initializations

    empirical result

    Initializing learnable adjacency matrices in GCNs with physical bone connections does not preserve topological skeleton inductive biases during training. Evaluating multiple GCN architectures on the NTU RGB+D 120 Cross-Subject (X-Sub) benchmark under four distinct adjacency matrix initialization strategies yields virtually identical top-1 classification accuracies:

    Adjacency Matrix Initialization Strategy
    Model Bone Connection (%) Identity Matrix (%) All Ones (%) Kaiming Uniform (%)
    DecouplingGCN 82.0 81.9 81.7 82.1
    SkeletonGCL 84.4 84.3 84.8 84.3
    CTR-GCN 84.9 85.0 84.8 84.8

    Because gradient-based optimization causes the adjacency matrices across layers to diverge heavily from the initial physical graph into arbitrary dense representations, the initial topological structure is catastrophically forgotten. Explicit topological encodings (such as Static Topological Encoding via shortest path distances) are necessary to maintain physical skeleton constraints.

  7. Knowl 7 — Action Recognition Benchmark Performance on NTU RGB+D, NTU RGB+D 120, and Northwestern-UCLA

    empirical result

    BlockGCN outperforms previous state-of-the-art skeleton-based action recognition models across NTU RGB+D 60, NTU RGB+D 120, and Northwestern-UCLA datasets under both standard 4-stream fusion (Joint + Bone + Joint Motion + Bone Motion) and single-stream (Joint only) modalities:

    NTU RGB+D 60 NTU RGB+D 120 NW-UCLA
    Method Modalities Params FLOPs X-Sub (%) X-View (%) X-Sub (%) X-Set (%) (%)
    DC-GCN+ADG J+B+JM+BM 4.9M 1.83G 90.8 96.6 86.5 88.1 95.3
    MS-G3D J+B+JM+BM 2.8M 5.22G 91.5 96.2 86.9 88.4 –
    MST-GCN J+B+JM+BM 12.0M – 91.5 96.6 87.5 88.8 –
    CTR-GCN J+B+JM+BM 1.5M 1.97G 92.4 96.4 88.9 90.4 96.5
    EfficientGCN-B4 J+B+JM+BM 2.0M 15.2G 91.7 95.7 88.3 89.1 –
    InfoGCN J+B+JM+BM 1.6M 1.84G 92.3 96.7 89.2 90.7 96.6
    FR Head J+B+JM+BM 2.0M – 92.8 96.8 89.5 90.9 96.8
    BlockGCN (Ours) J+B+JM+BM 1.3M 1.63G 93.1 97.0 90.3 91.5 96.9
    InfoGCN J 1.6M 1.84G 89.8 95.2 85.1 86.3 –
    HDGCN J 1.7M 1.77G – – 85.7 87.3 –
    FR Head J 2.0M – 90.3 95.3 85.5 87.3 –
    BlockGCN (Ours) J 1.3M 1.63G 90.9 95.4 86.9 88.2 95.5

    With 4-stream ensemble, BlockGCN achieves state-of-the-art results (90.3% X-Sub / 91.5% X-Set on NTU 120) with 35% fewer parameters (1.3M vs 2.0M) than FR Head. On the single joint modality, BlockGCN outperforms FR Head by 1.4% on NTU 120 X-Sub.

  8. Knowl 8 — Ablation Study of BlockGC and Topological Encoding Components

    empirical result

    An ablation study on the NTU RGB+D 120 Cross-Subject (X-Sub) benchmark using a single joint modality demonstrates the isolated contributions of BlockGC, learnable absolute positional embeddings (PE), static topological encoding, and dynamic topological encoding:

    Vanilla GC BlockGC Positional Embedding Dynamic Encoding Static Encoding Parameters Accuracy (%)
    ✓ – – – – 2.1M 85.2
    – ✓ – – – 1.2M 85.8
    – ✓ ✓ – – 1.2M 86.0
    – ✓ ✓ – ✓ 1.2M 86.2
    – ✓ ✓ ✓ – 1.3M 86.7
    – ✓ ✓ ✓ ✓ 1.3M 86.9

    Replacing Vanilla GC with BlockGC reduces parameters by 43% (from 2.1M to 1.2M) while increasing accuracy by +0.6%. Adding Static Topological Encoding improves accuracy to 86.2% with negligible parameter overhead, while Dynamic Topological Encoding provides a +0.7% boost to 86.7% for +0.1M parameters. Combining all components yields 86.9% accuracy (+1.7% over the baseline) with ~38% total parameter reduction.

  9. Knowl 9 — Empirical Comparison of Channel Grouping: BlockGC vs DecouplingGC

    empirical result

    Comparing BlockGC with DecouplingGC across different channel group counts K∈{4,8,16}K \in \{4, 8, 16\} on NTU RGB+D 120 X-Sub (single joint modality) demonstrates that BlockGC achieves higher accuracy with roughly half the parameters:

    Model Groups (KK) Parameters Accuracy (%)
    Vanilla GC (Baseline) 1 2.1M 85.2
    DecouplingGC 4 2.1M 85.5
    DecouplingGC 8 2.2M 85.6
    DecouplingGC 16 2.3M 85.4
    BlockGC (Ours) 4 1.2M 85.7
    BlockGC (Ours) 8 1.2M 85.8
    BlockGC (Ours) 16 1.2M 85.6

    Because DecouplingGC maintains a full d×dd \times d weight projection matrix and only ensembles adjacency matrices, its parameter count remains between 2.1M and 2.3M. In contrast, BlockGC's block diagonal projection reduces parameters to 1.2M across all group configurations while improving accuracy, peaking at 85.8% with K=8K=8 groups.

  10. Knowl 10 — Hyperparameter and Design Analysis for Topological Encodings

    empirical result

    Ablations on design dimensions for Static and Dynamic Topological Encodings on NTU RGB+D 120 X-Sub show:

    1. Feature-wise vs. Shared Encodings:

      • For Static Topological Encoding, a channel-shared encoding achieves 86.9% accuracy versus 86.7% for feature-wise encoding, indicating that discrete 1D graph distance requires only simple shared scalar parameter tables.
      • For Dynamic Topological Encoding, a feature-wise encoding achieves 86.9% accuracy versus 86.5% for shared encoding, demonstrating that continuous 3D Euclidean distances benefit from higher feature capacity.
    2. Barcode Vectorization Feature Dimension (d′d'):

      • d′=32d' = 32: 1.26M parameters, 86.5% accuracy
      • d′=64d' = 64: 1.32M parameters, 86.9% accuracy
      • d′=128d' = 128: 1.44M parameters, 86.4% accuracy Setting d′=64d'=64 provides the optimal trade-off between capacity and over-fitting.
    3. Number of Inserted Encoding Layers (out of 10 total layers):

      • 0 layers: 86.0% accuracy
      • 1 layer: 86.4% accuracy
      • 5 layers: 86.7% accuracy
      • 10 layers: 86.9% accuracy Providing topological encodings repeatedly across all hidden layers consistently benefits classification performance.

Coverage note — None was omitted; all contributed models, theoretical complexity formulations, empirical results, and ablation studies from the paper are represented.

References

  1. 1.Mehmet E Aktas, Esra Akbas, and Ahmed El Fatmaoui. Persistence homology of networks: methods and applications. Applied Network Science, 4(1):1–28, 2019.
  2. 2.Yuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li, Ying Deng, and Weiming Hu. Channel-wise topology refinement graph convolution for skeleton-based action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13359–13368, 2021.
  3. 3.Zhan Chen, Sicheng Li, Bing Yang, Qinghan Li, and Hong Liu. Multi-scale spatial temporal graph convolutional network for skeleton-based action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 1113–1122, 2021.
  4. 4.Ke Cheng, Yifan Zhang, Congqi Cao, Lei Shi, Jian Cheng, and Hanqing Lu. Decoupling gcn with dropgraph module for skeleton-based action recognition. In European Conference on Computer Vision, pages 536–553, 2020.
  5. 5.Hyung-gun Chi, Myoung Hoon Ha, Seunggeun Chi, Sang Wan Lee, Qixing Huang, and Karthik Ramani. Infogcn: Representation learning for human skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20186–20196, 2022.
  6. 6.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019.
  7. 7.Michael Defferrard, Xavier Bresson, and Pierre Van dergheynst. Convolutional neural networks on graphs with fast localized spectral filtering. Advances in Neural Information Processing systems, 29, 2016.
  8. 8.Josep Díaz, Jordi Petit, and Maria Serna. A survey of graph layout problems. ACM Computing Surveys (CSUR), 34(3):313–356, 2002.
  9. 9.Yong Du, Wei Wang, and Liang Wang. Hierarchical recurrent neural network for skeleton based action recognition. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1110–1118, 2015.
  10. 10.Haodong Duan, Yue Zhao, Kai Chen, Dahua Lin, and Bo Dai. Revisiting skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2969–2978, 2022.
  11. 11.Herbert Edelsbrunner. Persistent homology: theory and practice. 2013.
  12. 12.Lingling Gao, Yanli Ji, Yang Yang, and HengTao Shen. Global-local cross-view fisher discrimination for view-invariant action recognition. In Proceedings of the 30th ACM International Conference on Multimedia, pages 5255–5264, 2022.
  13. 13.Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. Neural message passing for quantum chemistry. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, pages 1263–1272, 2017.
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE International Conference on Computer Vision, pages 1026–1034, 2015.
  15. 15.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654, 2020.
  16. 16.Christoph Hofer, Florian Graf, Bastian Rieck, Marc Niethammer, and Roland Kwitt. Graph filtration learning. In International Conference on Machine Learning, pages 4314–4323. PMLR, 2020.
  17. 17.Xiaohu Huang, Hao Zhou, Bin Feng, Xinggang Wang, Wenyu Liu, Jian Wang, Haocheng Feng, Junyu Han, Errui Ding, and Jingdong Wang. Graph contrastive learning for skeleton-based action recognition. arXiv preprint arXiv:2301.10900, 2023.
  18. 18.Qiuhong Ke, Mohammed Bennamoun, Senjian An, Ferdous Sohel, and Farid Boussaid. A new representation of skeleton sequences for 3d action recognition. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4570–4579, 2017.
  19. 19.Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations, 2016.
  20. 20.Jungho Lee, Minhyeok Lee, Dogyoon Lee, and Sangyoun Lee. Hierarchically decomposed graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10444–10453, 2023.
  21. 21.Hao Li, Jinfa Huang, Peng Jin, Guoli Song, Qi Wu, and Jie Chen. Toward 3d spatial reasoning for human-like text-based visual question answering. arXiv preprint arXiv:2209.10326, 2022.
  22. 22.Hao Li, Jinfa Huang, Peng Jin, Guoli Song, Qi Wu, and Jie Chen. Weakly-supervised 3d spatial reasoning for text-based visual question answering. IEEE Transactions on Image Processing, 2023.
  23. 23.Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(10):2684–2701, 2019.
  24. 24.Mengyuan Liu, Hong Liu, and Chen Chen. Enhanced skeleton visualization for view invariant human action recognition. Pattern Recognition, 68(68):346–362, 2017.
  25. 25.Ziyu Liu, Hongwen Zhang, Zhenghao Chen, Zhiyong Wang, and Wanli Ouyang. Disentangling and unifying graph convolutions for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 143–152, 2020.
  26. 26.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021.
  27. 27.German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter. Continual lifelong learning with neural networks: A review. Neural Networks, 113:54–71, 2019.
  28. 28.Wei Peng, Xiaopeng Hong, Haoyu Chen, and Guoying Zhao. Learning graph convolutional network for skeleton-based human action recognition by neural searching. In Proceedings of the AAAI conference on artificial intelligence, pages 2669–2676, 2020.
  29. 29.Katharina Prasse, Steffen Jung, Yuxuan Zhou, and Margret Keuper. Local spherical harmonics improve skeleton-based hand action recognition. arXiv preprint arXiv:2308.10557, 2023.
  30. 30.Bastian Rieck, Christian Bock, and Karsten Borgwardt. A persistent weisfeiler-lehman procedure for graph classification. In International Conference on Machine Learning, pages 5448–5458. PMLR, 2019.
  31. 31.Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne Van Den Berg, Ivan Titov, and Max Welling. Modeling relational data with graph convolutional networks. In The Semantic Web: 15th International Conference, ESWC 2018, Heraklion, Crete, Greece, June 3–7, 2018, Proceedings 15, pages 593–607. Springer, 2018.
  32. 32.Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang. Ntu rgb+ d: A large scale dataset for 3d human activity analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1010–1019, 2016.
  33. 33.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155, 2018.
  34. 34.Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12026–12035, 2019.
  35. 35.Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng, and Jiaying Liu. An end-to-end spatio-temporal attention model for human action recognition from skeleton data. In Proceedings of the AAAI Conference on Artificial Intelligence, 2017.
  36. 36.Yi-Fan Song, Zhang Zhang, Caifeng Shan, and Liang Wang. Constructing stronger and faster baselines for skeleton-based action recognition. arXiv preprint arXiv:2106.15125, 2021.
  37. 37.Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. Mlp-mixer: An all-mlp architecture for vision. arXiv preprint arXiv:2105.01601, 2021.
  38. 38.Hugo Touvron, Piotr Bojanowski, Mathilde Caron, Matthieu Cord, Alaaeldin El-Nouby, Edouard Grave, Armand Joulin, Gabriel Synnaeve, Jakob Verbeek, and Herve Jégou. Resmlp: Feedforward networks for image classification with data-efficient training. arXiv preprint arXiv:2105.03404, 2021.
  39. 39.Petar Velicković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. Graph attention networks. arXiv preprint arXiv:1710.10903, 2017.
  40. 40.Jiang Wang, Xiaohan Nie, Yin Xia, Ying Wu, and Song-Chun Zhu. Cross-view action modeling, learning and recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2649–2656, 2014.
  41. 41.Kan Wu, Houwen Peng, Minghao Chen, Jianlong Fu, and Hongyang Chao. Rethinking and improving relative position encoding for vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10033–10041, 2021.
  42. 42.Hailun Xia and Xinkai Gao. Multi-scale mixed dense graph convolution network for skeleton-based action recognition. IEEE Access, 9:36475–36484, 2021.
  43. 43.Wangmeng Xiang, Chao Li, Yuxuan Zhou, Biao Wang, and Lei Zhang. Language supervised training for skeleton-based action recognition. arXiv preprint arXiv:2208.05318, 2022.
  44. 44.Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks. In International Conference on Learning Representations, 2018.
  45. 45.Sijie Yan, Yuanjun Xiong, and Dahua Lin. Spatial temporal graph convolutional networks for skeleton-based action recognition. In AAAI, pages 7444–7452, 2018.
  46. 46.Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. Do transformers really perform badly for graph representation? Advances in Neural Information Processing Systems, 34:28877–28888, 2021.
  47. 47.Pengfei Zhang, Cuiling Lan, Junliang Xing, Wenjun Zeng, Jianru Xue, and Nanning Zheng. View adaptive recurrent neural networks for high performance human action recognition from skeleton data. In Proceedings of the IEEE International Conference on Computer Vision, pages 2117–2126, 2017.
  48. 48.Pengfei Zhang, Cuiling Lan, Wenjun Zeng, Junliang Xing, Jianru Xue, and Nanning Zheng. Semantics-guided neural networks for efficient skeleton-based human action recognition. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1112–1121, 2020.
  49. 49.Qi Zhao and Yusu Wang. Learning metrics for persistence-based summaries and applications for graph classification. Advances in Neural Information Processing Systems, 32, 2019.
  50. 50.Huanyu Zhou, Qingjie Liu, and Yunhong Wang. Learning discriminative representations for skeleton based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10608–10617, 2023.
  51. 51.Yuxuan Zhou, Zhi-Qi Cheng, Chao Li, Yanwen Fang, Yifeng Geng, Xuansong Xie, and Margret Keuper. Hypergraph transformer for skeleton-based action recognition. arXiv preprint arXiv:2211.09590, 2022.
  52. 52.Yuxuan Zhou, Wangmeng Xiang, Chao Li, Biao Wang, Xihan Wei, Lei Zhang, Margret Keuper, and Xiansheng Hua. Sp-vit: Learning 2d spatial priors for vision transformers. arXiv preprint arXiv:2206.07662, 2022.

Citation

MLA
Zhou, Y., et al. “BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 2049–58, https://doi.org/10.1109/CVPR52733.2024.00200.
APA
Zhou, Y., Yan, X., Cheng, Z.-Q., Yan, Y., Dai, Q., & Hua, X.-S. (2024). BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2049–2058. https://doi.org/10.1109/CVPR52733.2024.00200
Chicago
Zhou, Y., X. Yan, Z.-Q. Cheng, Y. Yan, Q. Dai, and X.-S. Hua. 2024. “BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2049–58. https://doi.org/10.1109/CVPR52733.2024.00200.
Harvard
Zhou, Y. et al. (2024) “BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition”, 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 2049–2058. Available at: https://doi.org/10.1109/CVPR52733.2024.00200.
Vancouver
1. Zhou Y, Yan X, Cheng Z-Q, Yan Y, Dai Q, Hua X-S (2024) BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 2049–2058

BibTeX

@inproceedings{Zhou_2024, title={BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition}, url={http://dx.doi.org/10.1109/CVPR52733.2024.00200}, DOI={10.1109/cvpr52733.2024.00200}, booktitle={2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhou, Yuxuan and Yan, Xudong and Cheng, Zhi-Qi and Yan, Yan and Dai, Qi and Hua, Xian-Sheng}, year={2024}, month=June, pages={2049–2058} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE