NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding

Jun LiuAmir ShahroudyMauricio PerezGang WangLing-Yu DuanAlex C. Kot

article2019TPAMI1,792 citations

Presents NTU RGB+D 120, a large-scale benchmark of over 114,000 video samples across 120 classes and 106 subjects, establishing new standards for training deep learning models and evaluating one-shot 3D human action recognition.

Listen

Three-dimensional human activity analysis is essential for computer vision applications in surveillance, healthcare, and human-machine interaction. However, the field has been constrained by existing benchmarks that lack scale, subject variety, realistic action classes, and environmental diversity. Unlike standard video platforms, three-dimensional and depth-sensing data cannot be easily harvested from public online sources, which causes advanced data-driven models like deep neural networks to suffer from severe overfitting.

The article introduces NTU RGB+D 120, a large-scale benchmark designed to train and rigorously evaluate modern activity understanding models, and presents a new semantic framework for one-shot action recognition.

To construct the benchmark, the researchers recorded 114,480 video clips totaling over 8 million frames across four modalities: color video, depth maps, 3D skeleton joints, and infrared sequences. The collection spans 106 subjects aged 10 to 57 from 15 countries across 32 physical setups, 96 backgrounds, and 155 camera viewpoints. The article defines standardized cross-subject and cross-setup evaluation protocols, evaluates twelve existing deep learning and feature-based architectures, measures the impact of data volume and modality combinations, and evaluates an Action-Part Semantic Relevance-aware (APSR) framework that matches action descriptions to body parts via language embeddings.

The experimental findings demonstrate that deep neural models scale effectively with large datasets; for example, increasing training data from 20% to 100% boosted model accuracy from 40.6% to 62.4%. Among evaluated algorithms, the Body Pose Evolution Map achieved top performance at 64.6% cross-subject and 66.9% cross-setup accuracy. Modality fusion yielded substantial benefits, where combining color, depth, and skeleton data delivered 64.0% cross-subject accuracy compared to 58.5% for color alone, 48.7% for depth alone, and 55.7% for skeleton sequences alone. Furthermore, the proposed APSR framework attained 45.3% accuracy in one-shot recognition, outperforming conventional averaging and attention baselines.

These results confirm that scalable 3D action recognition requires both massive training samples and multimodal sensing. Modalities serve complementary roles: skeleton inputs provide view-invariant geometry across camera angles, whereas color and depth modalities provide crucial context for distinguishing object-heavy interactions that skeleton tracking alone confuses. The success of APSR indicates that language-guided priors can significantly reduce the costs of collecting data for rare or novel actions.

Organizations developing computer vision systems should adopt multimodal sensor configurations for robust performance and use the large-scale benchmark for network pre-training before fine-tuning on specialized target domains. Development teams should also explore semantic word embedding architectures to support rapid deployment on novel action classes without exhaustive re-annotation.

The study's results reflect laboratory-based sensor recordings with Microsoft Kinect v2, meaning performance may vary under consumer-grade hardware or unstructured real-world conditions. Furthermore, noisy skeleton tracking on fine-grained hand motions and persistent confusion between mirror-image actions (such as putting on versus taking off a shoe) remain technical challenges requiring targeted modeling advancements.

Cover for NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding

Abstract

Research on depth-based human activity analysis achieved outstanding performance and demonstrated the effectiveness of 3D representation for action recognition. The existing depth-based and RGB+D-based action recognition benchmarks have a number of limitations, including the lack of large-scale training samples, realistic number of distinct class categories, diversity in camera views, varied environmental conditions, and variety of human subjects. In this work, we introduce a large-scale dataset for RGB+D human action recognition, which is collected from 106 distinct subjects and contains more than 114 thousand video samples and 8 million frames. This dataset contains 120 different action classes including daily, mutual, and health-related activities. We evaluate the performance of a series of existing 3D activity analysis methods on this dataset, and show the advantage of applying deep learning methods for 3D-based human action recognition. Furthermore, we investigate a novel one-shot 3D activity recognition problem on our dataset, and a simple yet effective Action-Part Semantic Relevance-aware (APSR) framework is proposed for this task, which yields promising results for recognition of the novel action classes. We believe the introduction of this large-scale dataset will enable the community to apply, adapt, and develop various data-hungry learning techniques for depth-based and RGB+D-based human activity understanding. [The dataset is available at: this http URL]

Table of Contents

  • I Introduction
  • II Related work
  • II-A 3D Activity Analysis Datasets
  • II-B 3D Action Recognition Methods
  • III The NTU RGB+D 120 Dataset
  • III-A Dataset Structure
  • III-A1 Data Modalities
  • III-A2 Action Classes
  • III-A3 Subjects
  • III-A4 Collection Setups
  • III-B Benchmark Evaluations
  • III-B1 Cross-Subject Evaluation
  • III-B2 Cross-Setup Evaluation
  • IV APSR Framework for One-Shot 3D Action Recognition
  • IV-A One-Shot Recognition on NTU RGB+D 120
  • IV-B APSR Framework
  • V Experiments
  • V-A Experimental Evaluations of 3D Action Recognition
  • V-A1 Evaluation of state-of-the-art methods
  • V-A2 Evaluation of using different data modalities
  • V-A3 Evaluation of using different sizes of training set
  • V-A4 Detailed analysis according to data modalities
  • V-A5 Detailed analysis according to methods
  • V-B Experimental Evaluations of One-Shot Recognition
  • V-C Discussions
  • VI Conclusion
  • References

Knowls

  1. Knowl 1 — NTU RGB+D 120 Dataset Specification and Structure

    experimental setup

    The NTU RGB+D 120 dataset is a large-scale multimodal benchmark designed for 3D human activity understanding. It comprises 114,480 video samples spanning 120 distinct action classes performed by 106 subjects from 15 different countries.

    Key characteristics of the dataset include:

    • Subjects: 106 distinct human subjects aged between 10 and 57 years, with heights ranging from 1.3 m to 1.9 m.
    • Action Classes: 120 categories categorized into three broad groups:
      • 82 daily actions (e.g., eating, writing, sitting down, reading, moving objects).
      • 12 health-related actions (e.g., blowing nose, vomiting, staggering, falling down).
      • 26 two-person mutual actions (e.g., handshaking, pushing, hitting, hugging, wielding a knife towards another person).
    • Data Modalities (captured via Microsoft Kinect v2 sensors):
      1. RGB Video: recorded at 1920×10801920 \times 1080 pixel resolution.
      2. Depth Maps: sequences of 2D depth values in millimeters stored losslessly at 512×424512 \times 424 pixel resolution.
      3. 3D Joint Sequences (Skeleton): 3D spatial (X,Y,Z)(X, Y, Z) coordinates of 25 major body joints per tracked human body per frame, along with mapped 2D pixel coordinates on RGB and depth frames.
      4. Infrared (IR) Sequences: captured frame-by-frame at 512×424512 \times 424 pixel resolution.
    • Camera Configurations: 32 collection setups across 96 different indoor environmental backgrounds with significant illumination variation. In each setup, 3 simultaneous horizontal camera viewpoints (45-45^\circ, 00^\circ, +45+45^\circ) record actions performed facing left and right cameras (yielding 5 unique horizontal perspectives), with camera heights varying from 0.5 m to 2.7 m and camera-to-subject distances varying from 2.0 m to 4.5 m, totaling 155 distinct camera viewpoints.
  2. Knowl 2 — Cross-Subject and Cross-Setup Evaluation Protocols for NTU RGB+D 120

    experimental setup

    To provide standardized evaluation criteria on the NTU RGB+D 120 dataset, two distinct evaluation protocols are defined:

    1. Cross-Subject (X-Sub) Evaluation: The 106 subjects are split into two balanced groups of 53 subjects each for training and testing.

      • Training Subject IDs (53 subjects): 1, 2, 4, 5, 8, 9, 13, 14, 15, 16, 17, 18, 19, 25, 27, 28, 31, 34, 35, 38, 45, 46, 47, 49, 50, 52, 53, 54, 55, 56, 57, 58, 59, 70, 74, 78, 80, 81, 82, 83, 84, 85, 86, 89, 91, 92, 93, 94, 95, 97, 98, 100, 103.
      • Testing Subject IDs (53 subjects): The remaining 53 subjects.
    2. Cross-Setup (X-Setup) Evaluation: The 32 capture setups (varying camera height, distance, and environmental background) are split evenly based on setup identifiers:

      • Training Setups: All samples from the 16 even-numbered collection setup IDs (i.e., setups 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32).
      • Testing Setups: All samples from the 16 odd-numbered collection setup IDs (i.e., setups 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31).

    For both evaluation protocols, performance is measured as classification accuracy (percentage of correctly classified action clips).

  3. Knowl 3 — Action-Part Semantic Relevance-aware (APSR) Framework for One-Shot 3D Action Recognition

    model/method

    The Action-Part Semantic Relevance-aware (APSR) framework performs one-shot 3D skeleton action recognition on novel (unseen) action categories by incorporating semantic prior knowledge regarding the relevance of specific human body parts to specific action classes.

    The framework operates under a split-dataset setting:

    • Auxiliary Dataset: Contains video samples from known action classes used to train the feature extractor without any overlap with novel evaluation classes.
    • Evaluation Dataset: Contains novel action classes; for each novel category, exactly one video sample serves as the reference exemplar, and all remaining samples serve as query test samples.

    APSR addresses the inability of standard attention or data-driven weight-learning mechanisms to generalize to unseen action classes by:

    1. Deriving semantic relevance weights between textual descriptions of body parts and action class names using pre-trained distributed word embeddings.
    2. Utilizing a bidirectional 2D Spatio-Temporal LSTM (ST-LSTM) network to extract spatio-temporal features for every individual body part across all video frames.
    3. Weighting the training classification loss across body part units using their semantic relevance scores to focus feature learning on informative joints.
    4. Performing semantic relevance-weighted spatial-temporal feature pooling for exemplar matching during inference using cosine distance in the pooled feature space.
  4. Knowl 4 — Semantic Relevance Estimation and Normalization in the APSR Framework

    model/method

    In the Action-Part Semantic Relevance-aware (APSR) framework, prior knowledge regarding which body parts are informative for recognizing an action class cc is computed using continuous distributed word vector representations (Word2Vec).

    Let ecRd\mathbf{e}_c \in \mathbb{R}^d be the embedding vector for the text description of action class cc, obtained by averaging the dd-dimensional Word2Vec embeddings (d=300d = 300) of each word in the action name. Similarly, let epRd\mathbf{e}_p \in \mathbb{R}^d be the average word embedding for body part p{1,,P}p \in \{1, \dots, P\}, where PP is the total number of body joints (e.g., P=25P = 25).

    The unnormalized semantic relevance score rc,pr_{c,p} between action class cc and body part pp is defined as the rectified cosine similarity: rc,p=max(0,ecepec2ep2)r_{c,p} = \max\left(0, \frac{\mathbf{e}_c \cdot \mathbf{e}_p}{\|\mathbf{e}_c\|_2 \|\mathbf{e}_p\|_2}\right)

    The relevance scores are subsequently normalized across all PP body parts so that they sum to unity for each action category: sc,p=rc,pu=1Prc,us_{c,p} = \frac{r_{c,p}}{\sum_{u=1}^P r_{c,u}}

    The resulting vector Sc={sc,1,sc,2,,sc,P}S_c = \{s_{c,1}, s_{c,2}, \dots, s_{c,P}\} defines the action-part semantic relevance distribution for action class cc, satisfying p=1Psc,p=1\sum_{p=1}^P s_{c,p} = 1 and sc,p0s_{c,p} \ge 0 for all pp.

  5. Knowl 5 — Feature Generation Network and Semantic-Weighted Loss in APSR

    model/method

    In the Action-Part Semantic Relevance-aware (APSR) framework, body part features are extracted using a bidirectional 2D Spatio-Temporal Long Short-Term Memory (ST-LSTM) network that models context dependencies across temporal frames t{1,,T}t \in \{1, \dots, T\} and spatial body parts p{1,,P}p \in \{1, \dots, P\}.

    For each frame tt and body joint pp, the input is the 3D joint coordinate vector, and the network outputs a feature vector Fp,tRDF_{p,t} \in \mathbb{R}^D, yielding the full spatio-temporal feature set F={Fp,tp{1,,P},t{1,,T}}\mathcal{F} = \{F_{p,t} \mid p \in \{1, \dots, P\}, t \in \{1, \dots, T\}\}.

    During training on the auxiliary action dataset, each local spatio-temporal unit (p,t)(p, t) is connected to a shared softmax classifier predicting the action class probability distribution c^p,t\hat{c}_{p,t}. The network parameters are optimized end-to-end using a semantic-weighted negative log-likelihood loss: L=p=1Pt=1Tsc,pl(c,c^p,t)L = \sum_{p=1}^P \sum_{t=1}^T s_{c,p} \, l(c, \hat{c}_{p,t}) where cc is the ground-truth action class label, l(c,c^p,t)l(c, \hat{c}_{p,t}) is the negative log-likelihood loss at unit (p,t)(p, t), and sc,ps_{c,p} is the normalized semantic relevance weight of body part pp for action class cc. This formulation forces the network to penalize errors more heavily on the body parts that are semantically salient to the target action.

  6. Knowl 6 — One-Shot Action Inference via Semantic Relevance-Weighted Exemplar Matching

    algorithm

    In the Action-Part Semantic Relevance-aware (APSR) framework, classification of an unlabeled query video instance ii against a set of novel action class exemplars {Ωk}k=1K\{\Omega_k\}_{k=1}^K is performed via semantic relevance-weighted feature pooling and nearest neighbor search.

    Input: Query skeleton video instance ii with extracted features {Fp,t(i)}p=1,t=1P,T\{F_{p,t}^{(i)}\}_{p=1,t=1}^{P,T}; set of KK novel class exemplars {Ωk}k=1K\{\Omega_k\}_{k=1}^K with corresponding class labels cΩkc_{\Omega_k} and extracted features {Fp,t(Ωk)}p=1,t=1P,T\{F_{p,t}^{(\Omega_k)}\}_{p=1,t=1}^{P,T}; normalized semantic relevance vectors {ScΩk}k=1K\{S_{c_{\Omega_k}}\}_{k=1}^K where Sc={sc,p}p=1PS_{c} = \{s_{c,p}\}_{p=1}^P.
    Output: Predicted class label cc^* for query sample ii.
    for each exemplar Ωk\Omega_k where k=1,,Kk = 1, \dots, K do
        Compute aggregated exemplar representation:
        f(Ωk,ScΩk)=p=1Pt=1TscΩk,pFp,t(Ωk)f(\Omega_k, S_{c_{\Omega_k}}) = \sum_{p=1}^P \sum_{t=1}^T s_{c_{\Omega_k}, p} F_{p,t}^{(\Omega_k)}
        Compute class-conditioned query representation:
        f(i,ScΩk)=p=1Pt=1TscΩk,pFp,t(i)f(i, S_{c_{\Omega_k}}) = \sum_{p=1}^P \sum_{t=1}^T s_{c_{\Omega_k}, p} F_{p,t}^{(i)}
        Compute cosine distance between query and exemplar:
        D(i,Ωk)=1f(i,ScΩk)f(Ωk,ScΩk)f(i,ScΩk)2f(Ωk,ScΩk)2D(i, \Omega_k) = 1 - \frac{f(i, S_{c_{\Omega_k}}) \cdot f(\Omega_k, S_{c_{\Omega_k}})}{\|f(i, S_{c_{\Omega_k}})\|_2 \|f(\Omega_k, S_{c_{\Omega_k}})\|_2}
    end for
    Identify closest exemplar index k=argmink{1,,K}D(i,Ωk)k^* = \arg\min_{k \in \{1, \dots, K\}} D(i, \Omega_k)
    return c=cΩkc^* = c_{\Omega_{k^*}}

    Because novel action categories lack training samples to learn attention distributions, conditioning both the exemplar and the query feature aggregation on the novel class's semantic relevance weights ScΩkS_{c_{\Omega_k}} ensures that matching focuses on joints semantically pertinent to that specific target category.

  7. Knowl 7 — Benchmark Evaluation of 3D Action Recognition Methods on NTU RGB+D 120

    data/table

    Twelve representative 3D action recognition algorithms (spanning hand-crafted features, recurrent architectures, and convolutional networks) were evaluated under the standardized Cross-Subject (X-Sub) and Cross-Setup (X-Setup) protocols on the NTU RGB+D 120 benchmark dataset.

    Method Cross-Subject Accuracy (%) Cross-Setup Accuracy (%)
    Part-Aware LSTM 25.5 26.3
    Soft RNN 36.3 44.9
    Dynamic Skeleton 50.8 54.7
    Spatio-Temporal LSTM 55.7 57.9
    Internal Feature Fusion 58.2 60.9
    GCA-LSTM 58.3 59.2
    Multi-Task Learning Network 58.4 57.9
    FSNet 59.9 62.4
    Skeleton Visualization (Single Stream) 60.3 63.2
    Two-Stream Attention LSTM 61.2 63.3
    Multi-Task CNN with RotClips 62.2 61.8
    Body Pose Evolution Map 64.6 66.9

    The results demonstrate that:

    1. Deep convolutional network architectures that map skeleton sequences to image-like representations (e.g., Body Pose Evolution Map at 64.6% X-Sub / 66.9% X-Setup and Multi-Task CNN with RotClips at 62.2% / 61.8%) and attention-augmented recurrent models (e.g., Two-Stream Attention LSTM at 61.2% / 63.3%) achieve top performance.
    2. For almost all skeleton-based methods, cross-setup accuracy is higher than cross-subject accuracy because 3D skeleton coordinates provide view-invariant geometric configurations that generalize well across varied camera angles and environments, whereas subject biometric diversity introduces stronger intra-class variance.
  8. Knowl 8 — Multimodal Action Recognition and Modality Fusion on NTU RGB+D 120

    data/table

    Action recognition performance across individual sensor modalities (RGB video, depth video, and 3D skeleton sequences) and their multi-modal fusion combinations was evaluated on the NTU RGB+D 120 benchmark under both Cross-Subject (X-Sub) and Cross-Setup (X-Setup) protocols.

    Data Modality Cross-Subject Accuracy (%) Cross-Setup Accuracy (%)
    RGB Video 58.5 54.8
    Depth Video 48.7 40.1
    3D Skeleton Sequence 55.7 57.9
    RGB Video + Depth Video 61.9 59.2
    RGB Video + 3D Skeleton Sequence 61.2 63.1
    Depth Video + 3D Skeleton Sequence 59.2 61.2
    RGB Video + Depth Video + 3D Skeleton Sequence 64.0 66.1

    Key findings from modality comparisons:

    1. Modality Generalization Disparity: RGB and Depth modalities exhibit lower accuracy on Cross-Setup than Cross-Subject (RGB drops from 58.5% to 54.8%; Depth drops from 48.7% to 40.1%) because appearance and depth maps are sensitive to camera height/distance changes and background shifts. In contrast, 3D Skeleton sequences achieve higher Cross-Setup accuracy (57.9%) than Cross-Subject (55.7%) due to the inherent view-invariance of 3D joint coordinates.
    2. Complementarity and Fusion Gains: Fusing all three modalities achieves the highest recognition performance (64.0% X-Sub and 66.1% X-Setup), showing that visual appearance (color, texture), 3D surface geometry, and articulated skeleton dynamics provide complementary cues.
  9. Knowl 9 — Benchmark Results for One-Shot 3D Action Recognition and Auxiliary Set Scaling

    data/table

    One-shot 3D skeleton action recognition was evaluated on the NTU RGB+D 120 dataset using an auxiliary set of 100 base action classes for feature generation training and an evaluation set of 20 unseen novel action classes (with 1 exemplar per novel class).

    Comparing one-shot recognition methods on the 20 novel classes:

    Method Evaluation Accuracy (%)
    Average Pooling 42.9
    Fully Connected 42.1
    Attention Network 41.0
    APSR (Action-Part Semantic Relevance-aware) 45.3

    Evaluating the impact of auxiliary training set scale on APSR performance:

    Auxiliary Training Set Evaluation Accuracy (%)
    # Training Samples # Training Classes
    19,000 20 29.1
    38,000 40 34.8
    57,000 60 39.2
    76,000 80 42.8
    95,000 100 45.3

    Key takeaways:

    1. Data-driven attention mechanisms (Attention Network: 41.0%) underperform uniform Average Pooling (42.9%) on unseen classes because attention weights overfit to the base action classes. APSR uses semantic word embedding guidance to assign appropriate joint importance, outperforming all baselines (45.3%).
    2. Scaling the auxiliary training set size from 20 classes (19,000 samples) to 100 classes (95,000 samples) steadily increases one-shot accuracy from 29.1% to 45.3%, demonstrating the necessity of large-scale pre-training data for generalization.
  10. Knowl 10 — Modality-Specific Capabilities and Error Patterns in 3D Human Activity Analysis

    empirical result

    Empirical analysis of action-specific recognition accuracies and confusion matrices on the NTU RGB+D 120 dataset reveals characteristic advantages and limitations across modalities:

    1. Object Interactions vs. Geometric Posture:

      • Skeleton-only: Struggles with object-involved actions sharing similar bodily kinematics (e.g., confusing "play magic cube" with "counting money", and "touch other person's pocket (steal)" with "grab other person's stuff") because joint sensors do not encode object presence or identity.
      • RGB/Depth: Reliably identifies object-related actions (e.g., "carry things with other person" and "move heavy objects" are among the top 10 accurate classes) but confuses actions involving identical or visually similar objects (e.g., "ball up paper" vs. "fold paper", "reading" vs. "writing").
      • RGB vs. Depth distinction: Depth excels where 3D geometry disambiguates objects ("open a box" vs. "fold paper"), whereas RGB excels when color/texture provides discriminative cues ("put on jacket" vs. "put on bag/backpack").
    2. Fine-Grained Hand and Finger Motions: Skeleton data frequently confuses subtle hand gestures (e.g., "make ok sign", "make victory sign", "thumb up") because the skeleton model captures only three joints per hand and tracker estimation in distal extremities is noisy.

    3. Motion Speed and 3D Trajectory Orientation:

      • Skeleton and depth modalities accurately distinguish 3D directional kinematics (e.g., "kick backward" vs. "side kick"), whereas 2D RGB models confuse them due to perspective projection.
      • Frame-subsampled recurrent models (e.g., GCA-LSTM) lose speed cues, confusing "touch pocket (slow)" with "grab stuff (fast)", whereas temporal convolutional networks (e.g., FSNet) over all frames preserve speed discrimination.
    4. Symmetric Action Inversion: Actions with near-identical posture and object contexts differing primarily in temporal sequence direction (e.g., "take off a shoe" vs. "put on a shoe") exhibit high confusion across single modalities (52% RGB error, 65% depth error, 39% skeleton error), which reduces to 32% error only when fusing all three modalities.

Coverage note — Summaries of prior external datasets and standard external baseline network architectures were omitted to focus entirely on the contributed NTU RGB+D 120 dataset, evaluation protocols, the APSR one-shot framework, and comprehensive empirical analyses.

References

  1. 1.J. Han, L. Shao, D. Xu, and J. Shotton, “Enhanced computer vision with microsoft kinect sensor: A review,” IEEE Transactions on Cybernetics, 2013.
  2. 2.Y. Guo, M. Bennamoun, F. Sohel, M. Lu, and J. Wan, “3d object recognition in cluttered scenes with local surface features: a survey,” TPAMI, 2014.
  3. 3.S. Gupta, P. Arbelaez, R. Girshick, and J. Malik, “Indoor scene un- ´ derstanding with rgb-d images: Bottom-up segmentation, object detection and semantic segmentation,” IJCV, 2015.
  4. 4.J. Aggarwal and L. Xia, “Human activity recognition from 3d data: A review,” PR Letters, 2014.
  5. 5.H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “Hmdb: a large video database for human motion recognition,” in ICCV, 2011.
  6. 6.K. Soomro, A. R. Zamir, and M. Shah, “Ucf101: A dataset of 101 human actions classes from videos in the wild,” arXiv, 2012.
  7. 7.F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in CVPR, 2015.
  8. 8.W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev et al., “The kinetics human action video dataset,” arXiv, 2017.
  9. 9.Y. Du, W. Wang, and L. Wang, “Hierarchical recurrent neural network for skeleton based action recognition,” in CVPR, 2015.
  10. 10.P. Wang, W. Li, Z. Gao, J. Zhang, C. Tang, and P. Ogunbona, “Action recognition from depth maps using deep convolutional neural networks,” in THMS, 2015.
  11. 11.W. Li, Z. Zhang, and Z. Liu, “Action recognition based on a bag of 3d points,” in CVPR Workshops, 2010.
  12. 12.J. Sung, C. Ponce, B. Selman, and A. Saxena, “Human activity detection from rgbd images,” in AAAI Workshops, 2011.
  13. 13.B. Ni, G. Wang, and P. Moulin, “Rgbd-hudaact: A color-depth video database for human daily activity recognition,” in ICCV Workshops, 2011.
  14. 14.J. Wang, Z. Liu, Y. Wu, and J. Yuan, “Mining actionlet ensemble for action recognition with depth cameras,” in CVPR, 2012.
  15. 15.L. Xia, C.-C. Chen, and J. Aggarwal, “View invariant human action recognition using histograms of 3d joints,” in CVPR Workshops, 2012.
  16. 16.Z. Cheng, L. Qin, Y. Ye, Q. Huang, and Q. Tian, “Human daily action analysis with multi-view and color-depth data,” in ECCV Workshops, 2012.
  17. 17.H. S. Koppula, R. Gupta, and A. Saxena, “Learning human activities and object affordances from rgb-d videos,” IJRR, 2013.
  18. 18.O. Oreifej and Z. Liu, “Hon4d: Histogram of oriented 4d normals for activity recognition from depth sequences,” in CVPR, 2013.
  19. 19.P. Wei, Y. Zhao, N. Zheng, and S.-C. Zhu, “Modeling 4d humanobject interactions for event and object recognition,” in ICCV, 2013.
  20. 20.J. Wang, X. Nie, Y. Xia, Y. Wu, and S.-C. Zhu, “Cross-view action modeling, learning, and recognition,” in CVPR, 2014.
  21. 21.H. Rahmani, A. Mahmood, D. Q Huynh, and A. Mian, “Hopc: Histogram of oriented principal components of 3d pointclouds for action recognition,” in ECCV, 2014.
  22. 22.K. Wang, X. Wang, L. Lin, M. Wang, and W. Zuo, “3d human activity recognition with reconfigurable convolutional neural networks,” in ACM MM, 2014.
  23. 23.C. Chen, R. Jafari, and N. Kehtarnavaz, “Utd-mhad: A multimodal dataset for human action recognition utilizing a depth camera and a wearable inertial sensor,” in ICIP, 2015.
  24. 24.H. Rahmani, A. Mahmood, D. Huynh, and A. Mian, “Histogram of oriented principal components for cross-view action recognition,” TPAMI, 2016.
  25. 25.N. Xu, A. Liu, W. Nie, Y. Wong, F. Li, and Y. Su, “Multi-modal & multi-view & interactive benchmark dataset for human action recognition,” in ACM MM, 2015.
  26. 26.J.-F. Hu, W.-S. Zheng, J. Lai, and J. Zhang, “Jointly learning heterogeneous features for rgb-d activity recognition.” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 11, pp. 2186–2200, 2017.
  27. 27.H. Zamani and W. B. Croft, “Relevance-based word embedding,” in SIGIR, 2017.
  28. 28.T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in NIPS, 2013.
  29. 29.M. Ye, Q. Zhang, L. Wang, J. Zhu, R. Yang, and J. Gall, “A survey on human motion analysis from depth data,” in Time-of-flight and depth imaging. sensors, algorithms, and applications, 2013.
  30. 30.J. Liu, G. Wang, P. Hu, L.-Y. Duan, and A. C. Kot, “Global contextaware attention lstm networks for 3d action recognition,” in CVPR, 2017.
  31. 31.J. Zhang, W. Li, P. O. Ogunbona, P. Wang, and C. Tang, “Rgb-dbased action recognition datasets: A survey,” PR, 2016.
  32. 32.P. Wang, W. Li, P. Ogunbona, J. Wan, and S. Escalera, “Rgb-dbased human motion recognition with deep learning: A survey,” CVIU, 2018.
  33. 33.L. L. Presti and M. La Cascia, “3d skeleton-based human action classification: A survey,” PR, 2016.
  34. 34.L. Chen, H. Wei, and J. Ferryman, “A survey of human motion analysis using depth imagery,” PR Letters, 2013.
  35. 35.F. Han, B. Reily, W. Hoff, and H. Zhang, “Space-time representation of people based on 3d skeletal data: A review,” CVIU, 2017.
  36. 36.R. Lun and W. Zhao, “A survey of applications and human motion recognition with microsoft kinect,” IJPRAI, 2015.
  37. 37.Z. Zhang, “Microsoft kinect sensor and its effect,” IEEE MultiMedia, 2012.
  38. 38.A. Shahroudy, T. T. Ng, Q. Yang, and G. Wang, “Multimodal multipart learning for action recognition in depth videos,” TPAMI, 2016.
  39. 39.A. B. Tanfous, H. Drira, and B. B. Amor, “Coding kendall’s shape trajectories for 3d action recognition,” in CVPR, 2018.
  40. 40.C. Lu, J. Jia, and C.-K. Tang, “Range-sample depth feature for action recognition,” in CVPR, 2014.
  41. 41.J.-F. Hu, W.-S. Zheng, J. Lai, and J. Zhang, “Jointly learning heterogeneous features for rgb-d activity recognition,” in CVPR, 2015.
  42. 42.J. Luo, W. Wang, and H. Qi, “Group sparsity and geometry constrained dictionary learning for action recognition from depth maps,” in ICCV, 2013.
  43. 43.A. Shahroudy, T.-T. Ng, Y. Gong, and G. Wang, “Deep multimodal feature analysis for action recognition in rgb+d videos,” TPAMI, 2018.
  44. 44.Y. Kong and Y. Fu, “Bilinear heterogeneous information machine for rgb-d action recognition,” in CVPR, 2015.
  45. 45.V. Bloom, D. Makris, and V. Argyriou, “G3d: A gaming action dataset and real time action recognition evaluation framework,” in CVPR Workshops, 2012.
  46. 46.C. Liu, Y. Hu, Y. Li, S. Song, and J. Liu, “Pku-mmd: A large scale benchmark for continuous multi-modal human action understanding,” arXiv, 2017.
  47. 47.A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+d: A large scale dataset for 3d human activity analysis,” in CVPR, 2016.
  48. 48.J. Liu, A. Shahroudy, D. Xu, and G. Wang, “Spatio-temporal lstm with trust gates for 3d human action recognition,” in ECCV, 2016.
  49. 49.Q. Ke, M. Bennamoun, S. An., F. Sohel, and F. Boussaid, “A new representation of skeleton sequences for 3d action recognition,” in CVPR, 2017.
  50. 50.X. Yang and Y. Tian, “Super normal vector for activity recognition using depth sequences,” in CVPR, 2014.
  51. 51.J. Shotton, A. W. Fitzgibbon, M. Cook, T. Sharp, M. Finocchio, R. Moore, A. Kipman, and A. Blake, “Real-time human pose recognition in parts from single depth images.” in CVPR, 2011.
  52. 52.J. Liu, H. Ding, A. Shahroudy, L.-Y. Duan, X. Jiang, G. Wang, and A. C. Kot, “Feature boosting network for 3d pose estimation,” TPAMI, 2019.
  53. 53.G. Evangelidis, G. Singh, and R. Horaud, “Skeletal quads: Human action recognition using joint quadruples,” in ICPR, 2014.
  54. 54.R. Vemulapalli, F. Arrate, and R. Chellappa, “Human action recognition by representing 3d skeletons as points in a lie group,” in CVPR, 2014.
  55. 55.E. Ohn-Bar and M. Trivedi, “Joint angles similarities and hog2 for action recognition,” in CVPR Workshops, 2013.
  56. 56.J. Wang, Z. Liu, Y. Wu, and J. Yuan, “Learning actionlet ensemble for 3d human action recognition,” TPAMI, 2014.
  57. 57.A. Shahroudy, G. Wang, and T.-T. Ng, “Multi-modal feature fusion for action recognition in rgb-d sequences,” in ISCCSP, 2014.
  58. 58.H. Rahmani, A. Mahmood, D. Q. Huynh, and A. Mian, “Real time action recognition using histograms of depth gradients and random decision forests,” in WACV, 2014.
  59. 59.V. Veeriah, N. Zhuang, and G.-J. Qi, “Differential recurrent neural networks for action recognition,” in ICCV, 2015.
  60. 60.W. Zhu, C. Lan, J. Xing, W. Zeng, Y. Li, L. Shen, and X. Xie, “Cooccurrence feature learning for skeleton based action recognition using regularized deep lstm networks,” AAAI, 2016.
  61. 61.Z. Luo, B. Peng, D.-A. Huang, A. Alahi, and L. Fei-Fei, “Unsupervised learning of long-term motion dynamics for videos,” in CVPR, 2017.
  62. 62.Z. Huang, C. Wan, T. Probst, and L. Van Gool, “Deep learning on lie groups for skeleton-based action recognition,” in CVPR, 2017.
  63. 63.M. Zolfaghari, G. L. Oliveira, N. Sedaghat, and T. Brox, “Chained multi-stream networks exploiting pose, motion, and appearance for action classification and detection,” in ICCV, 2017.
  64. 64.Q. Ke, J. Liu, M. Bennamoun, H. Rahmani, S. An, F. Sohel, and F. Boussaid, “Global regularizer and temporal-aware crossentropy for skeleton-based early action recognition,” in ACCV, 2018.
  65. 65.T. S. Kim and A. Reiter, “Interpretable 3d human action analysis with temporal convolutional networks,” in CVPR Workshops, 2017.
  66. 66.H. Rahmani and M. Bennamoun, “Learning action recognition model from depth and skeleton videos,” in ICCV, 2017.
  67. 67.D. C. Luvizon, D. Picard, and H. Tabia, “2d/3d pose estimation and action recognition using multitask deep learning,” in CVPR, 2018.
  68. 68.J. Cavazza, P. Morerio, and V. Murino, “When kernel methods meet feature learning: Log-covariance network for action recognition from skeletal data,” in CVPR Workshops, 2017.
  69. 69.F. Baradel, C. Wolf, and J. Mille, “Pose-conditioned spatio-temporal attention for human action recognition,” arXiv:1703.10106, 2017.
  70. 70.F. Baradel, C. Wolf, J. Mille, and G. W. Taylor, “Glimpse clouds: Human activity recognition from unstructured feature points,” CVPR, 2018.
  71. 71.J. Wang, A. Cherian, F. Porikli, and S. Gould, “Video representation learning using discriminative pooling,” in CVPR, 2018.
  72. 72.J.-F. Hu, W.-S. Zheng, J. Pan, J. Lai, and J. Zhang, “Deep bilinear learning for rgb-d action recognition,” in ECCV, 2018.
  73. 73.S. Zhang, Y. Yang, J. Xiao, X. Liu, Y. Yang, D. Xie, and Y. Zhuang, “Fusing geometric features for skeleton-based action recognition using multilayer lstm networks,” TMM, 2018.
  74. 74.B. Zhang, J. Han, Z. Huang, J. Yang, and X. Zeng, “A real-time and hardware-efficient processor for skeleton-based action recognition with lightweight convolutional neural network,” TCS-II, 2019.
  75. 75.Q. Ke, J. Liu, M. Bennamoun, S. An, F. Sohel, and F. Boussaid, “Computer vision for human–machine interaction,” in Computer Vision for Assistive Healthcare, 2018.
  76. 76.M. Liu, C. Chen, and H. Liu, “3d action recognition using data visualization and convolutional neural networks,” in ICME, 2017.
  77. 77.Y. Tang, Y. Tian, J. Lu, P. Li, and J. Zhou, “Deep progressive reinforcement learning for skeleton-based action recognition,” in CVPR, 2018.
  78. 78.H. Wang and L. Wang, “Modeling temporal dynamics and spatial configurations of actions using two-stream recurrent neural networks,” in CVPR, 2017.
  79. 79.Z. Shi and T.-K. Kim, “Learning and refining of privileged information-based rnns for action recognition from depth sequences,” in CVPR, 2017.
  80. 80.P. Zhang, C. Lan, J. Xing, W. Zeng, J. Xue, and N. Zheng, “View adaptive recurrent neural networks for high performance human action recognition from skeleton data,” in ICCV, 2017.
  81. 81.P. Wang, W. Li, Z. Gao, Y. Zhang, C. Tang, and P. Ogunbona, “Scene flow to action map: A new representation for rgb-d based action recognition with convolutional neural networks,” in CVPR, 2017.
  82. 82.Q. Ke, S. An, M. Bennamoun, F. Sohel, and F. Boussaid, “Skeletonnet: Mining deep part features for 3-d action recognition,” SPL, 2017.
  83. 83.H. Rahmani and A. Mian, “Learning a non-linear knowledge transfer model for cross-view action recognition,” in CVPR, 2015.
  84. 84.G. Koch, R. Zemel, and R. Salakhutdinov, “Siamese neural networks for one-shot image recognition,” in ICML, 2015.
  85. 85.S. Ravi and H. Larochelle, “Optimization as a model for few-shot learning,” in ICLR, 2017.
  86. 86.H. Yang, X. He, and F. Porikli, “One-shot action localization by learning sequence matching network,” in CVPR, 2018.
  87. 87.L. Fei-Fei, R. Fergus, and P. Perona, “One-shot learning of object categories,” TPAMI, 2006.
  88. 88.P. Wang, L. Liu, C. Shen, Z. Huang, A. van den Hengel, and H. T. Shen, “Multi-attention network for one shot learning,” in CVPR, 2017.
  89. 89.O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra et al., “Matching networks for one shot learning,” in NIPS, 2016.
  90. 90.S. R. Fanello, I. Gori, G. Metta, and F. Odone, “One-shot learning for real-time action recognition,” in ICPRIA, 2013.
  91. 91.J. Wan, G. Guo, and S. Z. Li, “Explore efficient local features from rgb-d data for one-shot learning gesture recognition,” TPAMI, 2016.
  92. 92.J. Konecnˇ y and M. Hagara, “One-shot-learning gesture recogni- ` tion using hog-hof features,” JMLR, 2014.
  93. 93.H. Chen, G. Wang, J.-H. Xue, and L. He, “A novel hierarchical framework for human action recognition,” PR, 2016.
  94. 94.J. Liu, A. Shahroudy, D. Xu, A. C. Kot, and G. Wang, “Skeletonbased action recognition using spatio-temporal lstm network with trust gates,” TPAMI, 2017.
  95. 95.A. Graves, N. Jaitly, and A.-r. Mohamed, “Hybrid speech recognition with deep bidirectional lstm,” in ASRU, 2013.
  96. 96.J. Pennington, R. Socher, and C. Manning, “Glove: Global vectors for word representation,” in EMNLP, 2014.
  97. 97.C. De Boom, S. Van Canneyt, T. Demeester, and B. Dhoedt, “Representation learning for very short texts using weighted word embedding aggregation,” PRL, 2016.
  98. 98.C. Li, J. Cao, Z. Huang, L. Zhu, and H. T. Shen, “Leveraging weak semantic relevance for complex video event classification,” in ICCV, 2017.
  99. 99.Z. Yang, D. Yang, C. Dyer, X. He, A. Smola, and E. Hovy, “Hierarchical attention networks for document classification,” in ACL, 2016.
  100. 100.A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in CVPR, 2015.
  101. 101.W. Liu, Y. Wen, Z. Yu, M. Li, B. Raj, and L. Song, “Sphereface: Deep hypersphere embedding for face recognition,” in CVPR, 2017.
  102. 102.J.-F. Hu, W. Zheng, L. Ma, G. Wang, J. Lai, and J. Zhang, “Early action prediction by soft regression,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018.
  103. 103.J. Liu, A. Shahroudy, G. Wang, L.-Y. Duan, and A. C. Kot, “Skeleton-based online action prediction using scale selection network,” TPAMI, 2018.
  104. 104.M. Liu, H. Liu, and C. Chen, “Enhanced skeleton visualization for view invariant human action recognition,” PR, 2017.
  105. 105.J. Liu, G. Wang, L.-Y. Duan, K. Abdiyeva, and A. C. Kot, “Skeleton-based human action recognition with global contextaware attention lstm networks,” TIP, 2018.
  106. 106.Q. Ke, M. Bennamoun, S. An, F. Sohel, and F. Boussaid, “Learning clip representations for skeleton-based 3d action recognition,” TIP, 2018.
  107. 107.M. Liu and J. Yuan, “Recognizing human actions as the evolution of pose estimation maps,” in CVPR, 2018.
  108. 108.J. Liu, A. Shahroudy, G. Wang, L.-Y. Duan, and A. C. Kot, “Ssnet: Scale selection network for online 3d action prediction,” in CVPR, 2018.
  109. 109.A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv, 2017.
  110. 110.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” ICLR, 2015.
  111. 111.K. Simonyan and A. Zisserman., “Two-stream convolutional networks for action recognition in videos,” in NIPS, 2014.
  112. 112.S. Gupta, J. Hoffman, and J. Malik, “Cross modal distillation for supervision transfer,” in CVPR, 2016.

Citation

MLA
Liu, J., et al. “NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 10, 2020, pp. 2684–701, https://doi.org/10.1109/TPAMI.2019.2916873.
APA
Liu, J., Shahroudy, A., Perez, M., Wang, G., Duan, L.-Y., & Kot, A. C. (2020). NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(10), 2684–2701. https://doi.org/10.1109/TPAMI.2019.2916873
Chicago
Liu, J., A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot. 2020. “NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding”. IEEE Transactions on Pattern Analysis and Machine Intelligence 42 (10): 2684–2701. https://doi.org/10.1109/TPAMI.2019.2916873.
Harvard
Liu, J. et al. (2020) “NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(10), pp. 2684–2701. Available at: https://doi.org/10.1109/TPAMI.2019.2916873.
Vancouver
1. Liu J, Shahroudy A, Perez M, Wang G, Duan L-Y, Kot AC (2020) NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence 42:2684–2701

BibTeX

@article{Liu_2020, title={NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding}, volume={42}, ISSN={1939-3539}, url={http://dx.doi.org/10.1109/TPAMI.2019.2916873}, DOI={10.1109/tpami.2019.2916873}, number={10}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Liu, Jun and Shahroudy, Amir and Perez, Mauricio and Wang, Gang and Duan, Ling-Yu and Kot, Alex C.}, year={2020}, month=Oct, pages={2684–2701} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF