On the Benefits of 3D Pose and Tracking for Human Action Recognition

Jathushan RajasegaranGeorgios PavlakosAngjoo KanazawaChristoph FeichtenhoferJitendra Malik

article2023CVPR50 citations

Proposes a human-centric action recognition framework that tokenizes 3D body poses and contextual appearance along tracked trajectories, achieving significant performance gains on the AVA benchmark over conventional video-level methods.

Listen

Automated human action recognition in video is critical for applications ranging from autonomous systems and safety surveillance to digital media management. Most contemporary deep learning architectures analyze fixed spatial coordinates over time rather than following moving individuals. This fixed-location framing struggles to accurately capture fine-grained human motions, mutual interactions among individuals, and continuous actions across frame occlusions.

The article demonstrates the effectiveness of a trajectory-focused framework called Lagrangian Action Recognition with Tracking (LART). The main objective is to evaluate how explicitly tracking individuals over space and time, combined with three-dimensional (3D) body poses and visual appearance, improves action recognition accuracy relative to conventional video models.

The authors implemented a human-centric approach by extracting person trajectories from large-scale video data, totaling over one million tracks across approximately 900 hours of video from the AVA and Kinetics datasets. Using 3D tracking algorithms, each individual was modeled as a continuous trajectory incorporating 3D body joint angles, camera-relative 3D positions, and contextual visual appearance features. These human tokens were processed through a standard transformer network configured to jointly analyze the target individual's motion and the movements of other nearby people in the scene.

The findings show substantial performance improvements across standard benchmark tasks. First, the full model combining 3D pose, tracking, and contextual appearance achieved 42.3 mean average precision (mAP) on the AVA v2.2 benchmark, outperforming the prior state-of-the-art by 2.8 mAP. Second, when restricted entirely to 3D pose inputs without raw scene pixels, the system achieved 24.1 mAP, representing an improvement of 10.0 mAP over existing pose-only models. Third, tracking alone added an immediate 1.2 mAP performance gain by associating people across frames, while adding 3D pose contributed an additional 0.9 mAP gain. Fourth, tracking multiple individuals simultaneously improved classification in collaborative activities, driving large relative gains in actions like fighting, dancing, and hugging.

These results indicate that shifting from static grid-based video analysis to trajectory-based 3D tracking improves both detection accuracy and model robustness against visual occlusions. By explicitly modeling 3D spatial dynamics, systems can better interpret complex human-to-human interactions without requiring manually engineered relational rules. This improved accuracy helps lower false detection risks and enhances performance in safety-critical automated video monitoring.

Organizations developing video analytics pipelines should transition from single-frame bounding-box architectures to tracking-centric models that incorporate 3D body geometry. The primary implementation trade-off involves balancing the computational overhead of upstream 3D tracking pipelines against the achieved precision gains. Future development should prioritize incorporating detailed hand-pose models to enhance fine object manipulation tasks, as well as developing dedicated tracking representations for physical tools and manipulated objects.

Confidence in these findings is supported by rigorous evaluations on standard industry benchmarks. However, readers should note that the system's performance on fine-grained object manipulation remains lower than on full-body movements due to the lack of explicit object geometry in the current framework, which represents an area for targeted future enhancement.

arXiv: 2304.01199
Cover for On the Benefits of 3D Pose and Tracking for Human Action Recognition

Abstract

In this work we study the benefits of using tracking and 3D poses for action recognition. To achieve this, we take the Lagrangian view on analysing actions over a trajectory of human motion rather than at a fixed point in space. Taking this stand allows us to use the tracklets of people to predict their actions. In this spirit, first we show the benefits of using 3D pose to infer actions, and study person-person interactions. Subsequently, we propose a Lagrangian Action Recognition model by fusing 3D pose and contextualized appearance over tracklets. To this end, our method achieves state-of-the-art performance on the AVA v2.2 dataset on both pose only settings and on standard benchmark settings. When reasoning about the action using only pose cues, our pose model achieves +10.0 mAP gain over the corresponding state-of-the-art while our fused model has a gain of +2.8 mAP over the best state-of-the-art model. Code and results are available at: https://brjathu.github.io/LART

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Action Recognition with 3D Pose
  • 3.2. Actions from Appearance and 3D Pose
  • 4. Experiments
  • 4.1. Action Recognition with 3D Pose
  • 4.2. Actions from Appearance and 3D Pose
  • 4.3. Ablation Experiments
  • 4.4. Implementation details
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Lagrangian Action Recognition with Tracking (LART) Framework

    model/method

    Lagrangian Action Recognition with Tracking (LART) predicts human actions by following the trajectories of individuals across space-time (the Lagrangian viewpoint) rather than evaluating fixed spatial crops (the Eulerian viewpoint).

    Given a video of TT frames containing nn tracked persons, each individual i∈{1,2,…,n}i \in \{1, 2, \dots, n\} at time t∈{1,2,…,T}t \in \{1, 2, \dots, T\} is represented by a person-vector Hit={Pit,Qit}H_i^t = \{P_i^t, Q_i^t\}, where PitP_i^t is an explicit human-centric 3D geometric pose representation and QitQ_i^t is a contextualized appearance feature vector. The sequence of person-vectors across time defines an action-tube Φi={Hi1,Hi2,…,HiT}\Phi_i = \{H_i^1, H_i^2, \dots, H_i^T\}.

    A 16-layer, 16-head vanilla transformer with hidden embedding dimension d=512d=512 processes action tubes corresponding to a person-of-interest along with co-occurring context persons. A linear projection head maps the transformer output to binary action probabilities for multi-label action classification, optimized via binary cross-entropy loss.

  2. Knowl 2 — Person Vector Tokenization and Spatiotemporal-Identity Positional Encoding

    equation

    In the LART model, the person-vector HitH_i^t for person ii at frame tt combines 3D pose PitP_i^t and contextual appearance QitQ_i^t:

    Hit={Pit,Qit}={θit,ψit,Lit,Uit}H_i^t = \{P_i^t, Q_i^t\} = \{\theta_i^t, \psi_i^t, L_i^t, U_i^t\}

    where θit∈R23×3×3\theta_i^t \in \mathbb{R}^{23 \times 3 \times 3} denotes the SMPL relative 3D joint rotations, ψit∈R3×3\psi_i^t \in \mathbb{R}^{3 \times 3} is the global body orientation, Lit∈R3L_i^t \in \mathbb{R}^3 is the estimated 3D translation in camera space, and Uit∈RdfeatU_i^t \in \mathbb{R}^{d_{\text{feat}}} is the contextual appearance feature vector extracted from an MViT backbone across a temporal window of 2M2M frames centered at tt.

    Each vector HitH_i^t is projected linearly into a dd-dimensional embedding fproj(Hit)∈Rdf_{\text{proj}}(H_i^t) \in \mathbb{R}^d. For NN simultaneously modeled tracks, positional encoding PE(t,i,:)∈Rd\text{PE}(t, i, :) \in \mathbb{R}^d is added across time index tt and tracklet index i∈{0,1,…,N−1}i \in \{0, 1, \dots, N-1\} (where index 00 is assigned to the person-of-interest):

    PE(t,i,2r)=sin⁡(t/100004r/d),PE(t,i,2r+1)=cos⁡(t/100004r/d)\text{PE}(t, i, 2r) = \sin(t / 10000^{4r/d}), \quad \text{PE}(t, i, 2r + 1) = \cos(t / 10000^{4r/d})

    PE(t,i,2s+d/2)=sin⁡(i/100004s/d),PE(t,i,2s+d/2+1)=cos⁡(i/100004s/d)\text{PE}(t, i, 2s + d/2) = \sin(i / 10000^{4s/d}), \quad \text{PE}(t, i, 2s + d/2 + 1) = \cos(i / 10000^{4s/d})

    where r,s∈[0,d/4)r, s \in [0, d/4). The resulting (t+i×N)(t + i \times N)-th transformer token is:

    token(t+i×N)=fproj(Hit)+PE(t,i,:)\text{token}_{(t + i \times N)} = f_{\text{proj}}(H_i^t) + \text{PE}(t, i, :)

  3. Knowl 3 — State-of-the-Art Action Detection Results on AVA v2.2

    data/table

    On the AVA v2.2 benchmark, action detection performance is evaluated using mean Average Precision (mAP) across 60 action classes with a frame-level intersection-over-union (IoU) threshold of 0.5. LART incorporates features from MaskFeat-pretrained MViT backbones along with 3D pose and tracklet histories.

    Model Pretrain mAP
    SlowFast R101, 8×\times8 K400 23.8
    MViTv1-B, 64×\times3 K400 27.3
    SlowFast 16×\times8 +NL K600 27.5
    X3D-XL K600 27.4
    MViTv1-B-24, 32×\times3 K600 28.7
    Object Transformer K600 31.0
    ACAR R101, 8×\times8 +NL K600 31.4
    ACAR R101, 8×\times8 +NL K700 33.3
    MViT-L↑\uparrow312, 40×\times3 IN-21K+K400 31.6
    MaskFeat K400 37.5
    MaskFeat K600 38.8
    Video MAE K600 39.3
    Video MAE K400 39.5
    LART K400 42.3

    LART achieves 42.3 mAP with Kinetics-400 pretraining, outperforming the previous state-of-the-art Video MAE by +2.8 mAP. Fine-grained evaluations demonstrate performance improvements across 56 of the 60 AVA action classes, with gains exceeding +5 mAP for actions involving close person interaction or distinct whole-body dynamics (such as fighting, hugging, and climbing).

  4. Knowl 4 — Component Ablation on Tracking and 3D Pose Modalities

    data/table

    Ablation experiments isolate the individual benefits of tracking and 3D pose representations on the AVA v2.2 validation set using identical Mask R-CNN person detections. Performance is reported in mean Average Precision (mAP) overall, as well as on three sub-categories: Object Manipulation (OM), Person Interactions (PI), and Person Movement (PM).

    Model OM PI PM mAP
    MViT Baseline 32.2 41.1 58.6 40.2
    MViT + Tracking 33.4 43.0 59.3 41.4 (+1.2)
    MViT + Tracking + Pose (LART) 34.4 43.9 59.9 42.3 (+2.1)

    Extending single-frame MViT features across tracked temporal tubes contributes a +1.2 mAP gain over the standard mid-frame crop baseline. Adding explicit 3D SMPL pose and location features provides an additional +0.9 mAP gain, confirming that geometric 3D human pose offers orthogonal discriminative signals to raw 2D pixel features.

  5. Knowl 5 — Action Recognition from 3D Human Pose Alone

    data/table

    The pose-only model variant (LART-pose) evaluates human action understanding from 3D body pose without raw pixel appearance. Evaluated on the AVA v2.2 validation benchmark at frame IoU ≥0.5\ge 0.5, results are broken down across Object Manipulation (OM), Person Interactions (PI), Person Movement (PM), and overall mAP.

    Model Pose Type OM PI PM mAP
    PoTion 2D - - - 13.1
    JMRN 2D 7.1 17.2 27.6 14.1
    LART-pose SMPL 11.9 24.6 45.8 22.3
    LART-pose SMPL + Joints 13.3 25.9 48.7 24.1

    LART-pose achieves 24.1 mAP using only 3D SMPL parameters and 3D joint coordinates, outperforming 2D pose baselines (PoTion at 13.1 mAP and JMRN at 14.1 mAP) by +10.0 mAP. In the Person Movement category (e.g., walking, running, standing), LART-pose attains 48.7 mAP, reaching over 80% of the accuracy of appearance-based models (MaskFeat: 58.6 mAP) without observing scene context.

  6. Knowl 6 — Spatio-Temporal Action Localization on AVA-Kinetics

    data/table

    LART is evaluated on the AVA-Kinetics benchmark without model ensembling, following the standard spatio-temporal action localization protocol.

    Model mAP
    SlowFast 32.98
    ACAR 36.36
    RM 37.34
    LART 38.91

    LART achieves 38.91 mAP, outperforming the previous best single-model method (RM at 37.34 mAP) by +1.57 mAP.

  7. Knowl 7 — Multi-Track Context Modeling for Inter-Person Interactions

    model/method

    To capture person-person interactions and reciprocal actions (e.g., talking and listening, dancing, fighting), LART processes the action tube of the target actor alongside randomly sampled action tubes of other individuals in the video. For a target person ii, the model receives NN tracklets: the primary track Φi\Phi_i and N−1N - 1 context tracks {Φj∣j∈[N],j≠i}\{\Phi_j \mid j \in [N], j \neq i\}.

    F(Φi,{Φj∣j∈[N]};Θ)=Y^i\mathcal{F}(\Phi_i, \{\Phi_j \mid j \in [N]\}; \Theta) = \hat{Y}_i

    where Y^i={yi1,yi2,…,yiT}\hat{Y}_i = \{y_i^1, y_i^2, \dots, y_i^T\} contains action predictions for track ii across all time steps. Self-attention layers inside the transformer jointly attend over all N×TN \times T tokens, allowing the network to exploit relational cues (such as relative 3D distance Lit−LjtL_i^t - L_j^t and relative pose orientations).

    In pose-only evaluations, multi-person tracking (N=5N=5) increases Person Interaction performance by +2.4 mAP over single-person tracking (N=1N=1), yielding large relative gains on actions such as dancing (+39.8 mAP) and hugging (>+200% relative improvement).

  8. Knowl 8 — Large-Scale Pseudo-Supervised Pre-training and Fine-Tuning Pipeline

    experimental setup

    LART is trained via a two-stage training strategy leveraging automatically tracked video trajectories:

    1. Track Extraction: The PHALP tracker extracts 3D tracklets using Mask R-CNN detections across Kinetics-400 (686,000 tracks, 71.4 million bounding boxes) and AVA (320,000 tracks, 32.9 million bounding boxes), totaling over 1 million tracks and 900 hours of video with a mean track duration of 3.4 seconds.
    2. Pseudo-Supervision Pre-training: An MViT model pretrained with MaskFeat runs at 1Hz over every Kinetics-400 track to generate pseudo ground-truth action labels. Labels are propagated across 30-frame (1-second) windows. LART is pre-trained end-to-end on Kinetics-400 and AVA pseudo-labels for 30 epochs with a 0.4 mask token ratio.
    3. Fine-Tuning: The pre-trained weights are fine-tuned on AVA ground-truth labels using AdamW with a base learning rate of 0.001, β=(0.9,0.95)\beta = (0.9, 0.95), and a cosine annealing schedule with linear warmup.
    4. Inference: Missing detections in tracklets are in-filled with a learned mask token. Output predictions are average-pooled across a 12-frame sequence and evaluated at the center frame.

Coverage note — None was omitted; all primary architectural equations, pretraining protocols, benchmark comparisons (AVA v2.2, AVA-Kinetics), pose-only benchmarks, and ablation studies are fully represented.

References

  1. 1.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lućić, and Cordelia Schmid. ViViT: A video vision transformer. In ICCV, 2021.
  2. 2.Fabien Baradel, Thibault Groueix, Philippe Weinzaepfel, Romain Brégier, Yannis Kalantidis, and Grégory Rogez. Leveraging MoCap data for human mesh recovery. In 3DV, 2021.
  3. 3.Philipp Bergmann, Tim Meinhardt, and Laura Leal-Taixe. Tracking without bells and whistles. In ICCV, 2019.
  4. 4.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, 2021.
  5. 5.Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In ECCV, 2016.
  6. 6.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, 2017.
  7. 7.Vasileios Choutas, Lea Müller, Chun-Hao P Huang, Siyu Tang, Dimitrios Tzionas, and Michael J Black. Accurate 3D body shape regression using metric and semantic attributes. In CVPR, 2022.
  8. 8.Vasileios Choutas, Philippe Weinzaepfel, Jérôme Revaud, and Cordelia Schmid. PoTion: Pose motion representation for action recognition. In CVPR, 2018.
  9. 9.Piotr Dollár, Vincent Rabaud, Garrison Cottrell, and Serge Belongie. Behavior recognition via sparse spatio-temporal features. In 2005 IEEE international workshop on visual surveillance and performance evaluation of tracking and surveillance, 2005.
  10. 10.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. 2021.
  11. 11.Alexei A Efros, Alexander C Berg, Greg Mori, and Jitendra Malik. Recognizing action at a distance. In ICCV, 2003.
  12. 12.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. In ICCV, 2021.
  13. 13.Hao-Shu Fang, Shuqin Xie, Yu-Wing Tai, and Cewu Lu. RMPE: Regional multi-person pose estimation. In ICCV, 2017.
  14. 14.Christoph Feichtenhofer. X3D: Expanding architectures for efficient video recognition. In CVPR, 2020.
  15. 15.Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He. Masked autoencoders as spatiotemporal learners. In NeurIPS, 2022.
  16. 16.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In ICCV, 2019.
  17. 17.Yutong Feng, Jianwen Jiang, Ziyuan Huang, Zhiwu Qing, Xiang Wang, Shiwei Zhang, Mingqian Tang, and Yue Gao. Relation modeling in spatio-temporal action localization. arXiv preprint arXiv:2106.08061, 2021.
  18. 18.Georgia Gkioxari and Jitendra Malik. Finding action tubes. In CVPR, 2015.
  19. 19.Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Malik. Humans in 4D: Reconstructing and tracking humans with transformers. arXiv preprint (forthcoming), 2023.
  20. 20.Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik. AVA: A video dataset of spatio-temporally localized atomic visual actions. In CVPR, 2018.
  21. 21.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask R-CNN. In ICCV, 2017.
  22. 22.Gunnar Johansson. Visual perception of biological motion and a model for its analysis. Perception & psychophysics, 14(2):201–211, 1973.
  23. 23.Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In CVPR, 2018.
  24. 24.Angjoo Kanazawa, Jason Y Zhang, Panna Felsen, and Jitendra Malik. Learning 3D human dynamics from video. In CVPR, 2019.
  25. 25.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017.
  26. 26.Machiel Keestra. Understanding human action. Integrating meanings, mechanisms, causes, and contexts. 2015.
  27. 27.Alexander Klaser, Marcin Marszałek, and Cordelia Schmid. A spatio-temporal descriptor based on 3D-gradients. In BMVC, 2008.
  28. 28.Muhammed Kocabas, Nikos Athanasiou, and Michael J Black. VIBE: Video inference for human body pose and shape estimation. In CVPR, 2020.
  29. 29.Muhammed Kocabas, Chun-Hao P Huang, Otmar Hilliges, and Michael J Black. PARE: Part attention regressor for 3D human body estimation. In ICCV, 2021.
  30. 30.Muhammed Kocabas, Chun-Hao P Huang, Joachim Tesch, Lea Müller, Otmar Hilliges, and Michael J Black. SPEC: Seeing people in the wild with an estimated camera. In ICCV, 2021.
  31. 31.Nikos Kolotouros, Georgios Pavlakos, Michael J Black, and Kostas Daniilidis. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. In ICCV, 2019.
  32. 32.Nikos Kolotouros, Georgios Pavlakos, Dinesh Jayaraman, and Kostas Daniilidis. Probabilistic modeling for human mesh recovery. In ICCV, 2021.
  33. 33.Ang Li, Meghana Thotakuri, David A Ross, João Carreira, Alexander Vostrikov, and Andrew Zisserman. The ava-kinetics localized human actions video dataset. arXiv preprint arXiv:2005.00214, 2020.
  34. 34.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. MViTv2: Improved multiscale vision transformers for classification and detection. In CVPR, 2022.
  35. 35.Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. SMPL: A skinned multi-person linear model. ACM Transactions on Graphics (TOG), 34(6):1–16, 2015.
  36. 36.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  37. 37.Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, and Christoph Feichtenhofer. TrackFormer: Multi-object tracking with transformers. In CVPR, 2022.
  38. 38.Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. Video transformer network. In ICCV, 2021.
  39. 39.Junting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu, Jing Shao, and Hongsheng Li. Actor-context-actor relation network for spatio-temporal action localization. In CVPR, 2021.
  40. 40.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. Expressive body capture: 3D hands, face, and body from a single image. In CVPR, 2019.
  41. 41.Georgios Pavlakos, Jitendra Malik, and Angjoo Kanazawa. Human mesh recovery from multiple shots. In CVPR, 2022.
  42. 42.Jathushan Rajasegaran, Georgios Pavlakos, Angjoo Kanazawa, and Jitendra Malik. Tracking people with 3D representations. In NeurIPS, 2021.
  43. 43.Jathushan Rajasegaran, Georgios Pavlakos, Angjoo Kanazawa, and Jitendra Malik. Tracking people by predicting 3D appearance, location and pose. In CVPR, 2022.
  44. 44.Davis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang, Srinath Sridhar, and Leonidas J Guibas. HuMoR: 3D human motion model for robust pose estimation. In ICCV, 2021.
  45. 45.Anshul Shah, Shlok Mishra, Ankan Bansal, Jun-Cheng Chen, Rama Chellappa, and Abhinav Shrivastava. Pose and joint-aware action recognition. In WACV, 2022.
  46. 46.Karen Simonyan and Andrew Zisserman. Two-stream convolutional networks for action recognition in videos. NIPS, 2014.
  47. 47.Chen Sun, Abhinav Shrivastava, Carl Vondrick, Rahul Sukthankar, Kevin Murphy, and Cordelia Schmid. Relational action forecasting. In CVPR, 2019.
  48. 48.Graham W Taylor, Rob Fergus, Yann LeCun, and Christoph Bregler. Convolutional learning of spatio-temporal features. In ECCV, 2010.
  49. 49.Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training. In NeurIPS, 2022.
  50. 50.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3D convolutional networks. In ICCV, 2015.
  51. 51.Gõl Varol, Ivan Laptev, Cordelia Schmid, and Andrew Zisserman. Synthetic humans for action recognition from unseen viewpoints. International Journal of Computer Vision, 129(7):2264–2287, 2021.
  52. 52.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, 2017.
  53. 53.Heng Wang, A. Klaser, C. Schmid, and Cheng-Lin Liu. Action recognition by dense trajectories. In CVPR, 2011.
  54. 54.Heng Wang and Cordelia Schmid. Action recognition with improved trajectories. In ICCV, 2013.
  55. 55.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, 2018.
  56. 56.Xiaolong Wang and Abhinav Gupta. Videos as space-time region graphs. In ECCV, 2018.
  57. 57.Zelun Wang and Jyh-Charn Liu. Translating math formula images to latex sequences using deep neural networks with sequence-level training. International Journal on Document Analysis and Recognition (IJDAR), 24(1):63–75, 2021.
  58. 58.Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. arXiv preprint arXiv:2112.09133, 2021.
  59. 59.Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. In CVPR, 2022.
  60. 60.Philippe Weinzaepfel and Grégory Rogez. Mimetics: Towards understanding human actions out of context. IJCV, 2021.
  61. 61.Chao-Yuan Wu and Philipp Krähenbõhl. Towards long-form video understanding. In CVPR, 2021.
  62. 62.Yuliang Xiu, Jiefeng Li, Haoyu Wang, Yinghong Fang, and Cewu Lu. Pose Flow: Efficient online pose tracking. In BMVC, 2018.
  63. 63.An Yan, Yali Wang, Zhifeng Li, and Yu Qiao. PA3D: Pose-action 3D machine for video recognition. In CVPR, 2019.
  64. 64.Hongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang, Yebin Liu, Limin Wang, and Zhenan Sun. PyMAF: 3D human pose and shape regression with pyramidal mesh alignment feedback loop. In ICCV, 2021.
  65. 65.Yubo Zhang, Pavel Tokmakov, Martial Hebert, and Cordelia Schmid. A structured model for action detection. In CVPR, 2019.

Citation

MLA
Rajasegaran, J., et al. “On the Benefits of 3D Pose and Tracking for Human Action Recognition”. arXiv, 2023, http://arxiv.org/abs/2304.01199v2.
APA
Rajasegaran, J., Pavlakos, G., Kanazawa, A., Feichtenhofer, C., & Malik, J. (2023). On the Benefits of 3D Pose and Tracking for Human Action Recognition. arXiv. http://arxiv.org/abs/2304.01199v2
Chicago
Rajasegaran, J., G. Pavlakos, A. Kanazawa, C. Feichtenhofer, and J. Malik. 2023. “On the Benefits of 3D Pose and Tracking for Human Action Recognition”. arXiv. http://arxiv.org/abs/2304.01199v2.
Harvard
Rajasegaran, J. et al. (2023) “On the Benefits of 3D Pose and Tracking for Human Action Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.01199v2.
Vancouver
1. Rajasegaran J, Pavlakos G, Kanazawa A, Feichtenhofer C, Malik J (2023) On the Benefits of 3D Pose and Tracking for Human Action Recognition. arXiv

BibTeX

@article{rajasegaran2023the,
  title = {On the Benefits of 3D Pose and Tracking for Human Action Recognition},
  author = {Rajasegaran, Jathushan and Pavlakos, Georgios and Kanazawa, Angjoo and Feichtenhofer, Christoph and Malik, Jitendra},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.01199v2},
  eprint = {2304.01199}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE