JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection

Mahsa EhsanpourFatemeh Sadat SalehSilvio SavareseIan D. ReidHamid Rezatofighi

article2022CVPR69 citations

Presents JRDB-Act, a large-scale multimodal benchmark featuring over 2.8 million action labels alongside social group annotations and confidence ratings captured from a mobile robot, establishing an end-to-end framework for joint individual action and group activity detection in crowded real-world environments.

Listen

Deploying mobile robots and autonomous systems in crowded human environments requires machines to accurately perceive not only individual human movements but also social interactions and collective group behaviors. Existing action recognition benchmarks generally rely on static cameras or controlled settings, failing to capture the complexities faced by moving robotic platforms navigating crowded, unconstrained spaces.

The article introduces JRDB-Act, a large-scale multimodal dataset designed to benchmark individual action recognition, social group detection, and collective activity estimation from a mobile robot platform. It also presents and evaluates a unified learning pipeline tailored to handle the dataset’s dense crowds, moving perspective, and highly unbalanced real-world action distributions.

The dataset expands upon 64 minutes of sensory footage (spanning 54 campus video sequences captured with 360-degree cylindrical cameras and LiDAR) by contributing over 2.8 million dense spatio-temporal action annotations across 26 distinct daily action classes. Scenes feature an average density of 30 people per frame, and each annotation includes annotator confidence ratings ranging from easy to difficult. To model this data, the authors developed a baseline framework that integrates geometric and visual feature relations, an eigenvalue-based spectral loss for group cardinality estimation, and a partitioned loss mechanism to mitigate long-tailed class imbalances.

Evaluation reveals four core findings. First, the proposed framework improved overall social group detection accuracy to 59.2% average precision on ground-truth boxes, outperforming existing baseline methods by approximately 17.8 percentage points. Second, the loss partitioning approach boosted individual action recognition from 8.0% to 9.0% mean average precision, outperforming standard class-frequency weighting schemes. Third, when evaluating end-to-end system performance on test sequences with detected bounding boxes, better object localization significantly raised social group detection (from 29.1% to 31.5% average precision) but had negligible impact on action recognition, which remained near 5.4% mean average precision. Fourth, annotator difficulty tags substantially influence performance benchmarks, as models scored roughly 37% higher on social grouping when evaluated solely on visually distinct (easy) labels compared to full sets containing heavily occluded (difficult) targets.

These findings demonstrate that safe robotic navigation and human-robot interaction cannot rely on conventional video classification models. Action recognition in mobile crowd settings is limited by visual perspective, rapid camera motion, occlusion, and severe label imbalance rather than bounding box localization alone. Practical deployment requires perception architectures that can reason about geometric relationships and long-tailed distributions simultaneously.

For future development, the source recommends exploring multi-modal sensor fusion by directly integrating the available 3D LiDAR point-cloud data into feature extractors to enhance spatial reasoning. Researchers should also adopt partitioned loss formulations rather than standard cost-sensitive weighting when handling long-tailed behavioral data.

The primary limitations involve the overall low baseline scores on complex real-world actions (reaching 5.4% test mean average precision) and the fact that 38.6% of annotations involve difficult or partially occluded cases inferred from movement history. Stakeholders should view the presented framework as a foundational research baseline rather than an off-the-shelf system ready for immediate autonomous deployment.

arXiv: 2106.08827
Cover for JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection

Abstract

The availability of large-scale video action understanding datasets has facilitated advances in the interpretation of visual scenes containing people. However, learning to recognise human actions and their social interactions in an unconstrained real-world environment comprising numerous people, with potentially highly unbalanced and long-tailed distributed action labels from a stream of sensory data captured from a mobile robot platform remains a significant challenge, not least owing to the lack of a reflective large-scale dataset. In this paper, we introduce JRDB-Act, as an extension of the existing JRDB, which is captured by a social mobile manipulator and reflects a real distribution of human daily-life actions in a university campus environment. JRDB-Act has been densely annotated with atomic actions, comprises over 2.8M action labels, constituting a large-scale spatio-temporal action detection dataset. Each human bounding box is labeled with one pose-based action label and multiple (optional) interaction-based action labels. Moreover JRDB-Act provides social group annotation, conducive to the task of grouping individuals based on their interactions in the scene to infer their social activities (common activities in each social group). Each annotated label in JRDB-Act is tagged with the annotators’ confidence level which contributes to the development of reliable evaluation strategies. In order to demonstrate how one can effectively utilise such annotations, we develop an end-to-end trainable pipeline to learn and infer these tasks, i.e. individual action and social group detection. The data and the evaluation code will be publicly available at https://jrdb.erc.monash.edu/.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. The JRDB-Act Dataset
  • 4. Proposed Baseline
  • 5. Experiments
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — JRDB-Act Dataset Specification and Multi-Task Annotation Scheme

    definition

    JRDB-Act is a large-scale multimodal benchmark for spatio-temporal individual human action recognition, social group clustering, and group activity detection from a mobile robot viewpoint. Recorded across 54 indoor and outdoor scenes (64 minutes total) in a university campus via the JackRabbot mobile robotic platform, it comprises synchronized 360-degree cylindrical panoramic RGB video streams (from five stereo RGB cameras) and 3D point clouds (from two 16-array LiDAR sensors). The dataset features an average human density of 30 people per frame and is split at video level into 20 training sequences (1,419 one-second keyframes), 7 validation sequences (404 keyframes), and 27 test sequences (1,802 keyframes).

    Annotations are provided at 7 frames per second for over 2.4 million 2D bounding boxes and 1.8 million 3D bounding boxes. The action vocabulary contains 26 atomic action classes across three categories:

    1. Pose-based (11 classes): Standing, Walking, Sitting, Cycling, Going upstairs, Bending, Scooter riding, Going downstairs, Skating, Running, Lying.
    2. Human-Human Interaction (3 classes): Listening to someone, Talking to someone, Greeting gestures.
    3. Human-Object Interaction (12 classes): Holding something, Looking at robot, Looking into something, Looking at something, Typing, Interaction with door, Eating something, Talking on the phone, Pointing at something, Pushing, Reading, Writing.

    Each bounding box is annotated with exactly one mandatory pose-based label and zero or more optional interaction-based labels. Annotations are accompanied by confidence difficulty tags: easy (unambiguous visual evidence), moderate (minor uncertainty requiring probabilistic estimation), difficult (significant occlusion or distance requiring temporal trajectory cues), and impossible (fully occluded or out-of-range). Social groups are annotated by assigning shared group identifiers per frame, and the groundtruth social activity for each group is inferred as the most frequent action performed by its members.

  2. Knowl 2 — Eigenvalue and Cardinality Regularized Social Grouping Loss

    model/method

    To infer social group structures among NN detected individuals in a scene, a pairwise affinity matrix Aθ∈[0,1]N×NA_\theta \in [0, 1]^{N \times N} is predicted. For any bounding box pair (i,j)(i, j), a normalized visual feature distance DV(hiθ,hjθ)∈[0,1]D_V(h_i^\theta, h_j^\theta) \in [0, 1] and a normalized Generalized Intersection over Union (GIoU) geometric distance DG(i,j)∈[0,1]D_G(i, j) \in [0, 1] are computed, concatenated, and mapped to a scalar affinity score Ai,jθA_{i,j}^\theta using a Multi-Layer Perceptron (MLP).

    The social grouping loss function LGL_G optimizes three complementary objectives:

    LG=LBCE(Aθ,A^)+Leig(Lθ,L^)+LMSE([hθ∥∑iAiθ],GTcardinality)L_G = L_{BCE}(A_\theta, \hat{A}) + L_{eig}(L_\theta, \hat{L}) + L_{MSE}\left(\left[h_\theta \parallel \sum_{i} A_i^\theta\right], GT_{\text{cardinality}}\right)

    where:

    • LBCE(Aθ,A^)L_{BCE}(A_\theta, \hat{A}) is the binary cross-entropy loss between the predicted affinity matrix AθA_\theta and the binary groundtruth group connectivity matrix A^∈{0,1}N×N\hat{A} \in \{0, 1\}^{N \times N}.
    • Leig(Lθ,L^)L_{eig}(L_\theta, \hat{L}) enforces that the graph Laplacian Lθ=Dθ−AθL_\theta = D_\theta - A_\theta has the same number of zero eigenvalues as the groundtruth Laplacian L^=D^−A^\hat{L} = \hat{D} - \hat{A} (where zero eigenvalues correspond to connected components/social groups):
    Leig(θ)=e^TLθTLθe^+αexp⁡(−βtr⁡(LˉθTLˉθ))L_{eig}(\theta) = \hat{e}^T L_\theta^T L_\theta \hat{e} + \alpha \exp\left(-\beta \operatorname{tr}(\bar{L}_\theta^T \bar{L}_\theta)\right)

    where e^\hat{e} denotes the groundtruth eigenvector corresponding to zero eigenvalues, Lˉθ\bar{L}_\theta is the normalized Laplacian, and α,β\alpha, \beta are hyperparameters (set to α=1,β=1\alpha = 1, \beta = 1).

    • LMSEL_{MSE} is a mean squared error cardinality loss that predicts the total number of social groups in the scene by regressing the concatenation [⋅∥⋅][\cdot \parallel \cdot] of max-pooled visual features hθ=max⁡ihiθh_\theta = \max_i h_i^\theta and the sum of predicted pairwise affinities against the groundtruth group count GTcardinalityGT_{\text{cardinality}}.
  3. Knowl 3 — Hierarchical Loss Partitioning for Unbalanced Multi-Label Action Recognition

    model/method

    To mitigate severe performance degradation caused by long-tailed class imbalance across human daily-life actions, action categories are partitioned into multiple balanced subsets rather than trained with a monolithic cross-entropy or binary cross-entropy loss.

    Pose-based classes are divided into 3 disjoint partitions and interaction-based classes are divided into 4 disjoint partitions. Partitions are formed such that within any partition, the sample count of the least frequent class is at least 0.1×0.1 \times the sample count of the most frequent class in that partition. Each partition (except the terminal partition in the hierarchy) is augmented with an auxiliary Other class representing the presence of actions belonging to lower-frequency partitions.

    The action learning loss LActL_{Act} is formulated as:

    LAct=∑i=02λiLCE(Piθ,Pi)+∑j=03λjLBCE(Ijθ,Ij)L_{Act} = \sum_{i=0}^{2} \lambda_i L_{CE}(P_i^\theta, P^i) + \sum_{j=0}^{3} \lambda_j L_{BCE}(I_j^\theta, I^j)

    where LCEL_{CE} and LBCEL_{BCE} are cross-entropy and binary cross-entropy losses, PiθP_i^\theta and IjθI_j^\theta are predicted pose and interaction distributions for partition ii and jj, PiP^i and IjI^j are corresponding groundtruth vectors, and λi,λj\lambda_i, \lambda_j are balancing weights.

    During training, to prevent gradient corruption from all-zero label vectors, only partition heads containing an active groundtruth action are trained for a given sample, while other partition heads are masked. At inference time, predictions are decoded hierarchically: evaluation begins at the most frequent partition and proceeds down the hierarchy to the next partition if and only if the Other class is predicted.

  4. Knowl 4 — Two-Stage Baseline Training and Multi-Task Inference Protocol

    model/method

    The JRDB-Act baseline pipeline performs joint detection of individual actions, social groups, and per-group social activities using an end-to-end architecture:

    1. Feature Extraction: 15-frame video clips ending at the labeled keyframe are processed via an Inflated 3D ConvNet (I3D) backbone followed by self-attention and Graph Attention Network (GAT) modules to generate a 1024-dimensional spatio-temporal feature hiθh_i^\theta for each person bounding box ii.
    2. Group Representation: For action classification, individual features are fused with social context by concatenating hiθh_i^\theta with the max-pooled feature vector of their social group SGiθ=max⁡k∈Group(i)hkθSG_i^\theta = \max_{k \in \text{Group}(i)} h_k^\theta.
    3. Two-Stage Training: Optimization is performed using the Adam optimizer (β1=0.9,β2=0.999,ϵ=10−8\beta_1 = 0.9, \beta_2 = 0.999, \epsilon = 10^{-8}) with a batch size of 1.
      • Stage 1: Train the social grouping module under grouping loss LGL_G for 50 epochs at an initial learning rate of 10−410^{-4}, decaying by a factor of 10 upon validation loss plateau.
      • Stage 2: Fine-tune the full multi-task network under Ltotal=LG+LActL_{total} = L_G + L_{Act} for 50 epochs.
    4. Inference: Pairwise similarities AθA_\theta and predicted group cardinality are processed using self-tuning graph spectral clustering to segment individuals into discrete social groups. Group social activities are assigned by computing the mode of predicted individual member actions.
  5. Knowl 5 — Evaluation Protocol for Spatio-Temporal Action and Social Group Detection

    definition

    Evaluation is conducted on 1-second keyframes using Average Precision (AP) at an Intersection-over-Union (IoU) threshold of 0.5:

    • Social Grouping AP: For each detection threshold, detected bounding boxes are matched to groundtruth bounding boxes with IoU≥0.5\text{IoU} \ge 0.5 to form a set of True Positives (TP). Optimal bipartite matching is solved between predicted group IDs and groundtruth group IDs over the matched TP boxes. The number of true positive social group members is computed to calculate standard Average Precision (AP) separately for group sizes G1G_1 (1 member), G2G_2 (2 members), G3G_3 (3 members), G4G_4 (4 members), G5+G_{5+} (5 or more members), and overall average grouping AP.
    • Individual Action mAP: Mean Average Precision (mAP) computed over individual action categories across all detected person bounding boxes.
    • Social Activity Metrics:
      • Group Activity mAP1 (G-Act mAP1): Evaluates inferred group activity labels per bounding box independently of whether the social group clustering is correct.
      • Group Activity mAP2 (G-Act mAP2): Evaluates inferred group activity labels strictly, counting a bounding box as a true positive only if both its assigned social group and its social activity label are simultaneously correct.
  6. Knowl 6 — Ablation of Social Group Formation Modules on JRDB-Act Validation Set

    data/table

    Ablation results on the JRDB-Act validation set evaluate the impact of geometric GIoU features, learned cardinality regression, and eigenvalue loss on social group detection using groundtruth bounding boxes (with easy and moderate difficulty annotations):

    Method Grouping Loss Cardinality Geo Feat G1 AP G2 AP G3 AP G4 AP G5+ AP Overall AP
    Baseline1 BCE Heuristic - 8.0 29.3 37.5 65.4 67.0 41.4
    Baseline2 BCE Heuristic ✓ 26.1 57.0 61.2 63.0 53.7 52.2
    Baseline3 BCE MSE ✓ 79.6 63.0 43.7 56.9 40.7 56.8
    Ours BCE+EIGEN MSE ✓ 81.4 64.8 49.1 63.2 37.2 59.2

    Baseline 1 relies solely on visual features and an eigengap heuristic for cluster counts, which underestimates group counts and overmerges small groups (yielding low G1G_1 AP of 8.0% and artificially inflated G5+G_5+ AP). Incorporating geometric GIoU features (Baseline 2) improves overall AP from 41.4% to 52.2%. Replacing the heuristic with learned group cardinality regression (Baseline 3) boosts single-person group detection (G1G_1 AP) from 26.1% to 79.6%. Incorporating the eigenvalue regularizer (LeigL_{eig}) achieves the highest overall grouping AP of 59.2%.

  7. Knowl 7 — Ablation of Loss Balancing Strategies for Action Recognition on JRDB-Act Validation Set

    data/table

    Comparison of different loss formulations for handling long-tailed action distributions on the JRDB-Act validation set using groundtruth bounding boxes (evaluated on easy and moderate difficulty labels):

    Method Action mAP (%)
    Single CE + Single BCE [CE+BCE] 8.0
    Inverse-Frequency Weighted [W-CE+W-BCE] 8.1
    Partitioned Loss [M-CE+M-BCE] (Ours) 9.0

    Standard single-head cross-entropy and binary cross-entropy yield 8.0% mAP. Weighting loss terms by the normalized inverse class frequencies (W-CE+W-BCE) provides only a marginal improvement to 8.1% mAP. In contrast, partitioning action categories into balanced sub-groups with hierarchical inference (M-CE+M-BCE) achieves 9.0% mAP, demonstrating superior handling of long-tailed label distributions.

  8. Knowl 8 — Benchmark Evaluation on JRDB-Act Test Set Across Detection Backbones

    data/table

    Performance comparison on the JRDB-Act test set (1,802 keyframes) between the baseline architecture of Ehsanpour et al. (ECCV 2020) and the proposed model across two object detectors: Faster-RCNN (52.2% detection mAP on JRDB) and MMPAT (68.1% detection mAP on JRDB), evaluated on easy and moderate difficulty annotations:

    Method G1 AP G2 AP G3 AP G4 AP G5+ AP Overall AP Action mAP G-Act mAP1 G-Act mAP2
    Ehsanpour et al. + Faster-RCNN 9.5 24.3 21.2 39.8 10.8 21.1 4.4 3.5 1.3
    Ehsanpour et al. + MMPAT 11.8 27.5 22.4 38.8 24.6 25.0 4.9 3.5 1.3
    Ours + Faster-RCNN 42.5 40.8 23.1 25.6 13.4 29.1 5.3 4.4 3.4
    Ours + MMPAT 56.6 39.5 24.3 22.4 14.8 31.5 5.4 4.7 3.4

    The proposed framework outperforms the prior method across all tasks. Upgrading the detector from Faster-RCNN to MMPAT improves overall grouping AP from 29.1% to 31.5% (driven by a rise in G1G_1 AP from 42.5% to 56.6%), but produces only modest gains on Action mAP (5.3% to 5.4%) and G-Act mAP2 (3.4%), underscoring that action and social activity understanding are bottlenecked by camera motion, perspective distortion, and class imbalance rather than 2D localization alone.

  9. Knowl 9 — Impact of Annotator Confidence Difficulty Tags on Evaluation Metrics

    data/table

    Performance of the proposed model (with Faster-RCNN detections) on the JRDB-Act test set evaluated across different subsets of annotator difficulty levels:

    Evaluation Set G1 AP G2 AP G3 AP G4 AP G5+ AP Overall AP Action mAP G-Act mAP1 G-Act mAP2
    Easy Only [E] 44.4 42.7 27.1 28.4 13.9 31.3 5.7 4.4 3.5
    Easy + Moderate [E,M] 42.5 40.8 23.1 25.6 13.4 29.1 5.3 4.4 3.4
    Easy + Moderate + Difficult [E,M,D] 34.9 37.3 18.3 16.4 7.6 22.9 4.4 3.5 2.7

    Restricting evaluation strictly to Easy [E] annotations yields the highest performance across all metrics (31.3% overall grouping AP, 5.7% Action mAP, 3.5% G-Act mAP2). Adding Moderate annotations [E,M] introduces a minor drop (29.1% grouping AP, 5.3% Action mAP). Incorporating Difficult annotations [E,M,D] causes a steep drop (overall grouping AP falls to 22.9%, Action mAP to 4.4%, and G-Act mAP2 to 2.7%), confirming that heavily occluded or distant targets require explicit long-term reasoning.

  10. Knowl 10 — Empirical Group Size and Annotation Difficulty Distribution in JRDB-Act

    data/table

    Empirical statistical distribution of social group sizes and annotator confidence tags across JRDB-Act:

    • Social Group Size Distribution:

      • 1 member (G1G_1 / singletons): 75.5%
      • 2 members (G2G_2 / dyads): 16.6%
      • 3 members (G3G_3 / triads): 5.0%
      • 4 members (G4G_4 / tetrads): 1.2%
      • 5 or more members (G5+G_{5+}): 1.0% (with a maximum observed group size of 29 members)
    • Action Annotation Difficulty Distribution:

      • Easy: 38.6% (obvious visual cues)
      • Moderate: 21.1% (uncertainty requiring probabilistic inference)
      • Difficult: 40.3% (severe occlusion or distance requiring temporal trajectory history)

    Thus, only 59.7% of all action instances are identifiable from immediate visual cues alone (Easy and Moderate), while over 40% require temporal context.

Coverage note — None was omitted; all key contributions including dataset statistics, annotation schema, model formulation, loss equations, and experimental evaluations are represented.

References

  1. 1.Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. Youtube-8m: A large-scale video classification benchmark. arXiv preprint arXiv:1609.08675, 2016.
  2. 2.Yizhak Ben-Shabat, Xin Yu, Fatemeh Saleh, Dylan Campbell, Cristian Rodriguez-Opazo, Hongdong Li, and Stephen Gould. The ikea asm dataset: Understanding people assembling furniture through actions, objects and pose. In WACV, pages 847–859, 2021.
  3. 3.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In CVPR, pages 961–970, 2015.
  4. 4.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, pages 6299–6308, 2017.
  5. 5.Wongun Choi and Silvio Savarese. A unified framework for multi-target tracking and collective activity recognition. In ECCV, pages 215–230, 2012.
  6. 6.Wongun Choi and Silvio Savarese. Understanding collective activitiesof people from videos. IEEE transactions on pattern analysis and machine intelligence, 36(6):1242–1257, 2013.
  7. 7.Wongun Choi, Khuram Shahid, and Silvio Savarese. What are they doing?: Collective activity classification using spatio-temporal relationship among people. In ICCVW, pages 1282–1289, 2009.
  8. 8.Wongun Choi, Khuram Shahid, and Silvio Savarese. Learning context for collective activity recognition. In CVPR, pages 3273–3280, 2011.
  9. 9.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In ECCV, pages 720–736, 2018.
  10. 10.Zheng Dang, Kwang Moo Yi, Yinlin Hu, Fei Wang, Pascal Fua, and Mathieu Salzmann. Eigendecomposition-free training of deep networks with zero eigenvalue-based losses. In ECCV, pages 768–783, 2018.
  11. 11.Ali Diba, Mohsen Fayyaz, Vivek Sharma, Manohar Paluri, Jurgen Gall, Rainer Stiefelhagen, and Luc Van Gool. Large scale holistic video understanding. In ECCV, pages 593–610. Springer, 2020.
  12. 12.Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. Long-term recurrent convolutional networks for visual recognition and description. In CVPR, pages 2625–2634, 2015.
  13. 13.Mahsa Ehsanpour, Alireza Abedin, Fatemeh Saleh, Javen Shi, Ian Reid, and Hamid Rezatofighi. Joint learning of social groups, individuals action and sub-group activities in videos. In ECCV, 2020.
  14. 14.Mark Everingham, SM Ali Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge: A retrospective. International journal of computer vision, 111(1):98–136, 2015.
  15. 15.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In ICCV, pages 6202–6211, 2019.
  16. 16.Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman. A better baseline for ava. arXiv preprint arXiv:1807.10066, 2018.
  17. 17.Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman. Video action transformer network. In CVPR, pages 244–253, 2019.
  18. 18.Lena Gorelick, Moshe Blank, Eli Shechtman, Michal Irani, and Ronen Basri. Actions as space-time shapes. IEEE transactions on pattern analysis and machine intelligence, 29(12):2247–2253, 2007.
  19. 19.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The” something something” video database for learning and evaluating visual common sense. In ICCV, page 5, 2017.
  20. 20.Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al. Ava: A video dataset of spatio-temporally localized atomic visual actions. In CVPR, pages 6047–6056, 2018.
  21. 21.Yuhang He, Wentao Yu, Jie Han, Xing Wei, Xiaopeng Hong, and Yihong Gong. Know your surroundings: Panoramic multi-object tracking by multimodality collaboration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2969–2980, 2021.
  22. 22.Mostafa S Ibrahim, Srikanth Muralidharan, Zhiwei Deng, Arash Vahdat, and Greg Mori. A hierarchical deep temporal model for group activity recognition. In CVPR, pages 1971–1980, 2016.
  23. 23.Haroon Idrees, Amir R Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah. The thumos challenge on action recognition for videos “in the wild”. Computer Vision and Image Understanding, 155:1–23, 2017.
  24. 24.Hueihan Jhuang, Juergen Gall, Silvia Zuffi, Cordelia Schmid, and Michael J Black. Towards understanding action recognition. In ICCV, pages 3192–3199, 2013.
  25. 25.Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. Large-scale video classification with convolutional neural networks. In CVPR, pages 1725–1732, 2014.
  26. 26.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017.
  27. 27.Yan Ke, Rahul Sukthankar, and Martial Hebert. Efficient visual event detection using volumetric features. In ICCV, volume 1, pages 166–173, 2005.
  28. 28.Hildegard Kuehne, Hueihan Jhuang, Estıbaliz Garrote, Tomaso Poggio, and Thomas Serre. Hmdb: a large video database for human motion recognition. In ICCV, pages 2556–2563, 2011.
  29. 29.Ang Li, Meghana Thotakuri, David A. Ross, Joao Carreira, Alexander Vostrikov, and Andrew Zisserman. The ava-kinetics localized human actions video dataset, 2020.
  30. 30.Dong Li, Zhaofan Qiu, Qi Dai, Ting Yao, and Tao Mei. Recurrent tubelet proposal and recognition networks for action detection. In ECCV, pages 303–318, 2018.
  31. 31.Shuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang, Shinan Liu, Jun Hou, and Shuai Yi. Groupformer: Group activity recognition with clustered spatial-temporal transformer. In ICCV, 2021.
  32. 32.Wenbo Li, Ming-Ching Chang, and Siwei Lyu. Who did what at where and when: simultaneous multi-person tracking and activity recognition. arXiv preprint arXiv:1807.01253, 2018.
  33. 33.Yu Li, Tao Wang, Bingyi Kang, Sheng Tang, Chunfeng Wang, Jintao Li, and Jiashi Feng. Overcoming classifier imbalance for long-tail object detection with balanced group softmax. In CVPR, pages 10991–11000, 2020.
  34. 34.Marcin Marszałek, Ivan Laptev, and Cordelia Schmid. Actions in context. Citeseer, 2009.
  35. 35.Roberto Martın-Martın*, Mihir Patel*, Hamid Rezatofighi*, Abhijeet Shenoi, JunYoung Gwak, Nathan Dass, Alan Federman, Patrick Goebel, and Silvio Savarese. JRDB: A dataset and benchmark of egocentric robot visual perception of humans in built environments. IEEE transactions on pattern analysis and machine intelligence, 2021.
  36. 36.Andrew Y Ng, Michael I Jordan, and Yair Weiss. On spectral clustering: Analysis and an algorithm. In NIPS, pages 849–856, 2002.
  37. 37.Jamie Ray, Heng Wang, Du Tran, Yufei Wang, Matt Feiszli, Lorenzo Torresani, and Manohar Paluri. Scenes-objects-actions: A multi-task, multi-label video dataset. In ECCV, pages 635–651, 2018.
  38. 38.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence, 39(6):1137–1149, 2016.
  39. 39.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In CVPR, pages 658–666, 2019.
  40. 40.Mikel D Rodriguez, Javed Ahmed, and Mubarak Shah. Action mach a spatio-temporal maximum average correlation height filter for action recognition. In CVPR, pages 1–8, 2008.
  41. 41.Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, and Bernt Schiele. A database for fine grained activity detection of cooking activities. In CVPR, pages 1194–1201, 2012.
  42. 42.Christian Schuldt, Ivan Laptev, and Barbara Caputo. Recognizing human actions: a local svm approach. In ICPR, volume 3, pages 32–36, 2004.
  43. 43.Abhijeet Shenoi, Mihir Patel, JunYoung Gwak, Patrick Goebel, Amir Sadeghian, Hamid Rezatofighi, Roberto Martin-Martin, and Silvio Savarese. JRMOT: A real-time 3d multi-object tracker and a new large-scale dataset. In IROS, 2020.
  44. 44.Gunnar A Sigurdsson, Santosh Divvala, Ali Farhadi, and Abhinav Gupta. Asynchronous temporal fields for action recognition. In CVPR, pages 585–594, 2017.
  45. 45.Gunnar A Sigurdsson, Gul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. Hollywood in homes: Crowdsourcing data collection for activity understanding. In ECCV, pages 510–526, 2016.
  46. 46.Karen Simonyan and Andrew Zisserman. Two-stream convolutional networks for action recognition in videos. arXiv preprint arXiv:1406.2199, 2014.
  47. 47.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.
  48. 48.Chen Sun, Abhinav Shrivastava, Carl Vondrick, Kevin Murphy, Rahul Sukthankar, and Cordelia Schmid. Actor-centric relation network. In ECCV, pages 318–334, 2018.
  49. 49.Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In CVPR, pages 1207–1216, 2019.
  50. 50.Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick. Long-term feature banks for detailed video understanding. In CVPR, pages 284–293, 2019.
  51. 51.Junyuan Xie, Ross Girshick, and Ali Farhadi. Unsupervised deep embedding for clustering analysis. In ICML, pages 478–487, 2016.
  52. 52.Huijuan Xu, Abir Das, and Kate Saenko. R-c3d: Region convolutional 3d network for temporal activity detection. In CVPR, pages 5783–5792, 2017.
  53. 53.Serena Yeung, Olga Russakovsky, Ning Jin, Mykhaylo Andriluka, Greg Mori, and Li Fei-Fei. Every moment counts: Dense detailed labeling of actions in complex videos. International Journal of Computer Vision, 126(2-4):375–389, 2018.
  54. 54.Junsong Yuan, Zicheng Liu, and Ying Wu. Discriminative subvolume search for efficient action detection. In CVPR, pages 2442–2449, 2009.
  55. 55.Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. Beyond short snippets: Deep networks for video classification. In CVPR, pages 4694–4702, 2015.
  56. 56.Lihi Zelnik-Manor and Pietro Perona. Self-tuning spectral clustering. In NIPS, pages 1601–1608, 2005.
  57. 57.Hang Zhao, Antonio Torralba, Lorenzo Torresani, and Zhicheng Yan. Hacs: Human action clips and segments dataset for recognition and temporal localization. In ICCV, pages 8668–8678, 2019.
  58. 58.Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. Temporal relational reasoning in videos. In ECCV, pages 803–818, 2018.

Citation

MLA
Ehsanpour, M., et al. “JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection”. arXiv, 2021, http://arxiv.org/abs/2106.08827v2.
APA
Ehsanpour, M., Saleh, F., Savarese, S., Reid, I., & Rezatofighi, H. (2021). JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection. arXiv. http://arxiv.org/abs/2106.08827v2
Chicago
Ehsanpour, M., F. Saleh, S. Savarese, I. Reid, and H. Rezatofighi. 2021. “JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection”. arXiv. http://arxiv.org/abs/2106.08827v2.
Harvard
Ehsanpour, M. et al. (2021) “JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.08827v2.
Vancouver
1. Ehsanpour M, Saleh F, Savarese S, Reid I, Rezatofighi H (2021) JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection. arXiv

BibTeX

@article{ehsanpour2021jrdb,
  title = {JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection},
  author = {Ehsanpour, Mahsa and Saleh, Fatemeh and Savarese, Silvio and Reid, Ian and Rezatofighi, Hamid},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.08827v2},
  eprint = {2106.08827}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE