JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection
Mahsa EhsanpourFatemeh Sadat SalehSilvio SavareseIan D. ReidHamid Rezatofighi
Presents JRDB-Act, a large-scale multimodal benchmark featuring over 2.8 million action labels alongside social group annotations and confidence ratings captured from a mobile robot, establishing an end-to-end framework for joint individual action and group activity detection in crowded real-world environments.
Deploying mobile robots and autonomous systems in crowded human environments requires machines to accurately perceive not only individual human movements but also social interactions and collective group behaviors. Existing action recognition benchmarks generally rely on static cameras or controlled settings, failing to capture the complexities faced by moving robotic platforms navigating crowded, unconstrained spaces.
The article introduces JRDB-Act, a large-scale multimodal dataset designed to benchmark individual action recognition, social group detection, and collective activity estimation from a mobile robot platform. It also presents and evaluates a unified learning pipeline tailored to handle the dataset’s dense crowds, moving perspective, and highly unbalanced real-world action distributions.
The dataset expands upon 64 minutes of sensory footage (spanning 54 campus video sequences captured with 360-degree cylindrical cameras and LiDAR) by contributing over 2.8 million dense spatio-temporal action annotations across 26 distinct daily action classes. Scenes feature an average density of 30 people per frame, and each annotation includes annotator confidence ratings ranging from easy to difficult. To model this data, the authors developed a baseline framework that integrates geometric and visual feature relations, an eigenvalue-based spectral loss for group cardinality estimation, and a partitioned loss mechanism to mitigate long-tailed class imbalances.
Evaluation reveals four core findings. First, the proposed framework improved overall social group detection accuracy to 59.2% average precision on ground-truth boxes, outperforming existing baseline methods by approximately 17.8 percentage points. Second, the loss partitioning approach boosted individual action recognition from 8.0% to 9.0% mean average precision, outperforming standard class-frequency weighting schemes. Third, when evaluating end-to-end system performance on test sequences with detected bounding boxes, better object localization significantly raised social group detection (from 29.1% to 31.5% average precision) but had negligible impact on action recognition, which remained near 5.4% mean average precision. Fourth, annotator difficulty tags substantially influence performance benchmarks, as models scored roughly 37% higher on social grouping when evaluated solely on visually distinct (easy) labels compared to full sets containing heavily occluded (difficult) targets.
These findings demonstrate that safe robotic navigation and human-robot interaction cannot rely on conventional video classification models. Action recognition in mobile crowd settings is limited by visual perspective, rapid camera motion, occlusion, and severe label imbalance rather than bounding box localization alone. Practical deployment requires perception architectures that can reason about geometric relationships and long-tailed distributions simultaneously.
For future development, the source recommends exploring multi-modal sensor fusion by directly integrating the available 3D LiDAR point-cloud data into feature extractors to enhance spatial reasoning. Researchers should also adopt partitioned loss formulations rather than standard cost-sensitive weighting when handling long-tailed behavioral data.
The primary limitations involve the overall low baseline scores on complex real-world actions (reaching 5.4% test mean average precision) and the fact that 38.6% of annotations involve difficult or partially occluded cases inferred from movement history. Stakeholders should view the presented framework as a foundational research baseline rather than an off-the-shelf system ready for immediate autonomous deployment.
- Paper: Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks, Agrim Gupta et al. (2018). Introduces foundational methods for modeling pedestrian interactions and social dynamics in crowded spaces, providing key conceptual background for robotic social group and activity detection.
- Paper: MOT16: A Benchmark for Multi-Object Tracking, Anton Milan et al. (2016). Establishes standardized benchmarking and evaluation protocols for multi-object tracking in dense pedestrian crowds that inform spatio-temporal detection pipelines.
- Paper: Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset, João Carreira et al. (2017). Provides the foundational inflated 3D convolutional architecture (I3D) and large-scale video action recognition methodology utilized across modern spatio-temporal action understanding baselines.
- Paper: Ego4D: Around the World in 3,000 Hours of Egocentric Video, Kristen Grauman et al. (2021). Presents large-scale egocentric multimodal video benchmarking for social interactions and human activity, offering direct precedent for dynamic camera platforms in unscripted environments.
- Paper: Argoverse: 3D Tracking and Forecasting With Rich Maps, Ming-Fang Chang et al. (2019). Demonstrates multimodal perception and tracking using synchronized 360-degree cameras and LiDAR on moving robotic platforms in outdoor environments.
- Paper: ActivityNet: A large-scale video benchmark for human activity understanding, Fabian Caba Heilbron et al. (2015). Establishes core paradigms for untrimmed temporal activity detection and categorization across complex real-world daily human behaviors.
- Paper: Temporal Segment Networks: Towards Good Practices for Deep Action Recognition, Limin Wang et al. (2016). Introduces segment-based temporal modeling and training practices for video action recognition that underpin spatio-temporal action detection backbones.
- Paper: Pedestrian Detection: An Evaluation of the State of the Art, Piotr Dollár et al. (2012). Establishes standardized evaluation protocols for pedestrian localization and handling scale and occlusion from moving cameras.
- Paper: TAPVid-3D: A Benchmark for Tracking Any Point in 3D, Skanda Koppula et al. (2024). Extends long-range dynamic tracking in real-world mobile and multi-view video settings into full 3D point tracking across metric spatial scenes.
- Paper: HourVideo: 1-Hour Video-Language Understanding, Keshigeyan Chandrasegaran et al. (2024). Advances the evaluation of complex, long-horizon human activity understanding and multimodal temporal reasoning across extended real-world egocentric video streams.
- Paper: Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps, Yue Hu et al. (2022). Applies multi-sensor and multimodal spatial awareness to collaborative perception architectures across networked autonomous mobile platforms.
