topic
mixed reality
Mixed reality is an interactive computing environment that merges the physical world with digital content, enabling real and synthetic objects to coexist and interact in real time. Positioned along the continuum between completely real environments and fully immersive virtual reality, it encompasses technologies such as augmented reality and augmented virtuality. Mixed reality systems rely on spatial tracking, advanced display technologies, and three-dimensional interactive computer graphics to map physical surroundings and anchor responsive digital elements within the user environment, allowing natural perception and manipulation across both physical and virtual domains.
2 items

SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration
Dan Bohus, Sean Andrist, Ann Paradiso, Nick Saw, Tim Schoonbeek, Maia Stiber
Why you should read this
Introduces SigmaCollab, a multimodal dataset of interactive mixed-reality task assistance that pairs synchronized egocentric video, depth, audio, and gaze tracking to advance research on physically situated AI agents.
We introduce SigmaCollab, a dataset enabling research on physically situated human-AI collaboration. The dataset consists of a set of 85 sessions in which untrained participants were guided by a mixed-reality assistive AI agent in performing procedural tasks in the physical world. SigmaCollab includes a set of rich, multimodal data streams, such as the participant and system audio, egocentric camera views from the head-mounted device, depth maps, head, hand and gaze tracking information, as well as additional annotations performed post-hoc. While the dataset is relatively small in size (~ 14 hours), its application-driven and interactive nature brings to the fore novel research challenges for human-AI collaboration, and provides more realistic testing grounds for various AI models operating in this space. In future work, we plan to use the dataset to construct a set of benchmarks for physically situated collaboration in mixed-reality task assistive scenarios. SigmaCollab is available at this https URL.
Added
2026-09-30

Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, C. Sminchisescu
Why you should read this
Presents Human3.6M, a benchmark of 3.6 million accurate 3D poses paired with multi-view video, depth data, and high-resolution body scans, establishing standard evaluation protocols and scalable predictive baselines for 3D human pose estimation.
We introduce a new dataset, Human3.6M, of 3.6 Million accurate 3D Human poses, acquired by recording the performance of 5 female and 6 male subjects, under 4 different viewpoints, for training realistic human sensing systems and for evaluating the next generation of human pose estimation models and algorithms. Besides increasing the size of the datasets in the current state of the art by several orders of magnitude, we also aim to complement such datasets with a diverse set of motions and poses encountered as part of typical human activities (taking photos, talking on the phone, posing, greeting, eating, etc.), with additional synchronized image, human motion capture and time of flight (depth) data, and with accurate 3D body scans of all the subject actors involved. We also provide controlled mixed reality evaluation scenarios where 3D human models are animated using motion capture and inserted using correct 3D geometry, in complex real environments, viewed with moving cameras, and under occlusion. Finally, we provide a set of large scale statistical models and detailed evaluation baselines for the dataset illustrating its diversity and the scope for improvement by future work in the research community. Our experiments show that our best large scale model can leverage our full training set to obtain a 20% improvement in performance compared to a training set of the scale of the largest existing public dataset for this problem. Yet the potential for improvement by leveraging higher capacity, more complex models with our large dataset, is substantially vaster and should stimulate future research. The dataset together with code for the associated large-scale learning models, features, visualization tools, as well as the evaluation server, is available online at http://vision.imar.ro/human3.6m.
Added
2026-09-11
