Built independently by an author, for readers. Read the story and support ChapterPal

topic

mixed reality

Mixed reality is an interactive computing environment that merges the physical world with digital content, enabling real and synthetic objects to coexist and interact in real time. Positioned along the continuum between completely real environments and fully immersive virtual reality, it encompasses technologies such as augmented reality and augmented virtuality. Mixed reality systems rely on spatial tracking, advanced display technologies, and three-dimensional interactive computer graphics to map physical surroundings and anchor responsive digital elements within the user environment, allowing natural perception and manipulation across both physical and virtual domains.

2 items

SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration

SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration

Dan Bohus, Sean Andrist, Ann Paradiso, Nick Saw, Tim Schoonbeek, Maia Stiber

OrganizationsEindhoven University of TechnologyMicrosoft

Why you should read this

Introduces SigmaCollab, a multimodal dataset of interactive mixed-reality task assistance that pairs synchronized egocentric video, depth, audio, and gaze tracking to advance research on physically situated AI agents.

We introduce SigmaCollab, a dataset enabling research on physically situated human-AI collaboration. The dataset consists of a set of 85 sessions in which untrained participants were guided by a mixed-reality assistive AI agent in performing procedural tasks in the physical world. SigmaCollab includes a set of rich, multimodal data streams, such as the participant and system audio, egocentric camera views from the head-mounted device, depth maps, head, hand and gaze tracking information, as well as additional annotations performed post-hoc. While the dataset is relatively small in size (~ 14 hours), its application-driven and interactive nature brings to the fore novel research challenges for human-AI collaboration, and provides more realistic testing grounds for various AI models operating in this space. In future work, we plan to use the dataset to construct a set of benchmarks for physically situated collaboration in mixed-reality task assistive scenarios. SigmaCollab is available at this https URL.

Added

2026-09-30

Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments

Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments

Catalin Ionescu, Dragos Papava, Vlad Olaru, C. Sminchisescu

OrganizationsInstitute of Mathematics of the Romanian AcademyLund UniversityUniversity of Bonn

Why you should read this

Presents Human3.6M, a benchmark of 3.6 million accurate 3D poses paired with multi-view video, depth data, and high-resolution body scans, establishing standard evaluation protocols and scalable predictive baselines for 3D human pose estimation.

We introduce a new dataset, Human3.6M, of 3.6 Million accurate 3D Human poses, acquired by recording the performance of 5 female and 6 male subjects, under 4 different viewpoints, for training realistic human sensing systems and for evaluating the next generation of human pose estimation models and algorithms. Besides increasing the size of the datasets in the current state of the art by several orders of magnitude, we also aim to complement such datasets with a diverse set of motions and poses encountered as part of typical human activities (taking photos, talking on the phone, posing, greeting, eating, etc.), with additional synchronized image, human motion capture and time of flight (depth) data, and with accurate 3D body scans of all the subject actors involved. We also provide controlled mixed reality evaluation scenarios where 3D human models are animated using motion capture and inserted using correct 3D geometry, in complex real environments, viewed with moving cameras, and under occlusion. Finally, we provide a set of large scale statistical models and detailed evaluation baselines for the dataset illustrating its diversity and the scope for improvement by future work in the research community. Our experiments show that our best large scale model can leverage our full training set to obtain a 20% improvement in performance compared to a training set of the scale of the largest existing public dataset for this problem. Yet the potential for improvement by leveraging higher capacity, more complex models with our large dataset, is substantially vaster and should stimulate future research. The dataset together with code for the associated large-scale learning models, features, visualization tools, as well as the evaluation server, is available online at http://vision.imar.ro/human3.6m.

Added

2026-09-11