SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration
Dan BohusSean AndristAnn ParadisoNick SawTim SchoonbeekMaia Stiber
Introduces SigmaCollab, a multimodal dataset of interactive mixed-reality task assistance that pairs synchronized egocentric video, depth, audio, and gaze tracking to advance research on physically situated AI agents.
Building artificial intelligence systems capable of collaborating fluidly with humans in the physical world requires addressing challenges beyond traditional computer vision, such as real-time temporal coordination and understanding human cognitive states. Existing egocentric datasets typically capture solo activities or human-to-human guidance, failing to reflect how people naturally interact with an autonomous AI assistant. To bridge the gap between laboratory benchmarks and real-world deployment, the article introduces SIGMACOLLAB, an open-source, application-driven dataset designed to advance research on physically situated human-AI collaboration.
The dataset was created by having 21 untrained participants perform procedural tasks while receiving step-by-step guidance from SIGMA, a mixed-reality assistive AI running on a HoloLens 2 headset. The study captured rich, multimodal data streams across eight diverse physical tasks, including craft-making, computer hardware assembly, and beverage preparation. The final dataset spans 85 valid interactive sessions totaling roughly 13 hours and 45 minutes, containing 1,583 sub-step executions and 3,296 user utterances, alongside eight expert demonstration sessions.
The investigation revealed several key findings: First, overall task success reached 75.0% across the 85 sessions, though performance varied widely depending on task complexity—from 100% on PC hard-drive replacement to 40.0% on complex mocktail preparation. Second, runtime audio processing introduced substantial errors, showing an average voice activity detection error rate of 12.4% and an average word error rate of 20.2%, prompting the researchers to generate manual transcripts and post-hoc word-level alignments. Third, realistic interactive settings uncovered unique behavioral patterns, such as participants frequently talking to themselves during execution, which highlights the critical need for AI assistants to distinguish self-talk from actionable user queries.
These findings demonstrate that deploying AI agents in physical collaboration environments fundamentally shifts user behavior, introducing natural speech fragments, complex referential questions, and cognitive friction that static or human-human datasets fail to capture. To support accurate benchmarking and system improvements, researchers should leverage the dataset's synchronized multimodal streams—including eye gaze, 3D hand tracking, egocentric video, and post-processed transcripts—to train and evaluate models on proactive intervention, mistake detection, and cognitive state inference.
While the dataset provides strong ecological validity, its limitations include a relatively modest size of approximately 14 hours, a controlled laboratory setting, and minor latency variations across three different underlying language model deployments. Consequently, findings should be applied with caution when generalizing to open, unconstrained environments, and future work will focus on establishing standardized collaborative benchmarks and integrating refined interaction models back into the open-source SIGMA system.
- Paper: TEACh: Task-Driven Embodied Agents That Chat, Aishwarya Padmakumar et al. (2022). TEACh establishes how interactive dialogue can coordinate embodied agents through multi-step household tasks, providing a key predecessor for SigmaCollab’s study of collaboration in physical settings.
- Paper: Ego4D: Around the World in 3,000 Hours of Egocentric Video, Kristen Grauman et al. (2021). Ego4D develops large-scale first-person data collection with audio, gaze, and other streams, clarifying the egocentric sensing foundations that SigmaCollab adapts to interactive assistance.
- Paper: Scaling Egocentric Vision: The EPIC-KITCHENS Dataset, Dima Damen et al. (2018). EPIC-KITCHENS shows how head-mounted video and participant audio can capture natural activity in real environments, a useful precedent for SigmaCollab’s procedural-task recordings.
- Paper: From Gaze to Guidance: Interpreting and Adapting to Users' Cognitive Needs with Multimodal Gaze-Aware AI Assistants, Valdemar Danry et al. (2026). Building on SigmaCollab’s multimodal, physically situated assistance setting, this study tests whether gaze-aware AI can adapt guidance to users’ cognitive needs.
