SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration

Dan BohusSean AndristAnn ParadisoNick SawTim SchoonbeekMaia Stiber

article2025arXiv2 citations

Introduces SigmaCollab, a multimodal dataset of interactive mixed-reality task assistance that pairs synchronized egocentric video, depth, audio, and gaze tracking to advance research on physically situated AI agents.

Listen

Building artificial intelligence systems capable of collaborating fluidly with humans in the physical world requires addressing challenges beyond traditional computer vision, such as real-time temporal coordination and understanding human cognitive states. Existing egocentric datasets typically capture solo activities or human-to-human guidance, failing to reflect how people naturally interact with an autonomous AI assistant. To bridge the gap between laboratory benchmarks and real-world deployment, the article introduces SIGMACOLLAB, an open-source, application-driven dataset designed to advance research on physically situated human-AI collaboration.

The dataset was created by having 21 untrained participants perform procedural tasks while receiving step-by-step guidance from SIGMA, a mixed-reality assistive AI running on a HoloLens 2 headset. The study captured rich, multimodal data streams across eight diverse physical tasks, including craft-making, computer hardware assembly, and beverage preparation. The final dataset spans 85 valid interactive sessions totaling roughly 13 hours and 45 minutes, containing 1,583 sub-step executions and 3,296 user utterances, alongside eight expert demonstration sessions.

The investigation revealed several key findings: First, overall task success reached 75.0% across the 85 sessions, though performance varied widely depending on task complexity—from 100% on PC hard-drive replacement to 40.0% on complex mocktail preparation. Second, runtime audio processing introduced substantial errors, showing an average voice activity detection error rate of 12.4% and an average word error rate of 20.2%, prompting the researchers to generate manual transcripts and post-hoc word-level alignments. Third, realistic interactive settings uncovered unique behavioral patterns, such as participants frequently talking to themselves during execution, which highlights the critical need for AI assistants to distinguish self-talk from actionable user queries.

These findings demonstrate that deploying AI agents in physical collaboration environments fundamentally shifts user behavior, introducing natural speech fragments, complex referential questions, and cognitive friction that static or human-human datasets fail to capture. To support accurate benchmarking and system improvements, researchers should leverage the dataset's synchronized multimodal streams—including eye gaze, 3D hand tracking, egocentric video, and post-processed transcripts—to train and evaluate models on proactive intervention, mistake detection, and cognitive state inference.

While the dataset provides strong ecological validity, its limitations include a relatively modest size of approximately 14 hours, a controlled laboratory setting, and minor latency variations across three different underlying language model deployments. Consequently, findings should be applied with caution when generalizing to open, unconstrained environments, and future work will focus on establishing standardized collaborative benchmarks and integrating refined interaction models back into the open-source SIGMA system.

  • Paper: TEACh: Task-Driven Embodied Agents That Chat, Aishwarya Padmakumar et al. (2022). TEACh establishes how interactive dialogue can coordinate embodied agents through multi-step household tasks, providing a key predecessor for SigmaCollab’s study of collaboration in physical settings.
  • Paper: Ego4D: Around the World in 3,000 Hours of Egocentric Video, Kristen Grauman et al. (2021). Ego4D develops large-scale first-person data collection with audio, gaze, and other streams, clarifying the egocentric sensing foundations that SigmaCollab adapts to interactive assistance.
  • Paper: Scaling Egocentric Vision: The EPIC-KITCHENS Dataset, Dima Damen et al. (2018). EPIC-KITCHENS shows how head-mounted video and participant audio can capture natural activity in real environments, a useful precedent for SigmaCollab’s procedural-task recordings.
Cover for SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration

Abstract

We introduce SigmaCollab, a dataset enabling research on physically situated human-AI collaboration. The dataset consists of a set of 85 sessions in which untrained participants were guided by a mixed-reality assistive AI agent in performing procedural tasks in the physical world. SigmaCollab includes a set of rich, multimodal data streams, such as the participant and system audio, egocentric camera views from the head-mounted device, depth maps, head, hand and gaze tracking information, as well as additional annotations performed post-hoc. While the dataset is relatively small in size (~ 14 hours), its application-driven and interactive nature brings to the fore novel research challenges for human-AI collaboration, and provides more realistic testing grounds for various AI models operating in this space. In future work, we plan to use the dataset to construct a set of benchmarks for physically situated collaboration in mixed-reality task assistive scenarios. SigmaCollab is available at this https URL.

Citation

MLA
Bohus, D., et al. “SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration”. arXiv, 2025, http://arxiv.org/abs/2511.02560v1.
APA
Bohus, D., Andrist, S., Paradiso, A., Saw, N., Schoonbeek, T., & Stiber, M. (2025). SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration. arXiv. http://arxiv.org/abs/2511.02560v1
Chicago
Bohus, D., S. Andrist, A. Paradiso, N. Saw, T. Schoonbeek, and M. Stiber. 2025. “SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration”. arXiv. http://arxiv.org/abs/2511.02560v1.
Harvard
Bohus, D. et al. (2025) “SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2511.02560v1.
Vancouver
1. Bohus D, Andrist S, Paradiso A, Saw N, Schoonbeek T, Stiber M (2025) SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration. arXiv

BibTeX

@article{bohus2025sigmacollab,
  title = {SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration},
  author = {Bohus, Dan and Andrist, Sean and Paradiso, Ann and Saw, Nick and Schoonbeek, Tim and Stiber, Maia},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2511.02560v1},
  eprint = {2511.02560}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/