TEACh: Task-Driven Embodied Agents That Chat
Aishwarya PadmakumarJesse ThomasonAyush ShrivastavaPatrick LangeAnjali Narayan-ChenSpandana GellaRobinson PiramuthuGökhan TürDilek Hakkani-Tür
Introduces TEACh, a dataset and benchmark suite of over 3,000 simulated human-human dialogues that pairs freeform interactive communication with physical household object manipulation to advance conversational embodied AI agents.
Operating assistive robots in human domestic environments requires agents to understand complex instructions, navigate dynamic spaces, interact with household objects, and hold conversational dialogues to clarify ambiguities or correct mistakes. Prior research predominantly focused on navigation without dialogue or on planner-based demonstrations with rigid, turn-based single commands. The article addresses this gap by investigating how embodied agents can coordinate through natural, interactive communication to achieve multi-step, object-centric household goals.
The main objective of the article is to introduce and evaluate the Task-driven Embodied Agents that Chat (TEACh) dataset, an extensible task definition framework, and three distinct benchmark tasks designed to advance grounded language understanding and interactive execution in simulated environments.
To construct this resource, the authors paired human crowdworkers in the AI2-THOR virtual simulator across 12 distinct household chore types, spanning 30 rooms each across kitchens, living rooms, bedrooms, and bathrooms. One user acted as a Commander with privileged access to task requirements and search tools, while the other acted as a Follower executing physical actions and asking questions via unconstrained text chat. Out of 4,365 crowdsourced sessions collected at a cost of $105,000, 3,047 successful sessions were retained, comprising over 45,000 utterances and 3,429 unique vocabulary terms. Using this dataset, the authors established three benchmarks: Execution from Dialogue History (EDH), Trajectory from Dialogue (TfD), and Two-Agent Task Completion (TATC), and evaluated baseline machine learning models adapting the Episodic Transformer (E.T.) architecture alongside hand-crafted rule-based agents.
The findings demonstrate a significant performance gap between current artificial intelligence capabilities and human-level task execution. First, while human pairs achieved a 74.17% task success rate, the baseline E.T. model achieved a maximum success rate of only 15.62% in seen environments and 13.49% in unseen environments on the sub-trajectory EDH task. Second, on the full-trajectory TfD benchmark, model performance fell below 2% across all splits, reflecting severe difficulties with long-horizon prediction involving average trajectory lengths of roughly 130 steps. Third, unimodal ablations revealed that vision-only models performed comparably to multimodal models on seen validation splits, indicating that baseline neural architectures struggle to effectively exploit rich dialogue information. Finally, 150 hours of hand-crafting rule-based policies for the TATC task yielded an overall success rate of only 24.40%, with complex composite tasks such as preparing breakfast or sandwiches recording a 0.00% success rate.
These results demonstrate that standard planning algorithms and existing transformer architectures designed for simple, sequential benchmarks cannot scale to realistic, collaborative household tasks. Unconstrained human dialogue introduces conversational complexities—such as references to past actions, irrelevant pleasantries, and mid-task corrections—that existing models fail to parse and ground effectively. This gap represents a major risk for operational reliability and performance if developers rely on traditional robotic planning paradigms for household assistance.
The authors recommend that future research prioritize developing models capable of few-shot generalization across novel tasks, advancing speaker models to simulate Commander dialogue, and integrating human-in-the-loop evaluations. Automated systems require improved temporal alignment between dialogue acts and physical environment actions to resolve conversational context effectively.
The presented benchmarks have notable limitations: evaluations remain restricted to simulated environments, the initial baseline model omitted direct temporal alignment between utterances and actions, and data cleaning exclusions were required due to simulation replay issues. Nonetheless, the evidence strongly confirms that interactive, dialogue-grounded execution presents fundamental challenges that require new modeling approaches before reliable home robotics can be realized.
- Paper: Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments, Peter Anderson et al. (2017). This seminal benchmark establishes the foundations of vision-and-language navigation in simulated indoor environments, providing direct context for TEACh's interactive navigation and dialogue setting.
- Paper: MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling, Paweł Budzianowski et al. (2018). It provides essential background on Wizard-of-Oz multi-turn dialogue collection and task-oriented conversation modeling that TEACh adapts for embodied agent interaction.
- Paper: CoQA: A Conversational Question Answering Challenge, Siva Reddy et al. (2018). It introduces foundational methodologies for multi-turn conversational context tracking and information exchange between two participants that underpin TEACh's Commander-Follower dialogue setup.
- Paper: Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding, Gunnar A. Sigurdsson et al. (2016). It offers foundational methodologies for crowdsourcing complex, multi-step household action routines in simulated and real domestic spaces.
- Paper: PIQA: Reasoning about Physical Commonsense in Natural Language, Yonatan Bisk et al. (2019). It formalizes benchmarks for physical commonsense reasoning in goal-oriented everyday tasks, which embodied dialogue agents in TEACh must resolve during execution.
- Paper: Do As I Can, Not As I Say: Grounding Language in Robotic Affordances, Michael Ahn et al. (2022). It builds on the problem of grounding natural language in feasible actions by pairing large language models with affordance value functions for embodied physical control.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). It advances TEACh's interactive task execution paradigm by using continuous textual and perceptual feedback to achieve closed-loop replanning in embodied settings.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). It extends the execution of household instructions by exploring zero-shot task decomposition and action translation using pretrained language models in interactive environments.
- Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). It translates high-level task instructions into hierarchical, executable programmatic policies for embodied agents operating in domestic spaces.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). It develops an end-to-end multimodal language model that integrates continuous sensor inputs directly with language reasoning for embodied decision-making and planning.
- Paper: EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought, Yao Mu et al. (2023). It extends language-guided robotic planning by introducing embodied chain-of-thought vision-language pretraining to generate step-by-step action plans.
- Paper: CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Guohao Li et al. (2023). It formalizes autonomous multi-agent communicative problem solving through cooperative dialogue, expanding on the Commander-Follower paradigm established in TEACh.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). It scales vision-language models to output robotic action tokens directly, demonstrating how Internet-scale semantic knowledge transfers to physical task execution.
