TEACh: Task-Driven Embodied Agents That Chat

Aishwarya PadmakumarJesse ThomasonAyush ShrivastavaPatrick LangeAnjali Narayan-ChenSpandana GellaRobinson PiramuthuGökhan TürDilek Hakkani-Tür

article2022AAAI300 citations

Introduces TEACh, a dataset and benchmark suite of over 3,000 simulated human-human dialogues that pairs freeform interactive communication with physical household object manipulation to advance conversational embodied AI agents.

Listen

Operating assistive robots in human domestic environments requires agents to understand complex instructions, navigate dynamic spaces, interact with household objects, and hold conversational dialogues to clarify ambiguities or correct mistakes. Prior research predominantly focused on navigation without dialogue or on planner-based demonstrations with rigid, turn-based single commands. The article addresses this gap by investigating how embodied agents can coordinate through natural, interactive communication to achieve multi-step, object-centric household goals.

The main objective of the article is to introduce and evaluate the Task-driven Embodied Agents that Chat (TEACh) dataset, an extensible task definition framework, and three distinct benchmark tasks designed to advance grounded language understanding and interactive execution in simulated environments.

To construct this resource, the authors paired human crowdworkers in the AI2-THOR virtual simulator across 12 distinct household chore types, spanning 30 rooms each across kitchens, living rooms, bedrooms, and bathrooms. One user acted as a Commander with privileged access to task requirements and search tools, while the other acted as a Follower executing physical actions and asking questions via unconstrained text chat. Out of 4,365 crowdsourced sessions collected at a cost of $105,000, 3,047 successful sessions were retained, comprising over 45,000 utterances and 3,429 unique vocabulary terms. Using this dataset, the authors established three benchmarks: Execution from Dialogue History (EDH), Trajectory from Dialogue (TfD), and Two-Agent Task Completion (TATC), and evaluated baseline machine learning models adapting the Episodic Transformer (E.T.) architecture alongside hand-crafted rule-based agents.

The findings demonstrate a significant performance gap between current artificial intelligence capabilities and human-level task execution. First, while human pairs achieved a 74.17% task success rate, the baseline E.T. model achieved a maximum success rate of only 15.62% in seen environments and 13.49% in unseen environments on the sub-trajectory EDH task. Second, on the full-trajectory TfD benchmark, model performance fell below 2% across all splits, reflecting severe difficulties with long-horizon prediction involving average trajectory lengths of roughly 130 steps. Third, unimodal ablations revealed that vision-only models performed comparably to multimodal models on seen validation splits, indicating that baseline neural architectures struggle to effectively exploit rich dialogue information. Finally, 150 hours of hand-crafting rule-based policies for the TATC task yielded an overall success rate of only 24.40%, with complex composite tasks such as preparing breakfast or sandwiches recording a 0.00% success rate.

These results demonstrate that standard planning algorithms and existing transformer architectures designed for simple, sequential benchmarks cannot scale to realistic, collaborative household tasks. Unconstrained human dialogue introduces conversational complexities—such as references to past actions, irrelevant pleasantries, and mid-task corrections—that existing models fail to parse and ground effectively. This gap represents a major risk for operational reliability and performance if developers rely on traditional robotic planning paradigms for household assistance.

The authors recommend that future research prioritize developing models capable of few-shot generalization across novel tasks, advancing speaker models to simulate Commander dialogue, and integrating human-in-the-loop evaluations. Automated systems require improved temporal alignment between dialogue acts and physical environment actions to resolve conversational context effectively.

The presented benchmarks have notable limitations: evaluations remain restricted to simulated environments, the initial baseline model omitted direct temporal alignment between utterances and actions, and data cleaning exclusions were required due to simulation replay issues. Nonetheless, the evidence strongly confirms that interactive, dialogue-grounded execution presents fundamental challenges that require new modeling approaches before reliable home robotics can be realized.

arXiv: 2110.00534
Cover for TEACh: Task-Driven Embodied Agents That Chat

Abstract

Robots operating in human spaces must be able to engage in natural language interaction, both understanding and executing instructions, and using conversation to resolve ambiguity and correct mistakes. To study this, we introduce TEACh, a dataset of over 3,000 human–human, interactive dialogues to complete household tasks in simulation. A Commander with access to oracle information about a task communicates in natural language with a Follower. The Follower navigates through and interacts with the environment to complete tasks varying in complexity from MAKE COFFEE to PREPARE BREAKFAST, asking questions and getting additional information from the Commander. We propose three benchmarks using TEACh to study embodied intelligence challenges, and we evaluate initial models’ abilities in dialogue understanding, language grounding, and task execution.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The TEACh Dataset
  • 3.1 Household Tasks
  • 3.2 Gameplay Session Collection
  • 3.3 TEACh Statistics
  • 4 TEACh Benchmarks
  • 4.1 Execution from Dialogue History (EDH)
  • 4.2 Trajectory from Dialogue (TfD)
  • 4.3 Two-Agent Task Completion (TATC)
  • 5 Experiments and Results
  • 5.1 Follower Models for EDH and TfD
  • 5.2 Rule-based Agents for TATC
  • 6 Conclusions and Future Work
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — TEACh situated human-human dialogue dataset

    data/table

    TEACh is a dataset of human-human gameplay sessions in the AI2-THOR household simulator. A human Commander has oracle access to the task specification, object locations, and scene map, while a human Follower navigates and manipulates objects; the two agents communicate only through written English. Dialogue actions and Follower environment actions are interleaved in one session, so the data contains object references, action references, corrections after mistakes, incomplete instructions, irrelevant utterances, and non-turn-taking conversational overlap. The page-1 visualization illustrates this division of information: the Commander sees task and search information, whereas the Follower acts from an egocentric view and asks questions.

    The paper reports 4,365 crowdsourced sessions, of which 3,320 were successful according to the human-level collection success rate of 74.17%, at a total reported cost of $105k. It describes 3,047 successful replayable gameplay sessions in the released corpus and notes that some successful sessions were excluded from benchmark splits because of replay problems. The sessions span all 30 AI2-THOR kitchens and most of the 30 rooms in each of the living-room, bedroom, and bathroom categories. The successful sessions contain more than 45,000 utterances, averaging 8.40 Commander utterances and 5.25 Follower utterances per session; mean utterance lengths are 5.70 tokens for Commanders and 3.80 tokens for Followers, with a 3,429-token vocabulary.

  2. Knowl 2 — Extensible hierarchical task definition language

    model/method

    TEACh introduces a task definition language and simulator framework that specify household goals as object properties that must hold in the final environment state. For example, MAKE COFFEE is successful when some mug is clean and contains coffee. Parameterized templates support goals such as PUT ALL X ON Y, where XX may denote an object class such as forks or an abstract hypernym such as silverware. Determiners such as “a,” “all,” and explicit counts support definitions such as N SLICES OF X IN Y.

    Task definitions can be hierarchical: PREPARE BREAKFAST can contain MAKE COFFEE and MAKE PLATE OF TOAST as subtasks. The framework evaluates whether the simulator state satisfies the specified conditions and supplies template-based descriptions of tasks and subtasks to the Commander. The task language is intended to be extensible to additional tasks and potentially to other simulators.

  3. Knowl 3 — Human gameplay collection protocol and action interface

    experimental setup

    For each TEACh session, two vetted crowdworkers are assigned Commander and Follower roles after completing a tutorial. The Commander receives the task and required steps through template prompts that are hidden from the Follower, and can query object locations by text search or by selecting a task-relevant object. The Follower must infer task parameters and locations through free-form chat; only the Follower can interact with the environment. The page-3 collection interface shows the Commander’s top-down map and object-search views alongside the Follower’s egocentric scene and interaction target.

    Initial simulator states are randomized and retained only when task preconditions, such as the presence and reachability of relevant objects, are satisfied. A session stores an initial simulator state SiS_i, a sequence of actions A=(a1,a2,…)A=(a_1,a_2,\ldots), and a final state SfS_f. Follower actions include Forward, Backward, Turn Left, Turn Right, Look Up, Look Down, Strafe Left, Strafe Right, Pickup, Place, Open, Close, ToggleOn, ToggleOff, Slice, and Pour. Navigation is discrete. For object manipulation, the Follower supplies a relative egocentric coordinate (x,y)(x,y); the simulator wrapper examines the ground-truth segmentation mask in a 10×1010\times10 pixel patch around that coordinate and executes the requested action on a permissible object. Commander-only actions include Progress Check and SearchObject, and the Commander is represented as a disembodied camera.

  4. Knowl 4 — Three TEACh benchmark formulations

    model/method

    TEACh defines three benchmark tasks from the same sessions. EDH, or Execution from Dialogue History, evaluates a Follower that receives a partial dialogue-and-action history and must execute the remaining interaction trajectory. TfD, or Trajectory from Dialogue, evaluates a Follower that receives the complete dialogue history and must generate the entire sequence of its environmental actions. TATC, or Two-Agent Task Completion, gives only environment observations to learned Commander and Follower agents; the Commander must query task information with Progress Check and communicate it incrementally, while the Follower can respond through language and act in the environment.

    A session has initial state SiS_i, combined action sequence AA, final state SfS_f, dialogue subsequence ADA_D, and navigation/interaction subsequence AIA_I. Validation and test data are divided into seen rooms, whose AI2-THOR environments occur in training, and unseen rooms, which are absent from training. The reported split counts are:

    Could not parse LaTeX table
  5. Knowl 5 — EDH evaluation criteria and rollout limits

    equation

    An EDH instance is represented as (SE,AH,ARI,FE)(S^E,A_H,A_R^I,F^E), where SES^E is the instance’s initial simulator state, AHA_H is the observed action history, ARIA_R^I is the reference sequence of remaining navigation and interaction actions, and FEF^E is the set of expected state changes after those actions. The instance must contain at least one dialogue action in its history and at least one reference object interaction. A Stop action is appended to every reference sequence.

    After a model rollout, let E^\hat E be the resulting simulator state and let A^I\hat A_I be the inferred environmental-action sequence. Success is 1 if every expected state change in FEF^E appears in E^\hat E, and 0 otherwise. Goal-Condition Success is the fraction of expected state changes in FEF^E that appear in E^\hat E. For either metric mm, the trajectory-length-weighted score is

    TLW-m=m∣ARI∣max⁡(∣ARI∣,∣A^I∣).\mathrm{TLW}\text{-}m=m\frac{|A_R^I|}{\max(|A_R^I|,|\hat A_I|)}.

    Here ∣ARI∣|A_R^I| and ∣A^I∣|\hat A_I| are reference and inferred environmental-trajectory lengths. During inference, a Follower rollout stops when the model predicts Stop, reaches 1,000 steps, or accumulates 30 failed actions. TfD and TATC use the same state-change-based Success and Goal-Condition Success criteria, comparing the inferred final state with the reference final state.

  6. Knowl 6 — Adapted Episodic Transformer Follower baseline

    model/method

    The EDH and TfD baselines adapt the Episodic Transformer (E.T.) originally developed for ALFRED. A transformer language encoder processes dialogue, a ResNet-50 encodes visual observations, and two multimodal transformer layers fuse language, image, and action embeddings before predicting the next action and, for interactions, an object category. A Mask R-CNN detector pretrained on ALFRED predicts an egocentric object mask; the centroid of the predicted mask is converted into the relative coordinate expected by the TEACh simulator wrapper.

    For TEACh, the action head is replaced to predict the TEACh Follower action space. All dialogue utterances in the observed history are concatenated with separators as the language input, while previous environmental actions and their images are supplied in temporal order as action history. Dialogue and environmental actions are therefore not temporally aligned in the adapted input representation. The model predicts the complete reference action sequence and is trained with cross-entropy loss. For EDH, an optional history loss adds cross-entropy over both observed-history actions and remaining reference actions; for TfD, the environmental action history is empty. At rollout, each predicted action is executed in the simulator, the resulting image is appended to the input, and the next action is predicted iteratively. The authors also test initialization from ALFRED weights trained on base annotations or synthetic instructions, although changed vocabulary and action spaces require retraining of some layers.

  7. Knowl 7 — EDH and TfD baseline performance

    empirical result

    The adapted E.T. model substantially outperforms random and language-only baselines on EDH, but absolute performance remains low. The reported values are percentages; SR is Success Rate, GC is Goal-Condition Success, and bracketed values are trajectory-length-weighted versions of the corresponding metric.

    Could not parse LaTeX table

    For TfD, the random baseline scores 0.00 on every metric and split. E.T. obtains validation Seen SR 1.02 [0.17] and GC 1.42 [4.82], validation Unseen SR 0.48 [0.12] and GC 0.35 [0.59], test Seen SR 0.51 [0.23] and GC 1.60 [6.46], and test Unseen SR 0.17 [0.04] and GC 0.67 [2.50]. All E.T. EDH conditions are significantly better than Random and Lang-Only on SR and GC across splits using paired two-sided Welch tests with Bonferroni correction. E.T. improvements over Vision-Only are significant on unseen but not seen splits. ALFRED initialization does not yield statistically significant gains. TfD performance is nonzero but extremely poor, consistent with its much longer average trajectories of approximately 130 actions versus approximately 50 in ALFRED.

  8. Knowl 8 — TATC rule-based Commander-Follower construction

    algorithm

    The authors construct non-learning Commander and Follower agents as an initial TATC baseline. The Commander repeatedly executes Progress Check, identifies the next subgoal from the returned task state, and sends the Follower a templated sequence of navigation and interaction actions. Each subgoal type requires a manually written policy. For example, a PUT ALL X ON Y subgoal is handled by locating a particular XX instance, navigating to it, picking it up, navigating to a permissible YY, and placing XX on YY. Commander messages are simplified sequences of action names with a one-to-one mapping to Follower actions; interaction actions additionally contain screen coordinates. The Follower executes the received sequence in the simulator.

    The construction required approximately 150 hours of engineering. It has no learned language understanding, perception policy, or task policy; it relies on hand-authored subgoal templates and simulator-specific action rules. Policies were not successfully developed for BOIL POTATO, MAKE PLATE OF TOAST, MAKE SANDWICH, or PREPARE BREAKFAST.

  9. Knowl 9 — TATC rule-agent results expose compositional difficulty

    data/table

    The rule-based TATC agents solve some simple task types roughly half the time but fail entirely on several compositional tasks. Success is reported as a percentage; rule-agent and human columns give mean actions per session with standard deviation. The human action counts provide the demonstrated trajectory scale, whereas rule-agent counts show the cost of hand-crafted execution when a policy exists.

    Could not parse LaTeX table

    The authors conclude that expanding subgoal coverage and handling simulator corner cases could increase these scores, but the approximately 24.40% overall success after extensive manual engineering indicates that a planner-style solution that is effective for simpler navigation or ALFRED tasks is not a reasonable general solution for TEACh’s TATC setting. TATC scores are not directly comparable to EDH or TfD because TATC models both agents.

  10. Knowl 10 — TEACh task complexity and compositional scaling

    empirical result

    The 12 TEACh task types vary substantially in dialogue and action burden. The following values are means with standard deviations per session; “parameter variants” counts task-template parameterizations, “unique scenes” counts distinct scenes represented, “total sessions” is the number of successful sessions summarized for that task, and “all actions/session” includes dialogue and environment actions.

    Could not parse LaTeX table

    The progression from MAKE COFFEE to PREPARE BREAKFAST demonstrates that hierarchical and composite tasks require more dialogue and substantially longer action trajectories. The aggregate contains 438 parameter variants across 109 scenes, with averages of 13.67 utterances, 131.80 Follower actions, and 164.65 total actions per session.

  11. Knowl 11 — Limits of the initial modeling evidence

    limitation

    The initial experiments do not provide a learned end-to-end TATC model; TATC is assessed with manually engineered rule-based agents. The EDH and TfD baseline is adapted from ALFRED and therefore uses an input representation in which all dialogue is concatenated separately from environmental actions, without temporal alignment between the two modalities. Its weak TfD results and low EDH success rates show that the baseline does not solve long-horizon dialogue-grounded execution, especially in unseen rooms. The authors attribute the difficulty to TEACh’s multiple speakers, phatic and irrelevant utterances, dialogue anaphora, varied instruction granularity, object-state changes, and longer trajectories. They identify generalization to new task types, including few-shot task generalization, as necessary future capability rather than a demonstrated property of the presented models.

Coverage note — No substantial contributed material was deliberately omitted; appendix-only hyperparameters and implementation details were not expanded because they are supporting details rather than standalone headline contributions.

References

  1. 1.Abramson, J.; Ahuja, A.; Brussee, A.; Carnevale, F.; Cassin, M.; Clark, S.; Dudzik, A.; Georgiev, P.; Guy, A.; Harley, T.; Hill, F.; Hung, A.; Kenton, Z.; Landon, J.; Lillicrap, T.; Mathewson, K.; Muldal, A.; Santoro, A.; Savinov, N.; Varma, V.; Wayne, G.; Wong, N.; Yan, C.; and Zhu, R. 2020. Imitating Interactive Intelligence. arXiv.
  2. 2.Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sunderhauf, N.; Reid, I.; Gould, S.; and van den Hengel, A. 2018. Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  3. 3.Arumugam, D.; Karamcheti, S.; Gopalan, N.; Williams, E. C.; Rhee, M.; Wong, L. L.; and Tellex, S. 2018. Grounding Natural Language Instructions to Semantic Goal Representations for Abstraction and Generalization. Autonomous Robots.
  4. 4.Bisk, Y.; Holtzman, A.; Thomason, J.; Andreas, J.; Bengio, Y.; Chai, J.; Lapata, M.; Lazaridou, A.; May, J.; Nisnevich, A.; Pinto, N.; and Turian, J. 2020. Experience Grounds Language. In Empirical Methods in Natural Language Processing (EMNLP).
  5. 5.Bisk, Y.; Shih, K.; Choi, Y.; and Marcu, D. 2018. Learning Interpretable Spatial Operations in a Rich 3D Blocks World. In Proceedings of the Thirty Second AAAI Conference on Artificial Intelligence (AAAI), volume 32.
  6. 6.Blukis, V.; Misra, D.; Knepper, R. A.; and Artzi, Y. 2018. Mapping Navigation Instructions to Continuous Control Actions with Position Visitation Prediction. In Proceedings of the Conference on Robot Learning (CoRL).
  7. 7.Blukis, V.; Paxton, C.; Fox, D.; Garg, A.; and Artzi, Y. 2021. A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution. arXiv.
  8. 8.Chen, D.; and Mooney, R. 2011. Learning to Interpret Natural Language Navigation Instructions from Observations. In Proceedings of the Twenty Fifth AAAI Conference on Artificial Intelligence (AAAI), volume 25.
  9. 9.Chen, H.; Suhr, A.; Misra, D.; Snavely, N.; and Artzi, Y. 2019. Touchdown: Natural Language Navigation and Spatial Reasoning in Visual Street Environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 12538–12547.
  10. 10.Chi, T.-C.; Shen, M.; Eric, M.; Kim, S.; and Hakkani-Tur, D. 2020. Just Ask: An Interactive Learning Framework for Vision and Language Navigation. In Proceedings of the AAAI Conference on Artificial Intelligence.
  11. 11.Fried, D.; Hu, R.; Cirik, V.; Rohrbach, A.; Andreas, J.; Morency, L.-P.; Berg-Kirkpatrick, T.; Saenko, K.; Klein, D.; and Darrell, T. 2018. Speaker-Follower Models for Vision-and-Language Navigation. In Neural Information Processing Systems (NeurIPS).
  12. 12.Ghallab, M.; Howe, A.; Knoblock, C.; McDermott, D.; Ram, A.; Veloso, M.; Weld, D.; and Wilkins, D. 1998. PDDL The Planning Domain Definition Language. Yale Center for Computational Vision and Control.
  13. 13.Harnad, S. 1990. The Symbol Grounding Problem. Physica D: Nonlinear Phenomena, 42(1-3): 335–346.
  14. 14.He, K.; Gkioxari, G.; Dollar, P.; and Girshick, R. B. 2017. Mask R-CNN. International Conference on Computer Vision (ICCV).
  15. 15.Kim, B.; Bhambri, S.; Singh, K. P.; Mottaghi, R.; and Choi, J. 2021. Agent with the Big Picture: Perceiving Surroundings for Interactive Instruction Following. In Embodied AI Workshop CVPR.
  16. 16.Kim, H.; Zala, A.; Burri, G.; Tan, H.; and Bansal, M. 2020. ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments. In Findings of the Association for Computational Linguistics: EMNLP 2020.
  17. 17.Kollar, T.; Tellex, S.; Walter, M. R.; Huang, A.; Bachrach, A.; Hemachandra, S.; Brunskill, E.; Banerjee, A.; Roy, D.; Teller, S.; et al. 2013. Generalized Grounding Graphs: A Probabilistic Framework for Understanding Grounded Language. Journal of Artificial Intelligence Research.
  18. 18.Kolve, E.; Mottaghi, R.; Han, W.; VanderBilt, E.; Weihs, L.; Herrasti, A.; Gordon, D.; Zhu, Y.; Gupta, A.; and Farhadi, A. 2017. AI2-THOR: An Interactive 3D Environment for Visual AI. arXiv preprint arXiv:1712.05474.
  19. 19.MacMahon, M.; Stankiewicz, B.; and Kuipers, B. 2006. Walk the Talk: Connecting Language, Knowledge, and Action in Route Instructions. In Proceedings of the Twentieth AAAI Conference on Artifial Intelligence (AAAI), volume 20.
  20. 20.Matuszek, C.; Herbst, E.; Zettlemoyer, L.; and Fox, D. 2013. Learning to Parse Natural Language Commands to a Robot Control System. In Experimental Robotics, 403–415. Springer.
  21. 21.Mei, H.; Bansal, M.; and Walter, M. 2016. Listen, attend, and walk: Neural mapping of navigational instructions to action sequences. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (AAAI), volume 30.
  22. 22.Misra, D. K.; Bennett, A.; Blukis, V.; Niklasson, E.; Shatkhin, M.; and Artzi, Y. 2018. Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction. In Riloff, E.; Chiang, D.; Hockenmaier, J.; and Tsujii, J., eds., Empirical Methods in Natural (EMNLP).
  23. 23.Narayan-Chen, A.; Jayannavar, P.; and Hockenmaier, J. 2019. Collaborative Dialogue in Minecraft. In Association for Computational Linguistics (ACL).
  24. 24.Nguyen, K.; and Daume III, H. 2019. Help, Anna! Visual Navigation with Natural Multimodal Assistance via Retrospective Curiosity-Encouraging Imitation Learning. In Conference on Empirical Methods in Natural Language Processing and International Joint Conference on Natural Language Processing (EMNLP-IJCNLP).
  25. 25.Pashevich, A.; Schmid, C.; and Sun, C. 2021. Episodic Transformer for Vision-and-Language Navigation. arXiv preprint arXiv:2105.06453.
  26. 26.Puig, X.; Ra, K.; Boben, M.; Li, J.; Wang, T.; Fidler, S.; and Torralba, A. 2018. VirtualHome: Simulating Household Activities via Programs. In Computer Vision and Pattern Recognition (CVPR).
  27. 27.Roman, H. R.; Bisk, Y.; Thomason, J.; Celikyilmaz, A.; and Gao, J. 2020. RMM: A Recursive Mental Model for Dialog Navigation. In Findings of Empirical Methods in Natural Language Processing (EMNLP Findings).
  28. 28.Shah, R.; Wild, C.; Wang, S. H.; Alex, N.; Houghton, B.; Guss, W.; Mohanty, S.; Kanervisto, A.; Milani, S.; Topin, N.; et al. 2021. The MineRL BASALT Competition on Learning from Human Feedback. arXiv preprint arXiv:2107.01969.
  29. 29.Shridhar, M.; Thomason, J.; Gordon, D.; Bisk, Y.; Han, W.; Mottaghi, R.; Zettlemoyer, L.; and Fox, D. 2020. ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks. In Computer Vision and Pattern Recognition (CVPR).
  30. 30.Shrivastava, A.; Gopalakrishnan, K.; Liu, Y.; Piramuthu, R.; Tur, G.; Parikh, D.; and Hakkani-Tur, D. 2021. VIS-ITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator. arXiv preprint arXiv:2105.11589.
  31. 31.Suglia, A.; Gao, Q.; Thomason, J.; Thattai, G.; and Sukhatme, G. 2021. Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion. arXiv preprint arXiv:2108.04927.
  32. 32.Suhr, A.; Yan, C.; Schluger, J.; Yu, S.; Khader, H.; Mouallem, M.; Zhang, I.; and Artzi, Y. 2019. Executing Instructions in Situated Collaborative Interactions. In Empirical Methods in Natural Language Processing and International Joint Conference on Natural Language Processing (EMNLP-IJCNLP).
  33. 33.Tellex, S.; Knepper, R. A.; Li, A.; Roy, N.; and Rus, D. 2016. Asking for Help Using Inverse Semantics. In Robotics: Science and Systems Conference (RSS).
  34. 34.Thomason, J.; Gordon, D.; and Bisk, Y. 2018. Shifting the Baseline: Single Modality Performance on Visual Navigation & QA. arXiv preprint arXiv:1811.00613.
  35. 35.Thomason, J.; Murray, M.; Cakmak, M.; and Zettlemoyer, L. 2019. Vision-and-Dialog Navigation. In Conference on Robot Learning (CoRL).
  36. 36.Thomason, J.; Padmakumar, A.; Sinapov, J.; Walker, N.; Jiang, Y.; Yedidsion, H.; Hart, J.; Stone, P.; and Mooney, R. 2020. Jointly improving parsing and perception for natural language commands through human-robot dialog. Journal of Artificial Intelligence Research, 67: 327–374.
  37. 37.Zhang, Y.; and Chai, J. 2021. Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring. arXiv preprint arXiv:2106.03427.
  38. 38.Zhu, W.; Hu, H.; Chen, J.; Deng, Z.; Jain, V.; Ie, E.; and Sha, F. 2020. BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps. In Association for Computational Linguistics (ACL).

Citation

MLA
Padmakumar, A., et al. “TEACh: Task-driven Embodied Agents That Chat”. arXiv, 2021, http://arxiv.org/abs/2110.00534v3.
APA
Padmakumar, A., Thomason, J., Shrivastava, A., Lange, P., Narayan-Chen, A., Gella, S., Piramuthu, R., Tur, G., & Hakkani-Tur, D. (2021). TEACh: Task-driven Embodied Agents that Chat. arXiv. http://arxiv.org/abs/2110.00534v3
Chicago
Padmakumar, A., J. Thomason, A. Shrivastava, et al. 2021. “TEACh: Task-driven Embodied Agents That Chat”. arXiv. http://arxiv.org/abs/2110.00534v3.
Harvard
Padmakumar, A. et al. (2021) “TEACh: Task-driven Embodied Agents that Chat”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2110.00534v3.
Vancouver
1. Padmakumar A, Thomason J, Shrivastava A, Lange P, Narayan-Chen A, Gella S, Piramuthu R, Tur G, Hakkani-Tur D (2021) TEACh: Task-driven Embodied Agents that Chat. arXiv

BibTeX

@article{padmakumar2021teach,
  title = {TEACh: Task-driven Embodied Agents that Chat},
  author = {Padmakumar, Aishwarya and Thomason, Jesse and Shrivastava, Ayush and Lange, Patrick and Narayan-Chen, Anjali and Gella, Spandana and Piramuthu, Robinson and Tur, Gokhan and Hakkani-Tur, Dilek},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2110.00534v3},
  eprint = {2110.00534}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF