Core Challenges in Embodied Vision-Language Planning

Jonathan FrancisNariaki KitamuraFelix LabelleXiaopeng LuIngrid NavarroJean Oh

article2022JAIR64 citations

Presents a unifying taxonomy and comparative review of embodied vision-language planning methods, benchmarks, and simulation environments to clarify critical open challenges for deploying interactive agents in the physical world.

Listen

The field of embodied artificial intelligence is rapidly advancing toward building autonomous agents that can collaborate with humans in physical spaces. To operate effectively, these systems must combine visual perception, natural language understanding, and sequential decision-making to complete physical objectives such as household chores, search and rescue, and autonomous delivery. However, existing research has largely studied vision, language, and planning in isolation or through disparate, fragmented subtasks, leaving the overall landscape and real-world deployment challenges poorly understood.

The article establishes a unified taxonomy for Embodied Vision-Language Planning tasks and systematically surveys current algorithmic approaches, simulation platforms, datasets, and evaluation metrics. Through this comprehensive analysis, the article evaluates the state of the art and demonstrates the critical technical gaps that must be overcome to transition these artificial intelligence agents from virtual simulations to reliable real-world operations.

To conduct this evaluation, the authors categorize the field into five primary task families: vision-language navigation, embodied question answering, embodied object referral, vision and dialogue navigation, and embodied goal-directed manipulation. The article reviews widely used modeling paradigms—including supervised imitation learning, reinforcement learning, and multimodal transformer architectures—across standard simulation environments such as Matterport3D, AI Habitat, and AI2-THOR. It also analyzes evaluation metrics measuring success rates, path length, trajectory fidelity, and object interaction accuracy.

The review yields five key findings regarding current capabilities. First, while agents achieve high success rates in training environments, their performance drops substantially in unseen test environments, exposing severe overfitting and poor generalization. Second, recent diagnostic studies show that masking visual or textual inputs results in negligible performance drops, indicating that models often rely on dataset biases and statistical shortcuts rather than genuine cross-modal understanding. Third, hybrid training that combines supervised pre-training with reinforcement learning or reward shaping consistently outperforms single-paradigm methods by enabling error recovery. Fourth, standard success metrics are often insufficient; for example, simple destination-based metrics fail to reflect whether an agent followed safe or instruction-faithful paths, making path-similarity metrics like dynamic time warping more effective for evaluation. Fifth, most current implementations rely on static, turn-based dialogue and stationary worlds, completely ignoring the dynamic changes and interactive communication required in physical deployments.

These findings imply that high benchmark scores on public leaderboards do not translate to operational readiness. In safety-critical or cost-sensitive deployments, relying on current models introduces significant operational risks because agents struggle to adapt to unmapped obstacles, unexpected physical interventions, or ambiguous instructions. Furthermore, the disconnect between simulated high-level actions (such as teleporting between viewpoints) and low-level physical control poses severe sim-to-real transfer bottlenecks that delay commercial adoption.

To address these limitations, the article recommends prioritizing the development of dynamic simulation environments where external events occur independently of the agent's actions. Stakeholders and researchers should adopt standardized, cross-task evaluation suites to benchmark general skills—such as spatial reasoning and object grounding—rather than isolated single-task metrics. Additionally, future efforts must integrate structured commonsense knowledge bases and dynamic, multi-turn dialogue systems to ensure agents can clarify ambiguous instructions and reason about real-world physical constraints before deploying them into human environments.

The conclusions of the article are constrained by the current scope of the literature, which predominantly focuses on single-agent operations in simulated, indoor settings while excluding complex multi-agent dynamics and low-level physical robot hardware control. Nevertheless, the findings provide a highly credible, evidence-based assessment that serves as a vital strategic roadmap for transitioning embodied vision-language planning into robust, real-world robotic systems.

Cover for Core Challenges in Embodied Vision-Language Planning

Abstract

Recent advances in the areas of multimodal machine learning and artificial intelligence (AI) have led to the development of challenging tasks at the intersection of Computer Vision, Natural Language Processing, and Embodied AI. Whereas many approaches and previous survey pursuits have characterised one or two of these dimensions, there has not been a holistic analysis at the center of all three. Moreover, even when combinations of these topics are considered, more focus is placed on describing, e.g., current architectural methods, as opposed to also illustrating high-level challenges and opportunities for the field. In this survey paper, we discuss Embodied Vision-Language Planning (EVLP) tasks, a family of prominent embodied navigation and manipulation problems that jointly use computer vision and natural language. We propose a taxonomy to unify these tasks and provide an in-depth analysis and comparison of the new and current algorithmic approaches, metrics, simulated environments, as well as the datasets used for EVLP tasks. Finally, we present the core challenges that we believe new EVLP works should seek to address, and we advocate for task construction that enables model generalizability and furthers real-world deployment.

Table of Contents

  • 1. Introduction
  • 1.1 Scope of this Survey
  • 1.2 Intended Audience and Reading Guide
  • 1.3 Related Surveys
  • 2. Problem Definition
  • 2.1 Taxonomy
  • 2.2 Tasks in Embodied Vision-Language Planning
  • 2.2.1 VISION LANGUAGE NAVIGATION
  • 2.2.2 EMBODIED QUESTION ANSWERING
  • 2.2.3 EMBODIED OBJECT REFERRAL
  • 2.2.4 VISION AND DIALOGUE NAVIGATION
  • 2.2.5 EMBODIED GOAL-DIRECTED MANIPULATION
  • 3. Approaches
  • 3.1 Modeling Vision, Language, and Planning
  • 3.1.1 MODELING VISION
  • 3.1.2 MODELING LANGUAGE
  • 3.1.3 MODELING MULTIMODALITY
  • 3.1.4 MODELING ACTION-GENERATION AND PLANNING
  • 3.2 Learning Paradigms
  • 3.2.1 SUPERVISED LEARNING
  • 3.2.2 REINFORCEMENT LEARNING
  • 3.2.3 JOINT REINFORCEMENT AND SUPERVISED LEARNING
  • 3.3 Common Techniques
  • 3.3.1 DATA AUGMENTATION
  • 3.3.2 ADDITIONAL OBJECTIVES
  • 3.3.3 PRE-TRAINING
  • 3.3.4 MULTITASK LEARNING
  • 3.3.5 LEARNING AND OPTIMIZATION
  • 3.3.6 REWARD SHAPING
  • 4. Evaluation
  • 4.1 Metrics
  • 4.1.1 SUCCESS
  • 4.1.2 DISTANCE
  • 4.1.3 PATH-PATH SIMILARITY
  • 4.1.4 INSTRUCTION-BASED
  • 4.1.5 OBJECT REFERRAL
  • 4.2 Simulation Environments and Datasets
  • 4.2.1 SIMULATORS
  • 4.2.2 DATASETS
  • 5. Open Challenges in Embodied Vision-Language Planning
  • 5.1 New Directions in EVLP Research
  • 5.1.1 SOCIAL INTERACTION
  • 5.1.2 DYNAMIC ENVIRONMENTS
  • 5.1.3 CROSS-TASK ROBOT LEARNING
  • 5.2 Use of Domain Knowledge
  • 5.2.1 PRE-TRAINING
  • 5.2.2 COMMONSENSE KNOWLEDGE
  • 5.3 Agent Training Objectives
  • 5.4 Model Evaluation Paradigms
  • 5.4.1 SIMULATOR REALNESS
  • 5.4.2 DATASET REALNESS
  • 5.4.3 TESTS FOR GENERALISABILITY
  • 6. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — A three-part taxonomy distinguishes five embodied vision-language planning tasks

    definition

    The survey organizes embodied vision-language planning (EVLP) research along three branches: tasks, approaches, and evaluation. The task branch separates five benchmarked problem families by their objectives and interaction requirements:

    • Vision-language navigation (VLN): follow a language instruction through an environment to a goal location. The agent navigates but does not ordinarily identify a target object or change the environment.
    • Embodied question answering (EQA): navigate to gather visual evidence needed to answer a language question. An answer action, rather than reaching a uniquely specified destination, ends the episode; different routes or viewpoints may suffice.
    • Embodied object referral (EOR): use an instruction to navigate to and identify a referred object, typically by selecting its bounding box or mask at the final viewpoint.
    • Vision and dialogue navigation (VDN): navigate using sequential directives and/or interaction with an oracle or human collaborator to resolve uncertainty. Its action space can include requests for help as well as navigation.
    • Embodied goal-directed manipulation (EGM): carry out language-specified object interactions, often alongside navigation and multi-step planning. The agent must track object states and satisfy interaction constraints, such as whether an object can be picked up.

    Across these families, the tasks differ in whether object identification is required, whether the environment can be changed or another agent queried, and whether the main reasoning demand is instruction following, information gathering, spatial-semantic understanding, or manipulation.

  2. Knowl 2 — EVLP formalizes planning from partial visual and language histories

    definition

    A planning problem is represented by a state set SS, an action set AA, an initial state sini∈Ss_{\mathrm{ini}}\in S, and a goal state sgoal∈Ss_{\mathrm{goal}}\in S. A solution is a finite sequence of states and actions from the initial state to the goal, with actions admissible under the environment’s transition dynamics and the agent’s system constraints. Stateless tasks such as question answering fit the framework as single-step decision problems.

    In an EVLP problem, the agent may not know the full state space in advance. At time tt, its state estimate is based on the visual inputs v1,…,vtv_1,\ldots,v_t and language inputs l1,…,ltl_1,\ldots,l_t, where each viv_i belongs to the available vision-input set VV and each lil_i belongs to the language-input set LL:

    st=f(v1,…,vt,l1,…,lt).s_t=f(v_1,\ldots,v_t,l_1,\ldots,l_t).

    The agent predicts a solution sequence and aims to minimize its discrepancy from an admissible solution that satisfies the task. The framework allows inputs to arrive either with the initial task specification or during execution; visual observations, in particular, are often acquired online as the agent moves.

  3. Knowl 3 — EVLP systems combine visual-language representations with structured planning

    model/method

    The survey groups modeling choices into perception, language encoding, multimodal integration, and action generation. Vision is commonly encoded with pretrained convolutional networks; object-detector features can provide more explicit region-level representations than whole-image features. Language is commonly encoded with recurrent networks, particularly gated LSTMs or GRUs, or with transformers. Attention and multimodal transformers combine visual, linguistic, and agent-state information, supporting both modality fusion and grounding language elements in the scene.

    Planning methods reviewed in the survey include:

    • Mapping and exploration: build a more abstract representation, such as a metric map or topological graph, and explore unknown regions. Frontier-based exploration prioritizes unmapped areas; mapping can also support localization, reduce revisits, and enable global planning. Backtracking methods use progress or past observations to recover from confusion.
    • Search over candidate states or actions: greedy sequence prediction is common, while beam search retains multiple candidate sequences. Graph-based planners can update their map as new locations are observed.
    • Hierarchical planning: a higher-level planner predicts subgoals or waypoints and a lower-level controller produces actions to reach them. For tasks combining navigation and manipulation, task-and-motion planning separates symbolic task choices from geometric motion or manipulation planning.

    The survey treats these as complementary responses to different observability, horizon, and action-space demands, rather than as a single best architecture.

  4. Knowl 4 — EVLP training commonly combines demonstrations, interaction, and auxiliary signals

    model/method

    The survey distinguishes three learning paradigms. Supervised or imitation learning trains an agent to reproduce expert demonstrations, often shortest paths; approaches include behavior cloning, DAgger-style aggregation of expert-labeled trajectories, and training strategies that expose the learner to its own sampled states. Demonstrations can accelerate learning, but imitation can suffer from distribution shift when the agent reaches states absent from expert traces. Reinforcement learning (RL) trains through environment interaction and rewards; policy-gradient methods, including REINFORCE and actor-critic methods, are common, but reward specification, sample complexity, and slow convergence are recurring difficulties. Hybrid training uses demonstrations to initialize behavior and RL to improve it through interaction, which can help agents recover from errors beyond expert trajectories.

    Other recurring training interventions include back-translation to create synthetic instructions for sampled paths, environmental dropout to vary the visual context, pretraining and multitask learning to transfer representations, auxiliary objectives such as progress estimation or instruction-path matching, and reward shaping based on task progress or instruction-trajectory alignment. The survey does not identify a universally best combination of objectives: the appropriate choice depends on the available supervision and reward signals, and the field needs clearer evidence about which combinations transfer and generalize.

    The authors also identify external domain knowledge as an underdeveloped opportunity for EVLP. Pretraining can provide useful initialization, but poorly chosen pretraining tasks—including tasks with counterfactual or causally confusing examples—may hinder downstream learning. Structured commonsense knowledge could support reasoning about objects and environments, but its representation must fit the downstream task; its use in EVLP remains largely unexplored.

  5. Knowl 5 — EVLP evaluation separates success, efficiency, path fidelity, and object selection

    equation

    The survey groups metrics by what they measure: task completion, distance and travel efficiency, similarity to a reference path, language-path alignment, and object selection. Let P=(p1,…,pm)P=(p_1,\ldots,p_m) be a predicted path, R=(r1,…,rn)R=(r_1,\ldots,r_n) a reference path, d(x,y)d(x,y) the relevant distance between states, and dthd_{\mathrm{th}} a task-specific success threshold. The main navigation measures include:

    NE(P,R)=d(pm,rn),ONE(P,R)=min⁡p∈Pd(p,rn),PL(P)=∑i=1m−1d(pi,pi+1).\mathrm{NE}(P,R)=d(p_m,r_n),\qquad \mathrm{ONE}(P,R)=\min_{p\in P}d(p,r_n),\qquad \mathrm{PL}(P)=\sum_{i=1}^{m-1}d(p_i,p_{i+1}).

    Navigation error (NE) measures the final distance to the goal; oracle navigation error (ONE) measures the closest the trajectory came to it; path length (PL) measures travel. Success rate (SR) is the fraction of episodes ending within the threshold, while oracle success rate uses the nearest point on the trajectory. Success weighted by path length (SPL) combines success with the ratio of shortest start-to-goal distance to the longer of the predicted path length and that shortest distance.

    For instruction-following tasks, reaching the goal alone does not show whether the agent followed the described route. Coverage weighted by Length Score (CLS) combines path coverage and a penalty for paths shorter or longer than the reference. Normalized dynamic time warping (nDTW) measures path similarity while allowing temporal alignment; success weighted by nDTW (SDTW) additionally requires thresholded goal success. For object referral or manipulation, intersection over union (IoU) measures overlap between predicted and reference bounding boxes or masks. EQA commonly uses answer accuracy. The survey cautions that success thresholds and discretized action spaces can affect reported navigation success, and that endpoint-only metrics omit route fidelity and intermediate behavior.

  6. Knowl 6 — EVLP benchmarks span discrete indoor graphs, continuous navigation, and interactive manipulation

    experimental setup

    EVLP evaluation typically pairs a simulator, which supplies interactive observations and transitions, with a dataset, which specifies tasks and examples of desired behavior. The survey compares environments along dimensions including indoor versus outdoor scenes, synthetic versus photorealistic rendering, discrete versus continuous actions, supported modalities, and whether agents can change object states.

    Representative navigation benchmarks illustrate the range. R2R uses an indoor Matterport3D navigation graph and reports 7,189 trajectories and 21,567 instructions, with an average of 5 steps. RxR provides denser, multilingual instruction grounding; its dataset summary reports 16,522 trajectories and 126,000 instructions in English, Hindi, and Telugu. R2R variants include longer concatenated paths and finer-grained instruction alignment. VLN-CE moves navigation beyond a fixed graph in Habitat, while RoboVLN adds low-level velocity actions between higher-level actions. Outdoor benchmarks include StreetLearn-based tasks using urban panoramas.

    The other task families use distinct settings and annotations: EQA ranges from synthetic environments such as House3D and AI2-THOR to photorealistic Matterport3D-based environments; EOR benchmarks include indoor REVERIE with object annotations and outdoor Touchdown; and EGM includes interactive household or synthetic environments such as ALFRED, which provides demonstrations and multiple manipulation actions. These choices determine which visual inputs, action types, language forms, and interaction demands a benchmark can test.

  7. Knowl 7 — Real-world readiness requires more than success on current simulators and datasets

    limitation

    The survey argues that current evaluation can reward behavior that is inadequate for deployment. Simulators often simplify physical dynamics, use discretized navigation or abstract actions such as pick and place, and cover a limited range of scenes and events. Agents can therefore overfit to simulator transitions that do not match real execution. Task metrics can reinforce this problem when they record only end-goal success and ignore intermediate behavior, collisions, or inefficient movement.

    Dataset-based training has complementary weaknesses: demonstrations provide limited coverage of transitions and rare events; generated or templated instructions may not resemble real language; and datasets often omit sensory modalities available in the world. The authors recommend broader physical and semantic coverage, attention to simulation-to-real transfer and out-of-distribution behavior, metrics for intermediate actions and efficiency, and consistent reporting of dataset properties such as action spaces, instruction lengths, vocabulary sizes, collection procedures, and availability. They also advocate metrics that can be applied across environments so that performance reflects transferable skills rather than benchmark-specific correlations.

  8. Knowl 8 — Interactive EVLP needs social partners, not only prerecorded dialogue

    limitation

    Existing dialogue and assistance settings often provide pre-generated, static exchanges, limiting what an agent can learn about online collaboration. The survey identifies dynamic social interaction as an open requirement for real-world assistive and collaborative agents. Such settings should let agents maintain representations of other participants’ knowledge or conceptual schemas, use direct feedback about task progress, and generate questions or statements to request help, resolve ambiguous instructions, suggest alternatives, or report progress.

    The authors argue that tasks and simulators must support representations of other agents’ anticipated actions, mental states, and prior behavior. Without these capabilities, benchmarked dialogue does not fully test the shared understanding of space and task structure required for situated collaboration.

  9. Knowl 9 — EVLP benchmarks rarely test environments that change independently of the agent

    limitation

    Most contemporary EVLP tasks assume a stationary environment. The survey defines a dynamic environment as one that changes through events either within or outside the agent’s control. Existing robustness benchmarks may perturb perception or transition dynamics, but those perturbations can remain fixed during task execution; interactive benchmarks may allow changes caused by the agent without modeling independent changes in the world. Real settings also include other agents acting, physical events such as objects moving, and interactions whose availability depends on time. The authors call for tasks that include such changes, because planning under environmental uncertainty and non-stationarity remains comparatively unexplored.

  10. Knowl 10 — Generalization should be tested across domains, tasks, and longer horizons

    limitation

    The survey argues that evaluating task families separately fails to measure capabilities shared among them. It recommends benchmarks for transfer across tasks—for example, between instruction following, question answering, and manipulation—as well as explicit tests of transfer across environments. Existing seen/unseen splits do not systematically isolate changes in object distributions, visual backgrounds, or indoor layouts, so success on an unseen split does not by itself establish robust planning. Evaluation should also test longer and more complex instructions: path-concatenation datasets create longer routes, while path-decomposition datasets align finer-grained instructions with route segments.

    The surveyed diagnostic evidence raises doubts about whether current navigation models reliably ground language in vision: masking objects mentioned in instructions produced only a limited performance decrease, and the effect of masking object or directional information in text varied by dataset. The authors therefore advocate tests that distinguish genuine grounding and planning from reliance on spurious correlations, and that measure transfer to related tasks and longer horizons.

Coverage note — The survey’s exhaustive catalogue of individual model variants, simulators, datasets, and secondary language-generation metrics is deliberately compressed; those entries support the comparison but are less load-bearing than the taxonomy, evaluation framework, and challenges captured here.

References

  1. 1.Agrawal, A., Batra, D., Parikh, D., and Kembhavi, A. (2018). Don’t just assume; look and answer: Overcoming priors for visual question answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 4971–4980. IEEE Computer Society.
  2. 2.Akhtar, M. S., Chauhan, D., Ghosal, D., Poria, S., Ekbal, A., and Bhattacharyya, P. (2019). Multi-task learning for multi-modal emotion recognition and sentiment analysis. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 370–379, Minneapolis, Minnesota. Association for Computational Linguistics.
  3. 3.Alami, R., Chatila, R., Fleury, S., Ghallab, M., and Ingrand, F. (1998). An architecture for autonomy. The International Journal of Robotics Research, 17(4):315–337.
  4. 4.Anderson, P., Chang, A. X., Chaplot, D. S., Dosovitskiy, A., Gupta, S., Koltun, V., Kosecka, J., Malik, J., Mottaghi, R., Savva, M., and Zamir, A. R. (2018a). On evaluation of embodied navigation agents. CoRR, abs/1807.06757.
  5. 5.Anderson, P., Fernando, B., Johnson, M., and Gould, S. (2016). Spice: Semantic propositional image caption evaluation.
  6. 6.Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (2018b). Bottom-up and top-down attention for image captioning and visual question answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 6077–6086. IEEE Computer Society.
  7. 7.Anderson, P., Shrivastava, A., Parikh, D., Batra, D., and Lee, S. (2019). Chasing ghosts: Instruction following as bayesian state tracking. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R., editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 369–379.
  8. 8.Anderson, P., Shrivastava, A., Truong, J., Majumdar, A., Parikh, D., Batra, D., and Lee, S. (2020). Sim-to-real transfer for vision-and-language navigation. CoRR, abs/2011.03807.
  9. 9.Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sunderhauf, N., Reid, I. D., Gould, S., and van den Hengel, A. (2018c). Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 3674–3683. IEEE Computer Society.
  10. 10.Andreas, J., Rohrbach, M., Darrell, T., and Klein, D. (2016). Neural module networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 39–48. IEEE Computer Society.
  11. 11.Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. (2015). VQA: visual question answering. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pages 2425–2433. IEEE Computer Society.
  12. 12.Bahdanau, D., Cho, K., and Bengio, Y. (2015). Neural machine translation by jointly learning to align and translate. In Bengio, Y. and LeCun, Y., editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  13. 13.Baltrusaitis, T., Ahuja, C., and Morency, L. (2019). Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2):423–443.
  14. 14.Batra, D., Chang, A. X., Chernova, S., Davison, A. J., Deng, J., Koltun, V., Levine, S., Malik, J., Mordatch, I., Mottaghi, R., et al. (2020). Rearrangement: A challenge for embodied ai. arXiv preprint arXiv:2011.01975.
  15. 15.Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Kuttler, H., Lefrancq, A., Green, S., Valdes, V., Sadik, A., Schrittwieser, J., Anderson, K., York, S., Cant, M., Cain, A., Bolton, A., Gaffney, S., King, H., Hassabis, D., Legg, S., and Petersen, S. (2016). Deepmind lab. CoRR, abs/1612.03801.
  16. 16.Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2015). The arcade learning environment: An evaluation platform for general agents (extended abstract). In Yang, Q. and Wooldridge, M. J., editors, Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence, IJCAI 2015, Buenos Aires, Argentina, July 25-31, 2015, pages 4148–4152. AAAI Press.
  17. 17.Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009). Curriculum learning. In Danyluk, A. P., Bottou, L., and Littman, M. L., editors, Proceedings of the 26th Annual International Conference on Machine Learning, ICML 2009, Montreal, Quebec, Canada, June 14-18, 2009, volume 382 of ACM International Conference Proceeding Series, pages 41–48. ACM.
  18. 18.Bisk, Y., Holtzman, A., Thomason, J., Andreas, J., Bengio, Y., Chai, J., Lapata, M., Lazaridou, A., May, J., Nisnevich, A., Pinto, N., and Turian, J. (2020). Experience grounds language. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8718–8735, Online. Association for Computational Linguistics.
  19. 19.Black, P. E. et al. (1998). Dictionary of algorithms and data structures.
  20. 20.Blukis, V., Brukhim, N., Bennett, A., Knepper, R. A., and Artzi, Y. (2018a). Following high-level navigation instructions on a simulated quadcopter with imitation learning. In Kress-Gazit, H., Srinivasa, S. S., Howard, T., and Atanasov, N., editors, Robotics: Science and Systems XIV, Carnegie Mellon University, Pittsburgh, Pennsylvania, USA, June 26-30, 2018.
  21. 21.Blukis, V., Misra, D., Knepper, R. A., and Artzi, Y. (2018b). Mapping navigation instructions to continuous control actions with position-visitation prediction. In Conference on Robot Learning, pages 505–518. PMLR.
  22. 22.Blukis, V., Terme, Y., Niklasson, E., Knepper, R. A., and Artzi, Y. (2019). Learning to map natural language instructions to physical quadcopter control using simulated flight. In Conference on Robot Learning (CoRL).
  23. 23.Boularias, A., Duvallet, F., Oh, J., and Stentz, A. (2015). Grounding spatial relations for outdoor robot navigation. In 2015 IEEE International Conference on Robotics and Automation (ICRA), pages 1976–1982. IEEE.
  24. 24.Brodeur, S., Perez, E., Anand, A., Golemo, F., Celotti, L., Strub, F., Rouat, J., Larochelle, H., and Courville, A. C. (2018). Home: a household multimodal environment. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Workshop Track Proceedings. OpenReview.net.
  25. 25.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020). Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H., editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  26. 26.Burgard, W., Moors, M., Stachniss, C., and Schneider, F. E. (2005). Coordinated multi-robot exploration. IEEE Transactions on robotics, 21(3):376–386.
  27. 27.Cambon, S., Alami, R., and Gravot, F. (2009). A hybrid approach to intricate motion, manipulation and task planning. The International Journal of Robotics Research, 28(1):104–126.
  28. 28.Castelfranchi, C. (1998). Modelling social action for ai agents. Artificial intelligence, 103(1-2):157–182.
  29. 29.Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., and Zhang, Y. (2017). Matterport3d: Learning from RGB-D data in indoor environments. International Conference on 3D Vision (3DV).
  30. 30.Charalampous, K., Kostavelis, I., and Gasteratos, A. (2017). Recent trends in social aware robot navigation: A survey. Robotics and Autonomous Systems, 93:85–104.
  31. 31.Chattopadhyay, P., Hoffman, J., Mottaghi, R., and Kembhavi, A. (2021). Robustnav: Towards benchmarking robustness in embodied navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15691–15700.
  32. 32.Chen, B., Francis, J., Oh, J., Nyberg, E., and Herbert, S. L. (2021). Safe autonomous racing via approximate reachability on ego-vision.
  33. 33.Chen, B., Francis, J., Pritoni, M., Kar, S., and Berges, M. (2020a). Cohort: Coordination of heterogeneous thermostatically controlled loads for demand flexibility. In Proceedings of the 7th ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation.
  34. 34.Chen, C., Jain, U., Schissler, C., Gari, S. V. A., Al-Halah, Z., Ithapu, V. K., Robinson, P., and Grauman, K. (2020b). Soundspaces: Audio-visual navigation in 3d environments. In Vedaldi, A., Bischof, H., Brox, T., and Frahm, J., editors, Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part VI, volume 12351 of Lecture Notes in Computer Science, pages 17–36. Springer.
  35. 35.Chen, C., Jain, U., Schissler, C., Gari, S. V. A., Al-Halah, Z., Ithapu, V. K., Robinson, P., and Grauman, K. (2020c). Soundspaces: Audio-visual navigation in 3d environments. In ECCV.
  36. 36.Chen, C., Majumder, S., Al-Halah, Z., Gao, R., Ramakrishnan, S. K., and Grauman, K. (2020d). Learning to set waypoints for audio-visual navigation. arXiv preprint arXiv:2008.09622.
  37. 37.Chen, D., Zhou, B., Koltun, V., and Krahenbühl, P. (2019a). Learning by cheating.
  38. 38.Chen, H., Suhr, A., Misra, D., Snavely, N., and Artzi, Y. (2019b). TOUCHDOWN: natural language navigation and spatial reasoning in visual street environments. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 12538–12547. Computer Vision Foundation / IEEE.
  39. 39.Chi, T., Shen, M., Eric, M., Kim, S., and Hakkani-Tur, D. (2020). Just ask: An interactive learning framework for vision and language navigation. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 2459–2466. AAAI Press.
  40. 40.Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning phrase representations using RNN encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1724–1734, Doha, Qatar. Association for Computational Linguistics.
  41. 41.Choi, J. and Amir, E. (2009). Combining planning and motion planning. In 2009 IEEE International Conference on Robotics and Automation, pages 238–244. IEEE.
  42. 42.Cummins, M. and Newman, P. (2008). Fab-map: Probabilistic localization and mapping in the space of appearance. The International Journal of Robotics Research, 27(6):647–665.
  43. 43.Das, A., Carnevale, F., Merzic, H., Rimell, L., Schneider, R., Abramson, J., Hung, A., Ahuja, A., Clark, S., Wayne, G., and Hill, F. (2020). Probing emergent semantics in predictive agents via question answering. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 2376–2391. PMLR.
  44. 44.Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., and Batra, D. (2018a). Embodied question answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 1–10. IEEE Computer Society.
  45. 45.Das, A., Gkioxari, G., Lee, S., Parikh, D., and Batra, D. (2018b). Neural modular control for embodied question answering. arXiv preprint arXiv:1810.11181.
  46. 46.Dasari, S. and Gupta, A. (2020). Transformers for one-shot visual imitation.
  47. 47.de Haan, P., Jayaraman, D., and Levine, S. (2019). Causal confusion in imitation learning. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R., editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 11693–11704.
  48. 48.de Vries, H., Shuster, K., Batra, D., Parikh, D., Weston, J., and Kiela, D. (2018). Talk the walk: Navigating new york city through grounded dialogue.
  49. 49.Deitke, M., Han, W., Herrasti, A., Kembhavi, A., Kolve, E., Mottaghi, R., Salvador, J., Schwenk, D., VanderBilt, E., Wallingford, M., Weihs, L., Yatskar, M., and Farhadi, A. (2020). Robothor: An open simulation-to-real embodied AI platform. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 3161–3171. IEEE.
  50. 50.Deng, Z., Narasimhan, K., and Russakovsky, O. (2020). Evolving graphical planner: Contextual global planning for vision-and-language navigation. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H., editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  51. 51.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  52. 52.Duvallet, F. (2015). Natural language direction following for robots in unstructured unknown environments. Technical report, CARNEGIE-MELLON UNIV PITTSBURGH PA ROBOTICS INST.
  53. 53.Duvallet, F., Walter, M. R., Howard, T., Hemachandra, S., Oh, J., Teller, S., Roy, N., and Stentz, A. (2016). Inferring maps and behaviors from natural language instructions. In Experimental Robotics, pages 373–388. Springer.
  54. 54.Ehsani, K., Han, W., Herrasti, A., VanderBilt, E., Weihs, L., Kolve, E., Kembhavi, A., and Mottaghi, R. (2021). Manipulathor: A framework for visual object manipulation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 4497–4506. Computer Vision Foundation / IEEE.
  55. 55.Everingham, M., Gool, L. V., Williams, C. K. I., Winn, J. M., and Zisserman, A. (2010). The pascal visual object classes (VOC) challenge. Int. J. Comput. Vis., 88(2):303–338.
  56. 56.Fang, K., Toshev, A., Li, F., and Savarese, S. (2019). Scene memory transformer for embodied agents in long-horizon tasks. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 538–547. Computer Vision Foundation / IEEE.
  57. 57.Filliat, D. and Meyer, J.-A. (2003). Map-based navigation in mobile robots:: I. a review of localization strategies. Cognitive systems research, 4(4):243–282.
  58. 58.Francis, J., Chen, B., Ganju, S., Kathpal, S., Poonganam, J., Shivani, A., Vyas, V., Genc, S., Zhukov, I., Kumskoy, M., Oh, J., Nyberg, E., and Herbert, S. L. (2022). Learn-to-race challenge 2022: Benchmarking safe learning and cross-domain generalisation in autonomous racing.
  59. 59.Francis, J., Quintana, M., von Frankenberg, N., Munir, S., and Berges, M. (2019). Occutherm: Occupant thermal comfort inference using body shape information. In Proceedings of the 6th International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation, BuildSys ’19, New York, NY, USA. ACM.
  60. 60.Fried, D., Andreas, J., and Klein, D. (2018a). Unified pragmatic models for generating and following instructions. In Proceedings of NAACL-HLT, pages 1951–1963.
  61. 61.Fried, D., Chiu, J. T., and Klein, D. (2021). Reference-centric models for grounded collaborative dialogue. arXiv preprint arXiv:2109.05042.
  62. 62.Fried, D., Hu, R., Cirik, V., Rohrbach, A., Andreas, J., Morency, L., Berg-Kirkpatrick, T., Saenko, K., Klein, D., and Darrell, T. (2018b). Speaker-follower models for vision-and-language navigation. In Bengio, S., Wallach, H. M., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R., editors, Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montreal, Canada, pages 3318–3329.
  63. 63.Fu, T.-J., Wang, X. E., Peterson, M. F., Grafton, S. T., Eckstein, M. P., and Wang, W. Y. (2020). Counterfactual vision-and-language navigation via adversarial path sampler. In European Conference on Computer Vision, pages 71–86. Springer.
  64. 64.Gan, C., Schwartz, J., Alter, S., Schrimpf, M., Traer, J., Freitas, J. D., Kubilius, J., Bhandwaldar, A., Haber, N., Sano, M., Kim, K., Wang, E., Mrowca, D., Lingelbach, M., Curtis, A., Feigelis, K. T., Bear, D. M., Gutfreund, D., Cox, D. D., DiCarlo, J. J., McDermott, J. H., Tenenbaum, J. B., and Yamins, D. L. K. (2020). Threedworld: A platform for interactive multi-modal physical simulation. CoRR, abs/2007.04954.
  65. 65.Gan, C., Zhao, H., Chen, P., Cox, D. D., and Torralba, A. (2019). Self-supervised moving vehicle tracking with stereo sound. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 7052–7061. IEEE.
  66. 66.Garrett, C. R., Chitnis, R., Holladay, R., Kim, B., Silver, T., Kaelbling, L. P., and Lozano-Perez, T. (2021). Integrated task and motion planning. Annual review of control, robotics, and autonomous systems, 4:265–293.
  67. 67.Gordon, D., Kembhavi, A., Rastegari, M., Redmon, J., Fox, D., and Farhadi, A. (2018). IQA: visual question answering in interactive environments. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 4089–4098. IEEE Computer Society.
  68. 68.Hao, W., Li, C., Li, X., Carin, L., and Gao, J. (2020). Towards learning a generic agent for vision-and-language navigation via pre-training. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 13134–13143. IEEE.
  69. 69.Hart, P. E., Nilsson, N. J., and Raphael, B. (1968). A formal basis for the heuristic determination of minimum cost paths. IEEE transactions on Systems Science and Cybernetics, 4(2):100–107.
  70. 70.He, K., Gkioxari, G., Dollar, P., and Girshick, R. B. (2017). Mask R-CNN. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 2980–2988. IEEE Computer Society.
  71. 71.He, K., Zhang, X., Ren, S., and Sun, J. (2016). Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 770–778. IEEE Computer Society.
  72. 72.He, Z., Julian, R., Heiden, E., Zhang, H., Schaal, S., Lim, J. J., Sukhatme, G., and Hausman, K. (2018). Zero-shot skill composition and simulation-to-real transfer by learning task representations. arXiv preprint arXiv:1810.02422.
  73. 73.Herman, J., Francis, J., Ganju, S., Chen, B., Koul, A., Gupta, A., Skabelkin, A., Zhukov, I., Kumskoy, M., and Nyberg, E. (2021). Learn-to-race: A multimodal control environment for autonomous racing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9793–9802.
  74. 74.Hermann, K. M., Malinowski, M., Mirowski, P., Banki-Horvath, A., Anderson, K., and Hadsell, R. (2020). Learning to follow directions in street view. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 11773–11781. AAAI Press.
  75. 75.Hester, T., Vecerík, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., Dulac-Arnold, G., Agapiou, J. P., Leibo, J. Z., and Gruslys, A. (2018). Deep q-learning from demonstrations. In McIlraith, S. A. and Weinberger, K. Q., editors, Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2-7, 2018, pages 3223–3230. AAAI Press.
  76. 76.Hinton, G., Srivastava, N., and Swersky, K. (2012). Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.
  77. 77.Hochreiter, S. and Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8):1735–1780.
  78. 78.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. (2019). The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751.
  79. 79.Hong, Y., Rodriguez, C., Wu, Q., and Gould, S. (2020). Sub-instruction aware vision-and-language navigation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3360–3376, Online. Association for Computational Linguistics.
  80. 80.Hu, Y., Wang, W., Jia, H., Wang, Y., Chen, Y., Hao, J., Wu, F., and Fan, C. (2020). Learning to utilize shaping rewards: A new approach of reward shaping. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H., editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  81. 81.Huang, H., Jain, V., Mehta, H., Baldridge, J., and Ie, E. (2019). Multi-modal discriminative model for vision-and-language navigation. In Proceedings of the Combined Workshop on Spatial Language Understanding (SpLU) and Grounded Communication for Robotics (RoboNLP), pages 40–49, Minneapolis, Minnesota. Association for Computational Linguistics.
  82. 82.Ilharco, G., Jain, V., Ku, A., Ie, E., and Baldridge, J. (2019). General evaluation for instruction conditioned navigation using dynamic time warping. CoRR, abs/1907.05446.
  83. 83.Irshad, M. Z., Ma, C.-Y., and Kira, Z. (2021). Hierarchical cross-modal agent for robotics vision-and-language navigation. arXiv preprint arXiv:2104.10674.
  84. 84.Jain, V., Magalhaes, G., Ku, A., Vaswani, A., Ie, E., and Baldridge, J. (2019). Stay on the path: Instruction fidelity in vision-and-language navigation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1862–1872, Florence, Italy. Association for Computational Linguistics.
  85. 85.Jansen, P. (2020). Visually-grounded planning without vision: Language models infer detailed plans from high-level instructions. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4412–4417, Online. Association for Computational Linguistics.
  86. 86.Jeon, H. J., Losey, D. P., and Sadigh, D. (2020). Shared autonomy with learned latent actions. arXiv preprint arXiv:2005.03210.
  87. 87.Jiang, Z., Francis, J., Sahu, A. K., Munir, S., Shelton, C., Rowe, A., and Berges, M. (2018). Data-driven thermal model inference with armax, in smart environments, based on normalized mutual information. In 2018 Annual American Control Conference (ACC), pages 4634–4639.
  88. 88.Johnson, M., Hofmann, K., Hutton, T., and Bignell, D. (2016). The malmo platform for artificial intelligence experimentation. In Kambhampati, S., editor, Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence, IJCAI 2016, New York, NY, USA, 9-15 July 2016, pages 4246–4247. IJCAI/AAAI Press.
  89. 89.Kaelbling, L. P. and Lozano-Perez, T. (2010). Hierarchical planning in the now. In Workshops at the Twenty-Fourth AAAI Conference on Artificial Intelligence.
  90. 90.Kaelbling, L. P. and Lozano-Perez, T. (2013). Integrated task and motion planning in belief space. The International Journal of Robotics Research, 32(9-10):1194–1227.
  91. 91.Ke, L., Li, X., Bisk, Y., Holtzman, A., Gan, Z., Liu, J., Gao, J., Choi, Y., and Srinivasa, S. S. (2019). Tactical rewind: Self-correction via backtracking in vision-and-language navigation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 6741–6749. Computer Vision Foundation / IEEE.
  92. 92.Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaskowski, W. (2016). Vizdoom: A doom-based AI research platform for visual reinforcement learning. In IEEE Conference on Computational Intelligence and Games, CIG 2016, Santorini, Greece, September 20-23, 2016, pages 1–8. IEEE.
  93. 93.Khan, S., Naseer, M., Hayat, M., Zamir, S. W., Khan, F. S., and Shah, M. (2021). Transformers in vision: A survey. CoRR, abs/2101.01169.
  94. 94.Khatib, O. (1999). Mobile manipulation: The robotic assistant. Robotics and Autonomous Systems, 26(2-3):175–183.
  95. 95.Kim, H., Zala, A., Burri, G., Tan, H., and Bansal, M. (2020). ArraMon: A joint navigation-assembly instruction interpretation task in dynamic environments. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3910–3927, Online. Association for Computational Linguistics.
  96. 96.Kingma, D. P. and Ba, J. (2015). Adam: A method for stochastic optimization. In Bengio, Y. and LeCun, Y., editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  97. 97.Kolve, E., Mottaghi, R., Gordon, D., Zhu, Y., Gupta, A., and Farhadi, A. (2017). AI2-THOR: an interactive 3d environment for visual AI. CoRR, abs/1712.05474.
  98. 98.Krantz, J., Gokaslan, A., Batra, D., Lee, S., and Maksymets, O. (2021). Waypoint models for instruction-guided navigation in continuous environments. CoRR, abs/2110.02207.
  99. 99.Krantz, J., Wijmans, E., Majumdar, A., Batra, D., and Lee, S. (2020). Beyond the nav-graph: Vision-and-language navigation in continuous environments. CoRR, abs/2004.02857.
  100. 100.Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L., Shamma, D. A., Bernstein, M. S., and Li, F. (2016). Visual genome: Connecting language and vision using crowdsourced dense image annotations. CoRR, abs/1602.07332.
  101. 101.Ku, A., Anderson, P., Patel, R., Ie, E., and Baldridge, J. (2020). Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4392–4412, Online. Association for Computational Linguistics.
  102. 102.Ku, A., Anderson, P., Pont-Tuset, J., and Baldridge, J. (2021). Pangea: The panoramic graph environment annotation toolkit. CoRR, abs/2103.12703.
  103. 103.Lamb, A., Goyal, A., Zhang, Y., Zhang, S., Courville, A. C., and Bengio, Y. (2016). Professor forcing: A new algorithm for training recurrent networks. CoRR, abs/1610.09038.
  104. 104.Li, C., Xia, F., Mart´ın-Mart´ın, R., Lingelbach, M., Srivastava, S., Shen, B., Vainio, K., Gokmen, C., Dharan, G., Jain, T., Kurenkov, A., Liu, C. K., Gweon, H., Wu, J., Fei-Fei, L., and Savarese, S. (2021a). igibson 2.0: Object-centric simulation for robot learning of everyday household tasks. CoRR, abs/2108.03272.
  105. 105.Li, C., Xia, F., Martin-Martin, R., and Savarese, S. (2020a). Hrl4in: Hierarchical reinforcement learning for interactive navigation with mobile manipulators. In Conference on Robot Learning, pages 603–616. PMLR.
  106. 106.Li, D., Opazo, C. R., Yu, X., and Li, H. (2020b). Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. In IEEE Winter Conference on Applications of Computer Vision, WACV 2020, Snowmass Village, CO, USA, March 1-5, 2020, pages 1448–1458. IEEE.
  107. 107.Li, J., Yang, F., Tomizuka, M., and Choi, C. (2020c). Evolvegraph: Multi-agent trajectory prediction with dynamic relational reasoning. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H., editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  108. 108.Li, X., Li, C., Xia, Q., Bisk, Y., Celikyilmaz, A., Gao, J., Smith, N. A., and Choi, Y. (2019). Robust navigation with language pretraining and stochastic sampling. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1494–1499, Hong Kong, China. Association for Computational Linguistics.
  109. 109.Li, Y., Goel, P., Rajendra, V. K., Singh, H. S., Francis, J., Ma, K., Nyberg, E., and Oltramari, A. (2021b). Lexically-constrained text generation through commonsense knowledge extraction and injection. In Common Sense Knowledge Graphs at the 35th AAAI Conference on Artificial Intelligence (CSKGs@AAAI-21).
  110. 110.Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  111. 111.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. (2014). Microsoft coco: Common objects in context. In Fleet, D., Pajdla, T., Schiele, B., and Tuytelaars, T., editors, Computer Vision – ECCV 2014, pages 740–755, Cham. Springer International Publishing.
  112. 112.Lipton, Z. C., Berkowitz, J., and Elkan, C. (2015). A critical review of recurrent neural networks for sequence learning. arXiv preprint arXiv:1506.00019.
  113. 113.Lopez-Paz, D., Bottou, L., Scholkopf, B., and Vapnik, V. (2016). Unifying distillation and privileged information. In Bengio, Y. and LeCun, Y., editors, 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings.
  114. 114.Lu, J., Batra, D., Parikh, D., and Lee, S. (2019). Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R., editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13–23.
  115. 115.Lu, J., Goswami, V., Rohrbach, M., Parikh, D., and Lee, S. (2020a). 12-in-1: Multi-task vision and language representation learning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 10434–10443. IEEE.
  116. 116.Lu, X., West, P., Zellers, R., Bras, R. L., Bhagavatula, C., and Choi, Y. (2020b). Neurologic decoding: (un)supervised neural text generation with predicate logic constraints.
  117. 117.Luketina, J., Nardelli, N., Farquhar, G., Foerster, J. N., Andreas, J., Grefenstette, E., Whiteson, S., and Rocktaschel, T. (2019). A survey of reinforcement learning informed by natural language. In Kraus, S., editor, Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI 2019, Macao, China, August 10-16, 2019, pages 6309–6317. ijcai.org.
  118. 118.Ma, C., Lu, J., Wu, Z., AlRegib, G., Kira, Z., Socher, R., and Xiong, C. (2019a). Self-monitoring navigation agent via auxiliary progress estimation. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  119. 119.Ma, C., Wu, Z., AlRegib, G., Xiong, C., and Kira, Z. (2019b). The regretful agent: Heuristic-aided navigation through progress estimation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 6732–6740. Computer Vision Foundation / IEEE.
  120. 120.Ma, K., Francis, J., Lu, Q., Nyberg, E., and Oltramari, A. (2019c). Towards generalizable neuro-symbolic systems for commonsense question answering. In Proceedings of the First Workshop on Commonsense Inference in Natural Language Processing, pages 22–32, Hong Kong, China. Association for Computational Linguistics.
  121. 121.Ma, K., Ilievski, F., Francis, J., Bisk, Y., Nyberg, E., and Oltramari, A. (2021). Knowledge-driven data construction for zero-shot evaluation in commonsense question answering. In Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI-21).
  122. 122.Majumdar, A., Shrivastava, A., Lee, S., Anderson, P., Parikh, D., and Batra, D. (2020). Improving vision-and-language navigation with image-text pairs from the web. CoRR, abs/2004.14973.
  123. 123.Mavrogiannis, C., Baldini, F., Wang, A., Zhao, D., Trautman, P., Steinfeld, A., and Oh, J. (2021). Core challenges of social robot navigation: A survey. arXiv preprint arXiv:2103.05668.
  124. 124.Mehta, H., Artzi, Y., Baldridge, J., Ie, E., and Mirowski, P. (2020). Retouchdown: Adding touchdown to streetlearn as a shareable resource for language grounding tasks in street view.
  125. 125.Mirowski, P., Grimes, M. K., Malinowski, M., Hermann, K. M., Anderson, K., Teplyashin, D., Simonyan, K., Kavukcuoglu, K., Zisserman, A., and Hadsell, R. (2018). Learning to navigate in cities without a map. In Bengio, S., Wallach, H. M., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R., editors, Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montreal, Canada, pages 2424–2435.
  126. 126.Misra, D., Bennett, A., Blukis, V., Niklasson, E., Shatkhin, M., and Artzi, Y. (2018). Mapping instructions to actions in 3D environments with visual goal prediction. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2667–2678, Brussels, Belgium. Association for Computational Linguistics.
  127. 127.Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In Balcan, M. and Weinberger, K. Q., editors, Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, volume 48 of JMLR Workshop and Conference Proceedings, pages 1928–1937. JMLR.org.
  128. 128.Mogadala, A., Kalimuthu, M., and Klakow, D. (2019). Trends in integration of vision and language research: A survey of tasks, datasets, and methods. CoRR, abs/1907.09358.
  129. 129.Munir, S., Arora, R. S., Hesling, C., Li, J., Francis, J., Shelton, C., Martin, C., Rowe, A., and Berges, M. (2017). Real-time fine grained occupancy estimation using depth sensors on arm embedded platforms. In 2017 IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS), pages 295–306. IEEE.
  130. 130.Nachum, O., Gu, S., Lee, H., and Levine, S. (2018). Data-efficient hierarchical reinforcement learning. arXiv preprint arXiv:1805.08296.
  131. 131.Nau, D. S., Au, T.-C., Ilghami, O., Kuter, U., Murdock, J. W., Wu, D., and Yaman, F. (2003). Shop2: An htn planning system. Journal of artificial intelligence research, 20:379–404.
  132. 132.Newman, B. A., Aronson, R. M., Srinivasa, S. S., Kitani, K., and Admoni, H. (2018). Harmonic: A multimodal dataset of assistive human-robot collaboration. arXiv preprint arXiv:1807.11154.
  133. 133.Nguyen, D. and Okatani, T. (2019). Multi-task learning of hierarchical vision-language representation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 10492–10501. Computer Vision Foundation / IEEE.
  134. 134.Nguyen, K. and Daume III, H. (2019). Help, anna! visual navigation with natural multimodal assistance via retrospective curiosity-encouraging imitation learning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 684–695, Hong Kong, China. Association for Computational Linguistics.
  135. 135.Nguyen, K., Dey, D., Brockett, C., and Dolan, B. (2019a). Vision-based navigation with language-based assistance via imitation learning with indirect intervention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12527–12537.
  136. 136.Nguyen, K., Dey, D., Brockett, C., and Dolan, B. (2019b). Vision-based navigation with language-based assistance via imitation learning with indirect intervention. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 12527–12537. Computer Vision Foundation / IEEE.
  137. 137.Nichol, A., Pfau, V., Hesse, C., Klimov, O., and Schulman, J. (2018). Gotta learn fast: A new benchmark for generalization in rl. arXiv preprint arXiv:1804.03720.
  138. 138.Nilsson, N. J. et al. (1984). Shakey the robot.
  139. 139.Oh, J. H., Suppe, A., Duvallet, F., Boularias, A., Navarro-Serment, L. E., Hebert, M., Stentz, A., Vinokurov, J., Romero, O. J., Lebiere, C., and Dean, R. (2015). Toward mobile robots reasoning like humans. In Bonet, B. and Koenig, S., editors, Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, January 25-30, 2015, Austin, Texas, USA, pages 1371–1379. AAAI Press.
  140. 140.Oltramari*, A., Francis*, J., Henson, C., Ma, K., and Wickramarachchi, R. (2020). Neuro-symbolic architectures for context understanding.
  141. 141.Padilla, R., Netto, S. L., and da Silva, E. A. B. (2020). A survey on performance metrics for object-detection algorithms. In 2020 International Conference on Systems, Signals and Image Processing (IWSSIP), pages 237–242.
  142. 142.Padmakumar, A., Thomason, J., Shrivastava, A., Lange, P., Narayan-Chen, A., Gella, S., Piramuthu, R., Tur, G., and Hakkani-Tur, D. (2021). Teach: Task-driven embodied agents that chat. arXiv preprint arXiv:2110.00534.
  143. 143.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002). Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.
  144. 144.Park, S. H., Lee, G., Bhat, M., Seo, J., Kang, M., Francis, J., Jadhav, A. R., Liang, P. P., and Morency, L.-P. (2020). Diverse and admissible trajectory forecasting through multimodal context understanding. In European Conference on Computer Vision (ECCV).
  145. 145.Pennington, J., Socher, R., and Manning, C. (2014). GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  146. 146.Peters, M., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018). Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  147. 147.Puig, X., Ra, K., Boben, M., Li, J., Wang, T., Fidler, S., and Torralba, A. (2018). Virtualhome: Simulating household activities via programs. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 8494–8502. IEEE Computer Society.
  148. 148.Qi, Y., Wu, Q., Anderson, P., Wang, X., Wang, W. Y., Shen, C., and van den Hengel, A. (2020). REVERIE: remote embodied visual referring expression in real indoor environments. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 9979–9988. IEEE.
  149. 149.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Meila, M. and Zhang, T., editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR.
  150. 150.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019). Language models are unsupervised multitask learners.
  151. 151.Ravichander, A., Manzini, T., Grabmair, M., Neubig, G., Francis, J., and Nyberg, E. (2017). How would you say it? eliciting lexically diverse dialogue for supervised semantic parsing. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 374–383, Saarbrucken, Germany. Association for Computational Linguistics.
  152. 152.Rawat, W. and Wang, Z. (2017). Deep convolutional neural networks for image classification: A comprehensive review. Neural Computation, 29(9):2352–2449. PMID: 28599112.
  153. 153.Reddy, S., Dragan, A. D., and Levine, S. (2020). SQIL: imitation learning via reinforcement learning with sparse rewards. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  154. 154.Ross, S., Gordon, G. J., and Bagnell, D. (2011a). A reduction of imitation learning and structured prediction to no-regret online learning. In Gordon, G. J., Dunson, D. B., and Dud´ık, M., editors, Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, AISTATS 2011, Fort Lauderdale, USA, April 11-13, 2011, volume 15 of JMLR Proceedings, pages 627–635. JMLR.org.
  155. 155.Ross, S., Gordon, G. J., and Bagnell, J. A. (2011b). A reduction of imitation learning and structured prediction to no-regret online learning.
  156. 156.Sacerdoti, E. D. (1974). Planning in a hierarchy of abstraction spaces. Artificial intelligence, 5(2):115–135.
  157. 157.Savva, M., Chang, A. X., Dosovitskiy, A., Funkhouser, T. A., and Koltun, V. (2017). MINOS: multimodal indoor simulator for navigation in complex environments. CoRR, abs/1712.03931.
  158. 158.Savva, M., Malik, J., Parikh, D., Batra, D., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., and Koltun, V. (2019). Habitat: A platform for embodied AI research. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 9338–9346. IEEE.
  159. 159.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. CoRR, abs/1707.06347.
  160. 160.Sharma, P., Ding, N., Goodman, S., and Soricut, R. (2018). Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, Melbourne, Australia. Association for Computational Linguistics.
  161. 161.Sharma, P., Torralba, A., and Andreas, J. (2021). Skill induction and planning with latent language. arXiv preprint arXiv:2110.01517.
  162. 162.Shen, B., Xia, F., Li, C., Mart´ın-Mart´ın, R., Fan, L., Wang, G., Perez-D’Arpino, C., Buch, S., Srivastava, S., Tchapmi, L., Tchapmi, M., Vainio, K., Wong, J., Fei-Fei, L., and Savarese, S. (2021a). igibson 1.0: A simulation environment for interactive tasks in large realistic scenes. In IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2021, Prague, Czech Republic, September 27 - Oct. 1, 2021, pages 7520–7527. IEEE.
  163. 163.Shen, S., Li, L. H., Tan, H., Bansal, M., Rohrbach, A., Chang, K., Yao, Z., and Keutzer, K. (2021b). How much can CLIP benefit vision-and-language tasks? CoRR, abs/2107.06383.
  164. 164.Shin, A., Ishii, M., and Narihira, T. (2021). Perspectives and prospects on transformer architecture for cross-modal tasks with language and vision.
  165. 165.Shorten, C. and Khoshgoftaar, T. M. (2019). A survey on image data augmentation for deep learning. Journal of Big Data, 6(1):60.
  166. 166.Shridhar, M., Thomason, J., Gordon, D., Bisk, Y., Han, W., Mottaghi, R., Zettlemoyer, L., and Fox, D. (2020). ALFRED: A benchmark for interpreting grounded instructions for everyday tasks. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 10737–10746. IEEE.
  167. 167.Stentz, A. et al. (1995). The focussed dˆ* algorithm for real-time replanning. In IJCAI, volume 95, pages 1652–1659.
  168. 168.Suhr, A., Yan, C., Schluger, J., Yu, S., Khader, H., Mouallem, M., Zhang, I., and Artzi, Y. (2019). Executing instructions in situated collaborative interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2119–2130, Hong Kong, China. Association for Computational Linguistics.
  169. 169.Sun, F., Chang, Y., Wu, Y., and Lin, S. (2018). Designing non-greedy reinforcement learning agents with diminishing reward shaping. In Furman, J., Marchant, G. E., Price, H., and Rossi, F., editors, Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, AIES 2018, New Orleans, LA, USA, February 02-03, 2018, pages 297–302. ACM.
  170. 170.Sutton, R. S. and Barto, A. G. (2018). Reinforcement learning: An introduction. MIT press.
  171. 171.Syed, U. and Schapire, R. E. (2007). A game-theoretic approach to apprenticeship learning. In Platt, J. C., Koller, D., Singer, Y., and Roweis, S. T., editors, Advances in Neural Information Processing Systems 20, Proceedings of the Twenty-First Annual Conference on Neural Information Processing Systems, Vancouver, British Columbia, Canada, December 3-6, 2007, pages 1449–1456. Curran Associates, Inc.
  172. 172.Synnaeve, G., Nardelli, N., Auvolat, A., Chintala, S., Lacroix, T., Lin, Z., Richoux, F., and Usunier, N. (2016). Torchcraft: a library for machine learning research on real-time strategy games. CoRR, abs/1611.00625.
  173. 173.Szot, A., Clegg, A., Undersander, E., Wijmans, E., Zhao, Y., Turner, J., Maestre, N., Mukadam, M., Chaplot, D. S., Maksymets, O., et al. (2021). Habitat 2.0: Training home assistants to rearrange their habitat. Advances in Neural Information Processing Systems, 34.
  174. 174.Talmor, A., Elazar, Y., Goldberg, Y., and Berant, J. (2020). olmpics - on what language model pre-training captures. Trans. Assoc. Comput. Linguistics, 8:743–758.
  175. 175.Talmor, A., Herzig, J., Lourie, N., and Berant, J. (2019). CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota. Association for Computational Linguistics.
  176. 176.Tan, H. and Bansal, M. (2019). LXMERT: Learning cross-modality encoder representations from transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5100–5111, Hong Kong, China. Association for Computational Linguistics.
  177. 177.Tan, H., Yu, L., and Bansal, M. (2019). Learning to navigate unseen environments: Back translation with environmental dropout. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2610–2621, Minneapolis, Minnesota. Association for Computational Linguistics.
  178. 178.Tangiuchi, T., Mochihashi, D., Nagai, T., Uchida, S., Inoue, N., Kobayashi, I., Nakamura, T., Hagiwara, Y., Iwahashi, N., and Inamura, T. (2019). Survey on frontiers of language and robotics. Advanced Robotics, 33(15-16):700–730.
  179. 179.Tanner, M. A. (2012). Tools for statistical inference: observed data and data augmentation methods, volume 67. Springer Science & Business Media.
  180. 180.Tellex, S., Kollar, T., Dickerson, S., Walter, M. R., Banerjee, A. G., Teller, S. J., and Roy, N. (2011). Understanding natural language commands for robotic navigation and mobile manipulation. In Burgard, W. and Roth, D., editors, Proceedings of the Twenty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2011, San Francisco, California, USA, August 7-11, 2011. AAAI Press.
  181. 181.Thomason, J., Gordon, D., and Bisk, Y. (2019a). Shifting the baseline: Single modality performance on visual navigation & QA. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1977–1983, Minneapolis, Minnesota. Association for Computational Linguistics.
  182. 182.Thomason, J., Murray, M., Cakmak, M., and Zettlemoyer, L. (2019b). Vision-and-dialog navigation. CoRR, abs/1907.04957.
  183. 183.Thrun, S., Burgard, W., and Fox, D. (1998). A probabilistic approach to concurrent mapping and localization for mobile robots. Autonomous Robots, 5(3):253–271.
  184. 184.Tsai, C.-E. and Oh, J. (2020). A generative approach for socially compliant navigation. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 2160–2166. IEEE.
  185. 185.Uppal, S., Bhagat, S., Hazarika, D., Majumdar, N., Poria, S., Zimmermann, R., and Zadeh, A. (2020). Emerging trends of multimodal research in vision and language. arXiv preprint arXiv:2010.09522.
  186. 186.Valuations, E. (2015). A review on evaluation metrics for data classification evaluations.
  187. 187.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R., editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008.
  188. 188.Vedantam, R., Zitnick, C. L., and Parikh, D. (2015). Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pages 4566–4575. IEEE Computer Society.
  189. 189.Vemula, A., Muelling, K., and Oh, J. (2018). Social attention: Modeling attention in human crowds. In 2018 IEEE international Conference on Robotics and Automation (ICRA), pages 4601–4607. IEEE.
  190. 190.Wang, H., Wang, W., Shu, T., Liang, W., and Shen, J. (2020a). Active visual information gathering for vision-language navigation. CoRR, abs/2007.08037.
  191. 191.Wang, H., Wu, Q., and Shen, C. (2020b). Soft expert reward learning for vision-and-language navigation. In Vedaldi, A., Bischof, H., Brox, T., and Frahm, J., editors, Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part IX, volume 12354 of Lecture Notes in Computer Science, pages 126–141. Springer.
  192. 192.Wang, R., Ciliberto, C., Amadori, P. V., and Demiris, Y. (2019a). Random expert distillation: Imitation learning via expert policy support estimation. In Chaudhuri, K. and Salakhutdinov, R., editors, Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 6536–6544. PMLR.
  193. 193.Wang, X., Huang, Q., C¸ elikyilmaz, A., Gao, J., Shen, D., Wang, Y., Wang, W. Y., and Zhang, L. (2019b). Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 6629–6638. Computer Vision Foundation / IEEE.
  194. 194.Wang, X., Jain, V., Ie, E., Wang, W. Y., Kozareva, Z., and Ravi, S. (2019c). Natural language grounded multitask navigation. In Visually Grounded Interaction and Language (ViGIL), NeurIPS 2019 Workshop, Vancouver, Canada, December 13, 2019.
  195. 195.Wang, X., Jain, V., Ie, E., Wang, W. Y., Kozareva, Z., and Ravi, S. (2020c). Environment-agnostic multitask learning for natural language grounded navigation. arXiv preprint arXiv:2003.00443.
  196. 196.Wang, X., Xiong, W., Wang, H., and Wang, W. Y. (2018). Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation. In Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y., editors, Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part XVI, volume 11220 of Lecture Notes in Computer Science, pages 38–55. Springer.
  197. 197.Weihs, L., Salvador, J., Kotar, K., Jain, U., Zeng, K., Mottaghi, R., and Kembhavi, A. (2020). Allenact: A framework for embodied AI research. CoRR, abs/2008.12760.
  198. 198.Wijmans, E., Datta, S., Maksymets, O., Das, A., Gkioxari, G., Lee, S., Essa, I., Parikh, D., and Batra, D. (2019a). Embodied question answering in photorealistic environments with point cloud perception. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 6659–6668. Computer Vision Foundation / IEEE.
  199. 199.Wijmans, E., Kadian, A., Morcos, A., Lee, S., Essa, I., Parikh, D., Savva, M., and Batra, D. (2019b). Decentralized distributed PPO: solving pointgoal navigation. CoRR, abs/1911.00357.
  200. 200.Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach. Learn., 8:229–256.
  201. 201.Wolfe, J., Marthi, B., and Russell, S. (2010). Combined task and motion planning for mobile manipulation. In Twentieth International Conference on Automated Planning and Scheduling.
  202. 202.Wooldridge, M. and Jennings, N. R. (1995). Intelligent agents: Theory and practice. The knowledge engineering review, 10(2):115–152.
  203. 203.Wu, Y., Wu, Y., Gkioxari, G., and Tian, Y. (2018). Building generalizable agents with a realistic and rich 3d environment. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Workshop Track Proceedings. OpenReview.net.
  204. 204.Xia, F., Shen, W. B., Li, C., Kasimbeg, P., Tchapmi, M., Toshev, A., Mart´ın-Mart´ın, R., and Savarese, S. (2019). Interactive gibson: A benchmark for interactive navigation in cluttered environments. CoRR, abs/1910.14442.
  205. 205.Xia, F., Zamir, A. R., He, Z., Sax, A., Malik, J., and Savarese, S. (2018). Gibson env: Real-world perception for embodied agents. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 9068–9079. IEEE Computer Society.
  206. 206.Xiang, F., Qin, Y., Mo, K., Xia, Y., Zhu, H., Liu, F., Liu, M., Jiang, H., Yuan, Y., Wang, H., Yi, L., Chang, A. X., Guibas, L. J., and Su, H. (2020). SAPIEN: A simulated part-based interactive environment. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 11094–11104. IEEE.
  207. 207.Yan, A., Wang, X. E., Feng, J., Li, L., and Wang, W. Y. (2020). Beyond monolingual vision-language navigation.
  208. 208.Yan, C., Misra, D. K., Bennett, A., Walsman, A., Bisk, Y., and Artzi, Y. (2018). CHALET: cornell house agent learning environment. CoRR, abs/1801.07357.
  209. 209.Yang, S., Wang, Y., and Chu, X. (2020). A survey of deep learning techniques for neural machine translation. arXiv preprint arXiv:2002.07526.
  210. 210.Yang, W., Wang, X., Farhadi, A., Gupta, A., and Mottaghi, R. (2019a). Visual semantic navigation using scene priors. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  211. 211.Yang, Z., Dai, Z., Yang, Y., Carbonell, J. G., Salakhutdinov, R., and Le, Q. V. (2019b). Xlnet: Generalized autoregressive pretraining for language understanding. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R., editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 5754–5764.
  212. 212.Ye, X. and Yang, Y. (2020). From seeing to moving: A survey on learning for visual indoor navigation (vin).
  213. 213.Yu, F., Deng, Z., Narasimhan, K., and Russakovsky, O. (2020). Take the scenic route: Improving generalization in vision-and-language navigation. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR Workshops 2020, Seattle, WA, USA, June 14-19, 2020, pages 4000–4004. IEEE.
  214. 214.Yu, L., Chen, X., Gkioxari, G., Bansal, M., Berg, T. L., and Batra, D. (2019). Multi-target embodied question answering. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 6309–6318. Computer Vision Foundation / IEEE.
  215. 215.Zhang, Y., Tan, H., and Bansal, M. (2020). Diagnosing the environment bias in vision-and-language navigation. In Bessiere, C., editor, Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI 2020, pages 890–897. ijcai.org.
  216. 216.Zhang, Z., Li, Q., Huang, Z., Wu, J., Tenenbaum, J., and Freeman, B. (2017). Shape and material from sound. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R., editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 1278–1288.
  217. 217.Zhao, M., Anderson, P., Jain, V., Wang, S., Ku, A., Baldridge, J., and Ie, E. (2021). On the evaluation of vision-and-language navigation instructions. CoRR, abs/2101.10504.
  218. 218.Zhu, F., Zhu, Y., Chang, X., and Liang, X. (2020a). Vision-language navigation with self-supervised auxiliary reasoning tasks. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 10009–10019. IEEE.
  219. 219.Zhu, H., Neubig, G., and Bisk, Y. (2021a). Few-shot language coordination by modeling theory of mind. In International Conference on Machine Learning, pages 12901–12911. PMLR.
  220. 220.Zhu, W., Hu, H., Chen, J., Deng, Z., Jain, V., Ie, E., and Sha, F. (2020b). BabyWalk: Going farther in vision-and-language navigation by taking baby steps. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2539–2556, Online. Association for Computational Linguistics.
  221. 221.Zhu, W., Qi, Y., Narayana, P., Sone, K., Basu, S., Wang, X. E., Wu, Q., Eckstein, M. P., and Wang, W. Y. (2021b). Diagnosing vision-and-language navigation: What really matters. CoRR, abs/2103.16561.
  222. 222.Zhu, Y., Zhu, F., Zhan, Z., Lin, B., Jiao, J., Chang, X., and Liang, X. (2020c). Vision-dialog navigation by exploring cross-modal memory. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 10727–10736. IEEE.

Citation

MLA
Francis, J., et al. “Core Challenges in Embodied Vision-Language Planning”. Journal of Artificial Intelligence Research, vol. 74, 2022, pp. 459–515, https://doi.org/10.1613/JAIR.1.13646.
APA
Francis, J., Kitamura, N., Labelle, F., Lu, X., Navarro, I., & Oh, J. (2022). Core Challenges in Embodied Vision-Language Planning. Journal of Artificial Intelligence Research, 74, 459–515. https://doi.org/10.1613/JAIR.1.13646
Chicago
Francis, J., N. Kitamura, F. Labelle, X. Lu, I. Navarro, and J. Oh. 2022. “Core Challenges in Embodied Vision-Language Planning”. Journal of Artificial Intelligence Research 74: 459–515. https://doi.org/10.1613/JAIR.1.13646.
Harvard
Francis, J. et al. (2022) “Core Challenges in Embodied Vision-Language Planning”, Journal of Artificial Intelligence Research, 74, pp. 459–515. Available at: https://doi.org/10.1613/JAIR.1.13646.
Vancouver
1. Francis J, Kitamura N, Labelle F, Lu X, Navarro I, Oh J (2022) Core Challenges in Embodied Vision-Language Planning. Journal of Artificial Intelligence Research 74:459–515

BibTeX

@article{Francis_2022, title={Core Challenges in Embodied Vision-Language Planning}, volume={74}, ISSN={1076-9757}, url={http://dx.doi.org/10.1613/JAIR.1.13646}, DOI={10.1613/jair.1.13646}, journal={Journal of Artificial Intelligence Research}, publisher={AI Access Foundation}, author={Francis, Jonathan and Kitamura, Nariaki and Labelle, Felix and Lu, Xiaopeng and Navarro, Ingrid and Oh, Jean}, year={2022}, month=May, pages={459–515} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/