Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions

Jing GuEliana StefaniQi WuJesse ThomasonXin Wang

article2022ACL170 citations

Presents a comprehensive taxonomy of vision-and-language navigation benchmarks organized by communication complexity and task objectives, while systematically evaluating core modeling strategies, datasets, and open research challenges.

Listen

Building autonomous systems that understand human language, perceive physical environments, and carry out complex real-world tasks is a critical objective for artificial intelligence and robotics. Vision-and-language navigation addresses this need by investigating how embodied agents interpret natural instructions, navigate physical or simulated spaces, and interact with surroundings. The article systematically reviews the field by categorizing existing tasks, evaluation standards, and methodologies, while highlighting key challenges and strategic directions for future development.

The article establishes a taxonomy that classifies benchmarks across communication complexity and task objectives, ranging from single initial commands to continuous human dialogue, and from simple pathfinding to complex object manipulation. It also organizes technical solutions into four pillars: cross-modal representation learning, action strategy optimization, data-centric techniques, and prior environmental exploration. Across these domains, the analysis reveals that high-level semantic alignment, memory structures, and task-specific pretraining significantly outperform raw sensory models, improving success rates in unseen environments from around 20% in baseline models to upwards of 60–67% in modern systems. Furthermore, data augmentation and curriculum learning mitigate severe data scarcity, while prior exploration bridges the performance gap between known and unknown environments.

These findings indicate that while visual and linguistic reasoning algorithms are maturing rapidly, significant operational risks remain. Most current models rely on unrealistic assumptions—such as idealized teleportation between discrete nodes, perfect localization, static environments, and limited indoor scenes dominated by Western residential layouts. Transferring these simulated systems to continuous, real-world robotics will lead to substantial performance drops and safety hazards unless underlying dynamics, lighting variations, and sensor noise are resolved.

Organizations developing embodied AI should prioritize transitioning from discrete navigation graphs to continuous physical environments and invest in physical object manipulation beyond simple waypoint traversal. Additionally, development roadmaps must incorporate multi-agent and human-robot collaboration, rigorous data privacy safeguards, and geographically diverse testing environments. Because existing research is largely constrained to synthetic or residential simulations, stakeholders should maintain caution regarding immediate physical deployment and fund pilot validations in dynamic, real-world settings such as warehouses and healthcare facilities.

Cover for Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions

Abstract

A long-term goal of AI research is to build intelligent agents that can communicate with humans in natural language, perceive the environment, and perform real-world tasks. Vision-and-Language Navigation (VLN) is a fundamental and interdisciplinary research topic towards this goal, and receives increasing attention from natural language processing, computer vision, robotics, and machine learning communities. In this paper, we review contemporary studies in the emerging field of VLN, covering tasks, evaluation metrics, methods, etc. Through structured analysis of current progress and challenges, we highlight the limitations of current VLN and opportunities for future work. This paper serves as a thorough reference for the VLN research community.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Tasks and Datasets
  • 2.1 Initial Instruction
  • 2.2 Oracle Guidance
  • 2.3 Human Dialogue
  • 3 Evaluation
  • 4 VLN Methods
  • 4.1 Representation Learning
  • 4.1.1 Pretraining
  • 4.1.2 Semantic Understanding
  • 4.1.3 Graph Representation
  • 4.1.4 Memory-augmented Model
  • 4.1.5 Auxiliary Tasks
  • 4.2 Action Strategy Learning
  • 4.2.1 Reinforcement Learning
  • 4.2.2 Exploration during Navigation
  • 4.2.3 Navigation Planning
  • 4.2.4 Asking for Help
  • 4.3 Data-centric Learning
  • 4.3.1 Data Augmentation
  • 4.3.2 Curriculum Learning
  • 4.3.3 Multitask Learning
  • 4.3.4 Instruction Interpretation
  • 4.4 Prior Exploration
  • 5 Related Visual-and-Language Tasks
  • 6 Conclusion and Future Directions
  • References
  • A Dataset Details
  • B Simulator
  • C Room-to-Room Leaderboard

Knowls

  1. Knowl 1 — VLN tasks are organized by communication level and task objective

    definition

    Vision-and-Language Navigation (VLN) tasks can be classified along two independent dimensions. Communication complexity has three levels: an agent receives only an initial instruction; it can request additional oracle guidance, often by signaling that it needs help; or it can ask and answer free-form natural-language questions during navigation. Task objective also has three levels: fine-grained navigation follows a detailed route description; coarse-grained navigation finds a more distant goal from a less detailed description, requiring the agent to plan a path and possibly seek guidance; navigation with object interaction requires both reaching locations and changing or inspecting objects to complete the task. The task matrix therefore distinguishes navigation-only tasks from tasks requiring interaction, as well as passive instruction following from interactive communication.

  2. Knowl 2 — VLN benchmarks occupy distinct combinations of communication and task objective

    data/table

    The benchmark landscape maps datasets to the communication and objective categories of VLN. For initial-instruction fine-grained navigation, examples include Room-to-Room (R2R), Room-for-Room, Room-Across-Room, XL-R2R, RxR and its Landmark-RxR extension, VLNCE, TOUCHDOWN, StreetLearn, StreetNav, Talk2Nav, and LANI. Initial-instruction coarse-grained navigation includes RoomNav, EmbodiedQA, REVERIE, and SOON. Initial-instruction navigation with object interaction includes IQA, CHAI, and ALFRED. For oracle-guidance tasks, Just Ask covers fine-grained navigation, while VNLA and HANNA cover coarse-grained navigation; the survey identifies no object-interaction benchmark in this communication category. For free-form dialogue, coarse-grained examples include CVDN, RobotSlang, Talk the Walk, and CEREALBAR, while interaction-oriented examples include TEACh, Minecraft Collaborative Building, and DialFRED. The survey identifies no dialogue benchmark for fine-grained route navigation. This organization makes clear that benchmarks cover different combinations of language interaction, navigation precision, and physical task complexity.

  3. Knowl 3 — VLN evaluation distinguishes goal achievement from route fidelity

    definition

    VLN evaluation metrics measure either proximity to task completion or adherence to a desired route. Goal-oriented metrics include Success Rate (SR), the frequency of reaching the goal within an allowed distance; Goal Progress, the reduction in remaining distance to the target; Path Length (PL), the total distance traveled; and Shortest-Path Distance (SPD), the mean distance from the final agent location to the goal. Success weighted by Path Length (SPL) combines success with a preference for shorter paths. Success weighted by Edit Distance (SED) compares expert and agent action sequences or trajectories while also accounting for success and path cost. Oracle Navigation Error (ONE) measures the shortest distance from any visited path node to the goal, and Oracle Success Rate (OSR) records whether any visited node comes within a goal-distance threshold. Path-fidelity metrics evaluate similarity to an expert route: Path Coverage (PC) measures route coverage and Length Score (LS) measures path-length quality; Coverage weighted by LS (CLS) is their product. Normalized Dynamic Time Warping (nDTW) softly penalizes deviations between two paths, while Success weighted by nDTW (SDTW) applies the fidelity measure only to successful episodes.

  4. Knowl 4 — Representation learning connects language, visual observations, and navigation history

    model/method

    VLN representation methods address how agents relate instructions to observations and actions. Pretraining initializes visual or language encoders from single-modality models, uses joint vision-and-language representations, or trains on navigation-specific data such as instruction–path compatibility or image–text–action sequences. Semantic understanding selects useful features within a modality—for example, objects, route structure, or instruction tokens for objects and directions—and aligns relevant instruction content with observed scenes, objects, and candidate actions. Graph representations encode relations between textual entities and visual entities or track locations and navigable structure to guide trajectory or action prediction. Memory-augmented models preserve useful information from the navigation history; approaches range from recurrent states to separate memory modules and hierarchical encoding of views, panoramas, and their history. Auxiliary tasks add training objectives that help an agent estimate progress or task completion, explain prior actions, predict future decisions, or align instructions and visual observations.

  5. Knowl 5 — Action-strategy learning addresses sequential decisions, planning, exploration, and help-seeking

    model/method

    VLN action-strategy methods help an agent choose actions over a multi-step trajectory. Reinforcement learning treats navigation as sequential decision-making and addresses sparse end-of-episode success feedback with more informative rewards, including goal-oriented rewards, instruction-fidelity rewards, local instruction–landmark alignment, and signals based on path-fidelity metrics. Model-based approaches predict environmental state transitions; some methods alternate imitation and reinforcement learning. Exploration during navigation gathers information before or while acting, with approaches that trade off exploration against path length, decide when to backtrack, or predict future visual observations. Navigation planning predicts waypoints, future states or rewards, observations, or relevant neighboring views; it can also use instruction landmarks and direction cues to plan steps. Help-seeking determines when to ask an oracle, using action uncertainty or a separately trained decision model, and supports either signal-based requests or natural-language dialogue. Dialogue-capable systems may also need an oracle that generates or selects answers.

  6. Knowl 6 — Data-centric VLN methods expand, reorganize, and reuse training signals

    model/method

    Data-centric VLN methods make better use of limited training data or create additional training examples. Trajectory–instruction augmentation commonly uses a speaker model to generate instructions for paths; because generated pairs vary in quality, alignment scorers or adversarial discriminators can filter them. Environment augmentation creates variation by masking visual features across viewpoints, remixing parts of scenes, or substituting counterfactual visual features, and can support generation of more trajectory–instruction pairs. Curriculum learning gradually increases task difficulty, using properties such as instruction length or the number of rooms traversed to order training examples. Multitask learning transfers knowledge across related tasks, such as instruction following and navigation from dialogue history, or navigation and question answering. Instruction interpretation re-encodes instructions in multiple ways, identifies relevant object classes, or decomposes long instructions into shorter atomic units so an agent can track progress more clearly.

  7. Knowl 7 — Prior exploration adapts navigation to previously unseen environments

    model/method

    Prior exploration methods let a VLN agent inspect or adapt to an environment before carrying out instruction-following navigation there, with the aim of improving generalization to unseen environments. Described approaches include self-supervised imitation from the agent’s own previously successful behavior, sampling and augmenting paths using the test environment, and exploration restricted to the particular environment where the agent will operate. Graph-based approaches can build a map or overview of an unseen environment to guide later navigation. These methods may have access to extra test-environment observations, so their performance is not directly comparable with methods that do not use prior exploration.

  8. Knowl 8 — VLN datasets concentrate in a small set of simulated environments

    empirical result

    The surveyed benchmark collection contains more indoor than outdoor datasets, and its environment coverage is uneven. Matterport3D underlies many indoor benchmarks, including R2R and related datasets; House3D and AI2-THOR support other indoor tasks, with AI2-THOR particularly associated with object-interaction benchmarks. Habitat and Gibson also provide embodied-agent environments. Many photo-realistic outdoor benchmarks use Google Street View, while LANI and some drone-navigation tasks use a synthetic Unity3D environment. In common graph-based indoor benchmarks, agents move among predefined viewpoints, rather than navigating freely in continuous space. The survey’s benchmark and simulator inventories thus show both concentration in a few underlying environments and limited coverage of the range of real-world settings.

  9. Knowl 9 — VLN still has gaps in external knowledge, object manipulation, and environmental diversity

    limitation

    The survey identifies several unresolved capability and coverage gaps. Current navigation methods often do not explicitly use external knowledge, such as general information about objects or descriptions of homes. Agents may learn where to move and what to interact with without learning the final manipulation needed to complete a task—for example, locating a spoon is distinct from picking it up. The surveyed environments also lack diversity: many indoor scenes are American homes rather than settings such as warehouses or hospitals, and many outdoor datasets use major American cities. This limited coverage may hinder generalization to other layouts, regions, and cultures.

  10. Knowl 10 — Future VLN work should address realistic transfer, collaboration, privacy, and cultural coverage

    limitation

    The survey proposes several directions for extending VLN. Simulation-to-reality transfer needs to account for continuous robot motion, noisy localization, the absence of teleportation between viewpoints, and environments whose topology is not known in advance. Simulations could better match real sensors by combining 3D models with realistic imagery and could represent changing scenes, moving people or objects, lighting variation, and probabilistic transitions. Collaborative VLN could study coordination among multiple robots and between humans and robots, rather than treating the human only as an oracle. Privacy-preserving VLN is needed because agents may observe and store sensitive information during training and deployment. Multicultural VLN calls for environments from more regions and cultures, alongside multilingual data, to investigate whether agents generalize beyond the environments and languages currently represented.

Coverage note — The dated Room-to-Room leaderboard and exhaustive simulator descriptions are omitted: the leaderboard mixes evaluation modes that the survey says are not directly comparable, while the simulator inventory is descriptive rather than a separate methodological contribution.

References

  1. 1.Dong An, Yuankai Qi, Yan Huang, Qi Wu, Liang Wang, and Tieniu Tan. 2021. Neighbor-view enhanced model for vision and language navigation. arXiv preprint arXiv:2107.07201.
  2. 2.Peter Anderson, Angel Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, et al. 2018a. On evaluation of embodied navigation agents. arXiv preprint arXiv:1807.06757.
  3. 3.Peter Anderson, Ayush Shrivastava, Devi Parikh, Dhruv Batra, and Stefan Lee. 2019a. Chasing ghosts: Instruction following as bayesian state tracking. In Advances in Neural Information Processing Systems (NeurIPS).
  4. 4.Peter Anderson, Ayush Shrivastava, Devi Parikh, Dhruv Batra, and Stefan Lee. 2019b. Chasing ghosts: Instruction following as bayesian state tracking. Advances in Neural Information Processing Systems, 32:371–381.
  5. 5.Peter Anderson, Ayush Shrivastava, Joanne Truong, Arjun Majumdar, Devi Parikh, Dhruv Batra, and Stefan Lee. 2020. Sim-to-real transfer for vision-and-language navigation. In Conference on Robot Learning (CoRL).
  6. 6.Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. 2018b. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  7. 7.Shurjo Banerjee, Jesse Thomason, and Jason J. Corso. 2020. The RobotSlang Benchmark: Dialog-guided robot localization and navigation. In Conference on Robot Learning (CoRL).
  8. 8.Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pages 41–48.
  9. 9.Valts Blukis, Nataly Brukhim, Andrew Bennett, Ross A. Knepper, and Yoav Artzi. 2018. Following high-level navigation instructions on a simulated quadcopter with imitation learning. In Robotics: Science and Systems (RSS).
  10. 10.Valts Blukis, Yannick Terme, Eyvind Niklasson, Ross A. Knepper, and Yoav Artzi. 2019. Learning to map natural language instructions to physical quadcopter control using simulated flight. In Conference on Robot Learning (CoRL).
  11. 11.Valts Blukis, Yannick Terme, Eyvind Niklasson, Ross A. Knepper, and Yoav Artzi. 2020. Learning to map natural language instructions to physical quadcopter control using simulated flight. In Proceedings of the Conference on Robot Learning, volume 100 of Proceedings of Machine Learning Research, pages 1415–1438. PMLR.
  12. 12.Francisco Bonin-Font, Alberto Ortiz, and Gabriel Oliver. 2008. Visual navigation for mobile robots: A survey. Journal of intelligent and robotic systems, 53(3):263–296.
  13. 13.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. 2017. Matterport3D: Learning from RGB-D data in indoor environments. International Conference on 3D Vision (3DV).
  14. 14.Devendra Singh Chaplot, Lisa Lee, Ruslan Salakhutdinov, Devi Parikh, and Dhruv Batra. 2020. Embodied multimodal multitask learning. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20. International Joint Conferences on Artificial Intelligence Organization.
  15. 15.David Chen and Raymond Mooney. 2011. Learning to interpret natural language navigation instructions from observations. In AAAI Conference on Artificial Intelligence.
  16. 16.Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi. 2019. Touchdown: Natural language navigation and spatial reasoning in visual street environments. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12530–12539.
  17. 17.Kevin Chen, Junshen K Chen, Jo Chuang, Marynel Vázquez, and Silvio Savarese. 2021a. Topological planning with transformers for vision-and-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11276–11286.
  18. 18.Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, and Ivan Laptev. 2021b. History aware multimodal transformer for vision-and-language navigation. arXiv preprint arXiv:2110.13309.
  19. 19.Ta-Chung Chi, Minmin Shen, Mihail Eric, Seokhwan Kim, and Dilek Hakkani-tur. 2020. Just ask: An interactive learning framework for vision and language navigation. In AAAI Conference on Artificial Intelligence.
  20. 20.Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra. 2018. Embodied question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1–10.
  21. 21.Harm de Vries, Kurt Shuster, Dhruv Batra, Devi Parikh, Jason Weston, and Douwe Kiela. 2018. Talk the walk: Navigating new york city through grounded dialogue.
  22. 22.Zhiwei Deng, Karthik Narasimhan, and Olga Russakovsky. 2020. Evolving graphical planner: Contextual global planning for vision-and-language navigation. Advances in Neural Information Processing Systems, 2020-December.
  23. 23.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT (1).
  24. 24.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.
  25. 25.Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. 2006. Calibrating noise to sensitivity in private data analysis. In Theory of cryptography conference, pages 265–284. Springer.
  26. 26.Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. 2018. Speaker-follower models for vision-and-language navigation. In Neural Information Processing Systems (NeurIPS).
  27. 27.Justin Fu, Anoop Korattikara, Sergey Levine, and Sergio Guadarrama. 2019. From language to goals: Inverse reinforcement learning for vision-based instruction following. arXiv preprint arXiv:1902.07742.
  28. 28.Tsu-Jui Fu, Xin Eric Wang, Matthew Peterson, Scott Grafton, Miguel Eckstein, and William Yang Wang. 2020. Counterfactual vision-and-language navigation via adversarial path sampler. In European Conference on Computer Vision (ECCV).
  29. 29.Chen Gao, Jinyu Chen, Si Liu, Luting Wang, Qiong Zhang, and Qi Wu. 2021. Room-and-object aware knowledge reasoning for remote embodied referring expression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3064–3073.
  30. 30.Xiaofeng Gao, Qiaozi Gao, Ran Gong, Kaixiang Lin, Govind Thattai, and Gaurav S Sukhatme. 2022. Dialfred: Dialogue-enabled agents for embodied instruction following. arXiv preprint arXiv:2202.13330.
  31. 31.Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi. 2018. Iqa: Visual question answering in interactive environments. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4089–4098.
  32. 32.Pierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen, Ivan Laptev, and Cordelia Schmid. 2021. Airbert: In-domain pretraining for vision-and-language navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1634–1643.
  33. 33.Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao. 2020. Towards learning a generic agent for vision-and-language navigation via pretraining. Conference on Computer Vision and Pattern Recognition (CVPR).
  34. 34.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778.
  35. 35.Keji He, Yan Huang, Qi Wu, Jianhua Yang, Dong An, Shuanglin Sima, and Liang Wang. 2021. Landmark-rxr: Solving vision-and-language navigation with fine-grained alignment supervision. In NeurIPS.
  36. 36.Karl Moritz Hermann, Mateusz Malinowski, Piotr Mirowski, Andras Banki-Horvath, Keith Anderson, and Raia Hadsell. 2020. Learning to follow directions in street view. In AAAI Conference on Artificial Intelligence.
  37. 37.Yicong Hong, Cristian Rodriguez, Yuankai Qi, Qi Wu, and Stephen Gould. 2020a. Language and visual entity relationship graph for agent navigation. Advances in Neural Information Processing Systems, 33:7685–7696.
  38. 38.Yicong Hong, Cristian Rodriguez, Qi Wu, and Stephen Gould. 2020b. Sub-instruction aware vision-and-language navigation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3360–3376, Online. Association for Computational Linguistics.
  39. 39.Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez-Opazo, and Stephen Gould. 2021. Vln bert: A recurrent vision-and-language bert for navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1643–1653.
  40. 40.Ronghang Hu, Daniel Fried, Anna Rohrbach, Dan Klein, Trevor Darrell, and Kate Saenko. 2019. Are you looking? grounding to multiple modalities in vision-and-language navigation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6551–6557, Florence, Italy. Association for Computational Linguistics.
  41. 41.Haoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku, Gabriel Magalhaes, Jason Baldridge, and Eugene Ie. 2019. Transferable representation learning in vision-and-language navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
  42. 42.Gabriel Ilharco, Vihan Jain, Alexander Ku, Eugene Ie, and Jason Baldridge. 2019. General evaluation for instruction conditioned navigation using dynamic time warping. arXiv preprint arXiv:1907.05446.
  43. 43.Nikolai Ilinykh, Sina Zarrieß, and David Schlangen. 2019. Meetup! a corpus of joint activity dialogues in a visual environment. arXiv preprint arXiv:1907.05084.
  44. 44.Muhammad Zubair Irshad, Chih-Yao Ma, and Zsolt Kira. 2021. Hierarchical cross-modal agent for robotics vision-and-language navigation. arXiv preprint arXiv:2104.10674.
  45. 45.Vihan Jain, Gabriel Magalhaes, Alexander Ku, Ashish Vaswani, Eugene Ie, and Jason Baldridge. 2019. Stay on the path: Instruction fidelity in vision-and-language navigation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1862–1872, Florence, Italy. Association for Computational Linguistics.
  46. 46.Liyiming Ke, Xiujun Li, Yonatan Bisk, Ari Holtzman, Zhe Gan, Jingjing Liu, Jianfeng Gao, Yejin Choi, and Siddhartha Srinivasa. 2019. Tactical rewind: Self-correction via backtracking in vision-and-language navigation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  47. 47.Jing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge, and Peter Anderson. 2021. Pathdreamer: A world model for indoor navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 14738–14748.
  48. 48.Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi. 2017. AI2-THOR: An Interactive 3D Environment for Visual AI. arXiv.
  49. 49.Jakub Konecnˇ y, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon. 2016. Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492.
  50. 50.Jacob Krantz, Aaron Gokaslan, Dhruv Batra, Stefan Lee, and Oleksandr Maksymets. 2021. Waypoint models for instruction-guided navigation in continuous environments. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15162–15171.
  51. 51.Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee. 2020. Beyond the nav-graph: Vision-and-language navigation in continuous environments. In Computer Vision – ECCV 2020, pages 104–120, Cham. Springer International Publishing.
  52. 52.Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. 2020. Room-Across-Room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. In Conference on Empirical Methods for Natural Language Processing (EMNLP).
  53. 53.Federico Landi, Lorenzo Baraldi, Marcella Cornia, Massimiliano Corsini, and Rita Cucchiara. 2020. Perceive, transform, and act: Multi-modal attention networks for vision-and-language navigation.
  54. 54.Federico Landi, Lorenzo Baraldi, Massimiliano Corsini, and Rita Cucchiara. 2019. Embodied vision-and-language navigation with dynamic convolutional filters. In Proceedings of the British Machine Vision Conference.
  55. 55.Xiujun Li, Chunyuan Li, Qiaolin Xia, Yonatan Bisk, Asli Celikyilmaz, Jianfeng Gao, Noah Smith, and Yejin Choi. 2019. Robust navigation with language pretraining and stochastic sampling. In Empirical Methods in Natural Language Processing (EMNLP).
  56. 56.Xiangru Lin, Guanbin Li, and Yizhou Yu. 2021. Scene-intuitive agent for remote embodied visual grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7036–7045.
  57. 57.Chong Liu, Fengda Zhu, Xiaojun Chang, Xiaodan Liang, Zongyuan Ge, and Yi-Dong Shen. 2021. Vision-language navigation with random environmental mixup. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1644–1654.
  58. 58.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  59. 59.Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, Zsolt Kira, Richard Socher, and Caiming Xiong. 2019a. Self-monitoring navigation agent via auxiliary progress estimation. In International Conference on Learning Representations (ICLR).
  60. 60.Chih-Yao Ma, Zuxuan Wu, Ghassan AlRegib, Caiming Xiong, and Zsolt Kira. 2019b. The regretful agent: Heuristic-aided navigation through progress estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  61. 61.Matt MacMahon, Brian Stankiewicz, and Benjamin Kuipers. 2006. Walk the talk: Connecting language, knowledge, and action in route instructions. Def, 2(6):4.
  62. 62.Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra. 2020. Improving vision-and-language navigation with image-text pairs from the web. In Proceedings of the European Conference on Computer Vision (ECCV).
  63. 63.Manolis Savva*, Abhishek Kadian*, Oleksandr Maksymets*, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. 2019. Habitat: A Platform for Embodied AI Research. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
  64. 64.Harsh Mehta, Yoav Artzi, Jason Baldridge, Eugene Ie, and Piotr Mirowski. 2020. Retouchdown: Releasing touchdown on StreetLearn as a public resource for language grounding tasks in street view. In Proceedings of the Third International Workshop on Spatial Language Understanding, pages 56–62, Online. Association for Computational Linguistics.
  65. 65.Piotr Mirowski. 2019. Learning to navigate. In 1st International Workshop on Multimodal Understanding and Learning for Embodied Applications, MULEA ’19, page 25, New York, NY, USA. Association for Computing Machinery.
  66. 66.Piotr Mirowski, Andras Banki-Horvath, Keith Anderson, Denis Teplyashin, Karl Moritz Hermann, Mateusz Malinowski, Matthew Koichi Grimes, Karen Simonyan, Koray Kavukcuoglu, Andrew Zisserman, et al. 2019. The streetlearn environment and dataset. arXiv preprint arXiv:1903.01292.
  67. 67.Piotr Mirowski, Matthew Koichi Grimes, Mateusz Malinowski, Karl Moritz Hermann, Keith Anderson, Denis Teplyashin, Karen Simonyan, Koray Kavukcuoglu, Andrew Zisserman, and Raia Hadsell. 2018. Learning to navigate in cities without a map. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, page 2424–2435, Red Hook, NY, USA. Curran Associates Inc.
  68. 68.Dipendra Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, and Yoav Artzi. 2018. Mapping instructions to actions in 3d environments with visual goal prediction. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2667–2678.
  69. 69.Abhinav Moudgil, Arjun Majumdar, Harsh Agrawal, Stefan Lee, and Dhruv Batra. 2021. Soat: A scene- and object-aware transformer for vision-and-language navigation. In NeurIPS.
  70. 70.Anjali Narayan-Chen, Prashant Jayannavar, and Julia Hockenmaier. 2019. Collaborative dialogue in Minecraft. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy. Association for Computational Linguistics.
  71. 71.Khanh Nguyen, Yonatan Bisk, and Hal Daumé III au2. 2021a. Learning when and what to ask: a hierarchical reinforcement learning framework.
  72. 72.Khanh Nguyen and Hal Daumé III. 2019. Help, anna! visual navigation with natural multimodal assistance via retrospective curiosity-encouraging imitation learning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 684–695, Hong Kong, China. Association for Computational Linguistics.
  73. 73.Khanh Nguyen, Debadeepta Dey, Chris Brockett, and Bill Dolan. 2019. Vision-based navigation with language-based assistance via imitation learning with indirect intervention. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  74. 74.Van-Quang Nguyen, Masanori Suganuma, and Takayuki Okatani. 2021b. Look wide and interpret twice: Improving performance on interactive instruction-following tasks. arXiv preprint arXiv:2106.00596.
  75. 75.Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramithu, Gokhan Tur, and Dilek Hakkani-Tur. 2021. Teach: Task-driven embodied agents that chat.
  76. 76.Amin Parvaneh, Ehsan Abbasnejad, Damien Teney, Qinfeng Shi, and Anton van den Hengel. 2020. Counterfactual vision-and-language navigation: Unravelling the unseen. In NeurIPS.
  77. 77.Alexander Pashevich, Cordelia Schmid, and Chen Sun. 2021. Episodic transformer for vision-and-language navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15942–15952.
  78. 78.Tzuf Paz-Argaman and Reut Tsarfaty. 2019. Run through the streets: A new dataset and baseline models for realistic urban navigation. arXiv preprint arXiv:1909.08970.
  79. 79.Yuankai Qi, Zizheng Pan, Yicong Hong, Ming-Hsuan Yang, Anton van den Hengel, and Qi Wu. 2021. The road to know-where: An object-and-room informed sequential bert for indoor vision-language navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1655–1664.
  80. 80.Yuankai Qi, Zizheng Pan, Shengping Zhang, Anton van den Hengel, and Qi Wu. 2020a. Object-and-action aware model for visual language navigation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part X 16, pages 303–317. Springer.
  81. 81.Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. 2020b. Reverie: Remote embodied visual referring expression in real indoor environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  82. 82.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  83. 83.Homero Roman Roman, Yonatan Bisk, Jesse Thomason, Asli Celikyilmaz, and Jianfeng Gao. 2020. RMM: A recursive mental model for dialog navigation. In Findings of Empirical Methods in Natural Language Processing (EMNLP Findings).
  84. 84.Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020. ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  85. 85.Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser. 2017. Semantic scene completion from a single depth image. CVPR.
  86. 86.Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J. Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, Anton Clarkson, Mingfei Yan, Brian Budge, Yajie Yan, Xiaqing Pan, June Yon, Yuyang Zou, Kimberly Leon, Nigel Carter, Jesus Briales, Tyler Gillingham, Elias Mueggler, Luis Pesqueira, Manolis Savva, Dhruv Batra, Hauke M. Strasdat, Renzo De Nardi, Michael Goesele, Steven Lovegrove, and Richard Newcombe. 2019. The Replica dataset: A digital replica of indoor spaces. arXiv preprint arXiv:1906.05797.
  87. 87.Alane Suhr, Claudia Yan, Jack Schluger, Stanley Yu, Hadi Khader, Marwa Mouallem, Iris Zhang, and Yoav Artzi. 2019. Executing instructions in situated collaborative interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2119–2130, Hong Kong, China. Association for Computational Linguistics.
  88. 88.Q. Sun, Y. Zhuang, Z. Chen, Y. Fu, and X. Xue. 2021. Depth-guided adain and shift attention network for vision-and-language navigation. In 2021 IEEE International Conference on Multimedia and Expo (ICME), pages 1–6, Los Alamitos, CA, USA. IEEE Computer Society.
  89. 89.Andrew Szot, Alex Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra. 2021. Habitat 2.0: Training home assistants to rearrange their habitat. arXiv preprint arXiv:2106.14405.
  90. 90.Hao Tan, Licheng Yu, and Mohit Bansal. 2019. Learning to navigate unseen environments: Back translation with environmental dropout. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2610–2621, Minneapolis, Minnesota. Association for Computational Linguistics.
  91. 91.Sinan Tan, Mengmeng Ge, Di Guo, Huaping Liu, and Fuchun Sun. 2022. Self-supervised 3d semantic representation learning for vision-and-language navigation. arXiv preprint arXiv:2201.10788.
  92. 92.Stefanie Tellex, Thomas Kollar, Steven Dickerson, Matthew Walter, Ashis Banerjee, Seth Teller, and Nicholas Roy. 2011. Understanding natural language commands for robotic navigation and mobile manipulation. In AAAI Conference on Artificial Intelligence.
  93. 93.Jesse Thomason, Daniel Gordon, and Yonatan Bisk. 2019a. Shifting the baseline: Single modality performance on visual navigation & QA. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1977–1983, Minneapolis, Minnesota. Association for Computational Linguistics.
  94. 94.Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer. 2019b. Vision-and-dialog navigation. In Conference on Robot Learning (CoRL).
  95. 95.Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool. 2021. Talk2nav: Long-range vision-and-language navigation with dual attention and spatial memory. International Journal of Computer Vision, 129(1):246–266.
  96. 96.Adam Vogel and Dan Jurafsky. 2010. Learning to follow navigational directions. In Proceedings of the 48th annual meeting of the association for computational linguistics, pages 806–814.
  97. 97.Hanqing Wang, Wenguan Wang, Wei Liang, Caiming Xiong, and Jianbing Shen. 2021. Structured scene memory for vision-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8455–8464.
  98. 98.Hanqing Wang, Wenguan Wang, Tianmin Shu, Wei Liang, and Jianbing Shen. 2020a. Active visual information gathering for vision-language navigation. In European Conference on Computer Vision (ECCV).
  99. 99.Hu Wang, Qi Wu, and Chunhua Shen. 2020b. Soft expert reward learning for vision-and-language navigation. In European Conference on Computer Vision (ECCV’20).
  100. 100.Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Wang, and Lei Zhang. 2019. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In Proceedings of the CVF/IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA. CVF/IEEE.
  101. 101.Xin Wang, Wenhan Xiong, Hongmin Wang, and William Yang Wang. 2018. Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation. Proceedings of the European Conference on Computer Vision (ECCV 2018).
  102. 102.Xin Eric Wang, Vihan Jain, Eugene Ie, William Yang Wang, Zornitsa Kozareva, and Sujith Ravi. 2020c. Environment-agnostic multitask learning for natural language grounded navigation. In European Conference on Computer Vision (ECCV’20).
  103. 103.Erik Wijmans, Samyak Datta, Oleksandr Maksymets, Abhishek Das, Georgia Gkioxari, Stefan Lee, Irfan Essa, Devi Parikh, and Dhruv Batra. 2019a. Embodied Question Answering in Photorealistic Environments with Point Cloud Perception. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  104. 104.Erik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee, Irfan Essa, Devi Parikh, Manolis Savva, and Dhruv Batra. 2019b. Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames. In International Conference on Learning Representations.
  105. 105.Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian. 2018. Building generalizable agents with a realistic and rich 3d environment.
  106. 106.Fei Xia, Amir R. Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. 2018. Gibson env: Real-world perception for embodied agents. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  107. 107.Qiaolin Xia, Xiujun Li, Chunyuan Li, Yonatan Bisk, Zhifang Sui, Jianfeng Gao, Yejin Choi, and Noah A. Smith. 2020. Multi-view learning for vision-and-language navigation.
  108. 108.An Yan, Xin Eric Wang, Jiangtao Feng, Lei Li, and William Yang Wang. 2020. Cross-lingual vision-language navigation.
  109. 109.Jiwen Zhang, Zhongyu Wei, Jianqing Fan, and Jiajie Peng. 2021. Curriculum learning for vision-and-language navigation. In NeurIPS.
  110. 110.Weixia Zhang, Chao Ma, Qi Wu, and Xiaokang Yang. 2020a. Language-guided navigation via cross-modal grounding and alternate adversarial learning. IEEE Transactions on Circuits and Systems for Video Technology.
  111. 111.Yubo Zhang, Hao Tan, and Mohit Bansal. 2020b. Diagnosing the environment bias in vision-and-language navigation. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20, pages 890–897. International Joint Conferences on Artificial Intelligence Organization. Main track.
  112. 112.Ming Zhao, Peter Anderson, Vihan Jain, Su Wang, Alexander Ku, Jason Baldridge, and Eugene Ie. 2021. On the evaluation of vision-and-language navigation instructions. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1302–1316.
  113. 113.Xinzhe Zhou, Wei Liu, and Yadong Mu. 2021. Rethinking the spatial route prior in vision-and-language navigation.
  114. 114.Fengda Zhu, Xiwen Liang, Yi Zhu, Qizhi Yu, Xiaojun Chang, and Xiaodan Liang. 2021a. Soon: Scenario oriented object navigation with graph-based exploration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12689–12699.
  115. 115.Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang. 2020a. Vision-language navigation with self-supervised auxiliary reasoning tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  116. 116.Wang Zhu, Hexiang Hu, Jiacheng Chen, Zhiwei Deng, Vihan Jain, Eugene Ie, and Fei Sha. 2020b. BabyWalk: Going farther in vision-and-language navigation by taking baby steps. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2539–2556. Association for Computational Linguistics.
  117. 117.Wanrong Zhu, Yuankai Qi, Pradyumna Narayana, Kazoo Sone, Sugato Basu, Xin Eric Wang, Qi Wu, Miguel Eckstein, and William Yang Wang. 2021b. Diagnosing vision-and-language navigation: What really matters.
  118. 118.Yi Zhu, Yue Weng, Fengda Zhu, Xiaodan Liang, Qixiang Ye, Yutong Lu, and Jianbin Jiao. 2021c. Self-motivated communication agent for real-world vision-dialog navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1594–1603.
  119. 119.Yi Zhu, Fengda Zhu, Zhaohuan Zhan, Bingqian Lin, Jianbin Jiao, Xiaojun Chang, and Xiaodan Liang. 2020c. Vision-dialog navigation by exploring cross-modal memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10730–10739.
  120. 120.Yuke Zhu, Roozbeh Mottaghi, Eric Kolve, Joseph J Lim, Abhinav Gupta, Li Fei-Fei, and Ali Farhadi. 2017. Target-driven visual navigation in indoor scenes using deep reinforcement learning. In 2017 IEEE international conference on robotics and automation (ICRA), pages 3357–3364. IEEE.

Citation

MLA
Gu, J., et al. “Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7606–23, https://doi.org/10.18653/v1/2022.acl-long.524.
APA
Gu, J., Stefani, E., Wu, Q., Thomason, J., & Wang, X. E. (2022). Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7606–7623. https://doi.org/10.18653/v1/2022.acl-long.524
Chicago
Gu, J., E. Stefani, Q. Wu, J. Thomason, and X. E. Wang. 2022. “Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7606–23. https://doi.org/10.18653/v1/2022.acl-long.524.
Harvard
Gu, J. et al. (2022) “Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7606–7623. Available at: https://doi.org/10.18653/v1/2022.acl-long.524.
Vancouver
1. Gu J, Stefani E, Wu Q, Thomason J, Wang XE (2022) Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7606–7623

BibTeX

@inproceedings{gu-etal-2022-vision,
    title = "Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions",
    author = "Gu, Jing  and
      Stefani, Eliana  and
      Wu, Qi  and
      Thomason, Jesse  and
      Wang, Xin",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.524/",
    doi = "10.18653/v1/2022.acl-long.524",
    pages = "7606--7623"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/