Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation

Yicong HongZun WangQi WuStephen Gould

article2022CVPR137 citations

Proposes a candidate waypoint predictor that enables discrete Vision-and-Language Navigation models to operate directly in continuous environments, achieving state-of-the-art performance on the R2R-CE and RxR-CE benchmarks.

Listen

Vision-and-language navigation tasks require autonomous agents to interpret human language instructions and navigate through unfamiliar environments. While training agents in discrete spaces—where the environment is represented as a pre-mapped graph of locations—has led to significant algorithmic advances, these agents fail to operate in realistic, continuous spaces where an agent must execute low-level movements without a predefined map. Navigating in continuous spaces typically leaves agents struggling to identify obstacles and accessible paths, resulting in a substantial performance drop of approximately 20% compared to discrete navigation.

The article evaluates whether a visual candidate waypoints predictor can bridge this domain gap by generating accessible local destinations on the fly in continuous environments. This approach aims to demonstrate that agents designed for discrete navigation can be directly transferred to continuous settings using high-level directional actions instead of tedious low-level motor controls.

The authors developed a predictor using visual encoders and a spatial attention network trained on refined 3D indoor spatial graphs across 90 simulated building environments. This module predicts obstacle-free local candidate waypoints within a three-meter radius around the agent. The authors evaluated two distinct navigation architectures—a cross-modal matching agent and a visiolinguistic transformer agent—on continuous navigation benchmarks. They also applied a waypoint augmentation technique during training to diversify the paths and visual perspectives the agents experienced.

The key findings show that supplying predicted waypoints significantly improves continuous navigation. Incorporating waypoint prediction reduced the discrete-to-continuous performance gap by 11.76% in path efficiency for the cross-modal matching agent and by 18.24% for the transformer-based agent. Training with waypoint augmentation further enhanced generalization, allowing the simpler agent to reach a 40.80% success rate in unseen test spaces, matching or exceeding graph-based performance. Additionally, the approach reduced obstacle collision rates from 15% to 7% and established new state-of-the-art benchmarks on standard continuous navigation datasets (R2R-CE and RxR-CE), achieving a 38% to 42% success rate on unseen test paths.

These results demonstrate that the core challenge of continuous navigation is identifying navigable space rather than executing complex motor actions. Decoupling waypoint prediction from instruction following enables standard imitation learning to succeed without complex reinforcement learning pipelines. This dramatically lowers compute requirements, reducing training overhead from 64 graphics processors over five days to a single processor in roughly three and a half days while achieving superior navigation accuracy.

Teams developing embodied artificial intelligence and autonomous service robots should adopt local waypoint prediction modules to transition high-level planning models into continuous real-world environments. Future development should focus on extending waypoint predictors to support language-conditioned navigation and testing the framework across broader embodied tasks, such as object-goal and audio-visual navigation.

A current limitation is that the predictor may occasionally fail to identify valid pathways across challenging architectural features, such as stairs, leading to failed trajectories. Nevertheless, confidence in the primary findings remains high, as the methodology demonstrated strong, consistent improvements across multiple model architectures and standardized evaluation splits.

arXiv: 2203.02764
Cover for Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation

Abstract

Most existing works in vision-and-language navigation (VLN) focus on either discrete or continuous environments, training agents that cannot generalize across the two. Although learning to navigate in continuous spaces is closer to the real-world, training such an agent is significantly more difficult than training an agent in discrete spaces. However, recent advances in discrete VLN are challenging to translate to continuous VLN due to the domain gap. The fundamental difference between the two setups is that discrete navigation assumes prior knowledge of the connectivity graph of the environment, so that the agent can effectively transfer the problem of navigation with low-level controls to jumping from node to node with high-level actions by grounding to an image of a navigable direction.

To bridge the discrete-to-continuous gap, we propose a predictor to generate a set of candidate waypoints during navigation, so that agents designed with high-level actions can be transferred to and trained in continuous environments. We refine the connectivity graph of Matterport3D to fit the continuous Habitat-Matterport3D, and train the waypoints predictor with the refined graphs to produce accessible waypoints at each time step. Moreover, we demonstrate that the predicted waypoints can be augmented during training to diversify the views and paths, and therefore enhance agent's generalization ability.

Through extensive experiments we show that agents navigating in continuous environments with predicted waypoints perform significantly better than agents using low-level actions, which reduces the absolute discrete-to-continuous gap by 11.76% Success Weighted by Path Length (SPL) for the Cross-Modal Matching Agent and 18.24% SPL for the VLN○BERT. Our agents, trained with a simple imitation learning objective, outperform previous methods by a large margin, achieving new state-of-the-art results on the testing environments of the R2R-CE and RxR-CE datasets.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background
  • 3.1. Navigation Setups
  • 3.2. What is the Value of High-Level Actions?
  • 4. Candidate Waypoints Predictor
  • 4.1. Network Architecture and Processing
  • 4.2. Connectivity Graphs in Habitat-MP3D
  • 4.3. Confirmation of Predictor Performance
  • 5. Bridging the Discrete to Continuous Gap
  • 5.1. Setups
  • 5.2. Main Results
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Candidate Waypoints Predictor Architecture and Formulation

    model/method

    The candidate waypoints predictor estimates accessible local positions in continuous 3D environments from egocentric visual inputs to support high-level action selection without a pre-defined connectivity graph.

    At any agent position ptp_t, the environment is observed via an RGB panorama and a depth panorama, each discretized into 12 single-view images spaced by 30∘30^\circ with a 90∘90^\circ horizontal field of view. Visual features are extracted using two fixed ResNet-50 backbones: one pre-trained on ImageNet for RGB images yielding {v1rgb,…,v12rgb}\{v_1^{\text{rgb}}, \dots, v_{12}^{\text{rgb}}\}, and one pre-trained on PointGoal navigation for depth images yielding {v1d,…,v12d}\{v_1^{\text{d}}, \dots, v_{12}^{\text{d}}\}. For each direction i∈{1,…,12}i \in \{1, \dots, 12\}, the RGB and depth features are merged via a non-linear layer WmW_m:

    virgbd=Wm([virgb;vid])v_i^{\text{rgbd}} = W_m([v_i^{\text{rgb}}; v_i^{\text{d}}])

    The 12 merged representations are processed by a 2-layer Transformer with 12 self-attention heads. Self-attention for each token virgbdv_i^{\text{rgbd}} is constrained locally to its adjacent directional neighbors, attending only to vi−1rgbdv_{i-1}^{\text{rgbd}}, virgbdv_i^{\text{rgbd}}, and vi+1rgbdv_{i+1}^{\text{rgbd}}.

    The resulting output tokens v~irgbd\tilde{v}_i^{\text{rgbd}} are passed to a multi-layer perceptron classifier that predicts a polar waypoint probability heatmap of dimension 120×12120 \times 12, spanning 120 angular sectors (3∘3^\circ per sector over 360∘360^\circ) and 12 radial distance bins (0.25 m0.25\text{ m} to 3.00 m3.00\text{ m} in 0.25 m0.25\text{ m} steps). The network is trained by minimizing the mean squared error (MSE) against a ground-truth Gaussian heatmap. During inference, non-maximum suppression (NMS) over the predicted heatmap yields up to K=5K=5 candidate waypoint coordinates.

  2. Knowl 2 — Refined Ground-Truth Connectivity Graphs in Habitat-Matterport3D

    model/method

    To train the candidate waypoints predictor, the discrete connectivity graph GMP3DG^{\text{MP3D}} from the Matterport3D dataset is adapted into a continuous, collision-free graph G∗G^* in Habitat-Matterport3D (Habitat-MP3D).

    In GMP3DG^{\text{MP3D}}, graph edges frequently intersect 3D obstacle geometry and some nodes reside in inaccessible locations, making them unsuitable for learning navigability in continuous simulators. The refined graph G∗G^* constrains all nodes and edges to lie strictly within obstacle-free open space, inserting intermediate nodes where necessary to maintain global connectivity.

    Across 90 indoor environments, G∗G^* contains 13,358 nodes with an average node degree of 3.31 edges and an average edge length of 1.87 m1.87\text{ m}, compared to 10,559 nodes, average degree 4.07, and average edge length 2.26 m2.26\text{ m} in GMP3DG^{\text{MP3D}}.

    For each node g∈G∗g \in G^*, a local ground-truth polar heatmap of dimension 120×12120 \times 12 is constructed over a 3 m3\text{ m} radius. Each neighbor node is mapped to its corresponding angle-distance bin and represented as a 2D Gaussian distribution with radial standard deviation σr=1.75 m\sigma_r = 1.75\text{ m} and angular standard deviation σθ=15∘\sigma_\theta = 15^\circ to provide target tolerance.

  3. Knowl 3 — Waypoint Augmentation During Agent Policy Training

    model/method

    Waypoint augmentation is a data augmentation method applied when training navigation agents with predicted waypoints in continuous environments.

    During training with imitation learning, when the navigation agent selects a candidate high-level direction from the candidate waypoints predictor, the exact sub-goal location is randomly sampled from the local heatmap patch corresponding to that chosen view rather than deterministically selecting the NMS peak.

    This sampling perturbation shifts the target waypoint in continuous space, requiring the agent to traverse varying trajectory lengths, encounter diverse camera viewpoints, and navigate around different obstacle configurations to reach the sub-goal, thereby improving out-of-distribution generalization to unseen environments.

  4. Knowl 4 — Hierarchical Execution of High-Level Actions in Continuous VLN

    model/method

    Vision-and-language navigation in continuous environments (VLN-CE) using high-level actions is structured as a hierarchical process decoupled between high-level waypoint selection and low-level actuation.

    At step tt, the candidate waypoints predictor generates KK candidate reachable waypoints around the current agent pose ptp_t. For each candidate k∈{1,…,K}k \in \{1, \dots, K\}, a visual-directional feature vector fkp=[vkp;dkp]f_k^p = [v_k^p; d_k^p] is formed by concatenating visual feature vkpv_k^p pointing toward the candidate and relative directional encoding dkpd_k^p.

    A high-level discrete policy network (e.g., Cross-Modal Matching Agent or Recurrent VLN-BERT) updates its recurrent hidden state hth_t with attended language instruction features ctlangc_t^{\text{lang}} and prior action at−1a_{t-1}:

    ht=LSTM([ctlang;at−1],ht−1)h_t = \text{LSTM}([c_t^{\text{lang}}; a_{t-1}], h_{t-1})

    The probability of selecting candidate waypoint ii is determined via cross-modal attention:

    pi=Softmax([ht;ctlang;ctrgb]⊤Wpfip)p_i = \text{Softmax}\left([h_t; c_t^{\text{lang}}; c_t^{\text{rgb}}]^\top W_p f_i^p\right)

    where ctrgbc_t^{\text{rgb}} is the attended visual context and WpW_p is a learnable projection matrix. Once candidate ii is selected, its polar displacement relative to the agent is decomposed into a sequence of low-level primitive actions (turn left 15∘15^\circ, turn right 15∘15^\circ, move forward 0.25 m0.25\text{ m}, or stop) executed progressively in the continuous simulator until the agent reaches the sub-goal.

  5. Knowl 5 — Dissecting the Value of Navigable View Selection vs. Waypoint Teleportation

    empirical result

    A diagnostic study evaluates the independent contributions of two core mechanisms inherent to discrete graph navigation: navigable view selection (identifying unobstructed directions) and waypoint teleportation (skipping low-level step execution). Tests were conducted using the Reinforced Cross-Modal Matching (RCM / CMA) agent in Habitat-MP3D on R2R-CE validation splits.

    Action R2R-CE Val-Seen R2R-CE Val-Unseen
    Select Teleport NE ↓\downarrow nDTW ↑\uparrow SR ↑\uparrow SPL ↑\uparrow NE ↓\downarrow nDTW ↑\uparrow SR ↑\uparrow SPL ↑\uparrow Time (s) ↓\downarrow
    ✓ ✓ 6.28 58.37 41.48 36.58 6.51 55.41 39.32 33.89 0.55
    ✓ ×\times 5.70 61.06 46.31 43.20 6.38 56.27 37.72 35.02 1.79
    ×\times ✓ 8.10 43.19 23.36 19.85 7.62 47.69 28.32 24.64 0.78
    ×\times ×\times 7.20 50.67 29.53 27.29 7.54 49.19 27.29 24.97 2.61
    Select Dist (m) R2R-CE Val-Seen R2R-CE Val-Unseen
    NE ↓\downarrow nDTW ↑\uparrow SR ↑\uparrow SPL ↑\uparrow NE ↓\downarrow nDTW ↑\uparrow SR ↑\uparrow SPL ↑\uparrow Time (s) ↓\downarrow
    ×\times 0.25 7.20 50.67 29.53 27.29 7.54 49.19 27.29 24.97 2.61
    ×\times 1.00 7.21 52.52 29.66 27.19 7.51 50.14 25.47 23.51 1.28
    ×\times 2.00 7.60 47.90 24.16 22.25 8.06 45.74 22.96 20.59 1.18
    ×\times 3.00 7.66 48.28 23.62 21.73 7.87 46.29 21.37 19.39 0.90
    ✓ 0.25 6.10 58.24 39.33 38.26 6.52 55.26 32.02 31.15 1.58
    ✓ 1.00 6.88 53.09 36.11 33.91 6.85 52.44 34.81 32.51 0.46
    ✓ 2.00 6.56 52.83 38.52 35.62 6.96 50.25 33.05 30.35 0.27
    ✓ 3.00 6.75 51.85 32.21 29.28 7.00 49.35 31.00 28.12 0.20

    Navigable view selection accounts for the majority of the performance advantage: providing ground-truth navigable directions without teleportation achieves 37.72%37.72\% SR and 35.02%35.02\% SPL on Val-Unseen, comparable to full graph navigation (39.32%39.32\% SR, 33.89%33.89\% SPL). Increasing forward distance per step reduces decision frequency and inference time per episode (from 2.61 s2.61\text{ s} to 0.20 s0.20\text{ s}), but view selection remains the primary factor determining navigation accuracy.

  6. Knowl 6 — Comparative Evaluation of Candidate Waypoint Predictors

    empirical result

    The candidate waypoints predictor was evaluated against alternative architectures on the Matterport3D (MP3D) Train and Val-Unseen splits using four metrics: ∣Δ∣|\Delta| (absolute count error between predicted and target waypoints), %Open (percentage of predicted waypoints situated in collision-free open space), dCd_C (Chamfer distance in meters), and dHd_H (Hausdorff distance in meters).

    Model MP3D Train MP3D Val-Unseen
    ∣Δ∣|\Delta| %Open ↑\uparrow dCd_C (m) ↓\downarrow dHd_H (m) ↓\downarrow ∣Δ∣|\Delta| %Open ↑\uparrow dCd_C (m) ↓\downarrow dHd_H (m) ↓\downarrow
    Baseline (No Transformer) 1.29 81.73 1.13 2.17 1.37 80.18 1.08 2.06
    U-Net 1.15 63.60 1.05 2.10 1.21 52.54 1.01 2.00
    Ours (Transformer) 1.30 82.56 1.12 2.13 1.40 79.86 1.07 2.00

    The convolutional U-Net baseline achieves lower point distance metrics (dC=1.01 md_C = 1.01\text{ m}, dH=2.00 md_H = 2.00\text{ m}) but outputs nearly half of its predictions into obstacles on unseen scenes (52.54%52.54\% open space). The Transformer predictor maintains high obstacle-free prediction accuracy (79.86%79.86\% on Val-Unseen) while achieving point cloud distances superior to the non-Transformer baseline (dC=1.07 md_C = 1.07\text{ m}, dH=2.00 md_H = 2.00\text{ m}).

  7. Knowl 7 — Bridging the Discrete-to-Continuous Navigation Performance Gap on R2R-CE

    empirical result

    The impact of candidate waypoint prediction and waypoint augmentation on bridging the performance gap between discrete and continuous VLN was evaluated on R2R-CE using Cross-Modal Matching Agent (CMA) and Recurrent VLN-BERT (VLN↻\circlearrowrightBERT).

    Method Connectivity R2R-CE Val-Seen R2R-CE Val-Unseen
    Train Val NE ↓\downarrow nDTW ↑\uparrow SR ↑\uparrow SPL ↑\uparrow NE ↓\downarrow nDTW ↑\uparrow SR ↑\uparrow SPL ↑\uparrow
    CMA (Graph) Graph Graph 6.28 58.37 41.48 36.58 6.51 55.41 39.32 33.89
    CMA (Freeze Transfer) Graph Predictor 6.83 53.16 33.83 30.23 6.81 52.86 33.50 29.92
    CMA (Predictor Trained) Predictor Predictor 5.81 59.06 43.89 40.20 6.50 54.34 37.49 33.90
    CMA (Augmented) Augmented Predictor 5.36 61.46 48.05 43.89 6.19 55.56 40.80 36.73
    CMA (Low-Level Actions) Low-Level Low-Level 7.20 50.67 29.53 27.29 7.54 49.19 27.29 24.97
    VLN↻\circlearrowrightBERT (Graph) Graph Graph 4.63 61.46 52.21 42.53 5.22 57.71 48.89 40.36
    VLN↻\circlearrowrightBERT (Freeze Transfer) Graph Predictor 5.62 56.19 43.22 37.10 5.91 53.35 39.94 34.42
    VLN↻\circlearrowrightBERT (Predictor Trained) Predictor Predictor 5.06 59.46 49.40 43.43 5.68 53.50 42.96 38.53
    VLN↻\circlearrowrightBERT (Augmented) Augmented Predictor 5.07 59.18 52.35 46.08 5.52 54.20 45.19 39.91
    VLN↻\circlearrowrightBERT (Low-Level Actions) Low-Level Low-Level 7.40 48.07 26.85 25.19 7.66 47.71 23.19 21.67

    Compared to low-level control baselines, training with augmented predicted waypoints increases Val-Unseen SPL by 11.76%11.76\% for CMA (24.97%→36.73%24.97\% \to 36.73\%) and by 18.24%18.24\% for VLN↻\circlearrowrightBERT (21.67%→39.91%21.67\% \to 39.91\%), recovering 90%90\% of the discrete-to-continuous gap for VLN↻\circlearrowrightBERT and exceeding the ground-truth graph performance for CMA (36.73%36.73\% vs. 33.89%33.89\%).

  8. Knowl 8 — Benchmark Results on R2R-CE and RxR-CE Test Sets

    empirical result

    Continuous VLN agents trained using the candidate waypoints predictor and simple imitation learning were benchmarked on the test sets of R2R-CE and RxR-CE against prior state-of-the-art continuous navigation methods.

    Methods R2R-CE Val-Seen R2R-CE Val-Unseen R2R-CE Test-Unseen
    NE ↓\downarrow nDTW ↑\uparrow OSR ↑\uparrow SR ↑\uparrow SPL ↑\uparrow NE ↓\downarrow nDTW ↑\uparrow OSR ↑\uparrow SR ↑\uparrow SPL ↑\uparrow NE ↓\downarrow OSR ↑\uparrow SR ↑\uparrow SPL ↑\uparrow
    VLN-CE 7.12 54 46 37 35 7.37 51 40 32 30 7.91 36 28 25
    LAW – 58 – 40 37 – 54 – 35 31 – – – –
    SASRA 7.17 53 – 36 34 8.32 47 – 24 22 – – – –
    Waypoint Models 5.48 – 53 46 43 6.31 – 40 36 34 6.65 37 32 30
    Ours (CMA) 5.20 61 61 51 45 6.20 55 52 41 36 6.30 49 38 33
    Ours (VLN↻\circlearrowrightBERT) 5.02 58 59 50 44 5.74 54 53 44 39 5.89 51 42 36
    Methods RxR-CE Test-Unseen
    TL NE ↓\downarrow SR ↑\uparrow SPL ↑\uparrow nDTW ↑\uparrow SDTW ↑\uparrow
    VLN-CE 7.33 12.1 13.93 11.96 30.86 11.01
    Ours (CMA) 20.04 10.4 24.08 19.07 37.39 18.65
    Ours (VLN↻\circlearrowrightBERT) 20.09 10.4 24.85 19.61 37.30 19.05

    On R2R-CE Test-Unseen, the proposed model sets a new state-of-the-art with 42%42\% SR and 36%36\% SPL (compared to 32%32\% SR and 30%30\% SPL for Waypoint Models). On the multilingual RxR-CE Test-Unseen set, the model achieves 24.85%24.85\% SR and 19.61%19.61\% SPL (compared to 13.93%13.93\% SR and 11.96%11.96\% SPL for baseline VLN-CE). The model achieves these results via cross-entropy imitation learning trained on a single GPU in 3.5 days, compared to reinforcement learning methods requiring 64 GPUs over 5 days.

  9. Knowl 9 — Limitations of Language-Agnostic Candidate Waypoint Prediction

    limitation

    The candidate waypoints predictor is decoupled from the language instruction and agent trajectory history, predicting candidate waypoints purely from local sensory navigability.

    In environments requiring specialized navigation actions—such as locating and traversing staircases between floor levels—the predictor may fail to generate valid candidate waypoints in the target direction, leading to unrecoverable navigation failures. Navigation performance remains upper-bounded by the spatial accuracy and candidate recall of the local navigability predictor.

Coverage note — None was omitted; all key contributed methods (waypoint predictor architecture, refined Habitat-MP3D graph construction, waypoint augmentation), diagnostic analyses, gap-bridging ablations, test benchmark results, and limitations are fully covered.

References

  1. 1.Peter Anderson, Angel Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, et al. On evaluation of embodied navigation agents. arXiv preprint arXiv:1807.06757, 2018. 4, 7
  2. 2.Peter Anderson, Ayush Shrivastava, Joanne Truong, Arjun Majumdar, Devi Parikh, Dhruv Batra, and Stefan Lee. Sim-to-real transfer for vision-and-language navigation. In Conference on Robot Learning, pages 671–681. PMLR, 2021. 2, 3, 6, 8
  3. 3.Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3674–3683, 2018. 2, 3, 4, 6, 7
  4. 4.Dhruv Batra, Aaron Gokaslan, Aniruddha Kembhavi, Oleksandr Maksymets, Roozbeh Mottaghi, Manolis Savva, Alexander Toshev, and Erik Wijmans. Objectnav revisited: On evaluation of embodied agents navigating to objects. arXiv preprint arXiv:2006.13171, 2020. 8
  5. 5.Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. Scheduled sampling for sequence prediction with recurrent neural networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems-Volume 1, pages 1171–1179, 2015. 7
  6. 6.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niebner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. In 2017 International Conference on 3D Vision (3DV), pages 667–676. IEEE, 2017. 2, 3
  7. 7.Devendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta, and Ruslan Salakhutdinov. Learning to explore using active neural slam. In International Conference on Learning Representations, 2019. 3
  8. 8.Devendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta, and Saurabh Gupta. Neural topological slam for visual navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12875–12884, 2020. 3
  9. 9.Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman. Soundspaces: Audio-visual navigation in 3d environments. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16, pages 17–36. Springer, 2020. 8
  10. 10.Changan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao, Santhosh Kumar Ramakrishnan, and Kristen Grauman. Learning to set waypoints for audio-visual navigation. In International Conference on Learning Representations, 2020. 3
  11. 11.Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi. Touchdown: Natural language navigation and spatial reasoning in visual street environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 12538–12547, 2019. 2
  12. 12.Kevin Chen, Junshen K Chen, Jo Chuang, Marynel Vázquez, and Silvio Savarese. Topological planning with transformers for vision-and-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11276–11286, 2021. 3
  13. 13.Kevin Chen, Juan Pablo de Vicente, Gabriel Sepulveda, Fei Xia, Alvaro Soto, Marynel Vázquez, and Silvio Savarese. A behavioral approach to visual navigation with graph localization networks. arXiv preprint arXiv:1903.00445, 2019. 3
  14. 14.Tao Chen, Saurabh Gupta, and Abhinav Gupta. Learning exploration policies for navigation. In International Conference on Learning Representations, 2018. 3
  15. 15.Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra. Embodied question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1–10, 2018. 5
  16. 16.Zhiwei Deng, Karthik Narasimhan, and Olga Russakovsky. Evolving graphical planner: Contextual global planning for vision-and-language navigation. arXiv preprint arXiv:2007.05655, 2020. 2
  17. 17.Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. Speaker-follower models for vision-and-language navigation. In Advances in Neural Information Processing Systems, pages 3314–3325, 2018. 2, 3, 4
  18. 18.Tsu-Jui Fu, Xin Eric Wang, Matthew F Peterson, Scott T Grafton, Miguel P Eckstein, and William Yang Wang. Counterfactual vision-and-language navigation via adversarial path sampler. In European Conference on Computer Vision, pages 71–86. Springer, 2020. 2
  19. 19.Saurabh Gupta, James Davidson, Sergey Levine, Rahul Sukthankar, and Jitendra Malik. Cognitive mapping and planning for visual navigation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2616–2625, 2017. 3
  20. 20.Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao. Towards learning a generic agent for vision-and-language navigation via pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13137–13146, 2020. 2, 3, 7
  21. 21.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 4, 5
  22. 22.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997. 3
  23. 23.Yicong Hong, Cristian Rodriguez, Yuankai Qi, Qi Wu, and Stephen Gould. Language and visual entity relationship graph for agent navigation. Advances in Neural Information Processing Systems, 33, 2020. 2, 3
  24. 24.Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez-Opazo, and Stephen Gould. A recurrent vision-and-language bert for navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1643–1653, June 2021. 2, 3, 6, 7
  25. 25.Ronghang Hu, Daniel Fried, Anna Rohrbach, Dan Klein, Trevor Darrell, and Kate Saenko. Are you looking? grounding to multiple modalities in vision-and-language navigation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6551–6557, 2019. 3
  26. 26.Gabriel Ilharco, Vihan Jain, Alexander Ku, Eugene Ie, and Jason Baldridge. General evaluation for instruction conditioned navigation using dynamic time warping. arXiv preprint arXiv:1907.05446, 2019. 4, 7
  27. 27.Muhammad Zubair Irshad, Chih-Yao Ma, and Zsolt Kira. Hierarchical cross-modal agent for robotics vision-and-language navigation. arXiv preprint arXiv:2104.10674, 2021. 3
  28. 28.Muhammad Zubair Irshad, Niluthpol Chowdhury Mithun, Zachary Seymour, Han-Pang Chiu, Supun Samarasekera, and Rakesh Kumar. Sasra: Semantically-aware spatio-temporal reasoning agent for vision-and-language navigation in continuous environments. arXiv preprint arXiv:2108.11945, 2021. 3, 8
  29. 29.Liyiming Ke, Xiujun Li, Yonatan Bisk, Ari Holtzman, Zhe Gan, Jingjing Liu, Jianfeng Gao, Yejin Choi, and Siddhartha Srinivasa. Tactical rewind: Self-correction via backtracking in vision-and-language navigation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6741–6749, 2019. 2
  30. 30.Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi. Ai2-thor: An interactive 3d environment for visual ai. arXiv preprint arXiv:1712.05474, 2017. 3
  31. 31.Jacob Krantz, Aaron Gokaslan, Dhruv Batra, Stefan Lee, and Oleksandr Maksymets. Waypoint models for instruction-guided navigation in continuous environments. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15162–15171, 2021. 3, 8
  32. 32.Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee. Beyond the nav-graph: Vision-and-language navigation in continuous environments. In European Conference on Computer Vision, 2020. 2, 3, 4, 6, 7, 8
  33. 33.Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4392–4412, 2020. 2, 3, 6, 7
  34. 34.Xiujun Li, Chunyuan Li, Qiaolin Xia, Yonatan Bisk, Asli Celikyilmaz, Jianfeng Gao, Noah A Smith, and Yejin Choi. Robust navigation with language pretraining and stochastic sampling. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1494–1499, 2019. 3
  35. 35.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 6, 7
  36. 36.Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, Zsolt Kira, Richard Socher, and Caiming Xiong. Self-monitoring navigation agent via auxiliary progress estimation. In Proceedings of the International Conference on Learning Representations (ICLR), 2019. 2
  37. 37.Chih-Yao Ma, Zuxuan Wu, Ghassan AlRegib, Caiming Xiong, and Zsolt Kira. The regretful agent: Heuristic-aided navigation through progress estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6732–6740, 2019. 2
  38. 38.Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra. Improving vision-and-language navigation with image-text pairs from the web. In Proceedings of the European Conference on Computer Vision, 2020. 2, 3
  39. 39.Xiangyun Meng, Nathan Ratliff, Yu Xiang, and Dieter Fox. Scaling local control to large-scale topological navigation. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 672–678. IEEE, 2020. 3
  40. 40.Khanh Nguyen and Hal Daumé III. Help, anna! visual navigation with natural multimodal assistance via retrospective curiosity-encouraging imitation learning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 684–695, 2019. 2
  41. 41.Amin Parvaneh, Ehsan Abbasnejad, Damien Teney, Qinfeng Shi, and Anton van den Hengel. Counterfactual vision-and-language navigation: Unravelling the unseen. Advances in Neural Information Processing Systems, 33, 2020. 2
  42. 42.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32:8026–8037, 2019. 7
  43. 43.Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. Reverie: Remote embodied visual referring expression in real indoor environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9982–9991, 2020. 2
  44. 44.Sonia Raychaudhuri, Saim Wani, Shivansh Patel, Unnat Jain, and Angel Chang. Language-aligned waypoint (law) supervision for vision-and-language navigation in continuous environments. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4018–4028, 2021. 3, 8
  45. 45.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015. 6
  46. 46.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015. 5
  47. 47.Nikolay Savinov, Alexey Dosovitskiy, and Vladlen Koltun. Semi-parametric topological memory for navigation. In International Conference on Learning Representations, 2018. 3
  48. 48.Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al. Habitat: A platform for embodied ai research. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9339–9347, 2019. 2, 3, 4, 7, 8
  49. 49.Bokui Shen, Fei Xia, Chengshu Li, Roberto Martín-Martín, Linxi Fan, Guanzhi Wang, Shyamal Buch, Claudia D’Arpino, Sanjana Srivastava, Lyne P Tchapmi, et al. igibson, a simulation environment for interactive tasks in large realisticscenes. arXiv preprint arXiv:2012.02924, 2020. 3
  50. 50.Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10740–10749, 2020. 8
  51. 51.Hao Tan, Licheng Yu, and Mohit Bansal. Learning to navigate unseen environments: Back translation with environmental dropout. In Proceedings of NAACL-HLT, pages 2610–2621, 2019. 2
  52. 52.Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer. Vision-and-dialog navigation. In Conference on Robot Learning, pages 394–406, 2020. 2
  53. 53.Hanqing Wang, Wenguan Wang, Wei Liang, Caiming Xiong, and Jianbing Shen. Structured scene memory for vision-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8455–8464, 2021. 2
  54. 54.Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6629–6638, 2019. 2, 3, 4, 7
  55. 55.Erik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee, Irfan Essa, Devi Parikh, Manolis Savva, and Dhruv Batra. Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames. In International Conference on Learning Representations, 2019. 5, 8
  56. 56.Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian. Building generalizable agents with a realistic and rich 3d environment. arXiv preprint arXiv:1801.02209, 2018. 3
  57. 57.Fei Xia, Amir R Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese. Gibson env: Real-world perception for embodied agents. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 9068–9079, 2018. 3
  58. 58.Claudia Yan, Dipendra Misra, Andrew Bennnett, Aaron Walsman, Yonatan Bisk, and Yoav Artzi. Chalet: Cornell house agent learning environment. arXiv preprint arXiv:1801.07357, 2018. 3
  59. 59.Fengda Zhu, Xiwen Liang, Yi Zhu, Qizhi Yu, Xiaojun Chang, and Xiaodan Liang. Soon: Scenario oriented object navigation with graph-based exploration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12689–12699, 2021. 2

Citation

MLA
Hong, Y., et al. “Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation”. arXiv, 2022, http://arxiv.org/abs/2203.02764v1.
APA
Hong, Y., Wang, Z., Wu, Q., & Gould, S. (2022). Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation. arXiv. http://arxiv.org/abs/2203.02764v1
Chicago
Hong, Y., Z. Wang, Q. Wu, and S. Gould. 2022. “Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation”. arXiv. http://arxiv.org/abs/2203.02764v1.
Harvard
Hong, Y. et al. (2022) “Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.02764v1.
Vancouver
1. Hong Y, Wang Z, Wu Q, Gould S (2022) Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation. arXiv

BibTeX

@article{hong2022bridging,
  title = {Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation},
  author = {Hong, Yicong and Wang, Zun and Wu, Qi and Gould, Stephen},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.02764v1},
  eprint = {2203.02764}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE