Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation

Shizhe ChenPierre-Louis GuhurMakarand TapaswiCordelia SchmidIvan Laptev

article2022CVPR299 citations

Introduces DUET, a dual-scale graph transformer architecture that dynamically fuses coarse topological maps with fine-grained local visual observations to solve long-term action planning and instruction-following challenges in embodied AI.

Listen

Autonomous navigation guided by natural language is a foundational capability for mobile robots and virtual assistants operating in homes and workplaces. However, current systems struggle in previously unseen environments when users give high-level, goal-oriented commands such as finding and interacting with a remote object. Existing methods typically restrict an agent to making myopic step-by-step local moves or rely on condensed memory formats, which makes exploring large spaces inefficient, complicates backtracking, and lacks the fine-grained visual detail required to identify specific target objects.

The article develops and evaluates a new navigation framework called the Dual-scale graph Transformer (DUET). The primary objective is to demonstrate that dynamically combining coarse-scale global spatial reasoning with fine-scale local visual grounding substantially improves an autonomous agent's ability to plan long-term routes and locate target objects in unfamiliar environments.

To accomplish this, the authors construct an online topological map that tracks visited and unvisited navigable locations as the agent moves. DUET uses a dual-scale architecture powered by graph transformers: a coarse-scale encoder reasons over the global map and its structural connectivity to select long-range navigation targets, while a fine-scale encoder processes detailed panoramic and object features at the current location to evaluate immediate actions and pinpoint target objects. The system dynamically fuses predictions from both scales. Training combines pretraining on auxiliary multimodal tasks with policy learning guided by an interactive pseudo-demonstrator to correct errors during simulated exploration. The approach was evaluated on three established benchmarks: REVERIE and SOON for goal-oriented navigation, and R2R for detailed step-by-step navigation.

DUET achieves substantial performance gains across all benchmarks. On the unseen test split of the REVERIE benchmark, DUET increased the navigation success rate by 22.11 percentage points over the prior state of the art, rising from 30.40% to 52.51%, while improving target object grounding success penalized by path length from 13.08% to 22.06%. On the SOON benchmark, DUET improved the unseen test success rate from 12.90% to 33.44%, representing a gain of over 20 percentage points. On the step-by-step R2R dataset, DUET established a new state of the art by lifting the unseen test success rate from 65% to 69%. Ablation experiments confirmed that both the coarse-scale global map reasoning and fine-scale local object representations are essential, and that incorporating graph topology into self-attention directly improves navigation path efficiency.

These findings indicate that autonomous agents can overcome exploration bottlenecks and costly backtracking by decoupling long-term route planning from immediate visual grounding. For operational applications, this capability reduces navigation failures, shortens execution timelines in complex layouts, and enhances human-robot interaction by allowing agents to follow natural, high-level commands without requiring tedious step-by-step guidance.

Organizations developing embodied artificial intelligence and autonomous mobile systems should consider adopting dual-scale topological architectures and interactive demonstrator training strategies. Next steps include evaluating DUET in continuous, non-discrete physical environments and validating its performance on physical robotic platforms operating under real-world sensor noise and dynamic obstacles.

Readers should note that the evaluation is conducted within discrete, graph-based simulation environments with access to reliable orientation and location coordinates. While confidence in the benchmark improvements is very high due to consistent gains across multiple standardized datasets, stakeholders should exercise caution when translating these results directly to real-world hardware where mapping errors, camera blur, and moving obstacles may affect performance.

arXiv: 2202.11742
Cover for Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation

Abstract

Following language instructions to navigate in unseen environments is a challenging problem for autonomous embodied agents. The agent not only needs to ground languages in visual scenes, but also should explore the environment to reach its target. In this work, we propose a dual-scale graph transformer (DUET) for joint long-term action planning and fine-grained cross-modal understanding. We build a topological map on-the-fly to enable efficient exploration in global action space. To balance the complexity of large action space reasoning and fine-grained language grounding, we dynamically combine a fine-scale encoding over local observations and a coarse-scale encoding on a global map via graph transformers. The proposed approach, DUET, significantly outperforms state-of-the-art methods on goal-oriented vision-and-language navigation (VLN) benchmarks REVERIE and SOON. It also improves the success rate on the fine-grained VLN benchmark R2R.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Method
  • 3.1. Topological Mapping
  • 3.2. Global Action Planning
  • 3.2.1 Text Encoder
  • 3.2.2 Coarse-scale Cross-modal Encoder
  • 3.2.3 Fine-scale Cross-modal Encoder
  • 3.2.4 Dynamic Fusion
  • 3.3. Training and Inference
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Evaluation Metrics
  • 4.3. Implementation Details
  • 4.4. Ablation Study
  • 4.5. Comparison with State of the Art
  • References

Knowls

  1. Knowl 1 — Dual-Scale Graph Transformer Architecture for Vision-and-Language Navigation

    model/method

    The Dual-scale Graph Transformer (DUET) is an embodied navigation framework designed for discrete graph environments. It addresses the trade-off between global exploration across unvisited regions and fine-grained visual-linguistic grounding of local objects and scenes by operating across two distinct spatial scales.

    DUET consists of four primary modules:

    1. Text Encoder: A 9-layer transformer that encodes a sequence of instruction word tokens W={w1,…,wL}W = \{w_1, \dots, w_L\} combined with positional and text-type embeddings into contextual word embeddings W^={w^1,…,w^L}\hat{W} = \{\hat{w}_1, \dots, \hat{w}_L\} of dimension d=768d = 768.
    2. Panorama Encoder: A 2-layer transformer using self-attention to model spatial relations among panoramic visual patch tokens Rt={ri}i=1nR_t = \{r_i\}_{i=1}^n (extracted using ViT-B/16) and candidate object bounding box tokens Ot={oi}i=1mO_t = \{o_i\}_{i=1}^m extracted at current node VtV_t.
    3. Coarse-scale Cross-modal Encoder: A 4-layer graph-aware transformer that takes pooled node embeddings from an online topological map Gt=(Vt,Et)G_t = (V_t, E_t) alongside navigation step and relative location encodings, performing cross-attention over W^\hat{W} and graph-aware self-attention across nodes to predict navigation probabilities across all candidate locations in the global action space.
    4. Fine-scale Cross-modal Encoder: A 4-layer transformer that performs fine-grained cross-modal reasoning over unpooled visual/object tokens [r0;Rt;Ot][r_0; R_t; O_t] at the current node VtV_t, predicting local step actions in At={stop}∪N(Vt)\mathcal{A}_t = \{\text{stop}\} \cup \mathcal{N}(V_t) and scoring objects for final grounding.

    A dynamic fusion module predicts a data-dependent blending weight to combine action predictions from both scales into a unified global action score.

  2. Knowl 2 — Online Topological Map Construction and Node Feature Aggregation

    model/method

    In discrete vision-and-language navigation, the global environment connectivity graph G=(V,E)G = (V, E) is unknown a priori. DUET incrementally maintains an online topological map Gt=(Vt,Et)G_t = (V_t, E_t) with KtK_t nodes after tt navigation steps.

    At any step tt, the node set VtV_t contains three categories of nodes:

    1. Visited nodes: Locations previously traversed by the agent where complete panoramas were observed.
    2. Current node VtV_t: The current location of the agent, providing visual features Rt={ri}i=1nR_t = \{r_i\}_{i=1}^n and object features Ot={oi}i=1mO_t = \{o_i\}_{i=1}^m.
    3. Navigable nodes: Unvisited neighboring locations N(Vt)\mathcal{N}(V_t) that can be directly reached from visited nodes.

    Node features are constructed and updated through the following mechanisms:

    • Current node representation: Panoramic patch features RtR_t and object tokens OtO_t are processed via spatial self-attention: [Rt′,Ot′]=SelfAttn([Rt,Ot])[R_t', O_t'] = \text{SelfAttn}([R_t, O_t]) The pooled feature vector vtv_t for the current node is obtained by average pooling all vectors in Rt′R_t' and Ot′O_t'.
    • Navigable node representations: For unvisited candidate neighbors in N(Vt)\mathcal{N}(V_t), visual representations are initialized from the corresponding directional view embedding in Rt′R_t'. If a navigable node is visible from multiple visited vantage points across navigation time steps, its visual feature vector is the element-wise average of all observed partial view embeddings.
    • Structural and temporal embeddings: Each node representation viv_i is augmented by adding a navigation step encoding (the most recent visit timestamp for visited nodes, and 00 for unexplored navigable nodes) and an egocentric location encoding representing the relative heading, elevation, and distance of ViV_i with respect to the current node VtV_t.
    • Virtual stop node: A dedicated 'stop' token node v0v_0 is maintained in VtV_t and connected to all other nodes in EtE_t.
  3. Knowl 3 — Graph-Aware Self-Attention for Topological Map Encoding

    equation

    To encode graph connectivity and geometric layout into transformer representations, the coarse-scale cross-modal encoder applies Graph-Aware Self-Attention (GASA) over the node representations.

    Given a matrix of KtK_t node representations X∈RKt×dX \in \mathbb{R}^{K_t \times d} and a topological pairwise shortest-path distance matrix E∈RKt×KtE \in \mathbb{R}^{K_t \times K_t} derived from map edges EtE_t, GASA is computed as: GASA(X)=Softmax(XWq(XWk)Td+M)XWv\text{GASA}(X) = \text{Softmax}\left(\frac{X W_q (X W_k)^T}{\sqrt{d}} + M\right) X W_v where Wq,Wk,Wv∈Rd×dW_q, W_k, W_v \in \mathbb{R}^{d \times d} are learnable projection weight matrices, dd is the feature dimension, and M∈RKt×KtM \in \mathbb{R}^{K_t \times K_t} is a learnable relational bias matrix: M=EWe+beM = E W_e + b_e with learnable scalar weight We∈RW_e \in \mathbb{R} and bias be∈Rb_e \in \mathbb{R}.

    The distance bias MM ensures that attention weights between nodes reflect true topological proximity in the physical environment rather than purely visual feature similarity.

  4. Knowl 4 — Dynamic Dual-Scale Action Fusion Mechanism

    model/method

    DUET unifies global coarse-scale path reasoning with fine-scale local action reasoning by mapping local action predictions into the global action space and combining them via dynamic gating.

    1. Coarse-scale Prediction: The coarse-scale encoder produces action logits sic=FFN(v^i)s_i^c = \text{FFN}(\hat{v}_i) for all nodes Vi∈Vt∪{v0}V_i \in V_t \cup \{v_0\} over the global action space ⋃i=1tAi\bigcup_{i=1}^t \mathcal{A}_i. Scores for visited nodes are masked during navigation.
    2. Fine-scale Score Conversion: The fine-scale encoder generates scores sifs_i^f over the local action space At={stop}∪N(Vt)\mathcal{A}_t = \{\text{stop}\} \cup \mathcal{N}(V_t). To align local predictions with the full topological map, a backtracking score sbacks_{back} is computed by summing the scores of all visited neighboring nodes in N(Vt)\mathcal{N}(V_t): sback=∑Vj∈N(Vt)∩Visitedsjfs_{back} = \sum_{V_j \in \mathcal{N}(V_t) \cap \text{Visited}} s_j^f The global converted fine-scale score vector sf′s^{f\prime} is defined as: sif′={sif,if Vi∈{stop}∪N(Vt)sback,if Vi∈Vt∖N(Vt)s_i^{f\prime} = \begin{cases} s_i^f, & \text{if } V_i \in \{\text{stop}\} \cup \mathcal{N}(V_t) \\ s_{back}, & \text{if } V_i \in V_t \setminus \mathcal{N}(V_t) \end{cases}
    3. Dynamic Fusion: A scalar mixing coefficient σt∈[0,1]\sigma_t \in [0, 1] is dynamically estimated from the coarse-scale stop node representation v^0\hat{v}_0 and the fine-scale stop token representation r^0\hat{r}_0: σt=Sigmoid(FFN([v^0;r^0]))\sigma_t = \text{Sigmoid}(\text{FFN}([\hat{v}_0; \hat{r}_0]))
    4. Combined Action Scoring: The final action score for any node ViV_i is: si=σtsic+(1−σt)sif′s_i = \sigma_t s_i^c + (1 - \sigma_t) s_i^{f\prime}
  5. Knowl 5 — Policy Optimization via Pseudo Interactive Demonstrator

    algorithm

    To address exposure bias in imitation learning, DUET utilizes a Pseudo Interactive Demonstrator (PID) denoted π∗\pi^* to provide supervision on policy rollouts during training.

    Input: Natural language instruction WW, target destination node V∗V^*, ground-truth environment graph GG, policy parameters θ\theta, weight λ\lambda
    Output: Updated policy parameters θ\theta
    Sample trajectory P=(V1,V2,…,VT)P = (V_1, V_2, \dots, V_T) by rolling out current policy πθ\pi_\theta in GG
    for t=1t = 1 to TT do
        Identify unexplored navigable nodes NGt\mathcal{N}_{G_t} on explored topological map GtG_t
        atπ∗=arg⁡min⁡u∈NGt(distG(Vt,u)+distG(u,V∗))a_t^{\pi^*} = \arg\min_{u \in \mathcal{N}_{G_t}} \left(\text{dist}_G(V_t, u) + \text{dist}_G(u, V^*)\right)
        Compute step demonstrator loss: LPID(t)=−log⁡p(atπ∗∣W,P<t)L_{PID}^{(t)} = -\log p(a_t^{\pi^*} \mid W, P_{<t})
    end for
    LPID=∑t=1TLPID(t)L_{PID} = \sum_{t=1}^T L_{PID}^{(t)}
    Sample expert demonstration path P∗=(V1∗,…,VT′∗)P^* = (V_1^*, \dots, V_{T'}^*) with actions at∗a_t^* and ground-truth object o∗o^*
    LSAP=∑t=1T′−log⁡p(at∗∣W,P<t∗)L_{SAP} = \sum_{t=1}^{T'} -\log p(a_t^* \mid W, P_{<t}^*)
    LOG=−log⁡p(o∗∣W,PT′∗)L_{OG} = -\log p(o^* \mid W, P_{T'}^*)
    L=λLSAP+LPID+LOGL = \lambda L_{SAP} + L_{PID} + L_{OG}
    Update model parameters θ\theta using ∇θL\nabla_\theta L
  6. Knowl 6 — Test-Time Global Action Planning and Route Execution

    algorithm

    During inference in unseen environments, DUET plans navigation actions over the full topological map and executes multi-step shortest routes between non-adjacent nodes.

    Input: Instruction tokens WW, maximum step budget TmaxT_{max}
    Output: Final stopping node VstopV_{stop}, predicted target object index o∗o^*
    Initialize topological map G0=({V1},∅)G_0 = (\{V_1\}, \emptyset), current step counter t=1t = 1
    while t≤Tmaxt \le T_{max} do
        Extract panoramic image tokens RtR_t and object tokens OtO_t at current location VtV_t
        Update topological graph Gt=(Vt,Et)G_t = (V_t, E_t) with unvisited neighbors N(Vt)\mathcal{N}(V_t) and update node embeddings
        Compute action scores sis_i for all candidate nodes Vi∈Vt∪{v0}V_i \in V_t \cup \{v_0\} via dynamic dual-scale fusion
        at=arg⁡max⁡isia_t = \arg\max_i s_i
        if at==v0a_t == v_0 (stop action) then
            Vstop=VtV_{stop} = V_t
            break
        else
            Target node Vtarget=atV_{target} = a_t
            Compute shortest path (Vt,u1,u2,…,Vtarget)(V_t, u_1, u_2, \dots, V_{target}) in GtG_t using Floyd-Warshall algorithm
            Navigate agent sequentially along path to VtargetV_{target}
            t=t+path_lengtht = t + \text{path\_length}
        end if
    end while
    if agent reached TmaxT_{max} without selecting stop action then
        Vstop=arg⁡max⁡Vi∈Visitedp(stop at Vi)V_{stop} = \arg\max_{V_i \in \text{Visited}} p(\text{stop at } V_i)
        Navigate agent to VstopV_{stop} along shortest path in GtG_t
    end if
    Predict target object o∗=arg⁡max⁡jFFN(O^Vstop,j)o^* = \arg\max_j \text{FFN}(\hat{O}_{V_{stop}, j}) at node VstopV_{stop}
    return Vstop,o∗V_{stop}, o^*
  7. Knowl 7 — Performance on Goal-Oriented VLN Benchmarks REVERIE and SOON

    data/table

    DUET evaluated against previous state-of-the-art methods on two goal-oriented VLN benchmarks: REVERIE (high-level target room and object grounding with predefined bounding boxes) and SOON (detailed instructions without predefined object boxes, using BUTD object detector proposals).

    Evaluation metrics include Trajectory Length (TL, meters), Oracle Success Rate (OSR, %), Success Rate (SR, %), Success Rate weighted by Path Length (SPL, %), Remote Grounding Success (RGS, %), and RGS penalized by Path Length (RGSPL, %).

    Method REVERIE Val Unseen REVERIE Test Unseen
    TL OSR↑\uparrow SR↑\uparrow SPL↑\uparrow RGS↑\uparrow RGSPL↑\uparrow TL OSR↑\uparrow SR↑\uparrow SPL↑\uparrow RGS↑\uparrow RGSPL↑\uparrow
    Human - - - - - - 21.18 86.83 81.51 53.66 77.84 51.44
    Seq2Seq 11.07 8.07 4.20 2.84 2.16 1.63 10.89 6.88 3.99 3.09 2.00 1.58
    RCM 11.98 14.23 9.29 6.97 4.89 3.89 10.60 11.68 7.84 6.67 3.67 3.14
    SMNA 9.07 11.28 8.15 6.44 4.54 3.61 9.23 8.39 5.80 4.53 3.10 2.39
    FAST-MATTN 45.28 28.20 14.40 7.19 7.84 4.67 39.05 30.63 19.88 11.61 11.28 6.08
    SIA 41.53 44.67 31.53 16.28 22.41 11.56 48.61 44.56 30.80 14.85 19.02 9.20
    RecBERT 16.78 35.02 30.67 24.90 18.77 15.27 15.86 32.91 29.61 23.99 16.50 13.51
    Airbert 18.71 34.51 27.89 21.88 18.23 14.18 17.91 34.20 30.28 23.61 16.83 13.28
    HAMT 14.08 36.84 32.95 30.20 18.92 17.28 13.62 33.41 30.40 26.67 14.88 13.08
    DUET (Ours) 22.11 51.07 46.98 33.73 32.15 23.03 21.30 56.91 52.51 36.06 31.88 22.06
    Method SOON Val Unseen SOON Test Unseen
    TL OSR↑\uparrow SR↑\uparrow SPL↑\uparrow RGSPL↑\uparrow TL OSR↑\uparrow SR↑\uparrow SPL↑\uparrow RGSPL↑\uparrow
    GBE 28.96 28.54 19.52 13.34 1.16 27.88 21.45 12.90 9.23 0.45
    DUET (Ours) 36.20 50.91 36.28 22.58 3.75 41.83 43.00 33.44 21.42 4.17

    On REVERIE Val Unseen, DUET outperforms the previous best model HAMT by 14.03%14.03\% in SR (46.98%46.98\% vs 32.95%32.95\%), 3.53%3.53\% in SPL, and 5.75%5.75\% in RGSPL. On REVERIE Test Unseen, DUET achieves a 22.11%22.11\% improvement in SR (52.51%52.51\% vs 30.40%30.40\%) and an 8.98%8.98\% improvement in RGSPL (22.06%22.06\% vs 13.08%13.08\%). On the SOON Test Unseen split, DUET achieves 33.44%33.44\% SR and 21.42%21.42\% SPL, outperforming the previous graph-based exploration method GBE (12.90%12.90\% SR and 9.23%9.23\% SPL) by large margins.

  8. Knowl 8 — Performance Comparison on Fine-Grained VLN Benchmark R2R

    data/table

    DUET evaluated on the R2R benchmark, which provides fine-grained, step-by-step navigation instructions. Methods are grouped by memory architecture: Recurrent state ('Rec'), Sequence history ('Seq'), and Topological map ('Map').

    Memory Method R2R Val Unseen R2R Test Unseen
    TL↓\downarrow NE↓\downarrow SR↑\uparrow SPL↑\uparrow TL↓\downarrow NE↓\downarrow SR↑\uparrow SPL↑\uparrow
    Rec Seq2Seq 8.39 7.81 22 - 8.13 7.85 20 18
    SF - 6.62 35 - 14.82 6.62 35 28
    PRESS 10.36 5.28 49 45 10.77 5.49 49 45
    EnvDrop 10.70 5.22 52 48 11.66 5.23 51 47
    AuxRN - 5.28 55 50 - 5.15 55 51
    PREVALENT 10.19 4.71 58 53 10.51 5.30 54 51
    RelGraph 9.99 4.73 57 53 10.29 4.75 55 52
    RecBERT 12.01 3.93 63 57 12.35 4.09 63 57
    Seq HAMT 11.87 3.65 65 59 12.65 4.11 63 58
    HAMT-e2e 11.46 2.29 66 61 12.27 3.93 65 60
    Map EGP - 4.83 56 44 - 5.34 53 42
    GBE - 5.20 54 43 - 5.18 53 43
    SSM 20.7 4.32 62 45 20.4 4.57 61 46
    DUET-coarse 12.96 3.67 68 59 13.08 3.93 67 58
    DUET (Ours) 13.94 3.31 72 60 14.73 3.65 69 59

    DUET achieves 72%72\% SR on Val Unseen and 69%69\% SR on Test Unseen, surpassing sequence memory models like HAMT-e2e by 6%6\% and 4%4\%, while maintaining competitive SPL (60%60\% and 59%59\%) despite longer trajectory lengths caused by global backtracking. When using only the coarse-scale encoder (DUET-coarse), it achieves 68%68\% SR on Val Unseen, outperforming prior map-based approaches (EGP at 56%56\%, GBE at 54%54\%, and SSM at 62%62\%).

  9. Knowl 9 — Ablation of Visual Scale Encoders, Action Fusion, and Graph-Aware Attention

    data/table

    Ablation experiments on the REVERIE Val Unseen split analyze the individual contributions of fine- vs coarse-scale representations, fusion strategies, and Graph-Aware Self-Attention (GASA).

    Scale Fusion OSR↑\uparrow SR↑\uparrow SROSR↑\frac{\text{SR}}{\text{OSR}}\uparrow SPL↑\uparrow RGS↑\uparrow RGSPL↑\uparrow
    Fine - 30.96 28.86 93.22 23.57 20.39 16.64
    Coarse - 46.44 36.52 78.64 25.98 - -
    Multi Average 51.86 45.81 88.33 31.94 32.49 22.78
    Multi Dynamic 51.07 46.98 91.40 33.73 32.15 23.03
    Fusion GASA OSR↑\uparrow SR↑\uparrow SPL↑\uparrow RGS↑\uparrow RGSPL↑\uparrow
    Average ×\times 49.22 44.50 30.90 29.88 20.73
    ✓ 51.86 45.81 31.94 32.49 22.78
    Dynamic ×\times 49.25 45.24 32.88 29.91 21.57
    ✓ 51.07 46.98 33.73 32.15 23.03

    Key observations:

    1. Scale Trade-offs: The fine-scale encoder achieves high stop accuracy (SROSR=93.22%\frac{\text{SR}}{\text{OSR}} = 93.22\%) but poor exploration (30.96%30.96\% OSR) due to its local action space. The coarse-scale encoder enables broad exploration (46.44%46.44\% OSR) but lacks object features for target localization.
    2. Dynamic Fusion: Dynamic fusion outperforms uniform score averaging across all key metrics, improving SR from 45.81%45.81\% to 46.98%46.98\% and SPL from 31.94%31.94\% to 33.73%33.73\%.
    3. Graph-Aware Self-Attention (GASA): Integrating topological graph distances into self-attention yields consistent gains, improving SPL by +1.04%+1.04\% with average fusion and +0.85%+0.85\% with dynamic fusion.
  10. Knowl 10 — Ablation of Training Losses, Interactive Demonstrator, and Synthetic Augmentation

    data/table

    Ablations on REVERIE Val Unseen evaluate pretraining objectives, policy fine-tuning regimes, and synthetic speaker data augmentation.

    Pretrain Finetune OSR↑\uparrow SR↑\uparrow SPL↑\uparrow RGS↑\uparrow RGSPL↑\uparrow
    SAP OG Aux RL PID
    ✓ ×\times ×\times ×\times ×\times 38.45 35.30 24.55 - -
    ✓ ✓ ×\times ×\times ×\times 40.24 37.80 26.40 23.89 16.36
    ✓ ✓ ✓ ×\times ×\times 37.63 36.81 27.19 25.05 18.40
    ✓ ✓ ✓ ✓ ×\times 47.51 42.35 32.97 29.91 23.53
    ✓ ✓ ✓ ×\times ✓ 51.07 46.98 33.73 32.15 23.03
    PID Aug OSR↑\uparrow SR↑\uparrow SPL↑\uparrow RGS↑\uparrow RGSPL↑\uparrow
    ×\times ×\times 37.29 34.56 25.56 23.00 16.64
    ✓ 37.63 36.81 27.19 25.05 18.40
    ✓ ×\times 51.07 46.98 33.73 32.15 23.03
    ✓ 52.09 46.58 32.72 31.75 22.18

    Key observations:

    1. Auxiliary Pretraining Losses: Object grounding (OG) and auxiliary tasks (MLM, MRC) improve grounding performance from 16.36%16.36\% to 18.40%18.40\% RGSPL and navigation SPL from 24.55%24.55\% to 27.19%27.19\%.
    2. PID vs RL: Fine-tuning with the Pseudo Interactive Demonstrator (PID) outperforms standard A3C reinforcement learning across all metrics, improving SR from 42.35%42.35\% to 46.98%46.98\% and OSR from 47.51%47.51\% to 51.07%51.07\%.
    3. Synthetic Data Augmentation: Synthetic instructions improve pretraining (27.19%27.19\% vs 25.56%25.56\% SPL), but PID policy fine-tuning achieves superior SPL when trained strictly on clean human data (33.73%33.73\% vs 32.72%32.72\%).

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Peter Anderson, Angel Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, et al. On evaluation of embodied navigation agents. arXiv preprint arXiv:1807.06757, 2018. 1, 6
  2. 2.Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sunderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In CVPR, pages 3674–3683, 2018. 1, 2, 3, 6, 8
  3. 3.Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi. Touchdown: Natural language navigation and spatial reasoning in visual street environments. In CVPR, pages 12538–12547, 2019. 1, 2
  4. 4.Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. In EMNLP, pages 4392–4412, 2020. 1, 2
  5. 5.Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee. Beyond the nav-graph: Vision-and-language navigation in continuous environments. In ECCV, pages 104–120. Springer, 2020. 1, 2
  6. 6.Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In CVPR, pages 10740–10749, 2020. 1, 2
  7. 7.Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. Reverie: Remote embodied visual referring expression in real indoor environments. In CVPR, pages 9982–9991, 2020. 1, 3, 6, 8
  8. 8.Fengda Zhu, Xiwen Liang, Yi Zhu, Qizhi Yu, Xiaojun Chang, and Xiaodan Liang. Soon: Scenario oriented object navigation with graph-based exploration. In CVPR, pages 12689–12699, 2021. 1, 2, 3, 6, 7, 8
  9. 9.Muhammad Zubair Irshad, Chih-Yao Ma, and Zsolt Kira. Hierarchical cross-modal agent for robotics vision-and-language navigation. In ICRA, pages 13238–13246, 2021. 1, 2
  10. 10.Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. Speaker-follower models for vision-and-language navigation. In NeurIPS, pages 3318–3329, 2018. 1, 2, 3, 6, 8
  11. 11.Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, Zsolt Kira, Richard Socher, and Caiming Xiong. Self-monitoring navigation agent via auxiliary progress estimation. In ICLR, 2019. 1, 6, 8
  12. 12.Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In CVPR, pages 6629–6638, 2019. 1, 8
  13. 13.Hao Tan, Licheng Yu, and Mohit Bansal. Learning to navigate unseen environments: Back translation with environmental dropout. In NAACL, pages 2610–2621, 2019. 1, 2, 3, 8
  14. 14.Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez-Opazo, and Stephen Gould. Vln BERT: A recurrent vision-and-language BERT for navigation. In CVPR, pages 1643–1653, 2021. 1, 2, 3, 7, 8
  15. 15.Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, and Ivan Laptev. History aware multimodal transformer for vision-and-language navigation. In NeurIPS, 2021. 2, 3, 5, 7, 8
  16. 16.Alexander Pashevich, Cordelia Schmid, and Chen Sun. Episodic transformer for vision-and-language navigation. In ICCV, pages 15942–15952, 2021. 2, 3, 5
  17. 17.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, pages 5998–6008, 2017. 2, 4
  18. 18.Devendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta, and Saurabh Gupta. Neural topological SLAM for visual navigation. In CVPR, pages 12875–12884, 2020. 2
  19. 19.Zhiwei Deng, Karthik Narasimhan, and Olga Russakovsky. Evolving graphical planner: Contextual global planning for vision-and-language navigation. In NeurIPS, volume 33, 2020. 2, 3, 8
  20. 20.Hanqing Wang, Wenguan Wang, Wei Liang, Caiming Xiong, and Jianbing Shen. Structured scene memory for vision-language navigation. In CVPR, pages 8455–8464, 2021. 2, 3, 8
  21. 21.Vihan Jain, Gabriel Magalhaes, Alexander Ku, Ashish Vaswani, Eugene Ie, and Jason Baldridge. Stay on the path: Instruction fidelity in vision-and-language navigation. In ACL, pages 1862–1872, 2019. 2
  22. 22.Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra. Embodied question answering. In CVPR, pages 1–10, 2018. 2
  23. 23.Licheng Yu, Xinlei Chen, Georgia Gkioxari, Mohit Bansal, Tamara L Berg, and Dhruv Batra. Multi-target embodied question answering. In CVPR, pages 6309–6318, 2019. 2
  24. 24.Chih-Yao Ma, Zuxuan Wu, Ghassan AlRegib, Caiming Xiong, and Zsolt Kira. The regretful agent: Heuristic-aided navigation through progress estimation. In CVPR, pages 6732–6740, 2019. 2, 3
  25. 25.Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool. Talk2nav: Long-range vision-and-language navigation with dual attention and spatial memory. International Journal of Computer Vision, 129(1):246–266, 2021. 2
  26. 26.Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao. Towards learning a generic agent for vision-and-language navigation via pretraining. In CVPR, pages 13137–13146, 2020. 2, 5, 8
  27. 27.Xiujun Li, Chunyuan Li, Qiaolin Xia, Yonatan Bisk, Asli Celikyilmaz, Jianfeng Gao, Noah A Smith, and Yejin Choi. Robust navigation with language pretraining and stochastic sampling. In EMNLP, pages 1494–1499, 2019. 2, 8
  28. 28.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL, pages 4171–4186, 2019. 2, 4, 5
  29. 29.Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra. Improving vision-and-language navigation with image-text pairs from the web. In ECCV, pages 259–274. Springer, 2020. 2
  30. 30.Pierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen, Ivan Laptev, and Cordelia Schmid. Airbert: In-domain pretraining for vision-and-language navigation. In ICCV, pages 1634–1643, 2021. 2, 8
  31. 31.Sebastian Thrun. Probabilistic robotics. Communications of the ACM, 45(3):52–57, 2002. 3
  32. 32.Sebastian Thrun. Learning metric-topological maps for indoor mobile robot navigation. Artificial Intelligence, 99(1):21–71, 1998. 3
  33. 33.Albert S Huang, Abraham Bachrach, Peter Henry, Michael Krainin, Daniel Maturana, Dieter Fox, and Nicholas Roy. Visual odometry and mapping for autonomous flight using an rgb-d camera. In Robotics Research, pages 235–252. Springer, 2017. 3
  34. 34.Jingwei Zhang, Lei Tai, Ming Liu, Joschka Boedecker, and Wolfram Burgard. Neural SLAM: Learning to explore with external memory. arXiv preprint arXiv:1706.09520, 2017. 3
  35. 35.Saurabh Gupta, James Davidson, Sergey Levine, Rahul Sukthankar, and Jitendra Malik. Cognitive mapping and planning for visual navigation. In CVPR, pages 2616–2625, 2017. 3
  36. 36.Devendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta, and Ruslan Salakhutdinov. Learning to explore using active neural SLAM. In ICLR, 2020. 3
  37. 37.Peter Anderson, Ayush Shrivastava, Devi Parikh, Dhruv Batra, and Stefan Lee. Chasing ghosts: Instruction following as bayesian state tracking. NeurIPS, 32:371–381, 2019. 3
  38. 38.Nikolay Savinov, Alexey Dosovitskiy, and Vladlen Koltun. Semi-parametric topological memory for navigation. ICLR, 2018. 3
  39. 39.Kuan Fang, Alexander Toshev, Li Fei-Fei, and Silvio Savarese. Scene memory transformer for embodied agents in long-horizon tasks. In CVPR, pages 538–547, 2019. 3
  40. 40.Kevin Chen, Junshen K Chen, Jo Chuang, Marynel Vazquez, and Silvio Savarese. Topological planning with transformers for vision-and-language navigation. In CVPR, pages 11276–11286, 2021. 3
  41. 41.Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. Scheduled sampling for sequence prediction with recurrent neural networks. In NeurIPS, volume 28, 2015. 3
  42. 42.Stéphane Ross, Geoffrey Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In AISTATS, pages 627–635. JMLR Workshop and Conference Proceedings, 2011. 3, 5
  43. 43.Richard S Sutton and Andrew G Barto. Reinforcement learning: An introduction. MIT press, 2018. 3
  44. 44.Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In ICML, pages 1928–1937. PMLR, 2016. 3
  45. 45.Hu Wang, Qi Wu, and Chunhua Shen. Soft expert reward learning for vision-and-language navigation. In ECCV, pages 126–141. Springer, 2020. 3
  46. 46.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR, pages 6077–6086, 2018. 3, 6
  47. 47.Hao Tan and Mohit Bansal. LXMERT: Learning cross-modality encoder representations from transformers. In EMNLP, pages 5103–5114, 2019. 4, 5, 6
  48. 48.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In NeurIPS, volume 32, 2019. 5
  49. 49.Xiangru Lin, Guanbin Li, and Yizhou Yu. Scene-intuitive agent for remote embodied visual grounding. In CVPR, pages 7036–7045, 2021. 5, 8
  50. 50.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16× 16 words: Transformers for image recognition at scale. ICLR, 2020. 6
  51. 51.Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang. Vision-language navigation with self-supervised auxiliary reasoning tasks. In CVPR, pages 10012–10022, 2020. 8
  52. 52.Yicong Hong, Cristian Rodriguez, Yuankai Qi, Qi Wu, and Stephen Gould. Language and visual entity relationship graph for agent navigation. NeurIPS, 33:7685–7696, 2020. 8
  53. 53.Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. Show, attend and tell: Neural image caption generation with visual attention. In ICML, pages 2048–2057. PMLR, 2015.
  54. 54.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In EMNLP, pages 1532–1543, 2014.
  55. 55.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niebner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. In 3DV, pages 667–676. IEEE, 2017.
  56. 56.Federico Landi, Lorenzo Baraldi, Marcella Cornia, Massimiliano Corsini, and Rita Cucchiara. Perceive, transform, and act: Multi-modal attention networks for vision-and-language navigation. arXiv preprint arXiv:1911.12377, 2019.
  57. 57.Gabriel Ilharco, Vihan Jain, Alexander Ku, Eugene Ie, and Jason Baldridge. General evaluation for instruction conditioned navigation using dynamic time warping. In NeurIPS Workshop, 2019.

Citation

MLA
Chen, S., et al. “Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation”. arXiv, 2022, http://arxiv.org/abs/2202.11742v1.
APA
Chen, S., Guhur, P.-L., Tapaswi, M., Schmid, C., & Laptev, I. (2022). Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation. arXiv. http://arxiv.org/abs/2202.11742v1
Chicago
Chen, S., P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev. 2022. “Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation”. arXiv. http://arxiv.org/abs/2202.11742v1.
Harvard
Chen, S. et al. (2022) “Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2202.11742v1.
Vancouver
1. Chen S, Guhur P-L, Tapaswi M, Schmid C, Laptev I (2022) Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation. arXiv

BibTeX

@article{chen2022think,
  title = {Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language Navigation},
  author = {Chen, Shizhe and Guhur, Pierre-Louis and Tapaswi, Makarand and Schmid, Cordelia and Laptev, Ivan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2202.11742v1},
  eprint = {2202.11742}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE