Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation

Hanqing WangWei LiangJianbing ShenLuc Van GoolWenguan Wang

article2022CVPR86 citations

Proposes a cycle-consistent framework that jointly trains instruction-following and instruction-generation agents using counterfactual scene synthesis, allowing vision-language navigation models to learn effectively from both labeled trajectories and unlabeled paths.

Listen

Autonomous agents operating alongside humans in physical environments must be able to understand visual scenes and natural language instructions. While existing research in vision-language navigation has predominantly concentrated on training an agent to follow route descriptions (instruction following), far less focus has been given to the reverse capability: generating accurate, human-understandable route descriptions from visual paths (instruction generation). Traditional workflows treat instruction generation merely as a disconnected, isolated tool for generating synthetic training data. This isolation limits navigation robustness and prevents agents from effectively communicating, explaining actions, and collaborating with human partners during complex operations such as search-and-rescue.

The article introduces and evaluates a counterfactual cycle-consistent learning framework that jointly trains a route-following agent (follower) and an instruction-generating agent (speaker) alongside an environment-generating module (creator). The primary objective is to demonstrate that coupling instruction following and instruction generation in a closed learning loop, enhanced with counterfactual visual scenes, systematically improves the performance of both tasks across diverse navigation architectures.

The evaluated framework connects the follower and speaker so each evaluates the other: the speaker assesses whether a follower's path matches the given instruction, while the follower assesses whether a speaker's generated instruction produces the original path. Because this cycle-consistency mechanism relies on circular verification rather than aligned human labels, it seamlessly incorporates unlabeled navigation paths alongside standard labeled datasets. In addition, the creator module synthesizes counterfactual training environments by blending elements from alternative reference scenes while preserving critical navigation landmarks. The entire system was validated using standard benchmarks on the photo-realistic Room-to-Room dataset, testing across multiple baseline architectures and leading navigation models under both seen and previously unseen environments.

Key findings show significant, consistent performance gains across both navigation and text generation tasks. First, the proposed framework substantially boosted follower success rates across different model architectures, outperforming previous data augmentation methods by up to 4.8 percentage points in seen environments and up to 6.1 percentage points on unseen test environments. Second, when applied to existing top-performing benchmark followers, the framework consistently set new performance highs, raising the success rate of a leading memory-based navigation model to 62.2%. Third, the framework dramatically improved the linguistic quality of generated instructions, outperforming standalone speaker models across all standard language metrics. Finally, blind human user evaluations confirmed this language improvement, with human reviewers preferring the system's generated instructions over existing generation baselines by a 68.6% to 18.2% margin.

These results demonstrate that treating route following and instruction generation as interdependent tasks provides mutual regularization and resolves data quality issues inherent in traditional synthetic data pipelines. By enabling agents to generalize more effectively to unseen environments without requiring expensive manual annotations, this approach lowers the operational risk and cost of deploying autonomous robots. Technical leaders and engineering teams developing embodied artificial intelligence should integrate dual-task cycle consistency and counterfactual scene generation into their training pipelines rather than relying on isolated data augmentation steps. Future development should focus on extending these counterfactual training techniques to dynamic, real-time physical environments and evaluating their performance under noisy, real-world robotic sensor conditions.

Cover for Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation

Abstract

Since the rise of vision-language navigation (VLN), great progress has been made in instruction following – building a follower to navigate environments under the guidance of instructions. However, far less attention has been paid to the inverse task: instruction generation – learning a speaker to generate grounded descriptions for navigation routes. Existing VLN methods train a speaker independently and often treat it as a data augmentation tool to strengthen the follower, while ignoring rich cross-task relations. Here we describe an approach that learns the two tasks simultaneously and exploits their intrinsic correlations to boost the training of each: the follower judges whether the speaker-created instruction explains the original navigation route correctly, and vice versa. Without the need of aligned instruction-path pairs, such cycle-consistent learning scheme is complementary to task-specific training targets defined on labeled data, and can also be applied over unlabeled paths (sampled without paired instructions). Another agent, called creator is added to generate counterfactual environments. It greatly changes current scenes yet leaves novel items – which are vital for the execution of original instructions – unchanged. Thus more informative training scenes are synthesized and the three agents compose a powerful VLN learning system. Extensive experiments on a standard benchmark show that our approach improves the performance of various follower models and produces accurate navigation instructions.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Cycle-Consistent Learning for Instruction Following and Generation
  • 3.2. Counterfactual Environment Creation
  • 3.3. Implementation Details
  • 4. Experiment
  • 4.1. Performance on Instruction Following
  • 4.2. Performance on Instruction Generation
  • 4.3. Diagnostic Experiments
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Dual-Agent Cycle-Consistent Learning Formulation for Vision-Language Navigation

    model/method

    In Vision-Language Navigation (VLN), instruction following is framed as learning a follower agent f:E×X→Af: \mathcal{E} \times \mathcal{X} \to \mathcal{A} that maps visual observations of environment E∈EE \in \mathcal{E} and linguistic instruction X={xl}l=1L∈XX = \{x_l\}_{l=1}^L \in \mathcal{X} to a sequence of actions A={at}t=1T∈AA = \{a_t\}_{t=1}^T \in \mathcal{A}. The dual task of instruction generation is framed as learning a speaker agent s:E×A→Xs: \mathcal{E} \times \mathcal{A} \to \mathcal{X} that maps visual environment EE and path AA to an instruction XX.

    Rather than training ss and ff in isolation, the two tasks are trained jointly via cycle-consistency losses that treat each agent as an evaluator for the other. For an instruction XX in environment EE, the follower generates a trajectory A^∼P(A∣X;E)\hat{A} \sim P(A|X; E), which the speaker translates back into an instruction X^∼P(X∣A^;E)\hat{X} \sim P(X|\hat{A}; E). Conversely, for a trajectory AA in environment EE, the speaker generates an instruction X^∼P(X∣A;E)\hat{X} \sim P(X|A; E), which the follower executes to produce a path A^∼P(A∣X^;E)\hat{A} \sim P(A|\hat{X}; E). The negative log-likelihood cycle-consistency errors are defined as:

    ΔEA=−log⁡∑X^∈XP(X^∣A;E)P(A∣X^;E)\Delta_E^A = -\log \sum_{\hat{X} \in \mathcal{X}} P(\hat{X}|A; E) P(A|\hat{X}; E)

    ΔEX=−log⁡∑A^∈AP(A^∣X;E)P(X∣A^;E)\Delta_E^X = -\log \sum_{\hat{A} \in \mathcal{A}} P(\hat{A}|X; E) P(X|\hat{A}; E)

    Given a labeled dataset D={(En,Xn,An)}n=1N\mathcal{D} = \{(E_n, X_n, A_n)\}_{n=1}^N and an unlabeled dataset of randomly sampled trajectories without instructions U={(Em,Am′)}m=1M\mathcal{U} = \{(E_m, A'_m)\}_{m=1}^M, the total cycle-consistent loss is:

    Lcycle=1N∑(E,X,A)∈D(ΔEA+ΔEX)+1M∑(E,A′)∈UΔEA′\mathcal{L}_{cycle} = \frac{1}{N} \sum_{(E, X, A) \in \mathcal{D}} (\Delta_E^A + \Delta_E^X) + \frac{1}{M} \sum_{(E, A') \in \mathcal{U}} \Delta_E^{A'}

  2. Knowl 2 — Policy Gradient Optimization for Cycle-Consistent VLN Objectives

    equation

    Because direct marginalization over the combinatorial spaces of natural language instructions X\mathcal{X} and trajectory sequences A\mathcal{A} is intractable, the gradients of the cycle-consistency error ΔEA\Delta_E^A with respect to the follower parameters θf\theta^f and speaker parameters θs\theta^s are estimated via policy gradient sampling:

    ∂ΔEA∂θf≈−EX^∼P(⋅∣A;E)[∂log⁡P(A∣X^;E)∂θf]≈−∂log⁡P(A∣X^;E)∂θf\frac{\partial \Delta_E^A}{\partial \theta^f} \approx -\mathbb{E}_{\hat{X} \sim P(\cdot|A; E)} \left[ \frac{\partial \log P(A|\hat{X}; E)}{\partial \theta^f} \right] \approx - \frac{\partial \log P(A|\hat{X}; E)}{\partial \theta^f}

    ∂ΔEA∂θs≈−EX^∼P(⋅∣A;E)[log⁡P(A∣X^;E)∂log⁡P(X^∣A;E)∂θs]≈−(log⁡P(A∣X^;E)−bf)∂log⁡P(X^∣A;E)∂θs\frac{\partial \Delta_E^A}{\partial \theta^s} \approx -\mathbb{E}_{\hat{X} \sim P(\cdot|A; E)} \left[ \log P(A|\hat{X}; E) \frac{\partial \log P(\hat{X}|A; E)}{\partial \theta^s} \right] \approx -\left( \log P(A|\hat{X}; E) - b^f \right) \frac{\partial \log P(\hat{X}|A; E)}{\partial \theta^s}

    where X^∼P(⋅∣A;E)\hat{X} \sim P(\cdot|A; E) is an instruction sampled from the speaker given path AA in environment EE, and bfb^f is a moving average baseline of previous log⁡P(A∣X^;E)\log P(A|\hat{X}; E) values used to reduce variance during reinforcement learning. Identical policy gradient formulations are derived symmetrically for ∂ΔEX∂θf\frac{\partial \Delta_E^X}{\partial \theta^f} and ∂ΔEX∂θs\frac{\partial \Delta_E^X}{\partial \theta^s} by sampling trajectories A^∼P(⋅∣X;E)\hat{A} \sim P(\cdot|X; E).

  3. Knowl 3 — Counterfactual Environment Creator Mechanism

    model/method

    The creator agent c:E×X×A→Eˉc: \mathcal{E} \times \mathcal{X} \times \mathcal{A} \to \bar{\mathcal{E}} generates counterfactual visual environments Eˉ\bar{E} that alter visual scene layouts as much as possible while maintaining the visual cues necessary to execute the original instruction-path pair (X,A)(X, A).

    Given an environment-instruction-path tuple (E,X,A)(E, X, A), the creator first summarizes the tuple into a global trajectory descriptor u=hTcu = h_T^c using an LSTM:

    htc=LSTMc([Vt,at,X],ht−1c)h_t^c = \text{LSTM}^c([V_t, a_t, X], h_{t-1}^c)

    where VtV_t is the visual scene observation at step tt, ata_t is the action embedding, and XX is the linguistic instruction representation.

    To construct a counterfactual scene Vˉ∈Eˉ\bar{V} \in \bar{E} at feature level, visual elements V={vk}k=1KV = \{v_k\}_{k=1}^K from the original scene V∈EV \in E are blended with visual elements Vr={vkr}k=1KV^r = \{v_k^r\}_{k=1}^K from a reference scene Vr∈ErV^r \in E^r (sampled from the training set D\mathcal{D}):

    qk=softmax([v1r,v2r,…,vKr]⊤vk)q_k = \text{softmax}\left([v_1^r, v_2^r, \dots, v_K^r]^\top v_k\right)

    gk=[v1r,v2r,…,vKr]⋅qkg_k = [v_1^r, v_2^r, \dots, v_K^r] \cdot q_k

    λk=sigmoid(u⊤W3vk)\lambda_k = \text{sigmoid}\left(u^\top W_3 v_k\right)

    vˉk=λkvk+(1−λk)gk\bar{v}_k = \lambda_k v_k + (1 - \lambda_k) g_k

    Here qkq_k is the attention correlation vector between element vkv_k and reference elements VrV^r, gkg_k is the reference feature summary, W3W_3 is a learnable weight matrix, and λk∈[0,1]\lambda_k \in [0, 1] represents the retention gate indicating whether original visual element vkv_k is critical for completing trajectory AA under instruction XX or can be replaced by reference features gkg_k.

  4. Knowl 4 — Counterfactual Creator Training Objective and Adversarial Alignment

    equation

    The creator agent cc is trained to maximize the diversity of altered visual features while ensuring that the generated counterfactual environment Eˉ\bar{E} remains realistic and compatible with the original instruction-path tuple (X,A)(X, A). The creator objective function is:

    Lc=Lℓ2+Ladv=∥λ∥2+log⁡(1−d(Vˉ,X,A))\mathcal{L}^c = \mathcal{L}_{\ell_2} + \mathcal{L}_{adv} = \|\lambda\|_2 + \log\left(1 - d(\bar{V}, X, A)\right)

    where λ=[λk]k\lambda = [\lambda_k]_k is the vector of feature retention coefficients across visual elements, and the ℓ2\ell_2-norm loss Lℓ2=∥λ∥2\mathcal{L}_{\ell_2} = \|\lambda\|_2 induces sparsity in λ\lambda to maximize the number of features replaced by reference scene elements.

    The adversarial loss Ladv\mathcal{L}_{adv} trains the creator to fool an adversarial discriminator d:E×X×A→[0,1]d: \mathcal{E} \times \mathcal{X} \times \mathcal{A} \to [0, 1], which is trained to distinguish authentic scene-instruction-trajectory alignments from modified ones by minimizing d(Vˉ,X,A)−d(V,X,A)d(\bar{V}, X, A) - d(V, X, A). To prevent the discriminator from latching onto visual trajectory memorization rather than alignment logic, dd uses only geometric action embeddings for trajectory representation.

  5. Knowl 5 — Instruction Following Performance Comparison Across Baseline Architectures on R2R

    data/table

    The counterfactual cycle-consistent (CCC) framework was evaluated on the Room-to-Room (R2R) benchmark across three baseline navigation models (Seq2Seq, Speaker-Follower, and Reinforced Cross-Modal Matching [RCM]) against standard Back-Translation (BT) and Adversarial Path Sampling (APS). Performance is measured across Success Rate (SR, %, ↑\uparrow), Navigation Error (NE, meters, ↓\downarrow), Oracle Success Rate (OR, %, ↑\uparrow), and Success Rate Weighted by Path Length (SPL, %, ↑\uparrow) on validation seen, validation unseen, and test unseen splits.

    Model val seen val unseen test unseen
    SR↑\uparrow NE↓\downarrow OR↑\uparrow SPL↑\uparrow SR↑\uparrow NE↓\downarrow OR↑\uparrow SPL↑\uparrow SR↑\uparrow NE↓\downarrow OR↑\uparrow SPL↑\uparrow
    Seq2Seq 39.4 6.0 51.7 33.8 22.1 7.8 27.7 19.1 20.4 7.9 26.6 18.0
    + BT 43.7 5.3 58.1 37.2 22.6 7.7 28.9 19.9 21.0 7.8 26.2 18.8
    + APS 48.2 5.0 60.8 40.1 24.2 7.1 32.7 20.4 22.5 7.5 30.1 19.3
    + CCC 50.1 5.0 61.1 42.6 28.4 6.8 35.3 22.1 25.5 7.8 35.9 20.6
    Speaker-Follower 51.7 5.0 61.6 44.4 29.9 6.9 40.7 21.0 30.9 7.0 41.2 24.0
    + BT 66.4 3.7 74.2 59.8 36.1 6.6 46.6 28.8 34.8 6.6 43.4 29.2
    + APS 68.2 3.3 74.9 62.5 38.8 6.1 46.7 32.1 36.1 6.5 44.2 28.8
    + CCC 68.4 3.3 74.5 61.4 43.5 5.8 52.0 38.1 41.4 5.9 51.0 36.6
    RCM 47.0 5.7 53.8 44.3 35.0 6.8 43.0 31.4 35.9 6.7 43.5 33.1
    + BT 61.9 4.1 66.9 58.6 45.6 5.7 52.4 41.8 44.5 5.9 52.4 40.8
    + APS 63.2 3.9 69.3 59.5 47.7 5.4 56.6 42.8 45.1 5.8 53.9 40.9
    + CCC 68.0 3.4 77.5 62.1 50.4 5.2 57.8 46.4 51.0 5.3 57.2 48.2

    Integrating CCC consistently outperforms BT and APS across baseline models. On unseen environments, CCC yields substantial gains; for example, on test unseen, CCC improves RCM's SR from 35.9% to 51.0% and SPL from 33.1% to 48.2%.

  6. Knowl 6 — Benchmarking CCC Enhancement on Top-Performing VLN Followers

    data/table

    The CCC framework was applied to leading published Vision-Language Navigation followers (Environmental Dropout [E-Dropout], Active Perception, and Structured Scene Memory [SSM]) on the R2R test unseen benchmark.

    Models SR↑\uparrow NE↓\downarrow OR↑\uparrow SPL↑\uparrow
    Self-Monitoring 43.0 6.0 55.0 32.0
    Regretful 48.0 5.7 56.0 40.0
    OAAM 53.0 - 61.0 50.0
    Tactical Rewind 54.0 5.1 64.0 41.0
    AuxRN 55.0 5.2 62.0 51.0
    E-Dropout 48.0 5.6 58.0 44.0
    E-Dropout + CCC 52.2 5.1 59.8 46.9
    Active Perception 55.7 4.8 73.1 37.1
    Active Perception + CCC 60.6 4.3 71.4 41.3
    SSM 57.3 4.7 68.2 44.1
    SSM + CCC 62.2 4.3 72.3 49.2

    Adding CCC provides consistent improvements across all base models: E-Dropout achieves a +4.2 SR gain and +2.9 SPL gain; Active Perception gains +4.9 SR and +4.2 SPL; and SSM gains +4.9 SR (reaching 62.2%) and +5.1 SPL (reaching 49.2%).

  7. Knowl 7 — Evaluation of Instruction Generation Models and Human Evaluation on R2R

    data/table

    The instruction generation performance of speakers trained under the CCC framework (Ours-Seq2Seq, Ours-Speaker-Follower, Ours-RCM) was compared against the Back-Translation (BT) speaker and the Vision-Language Speaker (VLS) on R2R validation seen and validation unseen splits across standard NLP metrics: BLEU-1, BLEU-4, CIDEr, METEOR, ROUGE, and SPICE (primary metric).

    Model val seen val unseen
    Bleu-1↑\uparrow Bleu-4↑\uparrow CIDEr↑\uparrow Meteor↑\uparrow Rouge↑\uparrow SPICE↑\uparrow Bleu-1↑\uparrow Bleu-4↑\uparrow CIDEr↑\uparrow Meteor↑\uparrow Rouge↑\uparrow SPICE↑\uparrow
    BT Speaker 0.537 0.155 0.121 0.233 0.350 0.203 0.522 0.142 0.114 0.228 0.346 0.188
    VLS 0.549 0.157 0.137 0.228 0.352 0.214 0.548 0.159 0.132 0.231 0.357 0.197
    Ours-Seq2Seq 0.720 0.296 0.529 0.233 0.487 0.216 0.704 0.273 0.475 0.229 0.473 0.202
    Ours-Speaker-Follower 0.723 0.299 0.566 0.235 0.490 0.229 0.706 0.275 0.477 0.229 0.474 0.207
    Ours-RCM 0.728 0.287 0.543 0.236 0.493 0.231 0.708 0.272 0.461 0.231 0.477 0.214

    All three CCC speakers achieve substantial gains over previous instruction generation methods across both seen and unseen splits. Among the three, Ours-RCM achieves the highest SPICE scores (0.231 on val seen, 0.214 on val unseen).

    In human user studies evaluating 500 sampled paths on val unseen:

    1. Comparing the three CCC models, Ours-RCM was selected by 25 participants as generating the most meaningful instruction 40.5% of the time, versus 32.1% for Ours-Speaker-Follower and 27.4% for Ours-Seq2Seq.
    2. In a head-to-head comparison against prior baselines, Ours-RCM was preferred with a 68.6% selection rate, compared to 18.2% for VLS and 13.2% for BT Speaker.
  8. Knowl 8 — Ablation Analysis of Cycle-Consistency Losses and Counterfactual Synthesis

    data/table

    An ablation study on the validation unseen split of R2R (based on the Speaker-Follower baseline) assesses the contributions of individual cycle-consistency loss terms and counterfactual scene generation components on both instruction following and instruction generation.

    Model Component Instruction Following Instruction Generation
    SR↑\uparrow NE↓\downarrow OR↑\uparrow SPL↑\uparrow Bleu-1↑\uparrow Bleu-4↑\uparrow CIDEr↑\uparrow Meteor↑\uparrow Rouge↑\uparrow SPICE↑\uparrow
    Baseline - 29.9 6.9 40.7 21.0 0.522 0.142 0.114 0.228 0.346 0.188
    Cycle-Consistency ΔEA\Delta_E^A 33.2 6.7 44.7 23.7 0.694 0.269 0.446 0.228 0.471 0.192
    ΔEX\Delta_E^X 28.6 6.9 40.1 20.5 0.699 0.271 0.456 0.228 0.472 0.195
    ΔEA+ΔEX\Delta_E^A + \Delta_E^X 35.1 6.7 44.5 25.4 0.702 0.272 0.462 0.229 0.472 0.196
    ΔEA+ΔEX+ΔEA′\Delta_E^A + \Delta_E^X + \Delta_E^{A'} 37.7 6.6 45.9 26.7 0.703 0.272 0.467 0.229 0.471 0.198
    Counterfactual Env. w/o reference ErE^r 29.9 6.9 40.7 21.0 0.522 0.142 0.114 0.228 0.346 0.188
    w/ reference ErE^r 41.7 5.9 50.7 35.6 0.701 0.271 0.459 0.229 0.472 0.200
    Full Model ΔEA+ΔEX+ΔEA′+ΔEˉA+ΔEˉX\Delta_E^A + \Delta_E^X + \Delta_E^{A'} + \Delta_{\bar{E}}^A + \Delta_{\bar{E}}^X 43.5 5.8 52.0 38.1 0.706 0.275 0.477 0.229 0.474 0.207

    Key takeaways from the ablation data:

    1. Adding ΔEA\Delta_E^A improves follower performance (+3.3 SR), whereas adding ΔEX\Delta_E^X alone strongly boosts speaker metrics (e.g., CIDEr increases from 0.114 to 0.456). Combining both terms achieves mutual improvement (35.1 SR, 0.462 CIDEr).
    2. Incorporating unlabeled paths with ΔEA′\Delta_E^{A'} further lifts SR to 37.7 and SPL to 26.7.
    3. Synthesizing counterfactuals with a reference environment ErE^r provides a large individual boost (SR 41.7, SPL 35.6), whereas naive random masking without reference scenes yields no gain over baseline.
    4. Combining all components in the full CCC framework delivers the highest overall performance on both tasks (SR 43.5, SPL 38.1, SPICE 0.207).

Coverage note — None was omitted; all key methodology components, loss formulations, gradient derivations, and primary empirical evaluations across both tasks and ablation studies were captured.

References

  1. 1.Ehsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi, and Anton van den Hengel. Counterfactual vision and language learning. In CVPR, 2020.
  2. 2.Sanyam Agarwal, Devi Parikh, Dhruv Batra, Peter Anderson, and Stefan Lee. Visual landmark selection for generating grounded and interpretable navigation instructions. In CVPR Workshop, 2019.
  3. 3.Vedika Agarwal, Rakshith Shetty, and Mario Fritz. Towards causal VQA: Revealing and reducing spurious correlations by invariant and covariant semantic editing. In CVPR, 2020.
  4. 4.Gary L Allen. From knowledge to words to wayfinding: Issues in the production and comprehension of route directions. In International Conference on Spatial Information Theory, 1997.
  5. 5.Peter Anderson, Angel Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, et al. On evaluation of embodied navigation agents. arXiv preprint arXiv:1807.06757, 2018.
  6. 6.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. Spice: Semantic propositional image caption evaluation. In ECCV, 2016.
  7. 7.Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sunderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In CVPR, 2018.
  8. 8.Jacob Andreas and Dan Klein. Alignment-based compositional semantics for instruction following. In EMNLP, 2015.
  9. 9.Sean Andrist, Erin Spannan, and Bilge Mutlu. Rhetorical robots: making robots more effective speakers using linguistic cues of expertise. In International Conference on Human-Robot Interaction, 2013.
  10. 10.Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In ACL Workshop, 2005.
  11. 11.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3D: Learning from RGB-D data in indoor environments. 3DV, 2017.
  12. 12.David L Chen and Raymond J Mooney. Learning to interpret natural language navigation instructions from observations. In AAAI, 2011.
  13. 13.Kevin Chen, Junshen K. Chen, Jo Chuang, Marynel Vazquez, and Silvio Savarese. Topological planning with transformers for vision-and-language navigation. In CVPR, 2021.
  14. 14.Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang. Counterfactual samples synthesizing for robust visual question answering. In CVPR, 2020.
  15. 15.Heriberto Cuayáhuitl, Nina Dethlefs, Lutz Frommberger, Kai-Florian Richter, and John Bateman. Generating adaptive route instructions using hierarchical reinforcement learning. In International Conference on Spatial Cognition, 2010.
  16. 16.Amanda Cercas Curry, Dimitra Gkatzia, and Verena Rieser. Generating and evaluating landmark-based navigation instructions in virtual environments. In European Workshop on Natural Language Generation, 2015.
  17. 17.Robert Dale, Sabine Geldof, and J Prost. Using natural language generation in automatic route. Journal of Research and Practice in Information Technology, 36(3):23, 2004.
  18. 18.Andrea F Daniele, Mohit Bansal, and Matthew R Walter. Navigational instruction generation as inverse reinforcement learning with neural machine translation. In International Conference on Human-Robot Interaction, 2017.
  19. 19.Zhiwei Deng, Karthik Narasimhan, and Olga Russakovsky. Evolving graphical planner: Contextual global planning for vision-and-language navigation. In NeurIPS, 2021.
  20. 20.Mary T Dzindolet, Scott A Peterson, Regina A Pomranky, Linda G Pierce, and Hall P Beck. The role of trust in automation reliance. International Journal of Human-Computer Studies, 58(6):697–718, 2003.
  21. 21.Kai Epstude and Neal J Roese. The functional theory of counterfactual thinking. Personality and Social Psychology Review, 12(2):168–192, 2008.
  22. 22.Mary Ellen Foster. Natural language generation for social robotics: opportunities and challenges. Philosophical Transactions of the Royal Society B, 374(1771):20180027, 2019.
  23. 23.Daniel Fried, Jacob Andreas, and Dan Klein. Unified pragmatic models for generating and following instructions. arXiv preprint arXiv:1711.04987, 2017.
  24. 24.Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. Speaker-follower models for vision-and-language navigation. In NeurIPS, 2018.
  25. 25.Tsu-Jui Fu, Xin Wang, Matthew Peterson, Scott Grafton, Miguel Eckstein, and William Yang Wang. Counterfactual vision-and-language navigation via adversarial path sampling. In ECCV, 2020.
  26. 26.Vittorio Girotto, Donatella Ferrante, Stefania Pighin, and Michel Gonzalez. Postdecisional counterfactual thinking by actors and readers. Psychological Science, 2007.
  27. 27.Robert Goeddel and Edwin Olson. Dart: A particle-based method for generating easy-to-follow directions. In International Conference on Intelligent Robots and Systems, 2012.
  28. 28.Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee. Counterfactual visual explanations. In ICML, 2019.
  29. 29.Scott A Green, Mark Billinghurst, XiaoQi Chen, and J Geoffrey Chase. Human-robot collaboration: A literature review and augmented reality approach in design. International Journal of Advanced Robotic Systems, 5(1):1, 2008.
  30. 30.Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao. Towards learning a generic agent for vision-and-language navigation via pre-training. In CVPR, 2020.
  31. 31.Di He, Yingce Xia, Tao Qin, Liwei Wang, Nenghai Yu, Tie-Yan Liu, and Wei-Ying Ma. Dual learning for machine translation. In NeurIPS, 2016.
  32. 32.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  33. 33.Lisa Anne Hendricks, Ronghang Hu, Trevor Darrell, and Zeynep Akata. Grounding visual explanations. In ECCV, 2018.
  34. 34.Yicong Hong, Cristian Rodriguez-Opazo, Yuankai Qi, Qi Wu, and Stephen Gould. Language and visual entity relationship graph for agent navigation. NeurIPS, 2020.
  35. 35.Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez-Opazo, and Stephen Gould. Vln bert: A recurrent vision-and-language bert for navigation. In CVPR, 2021.
  36. 36.Ronghang Hu, Daniel Fried, Anna Rohrbach, Dan Klein, Trevor Darrell, and Kate Saenko. Are you looking? grounding to multiple modalities in vision-and-language navigation. In ACL, 2019.
  37. 37.Haoshuo Huang, Vihan Jain, Harsh Mehta, Alexander Ku, Gabriel Magalhaes, Jason Baldridge, and Eugene Ie. Transferable representation learning in vision-and-language navigation. In ICCV, 2019.
  38. 38.Liyiming Ke, Xiujun Li, Yonatan Bisk, Ari Holtzman, Zhe Gan, Jingjing Liu, Jianfeng Gao, Yejin Choi, and Siddhartha Srinivasa. Tactical rewind: Self-correction via backtracking in vision-and-language navigation. In CVPR, 2019.
  39. 39.Taeksoo Kim, Moonsu Cha, Hyunsoo Kim, Jung Kwon Lee, and Jiwon Kim. Learning to discover cross-domain relations with generative adversarial networks. In ICML, 2017.
  40. 40.Benjamin Kuipers. Modeling spatial knowledge. Cognitive Science, 2(2):129–153, 1978.
  41. 41.Matt Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva. Counterfactual fairness. In NeurIPS, 2017.
  42. 42.Yikang Li, Nan Duan, Bolei Zhou, Xiao Chu, Wanli Ouyang, Xiaogang Wang, and Ming Zhou. Visual question generation as dual task of visual question answering. In CVPR, 2018.
  43. 43.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, 2004.
  44. 44.Gary Look, Buddhika Kottahachchi, Robert Laddaga, and Howard Shrobe. A location representation for generating descriptive walking directions. In Proceedings of the International Conference on Intelligent User Interfaces, 2005.
  45. 45.Kristin L Lovelace, Mary Hegarty, and Daniel R Montello. Elements of good route directions in familiar and unfamiliar environments. In International Conference on Spatial Information Theory, 1999.
  46. 46.Xiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai, David J Crandall, and Steven CH Hoi. Learning video object segmentation from unlabeled videos. In CVPR, 2020.
  47. 47.Fuli Luo, Peng Li, Jie Zhou, Pengcheng Yang, Baobao Chang, Zhifang Sui, and Xu Sun. A dual reinforcement learning framework for unsupervised text style transfer. In IJCAI, 2019.
  48. 48.Kevin Lynch. The Image of the City. The MIT Press, 1960.
  49. 49.Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, Zsolt Kira, Richard Socher, and Caiming Xiong. Self-monitoring navigation agent via auxiliary progress estimation. In ICLR, 2019.
  50. 50.Chih-Yao Ma, Zuxuan Wu, Ghassan AlRegib, Caiming Xiong, and Zsolt Kira. The regretful agent: Heuristic-aided navigation through progress estimation. In CVPR, 2019.
  51. 51.Matt MacMahon, Brian Stankiewicz, and Benjamin Kuipers. Walk the talk: connecting language, knowledge, and action in route instructions. In AAAI, 2006.
  52. 52.Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra. Improving vision-and-language navigation with image-text pairs from the web. In ECCV, 2020.
  53. 53.Hongyuan Mei, Mohit Bansal, and Matthew R Walter. Listen, attend, and walk: Neural mapping of navigational instructions to action sequences. In AAAI, 2016.
  54. 54.Dipendra Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, and Yoav Artzi. Mapping instructions to actions in 3d environments with visual goal prediction. In EMNLP, 2018.
  55. 55.Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In ICML, 2016.
  56. 56.Lawrence Neal, Matthew Olson, Xiaoli Fern, Weng-Keen Wong, and Fuxin Li. Open set learning with counterfactual images. In ECCV, 2018.
  57. 57.Stefan Oßwald, Henrik Kretzschmar, Wolfram Burgard, and Cyrill Stachniss. Learning to give route directions from human demonstrations. In ICRA, 2014.
  58. 58.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002.
  59. 59.Amin Parvaneh, Ehsan Abbasnejad, Damien Teney, Qinfeng Shi, and Anton van den Hengel. Counterfactual vision-and-language navigation: Unravelling the unseen. In NeurIPS, 2020.
  60. 60.Yuankai Qi, Zizheng Pan, Shengping Zhang, Anton van den Hengel, and Qi Wu. Object-and-action aware model for visual language navigation. In ECCV, 2020.
  61. 61.Kai-Florian Richter and Matt Duckham. Simplest instructions: Finding easy-to-describe routes for navigation. In International Conference on Geographic Information Science, 2008.
  62. 62.Neal J Roese. Counterfactual thinking. Psychological Bulletin, 121(1):133, 1997.
  63. 63.Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D Manning. Generating semantically precise scene graphs from textual descriptions for improved image retrieval. In VL, 2015.
  64. 64.Meet Shah, Xinlei Chen, Marcus Rohrbach, and Devi Parikh. Cycle-consistency for robust visual question answering. In CVPR, 2019.
  65. 65.Kristina Striegnitz, Alexandre Denis, Andrew Gargett, Konstantina Garoufi, Alexander Koller, and Mariët Theune. Report on the second challenge on generating instructions in virtual environments (give-2.5). In European Workshop on Natural Language Generation, 2011.
  66. 66.Hao Tan, Licheng Yu, and Mohit Bansal. Learning to navigate unseen environments: Back translation with environmental dropout. In NAACL, 2019.
  67. 67.Duyu Tang, Nan Duan, Tao Qin, Zhao Yan, and Ming Zhou. Question answering and question generation as dual tasks. arXiv preprint arXiv:1706.02027, 2017.
  68. 68.Stefanie Tellex, Thomas Kollar, Steven Dickerson, Matthew R Walter, Ashis Gopal Banerjee, Seth Teller, and Nicholas Roy. Understanding natural language commands for robotic navigation and mobile manipulation. In AAAI, 2011.
  69. 69.Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer. Vision-and-dialog navigation. In Conference on Robot Learning, 2020.
  70. 70.Eric J Vanetti and Gary L Allen. Communicating environmental knowledge: The impact of verbal and spatial abilities on the production and comprehension of route directions. Environment and Behavior, 20(6):667–682, 1988.
  71. 71.Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In CVPR, 2015.
  72. 72.David Waller and Yvonne Lippa. Landmarks as beacons and associative cues: their role in route learning. Memory & Cognition, 35(5):910–924, 2007.
  73. 73.Hanqing Wang, Wenguan Wang, Wei Liang, Caiming Xiong, and Jianbing Shen. Structured scene memory for vision-language navigation. In CVPR, 2021.
  74. 74.Hanqing Wang, Wenguan Wang, Tianmin Shu, Wei Liang, and Jianbing Shen. Active visual information gathering for vision-language navigation. In ECCV, 2020.
  75. 75.Haiyang Wang, Wenguan Wang, Xizhou Zhu, Jifeng Dai, and Liwei Wang. Collaborative visual navigation. arXiv preprint arXiv:2107.01151, 2021.
  76. 76.Jijun Wang and Michael Lewis. Human control for cooperating robot teams. In International Conference on Human-Robot Interaction, 2007.
  77. 77.Ning Wang, Yibing Song, Chao Ma, Wengang Zhou, Wei Liu, and Houqiang Li. Unsupervised deep tracking. In CVPR, 2019.
  78. 78.Tan Wang, Jianqiang Huang, Hanwang Zhang, and Qianru Sun. Visual commonsense r-cnn. In CVPR, 2020.
  79. 79.Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In CVPR, 2019.
  80. 80.Xin Wang, Vihan Jain, Eugene Ie, William Yang Wang, Zornitsa Kozareva, and Sujith Ravi. Environment-agnostic multitask learning for natural language grounded navigation. In ECCV, 2020.
  81. 81.Xin Wang, Wenhan Xiong, Hongmin Wang, and William Yang Wang. Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation. In ECCV, 2018.
  82. 82.Yiren Wang, Yingce Xia, Tianyu He, Fei Tian, Tao Qin, ChengXiang Zhai, and Tie-Yan Liu. Multi-agent dual learning. In ICLR, 2019.
  83. 83.Shawn L Ward, Nora Newcombe, and Willis F Overton. Turn left at the church, or three miles north: A study of direction giving and sex differences. Environment and Behavior, 18(2):192–213, 1986.
  84. 84.James Woodward. Psychological studies of causal and counterfactual reasoning. Understanding Counterfactuals, Understanding Causation, 2011.
  85. 85.Yingce Xia, Tao Qin, Wei Chen, Jiang Bian, Nenghai Yu, and Tie-Yan Liu. Dual supervised learning. In ICML, 2017.
  86. 86.Yi Yang, Yueting Zhuang, and Yunhe Pan. Multiple knowledge representation for big data artificial intelligence: framework, applications, and case studies. Frontiers of Information Technology & Electronic Engineering, 2021.
  87. 87.Zhongqi Yue, Tan Wang, Hanwang Zhang, Qianru Sun, and Xian-Sheng Hua. Counterfactual zero-shot and open-set visual recognition. In CVPR, 2021.
  88. 88.Ming Zhao, Peter Anderson, Vihan Jain, Su Wang, Alexander Ku, Jason Baldridge, and Eugene Ie. On the evaluation of vision-and-language navigation instructions. In Conference of the European Chapter of the Association for Computational Linguistics, 2021.
  89. 89.Zhibing Zhao, Yingce Xia, Tao Qin, Lirong Xia, and Tie-Yan Liu. Dual learning: Theoretical study and an algorithmic extension. In Asian Conference on Machine Learning, 2020.
  90. 90.Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang. Vision-language navigation with self-supervised auxiliary reasoning tasks. In CVPR, 2020.
  91. 91.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV, 2017.

Citation

MLA
Wang, H., et al. “Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation”. arXiv, 2022, http://arxiv.org/abs/2203.16586v1.
APA
Wang, H., Liang, W., Shen, J., Gool, L. V., & Wang, W. (2022). Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation. arXiv. http://arxiv.org/abs/2203.16586v1
Chicago
Wang, H., W. Liang, J. Shen, L. V. Gool, and W. Wang. 2022. “Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation”. arXiv. http://arxiv.org/abs/2203.16586v1.
Harvard
Wang, H. et al. (2022) “Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.16586v1.
Vancouver
1. Wang H, Liang W, Shen J, Gool LV, Wang W (2022) Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation. arXiv

BibTeX

@article{wang2022counterfactual,
  title = {Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation},
  author = {Wang, Hanqing and Liang, Wei and Shen, Jianbing and Gool, Luc Van and Wang, Wenguan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.16586v1},
  eprint = {2203.16586}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE