PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation

Jialu LiMohit Bansal

article2023NeurIPS110 citations

Proposes a text-conditioned recursive diffusion outpainting method to synthesize diverse 360-degree panoramic rooms, overcoming the training environment scarcity in vision-and-language navigation and setting new state-of-the-art performance across multiple benchmarks.

Listen

Building autonomous systems that navigate real-world indoor spaces based on natural language instructions—such as home service robots—requires agents that can generalize reliably to new, unseen locations. A critical bottleneck in this domain is the scarcity of photorealistic 3D training environments. Existing benchmarks rely heavily on a small set of around 60 captured room environments, as collecting and annotating high-quality 3D spatial data is labor-intensive and costly. This limited exposure causes navigation models to overfit to known environments and struggle when deployed in unfamiliar layouts.

The article evaluates a new generative method, named PANOGEN, designed to create practically unlimited, diverse panoramic environments from text descriptions to improve navigation agents' ability to generalize. The research demonstrates how synthetic visual data and machine-generated navigation instructions can be integrated into model pre-training and fine-tuning pipelines to enhance performance without requiring manual human annotations.

The authors implemented a multi-stage approach using state-of-the-art vision and language models. First, an automated captioning model generated textual descriptions for 36 discrete camera angles within existing Matterport3D environments. Next, a text-to-image diffusion model generated an initial room image, and a recursive outpainting technique expanded the field of view by rotating camera angles to produce seamless, 360-degree panoramas that preserve realistic object relationships and layouts. In total, 7,644 synthetic panoramas (comprising 275,184 individual images) were generated. The authors evaluated two training strategies on a top-performing navigation agent: pre-training using new navigation instructions generated by a multi-modal speaker model, and fine-tuning where a portion of the original visual trajectory is randomly replaced with synthetic panoramas.

The key findings show substantial performance gains across major benchmark datasets. On the Room-to-Room benchmark, incorporating synthetic environments set a new state-of-the-art test performance, improving navigation success rate by 2.7 percentage points and path-efficiency-weighted success rate by 1.9 to 2.9 percentage points. On the Cooperative Vision-and-Dialog Navigation benchmark, which uses dialogue instructions requiring commonsense room understanding, the approach increased goal progress by 1.59 meters on the test leaderboard—a 28.5% relative improvement over prior leading systems. Additionally, ablation experiments indicated that replacing 30% of trajectory observations during fine-tuning yielded the best results, and performance scaled continuously as the volume of generated panoramic scans increased.

These results demonstrate that generative image outpainting can effectively bypass the physical and financial bottlenecks of collecting real-world 3D environments. By synthesizing realistic visual variations while maintaining sensible room logic, developers can equip navigation agents with broader commonsense spatial understanding at a fraction of standard data-collection costs. This reduces the deployment risk and time needed to adapt autonomous systems to new residential or commercial facilities.

Organizations developing embodied navigation and robotics systems should consider adopting synthetic panoramic augmentation and pre-training with automated instruction generation. Teams implementing this technique should calibrate the visual replacement ratio carefully during fine-tuning, as replacing too high a proportion of viewpoints can degrade alignment with ground-truth instructions. Further exploratory pilots could evaluate scaling the volume of generated environments even higher and testing consistency across consecutive multi-step trajectories.

Confidence in these findings is high based on consistent validation across multiple standard benchmarks and testing leaderboards. However, certain limitations exist. The diffusion models utilized were general-purpose vision generators rather than models fine-tuned specifically on architectural and room imagery, and the approach did not strictly enforce geometric consistency between consecutive viewpoints across a multi-step trajectory. Practical deployments to physical robotic hardware will require verification beyond simulation to account for real-world sensor dynamics and physics.

arXiv: 2305.19195
Cover for PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation

Abstract

Vision-and-Language Navigation requires the agent to follow language instructions to navigate through 3D environments. One main challenge in Vision-and-Language Navigation is the limited availability of photorealistic training environments, which makes it hard to generalize to new and unseen environments. To address this problem, we propose PANOGEN, a generation method that can potentially create an infinite number of diverse panoramic environments conditioned on text. Specifically, we collect room descriptions by captioning the room images in existing Matterport3D environments, and leverage a state-of-the-art text-to-image diffusion model to generate the new panoramic environments. We use recursive outpainting over the generated images to create consistent 360-degree panorama views. Our new panoramic environments share similar semantic information with the original environments by conditioning on text descriptions, which ensures the co-occurrence of objects in the panorama follows human intuition, and creates enough diversity in room appearance and layout with image outpainting. Lastly, we explore two ways of utilizing PANOGEN in VLN pre-training and fine-tuning. We generate instructions for paths in our PANOGEN environments with a speaker built on a pre-trained vision-and-language model for VLN pre-training, and augment the visual observation with our panoramic environments during agents’ fine-tuning to avoid overfitting to seen environments. Empirically, learning with our PANOGEN environments achieves the new state-of-the-art on the Room-to-Room, Room-for-Room, and CVDN datasets. Besides, we find that pre-training with our PANOGEN speaker data is especially effective for CVDN, which has under-specified instructions and needs commonsense knowledge to reach the target. Lastly, we show that the agent can benefit from training with more generated panoramic environments, suggesting promising results for scaling up the PANOGEN environments to enhance agents’ generalization to unseen environments.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 PANOGEN: Generating Panoramic Environments
  • 3.1 Room Description Collection
  • 3.2 Generating Panorama with Text-Conditioned Image Outpainting
  • 4 Utilizing Panoramic Environments for VLN Training
  • 4.1 Problem Setup
  • 4.2 VLN Training Procedures
  • 4.3 Generating Instructions for Paths in Panoramic Environments
  • 4.4 Learning from Panoramic Environments during Fine-tuning
  • 5 Experimental Setup
  • 5.1 Dataset and Evaluation Metrics
  • 5.2 Implementation Details
  • 6 Results and Analysis
  • 6.1 Test Set Results
  • 6.2 Qualitative Analysis of Panoramic Environment
  • 6.3 Effectiveness of Speaker Data in Pre-training
  • 6.4 Effectiveness of Panorama Replacement in Fine-tuning
  • 6.5 Impact of the Number of Panoramic Environments
  • 6.6 Comparison with Other Environment Augmentation Approaches
  • 6.7 Quantitative Evaluation of Generated Speaker Data
  • 7 Conclusion
  • 8 Acknowledgement
  • References
  • Appendix
  • A Datasets
  • B Implementation Details
  • C Ablation Performance on REVERIE Dataset
  • D Improving Consistency Across Steps in Panorama Environment
  • E Qualitative Example for PanoGen Environment
  • F Qualitative Example for Alignment between Speaker Data and PanoGen Environment
  • G Limitations and Broader Impacts
  • H Licenses

Knowls

  1. Knowl 1 — Text-conditioned recursive panorama generation

    model/method

    PANOGEN generates new panoramic environments from textual room descriptions while preserving plausible object co-occurrence and spatial continuity. Each Matterport3D training panorama is discretized into 36 views formed by 3 elevation levels and 12 heading angles, and each view is captioned independently with BLIP-2-FlanT5-xxL. A Stable Diffusion v2.1 base model generates one zero-elevation view from its caption. PANOGEN then rotates the virtual camera right, upward, and downward and uses Stable Diffusion v1.5 inpainting to outpaint each unseen neighboring view, conditioning on both the existing generated image and the caption of the nearby target view. The process is repeated until all 36 views are generated and the views are stitched into one panorama. Starting from the zero-elevation view provides a useful central scene containing more salient room content, while conditioning on the already generated view allows objects that cross view boundaries to retain their appearance. Applying this process to the R2R training panoramas produced 7,644 generated panoramas stitched from 275,184 images.

  2. Knowl 2 — Instruction generation for generated environments

    model/method

    Because PANOGEN changes the layout and appearance of the original environment, an instruction written for an original path is not necessarily aligned with the corresponding generated panorama. The authors therefore train an mPLUG-base vision-language speaker to generate new instructions for paths observed in PANOGEN environments. Each panorama is simplified to the single view faced by the agent at each step; CLIP-ViT/B-16 encodes the views, whose image patches are flattened and concatenated as the speaker input. The speaker is fine-tuned on R2R instruction-path data and then generates 4,675 instruction-trajectory pairs using PANOGEN observations. The generated pairs are added to R2R and Prevalent data during VLN pre-training. The mPLUG-generated instructions have CLIP image-text cosine similarity 0.2714 with PANOGEN trajectories, compared with 0.2669 for instructions generated by the EnvDrop speaker, and achieve BERTScore 71.8 against reference instructions versus 70.5 for EnvDrop. Replacing 30% of original trajectory views with PANOGEN views while retaining the original instructions yields similarity 0.2893 ± 0.0001, compared with 0.2845 for original instructions and original views, suggesting that moderate replacement does not reduce instruction-environment alignment.

  3. Knowl 3 — Panorama replacement during VLN fine-tuning

    model/method

    PANOGEN can be used directly as visual augmentation during VLN fine-tuning without generating additional instructions. At randomly selected time steps, the original panoramic observation in a trajectory is replaced by a PANOGEN panorama, while the original language instruction and trajectory supervision are retained. The generated panorama is conditioned on room captions, so its semantic content remains related to the original observation even though its appearance and layout differ. The replacement probability must remain moderate: excessive replacement can make the visual trajectory inconsistent with the instruction, whereas partial replacement exposes the agent to diverse environments and reduces overfitting to the limited Matterport3D training rooms.

  4. Knowl 4 — VLN training configuration used with PANOGEN

    experimental setup

    The navigation agent is based on the DUET dual-scale graph transformer, with CLIP-ViT/B-16 visual features rather than DUET's original ImageNet-pretrained ViT-B/16 features. During pre-training, the agent uses masked language modeling, instruction-trajectory matching, and single-action prediction. Instruction-trajectory matching uses one positive pair and four negatives: two other trajectories from the batch and two versions with shuffled observation order. During fine-tuning, the agent follows pseudo-interactive demonstrations in which trajectories are sampled according to the current policy. Pre-training uses batch size 64 for 150,000 iterations and fine-tuning uses batch size 8 for 40,000 iterations. The experiments cover R2R, R4R, and CVDN; the training split contains 61 environments, while unseen validation and test splits contain 11 and 18 environments, respectively. R2R and R4R emphasize instruction following and use success rate and SPL as primary metrics, whereas CVDN contains under-specified dialogue instructions and uses goal progress as its primary metric.

  5. Knowl 5 — State-of-the-art navigation results from PANOGEN environments

    empirical result

    Training with PANOGEN environments improves generalization to unseen rooms on the R2R and CVDN benchmarks. On R2R validation-unseen, PANOGEN obtains trajectory length 13.40, navigation error 3.03 m, success rate 74.2, and SPL 64.3. On the R2R test set, it obtains trajectory length 14.38, navigation error 3.31 m, success rate 71.7, and SPL 61.9. Relative to DUET on the R2R test leaderboard, these results improve success rate by 2.7 percentage points and SPL by 2.9 percentage points. On CVDN, PANOGEN obtains goal progress 5.93 on validation and 7.17 on test; the test result is 1.59 m higher than the previous reported best of 5.58, a relative gain of 28.5%. The authors attribute the particularly large CVDN improvement to the diverse commonsense room configurations learned from PANOGEN, which help with under-specified dialogue instructions.

  6. Knowl 6 — Effect of PANOGEN speaker data during pre-training

    empirical result

    On the R2R validation-unseen split, the DUET-CLIP baseline obtains trajectory length 12.92, navigation error 3.19 m, success rate 72.84, and SPL 63.37. Adding only PANOGEN environments without newly generated instructions gives 14.21, 2.99 m, 73.35, and 62.12, respectively. Adding EnvDrop-generated instructions gives 13.57, 3.05 m, 73.69, and 63.44. Adding the mPLUG-generated PANOGEN instructions gives 14.58, 2.85 m, 74.20, and 62.81. Thus, the mPLUG speaker improves R2R success rate over DUET-CLIP by 1.36 percentage points and improves navigation error by 0.34 m according to the reported metrics. On CVDN validation-unseen, the mPLUG variant obtains trajectory length 24.66 and goal progress 5.93, compared with 24.09 and 5.50 for DUET-CLIP; on R4R it obtains trajectory length 18.32, navigation error 6.12 m, success rate 45.78, and SPL 42.52, with an SPL improvement of 0.58 over the baseline. The results support using generated instructions, rather than generated environments alone, during pre-training.

  7. Knowl 7 — Optimal proportion of replaced observations

    empirical result

    On R2R validation-unseen, the DUET-CLIP model without PANOGEN replacement obtains trajectory length 12.92, navigation error 3.19 m, success rate 72.84, and SPL 63.37. Replacing 10% of trajectory viewpoints gives 13.16, 3.16 m, 72.84, and 63.24; replacing 30% gives 13.76, 2.99 m, 74.41, and 63.88; replacing 50% gives 13.03, 3.19 m, 72.84, and 63.84; and replacing 70% gives 12.62, 3.18 m, 72.33, and 63.93. The 30% setting provides the best success rate and the best combination of navigation error and SPL. The authors interpret the weaker results at larger replacement ratios as evidence that too many generated observations break alignment between the instruction and the visual trajectory. At the 30% setting, PANOGEN replacement also improves R2R success rate by 1.57 percentage points, CVDN goal progress by 0.13 m, and R4R SPL by 2.31 percentage points relative to DUET-CLIP.

  8. Knowl 8 — Performance scales with the number of generated environments

    empirical result

    Increasing the number of PANOGEN environments used during fine-tuning generally improves R2R validation-unseen performance. With 0 generated scans, the model obtains trajectory length 12.92, navigation error 3.19 m, success rate 72.84, and SPL 63.37. Using 10 scans gives 13.94, 3.00 m, 72.80, and 62.48; 30 scans gives 13.86, 3.05 m, 73.69, and 62.88; 61 scans gives 13.76, 2.99 m, 74.41, and 63.88; and 122 scans gives 13.40, 3.03 m, 74.20, and 64.27. The SPL increase from 61 to 122 scans is not saturated, supporting the authors' claim that generating more panoramic environments could further improve generalization.

  9. Knowl 9 — Comparison with existing environment augmentation methods

    empirical result

    On R2R validation-unseen, PANOGEN outperforms two environment-level augmentation baselines. EnvEdit, which changes object appearance and mixes edited examples into training, obtains trajectory length 13.61, navigation error 3.03 m, success rate 72.80, and SPL 63.17. EnvDrop, which applies dropout at the environment level during fine-tuning, obtains 13.28, 3.12 m, 72.58, and 62.40. PANOGEN obtains 13.40, 3.03 m, 74.20, and 64.27. Thus PANOGEN gives the highest success rate and SPL among the compared augmentation approaches while maintaining the same navigation error as EnvEdit.

  10. Knowl 10 — Stated limitations of PANOGEN

    limitation

    The panorama generator directly uses Stable Diffusion models trained for general image inpainting on the LAION-Aesthetics v2 5+ data rather than models further adapted to indoor room imagery; the authors therefore note that room-specific training could improve generation quality. The experiments focus on Vision-and-Language Navigation, although the same text-conditioned panoramic generation approach may also apply to other embodied tasks such as concept learning and grounding. The method also does not explicitly enforce consistency between generated views at different navigation steps, so temporal or cross-step consistency remains an incompletely addressed limitation.

Coverage note — Secondary appendix material was omitted from the top ten, including the REVERIE ablation and the exploratory cross-navigation-step consistency method, because these provide additional validation rather than core PANOGEN generation or training contributions.

References

  1. 1.P. Anderson, A. Chang, D. S. Chaplot, A. Dosovitskiy, S. Gupta, V. Koltun, J. Kosecka, J. Malik, R. Mottaghi, M. Savva, et al. On evaluation of embodied navigation agents. arXiv preprint arXiv:1807.06757, 2018.
  2. 2.P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3674–3683, 2018.
  3. 3.Y. Balaji, S. Nah, X. Huang, A. Vahdat, J. Song, K. Kreis, M. Aittala, T. Aila, S. Laine, B. Catanzaro, et al. ediffi: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022.
  4. 4.O. Bar-Tal, L. Yariv, Y. Lipman, and T. Dekel. Multidiffusion: Fusing diffusion paths for controlled image generation. arXiv preprint arXiv:2302.08113, 2, 2023.
  5. 5.T. Brooks, A. Holynski, and A. A. Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18392–18402, 2023.
  6. 6.M. Cha, Y. Gwon, and H. Kung. Adversarial nets with perceptual losses for text-to-image synthesis. In 2017 IEEE 27th international workshop on machine learning for signal processing (MLSP), pages 1–6. IEEE, 2017.
  7. 7.A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang. Matterport3d: Learning from rgb-d data in indoor environments. arXiv preprint arXiv:1709.06158, 2017.
  8. 8.H. Chen, A. Suhr, D. Misra, N. Snavely, and Y. Artzi. Touchdown: Natural language navigation and spatial reasoning in visual street environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12538–12547, 2019.
  9. 9.J. Chen, C. Gao, E. Meng, Q. Zhang, and S. Liu. Reinforced structured state-evolution for vision-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15450–15459, 2022.
  10. 10.S. Chen, P.-L. Guhur, C. Schmid, and I. Laptev. History aware multimodal transformer for vision-and-language navigation. Advances in neural information processing systems, 34:5834–5847, 2021.
  11. 11.S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev. Learning from unlabeled 3d environments for vision-and-language navigation. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXIX, pages 638–655. Springer, 2022.
  12. 12.S. Chen, P.-L. Guhur, M. Tapaswi, C. Schmid, and I. Laptev. Think global, act local: Dual-scale graph transformer for vision-and-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16537–16547, 2022.
  13. 13.A. Dosovitskiy, L. Beyer, A. Kolesnikov, L. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  14. 14.D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell. Speaker-follower models for vision-and-language navigation. Advances in Neural Information Processing Systems, 31, 2018.
  15. 15.C. Gao, X. Peng, M. Yan, H. Wang, L. Yang, H. Ren, H. Li, and S. Liu. Adaptive zone-aware hierarchical planner for vision-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14911–14920, 2023.
  16. 16.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial networks. Communications of the ACM, 63(11):139–144, 2020.
  17. 17.I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville. Improved training of wasserstein gans. Advances in neural information processing systems, 30, 2017.
  18. 18.W. Hao, C. Li, X. Li, L. Carin, and J. Gao. Towards learning a generic agent for vision-and-language navigation via pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13137–13146, 2020.
  19. 19.Y. Hong, C. Rodriguez-Opazo, Q. Wu, and S. Gould. Sub-instruction aware vision-and-language navigation. arXiv preprint arXiv:2004.02707, 2020.
  20. 20.Y. Hong, Q. Wu, Y. Qi, C. Rodriguez-Opazo, and S. Gould. A recurrent vision-and-language bert for navigation. arXiv preprint arXiv:2011.13922, 2020.
  21. 21.V. Jain, G. Magalhaes, A. Ku, A. Vaswani, E. Ie, and J. Baldridge. Stay on the path: Instruction fidelity in vision-and-language navigation. arXiv preprint arXiv:1905.12255, 2019.
  22. 22.M. Kang, J.-Y. Zhu, R. Zhang, J. Park, E. Shechtman, S. Paris, and T. Park. Scaling up gans for text-to-image synthesis. arXiv preprint arXiv:2303.05511, 2023.
  23. 23.H. Kim, J. Li, and M. Bansal. Ndh-full: Learning and evaluating navigational agents on full-length dialogue. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021.
  24. 24.J. Y. Koh, H. Lee, Y. Yang, J. Baldridge, and P. Anderson. Pathdreamer: A world model for indoor navigation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14738–14748, 2021.
  25. 25.J. Y. Koh, H. Agrawal, D. Batra, R. Tucker, A. Waters, H. Lee, Y. Yang, J. Baldridge, and P. Anderson. Simple and effective synthesis of indoor 3d scenes. arXiv preprint arXiv:2204.02960, 2022.
  26. 26.A. Ku, P. Anderson, R. Patel, E. Ie, and J. Baldridge. Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. arXiv preprint arXiv:2010.07954, 2020.
  27. 27.C. Li, H. Xu, J. Tian, W. Wang, M. Yan, B. Bi, J. Ye, H. Chen, G. Xu, Z. Cao, et al. mplug: Effective and efficient vision-language learning by cross-modal skip-connections. arXiv preprint arXiv:2205.12005, 2022.
  28. 28.J. Li and M. Bansal. Improving vision-and-language navigation by generating future-view image semantics. arXiv preprint arXiv:2304.04907, 2023.
  29. 29.J. Li, H. Tan, and M. Bansal. Improving cross-modal alignment in vision language navigation via syntactic information. arXiv preprint arXiv:2104.09580, 2021.
  30. 30.J. Li, H. Tan, and M. Bansal. Clear: Improving vision-language navigation with cross-lingual, environment-agnostic representations. arXiv preprint arXiv:2207.02185, 2022.
  31. 31.J. Li, H. Tan, and M. Bansal. Envedit: Environment editing for vision-and-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15407–15417, 2022.
  32. 32.J. Li, D. Li, S. Savarese, and S. Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023.
  33. 33.X. Li, C. Li, Q. Xia, Y. Bisk, A. Celikyilmaz, A. Gao, N. Smith, and Y. Choi. Robust navigation with language pretraining and stochastic sampling. arXiv preprint arXiv:1909.02244, 2019.
  34. 34.C. Liu, F. Zhu, X. Chang, X. Liang, Z. Ge, and Y.-D. Shen. Vision-language navigation with random environmental mixup. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1644–1654, 2021.
  35. 35.C.-Y. Ma, J. Lu, Z. Wu, G. AlRegib, Z. Kira, R. Socher, and C. Xiong. Self-monitoring navigation agent via auxiliary progress estimation. arXiv preprint arXiv:1901.03035, 2019.
  36. 36.A. Majumdar, A. Shrivastava, S. Lee, P. Anderson, D. Parikh, and D. Batra. Improving vision-and-language navigation with image-text pairs from the web. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16, pages 259–274. Springer, 2020.
  37. 37.L. Mescheder. On the convergence properties of gan training. arXiv preprint arXiv:1801.04406, 1:16, 2018.
  38. 38.L. Metz, B. Poole, D. Pfau, and J. Sohl-Dickstein. Unrolled generative adversarial networks. arXiv preprint arXiv:1611.02163, 2016.
  39. 39.K. Nguyen and H. Daumé III. Help, anna! visual navigation with natural multimodal assistance via retrospective curiosity-encouraging imitation learning. arXiv preprint arXiv:1909.01871, 2019.
  40. 40.A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  41. 41.A. Padmakumar, J. Thomason, A. Shrivastava, P. Lange, A. Narayan-Chen, S. Gella, R. Piramuthu, G. Tur, and D. Hakkani-Tur. Teach: Task-driven embodied agents that chat. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2017–2025, 2022.
  42. 42.T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2337–2346, 2019.
  43. 43.Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. v. d. Hengel. Reverie: Remote embodied visual referring expression in real indoor environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9982–9991, 2020.
  44. 44.Y. Qiao, Y. Qi, Y. Hong, Z. Yu, P. Wang, and Q. Wu. Hop: history-and-order aware pre-training for vision-and-language navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15418–15427, 2022.
  45. 45.A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, A. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  46. 46.S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. M. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, M. Savva, Y. Zhao, and D. Batra. Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/forum?id=-v4OuqNs5P.
  47. 47.R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, June 2022.
  48. 48.C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 35:36479–36494, 2022.
  49. 49.M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10740–10749, 2020.
  50. 50.L. Song, L. Cao, H. Xu, K. Kang, F. Tang, J. Yuan, and Y. Zhao. Roomdreamer: Text-driven 3d indoor scene synthesis with coherent geometry and texture. arXiv preprint arXiv:2305.11337, 2023.
  51. 51.H. Tan, L. Yu, and M. Bansal. Learning to navigate unseen environments: Back translation with environmental dropout. arXiv preprint arXiv:1904.04195, 2019.
  52. 52.J. Thomason, M. Murray, M. Cakmak, and L. Zettlemoyer. Vision-and-dialog navigation. In Conference on Robot Learning, pages 394–406. PMLR, 2020.
  53. 53.H. Wang, W. Wang, W. Liang, C. Xiong, and J. Shen. Structured scene memory for vision-language navigation. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition, pages 8455–8464, 2021.
  54. 54.S. Wang, C. Montgomery, J. Orbay, V. Birodkar, J. Faust, I. Gur, N. Jaques, A. Waters, J. Baldridge, and P. Anderson. Less is more: Generating grounded navigation instructions from landmarks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15428–15438, 2022.
  55. 55.X. Wang, Q. Huang, A. Celikyilmaz, J. Gao, D. Shen, Y.-F. Wang, W. Y. Wang, and L. Zhang. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6629–6638, 2019.
  56. 56.X. Wang, W. Wang, J. Shao, and Y. Yang. Lana: A language-capable navigator for instruction following and generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19048–19058, 2023.
  57. 57.X. E. Wang, V. Jain, E. Ie, W. Y. Wang, Z. Kozareva, and S. Ravi. Environment-agnostic multitask learning for natural language grounded navigation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIV 16, pages 413–430. Springer, 2020.
  58. 58.L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934, 2020.
  59. 59.Z. Yang, J. Dong, P. Liu, Y. Yang, and S. Yan. Very long natural scenery image prediction by outpainting. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10561–10570, 2019.
  60. 60.H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In Proceedings of the IEEE international conference on computer vision, pages 5907–5915, 2017.
  61. 61.H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas. Stackgan++: Realistic image synthesis with stacked generative adversarial networks. IEEE transactions on pattern analysis and machine intelligence, 41(8):1947–1962, 2018.
  62. 62.L. Zhang, Q. Chen, B. Hu, and S. Jiang. Text-guided neural image inpainting. In Proceedings of the 28th ACM International Conference on Multimedia, pages 1302–1310, 2020.
  63. 63.Y. Zhao, J. Chen, C. Gao, W. Wang, L. Yang, H. Ren, H. Xia, and S. Liu. Target-driven structured transformer planner for vision-language navigation. In Proceedings of the 30th ACM International Conference on Multimedia, pages 4194–4203, 2022.

Citation

MLA
Li, J., and M. Bansal. “PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 21878–94, https://proceedings.neurips.cc/paper_files/paper/2023/file/4522de4178bddb36b49aa26efad537cf-Paper-Conference.pdf.
APA
Li, J., & Bansal, M. (2023). PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation. Advances in Neural Information Processing Systems, 36, 21878–21894. https://proceedings.neurips.cc/paper_files/paper/2023/file/4522de4178bddb36b49aa26efad537cf-Paper-Conference.pdf
Chicago
Li, J., and M. Bansal. 2023. “PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation”. Advances in Neural Information Processing Systems 36: 21878–94. https://proceedings.neurips.cc/paper_files/paper/2023/file/4522de4178bddb36b49aa26efad537cf-Paper-Conference.pdf.
Harvard
Li, J. and Bansal, M. (2023) “PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 21878–21894. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/4522de4178bddb36b49aa26efad537cf-Paper-Conference.pdf.
Vancouver
1. Li J, Bansal M (2023) PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 21878–21894

BibTeX

@inproceedings{li2023panogen,
  title = {PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation},
  author = {Li, Jialu and Bansal, Mohit},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {21878-21894},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/4522de4178bddb36b49aa26efad537cf-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors