Learning to Model the World With Language

Jessy LinYuqing DuOlivia WatkinsDanijar HafnerPieter AbbeelDan KleinAnca D. Dragan

article2024ICML91 citations

Introduces Dynalang, an embodied agent that grounds diverse language like instructions, manuals, and environment descriptions by predicting future multimodal representations within a world model to plan actions from imagined rollouts.

Listen

Autonomous artificial intelligence agents operating in physical and simulated spaces must interpret varied human language to collaborate effectively. Standard reinforcement learning approaches typically limit language to direct commands, such as basic task instructions, and directly map those prompts to specific physical actions. However, human communication encompasses diverse expressions, including explanations of environmental dynamics, world state updates, and real-time corrections. Standard methods struggle to process this richer language because indirect statements often exhibit complex, weak statistical correlations with immediate optimal actions.

The article demonstrates that treating multimodal language understanding as a future prediction task enables embodied agents to utilize diverse linguistic inputs. It introduces and evaluates Dynalang, an agent framework that decouples learning a predictive generative world model from learning an action policy, allowing diverse text to ground naturally into visual predictions and future reward expectations.

To test this approach, the authors developed a multimodal world model that integrates visual frames and text tokens at each time step into compressed latent representations. The architecture utilizes a self-supervised objective to predict future latent states, reconstruct sensory inputs, and anticipate rewards. Action policies are subsequently trained entirely within the model’s imagined rollouts. Dynalang was evaluated across four distinct simulated settings: HomeGrid (a home-chore environment featuring diverse hints), Messenger (a multi-stage reasoning game with rule manuals), Habitat Vision-Language Navigation in Continuous Environments (navigating photorealistic home scans), and LangRoom (an embodied question-answering environment), alongside comparisons against standard model-free algorithms such as IMPALA and R2D2.

The evaluation revealed several critical findings. First, Dynalang effectively translated diverse language hints—including dynamics rules, future observation hints, and corrections—into substantial task performance gains in HomeGrid, whereas baseline models degraded in performance when exposed to diverse language. Second, on the Messenger benchmark, Dynalang successfully solved the most complex third stage using multi-hop reasoning over game manuals, while standard baselines and domain-specialized architectures failed completely. Third, in continuous vision-language navigation, Dynalang achieved a success rate near 30% from scratch, substantially outperforming model-free baselines. Fourth, by regularizing language actions with the world model's internal predictions, Dynalang scaled embodied question-answering to a 10,000-token vocabulary, matching small-vocabulary performance where unregularized policies failed. Finally, pretraining the world model on 500 million tokens of text-only data without actions or rewards significantly accelerated downstream task learning, outperforming models that relied on frozen external text representations.

These findings indicate that unifying language and visual reasoning under a self-supervised future prediction objective resolves a core bottleneck in embodied intelligence. Rather than requiring expensive, task-specific paired demonstration datasets, agents can leverage abundant offline text data and autonomously ground diverse language through online experience. This decoupling reduces the engineering complexity of multimodal training while improving policy robustness across varied operational conditions.

Organizations developing embodied AI should adopt generative multimodal world modeling frameworks when designing agents intended for dynamic, human-centric environments. System designers should prioritize token-level streaming inputs over static sentence-level conditioning and leverage offline domain-specific or general text corpora to pretrain world models before online deployment. Additional development is needed to close the remaining performance gap between reinforcement learning world models and specialized, demonstration-heavy navigation architectures in high-fidelity settings.

The findings are established across multiple varied simulation environments, providing high confidence in the core conceptual approach. However, users should note key boundaries: the experimental evaluations were conducted entirely in simulated environments rather than physical hardware, and performance on complex continuous navigation still trails specialized systems trained on human demonstrations. Physical deployments should await further validation on real-world robotic platforms.

Cover for Learning to Model the World With Language

Abstract

To interact with humans and act in the world, agents need to understand the range of language that people use and relate it to the visual world. While current agents can learn to execute simple language instructions, we aim to build agents that leverage diverse language—language like “this button turns on the TV” or “I put the bowls away”—that conveys general knowledge, describes the state of the world, provides interactive feedback, and more. Our key idea is that agents should interpret such diverse language as a signal that helps them predict the future: what they will observe, how the world will behave, and which situations will be rewarded. This perspective unifies language understanding with future prediction as a powerful self-supervised learning objective. We instantiate this in Dynalang, an agent that learns a multimodal world model to predict future text and image representations, and learns to act from imagined model rollouts. While current methods that learn language-conditioned policies degrade in performance with more diverse types of language, we show that Dynalang learns to leverage environment descriptions, game rules, and instructions to excel on tasks ranging from game-playing to navigating photorealistic home scans. Finally, we show that our method enables additional capabilities due to learning a generative model: Dynalang can be pretrained on text-only data, enabling learning from offline datasets, and generate language grounded in an environment.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Dynalang
  • 3.1 World Model Learning
  • 3.2 Policy Learning
  • 4 Experiments
  • 4.1 Aligning Language, Vision, and Action in a World Model
  • 4.2 HomeGrid: Language Hints
  • 4.3 Messenger: Game Manuals
  • 4.4 Vision-Language Navigation: Instruction Following
  • 4.5 LangRoom: Embodied Question Answering
  • 4.6 Text-only Pretraining
  • 5 Conclusion
  • Acknowledgments
  • Impact Statement
  • References
  • A World Model Learning
  • B Actor Critic Learning
  • C Detailed Related Work
  • D Environment Details
  • D.1 HomeGrid
  • D.2 Messenger
  • D.3 VLN-CE
  • D.4 LangRoom
  • E Text Pretraining: Text Generation Samples
  • F Qualitative Analysis
  • G HomeGrid Training Curves
  • H Additional Baseline Experiments
  • H.1 Token vs. Sentence Embeddings for Baselines
  • H.2 Model Scaling for Baselines
  • H.3 Auxiliary Reconstruction Loss for Baselines
  • I Model and Training Details
  • I.1 Baseline Hyperparameters
  • I.2 Dynalang Hyperparameters

Knowls

  1. Knowl 1 — Future prediction grounds diverse language for embodied action

    model/method

    Dynalang treats language as an observation modality that helps an agent predict what will happen, rather than mapping each utterance directly to an action. From its history of visual and textual observations and its own actions, the agent learns a generative world model that predicts future observations and rewards. A separate policy learns to act using imagined predictions from that model and task rewards. The paper’s central proposal is that predicting how language relates to future visual states, world dynamics, and rewards provides a self-supervised signal for grounding descriptions, prior knowledge, corrections, and instructions in experience.

  2. Knowl 2 — Dynalang’s multimodal world model and learning objectives

    equation

    At each time step, the world model combines an image xtx_t, a language token ltl_t, and recurrent state hth_t into a stochastic discrete latent representation ztz_t. Its recurrent sequence model uses the previous latent, recurrent state, and action at−1a_{t-1} to predict a prior distribution over the next latent; a decoder uses ztz_t and hth_t to reconstruct the image and language and predict reward and episode continuation. A GRU implements the recurrent model, with a strided CNN image encoder and decoder and MLPs for other components.

    Let qt=qϕ(zt∣xt,lt,ht)q_t=q_\phi(z_t\mid x_t,l_t,h_t) denote the encoder’s posterior distribution and pt=pθ(zt∣ht)p_t=p_\theta(z_t\mid h_t) the sequence model’s predictive prior. The representation objective is the sum of image reconstruction, language reconstruction, reward prediction, continuation prediction, and a KL regularizer: Lrepr=∥x^t−xt∥22+catxent⁡(l^t,lt)+catxent⁡(r^t,twohot⁡(rt))+binxent⁡(c^t,ct)+βregmax⁡(1,KL⁡(qt∥sg⁡(pt)))\mathcal{L}_{\mathrm{repr}}=\|\hat{x}_t-x_t\|_2^2+\operatorname{catxent}(\hat{l}_t,l_t)+\operatorname{catxent}(\hat{r}_t,\operatorname{twohot}(r_t))+\operatorname{binxent}(\hat{c}_t,c_t)+\beta_{\mathrm{reg}}\max(1,\operatorname{KL}(q_t\|\operatorname{sg}(p_t))). Here, hats denote decoder predictions, rtr_t is the environment reward, ctc_t indicates whether the episode continues, and sg⁡\operatorname{sg} stops gradients through its argument. The future-prediction objective is Lpred=βpredmax⁡(1,KL⁡(sg⁡(qt)∥pt))\mathcal{L}_{\mathrm{pred}}=\beta_{\mathrm{pred}}\max(1,\operatorname{KL}(\operatorname{sg}(q_t)\|p_t)). The model minimizes Lrepr+Lpred\mathcal{L}_{\mathrm{repr}}+\mathcal{L}_{\mathrm{pred}}; the paper uses βreg=0.1\beta_{\mathrm{reg}}=0.1 and βpred=0.5\beta_{\mathrm{pred}}=0.5. One-hot language is reconstructed with cross-entropy; pretrained language embeddings are reconstructed with squared error.

  3. Knowl 3 — Actions are learned from imagined latent rollouts

    model/method

    Dynalang trains an actor and critic on sequences imagined by the learned world model, rather than requiring the actor to learn only from real transitions. For each training batch, imagined rollouts of length T=15T=15 begin at latent representations computed from replay data; the actor samples actions, and the world model predicts subsequent latent states, rewards, and episode-continuation flags. The critic is trained by categorical regression toward two-hot encoded λ\lambda-return targets, and the actor is trained to favor actions with high returns while retaining an entropy regularizer. For an imagined trajectory, the return target is Rt=rt+γct((1−λ)Vt+1+λRt+1)R_t=r_t+\gamma c_t((1-\lambda)V_{t+1}+\lambda R_{t+1}), with terminal bootstrap RT=VTR_T=V_T, where rtr_t and ctc_t are imagined reward and continuation predictions, VtV_t is the critic estimate, γ\gamma is the reward discount, and λ\lambda controls the mixture of one-step bootstrap and later returns. The actor’s return advantage is normalized using an exponential moving average of the 5th-to-95th percentile return range. The policy and critic are MLPs.

  4. Knowl 4 — Language and images need not be temporally aligned

    model/method

    Dynalang consumes one image frame and one language token at each time step, using zero or padding inputs when a modality is absent, but it does not require a token to be semantically aligned with the frame at that time. The recurrent state can retain information from earlier inputs, so future prediction can use language and visual evidence even when they arrived at different times. This permits continuous, token-by-token language input while the agent acts and avoids the need for explicit temporal segmentation or language–image pairing.

  5. Knowl 5 — A simple per-timestep fusion design beat tested alternatives

    empirical result

    On Messenger Stage 1, the authors compared Dynalang with several ways of adding language to DreamerV3: conditioning the policy on a GRU encoding of the manual while leaving the world model language-free; feeding SentenceBERT sentence embeddings one sentence at a time; and fusing image and manual representations using a T5 image adapter, T5 cross-attention, fine-tuned T5 cross-attention, or two-way cross-attention. Dynalang achieved higher training scores than these alternatives through the 500,000-environment-step comparison, including approaches using pretrained T5, despite learning token representations from scratch. The result supports the paper’s choice to feed language token-by-token into the multimodal world model alongside visual inputs.

  6. Knowl 6 — HomeGrid tests grounding of state, dynamics, and correction hints

    empirical result

    HomeGrid is a partially observed visual gridworld with 38 tasks across five task types: find, get, clean up, rearrange, and open. Object and bin locations and bin-opening dynamics are randomized, and objects can move during an episode. Agents receive task instructions and, depending on the evaluation condition, token-by-token hints about future observations (such as an object’s room), environment dynamics (the correct action to open a bin), or corrections (such as “no, turn around”). Hints do not directly supervise the meaning of the utterance; without them, agents can in principle discover the same information through interaction. After 50 million environment steps, with two seeds, Dynalang scores higher with each hint type than with task instructions alone. IMPALA struggles to learn the tasks, while R2D2 can use task information and corrections but loses performance as language becomes more diverse; Dynalang improves with the additional language. Dynalang also achieves nontrivial performance with task-only language.

  7. Knowl 7 — Reading game manuals supports performance on Messenger

    empirical result

    In Messenger, agents receive a text manual describing randomized entity roles and movement dynamics and must retrieve and deliver a message while avoiding enemies. Success requires combining manual rules with visual observations of entity identities and behavior; the benchmark has three stages of increasing difficulty. Across the training comparisons, Dynalang outperformed language-conditioned IMPALA and R2D2 and the task-specific EMMA architecture. On the most difficult Stage 3, Dynalang learned nontrivial performance while the other compared methods failed to do so. Manuals are provided token-by-token before the episode, and Messenger’s symbolic observations and human-written templates require reasoning across visual and textual inputs.

  8. Knowl 8 — Text-only pretraining improves downstream Messenger learning

    empirical result

    The Dynalang world model can be pretrained without actions, rewards, or images by zeroing image and action inputs and disabling image, reward, and continuation decoder losses; it learns text representations and dynamics through its latent future-prediction objective rather than an explicit next-token loss. The paper tested pretraining on Messenger Stage 2 manuals and on TinyStories, a general-domain corpus of 2 million short stories containing approximately 500 million tokens. Downstream Stage 2 training began from scratch for each method rather than initializing from Stage 1. Even a small amount of in-domain pretraining closed much of the performance gap between learned one-hot token embeddings and pretrained T5 embeddings. TinyStories pretraining exceeded the final performance of the T5-embedding condition, which the authors suggest may reflect learning text dynamics offline rather than during environment interaction.

  9. Knowl 9 — Language generation enables embodied question answering in LangRoom

    empirical result

    LangRoom tests whether Dynalang can generate language grounded in visual information. The agent sees a partially observable room with four objects at fixed positions and randomized colors; it receives questions such as “what color is the ball?” and must move to the object to observe its color, then output the correct color token. The action space includes movement and language-token actions. Dynalang learned to take information-gathering actions and answer more questions accurately from task reward. With a 15-token language-action vocabulary, performance was comparable to using a 10,000-token vocabulary when the language actions were regularized toward the world model’s predicted next-token distribution: the policy distribution over language actions is KL-regularized against the stopped-gradient model prediction. Increasing the vocabulary to 10,000 without this prior failed to learn. This is a proof of concept for grounded language generation in a constrained QA task, not a demonstration of fluent open-ended speech.

  10. Knowl 10 — Dynalang learns low-level instruction following in photorealistic homes

    empirical result

    The paper evaluated Dynalang on VLN-CE, where an agent must follow a natural-language navigation instruction through Matterport3D home scenes using low-level discrete movement actions and must explicitly stop at the goal. The training set contains 10,819 unique instructions across 61 scenes; episodes sample an instruction and corresponding scene, and rewards include dense feedback based on distance to the goal plus success or penalty for stopping. Across training, Dynalang attained a higher instruction success rate than the model-free R2D2 baseline (Dynalang averaged over three seeds and R2D2 over two). The authors emphasize that, although Dynalang learned to ground instructions from scratch, its performance was not competitive with state-of-the-art VLN methods, many of which use expert demonstrations or navigation-specialized architectures.

Coverage note — No substantial contributed result is omitted; detailed per-environment hyperparameter tables and supplementary language-generation samples are left out because they provide implementation or qualitative supporting detail rather than distinct load-bearing contributions.

References

  1. 1.Abramson, J., Ahuja, A., Barr, I., Brussee, A., Carnevale, F., Cassin, M., Chhaparia, R., Clark, S., Damoc, B., Dudzik, A., et al. Imitating interactive intelligence. arXiv preprint arXiv:2012.05672, 2020.
  2. 2.Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N., Julian, R., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., and Yan, M. Do as I can and not as I say: Grounding language in robotic affordances. In arXiv preprint arXiv:2204.01691, 2022.
  3. 3.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: A visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022.
  4. 4.Ammanabrolu, P. and Riedl, M. O. Playing text-adventure games with graph-based deep reinforcement learning. arXiv preprint arXiv:1812.01628, 2018.
  5. 5.An, D., Wang, H., Wang, W., Wang, Z., Huang, Y., He, K., and Wang, L. Etpnav: Evolving topological planning for vision-language navigation in continuous environments. arXiv preprint arXiv:2304.03047, 2023.
  6. 6.Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I. D., Gould, S., and van den Hengel, A. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pp. 3674–3683. IEEE Computer Society, 2018. doi: 10.1109/CVPR.2018.00387. URL http://openaccess.thecvf.com/content_cvpr_2018/html/Anderson_Vision-and-Language_Navigation_Interpreting_CVPR_2018_paper.html.
  7. 7.Andreas, J. and Klein, D. Alignment-based compositional semantics for instruction following. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 1165–1174, Lisbon, Portugal, 2015. Association for Computational Linguistics. doi: 10.18653/v1/D15-1138. URL https://www.aclweb.org/anthology/D15-1138.
  8. 8.Andreas, J., Klein, D., and Levine, S. Modular multitask reinforcement learning with policy sketches. In Precup, D. and Teh, Y. W. (eds.), Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pp. 166–175. PMLR, 2017. URL http://proceedings.mlr.press/v70/andreas17a.html.
  9. 9.Bara, C.-P., CH-Wang, S., and Chai, J. MindCraft: Theory of mind modeling for situated dialogue in collaborative tasks. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 1112–1125, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.85. URL https://aclanthology.org/2021.emnlp-main.85.
  10. 10.Bengio, Y., Léonard, N., and Courville, A. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432, 2013.
  11. 11.Bisk, Y., Holtzman, A., Thomason, J., Andreas, J., Bengio, Y., Chai, J., Lapata, M., Lazaridou, A., May, J., Nisnevich, A., Pinto, N., and Turian, J. Experience grounds language. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 8718–8735, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.703. URL https://aclanthology.org/2020.emnlp-main.703.
  12. 12.Branavan, S., Zettlemoyer, L., and Barzilay, R. Reading between the lines: Learning to map high-level instructions to commands. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, pp. 1268–1277, Uppsala, Sweden, 2010. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/P10-1129.
  13. 13.Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023.
  14. 14.Carta, T., Romac, C., Wolf, T., Lamprier, S., Sigaud, O., and Oudeyer, P.-Y. Grounding large language models in interactive environments with online reinforcement learning. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 3676–3713. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/v202/carta23a.html.
  15. 15.Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., and Zhang, Y. Matterport3d: Learning from rgb-d data in indoor environments. arXiv preprint arXiv:1709.06158, 2017.
  16. 16.Chen, J., Guo, H., Yi, K., Li, B., and Elhoseiny, M. Visualgpt: Data-efficient adaptation of pretrained language models for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18030–18040, 2022.
  17. 17.Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. Learning phrase representations using rnn encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078, 2014.
  18. 18.Dagan, G., Keller, F., and Lascarides, A. Learning the effects of physical actions in a multi-modal environment. In Findings of the Association for Computational Linguistics: EACL 2023, pp. 133–148, Dubrovnik, Croatia, May 2023. Association for Computational Linguistics. URL https://aclanthology.org/2023.findings-eacl.10.
  19. 19.Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., and Batra, D. Embodied Question Answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  20. 20.Dasgupta, I., Kaeser-Chen, C., Marino, K., Ahuja, A., Babayan, S., Hill, F., and Fergus, R. Collaborating with language models for embodied reasoning. arXiv preprint arXiv:2302.00763, 2023.
  21. 21.Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023.
  22. 22.Du, Y., Watkins, O., Wang, Z., Colas, C., Darrell, T., Abbeel, P., Gupta, A., and Andreas, J. Guiding pretraining in reinforcement learning with large language models, 2023a.
  23. 23.Du, Y., Yang, M., Dai, B., Dai, H., Nachum, O., Tenenbaum, J. B., Schuurmans, D., and Abbeel, P. Learning universal policies via text-guided video generation. arXiv e-prints, pp. arXiv–2302, 2023b.
  24. 24.Eisenstein, J., Clarke, J., Goldwasser, D., and Roth, D. Reading to learn: Constructing features from semantic abstracts. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pp. 958–967, Singapore, 2009. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/D09-1100.
  25. 25.Eldan, R. and Li, Y. Tinystories: How small can language models be and still speak coherent english?, 2023.
  26. 26.Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al. Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures. In International conference on machine learning, pp. 1407–1416. PMLR, 2018.
  27. 27.Espeholt, L., Marinier, R., Stanczyk, P., Wang, K., and Michalski, M. Seed rl: Scalable and efficient deep-rl with accelerated central inference. arXiv preprint arXiv:1910.06591, 2019.
  28. 28.Guo, J., Li, J., Li, D., Tiong, A. M. H., Li, B., Tao, D., and Hoi, S. From images to textual prompts: Zero-shot visual question answering with frozen large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10867–10877, 2023.
  29. 29.Ha, D. and Schmidhuber, J. Recurrent world models facilitate policy evolution. In Advances in Neural Information Processing Systems 31, pp. 2451–2463. 2018.
  30. 30.Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J. Learning latent dynamics for planning from pixels. arXiv preprint arXiv:1811.04551, 2018.
  31. 31.Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J. Mastering atari with discrete world models. arXiv preprint arXiv:2010.02193, 2020.
  32. 32.Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2023.
  33. 33.Hanjie, A. W., Zhong, V., and Narasimhan, K. Grounding language to entities and dynamics for generalization in reinforcement learning. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021,18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pp. 4051–4062. PMLR, 2021. URL http://proceedings.mlr.press/v139/hanjie21a.html.
  34. 34.Hockett, C. F. and Hockett, C. D. The origin of speech. Sci. Am., 203(3):88–97, 1960.
  35. 35.Huang, W., Abbeel, P., Pathak, D., and Mordatch, I. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning. PMLR, 2022a.
  36. 36.Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., Sermanet, P., Brown, N., Jackson, T., Luu, L., Levine, S., Hausman, K., and Ichter, B. Inner monologue: Embodied reasoning through planning with language models. In arXiv preprint arXiv:2207.05608, 2022b.
  37. 37.Jiang, Y., Gu, S. S., Murphy, K. P., and Finn, C. Language as an abstraction for hierarchical deep reinforcement learning. Advances in Neural Information Processing Systems, 32, 2019.
  38. 38.Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L. Vima: General robot manipulation with multimodal prompts. arXiv preprint arXiv:2210.03094, 2022.
  39. 39.Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W. Recurrent experience replay in distributed reinforcement learning. In International conference on learning representations, 2019.
  40. 40.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  41. 41.Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M. Improved variational inference with inverse autoregressive flow. Advances in neural information processing systems, 29, 2016.
  42. 42.Krantz, J., Wijmans, E., Majumdar, A., Batra, D., and Lee, S. Beyond the nav-graph: Vision-and-language navigation in continuous environments. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVIII 16, pp. 104–120. Springer, 2020.
  43. 43.Krantz, J., Gokaslan, A., Batra, D., Lee, S., and Maksymets, O. Waypoint models for instruction-guided navigation in continuous environments. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15162–15171, 2021.
  44. 44.Ku, A., Anderson, P., Patel, R., Ie, E., and Baldridge, J. Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 4392–4412, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.356. URL https://aclanthology.org/2020.emnlp-main.356.
  45. 45.Li, B. Z., Nye, M., and Andreas, J. Implicit representations of meaning in neural language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 1813–1827, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.143. URL https://aclanthology.org/2021.acl-long.143.
  46. 46.Li, B. Z., Chen, W., Sharma, P., and Andreas, J. Lampp: Language models as probabilistic priors for perception and action. arXiv e-prints, 2023a.
  47. 47.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023b.
  48. 48.Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M. Emergent world representations: Exploring a sequence model trained on a synthetic task. In The Eleventh International Conference on Learning Representations, 2023c. URL https://openreview.net/forum?id=DeG07_TcZvT.
  49. 49.Li, S., Puig, X., Du, Y., Wang, C., Akyurek, E., Torralba, A., Andreas, J., and Mordatch, I. Pre-trained language models for interactive decision-making. arXiv preprint arXiv:2202.01771, 2022.
  50. 50.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. In NeurIPS, 2023.
  51. 51.Lu, J., Yang, J., Batra, D., and Parikh, D. Hierarchical question-image co-attention for visual question answering. In Lee, D., Sugiyama, M., Luxburg, U., Guyon, I., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc., 2016. URL https://proceedings.neurips.cc/paper_files/paper/2016/file/9dcb88e0137649590b755372b040afad-Paper.pdf.
  52. 52.Lu, J., Clark, C., Zellers, R., Mottaghi, R., and Kembhavi, A. Unified-io: A unified model for vision, language, and multi-modal tasks. arXiv preprint arXiv:2206.08916, 2022.
  53. 53.Luketina, J., Nardelli, N., Farquhar, G., Foerster, J. N., Andreas, J., Grefenstette, E., Whiteson, S., and Rocktäschel, T. A survey of reinforcement learning informed by natural language. In Kraus, S. (ed.), Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI 2019, Macao, China, August 10-16, 2019, pp. 6309–6317. ijcai.org, 2019. doi: 10.24963/ijcai.2019/880. URL https://doi.org/10.24963/ijcai.2019/880.
  54. 54.Lynch, C. and Sermanet, P. Language conditioned imitation learning over unstructured data. Robotics: Science and Systems, 2021. URL https://arxiv.org/abs/2005.07648.
  55. 55.Mirchandani, S., Karamcheti, S., and Sadigh, D. Ella: Exploration through learned language abstraction. Advances in Neural Information Processing Systems, 34:29529–29540, 2021.
  56. 56.Mu, J., Zhong, V., Raileanu, R., Jiang, M., Goodman, N. D., Rocktäschel, T., and Grefenstette, E. Improving intrinsic exploration with language abstractions. In NeurIPS, 2022.
  57. 57.Nair, S., Mitchell, E., Chen, K., Ichter, B., Savarese, S., and Finn, C. Learning language-conditioned robot behavior from offline data and crowd-sourced annotation. In Conference on Robot Learning, 2021. URL https://api.semanticscholar.org/CorpusID:237385309.
  58. 58.Narasimhan, K., Barzilay, R., and Jaakkola, T. Grounding language for transfer in deep reinforcement learning. Journal of Artificial Intelligence Research, 63:849–874, 2018.
  59. 59.Padmakumar, A., Thomason, J., Shrivastava, A., Lange, P., Narayan-Chen, A., Gella, S., Piramuthu, R., and Gokhan Tur and, D. H.-T. TEACh: Task-driven Embodied Agents that Chat. In Conference on Artificial Intelligence (AAAI), 2022. URL https://arxiv.org/abs/2110.00534.
  60. 60.Piantadosi, S. T. and Hill, F. Meaning without reference in large language models. ArXiv, abs/2208.02957, 2022.
  61. 61.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  62. 62.Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al. A generalist agent. arXiv preprint arXiv:2205.06175, 2022.
  63. 63.Reimers, N. and Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 11 2019. URL http://arxiv.org/abs/1908.10084.
  64. 64.Rezende, D. J., Mohamed, S., and Wierstra, D. Stochastic backpropagation and approximate inference in deep generative models. arXiv preprint arXiv:1401.4082, 2014.
  65. 65.Robine, J., Höftmann, M., Uelwer, T., and Harmeling, S. Transformer-based world models are happy with 100k interactions. arXiv preprint arXiv:2303.07109, 2023.
  66. 66.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  67. 67.Sharma, P., Torralba, A., and Andreas, J. Skill induction and planning with latent language. arXiv preprint arXiv:2110.01517, 2021.
  68. 68.Shridhar, M., Thomason, J., Gordon, D., Bisk, Y., Han, W., Mottaghi, R., Zettlemoyer, L., and Fox, D. ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020a. URL https://arxiv.org/abs/1912.01734.
  69. 69.Shridhar, M., Yuan, X., Côté, M.-A., Bisk, Y., Trischler, A., and Hausknecht, M. Alfworld: Aligning text and embodied environments for interactive learning. arXiv preprint arXiv:2010.03768, 2020b.
  70. 70.Shridhar, M., Manuelli, L., and Fox, D. Cliport: What and where pathways for robotic manipulation. In Conference on Robot Learning, pp. 894–906. PMLR, 2022.
  71. 71.Singh, I., Singh, G., and Modi, A. Pre-trained language models as prior knowledge for playing text-based games. arXiv preprint arXiv:2107.08408, 2021.
  72. 72.Sutton, R. S. Dyna, an integrated architecture for learning, planning, and reacting. ACM SIGART Bulletin, 2(4):160–163, 1991.
  73. 73.Sutton, R. S. and Barto, A. G. Reinforcement learning: An introduction. MIT press, 2018.
  74. 74.Tam, A. C., Rabinowitz, N. C., Lampinen, A. K., Roy, N. A., Chan, S. C. Y., Strouse, D., Wang, J., Banino, A., and Hill, F. Semantic exploration from language abstractions and pretrained representations. In NeurIPS, 2022.
  75. 75.Thomason, J., Murray, M., Cakmak, M., and Zettlemoyer, L. Vision-and-dialog navigation. In Conference on Robot Learning, 2019. URL https://api.semanticscholar.org/CorpusID:195886244.
  76. 76.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv: Arxiv-2305.16291, 2023.
  77. 77.Williams, R. J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3-4):229–256, 1992.
  78. 78.Winograd, T. Understanding natural language. Cognitive Psychology, 3(1):1–191, 1972. ISSN 0010-0285. doi: https://doi.org/10.1016/0010-0285(72)90002-3. URL https://www.sciencedirect.com/science/article/pii/0010028572900023.
  79. 79.Wu, Y., Min, S. Y., Prabhumoye, S., Bisk, Y., Salakhutdinov, R., Azaria, A., Mitchell, T., and Li, Y. SPRING: Studying papers and reasoning to play games. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=jU9qiRMDtR.
  80. 80.Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114, 2023.
  81. 81.Zhong, V., Rocktäschel, T., and Grefenstette, E. RTFM: generalising to new environment dynamics via reading. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=SJgob6NKvH.
  82. 82.Zhong, V., Mu, J., Zettlemoyer, L., Grefenstette, E., and Rocktaschel, T. Improving policy learning via language dynamics distillation. In Thirty-sixth Conference on Neural Information Processing Systems, 2022.
  83. 83.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences. ArXiv, abs/1909.08593, 2019.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/