Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
Xiao YuBaolin PengMichel GalleyHao ChengQianhui WuJanardhan KulkarniSuman NathZhou YuJianfeng Gao
Introduces Dyna-Mind, a training framework combining experience-grounded simulation reasoning with online reinforcement learning to equip AI agents with the ability to mentally evaluate future states before acting in complex, long-horizon interactive environments.
Artificial intelligence reasoning models have achieved strong results in static domains like mathematics and coding, but they frequently struggle in interactive, long-horizon tasks such as operating mobile phones, navigating graphical user interfaces, and solving multi-step text puzzles. These failures occur largely because current agents lack internal world models to simulate potential future states before committing to actions. The article evaluates and demonstrates Dyna-Mind, a two-stage training framework designed to teach vision-language agents how to mentally simulate environments and use those simulations to improve multi-step decision-making.
The framework integrates world simulation into an agent's reasoning process across two training phases. In the first stage, called Reasoning with Simulations, the system builds search trees from real environment rollouts, values each branch, and condenses these explored paths into single, structured reasoning traces used for supervised fine-tuning. In the second stage, an online reinforcement learning algorithm named Dyna-GRPO alternates between standard policy improvement and simulation refinement. During refinement steps, real future state feedback is provided to prompt the agent to correct its intermediate reasoning, optimizing both final task success and world simulation accuracy. The authors evaluated the approach on two text-based environments (Sokoban and ALFWorld) and one realistic graphical benchmark (AndroidWorld).
The evaluations yielded several key findings. First, simulation accuracy strongly correlates with final task success, showing that accurate internal foresight is a key driver of agent competence. Second, in text-based environments, Dyna-Mind models achieved an average success rate of 77.1% on Sokoban and 90.8% on ALFWorld, outperforming standard reinforcement learning baselines (which achieved 73.1% and 87.0% respectively) while using up to eleven times fewer output tokens than large reasoning models like DeepSeek-R1. Third, on the complex AndroidWorld benchmark, a 32-billion parameter model trained with Dyna-Mind achieved a 31.8% average success rate (40.7% in-distribution and 22.9% out-of-distribution), outperforming both the standard prompting baseline of 19.5% and standard reinforcement learning at 27.8%.
These findings indicate that autonomous agents do not require separate, computationally heavy search modules at runtime if world dynamics are baked directly into their internalized reasoning traces. By generating concise, simulation-grounded plans, agents reduce computational latency and token costs while making fewer irreversible operational mistakes. The article recommends adopting end-to-end simulation-guided training for autonomous agents and using environment interaction rollouts to refine agent foresight during reinforcement learning.
The approach currently faces limitations in visual environments. On AndroidWorld, performance remains constrained by underlying visual model errors, including misinterpreting complex interface icons and an inability to recover after sequential interface mistakes. While confidence is high that mental simulation improves agent planning across domains, scaling these techniques to real-world software will require stronger foundational visual perception.
- Paper: Reasoning with Language Model is Planning with World Model, Shibo Hao et al. (2023). This paper establishes the foundation for repurposing language models to act as internal world models for planning over simulated states, a core prerequisite conceptualized in Dyna-Mind's simulation framework.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). It provides the seminal paradigm of interleaving reasoning traces and environment actions in language models that Dyna-Mind builds upon and augments with internal simulations.
- Paper: Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models, Andy Zhou et al. (2024). This work unifies reasoning, acting, and tree-search planning with environment feedback, directly preceding Dyna-Mind's search-tree experience generation (ReSim).
- Paper: Dream to Control: Learning Behaviors by Latent Imagination, Danijar Hafner et al. (2019). It introduces the foundational concept of learning behaviors entirely through latent imagination in learned world models, which inspires Dyna-Mind's mental simulation mechanism.
- Paper: Recurrent World Models Facilitate Policy Evolution, David Ha et al. (2018). This foundational paper demonstrates how recurrent world models enable agents to simulate futures internally for policy optimization, motivating Dyna-Mind's vicarious trial-and-error reasoning.
- Paper: Reflexion: language agents with verbal reinforcement learning, Noah Shinn et al. (2023). It establishes iterative verbal self-reflection and learning from trial outcomes in interactive environments like ALFWorld, which Dyna-Mind formalizes into structured simulation training.
- Paper: Grounding Large Language Models in Interactive Environments with Online Reinforcement Learning, Thomas Carta et al. (2023). This work details grounding language models in interactive environments via online reinforcement learning, providing essential context for Dyna-Mind's second-stage RL optimization.
- Paper: DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning, A list of authors and their affiliations appears at the end of the paper (2025). It introduces Group Relative Policy Optimization (GRPO) for reasoning models, which Dyna-Mind adapts into Dyna-GRPO to leverage intermediate states and outcome rewards.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). This comprehensive survey organizes advanced agentic reasoning and self-evolving mechanisms, contextualizing Dyna-Mind's simulation-grounded planning within the broader agent landscape.
- Paper: Experiential Reinforcement Learning, Taiwei Shi et al. (2026). It builds upon experiential reinforcement learning in interactive puzzle environments like Sokoban by internalizing structured reflections into deployment policies.
- Paper: From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills, Zisu Huang et al. (2026). This paper investigates how raw interactive execution logs can be distilled into reusable agent skills across complex environments like ALFWorld, extending experience-driven reasoning.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). It analyzes the internal multi-perspective dynamics of reinforcement-learned reasoning traces, providing deeper mechanistic insight into how simulated reasoning trajectories emerge.
