Dyna-Mind: Learning to Simulate from Experience for Better AI Agents

Xiao YuBaolin PengMichel GalleyHao ChengQianhui WuJanardhan KulkarniSuman NathZhou YuJianfeng Gao

article2025arXiv8 citations

Introduces Dyna-Mind, a training framework combining experience-grounded simulation reasoning with online reinforcement learning to equip AI agents with the ability to mentally evaluate future states before acting in complex, long-horizon interactive environments.

Listen

Artificial intelligence reasoning models have achieved strong results in static domains like mathematics and coding, but they frequently struggle in interactive, long-horizon tasks such as operating mobile phones, navigating graphical user interfaces, and solving multi-step text puzzles. These failures occur largely because current agents lack internal world models to simulate potential future states before committing to actions. The article evaluates and demonstrates Dyna-Mind, a two-stage training framework designed to teach vision-language agents how to mentally simulate environments and use those simulations to improve multi-step decision-making.

The framework integrates world simulation into an agent's reasoning process across two training phases. In the first stage, called Reasoning with Simulations, the system builds search trees from real environment rollouts, values each branch, and condenses these explored paths into single, structured reasoning traces used for supervised fine-tuning. In the second stage, an online reinforcement learning algorithm named Dyna-GRPO alternates between standard policy improvement and simulation refinement. During refinement steps, real future state feedback is provided to prompt the agent to correct its intermediate reasoning, optimizing both final task success and world simulation accuracy. The authors evaluated the approach on two text-based environments (Sokoban and ALFWorld) and one realistic graphical benchmark (AndroidWorld).

The evaluations yielded several key findings. First, simulation accuracy strongly correlates with final task success, showing that accurate internal foresight is a key driver of agent competence. Second, in text-based environments, Dyna-Mind models achieved an average success rate of 77.1% on Sokoban and 90.8% on ALFWorld, outperforming standard reinforcement learning baselines (which achieved 73.1% and 87.0% respectively) while using up to eleven times fewer output tokens than large reasoning models like DeepSeek-R1. Third, on the complex AndroidWorld benchmark, a 32-billion parameter model trained with Dyna-Mind achieved a 31.8% average success rate (40.7% in-distribution and 22.9% out-of-distribution), outperforming both the standard prompting baseline of 19.5% and standard reinforcement learning at 27.8%.

These findings indicate that autonomous agents do not require separate, computationally heavy search modules at runtime if world dynamics are baked directly into their internalized reasoning traces. By generating concise, simulation-grounded plans, agents reduce computational latency and token costs while making fewer irreversible operational mistakes. The article recommends adopting end-to-end simulation-guided training for autonomous agents and using environment interaction rollouts to refine agent foresight during reinforcement learning.

The approach currently faces limitations in visual environments. On AndroidWorld, performance remains constrained by underlying visual model errors, including misinterpreting complex interface icons and an inability to recover after sequential interface mistakes. While confidence is high that mental simulation improves agent planning across domains, scaling these techniques to real-world software will require stronger foundational visual perception.

Cover for Dyna-Mind: Learning to Simulate from Experience for Better AI Agents

Abstract

Reasoning models have recently shown remarkable progress in domains such as math and coding. However, their expert-level abilities in math and coding contrast sharply with their performance in long-horizon, interactive tasks such as web navigation and computer/phone-use. Inspired by literature on human cognition, we argue that current AI agents need ''vicarious trial and error'' - the capacity to mentally simulate alternative futures before acting - in order to enhance their understanding and performance in complex interactive environments. We introduce Dyna-Mind, a two-stage training framework that explicitly teaches (V)LM agents to integrate such simulation into their reasoning. In stage 1, we introduce Reasoning with Simulations (ReSim), which trains the agent to generate structured reasoning traces from expanded search trees built from real experience gathered through environment interactions. ReSim thus grounds the agent's reasoning in faithful world dynamics and equips it with the ability to anticipate future states in its reasoning. In stage 2, we propose Dyna-GRPO, an online reinforcement learning method to further strengthen the agent's simulation and decision-making ability by using both outcome rewards and intermediate states as feedback from real rollouts. Experiments on two synthetic benchmarks (Sokoban and ALFWorld) and one realistic benchmark (AndroidWorld) demonstrate that (1) ReSim effectively infuses simulation ability into AI agents, and (2) Dyna-GRPO leverages outcome and interaction-level signals to learn better policies for long-horizon, planning-intensive tasks. Together, these results highlight the central role of simulation in enabling AI agents to reason, plan, and act more effectively in the ever more challenging environments.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Dyna-Mind
  • 3.1 Notation
  • 3.2 Reasoning with Simulations (ReSim)
  • 3.3 Dyna-GRPO
  • 4 Experiments
  • 4.1 Text Games
  • 4.1.1 Main Results
  • 4.1.2 Measuring Simulation Ability
  • 4.2 AndroidWorld
  • 4.2.1 Main Results
  • 5 Conclusion
  • References
  • A LLM Usage
  • B Ethics Statement
  • C Additional Algorithmic Details
  • D Additional Details on Text Games
  • D.1 Example Tasks and Actions
  • D.2 ReSim Implementation Details
  • D.3 Simulation Refinement Performance
  • D.4 Additional Training Details
  • D.5 Simulation Score Prompts
  • E Additional Details on AndroidWorld
  • E.1 Example Task and Actions in AndroidWorld
  • E.2 Additional Training Details
  • E.3 Other Implementation/Evaluation Details

Knowls

  1. Knowl 1 — Dyna-Mind Framework for Grounded Simulation in Language Model Agents

    model/method

    Dyna-Mind is a two-stage training framework designed to teach (visual) language model ((V)LM) agents to mentally simulate future environment states before selecting actions during interactive, multi-step tasks.

    The framework addresses the issue where (V)LM agents hallucinate environment dynamics in long-horizon planning tasks:

    1. Stage 1 (Reasoning with Simulations / RESIM): Synthesizes training data by constructing expanded search trees (via depth-first search or rollouts) grounded in real environment interactions. A (V)LM aggregates these search trees—including ground-truth intermediate states and discounted value estimates—into structured reasoning traces (aReSima^{\text{ReSim}}). The policy model is then trained via supervised fine-tuning (SFT) to generate these simulation-guided reasoning traces directly from observation histories without requiring tree search during inference.

    2. Stage 2 (Dyna-GRPO): An online reinforcement learning algorithm that alternates between simulation improvement and policy improvement. In simulation improvement, the agent generates simulation refinement rollouts (SIMROLLOUT) where proposed plans are executed to obtain true future states, prompting the model to refine its simulation. The model is trained on both unrefined trajectories and refined trajectories with a specialized advantage function. In policy improvement, standard rollouts without future-state information are optimized using Group Relative Policy Optimization (GRPO).

  2. Knowl 2 — Reasoning with Simulations (RESIM) Search Tree Aggregation and Distillation

    algorithm

    Reasoning with Simulations (RESIM) generates structured imitation learning data by executing search trees in the environment and condensing the exploration into a single simulation-guided reasoning trace per step.

    Input: Rollout policy πθ\pi_\theta, value function VνV_\nu, environment T\mathcal{T}, aggregation model MM, branching factor bb, search depth dd, maximum episode steps tmaxt_{\text{max}}, training branch count btrainb_{\text{train}}
    Output: Trajectory τ\tau of states and simulation-guided reasoning traces
    Initialize trajectory τ←{}\tau \leftarrow \{\}, time step t←0t \leftarrow 0, state s0∼Ts_0 \sim \mathcal{T}
    while task not done and t<tmaxt < t_{\text{max}} do
        Sample bb rollouts {τi}i=1b\{\tau^i\}_{i=1}^b of depth up to dd from state sts_t using policy πθ\pi_\theta
        Deduplicate sampled rollouts into unique set {τi}i=1b′\{\tau^i\}_{i=1}^{b'}
        Compute terminal state value estimates vi=Vν(st+di)v^i = V_\nu(s_{t+d}^i) for each rollout i∈{1,…,b′}i \in \{1, \dots, b'\}
        Identify best rollout τ∗=arg⁡max⁡τivi\tau^* = \arg\max_{\tau^i} v^i
        Form subset {τi}i=1btrain\{\tau^i\}_{i=1}^{b_{\text{train}}} consisting of τ∗\tau^* and btrain−1b_{\text{train}} - 1 subsampled rollouts from the remainder
        for each rollout i∈{1,…,btrain}i \in \{1, \dots, b_{\text{train}}\} do
            plani←summarize(M,τi,vi)\text{plan}^i \leftarrow \text{summarize}(M, \tau^i, v^i)
        end for
        aReSim←aggregate(M,st,{plani}i=1btrain)a^{\text{ReSim}} \leftarrow \text{aggregate}(M, s_t, \{\text{plan}^i\}_{i=1}^{b_{\text{train}}})
        Execute chosen action from aReSima^{\text{ReSim}} in T\mathcal{T} to transition st+1∼T(st,aReSim)s_{t+1} \sim \mathcal{T}(s_t, a^{\text{ReSim}})
        τ←τ∪{st,aReSim}\tau \leftarrow \tau \cup \{s_t, a^{\text{ReSim}}\}
        t←t+1t \leftarrow t + 1
    end while
    return τ\tau

    Following trajectory collection, the agent is trained via supervised fine-tuning (SFT) on the resulting sequence τ={s0,a0ReSim,s1,a1ReSim,…,sT,aTReSim}\tau = \{s_0, a_0^{\text{ReSim}}, s_1, a_1^{\text{ReSim}}, \dots, s_T, a_T^{\text{ReSim}}\} to predict atReSima_t^{\text{ReSim}} given the current state sts_t and up to hh historical state-action pairs.

  3. Knowl 3 — Simulation Refinement Rollout (SIMROLLOUT)

    algorithm

    SIMROLLOUT generates refined action trajectories by executing an agent's planned actions in the real environment to retrieve true intermediate states, then prompting the agent to refine its reasoning and action conditioned on these future ground-truth states.

    Input: Policy πθ\pi_\theta, environment T\mathcal{T}, group size GG, plan depth nn, maximum episode steps tmaxt_{\text{max}}
    Output: Refined trajectory set {τ′}\{\tau'\}, conditioned refined trajectory set {τrefine′}\{\tau'_{\text{refine}}\}
    Initialize sets {τ′}←{}\{\tau'\} \leftarrow \{\}, {τrefine′}←{}\{\tau'_{\text{refine}}\} \leftarrow \{\}
    repeat GG times
        Initialize episode buffers τ′←{}\tau' \leftarrow \{\}, τrefine′←{}\tau'_{\text{refine}} \leftarrow \{\}, t←0t \leftarrow 0, s0∼Ts_0 \sim \mathcal{T}
        while task not done and t<tmaxt < t_{\text{max}} do
            Sample initial response a∼πθ(⋅∣st)a \sim \pi_\theta(\cdot \mid s_t)
            Extract multi-step action plan {a^1,…,a^n}\{\hat{a}_1, \dots, \hat{a}_n\} from aa
            Execute plan in environment to obtain true future states: st+1∼T(st,a^1),…,st+n∼T(st+n−1,a^n)s_{t+1} \sim \mathcal{T}(s_t, \hat{a}_1), \dots, s_{t+n} \sim \mathcal{T}(s_{t+n-1}, \hat{a}_n)
            Construct refinement prompt: strefine={st⊕a⊕st+1⊕a^2⊕⋯⊕st+n}s_t^{\text{refine}} = \{s_t \oplus a \oplus s_{t+1} \oplus \hat{a}_2 \oplus \dots \oplus s_{t+n}\}
            Sample refined response arefine∼πθ(⋅∣strefine)a^{\text{refine}} \sim \pi_\theta(\cdot \mid s_t^{\text{refine}})
            τ′←τ′∪{st,arefine}\tau' \leftarrow \tau' \cup \{s_t, a^{\text{refine}}\}
            τrefine′←τrefine′∪{strefine,arefine}\tau'_{\text{refine}} \leftarrow \tau'_{\text{refine}} \cup \{s_t^{\text{refine}}, a^{\text{refine}}\}
            Execute arefinea^{\text{refine}} to obtain next environment state st+1∼T(st,arefine)s_{t+1} \sim \mathcal{T}(s_t, a^{\text{refine}})
            t←t+1t \leftarrow t + 1
        end while
        {τ′}←{τ′}∪{τ′}\{\tau'\} \leftarrow \{\tau'\} \cup \{\tau'\}
        {τrefine′}←{τrefine′}∪{τrefine′}\{\tau'_{\text{refine}}\} \leftarrow \{\tau'_{\text{refine}}\} \cup \{\tau'_{\text{refine}}\}
    until GG trajectories collected
    return {τ′},{τrefine′}\{\tau'\}, \{\tau'_{\text{refine}}\}
  4. Knowl 4 — Dyna-GRPO Online Reinforcement Learning Algorithm

    algorithm

    Dyna-GRPO strengthens an agent's simulation and decision-making capabilities without relying on external tree search during RL by iteratively alternating between simulation improvement (learning from future-state-guided rollouts) and direct policy improvement.

    Input: Policy πθ\pi_\theta, environment T\mathcal{T}, group size GG, training iterations NN, simulation steps nTn_T, policy steps nπn_\pi
    Output: Updated policy πθ\pi_\theta
    for iteration =1= 1 to NN do
        // Simulation Improvement Phase
        for step =1= 1 to nTn_T do
            {τ′},{τrefine′}←SimRollout(πθ,T,G/2)\{\tau'\}, \{\tau'_{\text{refine}}\} \leftarrow \text{SimRollout}(\pi_\theta, \mathcal{T}, G / 2)
            {τ}←Rollout(πθ,T,G/2)\{\tau\} \leftarrow \text{Rollout}(\pi_\theta, \mathcal{T}, G / 2)
            Update πθ\pi_\theta with GRPO on combined unconditioned group {τ}∪{τ′}\{\tau\} \cup \{\tau'\}
            Update πθ\pi_\theta with GRPO on future-conditioned group {τrefine′}\{\tau'_{\text{refine}}\} using advantage ArefineA_{\text{refine}}
        end for
        
        // Policy Improvement Phase
        for step =1= 1 to nπn_\pi do
            {τ}←Rollout(πθ,T,G)\{\tau\} \leftarrow \text{Rollout}(\pi_\theta, \mathcal{T}, G)
            Update πθ\pi_\theta with standard GRPO on {τ}\{\tau\}
        end for
    end for
    return πθ\pi_\theta
  5. Knowl 5 — Advantage Formulations and Objectives in Dyna-GRPO

    equation

    Dyna-GRPO builds upon Group Relative Policy Optimization (GRPO). The standard GRPO surrogate objective is defined as:

    JGRPO(θ)=Eτ∼πθold[1GT∑i=1G∑t=1Tmin⁡(ρθ(at(i))A(at(i)),clip(ρθ(at(i)),1−ϵ,1+ϵ)A(at(i)))−βDKL(πθ∥πθref)]\mathcal{J}_{\text{GRPO}}(\theta) = \mathbb{E}_{\tau \sim \pi_{\theta_{\text{old}}}} \left[ \frac{1}{GT} \sum_{i=1}^G \sum_{t=1}^T \min \left( \rho_\theta(a_t^{(i)}) A(a_t^{(i)}), \text{clip}(\rho_\theta(a_t^{(i)}), 1 - \epsilon, 1 + \epsilon) A(a_t^{(i)}) \right) - \beta D_{\text{KL}}(\pi_\theta \parallel \pi_{\theta_{\text{ref}}}) \right]

    where ρθ(a)=πθ(a∣s)πθref(a∣s)\rho_\theta(a) = \frac{\pi_\theta(a|s)}{\pi_{\theta_{\text{ref}}}(a|s)} is the importance sampling ratio, GG is the group size, TT is the trajectory horizon, ϵ\epsilon is the clipping parameter, β\beta is the KL divergence penalty coefficient, and πθref\pi_{\theta_{\text{ref}}} is the reference policy.

    The standard episode-level advantage function A(at(i))=AGRPO(τ(i))A(a_t^{(i)}) = A_{\text{GRPO}}(\tau^{(i)}) across a group of GG rollouts is normalized by group reward statistics:

    AGRPO(τ(i))=R(τ(i))−mean({R(τ(j))}j=1G)std({R(τ(j))}j=1G),R(τ(i))=∑t=1TR(st,at)A_{\text{GRPO}}(\tau^{(i)}) = \frac{R(\tau^{(i)}) - \text{mean}(\{R(\tau^{(j)})\}_{j=1}^G)}{\text{std}(\{R(\tau^{(j)})\}_{j=1}^G)}, \quad R(\tau^{(i)}) = \sum_{t=1}^T R(s_t, a_t)

    where R(st,at)=−0.1R(s_t, a_t) = -0.1 for non-terminal steps, and for terminal steps R(sT,aT)=10.0R(s_T, a_T) = 10.0 on success or 0.00.0 on failure.

    In the simulation improvement phase, the refined trajectory group {τrefine(i)}i=1G/2\{\tau^{(i)}_{\text{refine}}\}_{i=1}^{G/2} is trained using a specialized indicator advantage ArefineA_{\text{refine}}:

    Arefine(τrefine(i))={1.0,if τrefine(i) is correct and R(τrefine(i))>max⁡(Rˉ,Rˉrefine)0.0,otherwiseA_{\text{refine}}(\tau^{(i)}_{\text{refine}}) = \begin{cases} 1.0, & \text{if } \tau^{(i)}_{\text{refine}} \text{ is correct and } R(\tau^{(i)}_{\text{refine}}) > \max(\bar{R}, \bar{R}_{\text{refine}}) \\ 0.0, & \text{otherwise} \end{cases}

    where Rˉ=1G/2∑i=1G/2R(τ(i))\bar{R} = \frac{1}{G/2} \sum_{i=1}^{G/2} R(\tau^{(i)}) is the empirical mean reward of standard unconditioned rollouts, and Rˉrefine=1G/2∑i=1G/2R(τrefine(i))\bar{R}_{\text{refine}} = \frac{1}{G/2} \sum_{i=1}^{G/2} R(\tau^{(i)}_{\text{refine}}) is the mean reward of the SIMROLLOUT trajectories.

  6. Knowl 6 — Simulation Score Protocol for Evaluating Agent World Modeling

    definition

    The Simulation Score (Sim Score ∈[0,1]\in [0, 1]) quantifies the fidelity and utility of an agent's internal world simulation during its reasoning process.

    Given an agent's generated response at∼πθ(⋅∣st)a_t \sim \pi_\theta(\cdot|s_t) at state sts_t:

    1. An extraction LLM parses ata_t into an intended action plan (a^1,…,a^d)(\hat{a}_1, \dots, \hat{a}_d) and corresponding imagined next-state descriptions (s^t+1,…,s^t+d)(\hat{s}_{t+1}, \dots, \hat{s}_{t+d}).
    2. The action plan is executed in the actual environment to observe true next-states (st+1,…,st+d)(s_{t+1}, \dots, s_{t+d}).
    3. An independent judge LLM (e.g., Qwen3-235B-A22B-Instruct) evaluates the imagined states against ground-truth states across two sub-dimensions:
      • Correctness (maximum 0.30.3 points): Evaluates factual agreement between imagined coordinates, objects, and environment layout relative to the reference observation.
      • Progress (maximum 0.70.7 points): Evaluates whether the plan achieves meaningful task advancement (e.g., moving objects closer to targets, solving the task, or discovering required objects).
    4. The sum of correctness and progress scores yields a turn score ∈[0,1]\in [0, 1], which is averaged over all steps in a trajectory to obtain the trajectory-level Simulation Score.
  7. Knowl 7 — Task Success and Efficiency of Dyna-Mind on Sokoban and ALFWorld

    data/table

    Across both Sokoban (spatial planning) and ALFWorld (text-based embodied household tasks), RESIM and Dyna-Mind achieve high task success rates while maintaining significantly smaller token generation costs compared to standard reasoning model distillation (DISTILL(R1)). All Stage 1 and Stage 2 models use Qwen2.5-7B-Instruct as the base backbone.

    Method Gen. Token Sokoban ALFWorld
    ID OOD AVG ID OOD AVG
    REACT(Qwen2.5-7B-Instruct) 1.0x 25.8±1.825.8_{\pm 1.8} - - 35.4±1.935.4_{\pm 1.9} - -
    REACT(Qwen2.5-32B-Instruct) 2.7x 36.7±4.236.7_{\pm 4.2} - - 36.2±3.336.2_{\pm 3.3} - -
    REACT(GPT-4o) 1.5x 37.8±1.037.8_{\pm 1.0} - - 51.3±2.151.3_{\pm 2.1} - -
    REACT(Claude-3.7-Sonnet) 2.3x 70.3±1.270.3_{\pm 1.2} - - 46.1±1.046.1_{\pm 1.0} - -
    REACT(DeepSeek-V3) 2.5x 57.0±1.657.0_{\pm 1.6} - - 55.2±1.055.2_{\pm 1.0} - -
    REACT(DeepSeek-R1) 14.5x 96.6±0.296.6_{\pm 0.2} - - 62.5±0.562.5_{\pm 0.5} - -
    RESIM 2.0x 96.4±0.296.4_{\pm 0.2} - - 87.7±1.187.7_{\pm 1.1} - -
    Dyna-Think
    DIT(R1)+DDT(T^\hat{\mathcal{T}}) 24.2x 74.0±1.474.0_{\pm 1.4} 57.5±1.257.5_{\pm 1.2} 65.8±1.965.8_{\pm 1.9} 63.2±1.563.2_{\pm 1.5} 56.7±2.856.7_{\pm 2.8} 58.9±2.358.9_{\pm 2.3}
    Dyna-Mind Stage 1 (SFT)
    DISTILL(V3) 2.1x 49.2±1.149.2_{\pm 1.1} 34.4±1.334.4_{\pm 1.3} 41.8±1.141.8_{\pm 1.1} 58.9±1.158.9_{\pm 1.1} 56.7±1.056.7_{\pm 1.0} 57.8±1.257.8_{\pm 1.2}
    DISTILL(R1) 24.0x 72.5±2.972.5_{\pm 2.9} 57.0±1.957.0_{\pm 1.9} 64.8±2.564.8_{\pm 2.5} 59.4±1.559.4_{\pm 1.5} 54.2±3.954.2_{\pm 3.9} 56.8±3.556.8_{\pm 3.5}
    DISTILL(RESIM) 2.0x 71.9±1.571.9_{\pm 1.5} 55.5±1.655.5_{\pm 1.6} 63.7±1.963.7_{\pm 1.9} 78.9±2.178.9_{\pm 2.1} 69.3±1.369.3_{\pm 1.3} 74.1±1.874.1_{\pm 1.8}
    Dyna-Mind Stage 2 (RL)
    DISTILL(RESIM) + RLOO 2.2x 78.1±1.878.1_{\pm 1.8} 65.1±1.365.1_{\pm 1.3} 71.3±0.971.3_{\pm 0.9} 85.9±1.385.9_{\pm 1.3} 85.4±2.085.4_{\pm 2.0} 85.5±2.085.5_{\pm 2.0}
    DISTILL(RESIM) + GRPO 2.1x 79.1±1.379.1_{\pm 1.3} 67.8±0.667.8_{\pm 0.6} 73.1±1.473.1_{\pm 1.4} 87.0±3.287.0_{\pm 3.2} 87.1±1.187.1_{\pm 1.1} 87.0±1.887.0_{\pm 1.8}
    DISTILL(RESIM) + DYNA-GRPO 1.9x 82.5±1.582.5_{\pm 1.5} 70.1±1.670.1_{\pm 1.6} 77.1±1.777.1_{\pm 1.7} 92.5±0.892.5_{\pm 0.8} 89.1±1.389.1_{\pm 1.3} 90.8±0.990.8_{\pm 0.9}

    Key observations:

    • In ALFWorld, where DeepSeek-R1 struggles to simulate environment dynamics (62.5%62.5\% success), RESIM reaches 87.7%87.7\%, and DISTILL(RESIM) (74.1%74.1\% avg) outperforms DISTILL(R1) (56.8%56.8\% avg) while using 11×11\times fewer tokens.
    • Stage 2 online RL with DYNA-GRPO improves DISTILL(RESIM) to 77.1%77.1\% on Sokoban and 90.8%90.8\% on ALFWorld, exceeding standard GRPO and RLOO baselines.
  8. Knowl 8 — Empirical Correlation Between Simulation Ability and Agent Task Success

    data/table

    Evaluating agent simulation fidelity reveals a strong positive Spearman correlation (rsr_s) between Simulation Score and task success rate across both Sokoban and ALFWorld benchmarks.

    Method Sokoban ALFWorld
    Success Sim Score Success Sim Score
    REACT(Qwen2.5-7B-Instruct) 25.8±1.825.8_{\pm 1.8} 0.210.21 (r=0.64r=0.64) 35.4±1.935.4_{\pm 1.9} 0.180.18 (r=0.46r=0.46)
    REACT(DeepSeek-V3) 57.0±1.657.0_{\pm 1.6} 0.540.54 (r=0.81r=0.81) 55.2±1.055.2_{\pm 1.0} 0.350.35 (r=0.68r=0.68)
    REACT(DeepSeek-R1) 96.6±0.296.6_{\pm 0.2} 0.930.93 (r=0.96r=0.96) 62.5±0.562.5_{\pm 0.5} 0.360.36 (r=0.70r=0.70)
    RESIM 96.4±0.296.4_{\pm 0.2} 1.001.00 (−-) 87.7±1.187.7_{\pm 1.1} 1.001.00 (−-)
    Dyna-Think
    DIT(R1)+DDT(T^\hat{\mathcal{T}}) 74.0±1.474.0_{\pm 1.4} 0.620.62 (r=0.74r=0.74) 63.2±1.563.2_{\pm 1.5} 0.360.36 (r=0.76r=0.76)
    Dyna-Mind Stage 1 (SFT)
    DISTILL(R1) 72.5±2.972.5_{\pm 2.9} 0.610.61 (r=0.75r=0.75) 59.4±1.559.4_{\pm 1.5} 0.340.34 (r=0.77r=0.77)
    DISTILL(RESIM) 71.9±1.571.9_{\pm 1.5} 0.620.62 (r=0.78r=0.78) 78.9±2.178.9_{\pm 2.1} 0.370.37 (r=0.74r=0.74)
    Dyna-Mind Stage 2 (RL)
    DISTILL(RESIM) + GRPO 79.1±1.379.1_{\pm 1.3} 0.620.62 (r=0.65r=0.65) 87.0±3.287.0_{\pm 3.2} 0.380.38 (ρ=0.48\rho=0.48)
    DISTILL(RESIM) + DYNA-GRPO 82.5±1.582.5_{\pm 1.5} 0.670.67 (r=0.64r=0.64) 92.5±0.892.5_{\pm 0.8} 0.430.43 (r=0.55r=0.55)

    The measurements demonstrate:

    • DeepSeek-R1 achieves high simulation accuracy (0.930.93) and task success (96.6%96.6\%) in Sokoban, but drops to 0.360.36 Sim Score and 62.5%62.5\% success in ALFWorld where layout modeling is challenging.
    • DYNA-GRPO increases both the Simulation Score (e.g., from 0.370.37 to 0.430.43 on ALFWorld) and success rate (from 78.9%78.9\% to 92.5%92.5\% on ALFWorld ID), showing that improving the agent's internal simulation capability directly boosts downstream task performance.
  9. Knowl 9 — Performance of Dyna-Mind on the AndroidWorld Benchmark

    data/table

    Dyna-Mind evaluated on the realistic AndroidWorld GUI automation benchmark using screenshot-only observations and a maximum budget of 15 steps.

    Method Gen. Token AndroidWorld
    ID OOD AVG
    REACT(GPT-4o) 1.0x 5.1±0.25.1_{\pm 0.2} - -
    REACT(Qwen2.5-VL-7B-Instruct) 1.0x 5.3±0.25.3_{\pm 0.2} - -
    REACT(Qwen2.5-VL-72B-Instruct) 1.1x 19.5±0.419.5_{\pm 0.4} - -
    RESIM 2.1x 34.4±0.434.4_{\pm 0.4} - -
    Dyna-Mind Stage 1 (SFT)
    DISTILL-7B(Qwen2.5-VL-72B-Instruct) 1.0x 13.1±0.413.1_{\pm 0.4} 8.6±0.28.6_{\pm 0.2} 10.8±0.610.8_{\pm 0.6}
    DISTILL-7B(RESIM) 2.1x 21.1±0.421.1_{\pm 0.4} 10.2±0.610.2_{\pm 0.6} 15.7±0.815.7_{\pm 0.8}
    DISTILL-32B(RESIM) 2.0x 32.8±0.432.8_{\pm 0.4} 15.6±0.715.6_{\pm 0.7} 24.2±0.624.2_{\pm 0.6}
    Dyna-Mind Stage 2 (RL)
    DISTILL-32B(RESIM) + GRPO 2.1x 35.3±0.435.3_{\pm 0.4} 20.3±0.620.3_{\pm 0.6} 27.8±0.427.8_{\pm 0.4}
    DISTILL-32B(RESIM) + DYNA-GRPO 1.9x 40.7±1.040.7_{\pm 1.0} 22.9±1.022.9_{\pm 1.0} 31.8±1.031.8_{\pm 1.0}

    Key takeaways:

    • RESIM inference achieves 34.4%34.4\% on ID tasks, substantially exceeding direct prompting of Qwen2.5-VL-72B (19.5%19.5\%) and GPT-4o (5.1%5.1\%).
    • Distilling RESIM into Qwen2.5-VL-32B achieves 24.2%24.2\% average success across ID and OOD splits.
    • Stage 2 online RL with DYNA-GRPO achieves 31.8%31.8\% overall average (40.7%40.7\% ID, 22.9%22.9\% OOD), outperforming standard GRPO (27.8%27.8\%).
  10. Knowl 10 — Effect of Ground-Truth Next-State Feedback on Decision Making

    data/table

    Providing agents with intermediate ground-truth next-state observations during inference (via SIMROLLOUT) significantly boosts decision-making performance over standard REACT prompting across both open and proprietary base models.

    Base Model Method Sokoban ALFWorld
    Qwen2.5-7B-Instruct REACT 25.8±1.825.8_{\pm 1.8} 35.4±1.935.4_{\pm 1.9}
    SIMROLLOUT 30.0±1.430.0_{\pm 1.4} 39.1±1.639.1_{\pm 1.6}
    GPT-4o-2024-11-20 REACT 37.8±1.037.8_{\pm 1.0} 51.3±2.151.3_{\pm 2.1}
    SIMROLLOUT 41.4±1.241.4_{\pm 1.2} 64.8±2.564.8_{\pm 2.5}
    GPT-4.1 REACT 67.9±1.067.9_{\pm 1.0} 54.4±2.154.4_{\pm 2.1}
    SIMROLLOUT 71.1±1.371.1_{\pm 1.3} 67.9±2.067.9_{\pm 2.0}

    These results confirm that grounding the agent's thinking process in real next-state feedback enables error correction and reduces execution failures, providing the empirical foundation for using SIMROLLOUT to generate training data during online RL.

Coverage note — None was omitted; all primary methods, algorithms (ReSim, SimRollout, Dyna-GRPO), theoretical/equation formulations, evaluation definitions, and main empirical tables from both text and multimodal benchmarks are covered.

References

  1. 1.Saaket Agashe, Jiuzhou Han, Shuyu Gan, Jiachen Yang, Ang Li, and Xin Eric Wang. Agent s: An open agentic framework that uses computers like a human, 2024. URL https://arxiv.org/abs/2410.08164.
  2. 2.Saaket Agashe, Kyle Wong, Vincent Tu, Jiachen Yang, Ang Li, and Xin Eric Wang. Agent s2: A compositional generalist-specialist framework for computer use agents, 2025. URL https://arxiv.org/abs/2504.00906.
  3. 3.Anthropic. Claude 3.7 Sonnet and Claude Code. https://www.anthropic.com/news/claude-3-7-sonnet, 2025. Accessed: 2025-05-13.
  4. 4.Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025. URL https://arxiv.org/abs/2502.13923.
  5. 5.M.S. Bennett. A Brief History of Intelligence: Evolution, AI, and the Five Breakthroughs That Made Our Brains. HarperCollins, 2023. ISBN 9780063286368. URL https://books.google.com/books?id=tymCEAAAQBAJ.
  6. 6.Hyungjoo Chae, Namyoung Kim, Kai Tzu iunn Ong, Minju Gwak, Gwanwoo Song, Jihoon Kim, Sunghwan Kim, Dongha Lee, and Jinyoung Yeo. Web agents with world models: Learning and leveraging environment dynamics in web navigation, 2025. URL https://arxiv.org/abs/2410.13232.
  7. 7.Zehui Chen, Kuikun Liu, Qiuchen Wang, Wenwei Zhang, Jiangning Liu, Dahua Lin, Kai Chen, and Feng Zhao. Agent-flan: Designing data and methods of effective agent tuning for large language models, 2024. URL https://arxiv.org/abs/2403.12881.
  8. 8.Nathaniel D Daw and Peter Dayan. The algorithmic anatomy of model-based evaluation. Philosophical Transactions of the Royal Society B: Biological Sciences, 369(1655):20130478, 2014.
  9. 9.Nathaniel D. Daw, Yael Niv, and Peter Dayan. Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control. Nature Neuroscience, 8:1704–1711, 2005. URL https://api.semanticscholar.org/CorpusID:16385268.
  10. 10.DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Qu, Hui Li, Jianzhong Guo, Jiashi Li, Jiawei Wang, Jingchang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, J. L. Cai, Jiaqi Ni, Jian Liang, Jin Chen, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Liang Zhao, Litong Wang, Liyue Zhang, Lei Xu, Leyi Xia, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Meng Li, Miaojun Wang, Mingming Li, Ning Tian, Panpan Huang, Peng Zhang, Qiancheng Wang, Qinyu Chen, Qiushi Du, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, R. J. Chen, R. L. Jin, Ruyi Chen, and et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025a. URL https://arxiv.org/abs/2501.12948.
  11. 11.DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J. L. Cai, Jian Liang, Jianzhong Guo, Jiaqi Ni, Jiashi Li, Jiawei Wang, Jin Chen, Jingchang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, Junxiao Song, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Xu, Leyi Xia, Liang Zhao, Litong Wang, Liyue Zhang, Meng Li, Miaojun Wang, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingming Li, Ning Tian, Panpan Huang, Peiyi Wang, Peng Zhang, Qiancheng Wang, Qihao Zhu, Qinyu Chen, Qiushi Du, R. J. Chen, R. L. Jin, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runxin Xu, and et al. Deepseek-v3 technical report, 2025b. URL https://arxiv.org/abs/2412.19437.
  12. 12.Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web, 2023. URL https://arxiv.org/abs/2306.06070.
  13. 13.Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao. Deepresearch bench: A comprehensive benchmark for deep research agents, 2025. URL https://arxiv.org/abs/2506.11763.
  14. 14.Lang Feng, Zhenghai Xue, Tingcong Liu, and Bo An. Group-in-group policy optimization for llm agent training, 2025. URL https://arxiv.org/abs/2505.10978.
  15. 15.Boyu Gou, Zanming Huang, Yuting Ning, Yu Gu, Michael Lin, Weijian Qi, Andrei Kopanev, Botao Yu, Bernal Jiménez Gutiérrez, Yiheng Shu, Chan Hee Song, Jiaman Wu, Shijie Chen, Hanane Nour Moussa, Tianshu Zhang, Jian Xie, Yifei Li, Tianci Xue, Zeyi Liao, Kai Zhang, Boyuan Zheng, Zhaowei Cai, Viktor Rozgic, Morteza Ziyadi, Huan Sun, and Yu Su. Mind2web 2: Evaluating agentic search with agent-as-a-judge, 2025a. URL https://arxiv.org/abs/2506.21506.
  16. 16.Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. Navigating the digital world as humans do: Universal visual grounding for gui agents, 2025b. URL https://arxiv.org/abs/2410.05243.
  17. 17.Yu Gu, Kai Zhang, Yuting Ning, Boyuan Zheng, Boyu Gou, Tianci Xue, Cheng Chang, Sanjari Srivastava, Yanan Xie, Peng Qi, Huan Sun, and Yu Su. Is your llm secretly a world model of the internet? model-based planning for web agents, 2025. URL https://arxiv.org/abs/2411.06559.
  18. 18.Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023. URL https://arxiv.org/abs/2312.06674.
  19. 19.Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues?, 2024. URL https://arxiv.org/abs/2310.06770.
  20. 20.Jing Yu Koh, Stephen McAleer, Daniel Fried, and Ruslan Salakhutdinov. Tree search for language model agents, 2024. URL https://arxiv.org/abs/2407.01476.
  21. 21.Wouter Kool, Herke van Hoof, and Max Welling. Buy 4 REINFORCE samples, get a baseline for free!, 2019. URL https://openreview.net/forum?id=r1lgTGL5DE.
  22. 22.Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Exploring a sequence model trained on a synthetic task, 2024. URL https://arxiv.org/abs/2210.13382.
  23. 23.Haowei Liu, Xi Zhang, Haiyang Xu, Yuyang Wanyan, Junyang Wang, Ming Yan, Ji Zhang, Chunfeng Yuan, Changsheng Xu, Weiming Hu, and Fei Huang. Pc-agent: A hierarchical multi-agent collaboration framework for complex task automation on pc, 2025. URL https://arxiv.org/abs/2502.14282.
  24. 24.Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent linear representations in world models of self-supervised sequence models, 2023. URL https://arxiv.org/abs/2309.00941.
  25. 25.OpenAI. New and improved content moderation tooling. https://openai.com/index/new-and-improved-content-moderation-tooling/, 2022. Accessed: 2025-05-13.
  26. 26.OpenAI. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/, 2024. Accessed: 2024-09-28.
  27. 27.OpenAI. Introducing GPT-4.1 in the api. https://openai.com/index/gpt-4-1/, 2025. Accessed: 2025-09-17.
  28. 28.Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, Kam-Fai Wong, and Shang-Yu Su. Deep dyna-q: Integrating planning for task-completion dialogue policy learning, 2018. URL https://arxiv.org/abs/1801.06176.
  29. 29.Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, Wanjun Zhong, Kuanye Li, Jiale Yang, Yu Miao, Woyu Lin, Longxiang Liu, Xu Jiang, Qianli Ma, Jingyu Li, Xiaojun Xiao, Kai Cai, Chuang Li, Yaowei Zheng, Chaolin Jin, Chen Li, Xiao Zhou, Minchao Wang, Haoli Chen, Zhaojian Li, Haihua Yang, Haifeng Liu, Feng Lin, Tao Peng, Xin Liu, and Guang Shi. Ui-tars: Pioneering automated gui interaction with native agents, 2025. URL https://arxiv.org/abs/2501.12326.
  30. 30.Qwen, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. Qwen2.5 technical report, 2025. URL https://arxiv.org/abs/2412.15115.
  31. 31.Qwen Team. Qwen3 technical report, 2025. URL https://arxiv.org/abs/2505.09388.
  32. 32.Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William Bishop, Wei Li, Folawiyo Campbell-Ajala, Daniel Toyama, Robert Berry, Divya Tyamagundlu, Timothy Lillicrap, and Oriana Riva. Androidworld: A dynamic benchmarking environment for autonomous agents, 2025. URL https://arxiv.org/abs/2405.14573.
  33. 33.Max-Philipp B. Schrader. gym-sokoban. https://github.com/mpSchrader/gym-sokoban, 2018.
  34. 34.Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver. Mastering atari, go, chess and shogi by planning with a learned model. Nature, 588 (7839):604–609, December 2020. ISSN 1476-4687. doi: 10.1038/s41586-020-03051-4. URL http://dx.doi.org/10.1038/s41586-020-03051-4.
  35. 35.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024. URL https://arxiv.org/abs/2402.03300.
  36. 36.Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning, 2023. URL https://arxiv.org/abs/2303.11366.
  37. 37.Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton, Samy Bengio, and Mehrdad Farajtabar. The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity, 2025. URL https://arxiv.org/abs/2506.06941.
  38. 38.Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning, 2021. URL https://arxiv.org/abs/2010.03768.
  39. 39.Richard S. Sutton. Dyna, an integrated architecture for learning, planning, and reacting. SIGART Bull., 2(4):160–163, July 1991. ISSN 0163-5719. doi: 10.1145/122344.122377. URL https://doi.org/10.1145/122344.122377.
  40. 40.Edward C Tolman. Cognitive maps in rats and men. Psychological review, 55(4):189, 1948.
  41. 41.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models, 2023. URL https://arxiv.org/abs/2305.16291.
  42. 42.Xinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang, Tianbao Xie, Junli Wang, Jiaqi Deng, Xiaole Guo, Yiheng Xu, Chen Henry Wu, Zhennan Shen, Zhuokai Li, Ryan Li, Xiaochuan Li, Junda Chen, Boyuan Zheng, Peihang Li, Fangyu Lei, Ruisheng Cao, Yeqiao Fu, Dongchan Shin, Martin Shin, Jiarui Hu, Yuyan Wang, Jixuan Chen, Yuxiao Ye, Danyang Zhang, Dikang Du, Hao Hu, Huarong Chen, Zaida Zhou, Haotian Yao, Ziwei Chen, Qizheng Gu, Yipu Wang, Heng Wang, Diyi Yang, Victor Zhong, Flood Sung, Y. Charles, Zhilin Yang, and Tao Yu. Opencua: Open foundations for computer-use agents, 2025a. URL https://arxiv.org/abs/2508.09123.
  43. 43.Zihan Wang, Kangrui Wang, Qineng Wang, Pingyue Zhang, Linjie Li, Zhengyuan Yang, Xing Jin, Kefan Yu, Minh Nhat Nguyen, Licheng Liu, Eli Gottlieb, Yiping Lu, Kyunghyun Cho, Jiajun Wu, Li Fei-Fei, Lijuan Wang, Yejin Choi, and Manling Li. Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning, 2025b. URL https://arxiv.org/abs/2504.20073.
  44. 44.Yuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux, Lingming Zhang, Daniel Fried, Gabriel Synnaeve, Rishabh Singh, and Sida I. Wang. Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution, 2025a. URL https://arxiv.org/abs/2502.18449.
  45. 45.Zhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang, Qin Lu, Liang Qiu, Changlong Yu, Puyang Xu, Chao Zhang, Bing Yin, Hyokun Yun, and Lihong Li. Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning, 2025b. URL https://arxiv.org/abs/2505.16421.
  46. 46.Yuexin Wu, Xiujun Li, Jingjing Liu, Jianfeng Gao, and Yiming Yang. Switch-based active deep dyna-q: Efficient adaptive planning for task-completion dialogue policy learning, 2018. URL https://arxiv.org/abs/1811.07550.
  47. 47.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024. URL https://arxiv.org/abs/2404.07972.
  48. 48.Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, and Caiming Xiong. Aguvis: Unified pure vision agents for autonomous gui interaction, 2025. URL https://arxiv.org/abs/2412.04454.
  49. 49.John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering, 2024. URL https://arxiv.org/abs/2405.15793.
  50. 50.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models, 2023a. URL https://arxiv.org/abs/2305.10601.
  51. 51.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models, 2023b. URL https://arxiv.org/abs/2210.03629.
  52. 52.Xiao Yu, Maximillian Chen, and Zhou Yu. Prompt-based monte-carlo tree search for goal-oriented dialogue policy planning, 2023. URL https://arxiv.org/abs/2305.13660.
  53. 53.Xiao Yu, Baolin Peng, Vineeth Vajipey, Hao Cheng, Michel Galley, Jianfeng Gao, and Zhou Yu. Exact: Teaching ai agents to explore with reflective-mcts and exploratory learning, 2025a. URL https://arxiv.org/abs/2410.02052.
  54. 54.Xiao Yu, Baolin Peng, Ruize Xu, Michel Galley, Hao Cheng, Suman Nath, Jianfeng Gao, and Zhou Yu. Dyna-think: Synergizing reasoning, acting, and world model simulation in ai agents, 2025b. URL https://arxiv.org/abs/2506.00320.
  55. 55.Aohan Zeng, Mingdao Liu, Rui Lu, Bowen Wang, Xiao Liu, Yuxiao Dong, and Jie Tang. Agenttuning: Enabling generalized agent abilities for llms, 2023. URL https://arxiv.org/abs/2310.12823.
  56. 56.Jianguo Zhang, Tian Lan, Ming Zhu, Zuxin Liu, Thai Hoang, Shirley Kokane, Weiran Yao, Juntao Tan, Akshara Prabhakar, Haolin Chen, Zhiwei Liu, Yihao Feng, Tulika Awalgaonkar, Rithesh Murthy, Eric Hu, Zeyuan Chen, Ran Xu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Silvio Savarese, and Caiming Xiong. xlam: A family of large action models to empower ai agent systems, 2024. URL https://arxiv.org/abs/2409.03215.
  57. 57.Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. Gpt-4v(ision) is a generalist web agent, if grounded, 2024. URL https://arxiv.org/abs/2401.01614.
  58. 58.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https://arxiv.org/abs/2306.05685.
  59. 59.Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. Language agent tree search unifies reasoning acting and planning in language models, 2024a. URL https://arxiv.org/abs/2310.04406.
  60. 60.Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents, 2024b. URL https://arxiv.org/abs/2307.13854.
  61. 61.Lixin Zou, Long Xia, Pan Du, Zhuo Zhang, Ting Bai, Weidong Liu, Jian-Yun Nie, and Dawei Yin. Pseudo dyna-q: A reinforcement learning framework for interactive recommendation. In Proceedings of the 13th International Conference on Web Search and Data Mining, WSDM ’20, pp. 816–824, New York, NY, USA, 2020. Association for Computing Machinery. ISBN 9781450368223. doi: 10.1145/3336191.3371801. URL https://doi.org/10.1145/3336191.3371801.

Citation

MLA
Yu, X., et al. “Dyna-Mind: Learning to Simulate from Experience for Better AI Agents”. arXiv, 2025, http://arxiv.org/abs/2510.09577v1.
APA
Yu, X., Peng, B., Galley, M., Cheng, H., Wu, Q., Kulkarni, J., Nath, S., Yu, Z., & Gao, J. (2025). Dyna-Mind: Learning to Simulate from Experience for Better AI Agents. arXiv. http://arxiv.org/abs/2510.09577v1
Chicago
Yu, X., B. Peng, M. Galley, et al. 2025. “Dyna-Mind: Learning to Simulate from Experience for Better AI Agents”. arXiv. http://arxiv.org/abs/2510.09577v1.
Harvard
Yu, X. et al. (2025) “Dyna-Mind: Learning to Simulate from Experience for Better AI Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2510.09577v1.
Vancouver
1. Yu X, Peng B, Galley M, Cheng H, Wu Q, Kulkarni J, Nath S, Yu Z, Gao J (2025) Dyna-Mind: Learning to Simulate from Experience for Better AI Agents. arXiv

BibTeX

@article{yu2025dyna,
  title = {Dyna-Mind: Learning to Simulate from Experience for Better AI Agents},
  author = {Yu, Xiao and Peng, Baolin and Galley, Michel and Cheng, Hao and Wu, Qianhui and Kulkarni, Janardhan and Nath, Suman and Yu, Zhou and Gao, Jianfeng},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2510.09577v1},
  eprint = {2510.09577}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/