MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents

Xintao DingXinrui WangYifan YangHao WuShiqi JiangQianxi ZhangLiang MiHanxin ZhuKunxi LiYunxin Liu

article2026arXiv7 citations

Introduces a state-conditioned memory compilation framework for embodied agents that dynamically translates past experience into real-time text and latent guidance, boosting task success by up to 129% while cutting per-step latency by 60%.

Listen

Deploying autonomous embodied agents in persistent environments requires leveraging past experiences to handle long-horizon tasks and avoid repeating past errors. Standard architectures currently rely on Ahead-of-time Monolithic Memory Injection (AMMI), which retrieves past experiences and injects them as a static prompt context at the start of an episode. However, because embodied environments evolve continuously across visual and action streams, this static context quickly becomes misaligned with the agent's real-time state. Smaller and lightweight policy models cannot easily filter this bloated context, suffering from severe attention dilution that frequently degrades execution below no-memory baselines.

The article introduces and evaluates MemCompiler, a State-Conditioned Memory Compilation framework designed to dynamically translate stored past experiences into actionable, step-wise guidance. The primary objective is to demonstrate that compiling memory dynamically based on the agent's real-time state significantly improves task success and runtime efficiency compared to static upfront injection across open-source and closed-source foundation models.

The authors assess MemCompiler through extensive empirical evaluations across three simulated benchmark environments: AlfWorld, EmbodiedBench (EB-ALFRED and EB-Habitat), and ScienceWorld, covering household manipulation, visual navigation, and multi-step scientific reasoning. Testing spans four open-source models (including 14B, 27B, and 32B visual-language variants) and two closed-source systems (GPT-5.2 and Gemini-3-Flash). The framework employs a dedicated, lightweight Memory Compiler trained via supervised fine-tuning and reinforcement learning. This compiler tracks a structured Brief State covering task progress and environmental beliefs, then dynamically outputs targeted text guidance alongside latent perceptual tokens (Soft-Mem) directly into the execution policy model.

The findings establish that state-conditioned compilation consistently outperforms static memory injection. First, MemCompiler delivers substantial task success improvements across every tested open-source backbone, reaching relative gains of up to +110% on EB-ALFRED and +129% on ScienceWorld compared to unaugmented baselines. Second, traditional static AMMI baselines fail extensively on open-source executors, degrading performance in more than half of test settings due to attention dilution (for instance, dropping success by over 80% on certain benchmark tasks). Third, MemCompiler enables mid-sized open-source models to approach or exceed the performance of leading proprietary models operating without memory; for example, an open-source 27B model reached a 91.45% success rate on AlfWorld, matching Gemini-3-Flash. Fourth, dynamically passing only relevant guidance cuts executor input tokens by roughly 50% to 60% and reduces per-step execution latency from 0.30 seconds down to 0.12 seconds on AlfWorld.

These results indicate that the primary performance bottleneck for embodied agents is not model size or raw retrieval capability, but rather the timing, format, and relevance of how memory is delivered at runtime. By filtering and structuring guidance step by step, organizations can deploy smaller, cost-effective open-source models locally without suffering performance penalties or paying high inference costs associated with bloated context windows and frontier proprietary APIs.

Teams developing embodied AI and robotics systems should shift from static prompt stuffing toward modular, step-wise memory compilation architectures. System designers should adopt structured state tracking (such as belief state tracking) and dual text-latent channels to retain spatial and visual knowledge without cluttering text prompts. Further validation should focus on evaluating MemCompiler in physical hardware deployments, noisier real-world sensory environments, and large-scale, lifelong memory banks, as the current study is bounded by simulated environments, minimal curated memory repositories, and fixed policy backbones.

arXiv: 2605.07594air-embodied-brain/MemCompiler
  • Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). MemGPT establishes the foundational paradigm of managing external tiered context and dynamic memory for LLM agents, which MemCompiler directly addresses and redesigns to avoid static ahead-of-time injection.
  • Paper: MEMORYLLM: Towards Self-Updatable Large Language Models, Yu Wang et al. (2024). MemoryLLM introduces direct latent-space memory integration for language models, providing core conceptual groundwork for MemCompiler's latent Soft-Mem channel.
  • Paper: General Agentic Memory Via Deep Research, B. Y. Yan et al. (2025). This work explores dynamic agentic memory retrieval to replace static context pre-compression, framing the retrieval challenges that MemCompiler reframes as state-conditioned compilation.
  • Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). End-to-End Memory Networks supply the foundational neural memory reading and attention mechanisms over explicit memory stores that modern learned memory controllers build upon.
Cover for MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents

Abstract

Existing memory systems for embodied agents typically inject retrieved memory as static context at episode start, a paradigm we term Ahead-of-time Monolithic Memory Injection (AMMI). However, this static design quickly becomes misaligned with the agent's evolving state and may degrade lightweight executors below the no-memory baseline. To address this, we propose MemCompiler, which reframes memory utilization as State-Conditioned Memory Compilation. A learned Memory Compiler reads a structured Brief State capturing the agent's current execution state and dynamically selects and compiles only relevant memory into executable guidance. This guidance is delivered through a text channel and a latent Soft-Mem channel that preserves perceptual information not expressible in text. Across Alf World, EmbodiedBench, and ScienceWorld, MemCompiler consistently improves over no-memory across open-source backbones (up to +129%), matches or approaches frontier closed-source systems, and reduces per-step latency by 60%, demonstrating that state-aware memory compilation improves both effectiveness and efficiency.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Problem Formulation of MemCompiler
  • 3.2 Methodology of MemCompiler
  • 3.2.1 Overview and Motivation
  • 3.2.2 When to Compile
  • 3.2.3 How to Compile
  • 4 Experiment
  • 4.1 Experimental Setup
  • 4.2 Main Results
  • 5 Ablation
  • 6 Conclusion
  • References
  • A Progress Rate on EmbodiedBench
  • B Qualitative Case Studies
  • C Experimental Details
  • C.1 Benchmarks and Datasets Details
  • C.2 Dataset Construction and Annotation
  • C.3 Training Details and Hyperparameters
  • C.3.1 Supervised Fine-Tuning (SFT) Stage
  • C.3.2 Reinforcement Learning (GRPO) Stage
  • D Prompt Templates

Knowls

  1. Knowl 1 — State-Conditioned Memory Compilation Framework

    model/method

    State-Conditioned Memory Compilation (SCMC) is an embodied agent memory paradigm that replaces Ahead-of-time Monolithic Memory Injection (AMMI). In standard AMMI, a candidate task memory pool M=Retrieve(B,g)\mathcal{M} = \text{Retrieve}(\mathcal{B}, g) retrieved from a cross-episode memory bank B\mathcal{B} for goal gg is statically injected into the context of an executor policy π(at∣ot,M)\pi(a_t \mid o_t, \mathcal{M}) at episode start and held fixed, which degrades performance as the environment state evolves.

    SCMC retains M\mathcal{M} as an external source library and introduces a learned Memory Compiler policy πC\pi_C. At each timestep tt, the Memory Compiler reads the agent's runtime state st=(ot,bt)s_t = (o_t, b_t)—comprising visual observation ot∈RH×W×Co_t \in \mathbb{R}^{H \times W \times C} and a dynamically maintained Brief State btb_t—and selectively compiles state-appropriate executable guidance m∗,tm^{*,t}:

    m∗,t=πC(st,M)m^{*,t} = \pi_C(s_t, \mathcal{M})

    This formulation models the optimal memory entry as state-dependent:

    m∗,t=arg⁡max⁡m∈MEτ∼π(⋅∣st,m)[1(success(τ))]m^{*,t} = \arg\max_{m \in \mathcal{M}} \mathbb{E}_{\tau \sim \pi(\cdot \mid s_t, m)} [\mathbf{1}(\text{success}(\tau))]

    ensuring that the downstream Executor receives only concise, stage-relevant guidance aligned with its active subgoals.

  2. Knowl 2 — Brief State Representation and Dynamic Maintenance Operations

    model/method

    Brief State bt=ϕ(g,a0:t−1,o0:t−1)b_t = \phi(g, a_{0:t-1}, o_{0:t-1}) is a structured, task-relevant runtime context maintained and updated by the Memory Compiler throughout an embodied episode to provide explicit progress and environment signals.

    It consists of two structured dimensions:

    1. Task Progress State: encodes the overall goal gg, completed subgoals, current active subgoal, and pending subgoals to determine whether a candidate memory entry targets the immediate step.
    2. Environment Belief: records discovered object locations, agent status, and verified environmental facts to evaluate whether a memory entry's prerequisites hold in the current scene.

    When state modification is triggered, the state updates recursively via bt+1=Apply(Δbt,bt)b_{t+1} = \text{Apply}(\Delta b_t, b_t), where Δbt\Delta b_t executes one of four atomic operations:

    • CREATE: registers a newly discovered entity or environmental fact.
    • UPDATE: replaces outdated belief information with revised content.
    • DELETE: removes invalidated facts from memory.
    • FOLD: consolidates and summarizes completed subgoals into persistent beliefs once item count exceeds 10.
  3. Knowl 3 — Memory Compiler Four-Way Decision Space

    model/method

    At each decision step tt, given the runtime state st=(ot,bt)s_t = (o_t, b_t) and the candidate memory pool M\mathcal{M}, the Memory Compiler policy πC\pi_C chooses from a four-element discrete output space combining executor intervention and belief maintenance:

    πC(st,M)∈{m∗,t,Δbt,(m∗,t,Δbt),∅}\pi_C(s_t, \mathcal{M}) \in \{ m^{*,t}, \Delta b_t, (m^{*,t}, \Delta b_t), \emptyset \}

    • EXPERIENCE (m∗,tm^{*,t}): emits strategic memory guidance directly to the Executor without modifying the Brief State btb_t.
    • BRIEF (Δbt\Delta b_t): emits a belief state update only, updating btb_t for subsequent timesteps without intervening in the Executor's current action.
    • HYBRID ((m∗,t,Δbt)(m^{*,t}, \Delta b_t)): simultaneously outputs memory guidance m∗,tm^{*,t} and belief update Δbt\Delta b_t.
    • NOACTION (∅\emptyset): withholds guidance and state updates when the Executor is progressing normally and no entry in M\mathcal{M} exhibits sufficient utility for the current state sts_t.
  4. Knowl 4 — Soft-Mem Dual-Channel Memory Delivery Mechanism

    model/method

    Soft-Mem delivers compiled memory guidance m∗,tm^{*,t} through two complementary channels: an interpretable textual guidance channel mtext∗,tm^{*,t}_{\text{text}} and a continuous latent soft channel msoft∗,tm^{*,t}_{\text{soft}} designed to convey visual, spatial, and affordance features that natural language cannot readily express.

    After autoregressively generating mtext∗,tm^{*,t}_{\text{text}}, the Memory Compiler continues processing NN appended special [UNK] tokens. The final-layer hidden representations at these token positions, h∈RN×dMCh \in \mathbb{R}^{N \times d_{\text{MC}}}, attend over the multimodal task memory M\mathcal{M}. A trainable projection head fϕf_\phi maps hh to Gaussian parameters in the Executor's embedding space:

    (μ,log⁡σ)=fϕ(h),msoft∗,t=μ+σ⊙ϵ,ϵ∼N(0,I)(\boldsymbol{\mu}, \log \boldsymbol{\sigma}) = f_\phi(h), \quad m^{*,t}_{\text{soft}} = \boldsymbol{\mu} + \boldsymbol{\sigma} \odot \boldsymbol{\epsilon}, \quad \boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})

    The resulting soft tokens msoft∗,t∈RN×dbasem^{*,t}_{\text{soft}} \in \mathbb{R}^{N \times d_{\text{base}}} replace designated placeholder embeddings in the Executor's input sequence, supplying perceptual representations directly to the frozen Executor backbone.

  5. Knowl 5 — Supervised Fine-Tuning Objective for Dual-Channel Memory Compilation

    equation

    In the Supervised Fine-Tuning (SFT) phase of MemCompiler, the Memory Compiler and the Soft-Mem projection network fϕf_\phi are optimized using a multi-task loss function while the downstream Executor policy remains frozen:

    LSFT=Laction+Ltext+λsoftLsoft\mathcal{L}_{\text{SFT}} = \mathcal{L}_{\text{action}} + \mathcal{L}_{\text{text}} + \lambda_{\text{soft}} \mathcal{L}_{\text{soft}}

    where Ltext\mathcal{L}_{\text{text}} is the standard next-token prediction loss over teacher-annotated text guidance, and Laction\mathcal{L}_{\text{action}} is the cross-entropy action loss computed from the Executor's action prediction with injected soft tokens msoft∗,tm^{*,t}_{\text{soft}}. The soft-token loss Lsoft\mathcal{L}_{\text{soft}} is defined as:

    Lsoft=−1N∑i=1Nlog⁡pθ(e~i)+λ∣cos⁡(hˉsoft,sg[hˉtext])∣\mathcal{L}_{\text{soft}} = -\frac{1}{N}\sum_{i=1}^N \log p_\theta(\tilde{e}_i) + \lambda \left| \cos\left(\bar{h}_{\text{soft}}, \text{sg}[\bar{h}_{\text{text}}]\right) \right|

    where −1N∑i=1Nlog⁡pθ(e~i)-\frac{1}{N}\sum_{i=1}^N \log p_\theta(\tilde{e}_i) is a semantic alignment objective grounding the latent tokens in memory representations, and λ∣cos⁡(hˉsoft,sg[hˉtext])∣\lambda |\cos(\bar{h}_{\text{soft}}, \text{sg}[\bar{h}_{\text{text}}])| is an orthogonality constraint between mean soft hidden state hˉsoft\bar{h}_{\text{soft}} and stop-gradient mean text representation sg[hˉtext]\text{sg}[\bar{h}_{\text{text}}] that prevents the latent channel from collapsing into a copy of the textual channel.

  6. Knowl 6 — GRPO Policy Refinement for Memory Compilation

    model/method

    Following supervised fine-tuning, the Memory Compiler policy πθ\pi_\theta and Soft-Mem projection are optimized using Group Relative Policy Optimization (GRPO) while keeping the Executor frozen.

    For each training iteration, N=4N=4 tasks are sampled with K=4K=4 rollout trajectories per task, producing a group of G=16G=16 episodes. A sparse binary reward ri∈{0,1}r_i \in \{0, 1\} is awarded based exclusively on final task success (1.01.0 for success, 00 for failure), without step-level reward shaping.

    The policy objective is minimized according to:

    LGRPO=−1G∑i=1Gmin⁡(πθ(oi∣s)πθold(oi∣s)Ai,  clip(πθ(oi∣s)πθold(oi∣s),1−ϵlow,1+ϵhigh)Ai)\mathcal{L}_{\text{GRPO}} = -\frac{1}{G} \sum_{i=1}^G \min\left( \frac{\pi_\theta(o_i \mid s)}{\pi_{\theta_{\text{old}}}(o_i \mid s)} A_i, \; \text{clip}\left( \frac{\pi_\theta(o_i \mid s)}{\pi_{\theta_{\text{old}}}(o_i \mid s)}, 1-\epsilon_{\text{low}}, 1+\epsilon_{\text{high}} \right) A_i \right)

    where Ai=ri−mean(r1:G)std(r1:G)A_i = \frac{r_i - \text{mean}(r_{1:G})}{\text{std}(r_{1:G})} is the group-normalized advantage, and asymmetric clipping is applied with ϵlow=0.2\epsilon_{\text{low}} = 0.2 and ϵhigh=0.28\epsilon_{\text{high}} = 0.28 with no KL penalty (β=0\beta = 0).

  7. Knowl 7 — Performance of MemCompiler on EmbodiedBench

    data/table

    The table below details the Success Rate (%) across six capability dimensions—Base Ability (Base), Commonsense Reasoning (Com.), Complex Instruction Following (Comp.), Visual Recognition (Vis.), Spatial Reasoning (Spat.), and Long-Horizon Planning (Long)—on EmbodiedBench (EB-ALFRED and EB-Habitat).

    Executor Method EB-ALFRED (Avg / Base / Com / Comp / Vis / Spat / Long)
    Qwen-2.5-VL-32B No Mem 18.00 22 26 24 12 22 2
    Mem0 14.66 20 18 26 12 10 2
    G-Mem 7.67 4 14 6 16 6 0
    A-Mem 3.00 2 4 4 6 2 0
    LangMem 14.33 22 20 26 6 12 0
    MemGen 14.33 14 10 16 6 8 0
    Ours 35.33 42 50 44 38 34 4
    Qwen-3-VL-32B No Mem 19.00 24 24 28 24 14 0
    Mem0 30.33 40 42 40 26 26 8
    G-Mem 25.33 34 36 38 24 14 6
    A-Mem 32.33 46 50 34 30 30 4
    LangMem 31.33 38 44 44 28 26 8
    MemGen 14.22 22 18 16 6 20 4
    Ours 40.00 52 54 48 38 38 10
    Executor Method EB-Habitat (Avg / Base / Com / Comp / Vis / Spat / Long)
    Qwen-2.5-VL-32B No Mem 36.33 78 22 52 26 30 10
    Mem0 38.33 68 34 44 38 30 16
    G-Mem 14.67 38 12 8 8 20 2
    A-Mem 8.33 24 8 6 8 4 0
    LangMem 38.67 76 28 44 40 28 16
    MemGen 21.00 54 8 18 18 24 4
    Ours 54.67 84 54 56 60 36 38
    Qwen-3-VL-32B No Mem 34.67 78 26 28 34 28 14
    Mem0 38.33 84 24 26 42 34 20
    G-Mem 44.00 88 24 36 50 32 34
    A-Mem 36.67 86 12 38 24 36 24
    LangMem 38.33 86 26 28 40 32 18
    MemGen 25.67 50 14 22 36 22 10
    Ours 55.67 90 58 44 64 40 38

    While static AMMI baselines (Mem0, G-Mem, A-Mem, LangMem) often cause severe performance drops on open-source vision-language backbones (dropping up to -83.3% relative to No Mem), MemCompiler delivers consistent improvements (+96.2% on Qwen-2.5-VL-32B on EB-ALFRED and +110.0% on Qwen-3-VL-32B on EB-ALFRED), surpassing frontier closed-source baselines without memory on EB-Habitat (GPT-5.2 No Mem: 40.33%, Gemini-3-Flash No Mem: 51.00%).

  8. Knowl 8 — Performance Comparison on AlfWorld and ScienceWorld

    data/table

    The table below presents task success rates (%) on AlfWorld (134 unseen tasks across 120 rooms) and ScienceWorld (90 test instances across 16 tasks) across closed-source and open-source models.

    Executor Method AlfWorld (%) ScienceWorld (%)
    GPT-5.2 No Mem 70.89 25.55
    G-Mem 86.56 51.51
    Mem0 76.12 22.22
    LangMem 88.72 46.67
    A-Mem 92.54 50.00
    Gemini-3-Flash No Mem 92.54 45.55
    G-Mem 90.30 44.44
    Mem0 91.04 42.86
    LangMem 92.54 50.82
    A-Mem 97.01 47.78
    Qwen-2.5-14B No Mem 46.27 31.11
    G-Mem 67.16 32.22
    Mem0 12.69 27.78
    LangMem 36.57 15.56
    A-Mem 11.19 34.44
    MemGen 42.54 24.65
    Ours 82.16 46.67
    Qwen-3.5-27B No Mem 61.19 21.11
    G-Mem 74.62 32.22
    Mem0 62.69 26.67
    LangMem 62.69 26.67
    A-Mem 88.81 32.22
    MemGen 41.04 14.32
    Ours 91.45 48.44

    MemCompiler improves open-source model success by +77.6% (Qwen-2.5-14B) and +49.5% (Qwen-3.5-27B) on AlfWorld, and +50.0% and +129.0% on ScienceWorld. Qwen-3.5-27B equipped with MemCompiler achieves 91.45% on AlfWorld and 48.44% on ScienceWorld, matching or exceeding the zero-memory performance of frontier proprietary models (Gemini-3-Flash at 92.54% and 45.55%, respectively).

  9. Knowl 9 — Attention Dilution Phenomenon in Static Memory Delivery

    empirical result

    Tracking the Executor's mean pre-softmax attention logit over memory tokens across 30 steps on AlfWorld demonstrates the attention dilution effect inherent to Ahead-of-time Monolithic Memory Injection (AMMI):

    • Under AMMI, attention allocated to memory tokens decays monotonically as execution proceeds, falling by 86.7% from step 0 to step 29. Because the initial monolithic context cannot adapt to changing environmental observations and subgoal completions, static memory entries become increasingly irrelevant and are neglected by the Executor.
    • Under State-Conditioned Memory Compilation (SCMC), attention logits remain high and stable throughout all 30 steps, confirming that dynamic compilation delivers state-aligned memory at each step without suffering from attention dilution.
  10. Knowl 10 — Component Ablation Analysis of MemCompiler

    data/table

    An ablation study conducted with Qwen-2.5-14B on AlfWorld (Success Rate %) and ScienceWorld (Success Rate % / Partial Score) isolates the contribution of each core component of MemCompiler:

    Variant AlfWorld (%) ScienceWorld (%)
    MemCompiler (Full) 82.16 46.67 / 69.33
    w/o Soft-Mem (text only) 78.93 42.22 / 65.15
    w/o Brief State (raw obs as state) 73.94 38.50 / 60.67
    AMMI (static injection, same memory) 35.03 27.93 / 52.68
    No Memory 46.27 31.11 / 45.26

    Removing step-level dynamic compilation (AMMI static injection) causes the most severe drop (from 82.16% down to 35.03% on AlfWorld, falling well below the 46.27% no-memory baseline). Replacing the structured Brief State with raw observation histories degrades performance to 73.94%, and eliminating the Soft-Mem latent channel reduces success to 78.93% on AlfWorld and 42.22% on ScienceWorld.

  11. Knowl 11 — Soft-Mem Design Parameter Ablations

    data/table

    The impact of latent token count NN, orthogonality constraints, and projection architectures within Soft-Mem evaluated on EmbodiedBench EB-Habitat:

    Variant Average (%) Spatial (%) Visual (%)
    N=0N = 0 (text only) 47.11 32 42
    N=4N = 4 51.09 28 50
    N=8N = 8 52.83 34 56
    N=16N = 16 (default) 54.67 36 60
    N=32N = 32 53.17 36 50
    N=16N = 16, w/o orth. constraint 50.33 32 52
    N=16N = 16, linear proj. (no Gaussian) 46.23 30 48

    Key observations:

    1. Scaling the soft token count peaks at N=16N=16 (54.67% average), with larger counts (N=32N=32) experiencing diminishing returns and degradation on visual recognition (50%).
    2. Removing the orthogonality constraint causes a 4.34% average drop (and an 8% drop on Visual), confirming that unconstrained latent vectors collapse into redundant textual features.
    3. Replacing Gaussian reparameterization with a linear projection drops average success to 46.23%, as deterministic projections inhibit exploration during GRPO reinforcement learning.
  12. Knowl 12 — Inference Efficiency and Latency Reduction

    data/table

    Computational efficiency and latency comparison between AMMI and SCMC on the AlfWorld benchmark evaluated on an NVIDIA A100 GPU using vLLM:

    Method Compiler Latency (s) Executor Latency (s) Exec In Tokens Exec Out Tokens Compiler Total Tokens Success Rate (%)
    AMMI - 0.30 2481.1 18.2 - 35.03
    SCMC 0.093 0.12 1003.2 8.2 2081.4 82.16

    SCMC reduces Executor input tokens by 59.6% (from 2481.1 to 1003.2 tokens) and cuts Executor per-step latency by 60.0% (from 0.30s to 0.12s). Although SCMC introduces a 0.093s Memory Compiler execution step, the combined step latency remains lower than monolithic injection while dramatically increasing task success rate from 35.03% to 82.16%.

Coverage note — Specific benchmark-tailored prompt template texts (Appendix D) were omitted as implementation artifacts rather than core theoretical or architectural contributions.

References

  1. 1.Ayush Agrawal, Raghav Prabhakar, Anirudh Goyal, and Dianbo Liu. Physical reasoning and object planning for household embodied agents. arXiv preprint arXiv:2311.13577, 2023.
  2. 2.Zhaohan Feng, Ruiqi Xue, Lei Yuan, Yang Yu, Ning Ding, Meiqin Liu, Bingzhao Gao, Jian Sun, Xinhu Zheng, and Gang Wang. Multi-agent embodied ai: Advances and future directions. Science China Information Sciences, 69(5):151202, 2026.
  3. 3.Gemini Robotics Team, Abbas Abdolmaleki, Saminda Abeyruwan, Joshua Ainslie, Jean-Baptiste Alayrac, Montserrat Gonzalez Arenas, Ashwin Balakrishna, Nathan Batchelor, Alex Bewley, Jeff Bingham, et al. Gemini robotics 1.5: Pushing the frontier of generalist robots with advanced embodied reasoning, thinking, and motion transfer. arXiv preprint arXiv:2510.03342, 2025.
  4. 4.Marc Glocker, Peter Hönig, Matthias Hirschmanner, and Markus Vincze. Llm-empowered embodied agent for memory-augmented task planning in household robotics. arXiv preprint arXiv:2504.21716, 2025.
  5. 5.Andrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Singh Chaplot, Oleksandr Maksymets, et al. Habitat 2.0: Training home assistants to rearrange their habitat. Advances in neural information processing systems, 34:251–266, 2021.
  6. 6.Andrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure, Rin Metcalf, Walter Talbott, Natalie Mackraz, R Devon Hjelm, and Alexander T Toshev. Large language models as generalizable policies for embodied tasks. In The Twelfth International Conference on Learning Representations, 2023.
  7. 7.Kaizhi Zheng, Xiaotong Chen, Odest Chadwicke Jenkins, and Xin Wang. Vlmbench: A compositional benchmark for vision-and-language manipulation. Advances in Neural Information Processing Systems, 35:665–678, 2022.
  8. 8.Zouying Cao, Jiaji Deng, Li Yu, Weikang Zhou, Zhaoyang Liu, Bolin Ding, and Hai Zhao. Remember me, refine me: A dynamic procedural memory framework for experience-driven agent evolution. arXiv preprint arXiv:2512.10696, 2025.
  9. 9.Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Jinhe Bi, Kristian Kersting, Jeff Z Pan, et al. Memory-r1: Enhancing large language model agents to manage and utilize memories via reinforcement learning. arXiv preprint arXiv:2508.19828, 2025.
  10. 10.Siru Ouyang, Jun Yan, I Hsu, Yanfei Chen, Ke Jiang, Zifeng Wang, Rujun Han, Long T Le, Samira Daruki, Xiangru Tang, et al. Reasoningbank: Scaling agent self-evolving with reasoning memory. arXiv preprint arXiv:2509.25140, 2025.
  11. 11.Alireza Rezazadeh, Zichao Li, Wei Wei, and Yujia Bao. From isolated conversations to hierarchical schemas: Dynamic tree memory representation for llms. arXiv preprint arXiv:2410.14052, 2024.
  12. 12.Thomas Limbacher and Robert Legenstein. H-mem: Harnessing synaptic plasticity with hebbian memory networks. Advances in Neural Information Processing Systems, 33:21627–21637, 2020.
  13. 13.Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413, 2025.
  14. 14.Yaxiong Wu, Yongyue Zhang, Sheng Liang, and Yong Liu. Sgmem: Sentence graph memory for long-term conversational agents. arXiv preprint arXiv:2509.21212, 2025.
  15. 15.Mengkang Hu, Tianxing Chen, Qiguang Chen, Yao Mu, Wenqi Shao, and Ping Luo. Hiagent: Hierarchical working memory management for solving long-horizon agent tasks with large language model. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 32779–32798, 2025.
  16. 16.Lin Long, Yichen He, Wentao Ye, Yiyuan Pan, Yuan Lin, Hang Li, Junbo Zhao, and Wei Li. Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory. arXiv preprint arXiv:2508.09736, 2025.
  17. 17.OpenClaw. Openclaw: The ai that actually does things. https://openclaw.ai/, 2026. Accessed: 2026-05-04.
  18. 18.Anthropic. Claude code: Ai-powered coding assistant for developers. https://claude.com/product/claude-code, 2026. Accessed: 2026-05-04.
  19. 19.Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10740–10749, 2020.
  20. 20.Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Martín-Martín, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, et al. Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation. In Conference on Robot Learning, pages 80–93. PMLR, 2023.
  21. 21.Dongyoung Kim, Sumin Park, Huiwon Jang, Jinwoo Shin, Jaehyung Kim, and Younggyo Seo. Robot-r1: Reinforcement learning for enhanced embodied reasoning in robotics. arXiv preprint arXiv:2506.00070, 2025.
  22. 22.Soroush Nasiriany, Sepehr Nasiriany, Abhiram Maddukuri, and Yuke Zhu. Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots. arXiv preprint arXiv:2603.04356, 2026.
  23. 23.Sanjana Srivastava, Chengshu Li, Michael Lingelbach, Roberto Martín-Martín, Fei Xia, Kent Elliott Vainio, Zheng Lian, Cem Gokmen, Shyamal Buch, Karen Liu, et al. Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments. In Conference on robot learning, pages 477–490. PMLR, 2022.
  24. 24.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
  25. 25.Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, pages 1–22, 2023.
  26. 26.Junyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang. Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration. Advances in Neural Information Processing Systems, 37:2686–2710, 2024.
  27. 27.Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. Advances in neural information processing systems, 36:8634–8652, 2023.
  28. 28.Charles Packer, Vivian Fang, Shishir_G Patil, Kevin Lin, Sarah Wooders, and Joseph_E Gonzalez. Memgpt: towards llms as operating systems. 2023.
  29. 29.Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. Expel: Llm agents are experiential learners. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 19632–19642, 2024.
  30. 30.Bernal J Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. Hipporag: Neurobiologically inspired long-term memory for large language models. Advances in neural information processing systems, 37:59532–59569, 2024.
  31. 31.David Kadavy. Digital Zettelkasten: Principles, Methods, & Examples. Kadavy, Inc., 2021.
  32. 32.Sönke Ahrens. How to take smart notes: One simple technique to boost writing, learning and thinking. Sönke Ahrens, 2022.
  33. 33.Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. A-mem: Agentic memory for llm agents. arXiv preprint arXiv:2502.12110, 2025.
  34. 34.Anthropic. Introducing claude opus 4.6. https://www.anthropic.com/news/claude-opus-4-6, 2026. Online; accessed 2026-05-05.
  35. 35.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  36. 36.Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  37. 37.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023.
  38. 38.Zihao Wang, Shaofei Cai, Anji Liu, Yonggang Jin, Jinbing Hou, Bowei Zhang, Haowei Lin, Zhaofeng He, Zilong Zheng, Yaodong Yang, et al. Jarvis-1: Open-world multi-task agents with memory-augmented multimodal language models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(3):1894–1907, 2024.
  39. 39.Guibin Zhang, Muxin Fu, Guancheng Wan, Miao Yu, Kun Wang, and Shuicheng Yan. G-memory: Tracing hierarchical memory for multi-agent systems. arXiv preprint arXiv:2506.07398, 2025.
  40. 40.Tianxin Wei, Noveen Sachdeva, Benjamin Coleman, Zhankui He, Yuanchen Bei, Xuying Ning, Mengting Ai, Yunzhe Li, Jingrui He, Ed H Chi, et al. Evo-memory: Benchmarking llm agent test-time learning with self-evolving memory. arXiv preprint arXiv:2511.20857, 2025.
  41. 41.BY Yan, Chaofan Li, Hongjin Qian, Shuqi Lu, and Zheng Liu. General agentic memory via deep research. arXiv preprint arXiv:2511.18423, 2025.
  42. 42.Mingcong Lei, Honghao Cai, Zezhou Cui, Liangchen Tan, Junkun Hong, Gehan Hu, Shuangyu Zhu, Yimou Wu, Shaohan Jiang, Ge Wang, et al. Robomemory: A brain-inspired multi-memory agentic framework for lifelong learning in physical embodied systems. In NeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI, 2025.
  43. 43.Zeyu Zhang, Rui Li, Xiaoyan Zhao, Yang Zhang, Wenjie Wang, Xu Chen, and Tat-Seng Chua. Nextmem: Towards latent factual memory for llm-based agents. arXiv preprint arXiv:2603.15634, 2026.
  44. 44.Guibin Zhang, Muxin Fu, and Shuicheng Yan. Memgen: Weaving generative latent memory for self-evolving agents. arXiv preprint arXiv:2509.24704, 2025.
  45. 45.Xinlei Yu, Chengming Xu, Guibin Zhang, Zhangquan Chen, Yudong Zhang, Yongbo He, Peng-Tao Jiang, Jiangning Zhang, Xiaobin Hu, and Shuicheng Yan. Vismem: Latent vision memory unlocks potential of vision-language models. arXiv preprint arXiv:2511.11007, 2025.
  46. 46.Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. Alfworld: Aligning text and embodied environments for interactive learning. arXiv preprint arXiv:2010.03768, 2020.
  47. 47.Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, et al. Embodiedbench: Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents. arXiv preprint arXiv:2502.09560, 2025.
  48. 48.Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, and Prithviraj Ammanabrolu. Scienceworld: Is your agent smarter than a 5th grader? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11279–11298, 2022.
  49. 49.LangChain AI. Langmem: A framework for long-term memory in llm agents. https://github.com/langchain-ai/langmem, 2026. Accessed: 2026-05-04.
  50. 50.Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025.
  51. 51.Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025.
  52. 52.Marc-Alexandre Côté, Akos Kádár, Xingdi Yuan, Ben Kybartas, Tavian Barnes, Emery Fine, James Moore, Matthew Hausknecht, Layla El Asri, Mahmoud Adada, et al. Textworld: A learning environment for text-based games. In Workshop on Computer Games, pages 41–75. Springer, 2018.
  53. 53.Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, et al. Ai2-thor: An interactive 3d environment for visual ai. arXiv preprint arXiv:1712.05474, 2017.
  54. 54.Berk Calli, Arjun Singh, Aaron Walsman, Siddhartha Srinivasa, Pieter Abbeel, and Aaron M Dollar. The ycb object and model set: Towards common benchmarks for manipulation research. In 2015 international conference on advanced robotics (ICAR), pages 510–517. IEEE, 2015.
  55. 55.Yu Shang, Yu Li, Keyu Zhao, Likai Ma, Jiahe Liu, Fengli Xu, and Yong Li. Agentsquare: Automatic llm agent search in modular design space. In The Thirteenth International Conference on Learning Representations, 2024.
  56. 56.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. Iclr, 1(2):3, 2022.

Citation

MLA
Ding, X., et al. “MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents”. arXiv, 2026, https://doi.org/10.48550/arxiv.2605.07594.
APA
Ding, X., Wang, X., Yang, Y., Wu, H., Jiang, S., Zhang, Q., Mi, L., Zhu, H., Li, K., Liu, Y., Chen, Z., & Cao, T. (2026). MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents. arXiv. https://doi.org/10.48550/arxiv.2605.07594
Chicago
Ding, X., X. Wang, Y. Yang, et al. 2026. “MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2605.07594.
Harvard
Ding, X. et al. (2026) “MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents”. arXiv. Available at: https://doi.org/10.48550/arxiv.2605.07594.
Vancouver
1. Ding X, Wang X, Yang Y, et al (2026) MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents. https://doi.org/10.48550/arxiv.2605.07594

BibTeX

@misc{https://doi.org/10.48550/arxiv.2605.07594,
  doi = {10.48550/ARXIV.2605.07594},
  url = {https://arxiv.org/abs/2605.07594},
  author = {Ding, Xin and Wang, Xinrui and Yang, Yifan and Wu, Hao and Jiang, Shiqi and Zhang, Qianxi and Mi, Liang and Zhu, Hanxin and Li, Kun and Liu, Yunxin and Chen, Zhibo and Cao, Ting},
  keywords = {Robotics (cs.RO), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents},
  publisher = {arXiv},
  year = {2026},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/