MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents
Xintao DingXinrui WangYifan YangHao WuShiqi JiangQianxi ZhangLiang MiHanxin ZhuKunxi LiYunxin Liu
Introduces a state-conditioned memory compilation framework for embodied agents that dynamically translates past experience into real-time text and latent guidance, boosting task success by up to 129% while cutting per-step latency by 60%.
Deploying autonomous embodied agents in persistent environments requires leveraging past experiences to handle long-horizon tasks and avoid repeating past errors. Standard architectures currently rely on Ahead-of-time Monolithic Memory Injection (AMMI), which retrieves past experiences and injects them as a static prompt context at the start of an episode. However, because embodied environments evolve continuously across visual and action streams, this static context quickly becomes misaligned with the agent's real-time state. Smaller and lightweight policy models cannot easily filter this bloated context, suffering from severe attention dilution that frequently degrades execution below no-memory baselines.
The article introduces and evaluates MemCompiler, a State-Conditioned Memory Compilation framework designed to dynamically translate stored past experiences into actionable, step-wise guidance. The primary objective is to demonstrate that compiling memory dynamically based on the agent's real-time state significantly improves task success and runtime efficiency compared to static upfront injection across open-source and closed-source foundation models.
The authors assess MemCompiler through extensive empirical evaluations across three simulated benchmark environments: AlfWorld, EmbodiedBench (EB-ALFRED and EB-Habitat), and ScienceWorld, covering household manipulation, visual navigation, and multi-step scientific reasoning. Testing spans four open-source models (including 14B, 27B, and 32B visual-language variants) and two closed-source systems (GPT-5.2 and Gemini-3-Flash). The framework employs a dedicated, lightweight Memory Compiler trained via supervised fine-tuning and reinforcement learning. This compiler tracks a structured Brief State covering task progress and environmental beliefs, then dynamically outputs targeted text guidance alongside latent perceptual tokens (Soft-Mem) directly into the execution policy model.
The findings establish that state-conditioned compilation consistently outperforms static memory injection. First, MemCompiler delivers substantial task success improvements across every tested open-source backbone, reaching relative gains of up to +110% on EB-ALFRED and +129% on ScienceWorld compared to unaugmented baselines. Second, traditional static AMMI baselines fail extensively on open-source executors, degrading performance in more than half of test settings due to attention dilution (for instance, dropping success by over 80% on certain benchmark tasks). Third, MemCompiler enables mid-sized open-source models to approach or exceed the performance of leading proprietary models operating without memory; for example, an open-source 27B model reached a 91.45% success rate on AlfWorld, matching Gemini-3-Flash. Fourth, dynamically passing only relevant guidance cuts executor input tokens by roughly 50% to 60% and reduces per-step execution latency from 0.30 seconds down to 0.12 seconds on AlfWorld.
These results indicate that the primary performance bottleneck for embodied agents is not model size or raw retrieval capability, but rather the timing, format, and relevance of how memory is delivered at runtime. By filtering and structuring guidance step by step, organizations can deploy smaller, cost-effective open-source models locally without suffering performance penalties or paying high inference costs associated with bloated context windows and frontier proprietary APIs.
Teams developing embodied AI and robotics systems should shift from static prompt stuffing toward modular, step-wise memory compilation architectures. System designers should adopt structured state tracking (such as belief state tracking) and dual text-latent channels to retain spatial and visual knowledge without cluttering text prompts. Further validation should focus on evaluating MemCompiler in physical hardware deployments, noisier real-world sensory environments, and large-scale, lifelong memory banks, as the current study is bounded by simulated environments, minimal curated memory repositories, and fixed policy backbones.
- Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). MemGPT establishes the foundational paradigm of managing external tiered context and dynamic memory for LLM agents, which MemCompiler directly addresses and redesigns to avoid static ahead-of-time injection.
- Paper: MEMORYLLM: Towards Self-Updatable Large Language Models, Yu Wang et al. (2024). MemoryLLM introduces direct latent-space memory integration for language models, providing core conceptual groundwork for MemCompiler's latent Soft-Mem channel.
- Paper: General Agentic Memory Via Deep Research, B. Y. Yan et al. (2025). This work explores dynamic agentic memory retrieval to replace static context pre-compression, framing the retrieval challenges that MemCompiler reframes as state-conditioned compilation.
- Paper: End-To-End Memory Networks, Sainbayar Sukhbaatar et al. (2015). End-to-End Memory Networks supply the foundational neural memory reading and attention mechanisms over explicit memory stores that modern learned memory controllers build upon.
- Paper: SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills, Zhongxin Guo et al. (2026). SKILL-DISCO complements MemCompiler by discovering and compiling recurring agent execution traces into reusable procedural routines across shared embodied environments like ALFWorld.
- Paper: MemGym: a Long-Horizon Memory Environment for LLM Agents, Wujiang Xu et al. (2026). MemGym provides a standardized evaluation suite to benchmark the long-horizon memory and compression performance of dynamic agent architectures like MemCompiler.
- Paper: LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation, Dongge Han et al. (2026). LEGOMem extends dynamic memory compilation concepts to multi-agent workflow teams by modularizing and allocating procedural memory across orchestrator and task agents.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). This survey contextualizes state-aware memory compilation within broader efficiency and planning paradigms across modern LLM agent systems.
- Paper: Human-Inspired Memory Architecture for LLM Agents, Doga Kerestecioglu et al. (2026). This work builds on multi-tiered agent memory by introducing biologically inspired long-term consolidation and dynamic reconsolidation upon retrieval.
- Paper: MeMo: Memory as a Model, Ryan Wei Heng Quek et al. (2026). MEMO explores an alternative modular paradigm where dedicated memory models internalize and supply knowledge parametrically rather than through dynamic state-conditioned channels.
