LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation
Dongge HanCamille CouturierDaniel MadrigalXuchao ZhangVictor RuehleSaravan Rajmohan
Introduces LEGOMem, a modular procedural memory framework for multi-agent systems that distributes past execution trajectories across orchestrator and task agents, enabling teams of smaller language models to rival larger models in complex workflow automation.
Organizations are increasingly deploying teams of specialized artificial intelligence agents powered by large language models to automate complex, multi-step office workflows such as document editing, email handling, and calendar management. In typical architectures, a central orchestrator plans workflows and delegates subtasks to specialized software-tool agents. However, these systems currently operate in a stateless manner, solving every task from scratch and discarding execution experience. This inability to reuse operational knowledge leads to repeated errors and limits workflow reliability.
To solve this challenge, the article introduces and evaluates LEGOMem, a modular procedural memory framework designed specifically for multi-agent systems. The primary objective is to demonstrate how extracting, structuring, and reallocating procedural memory—knowledge of how to execute multi-step workflows—across orchestrators and task agents improves coordination, execution accuracy, and overall task success.
LEGOMem uses a two-phase retrieval-augmented approach. First, an offline curation process analyzes execution logs from successfully completed tasks and distills them into modular units: full-task memories containing high-level plans for the orchestrator, and subtask memories detailing specific tool interactions for individual agents. Second, during active task execution, the system retrieves relevant memories based on semantic similarity. The framework was evaluated across 300 multi-step automation tasks from the OfficeBench benchmark across three team configurations: teams using large language models exclusively (GPT-4o), hybrid teams (GPT-4o orchestrator with GPT-4o-mini task agents), and small-model teams (GPT-4o-mini exclusively).
The evaluation produced several critical findings. First, equipping agent teams with LEGOMem consistently boosted task success rates by over 12 percentage points across all configurations, improving full large-model teams from 45.83% to 58.44% and small-model teams from 24.78% to 38.16%. Second, modular procedural memory enabled smaller, lower-cost models to match or exceed the performance of larger, unaugmented models; for instance, the hybrid team with memory achieved a 50.22% success rate, outperforming the large-model team operating without memory. Third, memory placement experiments revealed that orchestrator memory is essential for effective task decomposition and planning, whereas providing memory solely to localized task agents led to substantial performance drops. Finally, procedural memory improved operational efficiency by reducing total execution steps by up to 16.2% on complex tasks and lowering step-level tool failure rates from 27.5% to 22.5%.
These findings have direct strategic implications for deploying automated enterprise workflows. Incorporating procedural memory allows organizations to deploy smaller, cost-effective language models while achieving performance levels previously requiring larger, computationally expensive models. Furthermore, lowering execution steps and action failures directly reduces operating costs, decreases latency, and mitigates the risk of erroneous automated software actions.
Organizations developing or deploying multi-agent workflow systems should implement role-aware procedural memory architectures that prioritize orchestrator-level planning guidance alongside execution-level tool support. For teams utilizing smaller agents, adopting dynamic retrieval or query-rewriting strategies provides additional performance gains by delivering fine-grained execution guidance.
Decision-makers should note that the article’s evaluations were conducted within a controlled benchmark involving simulated office applications (such as email, spreadsheets, and calendar tools) using successful execution traces. While confidence in the benchmark results is high, organizations should pilot the framework within their specific operational environments and explore expanding procedural memory systems to learn from failed executions in open-ended software ecosystems.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). Agent Workflow Memory introduces the foundational concept of extracting and reusing procedural workflows from action trajectories to guide subsequent agent execution.
- Paper: AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Qingyun Wu et al. (2023). AutoGen establishes the multi-agent conversational architecture that underpins collaborative agent orchestration and role specialization in workflow automation.
- Paper: MemGPT: Towards LLMs as Operating Systems, Charles Packer et al. (2023). MemGPT presents the tiered memory architecture and control loop for language models that informs structured procedural memory management.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). This survey provides essential background on the design spaces, memory modules, and task decomposition paradigms of LLM-based autonomous agents.
- Paper: ProAgent: Building Proactive Cooperative Agents with Large Language Models, Ceyao Zhang et al. (2024). ProAgent formulates cooperative multi-agent coordination with specialized memory and belief revision mechanisms for dynamic team planning.
- Paper: ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs, Yujia Qin et al. (2023). ToolLLM details the decision-making and tool-execution mechanics necessary for fine-grained task agents operating in workflow automation benchmarks.
- Paper: MemGym: a Long-Horizon Memory Environment for LLM Agents, Wujiang Xu et al. (2026). MemGym provides a standardized evaluation environment and reward modeling framework to rigorously benchmark the long-horizon memory dynamics explored in LEGOMem.
- Paper: Decentralized Multi-Agent Systems with Shared Context, Yuzhen Mao et al. (2026). DELM explores decentralized multi-agent coordination over shared context, providing an alternative paradigm to LEGOMem's orchestrator-centric procedural memory allocation.
- Paper: Learning to Orchestrate Agents in Natural Language with the Conductor, Stefan Nielsen et al. (2026). The Conductor builds upon multi-agent orchestration principles by training a dedicated model to dynamically assign subtasks and structure agent communication.
- Paper: Toward Efficient Agents: Memory, Tool learning, and Planning, Xiaofang Yang et al. (2026). This survey synthesizes efficiency trade-offs across memory, planning, and multi-agent coordination, placing modular procedural memory systems into a broader architectural perspective.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). This comprehensive survey maps the landscape of agentic reasoning, extending concepts of self-evolving memory structures and collective multi-agent collaboration.
- Paper: Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents, Shuo Ji et al. (2026). MRAgent extends agent memory retrieval by moving from static similarity lookups to dynamic, reasoning-guided graph reconstruction.
- Paper: Code as Agent Harness, Xuying Ning et al. (2026). Code as Agent Harness investigates using executable code representations as the shared structural harness for agent memory, planning, and multi-agent coordination.
