LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation

Dongge HanCamille CouturierDaniel MadrigalXuchao ZhangVictor RuehleSaravan Rajmohan

article2026Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems36 citations

Introduces LEGOMem, a modular procedural memory framework for multi-agent systems that distributes past execution trajectories across orchestrator and task agents, enabling teams of smaller language models to rival larger models in complex workflow automation.

Listen

Organizations are increasingly deploying teams of specialized artificial intelligence agents powered by large language models to automate complex, multi-step office workflows such as document editing, email handling, and calendar management. In typical architectures, a central orchestrator plans workflows and delegates subtasks to specialized software-tool agents. However, these systems currently operate in a stateless manner, solving every task from scratch and discarding execution experience. This inability to reuse operational knowledge leads to repeated errors and limits workflow reliability.

To solve this challenge, the article introduces and evaluates LEGOMem, a modular procedural memory framework designed specifically for multi-agent systems. The primary objective is to demonstrate how extracting, structuring, and reallocating procedural memory—knowledge of how to execute multi-step workflows—across orchestrators and task agents improves coordination, execution accuracy, and overall task success.

LEGOMem uses a two-phase retrieval-augmented approach. First, an offline curation process analyzes execution logs from successfully completed tasks and distills them into modular units: full-task memories containing high-level plans for the orchestrator, and subtask memories detailing specific tool interactions for individual agents. Second, during active task execution, the system retrieves relevant memories based on semantic similarity. The framework was evaluated across 300 multi-step automation tasks from the OfficeBench benchmark across three team configurations: teams using large language models exclusively (GPT-4o), hybrid teams (GPT-4o orchestrator with GPT-4o-mini task agents), and small-model teams (GPT-4o-mini exclusively).

The evaluation produced several critical findings. First, equipping agent teams with LEGOMem consistently boosted task success rates by over 12 percentage points across all configurations, improving full large-model teams from 45.83% to 58.44% and small-model teams from 24.78% to 38.16%. Second, modular procedural memory enabled smaller, lower-cost models to match or exceed the performance of larger, unaugmented models; for instance, the hybrid team with memory achieved a 50.22% success rate, outperforming the large-model team operating without memory. Third, memory placement experiments revealed that orchestrator memory is essential for effective task decomposition and planning, whereas providing memory solely to localized task agents led to substantial performance drops. Finally, procedural memory improved operational efficiency by reducing total execution steps by up to 16.2% on complex tasks and lowering step-level tool failure rates from 27.5% to 22.5%.

These findings have direct strategic implications for deploying automated enterprise workflows. Incorporating procedural memory allows organizations to deploy smaller, cost-effective language models while achieving performance levels previously requiring larger, computationally expensive models. Furthermore, lowering execution steps and action failures directly reduces operating costs, decreases latency, and mitigates the risk of erroneous automated software actions.

Organizations developing or deploying multi-agent workflow systems should implement role-aware procedural memory architectures that prioritize orchestrator-level planning guidance alongside execution-level tool support. For teams utilizing smaller agents, adopting dynamic retrieval or query-rewriting strategies provides additional performance gains by delivering fine-grained execution guidance.

Decision-makers should note that the article’s evaluations were conducted within a controlled benchmark involving simulated office applications (such as email, spreadsheets, and calendar tools) using successful execution traces. While confidence in the benchmark results is high, organizations should pilot the framework within their specific operational environments and explore expanding procedural memory systems to learn from failed executions in open-ended software ecosystems.

arXiv: 2510.04851
Cover for LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation

Abstract

We introduce LEGOMem, a modular procedural memory framework for multi-agent large language model (LLM) systems in workflow automation. LEGOMem decomposes past task trajectories into reusable memory units and flexibly allocates them across orchestrators and task agents to support planning and execution. To explore the design space of memory in multi-agent systems, we use LEGOMem as a lens and conduct a systematic study of procedural memory in multi-agent systems, examining where memory should be placed, how it should be retrieved, and which agents benefit most. Experiments on the OfficeBench benchmark show that orchestrator memory is critical for effective task decomposition and delegation, while fine-grained agent memory improves execution accuracy. We find that even teams composed of smaller language models can benefit substantially from procedural memory, narrowing the performance gap with stronger agents by leveraging prior execution traces for more accurate planning and tool use. These results position LEGOMem as both a practical framework for memory-augmented agent systems and a research tool for understanding memory design in multi-agent workflow automation.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems
  • 3.1 Problem formulation
  • 3.1.1 Multi-agent system for workflow automation.
  • 3.1.2 Multi-agent procedural memory.
  • 3.2 The LEGOMem framework
  • 3.2.1 Memory construction.
  • 3.2.2 Memory-augmented inference.
  • 3.3 LEGOMem variants
  • 4 Experiments
  • 4.1 Experimental setup
  • 4.1.1 Dataset and metrics
  • 4.1.2 Implementation details
  • 4.1.3 Memory curation and agent inference details
  • 4.2 Main results
  • 4.3 Ablations experiments
  • 4.3.1 Memory retrieval, allocation, and placement
  • 4.3.2 Effectiveness of adding reasoning in memory
  • 4.3.3 Effectiveness of memory on execution steps and failure rates
  • 5 Conclusion
  • References
  • A Prompts for Memory Curation

Knowls

  1. Knowl 1 — Modular Procedural Memory Architecture for Multi-Agent Systems

    model/method

    Multi-agent LLM systems for workflow automation typically consist of a central orchestrator agent AorchA_{\text{orch}} that decomposes high-level user tasks TT (with natural language description dd) into subtasks sts_t and delegates them to specialized tool-using task agents A={A1,…,Ak}A = \{A_1, \dots, A_k\} interacting with an external environment E\mathcal{E}. In stateless architectures, the orchestrator dynamically updates its state σt+1=f(σt,rt)\sigma_{t+1} = f(\sigma_t, r_t) from agent execution summaries rtr_t and tool observations oto_t, but discards operational knowledge across tasks.

    LEGOMem introduces modular procedural memory to capture and reuse multi-agent execution experience. It operates in two phases:

    1. Offline Memory Curation: Successful past execution logs are distilled by an LLM into structured memory units stored in a vector database M\mathcal{M} indexed by dense embeddings ϕ(⋅)\phi(\cdot). Each structured memory unit contains:

      • high_level_plan: The step-by-step plan executed and the assigned agent for each subtask.
      • subtasks: A sequence of localized subtask items, each specifying the responsible agent AjA_j, subtask description dsubtaskd_{\text{subtask}}, a think-action sequence of successful tool calls, and execution observations.
      • final_answer: The terminal answer produced for the user.
      • reflections: A concise synthesis of why the execution succeeded, potential pitfalls, and recovery strategies.
    2. Online Memory Allocation: Given a new task dnewd_{\text{new}}, semantically relevant full-task procedural memories are retrieved using embedding similarity ϕ(dnew)\phi(d_{\text{new}}) to guide orchestrator planning and decomposition, while matching subtask memories are routed to individual task agents to provide grounded, localized tool-use demonstrations.

  2. Knowl 2 — Vanilla LEGOMem Multi-Agent Execution Algorithm

    algorithm

    The Vanilla LEGOMem execution algorithm coordinates memory retrieval, initial planning, subtask delegation, and error recovery across a multi-agent workflow system.

    Input: Task description dnewd_{\text{new}}, procedural memory bank M\mathcal{M}, orchestrator AorchA_{\text{orch}}, task agent set A={A1,…,Ak}A = \{A_1, \dots, A_k\}, embedding model ϕ\phi, retrieval parameter KK.
    Output: Final response from orchestrator AorchA_{\text{orch}}.
    1. Compute task embedding ϕ(dnew)\phi(d_{\text{new}}).
    2. Query M\mathcal{M} to retrieve the top-KK most semantically similar full-task memories m={m1,…,mK}m = \{m_1, \dots, m_K\}.
    3. Extract subtask memory records {m11,…,mnK}\{m_1^1, \dots, m_n^K\} from mm and statically assign corresponding subtask memories to each respective task agent Aj∈AA_j \in A.
    4. Initialize the environment E\mathcal{E}.
    5. Augment AorchA_{\text{orch}} with the full-task memories mm, and prompt AorchA_{\text{orch}} to generate the initial execution plan π0\pi_0.
    6. while task is not completed do
    7. AorchA_{\text{orch}} selects the next task agent At∈AA_t \in A and generates the subtask instruction sts_t.
    8. Augment AtA_t with its allocated subtask memories.
    9. AtA_t generates and executes tool-use actions in E\mathcal{E}, receiving observation oto_t.
    10. AtA_t summarizes subtask execution into summary message rtr_t and transmits it to AorchA_{\text{orch}}.
    11. if progress stalls (e.g., detected repeating states or action loops) then
    12. AorchA_{\text{orch}} performs re-planning using memory context to update plan π′\pi'.
    13. end if
    14. end while
    15. return Final response from AorchA_{\text{orch}}.

    In this procedure, KK full-task memories (e.g., K=5K=5) provide high-level decomposition exemplars to the orchestrator, while extracted subtask units (e.g., 3 per agent) supply execution-level tool parameterization guidance directly to domain-specific agents.

  3. Knowl 3 — Finer-Grained Subtask Memory Retrieval: Dynamic and Query-Rewriting Strategies

    model/method

    Vanilla LEGOMem retrieves full-task memories based solely on the global task description dnewd_{\text{new}} and statically extracts subtask traces for task agents. When a new task shares high-level goals with retrieved cases but differs in specific subtask requirements, static extraction can yield mismatched subtask exemplars. To address this, two fine-grained subtask retrieval variants partition the global memory M\mathcal{M} into per-agent subtask memory stores {MAj∣Aj∈A}\{\mathcal{M}_{A_j} \mid A_j \in A\}:

    1. Dynamic Subtask Retrieval (LEGOMem-Dynamic): Performs just-in-time retrieval during online execution. When the orchestrator generates a subtask sts_t assigned to agent AtA_t, the system computes the embedding ϕ(st)\phi(s_t) and queries the specific agent's subtask store MAt\mathcal{M}_{A_t} at runtime. This retrieves only the most relevant subtask execution traces for AtA_t, minimizing noise from irrelevant subtasks at the cost of repeated vector similarity lookups during execution.

    2. Query-Rewriting Subtask Retrieval (LEGOMem-QueryRewrite): Shifts fine-grained retrieval to the pre-execution phase. After retrieving top-KK full-task memories for the orchestrator, a query rewriter LLM ψ\psi uses the memories and task description to generate a draft subtask plan πdraft′={s1′,s2′,…,sn′}\pi'_{\text{draft}} = \{s'_1, s'_2, \dots, s'_n\}. Each draft subtask sj′s'_j is embedded via ϕ(sj′)\phi(s'_j) to pre-fetch relevant subtask memories from MAj\mathcal{M}_{A_j} for all candidate agents before execution begins, avoiding runtime retrieval latency while retaining fine-grained subtask relevance.

  4. Knowl 4 — Experimental Benchmark Setup for Multi-Agent Workflow Automation

    experimental setup

    LEGOMem and procedural memory baselines are evaluated on the OfficeBench workflow automation benchmark, consisting of 300 multi-step interactive tasks divided into 148 training tasks (used for offline memory curation) and 152 test tasks (for evaluation). Tasks span three application complexity levels:

    • Level 1: Single-application workflows.
    • Level 2: Two-application workflows requiring coordination across two domains.
    • Level 3: Multi-application workflows involving three or more tools.

    Agents interact with Word, Excel, Calendar, Email, System, and OCR-PDF applications encapsulated in Docker containers via structured API tool calls. The vision-language model Phi-3.5-mini is used for image parsing in the OCR application, and text-embedding-3-large paired with FAISS is used for vector memory storage and retrieval.

    Agent team configurations comprise:

    • LLM team: GPT-4o for both orchestrator and all task agents.
    • Hybrid team: GPT-4o for the orchestrator and GPT-4o-mini for task agents.
    • SLM team: GPT-4o-mini for both orchestrator and task agents.

    The evaluation metric is task success rate (percentage of tasks solved correctly), programmatically verified against the final environment state (e.g., exact or fuzzy keyword matches on modified spreadsheet cells, calendar events, sent emails, and generated documents). Memory-augmented systems retrieve 5 full-task memories for the orchestrator and 3 subtask memories per task agent, curated from 93 successful trajectories (comprising 250 extracted subtask memories) obtained by running the LLM team on the training set.

  5. Knowl 5 — Workflow Success Rates Across Multi-Agent Team Configurations and Memory Strategies

    data/table

    Task success rates (%) on the OfficeBench benchmark compare memory-augmented multi-agent teams against memory-less execution and multi-agent adaptations of single-agent procedural memory methods (Synapse, which injects raw action trajectories, and Agent Workflow Memory [AWM], which injects clustered subtask summaries). Results represent means across three random seeds.

    LLM team (GPT-4o) Hybrid team (GPT-4o + Mini) SLM team (GPT-4o-mini)
    Method L1 L2 L3 Overall L1 L2 L3 Overall L1 L2 L3 Overall
    No memory 49.31 58.52 33.33 45.83 45.14 48.89 16.95 35.31 36.81 34.81 7.34 24.78
    Synapse 59.72 75.56 43.50 58.11 46.53 68.15 29.94 46.49 36.81 42.22 20.90 32.24
    AWM 54.17 58.52 35.03 48.03 43.75 55.56 18.64 37.50 35.42 36.30 12.99 26.97
    LEGOMem (Vanilla) 57.99 73.33 47.46 58.44 49.31 62.22 36.16 48.03 38.89 54.07 25.42 38.16
    LEGOMem-Dynamic 56.25 75.56 43.79 57.12 44.44 65.93 36.16 47.59 38.89 50.37 27.12 37.72
    LEGOMem-QueryRewrite 54.17 72.59 42.94 55.26 47.22 66.67 40.11 50.22 36.81 48.89 26.55 36.40

    LEGOMem variants outperform the memory-less baseline by +12.61%, +12.72%, and +13.38% overall on LLM, Hybrid, and SLM teams respectively. Notably, procedural memory enables smaller agent configurations to match or surpass larger memory-less baselines: the Hybrid team with LEGOMem-QueryRewrite (50.22%) exceeds the memory-less LLM team (45.83%), and the SLM team with Vanilla LEGOMem (38.16%) exceeds the memory-less Hybrid team (35.31%).

  6. Knowl 6 — Ablation of Procedural Memory Placement in Multi-Agent Hierarchies

    data/table

    Evaluating different memory placement configurations on OfficeBench demonstrates the comparative necessity of orchestrator-level memory versus task-agent-level memory across task difficulty levels (Level 1, Level 2, Level 3, and Overall success rate in %).

    LLM variants Hybrid (LLM + SLM) variants
    Memory Placement L1 L2 L3 Overall L1 L2 L3 Overall
    Orchestrator + Agent memory
    LEGOMem 57.99 73.33 47.46 58.44 49.31 62.22 36.16 48.03
    LEGOMem-Dynamic 56.25 75.56 43.79 57.12 44.44 65.93 36.16 47.59
    LEGOMem-QueryRewrite 54.17 72.59 42.94 55.26 47.22 66.67 40.11 50.22
    Orchestrator (planning) + Agent
    LEGOMem 54.86 76.30 35.03 53.51 45.14 63.70 30.51 44.96
    LEGOMem-Dynamic 54.86 73.33 41.81 55.26 46.53 64.44 32.77 46.49
    LEGOMem-QueryRewrite 51.39 70.37 42.94 53.73 49.31 59.26 35.59 46.93
    Orchestrator memory only
    LEGOMem 51.39 74.07 38.98 53.29 45.83 68.89 32.77 47.59
    Task Agent memory only
    LEGOMem 50.00 63.70 38.98 49.78 44.44 46.67 19.21 35.31
    LEGOMem-Dynamic 49.31 62.96 38.98 49.34 47.22 55.56 23.16 40.35
    LEGOMem-QueryRewrite 54.86 66.67 35.03 50.66 44.44 54.81 24.29 39.69
    No memory 49.31 58.52 33.33 45.83 45.14 48.89 16.95 35.31

    Key empirical observations:

    1. Dominance of Orchestrator Memory: Removing memory from the orchestrator (Task Agent memory only) produces a steep drop in overall success rate (from 58.44% to 49.78% in LLM teams, and from 48.03% to 35.31% in Hybrid teams under vanilla LEGOMem), showing that local execution memory without global orchestration memory is insufficient.
    2. Value of Fine-Grained Retrieval for Smaller Agents: In the Task Agent memory only setting for Hybrid teams, LEGOMem-Dynamic (40.35%) and LEGOMem-QueryRewrite (39.69%) outperform vanilla LEGOMem (35.31%) by 4–5%, demonstrating that dynamic and rewritten retrieval provide more relevant context to smaller models that lack strong orchestrator guidance.
  7. Knowl 7 — Effect of Reasoning Traces in Procedural Memory Units

    empirical result

    An ablation study on whether augmenting structured procedural memories with explicit <think> reasoning traces improves multi-agent workflow performance reveals that the benefit of reasoning steps inside memory exemplars is marginal:

    LLM team (GPT-4o) Hybrid team (GPT-4o + Mini)
    Memory Variant L1 L2 L3 Overall L1 L2 L3 Overall
    No reasoning
    LEGOMem 54.17 75.93 43.22 56.36 48.61 68.15 36.72 49.78
    LEGOMem-Dynamic 56.94 73.33 44.63 57.02 50.00 68.15 36.72 50.22
    LEGOMem-QueryRewrite 48.61 69.63 48.02 54.61 37.50 65.19 36.16 45.18
    With reasoning
    LEGOMem 57.99 73.33 47.46 58.44 49.31 62.22 36.16 48.03
    LEGOMem-Dynamic 56.25 75.56 43.79 57.12 44.44 65.93 36.16 47.59
    LEGOMem-QueryRewrite 54.17 72.59 42.94 55.26 47.22 66.67 40.11 50.22

    Overall task success rates change by less than 2 percentage points between reasoning-augmented and reasoning-free memory units across LLM and Hybrid configurations (e.g., Vanilla LEGOMem shifts from 56.36% to 58.44% on LLM teams, and from 49.78% to 48.03% on Hybrid teams). This indicates that the modular decomposition of trajectories into structured subtask and tool-call actions provides the essential procedural guidance, making detailed reasoning tokens largely redundant.

  8. Knowl 8 — Reduction in Execution Steps and Tool-Use Failure Rates via Procedural Memory

    empirical result

    Equipping multi-agent systems with procedural memory improves execution efficiency and operational reliability across task difficulty levels (evaluated on LLM teams using GPT-4o):

    • Execution Step Reduction: On complex multi-application Level 3 tasks, full LEGOMem reduces the average number of execution steps from 26.5 (no memory) to 22.2 steps, representing a 16.2% reduction in required interaction steps. In comparison, orchestrator-only memory requires 21.4 steps, while task-agent-only memory requires 23.6 steps (and 16.6 steps vs 12.8 steps for orchestrator-only on Level 2 tasks), demonstrating that orchestrator planning memory is the primary driver of execution conciseness.
    • Per-Step Failure Rate Reduction: Procedural memory reduces the rate of failed environment actions (caused by malformed tool calls, invalid parameters, or erroneous API operations). For Level 3 tasks, the per-step failure rate decreases from 0.275 in the memory-less system to 0.225 with full LEGOMem (orchestrator-only memory reaches 0.261 and task-agent-only reaches 0.262). Across Level 1 and Level 2 tasks, full LEGOMem similarly suppresses step failure rates compared to memory-less baselines (Level 2: 0.161 vs 0.220).
  9. Knowl 9 — Limitations and Scope of the LEGOMem Framework

    limitation

    The LEGOMem framework and its evaluation have the following stated limitations:

    1. Exclusive Reliance on Successful Trajectories: The memory curation pipeline distills memory units solely from successful past execution logs. It does not incorporate continual learning from failed trajectories, missing opportunities to learn explicit error-avoidance behaviors or negative constraints.
    2. Controlled Benchmark Environment: Evaluation is restricted to the simulated OfficeBench productivity suite environment (comprising six defined office and OS applications), rather than open-ended, real-world operating systems or dynamic web ecosystems.
    3. Static Agent and Tool Ecosystem: The task agent roster and tool API schemas are fixed during execution, without supporting open-ended dynamic tool creation or fluid multi-agent role discovery.

Coverage note — None was omitted; all key architectural components, retrieval algorithms, experimental setups, empirical data tables, ablations, and stated limitations are fully covered.

References

  1. 1.Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. 2024. Phi-4 technical report. arXiv preprint arXiv:2412.08905 (2024).
  2. 2.Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. 2022. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691 (2022).
  3. 3.Ruisheng Cao, Fangyu Lei, Haoyuan Wu, Jixuan Chen, Yeqiao Fu, Hongcheng Gao, Xinzhuang Xiong, Hanchong Zhang, Wenjing Hu, Yuchen Mao, et al. 2024. Spider2-v: How far are multimodal agents from automating data science and engineering workflows? Advances in Neural Information Processing Systems 37 (2024), 107703–107744.
  4. 4.Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chen Qian, Chi-Min Chan, Yujia Qin, Yaxi Lu, Ruobing Xie, et al. 2023. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors in agents. arXiv preprint arXiv:2308.10848 2, 4 (2023), 6.
  5. 5.Yuheng Cheng, Ceyao Zhang, Zhengwen Zhang, Xiangrui Meng, Sirui Hong, Wenhao Li, Zihao Wang, Zekai Wang, Feng Yin, Junhua Zhao, et al. 2024. Exploring large language model based intelligent agents: Definitions, methods, and prospects. arXiv preprint arXiv:2401.03428 (2024).
  6. 6.Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. 2025. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. https://doi.org/10.48550/arXiv.2504.19413 arXiv:2504.19413 [cs].
  7. 7.Yufan Dang, Chen Qian, Xueheng Luo, Jingru Fan, Zihao Xie, Ruijie Shi, Weize Chen, Cheng Yang, Xiaoyin Che, Ye Tian, et al. 2025. Multi-Agent Collaboration via Evolving Orchestration. arXiv preprint arXiv:2505.19591 (2025).
  8. 8.Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024. The faiss library. arXiv preprint arXiv:2401.08281 (2024).
  9. 9.Adam Fourney, Gagan Bansal, Hussein Mozannar, Cheng Tan, Eduardo Salinas, Friederike Niedtner, Grace Proebsting, Griffin Bassman, Jack Gerrits, Jacob Alber, et al. 2024. Magentic-one: A generalist multi-agent system for solving complex tasks. arXiv preprint arXiv:2411.04468 (2024).
  10. 10.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 2, 1 (2023).
  11. 11.Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024).
  12. 12.Minki Kang, Wei-Ning Chen, Dongge Han, Huseyin A Inan, Lukas Wutschitz, Yanzhi Chen, Robert Sim, and Saravan Rajmohan. 2025. ACON: Optimizing Context Compression for Long-horizon LLM Agents. arXiv preprint arXiv:2510.00615 (2025).
  13. 13.Sehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee, Michael W Mahoney, Kurt Keutzer, and Amir Gholami. 2024. An llm compiler for parallel function calling. In Forty-first International Conference on Machine Learning.
  14. 14.Dongjun Lee, Juyong Lee, Kyuyoung Kim, Jihoon Tack, Jinwoo Shin, Yee Whye Teh, and Kimin Lee. 2025. Learning to contextualize web pages for enhanced decision making by LLM agents. arXiv preprint arXiv:2503.10689 (2025).
  15. 15.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems 33 (2020), 9459–9474.
  16. 16.Zhicong Li, Jiahao Wang, Hangyu Mao, ZhiShu Jiang, Zhongxia Chen, Du Jiazhen, Fuzheng Zhang, Di ZHANG, and Yong Liu. [n.d.]. DMQR-RAG: Diverse MultiQuery Rewriting in Retrieval-Augmented Generation. ([n. d.]).
  17. 17.Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query rewriting in retrieval-augmented large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 5303–5315.
  18. 18.Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. 2024. Evaluating very long-term conversational memory of llm agents. arXiv preprint arXiv:2402.17753 (2024).
  19. 19.Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023. Gaia: a benchmark for general ai assistants. In The Twelfth International Conference on Learning Representations.
  20. 20.Krishan Rana, Jesse Haviland, Sourav Garg, Jad Abou-Chakra, Ian Reid, and Niko Suenderhauf. 2023. Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning. arXiv preprint arXiv:2307.06135 (2023).
  21. 21.Preston Rasmussen, Pavlo Paliychuk, Travis Beauvais, Jack Ryan, and Daniel Chalef. 2025. Zep: a temporal knowledge graph architecture for agent memory. arXiv preprint arXiv:2501.13956 (2025).
  22. 22.Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M Sadler, Wei-Lun Chao, and Yu Su. 2023. Llm-planner: Few-shot grounded planning for embodied agents with large language models. In Proceedings of the IEEE/CVF international conference on computer vision. 2998–3009.
  23. 23.Peter Stone and Manuela Veloso. 2000. Multiagent systems: A survey from a machine learning perspective. Autonomous Robots 8, 3 (2000), 345–383.
  24. 24.Haoran Sun and Shaoning Zeng. 2025. Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents. arXiv preprint arXiv:2507.22925 (2025).
  25. 25.Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science 18, 6 (2024), 186345.
  26. 26.Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023. Plan-and-solve prompting: Improving zero-shot chainof-thought reasoning by large language models. arXiv preprint arXiv:2305.04091 (2023).
  27. 27.Weixuan Wang, Dongge Han, Daniel Madrigal Diaz, Jin Xu, Victor Rühle, and Saravan Rajmohan. 2025. OdysseyBench: Evaluating LLM Agents on LongHorizon Complex Office Application Workflows. arXiv preprint arXiv:2508.09124 (2025).
  28. 28.Zilong Wang, Yuedong Cui, Li Zhong, Zimin Zhang, Da Yin, Bill Yuchen Lin, and Jingbo Shang. 2024. Officebench: Benchmarking language agents across multiple applications for office automation. arXiv preprint arXiv:2407.19056 (2024).
  29. 29.Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. 2024. Agent Workflow Memory. https://doi.org/10.48550/arXiv.2409.07429 arXiv:2409.07429 [cs].
  30. 30.Michael Wooldridge. 2009. An Introduction to MultiAgent Systems (2nd ed.). Wiley Publishing.
  31. 31.Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. 2024. Longmemeval: Benchmarking chat assistants on long-term interactive memory. arXiv preprint arXiv:2410.10813 (2024).
  32. 32.Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, et al. 2024. Autogen: Enabling next-gen LLM applications via multi-agent conversations. In First Conference on Language Modeling.
  33. 33.Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. 2024. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems 37 (2024), 52040–52094.
  34. 34.Wujiang Xu, Kai Mei, Hang Gao, Juntao Tan, Zujie Liang, and Yongfeng Zhang. 2025. A-MEM: Agentic Memory for LLM Agents. https://doi.org/10.48550/arXiv.2502.12110 arXiv:2502.12110 [cs].
  35. 35.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR).
  36. 36.Chaoyun Zhang, Liqun Li, Shilin He, Xu Zhang, Bo Qiao, Si Qin, Minghua Ma, Yu Kang, Qingwei Lin, Saravan Rajmohan, et al. 2024. Ufo: A ui-focused agent for windows os interaction. arXiv preprint arXiv:2402.07939 (2024).
  37. 37.Longtao Zheng, Rundong Wang, Xinrun Wang, and Bo An. 2023. Synapse: Trajectory-as-exemplar prompting with memory for computer control. arXiv preprint arXiv:2306.07863 (2023).
  38. 38.Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. 2024. Memorybank: Enhancing large language models with long-term memory. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 19724–19731.
  39. 39.Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. 2023. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854 (2023).
  40. 40.Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, and Paul Pu Liang. 2025. MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents. arXiv preprint arXiv:2506.15841 (2025).

Citation

MLA
Han, D., et al. “LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation”. arXiv, 2025, http://arxiv.org/abs/2510.04851v1.
APA
Han, D., Couturier, C., Diaz, D. M., Zhang, X., Rühle, V., & Rajmohan, S. (2025). LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation. arXiv. http://arxiv.org/abs/2510.04851v1
Chicago
Han, D., C. Couturier, D. M. Diaz, X. Zhang, V. Rühle, and S. Rajmohan. 2025. “LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation”. arXiv. http://arxiv.org/abs/2510.04851v1.
Harvard
Han, D. et al. (2025) “LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2510.04851v1.
Vancouver
1. Han D, Couturier C, Diaz DM, Zhang X, Rühle V, Rajmohan S (2025) LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation. arXiv

BibTeX

@article{han2025legomem,
  title = {LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation},
  author = {Han, Dongge and Couturier, Camille and Diaz, Daniel Madrigal and Zhang, Xuchao and Rühle, Victor and Rajmohan, Saravan},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2510.04851v1},
  eprint = {2510.04851}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/