SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills
Zhongxin GuoDanrui QiHanwen GuPeng ChengYongqiang Xiong
Introduces SkillDisCo, a framework that distills successful agent execution paths into reusable control-flow subgraphs and compiles them into executable procedural skills to reduce redundant reasoning and improve task success rates across complex interactive benchmarks.
Autonomous artificial intelligence agents often solve interactive tasks independently from scratch. This practice causes agents to repeatedly discover identical low-level action sequences, resulting in high computational costs, lengthy execution traces, and brittle performance. While prior methods extract textual workflows or raw scripts from previous executions, these approaches lack explicit structural representations and frequently produce fragmented, redundant, and error-prone skill libraries. The article addresses this operational challenge by introducing SKILL-DISCO, a framework designed to discover reusable procedural skills across successful execution records and compile them into verifiable, executable software routines.
To establish these skills, the approach models deterministic execution environments as state machines and represents skills as parameterized control-flow subgraphs. The framework operates in two distinct phases: distillation and compilation. During distillation, the system normalizes raw execution logs into intermediate code representations, extracts operations aligned with intermediate subgoals, and clusters them across multiple execution traces to identify recurring execution patterns. During compilation, these high-coverage clusters are converted into typed interface specifications, synthesized into standalone Python programs, and verified against held-out test tasks. The evaluation tested this framework on the ALFWorld household task benchmark and the WebArena realistic web-navigation environment across diverse language model architectures and scales.
The experimental findings show substantial improvements in both operational performance and reliability. First, SKILL-DISCO achieved the highest overall task success rates across all tested configurations, boosting success from 82.0% to 92.4% on ALFWorld ReAct benchmarks and from 23.9% to 29.1% on WebArena ReAct evaluations, while outperforming existing skill-induction baselines. Second, the system markedly improved execution efficiency, reducing agent interaction turns by 11.3% to 54.5% on ALFWorld and by 13.1% to 22.0% on WebArena as low-level actions were consolidated into reusable skills. Third, the resulting skill libraries remained exceptionally compact—generating only 5 skills for ALFWorld and 20 for WebArena compared to 110 and 146 skills produced by prior methods—while eliminating skill execution failure rates entirely on ALFWorld (from 75.3% down to 0.0%) and reducing them from 33.9% to 21.5% on WebArena. Finally, skills compiled using larger frontier models successfully transferred to smaller open-source models, enabling smaller architectures such as Qwen3.5-9B to achieve a 98.5% success rate on ALFWorld, exceeding the standalone performance of larger induction models.
These findings demonstrate that distilling and compiling execution traces into validated, executable routines significantly reduces reasoning costs, lowers API expenses, and improves system reliability. Organizations deploying autonomous interactive agents can utilize stronger frontier models offline to induce structured skill libraries, which can then be executed in production by substantially smaller, cheaper models without sacrificing performance. This approach provides a practical pathway to mitigate deployment costs while enhancing operational safety and execution predictability.
Organizations developing agent workflows should consider deploying distillation and compilation frameworks to eliminate redundant reasoning in repetitive procedural domains. Prior to full-scale adoption, engineering teams should conduct targeted pilot evaluations on specific operational workflows to verify that task environments adhere to deterministic state transitions. Additionally, practitioners should establish automated test suites to maintain and re-verify compiled skill libraries whenever target environments or tool interfaces change.
The findings are subject to specific operational boundaries. The framework relies strictly on the availability of successful execution traces, meaning it cannot extract skills in sparse environments where initial agent success is unattainable. Furthermore, the approach applies specifically to structured procedural workflows such as web automation and tool interaction, offering no direct benefits for open-ended generation or unstructured linguistic tasks. Within these defined operational parameters, there is high confidence that the framework delivers consistent gains in task reliability and execution efficiency.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). Agent Workflow Memory establishes the baseline paradigm of extracting abstract, reusable procedural workflows from successful digital agent trajectories to guide subsequent tasks across WebArena.
- Paper: Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning, Rodrigo Toro Icarte et al. (2022). Reward Machines provides the theoretical foundation for modeling procedural tasks as structured finite-state automata with discrete transitions and sub-policy decompositions.
- Paper: Automaton-Guided Curriculum Generation for Reinforcement Learning Agents, Yash Shukla et al. (2023). AGCL introduces techniques for formulating sequential sub-goals and agent execution structures via deterministic finite automata.
- Paper: Learning by Distilling Context, Charlie Snell et al. (2022). Learning by Distilling Context introduces the core methodology of distilling multi-step prompt context and execution histories into compact, efficient representations.
- Paper: SWE-Exp: Experience-Driven Software Issue Resolution, Silin Chen et al. (2025). SWE-Exp demonstrates how empirical problem-solving experiences can be distilled into reusable memory units to avoid repeated exploration costs.
- Paper: From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills, Zisu Huang et al. (2026). This systematic study evaluates the end-to-end lifecycle, transferability, and failure modes of model-generated agent skills across domains including ALFWorld.
- Paper: Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills, Agamdeep Singh et al.. This work explores amortizing the computational reasoning premium by distilling multi-step agent trajectories into reusable natural language skill files on benchmarks like ALFWorld.
- Paper: LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation, Dongge Han et al. (2026). LEGOMem extends the concept of distilling procedural workflows into modular, reusable subtask memories to multi-agent collaborative systems.
- Paper: Code as Agent Harness, Xuying Ning et al. (2026). This survey provides an overarching framework for using executable code, harnesses, and programmatic policies as grounded substrates for reusable agent skills.
- Paper: Sample-Efficient Learning from Agent Experience, Chenhui Gou et al. (2026). This work explores an alternative path for reusing agent experience by consolidating trajectory-induced behaviors directly into model weights via experience distillation.
- Paper: Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking, Jiwan Chung et al. (2026). WEBSTEP applies state-machine tracking to evaluate the fine-grained, step-level procedural execution of interactive web agents.
- Paper: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models, Qizheng Zhang et al. (2026). Agentic Context Engineering investigates evolving structured contextual playbooks and strategies dynamically from execution traces over time.
