SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills

Zhongxin GuoDanrui QiHanwen GuPeng ChengYongqiang Xiong

article2026arXiv5 citations

Introduces SkillDisCo, a framework that distills successful agent execution paths into reusable control-flow subgraphs and compiles them into executable procedural skills to reduce redundant reasoning and improve task success rates across complex interactive benchmarks.

Listen

Autonomous artificial intelligence agents often solve interactive tasks independently from scratch. This practice causes agents to repeatedly discover identical low-level action sequences, resulting in high computational costs, lengthy execution traces, and brittle performance. While prior methods extract textual workflows or raw scripts from previous executions, these approaches lack explicit structural representations and frequently produce fragmented, redundant, and error-prone skill libraries. The article addresses this operational challenge by introducing SKILL-DISCO, a framework designed to discover reusable procedural skills across successful execution records and compile them into verifiable, executable software routines.

To establish these skills, the approach models deterministic execution environments as state machines and represents skills as parameterized control-flow subgraphs. The framework operates in two distinct phases: distillation and compilation. During distillation, the system normalizes raw execution logs into intermediate code representations, extracts operations aligned with intermediate subgoals, and clusters them across multiple execution traces to identify recurring execution patterns. During compilation, these high-coverage clusters are converted into typed interface specifications, synthesized into standalone Python programs, and verified against held-out test tasks. The evaluation tested this framework on the ALFWorld household task benchmark and the WebArena realistic web-navigation environment across diverse language model architectures and scales.

The experimental findings show substantial improvements in both operational performance and reliability. First, SKILL-DISCO achieved the highest overall task success rates across all tested configurations, boosting success from 82.0% to 92.4% on ALFWorld ReAct benchmarks and from 23.9% to 29.1% on WebArena ReAct evaluations, while outperforming existing skill-induction baselines. Second, the system markedly improved execution efficiency, reducing agent interaction turns by 11.3% to 54.5% on ALFWorld and by 13.1% to 22.0% on WebArena as low-level actions were consolidated into reusable skills. Third, the resulting skill libraries remained exceptionally compact—generating only 5 skills for ALFWorld and 20 for WebArena compared to 110 and 146 skills produced by prior methods—while eliminating skill execution failure rates entirely on ALFWorld (from 75.3% down to 0.0%) and reducing them from 33.9% to 21.5% on WebArena. Finally, skills compiled using larger frontier models successfully transferred to smaller open-source models, enabling smaller architectures such as Qwen3.5-9B to achieve a 98.5% success rate on ALFWorld, exceeding the standalone performance of larger induction models.

These findings demonstrate that distilling and compiling execution traces into validated, executable routines significantly reduces reasoning costs, lowers API expenses, and improves system reliability. Organizations deploying autonomous interactive agents can utilize stronger frontier models offline to induce structured skill libraries, which can then be executed in production by substantially smaller, cheaper models without sacrificing performance. This approach provides a practical pathway to mitigate deployment costs while enhancing operational safety and execution predictability.

Organizations developing agent workflows should consider deploying distillation and compilation frameworks to eliminate redundant reasoning in repetitive procedural domains. Prior to full-scale adoption, engineering teams should conduct targeted pilot evaluations on specific operational workflows to verify that task environments adhere to deterministic state transitions. Additionally, practitioners should establish automated test suites to maintain and re-verify compiled skill libraries whenever target environments or tool interfaces change.

The findings are subject to specific operational boundaries. The framework relies strictly on the availability of successful execution traces, meaning it cannot extract skills in sparse environments where initial agent success is unattainable. Furthermore, the approach applies specifically to structured procedural workflows such as web automation and tool interaction, offering no direct benefits for open-ended generation or unstructured linguistic tasks. Within these defined operational parameters, there is high confidence that the framework delivers consistent gains in task reliability and execution efficiency.

arXiv: 2606.26669
Cover for SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills

Abstract

Agents often repeatedly solve similar task instances from scratch, leading to unnecessary reasoning cost and long execution traces. Prior work has explored workflow reuse and executable skill induction, but it remains unclear which task scenarios admit procedural skills and how the shared procedural structure should be represented across successful traces. We study this problem in FSM-defined scenarios, where successful traces can be viewed as paths in an unknown transition graph, and formulate procedural skills as reusable parameterized control-flow subgraphs. Based on this view, we introduce SkillDisCo, a distillation-and-compilation framework that distills reusable PFSM subgraphs from successful traces and compiles them into callable, executable, and verifiable procedural skills. Experiments on ALFWorld and WebArena show that SkillDisCo improves success rates and reduces agent turns across benchmarks and model scales, demonstrating the benefits of representing shared experience as reusable execution structures.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 FSM-Defined Scenarios
  • 2.2 PFSM-Based Procedural Skill Discovery
  • 3 The Skill-DisCo Framework
  • 3.1 Framework Overview
  • 3.2 The Distillation Phase
  • 3.3 The Compilation Phase
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 End-to-End Performance
  • 4.3 Cross-Model Skill Transferability
  • 4.4 Skill Usage and Execution Reliability
  • 4.5 Ablation Study
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Full Results: Token Usage and Inference Cost
  • A.1 API Pricing Assumptions
  • A.2 Full Baseline Comparison with Token and Cost Statistics
  • A.3 Per-Episode Token and Cost Breakdown
  • A.4 Full Ablation Breakdown
  • B Pipeline Stage Prompts
  • B.1 Stage 1 — Trace →\to Program Transpilation
  • B.2 Stage 2 — Semantic Operation Extraction
  • B.3 Stage 3 — Operation Clustering (Two Passes)
  • B.4 Stage 4 — Skill Contract Definition
  • B.5 Stage 5 — Skill Synthesis
  • C Skill Library Examples

Knowls

  1. Knowl 1 — Parameterized Finite-State Machine Formulation of Procedural Skills

    definition

    Procedural skill discovery for interactive agents operating in deterministic environments is formalized using Parameterized Finite-State Machines (PFSMs):

    • FSM-Defined Scenario: An environment whose execution dynamics are defined by a finite-state machine M=(S,A,δ,S0,Sgoal)M = (S, A, \delta, S_0, S_{\text{goal}}), where SS is a finite state space, AA is a finite set of primitive actions (primitive operators op=(X,Y,Pre,Post)op = (\mathcal{X}, \mathcal{Y}, \text{Pre}, \text{Post})), δ:S×A→S\delta : S \times A \to S is a deterministic transition function, S0⊆SS_0 \subseteq S is the initial state set, and Sgoal⊆SS_{\text{goal}} \subseteq S is the set of goal states. The underlying transition graph is G∗=(V∗,E∗)G^* = (V^*, E^*) where V∗=SV^* = S and E∗={(s,a,s′)∣s,s′∈S,a∈A,δ(s,a)=s′}E^* = \{(s, a, s') \mid s, s' \in S, a \in A, \delta(s, a) = s'\}.

    • Successful Agent Trace: An execution sequence τ=(o0,a0,o1,a1,…,aT−1,oT)\tau = (o_0, a_0, o_1, a_1, \dots, a_{T-1}, o_T) where oto_t is the observation generated from environment state sts_t, at∈Aa_t \in A is the primitive action, and st+1=δ(st,at)s_{t+1} = \delta(s_t, a_t), with s0∈S0s_0 \in S_0 and sT∈Sgoals_T \in S_{\text{goal}}.

    • Parameterized Finite-State Machine (PFSM): Defined as M~=(S~,A~,Θ,δ~,S~0,S~goal)\tilde{M} = (\tilde{S}, \tilde{A}, \Theta, \tilde{\delta}, \tilde{S}_0, \tilde{S}_{\text{goal}}), where S~\tilde{S} is a finite set of abstract parameterized states, A~\tilde{A} is a finite set of parameterized action schemas, Θ\Theta is the parameter space, and δ~:S~×A~×Θ→S~\tilde{\delta} : \tilde{S} \times \tilde{A} \times \Theta \to \tilde{S} is a deterministic parameterized transition function. An assignment θ∈Θ\theta \in \Theta binds parameters to concrete entities (e.g., locations, objects), instantiating a parameterized transition into a concrete FSM transition.

    • Procedural Skill Discovery: Given a set of successful traces T+={τ1,τ2,…,τN}\mathcal{T}^+ = \{\tau_1, \tau_2, \dots, \tau_N\} and a lifting function ϕ:τi↦G~i\phi : \tau_i \mapsto \tilde{G}_i mapping each concrete trace to a parameterized trace graph G~i\tilde{G}_i, the objective is to discover a skill library K={K1,K2,…,Km}\mathcal{K} = \{K_1, K_2, \dots, K_m\} where each procedural skill KjK_j is a reusable parameterized control-flow subgraph that matches a subset of {G~i}i=1N\{\tilde{G}_i\}_{i=1}^N under parameter binding (Kj⪯G~iK_j \preceq \tilde{G}_i).

  2. Knowl 2 — The SKILL-DISCO Skill Distillation and Compilation Pipeline

    algorithm

    SKILL-DISCO discovers reusable procedural skills from successful agent execution traces via a two-phase pipeline comprising distillation (approximating PFSM subgraph extraction and clustering) and compilation (producing verified, executable Python skills).

    Input: Set of successful traces T+={τ1,τ2,…,τN}\mathcal{T}^+ = \{\tau_1, \tau_2, \dots, \tau_N\}, retry limit RR
    Output: Executable and verified skill library K\mathcal{K}
    // Phase 1: Distillation
    for each trace τi∈T+\tau_i \in \mathcal{T}^+ do
        // Stage 1: Trace Normalization
        pi←NormalizeTrace(τi)p_i \leftarrow \text{NormalizeTrace}(\tau_i) // Transpile trace to a Python-like intermediate program with explicit control flow
        
        // Stage 2: Subgoal-level Operation Extraction
        Oi←ExtractOperations(pi)\mathcal{O}_i \leftarrow \text{ExtractOperations}(p_i) // Extract operations o=(ν,σ,u,c)o = (\nu, \sigma, u, c) where ∣u∣≥2|u| \ge 2
    end for
    Omulti←⋃i{o∈Oi:∣o∣≥2}\mathcal{O}_{\text{multi}} \leftarrow \bigcup_{i} \{o \in \mathcal{O}_i : |o| \ge 2\}
    // Stage 3: Procedural Skill Consolidation
    C←TwoPassClustering(Omulti)\mathcal{C} \leftarrow \text{TwoPassClustering}(\mathcal{O}_{\text{multi}}) // Cluster by shared control flow and compute reusability rkr_k
    Chigh←{ck∈C:rk≥threshold}\mathcal{C}_{\text{high}} \leftarrow \{c_k \in \mathcal{C} : r_k \ge \text{threshold}\}
    // Phase 2: Compilation
    K←∅\mathcal{K} \leftarrow \emptyset
    for each cluster ck∈Chighc_k \in \mathcal{C}_{\text{high}} do
        // Stage 4: Skill Contract Definition
        speck←DefineContract(ck)\text{spec}_k \leftarrow \text{DefineContract}(c_k) // Generate typed signature, docstring, pre/postconditions, and side effects
        
        // Stage 5: Skill Synthesis & Verification
        for retry =1= 1 to RR do
            fk←SynthesizePythonCode(speck,ck)f_k \leftarrow \text{SynthesizePythonCode}(\text{spec}_k, c_k)
            passed←VerifyOnHeldOutTasks(fk,speck)\text{passed} \leftarrow \text{VerifyOnHeldOutTasks}(f_k, \text{spec}_k)
            if passed\text{passed} then
                K←K∪{(speck,fk)}\mathcal{K} \leftarrow \mathcal{K} \cup \{(\text{spec}_k, f_k)\}
                break
            end if
        end for
    end for
    return K\mathcal{K}
  3. Knowl 3 — Reusability Score for PFSM Subgraph Consolidation

    equation

    In the distillation phase of SKILL-DISCO, subgoal-level operations extracted across multiple trajectories are grouped by parameterized control-flow equivalence. For each candidate skill cluster ckc_k approximating a parameterized finite-state machine (PFSM) subgraph KkK_k, its reusability score rkr_k over a set of NN successful traces T+={τ1,…,τN}\mathcal{T}^+ = \{\tau_1, \dots, \tau_N\} is estimated as:

    rk≈1N∑i=1NI[Kk⪯G~i]r_k \approx \frac{1}{N} \sum_{i=1}^{N} \mathbb{I}[K_k \preceq \tilde{G}_i]

    where G~i=ϕ(τi)\tilde{G}_i = \phi(\tau_i) is the parameterized trace graph lifted from trace τi\tau_i, I[⋅]\mathbb{I}[\cdot] is the indicator function, and Kk⪯G~iK_k \preceq \tilde{G}_i denotes that the PFSM subgraph approximated by cluster ckc_k matches a subgraph of G~i\tilde{G}_i under some parameter binding θ∈Θ\theta \in \Theta. Clusters with high reusability scores are prioritized for compilation into callable procedural skills.

  4. Knowl 4 — SKILL-DISCO Skill Contract Specification and Execution Protocol

    model/method

    In SKILL-DISCO, distilled PFSM clusters are converted into executable, typed Python functions adhering to an explicit interface contract.

    Skill Specification Structure:

    1. Signature: Action-oriented function name, typed domain parameters with default values, and a structured return type Dict.
    2. Description: Machine-readable docstring and LLM-facing guidance on when and how to invoke the skill.
    3. Behavioral Requirements: Preconditions, postconditions, and declared side effects, grounded in a canonical action sequence template derived from successful traces.
    4. Metadata: Abstraction level (primitive, composite, or workflow), estimated primitive actions saved per call, and confidence score inherited from cluster reusability rkr_k.

    Runtime Execution Protocol: A compiled skill function executes standalone in the workload environment via env.step(action) using deterministic control flow (loops, conditionals) that branches directly on the latest observation. Upon completion, it returns a dictionary with the following schema:

    • success: Boolean indicating whether the sub-goal was achieved.
    • observation: The latest environment observation string.
    • available_actions: The list of valid actions at the final step.
    • process_trace: A chronological list of (action, observation) tuples executed during the invocation, preventing caller blind spots.
  5. Knowl 5 — End-to-End Performance on ALFWorld and WebArena

    data/table

    When evaluated on ALFWorld and WebArena, integrating SKILL-DISCO into interactive base agents (ReAct and CodeAct) using a GPT-4o backbone outperforms both vanilla base agents and prior offline skill-induction methods (AWM_offline and ASI_offline):

    Method SR (%) ↑\uparrow Avg. Turns ↓\downarrow In-Tok. (Cached) Out-Tok. Cost ($)
    (a) ALFWorld (134 unseen tasks)
    ReAct 82.00 19.29 35,600 (22,200) 626 0.0450
    CodeAct 96.27 3.63 7,721 (6,232) 562 0.0109
    AWMoffline\text{AWM}_{\text{offline}} 54.48 11.34 38,908 (22,923) 275 0.0484
    ASIoffline\text{ASI}_{\text{offline}} 47.01 11.43 68,945 (54,175) 286 0.0533
    SKILL-DISCO + ReAct 92.40 (+12.7%) 8.78 (−-54.5%) 22,100 (11,300) 607 0.0360
    SKILL-DISCO + CodeAct 99.25 (+3.1%) 3.22 (−-11.3%) 8,468 (6,303) 368 0.0107
    (b) WebArena (406 evaluation tasks)
    ReAct 23.89 5.88 50,914 (8,605) 728 0.1152
    CodeAct 19.95 10.86 75,962 (24,202) 806 0.1435
    AWMoffline\text{AWM}_{\text{offline}} 21.18 5.92 49,025 (11,748) 564 0.1018
    ASIoffline\text{ASI}_{\text{offline}} 24.63 5.71 56,662 (9,247) 603 0.1269
    SKILL-DISCO + ReAct 29.06 (+21.6%) 5.11 (−-13.1%) 51,068 (8,582) 561 0.1215
    SKILL-DISCO + CodeAct 22.91 (+14.8%) 8.47 (−-22.0%) 85,649 (27,078) 736 0.1606

    On ALFWorld, SKILL-DISCO + CodeAct improves success rate from 96.27% to 99.25% while decreasing average turns from 3.63 to 3.22 (−11.3%-11.3\%) and output tokens by 34.5%34.5\%. On WebArena, SKILL-DISCO + ReAct reaches 29.06% success rate (+21.6%+21.6\% relative over ReAct and outperforming ASI_offline's 24.63%) with a 13.1%13.1\% reduction in turns.

  6. Knowl 6 — Cross-Model and Cross-Scale Transferability of Compiled Skills

    data/table

    A skill library compiled once using GPT-4o transfers zero-shot across different model families and scales without re-induction. On ALFWorld (CodeAct backbone, 134 unseen tasks) and WebArena (ReAct backbone, 406 evaluation tasks), all evaluated models benefit from the shared skill library:

    Success Rate (%) ↑\uparrow Avg. Agent Turns ↓\downarrow
    Target Model Vanilla +Skill (Δ\Delta) Vanilla +Skill (Δ\Delta)
    (a) ALFWorld (CodeAct backbone)
    Qwen3.5-4B 61.2 91.0 (+48.8%) 9.0 5.0 (−-44.4%)
    Qwen3.5-9B 54.5 98.5 (+80.8%) 10.8 3.4 (−-68.2%)
    GPT-4o 96.3 99.3 (+3.1%) 3.6 3.2 (−-11.3%)
    GPT-4o-mini 71.6 94.8 (+32.3%) 7.1 4.4 (−-38.0%)
    GPT-5-chat 91.0 98.5 (+8.2%) 4.7 3.0 (−-34.6%)
    GPT-5-mini 91.8 98.5 (+7.3%) 4.9 3.3 (−-33.3%)
    (b) WebArena (ReAct backbone)
    Qwen3.5-4B 10.1 18.7 (+85.3%) 8.6 6.8 (−-20.4%)
    Qwen3.5-9B 21.2 23.9 (+12.8%) 6.8 6.0 (−-12.0%)
    GPT-4o 23.9 29.1 (+21.6%) 5.9 5.1 (−-13.1%)
    GPT-4o-mini 18.0 18.0 (±\pm0.0%) 7.7 7.1 (−-7.7%)
    GPT-5-chat 32.5 37.0 (+13.7%) 5.3 4.8 (−-8.6%)
    GPT-5-mini 32.5 33.7 (+3.8%) 5.4 4.8 (−-10.1%)

    Smaller open-source models experience substantial gains: Qwen3.5-9B on ALFWorld increases success rate by +80.8%+80.8\% relative (54.5% to 98.5%), surpassing the vanilla inducer model (GPT-4o at 96.3%) while reducing turns by 68.2%68.2\%. On WebArena, Qwen3.5-4B gains +85.3%+85.3\% relative in success rate.

  7. Knowl 7 — Skill Library Compactness, Step Compression, and Execution Reliability

    data/table

    Compared to per-trajectory skill induction (ASI_offline), SKILL-DISCO compiles a substantially smaller skill library that compresses more primitive steps per invocation and executes with significantly lower error rates:

    Benchmark Method #Sk Turns ↓\downarrow Sk. turns Avg. prim. Max prim. Err. (%) ↓\downarrow
    ALFWorld ASIoffline\text{ASI}_{\text{offline}} 110 11.4 1.4 0.2 10 75.3
    SKILL-DISCO 5 3.2 2.6 5.6 33 0.0
    WebArena ASIoffline\text{ASI}_{\text{offline}} 146 5.5 1.3 2.5 8 33.9
    SKILL-DISCO 20 5.1 1.0 2.8 36 21.5

    Key metrics:

    • Library Size (#Sk): SKILL-DISCO produces 5 skills on ALFWorld (vs. 110 in ASI_offline) and 20 skills on WebArena (vs. 146 in ASI_offline).
    • Step Compression: On ALFWorld, each SKILL-DISCO skill call collapses an average of 5.6 primitive steps (maximum 33 steps), compared to 0.2 average steps (maximum 10) in ASI_offline. On WebArena, maximum step compression reaches 36 steps (vs. 8 in ASI_offline).
    • Execution Error Rate (Err.): Execution failures drop from 75.3% to 0.0% on ALFWorld and from 33.9% to 21.5% on WebArena.
  8. Knowl 8 — Ablation Analysis of Distillation and Compilation Stages

    data/table

    An ablation study on ALFWorld (CodeAct backbone with GPT-4o) evaluates the distinct roles of the distillation and compilation phases in SKILL-DISCO:

    Configuration #Sk SR (%) Turns In-Tok. (Cached) Out-Tok. Cost ($)
    Vanilla (no skills) 0 96.27 3.63 7,720 (6,231) 561 0.01090
    A1: no distill 43 52.99 11.51 107,528 (97,901) 994 0.05849
    A2: no compile 0 97.01 3.59 19,083 (17,393) 559 0.01417
    Full SKILL-DISCO 5 99.25 3.22 8,468 (6,302) 368 0.01067
    • A1: No Distillation (per-trace induction): Inducing skills from individual successful traces without cross-trace clustering yields 43 skills. The resulting redundancy degrades the success rate to 52.99% (a 46.6%46.6\% relative drop below the full pipeline, performing worse than the 96.27% vanilla agent), increases turns to 11.51, and increases per-episode cost to $0.05849.
    • A2: No Compilation (natural language procedures): Shipping procedural summaries as text descriptions rather than compiled code achieves 97.01% success rate and 3.59 turns, but increases prompt input tokens from 8,468 to 19,083 and per-episode cost to $0.01417 due to in-context procedural descriptions.
    • Full SKILL-DISCO: Achieves the highest success rate (99.25%), fewest turns (3.22), and lowest inference cost ($0.01067).
  9. Knowl 9 — Experimental Induction and Evaluation Setup on ALFWorld and WebArena

    experimental setup

    SKILL-DISCO and baseline methods are evaluated under strict induction/evaluation splits across two interactive agent benchmarks:

    1. ALFWorld: Text-based household simulator featuring navigation, search, and object manipulation. The induction set contains 200 tasks sampled from the official training split; evaluation is conducted on the 134 tasks of the official unseen test split.
    2. WebArena: Interactive web environment covering simulated multi-page websites. The 812 benchmark tasks are partitioned evenly into 406 induction tasks and 406 held-out evaluation tasks.

    Agent Architectures and Baseline Setup:

    • Base agents: ReAct (interleaved reasoning and text action generation) and CodeAct (direct Python code execution).
    • Offline skill induction baselines: AWMoffline\text{AWM}_{\text{offline}} (workflow memory) and ASIoffline\text{ASI}_{\text{offline}} (programmatic skill induction), evaluated on identical induction and evaluation splits.
    • Deployment: Skill signatures and descriptions are injected into the agent system prompt, allowing the LLM to choose dynamically between calling a compiled procedural skill or emitting primitive actions.
  10. Knowl 10 — Limitations of Procedural Skill Discovery via SKILL-DISCO

    limitation

    The SKILL-DISCO framework has three primary limitations:

    1. Restricted to Procedural Tasks: SKILL-DISCO is specifically designed for environments with reusable, state-transition control flows (e.g., embodied navigation, web automation, tool execution). It provides no utility for open-ended natural language generation or single-step reading comprehension where performance depends on linguistic understanding rather than procedural subroutines.
    2. Dependence on Compiler LLM Capability: The distillation and synthesis stages rely on the reasoning and code-generation ability of the LLM compiler. Weak inducer models generate incorrect or overly specific skills, reducing library quality.
    3. Requirement of Pre-Existing Successful Traces: Distillation requires an initial corpus of successful execution traces (T+\mathcal{T}^+). In highly complex or low-success domains where frontier models rarely succeed, the available trace corpus may be too sparse to extract high-coverage reusable skills.

Coverage note — None was omitted; the knowls fully capture the PFSM formulation, the distillation-and-compilation algorithm, reusability metric, contract protocol, main experimental results on ALFWorld and WebArena, cross-model transferability, library statistics, ablations, experimental setup, and limitations.

References

  1. 1.Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou. 2023. Large language models as tool makers. arXiv preprint arXiv:2305.17126.
  2. 2.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. PAL: Program-aided language models. arXiv preprint arXiv:2211.10435.
  3. 3.Shengran Hu, Cong Lu, and Jeff Clune. 2024. Automated design of agentic systems (ADAS). arXiv preprint arXiv:2408.08435.
  4. 4.Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. 2023. Code as policies: Language model programs for embodied control. In IEEE International Conference on Robotics and Automation (ICRA), pages 9493–9500.
  5. 5.Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, and 3 others. 2023. AgentBench: Evaluating LLMs as agents. arXiv preprint arXiv:2308.03688.
  6. 6.Jingwei Ni, Yihao Liu, Xinpeng Liu, Yutao Sun, Mengyu Zhou, Pengyu Cheng, Dexin Wang, Erchao Zhao, Xiaoxi Jiang, and Guanjun Jiang. 2026. Trace2Skill: Distill trajectory-local lessons into transferable agent skills. arXiv preprint arXiv:2603.25158. Work in progress.
  7. 7.Cheng Qian, Chi Han, Yi R. Fung, Yujia Qin, Zhiyuan Liu, and Heng Ji. 2023. CREATOR: Tool creation for disentangling abstract and concrete reasoning of large language models. arXiv preprint arXiv:2305.14318.
  8. 8.Jieyi Qiu, Xiang Juan, Yan Wang, Lei Yang, Xin Qi, Tao Zhang, Jun Guo, Yun Lu, Zheng Yao, Mei Wang, and 1 others. 2025. AgentDistill: Training-free agent distillation with generalizable MCP boxes. arXiv preprint arXiv:2506.14728.
  9. 9.Yu Shang, Yu Li, Keyu Zhao, Likai Ma, Jiahe Liu, Fengli Xu, and Yong Li. 2024. AgentSquare: Automatic LLM agent search in modular design space. arXiv preprint arXiv:2410.06153.
  10. 10.Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36.
  11. 11.Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2021. ALFWorld: Aligning text and embodied environments for interactive learning. In International Conference on Learning Representations.
  12. 12.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291.
  13. 13.Renxi Wang, Xudong Han, Lei Ji, Shu Wang, Timothy Baldwin, and Haonan Li. 2024a. ToolGen: Unified tool retrieval and calling via generation. arXiv preprint arXiv:2410.03439.
  14. 14.Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji. 2024b. Executable code actions elicit better llm agents. In ICML.
  15. 15.Zora Zhiruo Wang, Apurva Gandhi, Graham Neubig, and Daniel Fried. 2025. Inducing programmatic skills for agentic tasks. arXiv preprint arXiv:2504.06821.
  16. 16.Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. 2024c. Agent workflow memory. arXiv preprint arXiv:2409.07429.
  17. 17.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35.
  18. 18.John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik R Narasimhan, and Ofir Press. 2024. SWE-agent: Agent-computer interfaces enable automated software engineering. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
  19. 19.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations.
  20. 20.Guibin Zhang, Luyang Niu, Junfeng Fang, Kun Wang, Lei Bai, and Xiang Wang. 2025. Multi-agent architecture search via agentic supernet. arXiv preprint arXiv:2502.04180.
  21. 21.Jiayi Zhang, Jinyu Xiang, Zhaoyang Yu, Fengwei Teng, Xionghui Chen, Jiaqi Chen, Mingchen Zhuge, Xin Cheng, Sirui Hong, and 1 others. 2024. AFlow: Automating agentic workflow generation. arXiv preprint arXiv:2410.10762.
  22. 22.Boyuan Zheng, Michael Y. Fatemi, Xiaolong Jin, Zora Zhiruo Wang, Apurva Gandhi, Yueqi Song, Yu Gu, Jayanth Srinivasa, and 1 others. 2025. SkillWeaver: Web agents can self-improve by discovering and honing skills. arXiv preprint arXiv:2504.07079.
  23. 23.Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024. WebArena: A realistic web environment for building autonomous agents. In International Conference on Learning Representations.

Citation

MLA
Guo, Z., et al. “SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills”. arXiv, 2026, http://arxiv.org/abs/2606.26669v1.
APA
Guo, Z., Qi, D., Gu, H., Cheng, P., & Xiong, Y. (2026). SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills. arXiv. http://arxiv.org/abs/2606.26669v1
Chicago
Guo, Z., D. Qi, H. Gu, P. Cheng, and Y. Xiong. 2026. “SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills”. arXiv. http://arxiv.org/abs/2606.26669v1.
Harvard
Guo, Z. et al. (2026) “SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2606.26669v1.
Vancouver
1. Guo Z, Qi D, Gu H, Cheng P, Xiong Y (2026) SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills. arXiv

BibTeX

@article{guo2026skill,
  title = {SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills},
  author = {Guo, Zhongxin and Qi, Danrui and Gu, Hanwen and Cheng, Peng and Xiong, Yongqiang},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2606.26669v1},
  eprint = {2606.26669}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/