Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments

Hongjin SuRuoxi SunJinsung YoonPengcheng YinTao YuSercan . Arik

article2025ICLR99 citations

Proposes an automated framework that generates synthetic environment trajectories and task instructions from documentation via backward construction, substantially boosting large language model performance across complex coding, web, and desktop benchmarks without human annotation.

Listen

Autonomous digital agents driven by large language models often struggle to perform multi-step tasks across complex software, coding, and web environments. Adapting these models to new environments typically requires expensive human annotations or relies on brittle prompting techniques that fail to capture the nuanced dynamics of digital interfaces.

The article introduces and evaluates Learn-by-interact, a data-centric framework designed to adapt language model agents to new environments autonomously without human labeling. The approach generates realistic task instructions from standard resources like documentation, executes them to collect interaction histories, and fixes instruction-action misalignments by synthesizing new objectives from the resulting trajectories—a process termed backward construction. The resulting synthetic data are then leveraged either as demonstration examples through a multi-tiered agentic retrieval mechanism during inference or directly for model fine-tuning across four leading benchmark environments: SWE-bench, WebArena, OSWorld, and Spider2-V.

The findings show that Learn-by-interact consistently delivers state-of-the-art performance across all four benchmarks. In training-free inference, it boosted task resolution rates by up to 12.2 percentage points with Claude-3.5-sonnet and nearly doubled performance on OSWorld from 12.4% to 22.5%. When used for model fine-tuning, the generated data raised the performance of Codestral-22B on WebArena from 4.7% to 24.2%, an improvement of 19.5 percentage points. Furthermore, combining observation-based and model-based retrieval proved markedly superior to conventional document retrieval, and backward construction yielded up to a 14.0 percentage point improvement in training by eliminating noisy and misaligned interaction paths.

These results demonstrate that high-quality synthetic interaction data can effectively bypass the traditional data-annotation bottleneck, significantly enhancing agent accuracy while maintaining operational efficiency. Unlike complex search methods that multiply token usage and latency during deployment, Learn-by-interact shifts computational costs upstream into the data-generation phase, resulting in faster and cheaper inference in production environments.

Organizations deploying automated software and web agents should adopt autonomous trajectory synthesis and backward construction pipelines instead of relying solely on standard documentation retrieval or manual trajectory logging. For downstream applications, teams should prioritize fine-tuning smaller, domain-adapted models or implementing hybrid retrieval that queries both interface states and operational intent.

Decision-makers should note that the initial data generation and filtering process requires substantial computational resources and multiple model calls, and its effectiveness depends on the accessibility of basic technical documentation or software manuals. Nonetheless, the consistent gains across coding, operating systems, and web applications provide high confidence in the framework's core methodologies for real-world agent adaptation.

Cover for Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments

Abstract

Autonomous agents powered by large language models (LLMs) have the potential to enhance human capabilities, assisting with digital tasks from sending emails to performing data analysis. The abilities of existing LLMs at such tasks are often hindered by the lack of high-quality agent data from the corresponding environments they interact with. We propose Learn-by-interact, a data-centric framework to adapt LLM agents to any given environments without human annotations. Learn-by-interact synthesizes trajectories of agent-environment interactions based on documentations, and constructs instructions by summarizing or abstracting the interaction histories, a process called backward construction. We assess the quality of our synthetic data by using them in both training-based scenarios and training-free in-context learning (ICL), where we craft innovative retrieval approaches optimized for agents. Extensive experiments on SWE-bench, WebArena, OSWorld and Spider2-V spanning across realistic coding, web, and desktop environments show the effectiveness of Learn-by-interact in various downstream agentic tasks -- baseline results are improved by up to 12.2% for ICL with Claude-3.5 and 19.5% for training with Codestral-22B. We further demonstrate the critical role of backward construction, which provides up to 14.0% improvement for training. Our ablation studies demonstrate the efficiency provided by our synthesized data in ICL and the superiority of our retrieval pipeline over alternative approaches like conventional retrieval-augmented generation (RAG). We expect that Learn-by-interact will serve as a foundation for agent data synthesis as LLMs are increasingly deployed at real-world environments.

Table of Contents

  • 1 Introduction
  • 2 Learn-by-interact
  • 2.1 Task formulation
  • 2.2 Agentic data synthesis
  • 2.3 Filtering
  • 2.4 Adaptation
  • 3 Experiments
  • 3.1 Baselines
  • 3.2 Datasets
  • 3.3 Settings
  • 3.4 Evaluation
  • 3.5 Results
  • 3.5.1 Training-free Evaluation
  • 3.5.2 Training-based Evaluation
  • 4 Analysis
  • 4.1 Inference Efficiency
  • 4.2 The Impact of Retrieval
  • 4.3 Data granularity
  • 4.4 Scaling Laws
  • 5 Related work
  • 6 Conclusion
  • 7 Limitations
  • References
  • A Baseline implementations
  • B Dataset examples
  • C Experimental settings
  • D Document sources
  • E Synthesized data examples
  • F Case study on filtered examples
  • G Synthesized data from environments
  • H Cross-website generalization

Knowls

  1. Knowl 1 — Learn-by-Interact Agent Data Synthesis Pipeline

    algorithm

    The Learn-by-Interact data synthesis pipeline automatically generates high-quality, environment-specific agent trajectories without human annotations. It leverages standard documentation via self-instruct to generate candidate tasks, executes multi-step rollouts within the target digital environment, derives aligned instruction-trajectory pairs across all sub-trajectories via backward construction, and filters noisy instances.

    Input: Large language model LLMLLM, environment EE, documentation corpus DocDoc, instructions per document NN, data filter FF
    Output: Filtered dataset of instruction-trajectory pairs DD
    D←[]D \leftarrow []
    for each document d∈Docd \in Doc do
        Instructions←LLM(d,N)Instructions \leftarrow LLM(d, N)
        for each instruction I∈InstructionsI \in Instructions do
            E.reset()E.\text{reset}()
            T←[]T \leftarrow []
            while not E.finished()E.\text{finished}() do
                o←E.get_observation()o \leftarrow E.\text{get\_observation}()
                a←LLM(I,T,o)a \leftarrow LLM(I, T, o)
                T←T+[o,a]T \leftarrow T + [o, a]
            end while
            T.append(E.get_observation())T.\text{append}(E.\text{get\_observation}())
            for i←0i \leftarrow 0 to len(T)−3\text{len}(T) - 3 step 2 do
                for j←i+2j \leftarrow i + 2 to len(T)−1\text{len}(T) - 1 step 2 do
                    T′←T[i:j+1]T' \leftarrow T[i : j + 1]
                    I′←LLM(T′)I' \leftarrow LLM(T')
                    D.append([I′,T′])D.\text{append}([I', T'])
                end for
            end for
        end for
    end for
    D←F(D)D \leftarrow F(D)
    return DD

    Given a trajectory T=(o0,a1,o1,…,an,on)T = (o_0, a_1, o_1, \dots, a_n, o_n), decomposing TT into contiguous sub-trajectories T′=(oi,ai+1,…,aj,oj)T' = (o_i, a_{i+1}, \dots, a_j, o_j) for all 0≤i<j≤n0 \le i < j \le n creates a quadratic number of training examples relative to sequence length.

  2. Knowl 2 — Backward Construction for Agent Instruction Synthesis

    model/method

    When LLM agents execute multi-step trajectories in complex environments, action errors frequently cause the trajectory T=(o0,a1,o1,…,an,on)T = (o_0, a_1, o_1, \dots, a_n, o_n) to diverge from the original task instruction II, producing noisy training pairs (I,T)(I, T).

    Backward construction eliminates this misalignment by taking the executed sub-trajectory T′=(oi,ai+1,…,aj,oj)T' = (o_i, a_{i+1}, \dots, a_j, o_j) (where 0≤i<j≤n0 \le i < j \le n) and using an LLM to generate a new instruction I′I' that accurately characterizes the actions actually performed. Backward construction synthesizes two types of instructions:

    1. Sub-step Summaries: For short sub-trajectories, the model generates an instruction summarizing the exact UI interactions and state transitions (e.g., replicating a button click or parameter change).
    2. Goal Abstractions: For full or multi-step trajectories, the model generates a high-level user goal that matches what was achieved by the trajectory rather than what was originally attempted.

    This process converts erroneous or uncompleted rollouts into valid, aligned instruction-trajectory pairs while quadratically expanding the number of usable training examples per rollout.

  3. Knowl 3 — Agentic Retrieval for In-Context Learning

    algorithm

    Agentic retrieval dynamically supplies demonstration trajectories to an LLM agent during multi-turn interactions. It combines an observation-based lexical retriever with an LLM-driven dense query retriever at every step of execution.

    Input: Large language model LLMLLM, environment EE, synthesized dataset DD, BM25 retriever BM25BM25, dense retriever RMRM, task instruction II, observation quota m1m_1, model query quota m2m_2
    Output: Execution history HH
    H←[]H \leftarrow []
    R←[]R \leftarrow []
    while not E.finished()E.\text{finished}() do
        o←E.get_observation()o \leftarrow E.\text{get\_observation}()
        R←BM25(o,D,m1)R \leftarrow BM25(o, D, m_1)
        q←LLM(I,H,o)q \leftarrow LLM(I, H, o)
        R←R+RM(q,D,m2,R)R \leftarrow R + RM(q, D, m_2, R)
        a←LLM(I,H,o,R)a \leftarrow LLM(I, H, o, R)
        H←H+[o,a]H \leftarrow H + [o, a]
    end while
    return HH

    In observation-based retrieval, the current environment observation oo is matched via BM25 against state observations in the synthesized trajectory dataset DD up to m1=5m_1 = 5 examples, retrieving past experiences that encountered the same UI state. In model-based retrieval, the agent prompts an LLM to write a query qq conditioned on the global instruction II, history HH, and current state oo, retrieving up to m2=5m_2 = 5 non-duplicate examples via dense embeddings.

  4. Knowl 4 — Filtering Criteria for Synthesized Trajectories

    model/method

    To prune inferior synthesized instruction-trajectory pairs (I′,T′)(I', T'), a two-tier filtering strategy is applied:

    1. Duplicate State Removal: Consecutive identical state-action steps (ai,oi)=(ai−1,oi−1)(a_i, o_i) = (a_{i-1}, o_{i-1}) are removed from sub-trajectories T′T', eliminating no-op steps caused by invalid UI actions or environment inactivity.
    2. LLM Committee Check: The candidate pair (I′,T′)(I', T') is evaluated by a committee of LLMs (Gemini-1.5-pro and Claude-3.5-sonnet). A pair is accepted only if all committee models agree that it meets all four criteria:
      • Alignment: The trajectory T′T' successfully fulfills the objective in I′I'.
      • Coherence: Each action is logically justified by the preceding observation and consistent with I′I'.
      • Naturalness: The trajectory reflects plausible real-world human behavior.
      • Reasonableness: The solution avoids circular state transitions, excessive detours, or unnatural simplifications.
  5. Knowl 5 — In-Context Learning Performance Across Realistic Agent Benchmarks

    empirical result

    The effectiveness of Learn-by-Interact with agentic retrieval in a training-free in-context learning setting was evaluated across four interactive benchmarks: SWE-bench (software engineering), WebArena (web navigation), OSWorld (desktop operating system tasks), and Spider2-V (multimodal data engineering workflows). Models evaluated include Gemini-1.5-pro and Claude-3.5-sonnet, measured by task resolution percentage (% resolved).

    Could not parse LaTeX table

    Learn-by-Interact outperforms the baseline, static document RAG, unaligned data distillation, Reflexion, and Language Agent Tree Search (LATS) across all benchmarks, achieving up to a +12.2% gain on WebArena and nearly doubling Claude-3.5-sonnet's performance on OSWorld (from 12.4% to 22.5%).

  6. Knowl 6 — Supervised Fine-Tuning Performance on Synthesized Agent Trajectories

    empirical result

    Open-weight models (Codegemma-7B and Codestral-22B) were fine-tuned with LoRA on next-action prediction pairs formatted from Learn-by-Interact synthetic data. Evaluation was performed on WebArena and OSWorld without test-time retrieval, reporting percentage of resolved tasks.

    Could not parse LaTeX table

    Fine-tuning Codestral-22B on Learn-by-Interact data improves task success on WebArena from 4.7% to 24.2% (+19.5%) and on OSWorld from 2.2% to 11.7% (+9.5%), significantly exceeding standard data distillation. Supervised fine-tuning of smaller models yields larger performance improvements than providing the synthetic data as in-context demonstrations.

  7. Knowl 7 — Inference Efficiency: Learn-by-Interact vs Test-Time Search Methods

    empirical result

    Across SWE-bench, WebArena, OSWorld, and Spider2-V using Claude-3.5-sonnet, test-time exploration approaches like Reflexion and Language Agent Tree Search (LATS) incur heavy compute overheads: LATS requires nearly 4×4\times more tokens per instance (over 200k tokens vs ~55k for the baseline) and over 35 LLM calls per instance to attain an average 2.5% accuracy improvement.

    In contrast, Learn-by-Interact achieves an average accuracy of ~37% (vs ~28% for baseline) while using fewer LLM calls per task instance than the baseline (~7 calls vs ~10 calls) and only slightly more tokens (~75k tokens). By retrieving targeted environment interaction demonstrations, the agent completes tasks in fewer total interactive steps, bypassing the runtime latency and compute trade-offs of online tree search.

  8. Knowl 8 — Ablation of Agentic Retrieval Paradigms

    empirical result

    The relative contributions of different retrieval mechanisms within agentic retrieval were evaluated using Gemini-1.5-pro and Claude-3.5-sonnet across four benchmarks, measuring task completion rate (% resolved):

    Could not parse LaTeX table

    Instruction-only retrieval provides the smallest performance gain. Observation-based retrieval (matching current environment state oo) substantially outperforms instruction retrieval (e.g., reaching 14.6% vs 10.2% on Spider2-V for Gemini). Model-based query generation provides further gains, and combining observation-based with model-based retrieval achieves the highest overall accuracy.

  9. Knowl 9 — Effect of Trajectory Granularity on Agent Adaptation

    empirical result

    Synthesized data was partitioned into three trajectory length granularities:

    • Short: <5< 5 interaction steps
    • Medium: 5≤steps<105 \le \text{steps} < 10
    • Long: ≥10\ge 10 interaction steps

    Each single granularity and combination was sub-sampled to an identical data budget of 200M tokens. Performance was measured via in-context learning with Claude-3.5-sonnet and fine-tuning with Codestral-22B.

    Could not parse LaTeX table

    Short trajectories yield higher individual performance gains than medium or long trajectories due to the high reusability of sub-step workflows. Combining short, medium, and long trajectories achieves the best results, showing the complementary value of combining atomic sub-actions and macro-plans.

  10. Knowl 10 — Cross-Website Generalization and Documentation Conditioning Ablations

    empirical result

    Two experimental analyses evaluate generalization and external resource conditioning:

    1. Conditioning on External Documentation: On WebArena with Claude-3.5-sonnet, synthesizing 10k examples conditioned on documentation/tutorials achieves a 48.0% success rate, compared to 39.6% when instructions are generated purely from the environment interface without external references (and 35.8% baseline). External documentation guides task synthesis to align with representative human usage.
    2. Cross-Website Skill Transfer: On WebArena, Content Management Systems (CMS) was held out as a test domain while training or retrieving solely from non-CMS websites.
    Could not parse LaTeX table

    Even when CMS synthetic data is entirely excluded, Codestral-22B improves from 3.3% to 12.6% on CMS tasks, indicating strong cross-website transfer of learned interactive web navigation actions.

  11. Knowl 11 — Limitations of Learn-by-Interact

    limitation

    The Learn-by-Interact data synthesis framework exhibits two main limitations:

    1. Compute and Inference Costs During Data Construction: Autonomous rollout generation, backward construction of sub-trajectory instructions, and multi-model committee filtering require a large volume of LLM API calls before downstream adaptation.
    2. Reliance on Quality Reference Documentation: Task proposal relies on external documentation, tutorials, or user manuals to cover realistic use cases. In proprietary or newly created environments lacking accessible reference documentation, instruction synthesis requires unconstrained exploration, which is harder to control for task diversity and user intent alignment.

Coverage note — No substantial contributed material was omitted; the knowls cover the full synthesis pipeline, backward construction, agentic retrieval, filtering, main ICL/SFT benchmarks, efficiency, granularity, retrieval ablations, generalization, and limitations.

References

  1. 1.J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.R. Aksitov, S. Miryoosefi, Z. Li, D. Li, S. Babayan, K. Kopparapu, Z. Fisher, R. Guo, S. Prakash, P. Srinivasan, et al. Rest meets react: Self-improvement for multi-step reasoning llm agent. arXiv preprint arXiv:2312.10003, 2023.
  3. 3.Anthropic. Introducing claude 3.5 sonnet, 2024. URL https://www.anthropic.com/news/claude-3-5-sonnet.
  4. 4.P. J. Ball, L. Smith, I. Kostrikov, and S. Levine. Efficient online reinforcement learning with offline data. In International Conference on Machine Learning, pages 1577–1594. PMLR, 2023.
  5. 5.R. Cao, F. Lei, H. Wu, J. Chen, Y. Fu, H. Gao, X. Xiong, H. Zhang, Y. Mao, W. Hu, et al. Spider2-v: How far are multimodal agents from automating data science and engineering workflows? arXiv preprint arXiv:2407.10956, 2024.
  6. 6.B. Chen, C. Shu, E. Shareghi, N. Collier, K. Narasimhan, and S. Yao. Fireact: Toward language agent fine-tuning. arXiv preprint arXiv:2310.05915, 2023.
  7. 7.D. Chen, S. Lin, M. Zeng, D. Zan, J.-G. Wang, A. Cheshkov, J. Sun, H. Yu, G. Dong, A. Aliev, et al. Coder: Issue resolving with multi-agent and task graphs. arXiv preprint arXiv:2406.01304, 2024a.
  8. 8.Z. Chen, K. Liu, Q. Wang, W. Zhang, J. Liu, D. Lin, K. Chen, and F. Zhao. Agent-flan: Designing data and methods of effective agent tuning for large language models. arXiv preprint arXiv:2403.12881, 2024b.
  9. 9.R. Coulom. Efficient selectivity and backup operators in monte-carlo tree search. In International conference on computers and games, pages 72–83. Springer, 2006.
  10. 10.X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems, 36, 2024.
  11. 11.A. Drouin, M. Gasse, M. Caccia, I. H. Laradji, M. Del Verme, T. Marty, D. Vazquez, N. Chapados, and A. Lacoste. WorkArena: How capable are web agents at solving common knowledge work tasks? In R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp, editors, Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 11642–11662. PMLR, 21–27 Jul 2024. URL https://proceedings.mlr.press/v235/drouin24a.html.
  12. 12.C. Gulcehre, T. L. Paine, S. Srinivasan, K. Konyushkova, L. Weerts, A. Sharma, A. Siddhant, A. Ahern, M. Wang, C. Gu, et al. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998, 2023.
  13. 13.I. Gur, H. Furuta, A. Huang, M. Safdari, Y. Matsuo, D. Eck, and A. Faust. A real-world webagent with planning, long context understanding, and program synthesis. arXiv preprint arXiv:2307.12856, 2023.
  14. 14.E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  15. 15.M. Hu, P. Zhao, C. Xu, Q. Sun, J. Lou, Q. Lin, P. Luo, S. Rajmohan, and D. Zhang. Agentgen: Enhancing planning abilities for large language model based agent via environment and task generation. arXiv preprint arXiv:2408.00764, 2024.
  16. 16.W. Huang, P. Abbeel, D. Pathak, and I. Mordatch. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International conference on machine learning, pages 9118–9147. PMLR, 2022.
  17. 17.C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770, 2023.
  18. 18.A. Keipour. Physical interaction and manipulation of the environment using aerial robots. arXiv preprint arXiv:2207.02856, 2022.
  19. 19.L. Kocsis and C. Szepesvári. Bandit based monte-carlo planning. In European conference on machine learning, pages 282–293. Springer, 2006.
  20. 20.J. Y. Koh, R. Lo, L. Jang, V. Duvvur, M. C. Lim, P.-Y. Huang, G. Neubig, S. Zhou, R. Salakhutdinov, and D. Fried. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. arXiv e-prints, pages arXiv–2401, 2024.
  21. 21.Y. Li, J. He, X. Zhou, Y. Zhang, and J. Baldridge. Mapping natural language instructions to mobile ui action sequences. arXiv preprint arXiv:2005.03776, 2020.
  22. 22.Y. Liu, K. Shi, K. S. He, L. Ye, A. R. Fabbri, P. Liu, D. Radev, and A. Cohan. On learning to summarize with large language models as references. arXiv preprint arXiv:2305.14239, 2023.
  23. 23.A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36, 2024.
  24. 24.O. Nachum, S. S. Gu, H. Lee, and S. Levine. Data-efficient hierarchical reinforcement learning. Advances in neural information processing systems, 31, 2018.
  25. 25.X. Pu, M. Gao, and X. Wan. Summarization is (almost) dead. arXiv preprint arXiv:2309.09558, 2023.
  26. 26.M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  27. 27.M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. Courville, and P. Bachman. Data-efficient reinforcement learning with self-predictive representations. arXiv preprint arXiv:2007.05929, 2020.
  28. 28.M. Schwarzer, N. Rajkumar, M. Noukhovitch, A. Anand, L. Charlin, R. D. Hjelm, P. Bachman, and A. C. Courville. Pretraining representations for data-efficient reinforcement learning. Advances in Neural Information Processing Systems, 34:12686–12699, 2021.
  29. 29.N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36, 2024.
  30. 30.H. Su, J. Kasai, C. H. Wu, W. Shi, T. Wang, J. Xin, R. Zhang, M. Ostendorf, L. Zettlemoyer, N. A. Smith, et al. Selective annotation makes language models better few-shot learners. arXiv preprint arXiv:2209.01975, 2022.
  31. 31.C. Team. Codegemma: Open code models based on gemma. arXiv preprint arXiv:2406.11409, 2024a.
  32. 32.T. M. A. Team. Codestral: Hello, world!, 2024b. URL https://mistral.ai/news/codestral/.
  33. 33.P. Thomas and E. Brunskill. Data-efficient off-policy policy evaluation for reinforcement learning. In International Conference on Machine Learning, pages 2139–2148. PMLR, 2016.
  34. 34.G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023a.
  35. 35.R. Wang, P. Jansen, M.-A. Côté, and P. Ammanabrolu. Scienceworld: Is your agent smarter than a 5th grader? arXiv e-prints, pages arXiv–2203, 2022a.
  36. 36.X. Wang, Y. Chen, L. Yuan, Y. Zhang, Y. Li, H. Peng, and H. Ji. Executable code actions elicit better llm agents. arXiv preprint arXiv:2402.01030, 2024a.
  37. 37.X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y. Song, B. Li, J. Singh, H. H. Tran, F. Li, R. Ma, M. Zheng, B. Qian, Y. Shao, N. Muennighoff, Y. Zhang, B. Hui, J. Lin, R. Brennan, H. Peng, H. Ji, and G. Neubig. OpenHands: An Open Platform for AI Software Developers as Generalist Agents, 2024b. URL https://arxiv.org/abs/2407.16741.
  38. 38.Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi. Self-instruct: Aligning language models with self-generated instructions. arXiv preprint arXiv:2212.10560, 2022b.
  39. 39.Z. Wang, S. Cai, G. Chen, A. Liu, X. Ma, and Y. Liang. Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents. arXiv preprint arXiv:2302.01560, 2023b.
  40. 40.T. Xie, F. Zhou, Z. Cheng, P. Shi, L. Weng, Y. Liu, T. J. Hua, J. Zhao, Q. Liu, C. Liu, et al. Openagents: An open platform for language agents in the wild. arXiv preprint arXiv:2310.10634, 2023.
  41. 41.T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. arXiv preprint arXiv:2404.07972, 2024.
  42. 42.J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press. Swe-agent: Agent-computer interfaces enable automated software engineering. arXiv preprint arXiv:2405.15793, 2024.
  43. 43.Z. Yang, J. Liu, Y. Han, X. Chen, Z. Huang, B. Fu, and G. Yu. Appagent: Multimodal agents as smartphone users. arXiv preprint arXiv:2312.13771, 2023.
  44. 44.S. Yao, H. Chen, J. Yang, and K. Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems, 35:20744–20757, 2022a.
  45. 45.S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022b.
  46. 46.D. Yin, F. Brahman, A. Ravichander, K. Chandu, K.-W. Chang, Y. Choi, and B. Y. Lin. Lumos: Learning agents with unified data, modular design, and open-source llms. arXiv preprint arXiv:2311.05657, 2023.
  47. 47.A. Zeng, M. Liu, R. Lu, B. Wang, X. Liu, Y. Dong, and J. Tang. Agenttuning: Enabling generalized agent abilities for llms. arXiv preprint arXiv:2310.12823, 2023.
  48. 48.Z. Zhan and A. Zhang. You only look at screens: Multimodal chain-of-action agents. arXiv preprint arXiv:2309.11436, 2023.
  49. 49.J. Zhang, Y. Yu, M. Liao, W. Li, J. Wu, and Z. Wei. Ui-hawk: Unleashing the screen stream understanding for gui agents. arXiv preprint, 2024.
  50. 50.Z. Zhao, K. Ma, W. Chai, X. Wang, K. Chen, D. Guo, Y. Zhang, H. Wang, and G. Wang. Do we really need a complex agent system? distill embodied agent into a single model. arXiv preprint arXiv:2404.04619, 2024.
  51. 51.A. Zhou, K. Yan, M. Shlapentokh-Rothman, H. Wang, and Y.-X. Wang. Language agent tree search unifies reasoning acting and planning in language models. arXiv preprint arXiv:2310.04406, 2023a.
  52. 52.S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, et al. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854, 2023b.

Citation

MLA
Su, H., et al. “Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments”. arXiv, 2025, http://arxiv.org/abs/2501.10893v1.
APA
Su, H., Sun, R., Yoon, J., Yin, P., Yu, T., & Arık, S. Ö. (2025). Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments. arXiv. http://arxiv.org/abs/2501.10893v1
Chicago
Su, H., R. Sun, J. Yoon, P. Yin, T. Yu, and S. Ö. Arık. 2025. “Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments”. arXiv. http://arxiv.org/abs/2501.10893v1.
Harvard
Su, H. et al. (2025) “Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.10893v1.
Vancouver
1. Su H, Sun R, Yoon J, Yin P, Yu T, Arık SÖ (2025) Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments. arXiv

BibTeX

@article{su2025learn,
  title = {Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments},
  author = {Su, Hongjin and Sun, Ruoxi and Yoon, Jinsung and Yin, Pengcheng and Yu, Tao and Arık, Sercan Ö.},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.10893v1},
  eprint = {2501.10893}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission