Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments
Hongjin SuRuoxi SunJinsung YoonPengcheng YinTao YuSercan . Arik
Proposes an automated framework that generates synthetic environment trajectories and task instructions from documentation via backward construction, substantially boosting large language model performance across complex coding, web, and desktop benchmarks without human annotation.
Autonomous digital agents driven by large language models often struggle to perform multi-step tasks across complex software, coding, and web environments. Adapting these models to new environments typically requires expensive human annotations or relies on brittle prompting techniques that fail to capture the nuanced dynamics of digital interfaces.
The article introduces and evaluates Learn-by-interact, a data-centric framework designed to adapt language model agents to new environments autonomously without human labeling. The approach generates realistic task instructions from standard resources like documentation, executes them to collect interaction histories, and fixes instruction-action misalignments by synthesizing new objectives from the resulting trajectories—a process termed backward construction. The resulting synthetic data are then leveraged either as demonstration examples through a multi-tiered agentic retrieval mechanism during inference or directly for model fine-tuning across four leading benchmark environments: SWE-bench, WebArena, OSWorld, and Spider2-V.
The findings show that Learn-by-interact consistently delivers state-of-the-art performance across all four benchmarks. In training-free inference, it boosted task resolution rates by up to 12.2 percentage points with Claude-3.5-sonnet and nearly doubled performance on OSWorld from 12.4% to 22.5%. When used for model fine-tuning, the generated data raised the performance of Codestral-22B on WebArena from 4.7% to 24.2%, an improvement of 19.5 percentage points. Furthermore, combining observation-based and model-based retrieval proved markedly superior to conventional document retrieval, and backward construction yielded up to a 14.0 percentage point improvement in training by eliminating noisy and misaligned interaction paths.
These results demonstrate that high-quality synthetic interaction data can effectively bypass the traditional data-annotation bottleneck, significantly enhancing agent accuracy while maintaining operational efficiency. Unlike complex search methods that multiply token usage and latency during deployment, Learn-by-interact shifts computational costs upstream into the data-generation phase, resulting in faster and cheaper inference in production environments.
Organizations deploying automated software and web agents should adopt autonomous trajectory synthesis and backward construction pipelines instead of relying solely on standard documentation retrieval or manual trajectory logging. For downstream applications, teams should prioritize fine-tuning smaller, domain-adapted models or implementing hybrid retrieval that queries both interface states and operational intent.
Decision-makers should note that the initial data generation and filtering process requires substantial computational resources and multiple model calls, and its effectiveness depends on the accessibility of basic technical documentation or software manuals. Nonetheless, the consistent gains across coding, operating systems, and web applications provide high confidence in the framework's core methodologies for real-world agent adaptation.
- Paper: WebArena: A Realistic Web Environment for Building Autonomous Agents, Shuyan Zhou et al. (2023). It establishes WebArena, one of the primary realistic evaluation environments used to benchmark and validate the autonomous agents in Learn-by-interact.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). It introduces agent-computer interfaces and the interactive software engineering paradigms evaluated on SWE-bench that underpin autonomous agent execution.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). It defines the foundational ReAct framework for interleaving reasoning and environment interaction traces that modern digital agents execute.
- Paper: Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models, Zhihong Shao et al. (2023). It introduces backward reasoning and synthesis techniques to construct demonstrations, providing conceptual foundations for backward trajectory construction.
- Paper: Self-Instruct: Aligning Language Models with Self-Generated Instructions, Yizhong Wang et al. (2023). It establishes the core methodology for bootstrapping instruction-tuning datasets autonomously using a model's self-generated outputs.
- Paper: Mind2Web: Towards a Generalist Agent for the Web, Xiang Deng et al. (2023). It formalizes realistic multi-step web interaction tasks and benchmarks that motivated data-centric adaptation for generalist digital agents.
- Paper: Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models, Andy Zhou et al. (2024). It provides the decision-tree and environment reflection mechanisms that contrast with Learn-by-interact's upstream data generation strategy.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). It extends autonomous agent adaptation by extracting reusable procedural workflows into memory to guide multi-step web navigation tasks.
- Paper: WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration, Yao Zhang et al. (2025). It builds on multi-agent execution in realistic web environments like WebArena by combining global planning with strategic, reflection-guided tree search.
- Paper: Sample-Efficient Learning from Agent Experience, Chenhui Gou et al. (2026). It advances the paradigm of learning from interactive agent experience by distilling trajectory gains directly into model weights with high sample efficiency.
- Paper: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models, Qizheng Zhang et al. (2026). It develops an evolving context engineering playbook that continuously adapts agent strategies from interaction feedback without weight updates.
- Paper: daVinci-Dev: Agent-native Mid-training for Software Engineering, Ji Zeng et al. (2026). It scales agentic adaptation to mid-training foundation models directly on large-scale interactive software engineering trajectories.
- Paper: Agent-as-a-Judge: Evaluate Agents with Agents, Mingchen Zhuge et al. (2025). It introduces agentic judges equipped with environment tools to automatically evaluate complex multi-step trajectories generated by coding agents.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). It synthesizes emerging paradigms in agentic reasoning, providing a broader framework for how agents learn through environmental interaction and feedback.
