SWE-Exp: Experience-Driven Software Issue Resolution
Silin ChenShaoxin LinXiaodong GuYuling ShiHeng LianLongfei YunDong ChenWei SunLinbo CaoQianxiang Wang
Proposes an experience-driven framework that enables automated software engineering agents to learn from both successful and failed past repairs, achieving a 73.0% Pass@1 resolution rate on SWE-bench Verified.
Automated software repair agents powered by large language models have advanced significantly, but they operate primarily as stateless problem-solvers that treat each issue in isolation. Because these systems fail to retain insights from previous repair sessions, they frequently repeat failed exploration strategies, miss opportunities to apply proven fixes across similar contexts, and generate fragile patches that address surface-level symptoms rather than root causes.
The article demonstrates and evaluates SWE-Exp, an experience-driven framework designed to enable continuous cross-repository learning for automated software engineering agents. The system extracts concise, reusable diagnostic perspectives and code modification strategies from past problem-solving attempts to systematically guide future repair tasks.
To establish credibility and avoid data leakage, the authors evaluate the framework on the benchmark dataset of 500 human-verified GitHub issues, strictly excluding past experiences from the same code repository or future time periods. The approach organizes extracted insights into a searchable experience bank, retrieves relevant guidance using semantic matching and automated reranking, and executes fixes through a coordinated dual-agent structure where a high-level planning agent directs a low-level execution agent.
The evaluation yields several key findings in order of importance:
- SWE-Exp achieves state-of-the-art problem resolution rates on the verified benchmark, reaching a 73.0% success rate on the first attempt using Claude 4 Sonnet and 42.0% using DeepSeek-V3, outperforming both non-experience and alternative memory-based baselines under identical model setups.
- Cross-repository experience abstraction and intelligent reranking are essential to performance; removing distilled experience extraction causes the largest drop in success rate (6.0 percentage points), while omitting diagnostic comprehension experiences drops success by 3.2 percentage points.
- Experience quality outweighs quantity: providing a single highly relevant experience yields optimal performance, whereas injecting multiple experiences degrades results due to conflicting information and cognitive overload.
- Expanding the overall experience bank demonstrates clear scaling benefits up to approximately 300 experiences, after which performance gains plateau as core bug patterns are saturated.
- The performance improvements require negligible operational overhead, adding only about $0.01 in computing cost and 37 seconds of retrieval time per issue.
These findings indicate that automated coding systems can transition from trial-and-error exploration to strategic, knowledge-guided problem-solving without incurring substantial runtime or financial costs. By shifting from symptom-level workarounds to principled root-cause fixes, the approach lowers software maintenance risk, improves patch reliability, and accelerates development timelines.
Organizations developing automated code repair systems should adopt structured experience extraction and maintain cross-repository knowledge banks, while ensuring retrieval mechanisms supply focused, single-case guidance rather than dense collections of past attempts. Future development should focus on integrating automated applicability scoring and formal code verification to prevent agents from applying irrelevant patterns to novel problems.
The findings are subject to specific limitations, notably that the evaluation was conducted entirely on Python-based repositories and depends on the quality of automated experience extraction. Readers can maintain high confidence in the demonstrated performance gains for benchmark-aligned tasks, though cautious validation is advised when extending the framework to other programming languages or highly unfamiliar code environments.
- Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Carlos E. Jimenez et al. (2024). This benchmark defines SWE-bench and establishes the standard task of resolving repository-level software engineering issues via execution-based validation.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). This foundational paper introduces agent-computer interfaces that enable autonomous LLM agents to interact with real codebases for software issue resolution.
- Paper: Reflexion: language agents with verbal reinforcement learning, Noah Shinn et al. (2023). This work establishes the paradigm of verbal reinforcement learning and episodic memory from execution feedback, which directly informs SWE-Exp's experience bank design.
- Paper: Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models, Andy Zhou et al. (2024). This study demonstrates how Monte Carlo Tree Search and reflection can be integrated into language agents, setting up the baseline exploration techniques SWE-Exp aims to improve.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). This paper presents the core concepts of iterative code self-debugging via execution feedback that underlying software repair agents build upon.
- Paper: Sample-Efficient Learning from Agent Experience, Chenhui Gou et al. (2026). This paper explores distilling multi-trial agent experiences directly into model weights to retain trajectory-level knowledge without permanent context overhead.
- Paper: AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation, Priyam Sahoo et al. (2026). This work analyzes agent trajectories on SWE-bench to diagnose weak or accidental problem-solving processes versus principled engineering workflows.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). This research generalizes procedural workflow memory extraction to dynamic web navigation environments.
- Paper: LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation, Dongge Han et al. (2026). This paper extends procedural memory architectures to multi-agent LLM systems by curating modular task and subtask plans from prior executions.
- Paper: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models, Qizheng Zhang et al. (2026). This work provides a systematic framework for continuously evolving structured playbooks from successes and failures to prevent context collapse.
- Paper: Human-Inspired Memory Architecture for LLM Agents, Doga Kerestecioglu et al. (2026). This paper introduces a tiered, biologically inspired memory architecture to support long-term consolidation and retention over software tracking histories.
- Paper: MemGym: a Long-Horizon Memory Environment for LLM Agents, Wujiang Xu et al. (2026). This study isolates and benchmarks memory mechanisms across long-horizon interactive agent environments, including repository-level coding.
