SWE-Exp: Experience-Driven Software Issue Resolution

Silin ChenShaoxin LinXiaodong GuYuling ShiHeng LianLongfei YunDong ChenWei SunLinbo CaoQianxiang Wang

article2025arXiv73 citations

Proposes an experience-driven framework that enables automated software engineering agents to learn from both successful and failed past repairs, achieving a 73.0% Pass@1 resolution rate on SWE-bench Verified.

Listen

Automated software repair agents powered by large language models have advanced significantly, but they operate primarily as stateless problem-solvers that treat each issue in isolation. Because these systems fail to retain insights from previous repair sessions, they frequently repeat failed exploration strategies, miss opportunities to apply proven fixes across similar contexts, and generate fragile patches that address surface-level symptoms rather than root causes.

The article demonstrates and evaluates SWE-Exp, an experience-driven framework designed to enable continuous cross-repository learning for automated software engineering agents. The system extracts concise, reusable diagnostic perspectives and code modification strategies from past problem-solving attempts to systematically guide future repair tasks.

To establish credibility and avoid data leakage, the authors evaluate the framework on the benchmark dataset of 500 human-verified GitHub issues, strictly excluding past experiences from the same code repository or future time periods. The approach organizes extracted insights into a searchable experience bank, retrieves relevant guidance using semantic matching and automated reranking, and executes fixes through a coordinated dual-agent structure where a high-level planning agent directs a low-level execution agent.

The evaluation yields several key findings in order of importance:

  1. SWE-Exp achieves state-of-the-art problem resolution rates on the verified benchmark, reaching a 73.0% success rate on the first attempt using Claude 4 Sonnet and 42.0% using DeepSeek-V3, outperforming both non-experience and alternative memory-based baselines under identical model setups.
  2. Cross-repository experience abstraction and intelligent reranking are essential to performance; removing distilled experience extraction causes the largest drop in success rate (6.0 percentage points), while omitting diagnostic comprehension experiences drops success by 3.2 percentage points.
  3. Experience quality outweighs quantity: providing a single highly relevant experience yields optimal performance, whereas injecting multiple experiences degrades results due to conflicting information and cognitive overload.
  4. Expanding the overall experience bank demonstrates clear scaling benefits up to approximately 300 experiences, after which performance gains plateau as core bug patterns are saturated.
  5. The performance improvements require negligible operational overhead, adding only about $0.01 in computing cost and 37 seconds of retrieval time per issue.

These findings indicate that automated coding systems can transition from trial-and-error exploration to strategic, knowledge-guided problem-solving without incurring substantial runtime or financial costs. By shifting from symptom-level workarounds to principled root-cause fixes, the approach lowers software maintenance risk, improves patch reliability, and accelerates development timelines.

Organizations developing automated code repair systems should adopt structured experience extraction and maintain cross-repository knowledge banks, while ensuring retrieval mechanisms supply focused, single-case guidance rather than dense collections of past attempts. Future development should focus on integrating automated applicability scoring and formal code verification to prevent agents from applying irrelevant patterns to novel problems.

The findings are subject to specific limitations, notably that the evaluation was conducted entirely on Python-based repositories and depends on the quality of automated experience extraction. Readers can maintain high confidence in the demonstrated performance gains for benchmark-aligned tasks, though cautious validation is advised when extending the framework to other programming languages or highly unfamiliar code environments.

Cover for SWE-Exp: Experience-Driven Software Issue Resolution

Abstract

Recent advances in large language model (LLM) agents have shown remarkable progress in software issue resolution, leveraging advanced techniques such as multi-agent collaboration and Monte Carlo Tree Search (MCTS). However, current agents act as memoryless explorers - treating each problem separately without retaining or reusing knowledge from previous repair experiences. This leads to redundant exploration of failed trajectories and missed chances to adapt successful issue resolution methods to similar problems. To address this problem, we introduce SWE-Exp, an experience-enhanced approach that distills concise and actionable experience from prior agent trajectories, enabling continuous learning across issues. Our method introduces a multi-faceted experience bank that captures both successful and failed repair attempts. Specifically, it extracts reusable issue resolution knowledge at different levels - from high-level problem comprehension to specific code changes. Experiments show that SWE-Exp achieves a Pass@1 resolution rate of 73.0% on SWE-Bench Verified using the state-of-the-art LLM Claude 4 Sonnet, significantly outperforming prior results under other agent frameworks. Our approach establishes a new paradigm in which automated software engineering agents systematically accumulate and leverage repair expertise, fundamentally shifting from trial-and-error exploration to strategic, experience-driven issue resolution.

Table of Contents

  • 1 Introduction
  • 2 Motivation
  • 3 Methodology
  • 3.1 Conceptual Framework
  • 3.2 Trajectories Collection
  • 3.3 Experiences Extraction
  • 3.3.1 Experience Representation.
  • 3.3.2 Offline Embedding and Storage
  • 3.3.3 Multi-facet Categorization
  • 3.4 Experiences Reuse
  • 3.4.1 Experience Retrieval
  • 3.4.2 Agent Role Separation
  • 4 Experimental Setup
  • 4.1 Research Questions
  • 4.2 Datasets
  • 4.3 Baselines
  • 4.4 Implementation Details
  • 5 Results
  • 5.1 RQ1: Effectiveness
  • 5.2 RQ2: Ablation Study
  • 5.3 RQ3: Impact of Hyperparameters
  • 5.4 Case Study
  • 6 Discussion
  • 6.1 Data Leakage
  • 6.2 Quality of Extracted Experience
  • 6.3 Cost Analysis
  • 6.4 Limitations and Future Directions
  • 7 Threats to Validity
  • 8 Related Work
  • 8.1 Repository-Level Issue Resolution
  • 8.2 Experience Enhanced AI Agents
  • 9 Conclusion
  • References
  • A Hyperparameters of MCTS
  • B Prompt Templates
  • B.1 Instructor
  • B.2 Assistant
  • B.3 Issue Agent
  • B.4 Issue Comprehension ExpAgent
  • B.4.1 Successful Experience Extraction Prompt
  • B.4.2 Failed Experience Extraction Prompt
  • B.5 Modification ExpAgent
  • B.6 RerankAgent
  • B.7 Reuser
  • B.7.1 Reuse Comprehension Experience Prompt
  • B.7.2 Reuse Modification Experience Prompt

Knowls

  1. Knowl 1 — SWE-Exp Framework for Experience-Driven Software Issue Resolution

    model/method

    SWE-Exp is an automated program repair framework designed to replace memoryless, trial-and-error repository issue exploration with continuous, cross-issue experiential learning. For an issue instance pp, standard approaches generate an action trajectory τ=(s0,a0,s1,a1,… )\tau = (s_0, a_0, s_1, a_1, \dots), where sts_t denotes the repository state at time step tt and ata_t is the action taken, treating previous trajectories H={τ1,τ2,… }\mathcal{H} = \{\tau_1, \tau_2, \dots\} independently. In contrast, SWE-Exp maintains a persistent experience bank Bexp\mathcal{B}_{\text{exp}} containing distilled knowledge from past resolution trajectories across multiple repositories.

    When presented with a new issue p′p', SWE-Exp operates across a four-stage pipeline:

    1. Trajectory Collection: Collects execution traces from both successful and failed repair attempts across diverse codebases, annotating failed attempts with specific failure modes (such as localization errors, flawed modification strategies, or incorrect problem comprehension).
    2. Experience Extraction: An offline extraction agent transforms raw trajectories into structured, reusable knowledge units organized into abstract problem comprehension patterns and code modification strategies.
    3. Experience Retrieval and Reranking: When resolving a new issue, the system retrieves the top N=10N=10 candidate experiences from Bexp\mathcal{B}_{\text{exp}} via dense vector similarity over error types and descriptions, after which a specialized reranking agent selects the single most relevant experience (k=1k=1), denoted E′=Retrieve(p′,Bexp)E' = \text{Retrieve}(p', \mathcal{B}_{\text{exp}}).
    4. Experience-Conditioned Tree Search: The agent executes Monte Carlo Tree Search (MCTS) conditioned on E′E', using a hierarchical dual-agent system (an Instructor planner and an Assistant executor) to explore the repository and synthesize patches.
  2. Knowl 2 — Multi-Facet Experience Representation and Experience Bank Indexing

    model/method

    In SWE-Exp, raw agent trajectories are represented as sequences of tuples ⟨(dt,at,st+1,ft)⟩t=0N\langle (d_t, a_t, s_{t+1}, f_t) \rangle_{t=0}^{N}, where dtd_t denotes high-level directives, ata_t represents low-level actions, st+1s_{t+1} captures resulting repository states, and ftf_t denotes environment feedback. An offline Experiencer agent distills these trajectories into structured experience dictionaries comprising two distinct facets:

    1. Comprehension Experiences (Perspectives): Abstract conceptual reasoning patterns that encode how an issue was diagnosed. They specify high-level symptoms, diagnostic hypotheses, and entry-point reasoning without hardcoding repository-specific entity names.
    2. Modification Experiences (Modifications): Generalized strategies for code modifications derived from successful patches. They capture code-level repair principles, such as parameter-handling logic, defensive copying, separating validation from default assignment, and mitigating unintended regression side effects.

    To facilitate fast and semantically relevant retrieval, each extracted experience is embedded using the Multilingual-E5-Large model and stored in the vector database (Experience Bank Bexp\mathcal{B}_{\text{exp}}) indexed by two metadata fields:

    • Issue Type: A generalized descriptive label categorizing the error class (e.g., VariableReferenceError, AttributeError).
    • Description: A generalized natural language summary outlining the typical conditions and operational scenarios under which this error occurs.
  3. Knowl 3 — Hierarchical Dual-Agent Architecture for Action Decoupling in MCTS

    model/method

    To prevent Monte Carlo Tree Search (MCTS) agents from over-relying on code exploration commands (such as searches) while neglecting actual code modifications, SWE-Exp decouples high-level strategic reasoning from low-level execution via a hierarchical dual-agent architecture:

    • Instructor Agent (High-Level Planner): Analyzes the interaction history and retrieved comprehension experiences (provided as background context). At each decision point, the Instructor generates a structured JSON object specifying:
      1. thoughts: Summary of current repository state, exploration history, and reasoning.
      2. instructions: A specific, goal-oriented task for the Assistant without prescribing low-level tool invocation mechanics.
      3. context: File paths, class names, function names, code blocks, or line ranges extracted from real environment feedback.
      4. type: The required high-level action category, selected from search, view, modify, or finish.
    • Assistant Agent (Low-Level Executor): Receives the Instructor's directive along with retrieved modification experiences to formulate and execute concrete environment actions. The Assistant's action space includes:
      • Code localization: FindClass, FindFunction, FindCodeSnippet, SemanticSearch.
      • Code inspection: ViewCode.
      • Code modification: StringReplace, CreateFile.
      • Completion: Finish.

    This role separation streamlines strategic planning by relieving the planner of low-level tool argument handling and allows modification experiences to directly guide code editing.

  4. Knowl 4 — Temporal and Cross-Repository Data Leakage Prevention Protocol

    model/method

    To prevent knowledge leakage and ensure rigorous evaluation on benchmark datasets such as SWE-Bench, SWE-Exp enforces two strict retrieval constraints during experience selection:

    1. Cross-Repository Boundary Enforcement: For any target issue instance, all experiences and trajectories originating from the same software repository are strictly excluded from the candidate retrieval pool.
    2. Temporal Ordering Constraint: All repair trajectories generated chronologically after the creation timestamp of the target issue instance are excluded.

    These constraints ensure that the agent cannot access repository-specific artifacts, memorized bug fixes, or ground-truth-related shortcuts, verifying that performance gains arise strictly from the transfer of generalized diagnostic and repair patterns across heterogeneous codebases.

  5. Knowl 5 — Issue Resolution Performance on SWE-Bench-Verified

    data/table

    The issue resolution effectiveness of SWE-Exp was evaluated on the SWE-Bench-Verified benchmark (comprising 500 human-verified GitHub issues) using the Pass@1 metric (the percentage of issues resolved correctly on the first attempt). Evaluations compared SWE-Exp against agentic and non-agentic baselines across both open-source (DeepSeek-V3-0324) and closed-source (Claude-4 Sonnet 20250514) foundation models.

    Method Model Pass@1
    Agentless DeepSeek-V3-0324 36.6%
    SWE-Agent DeepSeek-V3-0324 38.8%
    SWE-Agent Claude-4 sonnet (20250514) 66.6%
    SWE-Search DeepSeek-V3-0324 35.4%
    SWE-Search Claude-4 sonnet (20250514) 70.8%
    OpenHands DeepSeek-V3-0324 38.8%
    OpenHands Claude-4 sonnet (20250514) 70.4%
    EvoCoder DeepSeek-V3-0324 38.0%
    EvoCoder Claude-4 sonnet (20250514) 67.8%
    SAGE Claude-4 sonnet (20250514) 71.6%
    SWE-Exp DeepSeek-V3-0324 42.0%
    SWE-Exp Claude-4 sonnet (20250514) 73.0%

    Under DeepSeek-V3-0324, SWE-Exp achieved 42.0% Pass@1, outperforming SWE-Search (35.4%) and other baselines (36.6%–38.8%). Under Claude-4 Sonnet, SWE-Exp achieved 73.0% Pass@1, surpassing SWE-Search (70.8%) and experience-driven continuous learning baselines EvoCoder (67.8%) and SAGE (71.6%).

  6. Knowl 6 — Ablation Study of SWE-Exp Components

    data/table

    An ablation study evaluated the individual contributions of the core components in SWE-Exp using the DeepSeek-V3-0324 backbone on the SWE-Bench-Verified benchmark (full framework Pass@1 of 42.0%).

    Method Pass@1 Δ\Delta
    SWE-Exp 42.0% –
    w/o Comprehension Experience 38.8% -3.2%
    w/o Modification Experience 39.4% -2.6%
    w/o Dual-Agent 39.8% -2.2%
    w/o Experiences Extraction 36.0% -6.0%
    w/o LLM Reranking 38.2% -3.8%

    Key takeaways:

    • Removing Experiences Extraction (which replaces abstracted experiences with raw problem statements and golden patches directly as in-context demonstrations) produced the largest degradation (-6.0%), confirming that abstracting reusable principles across repositories is essential for effective knowledge transfer.
    • Disabling LLM Reranking (relying strictly on embedding cosine similarity) reduced Pass@1 by -3.8%, showing the importance of LLM-based context-aware relevance filtering.
    • Removing Comprehension Experiences led to a -3.2% decrease, while removing Modification Experiences resulted in a -2.6% decrease, showing their complementary roles.
    • Collapsing the Dual-Agent Architecture into a single monolithic agent caused a -2.2% drop, confirming the value of separating strategic planning from tactical tool execution.
  7. Knowl 7 — Impact of Retrieved Experience Count and Experience Bank Size

    empirical result

    Empirical sensitivity experiments on SWE-Bench-Verified using the DeepSeek-V3-0324 backbone demonstrated the following effects of experience retrieval parameters:

    1. Number of Retrieved Experiences per Issue (kk):

      • Testing k∈{0,1,2,3,4}k \in \{0, 1, 2, 3, 4\} on SWE-Bench-Verified yielded Pass@1 scores of:
        • k=0k=0 (no experience): 37.8%
        • k=1k=1: 42.0% (optimal peak)
        • k=2k=2: 40.4%
        • k=3k=3: 40.2%
        • k=4k=4: 39.6%
      • Performance exhibits an increase-then-decline pattern. A single (k=1k=1) highly relevant experience provides optimal guidance, whereas retrieving multiple experiences (k≥2k \ge 2) introduces cognitive overhead, conflicting directives, and distraction.
    2. Experience Bank Size Scaling:

      • Evaluating experience bank capacity in increments of 100 on a representative 75-issue subset (25 issues each from Django, SymPy, and Sphinx) yielded Pass@1 scores of:
        • 0 experiences: 46.7%
        • 100 experiences: 49.3%
        • 200 experiences: 53.3%
        • 300 experiences: 57.3%
        • 400 experiences: 56.0%
        • 500 experiences: 57.3%
      • Performance increases steadily up to approximately 300 experiences, beyond which resolution rates plateau as common bug patterns and repair templates reach saturation.
  8. Knowl 8 — Computational and Financial Cost Overhead of SWE-Exp

    data/table

    The computational, financial, and time costs of SWE-Exp were compared against the baseline SWE-Search framework on SWE-Bench-Verified using the DeepSeek-V3-0324 backbone.

    Metrics SWE-Search SWE-Exp
    Average Token Costs 189.1K 203.3K
    Average API Costs $0.12 $0.13
    Average Wall Time 12min37s 15min49s
    Average Retrieval and Rerank time 0s 37.5s

    The dual-agent architecture and experience retrieval mechanism in SWE-Exp increase average token consumption by 7.5% (203.3K vs. 189.1K) and average API usage cost by $0.01 per instance ($0.13 vs. $0.12). Average wall-clock execution time increases by 3 minutes and 12 seconds per instance, of which experience retrieval and reranking account for 37.5 seconds.

  9. Knowl 9 — Vulnerability to Diagnostic Anchoring and Context Distraction in Experience Reuse

    limitation

    The integration of retrieved experiences introduces specific failure modes and vulnerabilities:

    1. Diagnostic Anchoring and Misleading Thoughts: In the problem comprehension stage, retrieved experiences can bias the Instructor agent toward inappropriate hypotheses. If the Instructor is forced to explicitly cite experiences in its reasoning, it may continue relying on past patterns even when codebase exploration provides contradictory evidence. Supplying experiences as background message context rather than mandatory instructions mitigates this dependency but does not eliminate it.
    2. Context Congestion from Multiple Experiences: Increasing the number of past experiences retrieved for a single instance introduces conflicting or irrelevant guidelines, causing the agent to lose focus on the specific requirements of the target issue.
    3. Lack of Dynamic Applicability Verification: SWE-Exp does not include a dynamic mechanism or confidence scoring model to evaluate whether a retrieved experience is truly applicable to the target codebase's architectural context, occasionally leading to negative transfer.

Coverage note — Specific prompt templates from Appendix B and the single-instance qualitative case study on django-11964 were deliberately omitted as standalone knowls because their core concepts are fully represented in the methodology, dual-agent architecture, and empirical findings.

References

  1. 1.Anthropic. 2025. Introducing Claude 4. https://www.anthropic.com/news/claude-4
  2. 2.Antonis Antoniades, Albert Örwall, Kexun Zhang, Yuxi Xie, Anirudh Goyal, and William Wang. 2024. SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement. arXiv:2410.20285
  3. 3.Dong Chen, Shaoxin Lin, Muhan Zeng, Daoguang Zan, Jian-Gang Wang, Anton Cheshkov, Jun Sun, Hao Yu, Guoliang Dong, Artem Aliev, Jie Wang, Xiao Cheng, Guangtai Liang, Yuchi Ma, Pan Bian, Tao Xie, and Qianxiang Wang. 2024. CodeR: Issue Resolving with Multi-Agent and Task Graphs. arXiv:2406.01304 [cs]
  4. 4.Zhi Chen, Wei Ma, and Lingxiao Jiang. 2025. Unveiling Pitfalls: Understanding Why AI-driven Code Agents Fail at GitHub Issue Resolution. arXiv:2503.12374 [cs]
  5. 5.Zhaoling Chen, Xiangru Tang, Gangda Deng, Fang Wu, Jialong Wu, Zhiwei Jiang, Viktor Prasanna, Arman Cohan, and Xingyao Wang. 2025. LocAgent: Graph-Guided LLM Agents for Code Localization. arXiv:2503.09089 [cs]
  6. 6.Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, Nicholas Thumiger, Aditya Desai, Ion Stoica, Ana Klimovic, Graham Neubig, and Joseph E. Gonzalez. 2025. The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks.
  7. 7.DeepSeek-AI. 2025. DeepSeek-V3 Technical Report. arXiv:2412.19437 [cs]
  8. 8.Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2024. Improving Factuality and Reasoning in Language Models through Multiagent Debate. In Proceedings of the 41st International Conference on Machine Learning (ICML'24, Vol. 235). JMLR.org, Vienna, Austria, 11733–11763.
  9. 9.Ryan Ehrlich, Bradley Brown, Jordan Juravsky, Ronald Clark, Christopher Ré, and Azalia Mirhoseini. 2025. CodeMonkeys: Scaling Test-Time Compute for Software Engineering.
  10. 10.Erhu Feng, Wenbo Zhou, Zibin Liu, Le Chen, Yunpeng Dong, Cheng Zhang, Yisheng Zhao, Dong Du, Zhichao Hua, Yubin Xia, and Haibo Chen. 2025. Get Experience from Practice: LLM Agents with Record & Replay. arXiv:2505.17716 [cs]
  11. 11.Yao Fu, Dong-Ki Kim, Jaekyeom Kim, Sungryull Sohn, Lajanugen Logeswaran, Kyunghoon Bae, and Honglak Lee. 2024. AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents. arXiv:2403.08978 [cs]
  12. 12.Hiroaki Hayashi, Bo Pang, Wenting Zhao, Ye Liu, Akash Gokul, Srijan Bansal, Caiming Xiong, Semih Yavuz, and Yingbo Zhou. 2025. Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement. arXiv preprint arXiv:2511.05931 (2025).
  13. 13.Dong Huang, Qingwen Bu, Jie M. Zhang, Michael Luck, and Heming Cui. 2023. AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation. arXiv:2312.13010
  14. 14.Mingjian Jiang, Yangjun Ruan, Luis Lastras, Pavan Kapanipathi, and Tatsunori Hashimoto. 2025. Putting It All into Context: Simplifying Agents with LCLMs. arXiv:2505.08120 [cs]
  15. 15.Zhonghao Jiang, Xiaoxue Ren, Meng Yan, Wei Jiang, Yong Li, and Zhongxin Liu. 2025. CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching. arXiv:2503.22424 [cs]
  16. 16.Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R. Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real-world Github Issues?. In ICLR.
  17. 17.Han Li, Yuling Shi, Shaoxin Lin, Xiaodong Gu, Heng Lian, Xin Wang, Yantao Jia, Tao Huang, and Qianxiang Wang. 2025. Swe-debate: Competitive multi-agent debate for software issue resolution. arXiv preprint arXiv:2507.23348 (2025).
  18. 18.Yalan Lin, Yingwei Ma, Rongyu Cao, Binhua Li, Fei Huang, Xiaodong Gu, and Yongbin Li. 2024. LLMs as Continuous Learners: Improving the Reproduction of Defective Code in Software Issues. arXiv:2411.13941 [cs]
  19. 19.Lei Liu, Xiaoyan Yang, Yue Shen, Binbin Hu, Zhiqiang Zhang, Jinjie Gu, and Guannan Zhang. 2023. Think-in-Memory: Recalling and Post-thinking Enable LLMs with Long-Term Memory. arXiv:2311.08719 [cs]
  20. 20.Weijie Lv, Xuan Xia, and Sheng-Jun Huang. 2024. CodeACT: Code Adaptive Compute-efficient Tuning Framework for Code LLMs. arXiv:2408.02193 [cs]
  21. 21.Yingwei Ma, Binhua Li, Yihong Dong, Xue Jiang, Rongyu Cao, Jue Chen, Fei Huang, and Yongbin Li. 2025. Thinking Longer, Not Larger: Enhancing Software Engineering Agents via Scaling Test-Time Compute. arXiv:2503.23803 [cs]
  22. 22.Yingwei Ma and Yue Liu. 2025. Improving Automated Issue Resolution via Comprehensive Repository Exploration. In ICLR 2025 Third Workshop on Deep Learning for Code.
  23. 23.Zexiong Ma, Chao Peng, Pengfei Gao, Xiangxin Meng, Yanzhen Zou, and Bing Xie. 2025. SoRFT: Issue Resolving with Subtask-oriented Reinforced Fine-Tuning. arXiv:2502.20127 [cs]
  24. 24.OpenAI. 2024. Introducing SWE-bench Verified. https://openai.com/index/introducing-swe-bench-verified/.
  25. 25.OpenAI. 2024. Memory and New Controls for ChatGPT. https://openai.com/index/memory-and-new-controls-for-chatgpt/.
  26. 26.Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. 2024. MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560 [cs]
  27. 27.Jiayi Pan, Xingyao Wang, Graham Neubig, Navdeep Jaitly, Heng Ji, Alane Suhr, and Yizhe Zhang. 2024. Training Software Engineering Agents and Verifiers with SWE-Gym.
  28. 28.Weihan Peng, Yuling Shi, Yuhang Wang, Xinyun Zhang, Beijun Shen, and Xiaodong Gu. 2025. SWE-QA: Can Language Models Answer Repository-level Code Questions? arXiv preprint arXiv:2509.14635 (2025).
  29. 29.Minh V. T. Pham, Huy N. Phan, Hoang N. Phan, Cuong Le Chi, Tien N. Nguyen, and Nghi D. Q. Bui. 2025. SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs. arXiv:2504.14757 [cs]
  30. 30.Chen Qian, Yufan Dang, Jiahao Li, Wei Liu, Zihao Xie, Yifei Wang, Weize Chen, Cheng Yang, Xin Cong, Xiaoyin Che, et al. 2023. Experiential co-learning of software-developing agents. arXiv preprint arXiv:2312.17025 (2023).
  31. 31.Maxime Robeyns, Martin Szummer, and Laurence Aitchison. 2025. A Self-Improving Coding Agent. arXiv:2504.15228 [cs]
  32. 32.Haifeng Ruan, Yuntong Zhang, and Abhik Roychoudhury. 2024. SpecRover: Code Intent Extraction via LLMs. arXiv:2408.02232 [cs]
  33. 33.Yuchen Shao, Yuheng Huang, Jiawei Shen, Lei Ma, Ting Su, and Chengcheng Wan. 2025. Are LLMs Correctly Integrated into Software Systems?. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, 1178–1190.
  34. 34.Yuling Shi, Yichun Qian, Hongyu Zhang, Beijun Shen, and Xiaodong Gu. 2025. LongCodeZip: Compress Long Context for Code Language Models. arXiv preprint arXiv:2510.00446 (2025).
  35. 35.Yuling Shi, Songsong Wang, Chengcheng Wan, and Xiaodong Gu. 2024. From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging. arXiv:2410.01215 [cs]
  36. 36.Yuling Shi, Hongyu Zhang, Chengcheng Wan, and Xiaodong Gu. 2024. Between Lines of Code: Unraveling the Distinct Patterns of Machine and Human Programmers. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE Computer Society, 51–62.
  37. 37.Xiangru Tang, Tianrui Qin, Tianhao Peng, Ziyang Zhou, Daniel Shao, Tingting Du, Xinming Wei, Peng Xia, Fang Wu, He Zhu, et al. 2025. Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving. arXiv preprint arXiv:2507.06229 (2025).
  38. 38.Wenyi Wang, Piotr Piękos, Li Nanbo, Firas Laakom, Yimeng Chen, Mateusz Ostaszewski, Mingchen Zhuge, and Jürgen Schmidhuber. 2025. Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine. arXiv:2510.21614 [cs.AI] https://arxiv.org/abs/2510.21614
  39. 39.Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brennan, Hao Peng, Heng Ji, and Graham Neubig. 2024. OpenHands: An Open Platform for AI Software Developers as Generalist Agents. arXiv:2407.16741
  40. 40.Yibo Wang, Zhihao Peng, Ying Wang, Zhao Wei, Hai Yu, and Zhiliang Zhu. 2025. MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue Resolution. arXiv:2506.12728 [cs]
  41. 41.Yuhang Wang, Yuling Shi, Mo Yang, Rongrui Zhang, Shilin He, Heng Lian, Yuting Chen, Siyu Ye, Kai Cai, and Xiaodong Gu. 2026. SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents. arXiv preprint arXiv:2601.16746 (2026).
  42. 42.Yifei Wang, Feng Xiong, Yong Wang, Linjing Li, Xiangxiang Chu, and Daniel Dajun Zeng. 2025. Position bias mitigates position bias: Mitigate position bias through inter-position knowledge distillation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 1495–1512.
  43. 43.Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. 2024. Agent workflow memory. arXiv preprint arXiv:2409.07429 (2024).
  44. 44.Yuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux, Lingming Zhang, Daniel Fried, Gabriel Synnaeve, Rishabh Singh, and Sida I. Wang. 2025. SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution. arXiv:2502.18449 [cs]
  45. 45.Rebecca Westhäußer, Frederik Berenz, Wolfgang Minker, and Sebastian Zepf. 2025. CAIM: Development and Evaluation of a Cognitive AI Memory Framework for Long-Term Interaction with Intelligent Agents. arXiv:2505.13044 [cs]
  46. 46.Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, and Chi Wang. 2023. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv:2308.08155 [cs]
  47. 47.Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang. 2024. Agentless: Demystifying LLM-based Software Engineering Agents. arXiv:2407.01489
  48. 48.Chunqiu Steven Xia, Zhe Wang, Yan Yang, Yuxiang Wei, and Lingming Zhang. 2025. Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly? arXiv preprint arXiv:2511.13646 (2025).
  49. 49.Yuanzhen Xie, Tao Xie, Mingxiong Lin, WenTao Wei, Chenglin Li, Beibei Kong, Lei Chen, Chengxiang Zhuo, Bo Hu, and Zang Li. 2023. OlaGPT: Empowering LLMs With Human-like Problem-Solving Abilities. arXiv:2305.16334 [cs]
  50. 50.John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik R. Narasimhan, and Ofir Press. 2024. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
  51. 51.John Yang, Kilian Leret, Carlos E. Jimenez, Alexander Wettig, Kabir Khandpur, Yanzhe Zhang, Binyuan Hui, Ofir Press, Ludwig Schmidt, and Diyi Yang. 2025. SWE-smith: Scaling Data for Software Engineering Agents.
  52. 52.Xiao Yu, Baolin Peng, Vineeth Vajipey, Hao Cheng, Michel Galley, Jianfeng Gao, and Zhou Yu. 2025. ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning. arXiv:2410.02052 [cs]
  53. 53.Jusheng Zhang, Yijia Fan, Wenjun Lin, Ruiqi Chen, Haoyi Jiang, Wenhao Chai, Jian Wang, and Keze Wang. 2025. GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning. arXiv preprint arXiv:2505.23399 (2025).
  54. 54.Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, and Jeff Clune. 2025. Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents. arXiv preprint arXiv:2505.22954 (2025).
  55. 55.Jusheng Zhang, Zimeng Huang, Yijia Fan, Ningyuan Liu, Mingyan Li, Zhuojie Yang, Jiawei Yao, Jian Wang, and Keze Wang. 2025. KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems. In Forty-second International Conference on Machine Learning. https://openreview.net/forum?id=AKvy9a4jho
  56. 56.Linghao Zhang, Shilin He, Chaoyun Zhang, Yu Kang, Bowen Li, Chengxing Xie, Junhao Wang, Maoquan Wang, Yufan Huang, Shengyu Fu, Elsie Nallipogu, Qingwei Lin, Yingnong Dang, Saravan Rajmohan, and Dongmei Zhang. 2025. SWE-bench Goes Live! arXiv:2505.23419 [cs]
  57. 57.Wenqi Zhang, Ke Tang, Hai Wu, Mengna Wang, Yongliang Shen, Guiyang Hou, Zeqi Tan, Peng Li, Yueting Zhuang, and Weiming Lu. 2024. Agent-pro: Learning to evolve via policy-level reflection and optimization. arXiv preprint arXiv:2402.17574 (2024).
  58. 58.Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury. 2024. AutoCodeRover: Autonomous Program Improvement. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2024). Association for Computing Machinery, New York, NY, USA, 1592–1604.
  59. 59.Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. 2024. ExpeL: LLM Agents Are Experiential Learners. Proceedings of the AAAI Conference on Artificial Intelligence 38, 17 (March 2024), 19632–19642.
  60. 60.Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. 2024. MemoryBank: Enhancing Large Language Models with Long-Term Memory. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence (AAAI'24/IAAI'24/EAAI'24, Vol. 38). AAAI Press, 19724–19731.

Citation

MLA
Chen, S., et al. “SWE-Exp: Experience-Driven Software Issue Resolution”. arXiv, 2025, http://arxiv.org/abs/2507.23361v2.
APA
Chen, S., Lin, S., Shi, Y., Lian, H., Gu, X., Yun, L., Chen, D., Cao, L., Liu, J., Xia, N., & Wang, Q. (2025). SWE-Exp: Experience-Driven Software Issue Resolution. arXiv. http://arxiv.org/abs/2507.23361v2
Chicago
Chen, S., S. Lin, Y. Shi, et al. 2025. “SWE-Exp: Experience-Driven Software Issue Resolution”. arXiv. http://arxiv.org/abs/2507.23361v2.
Harvard
Chen, S. et al. (2025) “SWE-Exp: Experience-Driven Software Issue Resolution”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2507.23361v2.
Vancouver
1. Chen S, Lin S, Shi Y, et al (2025) SWE-Exp: Experience-Driven Software Issue Resolution. arXiv

BibTeX

@article{chen2025swe,
  title = {SWE-Exp: Experience-Driven Software Issue Resolution},
  author = {Chen, Silin and Lin, Shaoxin and Shi, Yuling and Lian, Heng and Gu, Xiaodong and Yun, Longfei and Chen, Dong and Cao, Lin and Liu, Jiyang and Xia, Nu and Wang, Qianxiang},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2507.23361v2},
  eprint = {2507.23361}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/