Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning

Jiaxing QiZhongzhi LuanHongyu ZhangShaohan HuangCarol J. FungYongxin TongHailong YangDepei Qian

article2026arXiv2 citations

Demonstrates that leading large language models fail to select valid recovery actions for up to 60% of correctly diagnosed microservice failures, introducing the R2Act benchmark to evaluate post-diagnosis incident remediation.

Listen

Modern cloud-native software architectures rely heavily on microservices, where complex interdependencies make rapid incident response both critical and difficult. While artificial intelligence and large language models are increasingly deployed to analyze operational logs and diagnose system failures, successful incident resolution requires taking the correct operational remediation step rather than merely identifying what broke. Existing industry benchmarks evaluate how well models localize root causes, but they overlook whether automated systems can translate those diagnoses into valid, executable recovery actions.

The article introduces and evaluates R2Act, an evaluation framework and benchmark designed to assess diagnosis-to-action reasoning during microservice failures. The objective is to demonstrate whether automated diagnostic methods and large language models can select valid recovery operations and appropriate targets under realistic system constraints.

To evaluate this capability, the authors constructed a benchmark of 302 quality-audited incidents using a standard multi-service application deployed on a container management platform. The benchmark spans six distinct service roles and eight fault categories, incorporating synchronized multi-modal data including over 12 million log records, platform events, and system metrics. The evaluation compared heuristic, supervised, specialized root-cause analysis, and large language model techniques, testing both offline plan validity and live system recovery replay.

The findings show a significant gap between diagnosis and recovery. While top retrieval-augmented language models achieved 91.4% to 99.7% accuracy in identifying the faulty service, their recovery action validity reached only 36.8% to 60.3%. Even when models correctly diagnosed both the faulty service and the exact failure type, they still selected invalid recovery actions in 39.5% to 62.0% of cases. An analysis of failure causes revealed that 67.9% of invalid plans stemmed from selecting the wrong operational action, while 29.7% failed due to invalid plan structures. Failures were heavily concentrated in network domain name, routing, and memory issues, whereas routine service restarts and scaling were handled more effectively. In live execution tests on a representative model, only 146 of 302 predictions (48.3%) successfully met validity criteria and restored service health.

These results indicate that automated operational tools cannot be trusted to remediate systems safely based solely on diagnostic accuracy. Deploying automated self-healing systems that simply attach generic actions to root-cause reports creates operational risks, downtime, and potential service disruption. Successful incident mitigation depends on fine-grained operational semantics, such as configuration scopes, resource limits, and network dependencies, which current models fail to manage reliably.

Organizations developing or deploying automated operational agents should avoid fully autonomous remediation until systems incorporate explicit action-space modeling and policy constraints. Engineering teams should implement validation gates that verify operational preconditions before actions execute live. Future technical efforts must focus on providing models with deeper operational context, configuration awareness, and live execution validation rather than solely optimizing diagnostic accuracy.

While the findings provide high confidence regarding the evaluated failure modes, the study is limited to a single multi-service architecture within a controlled test environment. Reader caution is warranted when generalizing these results to diverse enterprise production environments with custom architectures, larger service graphs, or distinct organizational response policies.

Cover for Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning

Abstract

Large language models (LLMs) are increasingly used to interpret operational evidence and assist incident response in cloud-native microservice systems. However, recovery-oriented use cases require more than identifying a root cause. After observing symptoms and diagnosing a fault, an operator or agent must translate the diagnosis into a concrete recovery action, apply it to an admissible target, and verify that service health has been restored. Existing RCA and log-analysis evaluations are well-suited to diagnosis, but they do not characterize this subsequent action decision. This paper presents R2Act, a recovery-action evaluation framework for post-diagnosis incident response. R2Act defines an incident schema, quality gate, action-space representation, recovery-validity metrics, offline evaluator, and live-replay protocol. We instantiate the framework as a benchmark dataset of 302 quality-audited Kubernetes incidents from \system. Each incident provides synchronized multi-modal observations, root-cause labels, an incident-specific action space, and annotated valid and invalid recovery plans. We evaluate heuristic, supervised, RCA-oriented, deep log, and LLM-based methods. The strongest RAG-based LLMs reach 91.4%--99.7% root-cause service accuracy, yet their recovery validity remains only 36.8%--60.3%. Even when both the root-cause service and fault type are correct, recovery-oriented methods still choose invalid actions for 39.5%--62.0% of correctly diagnosed incidents. Overall, this work reveals that many recovery failures arise not from missing diagnostic knowledge, but from the difficulty of translating diagnostic evidence into valid recovery actions and admissible targets. This work provides a reproducible, simplified starting point for research and evaluation.

Table of Contents

  • I Introduction
  • II Related Work and Motivation
  • II-A Related Work
  • II-B Motivation
  • III Overview
  • III-A Benchmark Design
  • III-B Construction and Quality Audit
  • III-C Benchmark Records and Recovery Semantics
  • IV Dataset Characterization
  • IV-A Dataset Scope and Quality Audit
  • IV-B Service and Fault-Type Composition
  • IV-C Recovery Labels, Observation Scale, and Cell Coverage
  • V Evaluation Design
  • VI Results
  • VI-A RQ1: Diagnosis vs. Recovery-Action Validity
  • VI-B RQ2: Modality Effects on Diagnosis and Recovery
  • VI-C RQ3: Remaining Recovery Errors After Correct Diagnosis
  • VI-D RQ4: Error Sources in Recovery-Action Selection
  • VI-E RQ5: Live Replay Check
  • VII Discussion
  • VIII Conclusion
  • References

Knowls

  1. Knowl 1 — Diagnosis-to-Action Gap in Microservice Incident Response

    empirical result

    Across heuristic baselines, supervised classifiers, root cause analysis (RCA) techniques, deep log models, and large language models (LLMs), diagnostic correctness does not imply recovery-action validity. While top retrieval-augmented generation (RAG) LLMs achieve root-cause service accuracy between 91.4%91.4\% and 99.7%99.7\% and root-cause fault type accuracy up to 100.0%100.0\%, their recovery validity ranges from only 36.8%36.8\% to 60.3%60.3\%.

    Method RCA-S ↑\uparrow RCA-T ↑\uparrow Action ↑\uparrow Target ↑\uparrow Exact ↑\uparrow Valid ↑\uparrow No-op ↓\downarrow
    Recovery-oriented methods
    LogPrompt 0.603 0.162 0.228 0.513 0.000 0.109 0.470
    LogRAG 0.616 0.142 0.228 0.546 0.000 0.123 0.447
    OpenRCA-D 0.563 0.175 0.179 0.464 0.089 0.142 0.510
    Qwen-RAG 0.990 0.957 0.719 0.709 0.258 0.483 0.119
    Kimi-RAG 0.990 1.000 0.566 0.712 0.245 0.517 0.116
    MiniMax-RAG 0.983 0.937 0.434 0.705 0.195 0.368 0.222
    DeepSeek-RAG 0.914 0.934 0.838 0.666 0.222 0.530 0.103
    GLM-ZS 0.997 0.993 0.321 0.712 0.073 0.315 0.288
    GLM-RAG 0.997 0.997 0.722 0.709 0.242 0.603 0.043
    RCA-only + fixed mapper
    Keyword 0.116 0.063 0.149 0.103 0.026 0.089 0.775
    Rule 0.338 0.228 0.298 0.278 0.139 0.222 0.642
    NB 0.679 0.536 0.639 0.444 0.318 0.467 0.281
    Epsilon 0.149 0.053 0.166 0.096 0.013 0.013 0.752
    RCD 0.149 0.053 0.166 0.096 0.013 0.013 0.752
    Random Walk 0.149 0.053 0.166 0.096 0.013 0.013 0.752
    BARO 0.136 0.122 0.291 0.000 0.026 0.281 0.709
    CIRCA 0.149 0.053 0.166 0.096 0.013 0.013 0.752
    LogFormer 0.381 0.116 0.186 0.106 0.086 0.182 0.755
    OneLog 0.533 0.086 0.169 0.166 0.109 0.152 0.732

    In this benchmark evaluation of 302 quality-audited incidents:

    • RCA-S (Root-Cause Service Accuracy) measures the fraction of incidents where the predicted faulty service matches ground truth.
    • RCA-T (Root-Cause Fault Type Accuracy) measures the fraction where the predicted fault mechanism matches ground truth.
    • Action Hit checks whether the predicted operation matches a valid operation.
    • Target Hit checks whether the predicted target service matches a valid target service.
    • Exact Match requires both predicted operation and target service to strictly match the gold recovery plan.
    • Valid (Recovery Validity) measures whether the predicted recovery plan belongs to the set of incident-specific valid recovery plans.
    • No-op indicates the fraction of instances where no executable recovery action is returned.

    Methods frequently identify the correct target service while failing to propose a valid operation or admissible plan parameters.

  2. Knowl 2 — R2Act Incident and Multi-Modal Observation Representation

    definition

    In the R2Act framework for evaluating post-diagnosis microservice failure recovery, each incident xix_i is formalized as a four-tuple:

    xi=(Oi,Mi,Yi,Ai)x_i = (O_i, M_i, Y_i, A_i)

    where:

    • OiO_i is the synchronized multi-modal observation window captured across pre-fault, during-fault, and post-fault phases.
    • MiM_i denotes campaign metadata (e.g., target application, fault intensity, load profile, injection duration, repetition ID).
    • YiY_i contains ground-truth diagnostic labels (root-cause service and root-cause fault type) and valid/invalid recovery action labels.
    • AiA_i represents the incident-specific allowed recovery action space defining admissible operations, target services, and required fields under current cluster constraints.

    The observation component OiO_i is decomposed as:

    Oi={Oilog,Oievent,Oimetric,Oistate,Oichaos}O_i = \{O^{log}_i, O^{event}_i, O^{metric}_i, O^{state}_i, O^{chaos}_i\}

    where OilogO^{log}_i contains service logs from Kubernetes pods, OieventO^{event}_i contains Kubernetes cluster events, OimetricO^{metric}_i denotes Prometheus metric time series (resource- and service-level indicators), OistateO^{state}_i captures pod and deployment specifications/statuses, and OichaosO^{chaos}_i records Chaos Mesh fault injection status.

  3. Knowl 3 — Incident Quality Gate Criterion

    equation

    To prevent incomplete or corrupted observations from entering benchmark evaluations, R2Act applies a quality gate qi∈{0,1}q_i \in \{0, 1\} to every collected incident run:

    qi=∏m∈MI[Oim≠∅]⋅I[Yi is valid]⋅I[Ai is valid]q_i = \prod_{m \in \mathcal{M}} \mathbb{I}[O_i^m \neq \emptyset] \cdot \mathbb{I}[Y_i \text{ is valid}] \cdot \mathbb{I}[A_i \text{ is valid}]

    where:

    • M={log,event,metric,state,chaos}\mathcal{M} = \{log, event, metric, state, chaos\} is the set of required observation modalities.
    • I[⋅]\mathbb{I}[\cdot] is the indicator function evaluating to 1 if the condition is satisfied and 0 otherwise.
    • OimO_i^m is the observation data for modality mm across pre-fault, during-fault, and post-fault snapshots.
    • YiY_i is valid if the incident annotation JSON contains complete root-cause service, fault type, intensity, and recovery labels.
    • AiA_i is valid if the incident-specific recovery action space can be constructed.

    Only runs with qi=1q_i = 1 are admitted into the primary evaluation dataset (302 incidents), while runs with qi=0q_i = 0 (71 runs) are relegated to the audit trail as warning, failed, or legacy pilot runs.

  4. Knowl 4 — Recovery Plan Validity Formulation

    definition

    In microservice failure recovery, recovery plans are represented as typed operation-target decisions. A recovery plan p^i\hat{p}_i specifies an operation type (e.g., restart service, scale out, roll back configuration, repair DNS), a target service, and operation-specific parameters (such as configuration scope, route rule, or memory limit).

    The validity of a predicted recovery plan p^i\hat{p}_i is formally evaluated against the incident-specific action space AiA_i, the set of valid gold recovery plans Pi+P^+_i, and the set of invalid recovery plans Pi−P^-_i:

    valid(p^i)=I[p^i∈Pi+∧p^i∉Pi−∧p^i∈Ai]\text{valid}(\hat{p}_i) = \mathbb{I}[\hat{p}_i \in P^+_i \wedge \hat{p}_i \notin P^-_i \wedge \hat{p}_i \in A_i]

    where:

    • Pi+P^+_i is a set of semantically equivalent valid recovery plans that restore health without violating operational constraints.
    • Pi−P^-_i is a set of known invalid plans (e.g., incorrect operations, invalid dependency targets, out-of-scope configurations, or no-op actions).
    • AiA_i is the incident-specific admissible action space enumerated under the injected fault and live deployment topology.
  5. Knowl 5 — Post-Diagnosis Invalid-Action Rate ($E_C$)

    equation

    To evaluate recovery planning independent of diagnostic errors, performance is conditioned on incidents where root-cause analysis (RCA) is fully correct.

    Let CC denote the event that both the root-cause service and fault type are correctly identified (C=[s^i=si∗∧t^i=ti∗]C = [\hat{s}_i = s_i^* \wedge \hat{t}_i = t_i^*]), and let VV denote the event that the predicted recovery plan is valid (V=[valid(p^i)=1]V = [\text{valid}(\hat{p}_i) = 1]). Conditional recovery validity VCV_C and the invalid-action rate after correct diagnosis ECE_C are defined as:

    VC=Pr⁡(V∣C)V_C = \Pr(V \mid C)

    EC=Pr⁡(¬V∣C)=1−VCE_C = \Pr(\neg V \mid C) = 1 - V_C

    ECE_C isolates the post-diagnosis action-decision residual error: it quantifies the probability that a method fails to select a valid recovery operation, target, or parameter given an accurate diagnostic foundation.

  6. Knowl 6 — Recovery Failure Persistence After Correct Diagnosis

    empirical result

    Conditioning recovery action evaluation on correct diagnostic identification (CC) demonstrates that accurate RCA is insufficient for successful recovery planning. Even when both the root-cause service and fault type are correct, zero-shot (ZS) and few-shot (FS) LLMs exhibit invalid-action rates ECE_C between 0.5910.591 and 0.7530.753, while RAG-based LLMs exhibit ECE_C between 0.3950.395 and 0.6200.620.

    Method Pr⁡(C)\Pr(C) VCV_C ECE_C Method Pr⁡(C)\Pr(C) VCV_C ECE_C
    Keyword 0.000 – – MiniMax-RAG 0.924 0.380 0.620
    Rule-L 0.010 1.000 0.000 DeepSeek-ZS 0.805 0.337 0.663
    Rule-LE 0.126 1.000 0.000 DeepSeek-FS 0.728 0.409 0.591
    Rule-LEM 0.212 0.984 0.016 DeepSeek-RAG 0.897 0.565 0.435
    LogPrompt 0.149 0.511 0.489 GLM-ZS 0.990 0.318 0.682
    LogRAG 0.132 0.675 0.325 GLM-FS 0.944 0.354 0.646
    OpenRCA-D 0.149 0.778 0.222 GLM-RAG 0.997 0.605 0.395
    OpenRCA-C 0.000 – – NB 0.434 0.832 0.168
    Qwen-ZS 0.891 0.335 0.665 LogFormer 0.056 1.000 0.000
    Qwen-FS 0.758 0.354 0.646 OneLog 0.060 1.000 0.000
    Qwen-RAG 0.947 0.500 0.500 Diagnostic control
    Kimi-ZS 0.993 0.320 0.680 Gold RCA + Mapper 1.000 0.599 0.401
    Kimi-FS 0.907 0.380 0.620
    Kimi-RAG 0.990 0.522 0.478
    MiniMax-ZS 0.897 0.247 0.753
    MiniMax-FS 0.825 0.285 0.715

    Applying a fixed deterministic mapping policy to perfect diagnosis (Gold RCA + Mapper) yields VC=0.599V_C = 0.599 and EC=0.401E_C = 0.401, demonstrating that coarse service and fault-type labels do not capture operational parameters (e.g., dependency endpoints, routing rules, or resource limits).

  7. Knowl 7 — Deterministic RCA-to-Action Mapping Policy

    model/method

    To evaluate the downstream recoverability of RCA-only methods (which output only predicted root-cause service s^i\hat{s}_i and fault type t^i\hat{t}_i), R2Act employs a deterministic mapping rule p^i=g(s^i,t^i)\hat{p}_i = g(\hat{s}_i, \hat{t}_i) to construct structured recovery plans within the incident-specific action space AiA_i:

    1. Pod Failure & Network Delay: Triggers a Restart service operation targeted at s^i\hat{s}_i.
    2. CPU Saturation: Triggers a Scale out operation specifying target service s^i\hat{s}_i and replica counts.
    3. Memory Pressure: Triggers an Increase memory limit operation specifying target service s^i\hat{s}_i and memory resource limits.
    4. HTTP Error & HTTP Abort: Triggers a Roll back configuration operation specifying target service s^i\hat{s}_i and the route or configuration scope.
    5. DNS Fault & DNS Randomization: Triggers a Repair DNS/dependency operation targeting the affected service dependency endpoint.

    Required parameter fields (e.g., dependency endpoints or configuration scopes) are populated from admissible candidates within AiA_i. The mapper establishes a baseline control quantifying how far coarse diagnosis can go without fine-grained action reasoning.

  8. Knowl 8 — Taxonomy and Sources of Recovery-Action Selection Errors

    empirical result

    An analysis of 755 invalid recovery predictions produced by RAG-based LLM backbones categorizes the failure mechanisms into four distinct classes:

    Error Type Count Ratio Interpretation
    Wrong operation 513 67.9% The method selects an operation invalid for the incident.
    Invalid plan structure 224 29.7% The operation is plausible, but fails plan semantics/fields.
    Wrong target/dependency 10 1.3% The operation is correct, but target service/dependency is invalid.
    No-op or empty plan 8 1.1% The method produces no executable recovery action.

    Recovery difficulty is heavily concentrated in specific fault mechanisms:

    • DNS Faults (DNS-F, DNS-R): LLMs achieve 0.880.88--1.001.00 RCA service accuracy, but 0.000.00 recovery validity across all models because they prescribe service restarts instead of repairing dependency-resolution paths.
    • HTTP Faults (HTTP-E, HTTP-A): RCA service accuracy is 0.820.82--1.001.00, while recovery validity fluctuates between 0.150.15 and 1.001.00 due to errors in identifying routing and configuration scopes.
    • Memory Pressure: RCA service accuracy reaches 0.970.97--1.001.00, but recovery validity remains low (0.290.29--0.650.65) due to failures in specifying resource-limit modification parameters.
    • Pod Failure, CPU Saturation, Network Delay: Attain high recovery validity (0.920.92--1.001.00) because their mitigations map directly to standard single-action routines (service restart or scale-out).
  9. Knowl 9 — Telemetry Modality Ablation on Recovery Validity

    empirical result

    Ablating observation modalities across (1) Logs only, (2) Logs + Events, and (3) Logs + Events + Metrics demonstrates that richer multi-modal telemetry improves diagnostic accuracy and recovery validity, but cannot eliminate the diagnosis-to-action gap.

    Logs Logs+Events Logs+Events+Metrics
    Method RCA-S RCA-T Recov. No-op RCA-S RCA-T Recov. No-op RCA-S RCA-T Recov. No-op
    Rule 3.3 10.3 12.9 72.5 34.1 16.6 23.5 65.5 33.8 22.9 22.2 64.2
    NB 17.9 16.2 6.6 61.9 64.9 42.0 32.4 30.4 68.6 52.7 46.0 28.8
    LogRAG 12.6 11.9 2.6 65.5 59.6 13.9 10.3 46.0 59.3 13.9 11.9 46.0
    LogPrompt 15.6 11.6 2.3 68.9 59.6 13.3 8.3 47.0 61.6 16.3 12.3 46.0
    OpenRCA-D 12.6 10.9 4.3 75.5 52.7 13.9 11.3 55.6 55.6 15.2 12.2 52.6
    Qwen-RAG 93.1 93.4 32.1 24.8 99.3 96.7 43.4 17.6 99.0 96.4 48.0 11.9
    Kimi-RAG 97.4 100.0 54.6 9.9 99.3 100.0 56.0 8.3 99.0 100.0 51.0 12.6
    MiniMax-RAG 71.8 69.8 23.9 43.5 98.3 96.7 32.8 25.5 97.7 92.4 36.1 22.9
    DeepSeek-RAG 85.7 97.3 54.5 8.3 94.7 96.7 54.3 9.3 94.1 96.7 54.7 7.0
    GLM-RAG 99.7 100.0 61.9 1.0 100.0 100.0 60.6 3.3 100.0 100.0 61.9 2.0

    All metrics are reported as 5-fold cross-validation means in percentage. While adding Kubernetes events and Prometheus metrics raises Naive Bayes (NB) recovery from 6.6%6.6\% to 46.0%46.0\% and reduces no-op rates, specialized LLM methods like LogPrompt and LogRAG remain near 12%12\% recovery validity despite attaining >59%>59\% RCA-Service accuracy.

  10. Knowl 10 — Validity-Gated Live Execution Replay Outcomes

    empirical result

    Live execution of model-generated recovery plans on a Kubernetes testbed demonstrates consistent alignment between offline validity and live service restoration. Replaying all 302 benchmark incidents using Qwen-RAG predictions under live Chaos Mesh fault injections yields:

    Fault type Plan RV-valid Rec. Replay RV
    CPU saturation 36 36 36 100.0%
    DNS fault 48 0 0 0.0%
    DNS random 48 0 0 0.0%
    HTTP abort 39 23 23 59.0%
    HTTP error 48 27 27 56.2%
    Memory pressure 34 11 11 32.4%
    Network delay 13 13 13 100.0%
    Pod failure 36 36 36 100.0%
    Total 302 146 146 48.3%

    In this evaluation:

    • Plan: Total planned test attempts (302302).
    • RV-valid: Model predictions satisfying offline recovery-validity constraints (146146).
    • Rec.: Number of offline-valid predictions that successfully restore cluster health within a 75-second post-recovery observation window (146146).
    • Replay RV: Validity-gated replay recovery rate (146/302=48.3%146 / 302 = 48.3\%).

    Every plan that satisfied the offline validity criteria successfully restored health upon execution, confirming that recovery failures stem from incomplete or invalid operation-target planning rather than executor or environment faults.

  11. Knowl 11 — R2Act Benchmark Dataset Scope and Composition

    experimental setup

    The R2Act benchmark dataset consists of 302 quality-audited failure incidents collected on a Kubernetes deployment of the Google Online Boutique microservices application (12.6M total log records, averaging 41,736 logs per incident):

    • Service Roles (6 services, 44–55 incidents each): frontend (entry point, 14.6%), checkoutservice (orchestration, 17.2%), recommendationservice (computation, 15.6%), cartservice (state-related, 16.6%), productcatalogservice (catalog read, 17.9%), paymentservice (transaction, 18.2%).
    • Fault Categories (8 mechanisms across moderate and severe intensities):
      1. Kubernetes Lifecycle & Resource: Pod failure (11.9%, 36 incidents), CPU saturation (11.9%, 36 incidents), Memory pressure (11.3%, 34 incidents).
      2. Network & Application-Layer: Network delay (4.3%, 13 incidents), HTTP error (15.9%, 48 incidents), HTTP abort (12.9%, 39 incidents).
      3. Dependency Resolution: DNS fault (15.9%, 48 incidents), DNS randomization (15.9%, 48 incidents).
    • Service-Fault Cells: 44 non-empty service-fault combinations with 3 to 8 incidents per cell.
    • Action Labels: 7 candidate recovery operations in the action space, 5 gold operations utilized in valid recovery plans (96 dependency restarts, 87 config rollbacks, 61 service restarts, 36 scale-outs, 22 memory limit increases), and 604 annotated invalid plans (2 per incident).

Coverage note — None was omitted. All contributed formulations, benchmark specifications, empirical evaluations (RQ1–RQ5), error taxonomies, and live replay experiments are represented.

References

  1. 1.K. Sarda, Z. Namrud, M. Litoiu, L. Shwartz, and I. Watts, —Leveraging large language models for the auto-remediation of microservice applications: An experimental study,— in Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, 2024, pp. 163–174.
  2. 2.Z. Zhong, R. Fu, M. Ma, S. Zhang, Y. Sun, C. Bansal, and D. Pei, —Llm-enhanced failure localization in microservices: Integrating multi-modal data and expert interpretation,— IEEE Transactions on Services Computing, pp. 1–14, 2026.
  3. 3.X. Zhou, X. Peng, T. Xie, J. Sun, C. Xu, C. Ji, and W. Zhao, —Benchmarking microservice systems for software engineering research,— in Proceedings of the 40th International Conference on Software Engineering: Companion Proceeedings, 2018, pp. 323–324.
  4. 4.D. Liu, C. He, X. Peng, F. Lin, C. Zhang, S. Gong, Z. Li, J. Ou, and Z. Wu, —Microhecl: High-efficient root cause localization in large-scale microservice systems,— in 2021 IEEE/ACM 43rd International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 2021, pp. 338–347.
  5. 5.Y. Sun, Z. Lin, B. Shi, S. Zhang, S. Ma, P. Jin, Z. Zhong, L. Pan, Y. Guo, and D. Pei, —Interpretable failure localization for microservice systems based on graph autoencoder,— ACM Transactions on Software Engineering and Methodology, vol. 34, no. 2, pp. 1–28, 2025.
  6. 6.J. Zhu, S. He, P. He, J. Liu, and M. R. Lyu, —Loghub: A large collection of system log datasets for ai-driven log analytics,— in 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 2023, pp. 355–366.
  7. 7.Z. Jiang, J. Liu, J. Huang, Y. Li, Y. Huo, J. Gu, Z. Chen, J. Zhu, and M. R. Lyu, —A large-scale evaluation for log parsing techniques: How far are we?— in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, 2024, pp. 223–234.
  8. 8.T. Cui, S. Ma, Z. Chen, T. Xiao, C. Zhao, S. Tao, Y. Liu, S. Zhang, D. Lin, C. Liu et al., —Logeval: A comprehensive benchmark suite for llms in log analysis,— Empirical Software Engineering, vol. 30, no. 6, p. 173, 2025.
  9. 9.Y. Wang, Z. Zhu, Q. Fu, Y. Ma, and P. He, —Mrca: Metric-level root cause analysis for microservices via multi-modal data,— in Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, ser. ASE '24. New York, NY, USA: Association for Computing Machinery, 2024, p. 1057–1068. [Online]. Available: https://doi.org/10.1145/3691620.3695485
  10. 10.L. Pham, H. Zhang, H. Ha, F. Salim, and X. Zhang, —Rcaeval: a benchmark for root cause analysis of microservice systems with telemetry data,— in Companion Proceedings of the ACM on Web Conference 2025, 2025, pp. 777–780.
  11. 11.W. Xu, J. Luo, T. Huang, K. Sui, J. Geng, Q. Ma, I. Akasaka, X. Shi, J. Tang, and P. Cai, —Logsage: An llm-based framework for ci/cd failure detection and remediation with industrial validation,— in 2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2025, pp. 3742–3753.
  12. 12.M. E. Barnes, T. A. Ghaleb, and S. Hassan, —Logsieve: Task-aware ci log reduction for sustainable llm-based analysis,— arXiv preprint arXiv:2601.20148, 2026.
  13. 13.J. Xu, Q. Zhang, Z. Zhong, S. He, C. Zhang, Q. Lin, D. Pei, P. He, D. Zhang, and Q. Zhang, —Openrca: Can large language models locate the root cause of software failures?— in The thirteenth international conference on learning representations, 2025.
  14. 14.Z. Wang, Z. Liu, Y. Zhang, A. Zhong, J. Wang, F. Yin, L. Fan, L. Wu, and Q. Wen, —Rcagent: Cloud root cause analysis by autonomous agents with tool-augmented large language models,— arXiv preprint arXiv:2310.16340, 2023.
  15. 15.Y. Chen, H. Xie, M. Ma, Y. Kang, X. Gao, L. Shi, Y. Cao, X. Gao, H. Fan, M. Wen et al., —Automatic root cause analysis via large language models for cloud incidents,— in Proceedings of the Nineteenth European Conference on Computer Systems, 2024, pp. 674–688.
  16. 16.A. Bucchiarone, C. Guidi, I. Lanese, N. Bencomo, and J. Spillner, —A mape-k approach to autonomic microservices,— in 2022 IEEE 19th International Conference on Software Architecture Companion (ICSA-C). IEEE, 2022, pp. 100–103.
  17. 17.L. Zhang, Y. Zhai, T. Jia, C. Duan, M. He, L. Pan, Z. Liu, B. Ding, and Y. Li, —Microremed: Benchmarking llms in microservices remediation,— arXiv preprint arXiv:2511.01166, 2025.
  18. 18.C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan, —Swe-bench: Can language models resolve real-world github issues?— in International Conference on Learning Representations, vol. 2024, 2024, pp. 54 107–54 157.
  19. 19.M. H. M. Bhuiyan, A. S. Parthasarathy, N. Vasilakis, M. Pradel, and C.-A. Staicu, —Secbench. js: An executable security benchmark suite for server-side javascript,— in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 2023, pp. 1059–1070.
  20. 20.L. Pham, H. Ha, and H. Zhang, —Root cause analysis for microservice system based on causal inference: How far are we?— in Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, 2024, pp. 706–715.
  21. 21.C. Liu, W. Yang, H. Mittal, M. Singh, D. Sahoo, and S. C. Hoi, —Pyrca: A library for metric-based root cause analysis,— arXiv preprint arXiv:2306.11417, 2023.
  22. 22.Y. Liu, S. Tao, W. Meng, F. Yao, X. Zhao, and H. Yang, —Logprompt: Prompt engineering towards zero-shot and interpretable log analysis,— in Proceedings of the 2024 IEEE/ACM 46th international conference on software engineering: Companion proceedings, 2024, pp. 364–365.
  23. 23.W. Zhang, Q. Zhang, E. Yu, Y. Ren, Y. Meng, M. Qiu, and J. Wang, —Leveraging rag-enhanced large language model for semi-supervised log anomaly detection,— in 2024 IEEE 35th International Symposium on Software Reliability Engineering (ISSRE). IEEE, 2024, pp. 168–179.
  24. 24.Y. Gao, Z. Cai, and B. Yang, —Rcaflow: A workflow-informed hierarchical planning multi-agent system for root cause analysis,— in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 1, 2026, pp. 300–308.
  25. 25.R. Chen, Y. Pu, J. Xin, J. Wang, X. Liao, K. Zhang, and W. Wu, —Grace: A strategic llm-enhanced graph reinforcement learning framework for adaptive fault recovery in microservice systems,— in Service-Oriented Computing: 23rd International Conference, ICSOC 2025, Shenzhen, China, December 1–4, 2025, Proceedings, Part I. Berlin, Heidelberg: Springer-Verlag, 2025, p. 155–170. [Online]. Available: https://doi.org/10.1007/978-981-95-5012-8_12
  26. 26.T. Ahmed, S. Ghosh, C. Bansal, T. Zimmermann, X. Zhang, and S. Rajmohan, —Recommending root-cause and mitigation steps for cloud incidents using large language models,— in Proceedings of the 45th IEEE/ACM International Conference on Software Engineering, 2023, pp. 1737–1749.
  27. 27.E. Malul, Y. Meidan, D. Mimran, Y. Elovici, and A. Shabtai, —Genkubesec: Llm-based kubernetes misconfiguration detection, localization, reasoning, and remediation,— arXiv preprint arXiv:2405.19954, 2024.
  28. 28.W. Zhang, Z. Yang, F. Peng, L. Zhang, Y. Chen, and R. Chen, —Galr: Graph-based root cause localization and llm-assisted recovery for microservice systems,— Electronics, vol. 15, no. 1, p. 243, 2026.
  29. 29.H. Guo, J. Yang, J. Liu, J. Bai, B. Wang, Z. Li, T. Zheng, B. Zhang, J. Peng, and Q. Tian, —Logformer: A pre-train and tuning pipeline for log anomaly detection,— in Proceedings of the AAAI conference on artificial intelligence, vol. 38, no. 1, 2024, pp. 135–143.
  30. 30.Google Cloud, —Online boutique,— https://github.com/GoogleCloudPlatform/microservices-demo, 2026, accessed 2026-05-17.
  31. 31.The Kubernetes Authors, —Kubernetes documentation,— https://kubernetes.io/docs/, 2026, accessed 2026-05-17.
  32. 32.Prometheus Authors, —Prometheus monitoring system,— https://prometheus.io/docs/, 2026, accessed 2026-05-17.
  33. 33.Chaos Mesh Authors, —Chaos mesh documentation,— https://chaos-mesh.org/docs/, 2026, accessed 2026-05-17.
  34. 34.S. Hashemi and M. Mantyla, —Onelog: towards end-to-end software log anomaly detection,— Automated Software Engineering, vol. 31, no. 2, p. 37, 2024.

Citation

MLA
Qi, J., et al. “Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning”. arXiv, 2026, http://arxiv.org/abs/2607.04623v1.
APA
Qi, J., Luan, Z., Zhang, H., Huang, S., Fung, C., Tong, Y., Yang, H., & Qian, D. (2026). Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning. arXiv. http://arxiv.org/abs/2607.04623v1
Chicago
Qi, J., Z. Luan, H. Zhang, et al. 2026. “Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning”. arXiv. http://arxiv.org/abs/2607.04623v1.
Harvard
Qi, J. et al. (2026) “Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2607.04623v1.
Vancouver
1. Qi J, Luan Z, Zhang H, Huang S, Fung C, Tong Y, Yang H, Qian D (2026) Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning. arXiv

BibTeX

@article{qi2026can,
  title = {Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning},
  author = {Qi, Jiaxing and Luan, Zhongzhi and Zhang, Hongyu and Huang, Shaohan and Fung, Carol and Tong, Yongxin and Yang, Hailong and Qian, Depei},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2607.04623v1},
  eprint = {2607.04623}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/