Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning
Jiaxing QiZhongzhi LuanHongyu ZhangShaohan HuangCarol J. FungYongxin TongHailong YangDepei Qian
Demonstrates that leading large language models fail to select valid recovery actions for up to 60% of correctly diagnosed microservice failures, introducing the R2Act benchmark to evaluate post-diagnosis incident remediation.
Modern cloud-native software architectures rely heavily on microservices, where complex interdependencies make rapid incident response both critical and difficult. While artificial intelligence and large language models are increasingly deployed to analyze operational logs and diagnose system failures, successful incident resolution requires taking the correct operational remediation step rather than merely identifying what broke. Existing industry benchmarks evaluate how well models localize root causes, but they overlook whether automated systems can translate those diagnoses into valid, executable recovery actions.
The article introduces and evaluates R2Act, an evaluation framework and benchmark designed to assess diagnosis-to-action reasoning during microservice failures. The objective is to demonstrate whether automated diagnostic methods and large language models can select valid recovery operations and appropriate targets under realistic system constraints.
To evaluate this capability, the authors constructed a benchmark of 302 quality-audited incidents using a standard multi-service application deployed on a container management platform. The benchmark spans six distinct service roles and eight fault categories, incorporating synchronized multi-modal data including over 12 million log records, platform events, and system metrics. The evaluation compared heuristic, supervised, specialized root-cause analysis, and large language model techniques, testing both offline plan validity and live system recovery replay.
The findings show a significant gap between diagnosis and recovery. While top retrieval-augmented language models achieved 91.4% to 99.7% accuracy in identifying the faulty service, their recovery action validity reached only 36.8% to 60.3%. Even when models correctly diagnosed both the faulty service and the exact failure type, they still selected invalid recovery actions in 39.5% to 62.0% of cases. An analysis of failure causes revealed that 67.9% of invalid plans stemmed from selecting the wrong operational action, while 29.7% failed due to invalid plan structures. Failures were heavily concentrated in network domain name, routing, and memory issues, whereas routine service restarts and scaling were handled more effectively. In live execution tests on a representative model, only 146 of 302 predictions (48.3%) successfully met validity criteria and restored service health.
These results indicate that automated operational tools cannot be trusted to remediate systems safely based solely on diagnostic accuracy. Deploying automated self-healing systems that simply attach generic actions to root-cause reports creates operational risks, downtime, and potential service disruption. Successful incident mitigation depends on fine-grained operational semantics, such as configuration scopes, resource limits, and network dependencies, which current models fail to manage reliably.
Organizations developing or deploying automated operational agents should avoid fully autonomous remediation until systems incorporate explicit action-space modeling and policy constraints. Engineering teams should implement validation gates that verify operational preconditions before actions execute live. Future technical efforts must focus on providing models with deeper operational context, configuration awareness, and live execution validation rather than solely optimizing diagnostic accuracy.
While the findings provide high confidence regarding the evaluated failure modes, the study is limited to a single multi-service architecture within a controlled test environment. Reader caution is warranted when generalizing these results to diverse enterprise production environments with custom architectures, larger service graphs, or distinct organizational response policies.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). This foundational work establishes the ReAct paradigm for interleaving LLM reasoning traces with concrete environment actions, which the source builds upon to evaluate whether models can translate microservice diagnoses into valid recovery executions.
- Paper: Causal structure-based root cause analysis of outliers, Kailash Budhathoki et al. (2022). This paper establishes causal structure modeling for root cause analysis in cloud environments, providing the foundational diagnostic principles that the source demonstrates are necessary but insufficient for executing valid recovery actions.
- Paper: Detecting large-scale system problems by mining console logs, Wei Xu et al. (2009). This classic work defines telemetry and console log mining for automated system fault detection, providing the baseline operational log analysis foundations evaluated in the source's benchmark.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). This survey provides the structural architectural principles of LLM-based autonomous planning and action modules that the source directly stress-tests in cloud-native microservice recovery settings.
- Paper: Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement, Weimin Xiong et al. (2024). This study introduces step-level process supervision to prevent LLM agents from failing during multi-step interactive executions, addressing the intermediate action failures evaluated by the source.
- Paper: RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs, Yue Yu et al. (2024). This paper formalizes context reranking and retrieval-augmented generation architectures, which represent the exact class of top-performing diagnostic models evaluated in the source.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). This survey provides a comprehensive taxonomy of agentic planning, tool execution, and dynamic environment interactions, offering broader design paradigms to address the diagnosis-to-action reasoning gaps highlighted by the source.
- Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). This paper presents a diagnostic guardrail framework to trace root causes and prevent unsafe or inadmissible action executions along an AI agent's trajectory, directly addressing the invalid recovery action problem exposed in the source.
- Paper: AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation, Priyam Sahoo et al. (2026). This work analyzes process-level trajectory quality versus superficial outcomes in autonomous agents, extending the source's focus on separating diagnostic accuracy from admissible execution validity.
- Paper: daVinci-Dev: Agent-native Mid-training for Software Engineering, Ji Zeng et al. (2026). This work explores agent-native mid-training in simulated execution environments to help models better map problem diagnosis to valid, multi-step actions.
