Intent-based System Design and Operation
Vaastav AnandYichen LiAlok KumbhareCeline IrveneChetan BansalGagan SomashekarJonathan MacePedro Las-CasasRicardo BianchiniRodrigo Fonseca
Proposes intent as a foundational abstraction that encodes high-level functional and operational goals to automate cloud system design, implementation, runtime management, and evolution.
Modern cloud computing relies heavily on distributed microservice architectures that require extensive, continuous manual effort to design, deploy, operate, and maintain. As these systems grow in scale and complexity, human operators struggle to track billions of runtime traces, petabytes of operational logs, and dynamic environment changes. The article proposes an overarching framework for intent-based system design and operation, demonstrating how combining standardized cloud tooling with large language models can automate the entire software lifecycle and produce self-managing cloud services.
The authors conceptualize a human-in-the-loop architectural vision centered on "intent," which encapsulates high-level functional features, operational service-level agreements, and refinement goals. By evaluating case studies across microservice generation, real-time observability, automated troubleshooting, and dynamic resilience, the article synthesizes existing tools—such as Cerulean for hierarchical system generation and Llexus for automated runbook execution—into a unified operational blueprint.
The article delivers several key findings. First, adopting intent as a primary abstraction bridges high-level stakeholder requirements and low-level code generation, significantly reducing manual developer effort. Second, hierarchical generation enables large language models to construct both microservice business logic and corresponding end-to-end test suites directly from user specifications. Third, unifying runtime observability data (metrics, logs, and traces) with static domain knowledge (code and documentation) into event graphs provides real-time context awareness, preventing costly and inefficient data processing. Fourth, dynamic operational models can automate incident mitigation and reproduce complex runtime failures, such as metastable bottlenecks, allowing systems to autonomously generate candidate code fixes.
These findings suggest that organizations can achieve notable gains in productivity, system availability, and recovery speeds while reducing the risk of human error in complex cloud operations. However, realizing fully autonomous cloud systems requires addressing critical limitations, such as model hallucinations, non-deterministic outputs, context-window constraints, and poor action selection. The authors recommend that engineering teams evolve development workflows to support iterative autonomy, integrate formal verification tools to validate generated code, and maintain careful human oversight rather than granting complete, unchecked autonomy.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Its taxonomy of LLM-agent architectures and capabilities provides the conceptual foundation for following the source’s proposed agentic cloud operations.
- Paper: Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy, Ben Shneiderman (2020). Its framework for combining automation with human control clarifies the human-in-the-loop design principle that underpins the source’s autonomy recommendations.
- Paper: Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning, Jiaxing Qi et al. (2026). It advances the source’s automated troubleshooting agenda by evaluating whether LLM diagnoses can be turned into valid, executable microservice recovery actions.
