Intent-based System Design and Operation

Vaastav AnandYichen LiAlok KumbhareCeline IrveneChetan BansalGagan SomashekarJonathan MacePedro Las-CasasRicardo BianchiniRodrigo Fonseca

article2025SOSP2 citations

Proposes intent as a foundational abstraction that encodes high-level functional and operational goals to automate cloud system design, implementation, runtime management, and evolution.

Listen

Modern cloud computing relies heavily on distributed microservice architectures that require extensive, continuous manual effort to design, deploy, operate, and maintain. As these systems grow in scale and complexity, human operators struggle to track billions of runtime traces, petabytes of operational logs, and dynamic environment changes. The article proposes an overarching framework for intent-based system design and operation, demonstrating how combining standardized cloud tooling with large language models can automate the entire software lifecycle and produce self-managing cloud services.

The authors conceptualize a human-in-the-loop architectural vision centered on "intent," which encapsulates high-level functional features, operational service-level agreements, and refinement goals. By evaluating case studies across microservice generation, real-time observability, automated troubleshooting, and dynamic resilience, the article synthesizes existing tools—such as Cerulean for hierarchical system generation and Llexus for automated runbook execution—into a unified operational blueprint.

The article delivers several key findings. First, adopting intent as a primary abstraction bridges high-level stakeholder requirements and low-level code generation, significantly reducing manual developer effort. Second, hierarchical generation enables large language models to construct both microservice business logic and corresponding end-to-end test suites directly from user specifications. Third, unifying runtime observability data (metrics, logs, and traces) with static domain knowledge (code and documentation) into event graphs provides real-time context awareness, preventing costly and inefficient data processing. Fourth, dynamic operational models can automate incident mitigation and reproduce complex runtime failures, such as metastable bottlenecks, allowing systems to autonomously generate candidate code fixes.

These findings suggest that organizations can achieve notable gains in productivity, system availability, and recovery speeds while reducing the risk of human error in complex cloud operations. However, realizing fully autonomous cloud systems requires addressing critical limitations, such as model hallucinations, non-deterministic outputs, context-window constraints, and poor action selection. The authors recommend that engineering teams evolve development workflows to support iterative autonomy, integrate formal verification tools to validate generated code, and maintain careful human oversight rather than granting complete, unchecked autonomy.

arXiv: 2502.05984
Cover for Intent-based System Design and Operation

Abstract

Cloud systems are the backbone of today's computing industry. Yet, these systems remain complicated to design, build, operate, and improve. All these tasks require significant manual effort by both developers and operators of these systems.

To reduce this manual burden, in this paper we set forth a vision for achieving holistic automation, intent-based system design and operation. We propose intent as a new abstraction within the context of system design and operation. Intent encodes the functional and operational requirements of the system at a high-level, which can be used to automate design, implementation, operation, and evolution of systems. We detail our vision of intent-based system design, highlight its four key components, and provide a roadmap for the community to enable autonomous systems.

Table of Contents

  • 1 Introduction
  • 2 Intent for Cloud System Design
  • 2.1 Manifesting Intent
  • 2.2 Challenges
  • 3 Intent-based self-managing cloud systems
  • 3.1 Distributed System Design
  • 3.1.1 Use Case: Generating Microservices
  • 3.2 Real-Time Context Awareness
  • 3.2.1 Use Case: Behavior Comprehension
  • 3.3 System Operation
  • 3.3.1 Use Case: Automated Incident Management
  • 3.4 System Improvement
  • 3.4.1 Use Case: Mitigating Metastable Failures
  • 4 Future Directions
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Intent comprises functional, operational, and refinement requirements

    definition

    For a cloud system, intent is the user’s potentially changing, high-level statement of what the system should do and how it should operate. Functional intent specifies features, security requirements, and design requirements; for example, a hotel-reservation service might need to support hotel search, reservations, payments, and reviews. Operational intent specifies operating properties used to define service-level objectives or agreements (SLOs/SLAs), observability, deviation detection, and mitigation; an example is maintaining a 100 ms 99th-percentile latency under high load. Refinement intent is the change from an existing intent to a revised one, and provides the target for improving the system.

  2. Knowl 2 — Intent-based system design and operation is a proposed lifecycle architecture

    model/method

    The paper proposes a cloud-system lifecycle organized around four connected capabilities: automated design, real-time context awareness, autonomous operation, and continuous improvement. Functional and operational intent guide design and operation; design produces a system, and operation compares its behavior with a desired state. Context awareness combines runtime information with system knowledge to inform operation and improvement, while patches or redesigned implementations can be used to address deviations. The proposal uses large language models (LLMs) to turn high-level intent into concrete actions, with those actions applied through automation tools and human oversight retained at key points. The approach is intended to span requirements, design, implementation, testing, deployment, and maintenance rather than automate only one operational task.

  3. Knowl 3 — Generated distributed-system designs require correctness, explainability, and performance guarantees

    model/method

    The proposed LLM-based design process starts from stakeholder requirements: functional features and desired architecture patterns define functional intent, while behavioral properties define operational intent. A generated distributed system should satisfy three classes of guarantees: correctness, relative to user requirements, test suites, and formal specifications; explainability, so people can understand generated code and use accompanying artifacts to inspect the system; and performance, including scalability, service-level objectives, and avoidance of emergent misbehavior. These are requirements for reliable generation, not guarantees shown to have been achieved by the paper.

  4. Knowl 4 — Microservice generation can be extended with end-to-end tests and workload generators

    model/method

    The paper proposes extending a hierarchical LLM-based microservice-generation process with two components. An end-to-end test generator extracts use cases from functional intent, creates an implementation plan as API calls to frontend services using interfaces generated in an earlier stage, and turns the plan into executable tests. The resulting tests can run as conventional Go tests or be compiled as black-box tests against a deployed system. A workload generator uses a workload description from operational intent and the generated interfaces to produce a process that exercises the target workload against the deployed system; if the description is absent, the user is prompted to provide one. The paper also identifies formal-model generation and use as a possible future extension for stronger verification, not as a component already provided.

  5. Knowl 5 — Real-time context awareness joins intent-relevant runtime signals with system knowledge

    model/method

    The proposed context-awareness framework continuously constructs an intent-relevant view of a cloud system’s operating state. It combines runtime information—metrics, logs, traces, and monitors—with domain knowledge such as source code, documentation, and operational or troubleshooting guidelines. User intent helps determine which of the system’s large volume of runtime and domain data is relevant. The resulting contextualized, summarized knowledge is meant to support automated operations and improvement as well as developers’ analysis tasks.

  6. Knowl 6 — Runtime observability can be represented as loosely unified event graphs

    model/method

    For intent-guided retrieval and modeling, the paper proposes representing multimodal runtime data as loosely unified event graphs. Metric events represent anomalies or deviations from normal behavior; log patterns are mined, with noteworthy subsequences treated as events; and traces connect events to express interactions and dependencies among components. The framework also enriches this runtime representation with domain knowledge by extracting documentation relevant to terms in logs and alerts and incorporating source code associated with logs. Pattern mining and comprehension generation are triggered on demand to limit unnecessary costs, and the resulting comprehension can be shared with users and system components.

  7. Knowl 7 — An operational model translates operational intent into desired state and response policy

    definition

    The proposed ops model is a representation derived from a system’s operational intent, including its SLOs and SLAs. It specifies the system’s desired state, identifies and prioritizes operational risks and vulnerabilities that threaten service objectives, defines the observability needed for detection and diagnosis (metrics, monitors, logs, and traces), and describes mitigations and countermeasures for failures. The model is intended to guide operation alongside real-time context about the system’s condition. It is not fixed: it can evolve as the system and user requirements change.

  8. Knowl 8 — Incident response can combine the ops model with live context to generate mitigation instructions

    model/method

    The paper envisions using the ops model and real-time context comprehension together to automate incident management. The ops model supplies operational requirements, relevant observability, and potential mitigations; context comprehension supplies an up-to-date understanding of the system state. Together, they could produce actionable instructions for a tool such as Llexus, which turns troubleshooting guidance into executable plans for incident mitigation and resolution. This proposed route aims to reduce dependence on human-curated troubleshooting guides with complete, high-quality coverage; the paper does not report an evaluation of the combined approach.

  9. Knowl 9 — Continuous improvement uses intent violations to trigger reconfiguration or redesign

    model/method

    The proposed improvement loop detects departures from functional or operational intent by combining intent at relevant abstraction levels with real-time context. Violations may arise from previously unseen bugs, workload changes, metastable failures, or a change in user intent. The system then forms refinement intent and responds either by reconfiguring the deployed system or by redesigning and regenerating part or all of its implementation through hierarchical generation. For a metastable failure, the proposal uses operational and context data to construct trigger scenarios, reproduces the failure in a controlled setting, generates candidate designs, and reruns the scenarios to check whether a candidate avoids the metastable state while meeting functional and operational intent; this cycle continues until a suitable design is found.

  10. Knowl 10 — Reliability, context, action selection, and safe adaptation remain open challenges

    limitation

    The paper presents a vision and research roadmap, not an evaluated autonomous-system implementation. It identifies risks that constrain the proposal: LLM hallucinations and weak numerical or logical reasoning can produce incorrect outputs; generated actions need explanations and validation for human oversight; supplying relevant context is difficult when systems contain extensive code, documentation, and operational data; and instruction inconsistency plus many possible actions complicate correct action selection. Online changes also require safe dynamic reconfiguration without taking the system offline. Broader open problems include machine-readable service and operations models, deterministic and verifiable generation, reasoning under uncertainty or conflicting intents, balancing user burden with meaningful human participation, and guardrails for reliable, secure short- and long-term self-healing.

Coverage note — No substantial contributed material is omitted: the future-work agenda and principal limitations are captured in the limitation knowl. The paper presents a vision and roadmap rather than quantitative experiments, so no empirical-result knowl is included.

References

  1. 1.Dapr: Distributed application runtime. https://dapr.io/.
  2. 2.P. Abrahamsson, O. Salo, J. Ronkainen, and J. Warsta. Agile software development methods: Review and analysis. arXiv preprint arXiv:1709.08439, 2017.
  3. 3.T. Ahmed, S. Ghosh, C. Bansal, T. Zimmermann, X. Zhang, and S. Rajmohan. Recommending root-cause and mitigation steps for cloud incidents using large language models. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), pages 1737–1749. IEEE, 2023.
  4. 4.V. Anand, D. Garg, A. Kaufmann, and J. Mace. Blueprint: A toolchain for highly-reconfigurable microservice applications. In Proceedings of the 29th Symposium on Operating Systems Principles, pages 482–497, 2023.
  5. 5.V. Anand, A. Kumbhare, C. Irvene, C. Bansal, G. Somashekar, J. Mace, P. Las-Casas, and R. Fonseca. Automated service design with cerulean. To appear in 6th International Workshop on Cloud Intelligence / AIOps (AIOps ’25), 2025.
  6. 6.V. Anand, P. Las-Casas, R. Fonseca, and A. Kaufmann. Towards using llms for distributed trace comparison. To appear in 6th International Workshop on Cloud Intelligence / AIOps (AIOps ’25), 2025.
  7. 7.Asim. How a single chatgpt mistake cost us $10,000+. Accessed 9th June, 2024 from https://web.archive.org/web/20240610032818/https://asim.bearblog.dev/how-a-single-chatgpt-mistake-cost-us-10000/, 2024.
  8. 8.N. Bronson, A. Aghayev, A. Charapko, and T. Zhu. Metastable failures in distributed systems. In Proceedings of the Workshop on Hot Topics in Operating Systems, pages 221–227, 2021.
  9. 9.Y. Chen, H. Xie, M. Ma, Y. Kang, X. Gao, L. Shi, Y. Cao, X. Gao, H. Fan, M. Wen, et al. Empowering practical root cause analysis by large language models for cloud incidents. arXiv preprint arXiv:2305.15778, 2023.
  10. 10.A. Clemm, L. Ciavaglia, L. Z. Granville, and J. Tantsura. Rfc 9315: Intent-based networking - concepts and definitions, 2022.
  11. 11.A. Cockcroft. The evolution of microservices. (April 2016). Retrieved October 2020 from https://www.slideshare.net/adriancockcroft/evolution-of-microservices-craft-conference, 2016.
  12. 12.A. Cockcroft. Microservices workshop: Why, what, and how to get there. (April 2016). Retrieved October 2020 from https://www.slideshare.net/adriancockcroft/microservices-workshop-craft-conference, 2016.
  13. 13.V. Ganatra, A. Parayil, S. Ghosh, Y. Kang, M. Ma, C. Bansal, S. Nath, and J. Mace. Detection is better than cure: A cloud incidents perspective. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pages 1891–1902, 2023.
  14. 14.S. Ghemawat, R. Grandl, S. Petrovic, M. Whittaker, P. Patel, I. Posva, and A. Vahdat. Towards modern development of cloud applications. In Proceedings of the 19th Workshop on Hot Topics in Operating Systems, pages 110–117, 2023.
  15. 15.S. Ghosh, M. Shetty, C. Bansal, and S. Nath. How to fight production incidents? an empirical study on a large-scale cloud service. In Proceedings of the 13th Symposium on Cloud Computing, pages 126–141, 2022.
  16. 16.A. Gluck. Introducing domain-oriented microservice architecture. Accessed June 2024 from https://www.uber.com/blog/microservice-architecture/, 2020.
  17. 17.D. Goel, F. Husain, A. Singh, S. Ghosh, A. Parayil, C. Bansal, X. Zhang, and S. Rajmohan. X-lifecycle learning for cloud incident management using llms. arXiv preprint arXiv:2404.03662, 2024.
  18. 18.E. Haddad. Service-oriented architecture: Scaling the uber engineering codebase as we grow. (September 2015). Retrieved October 2020 from https://eng.uber.com/service-oriented-architecture/, 2015.
  19. 19.P. Hamadanian, B. Arzani, S. Fouladi, S. K. R. Kakarla, R. Fonseca, D. Billor, A. Cheema, E. Nkposong, and R. Chandra. A holistic view of ai-driven network incident management. In Proceedings of the 22nd ACM Workshop on Hot Topics in Networks, pages 180–188, 2023.
  20. 20.M. Hashemi. The infrastructure behind twitter : Scale. (January 2017). Retrieved February 2021 from https://blog.twitter.com/engineering/en_us/topics/infrastructure/2017/the-infrastructure-behind-twitter-scale.html, 2017.
  21. 21.L. Huang, M. Magnusson, A. B. Muralikrishna, S. Estyak, R. Isaacs, A. Aghayev, T. Zhu, and A. Charapko. Metastable failures in the wild. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), pages 73–90, 2022.
  22. 22.L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232, 2023.
  23. 23.Y. Jiang, C. Zhang, S. He, Z. Yang, M. Ma, S. Qin, Y. Kang, Y. Dang, S. Rajmohan, Q. Lin, et al. Xpert: Empowering incident management with query recommendations via large language models. arXiv preprint arXiv:2312.11988, 2023.
  24. 24.J. Kaldor, J. Mace, M. Bejda, E. Gao, W. Kuropatwa, J. O’Neill, K. W. Ong, B. Schaller, P. Shan, B. Viscomi, et al. Canopy: An end-to-end performance tracing and analysis system. In Proceedings of the 26th symposium on operating systems principles, pages 34–50, 2017.
  25. 25.B. Lampson. Hints and principles for computer system design. arXiv preprint arXiv:2011.02455, 2020.
  26. 26.B. W. Lampson. Hints for computer system design. In Proceedings of the ninth ACM symposium on Operating systems principles, pages 33–48, 1983.
  27. 27.P. Las-Casas, A. Kumbhare, R. Fonseca, and S. Agarwal. Llexus: an ai agent system for incident management. SIGOPS Oper. Syst. Rev., 58(1), 2024. To appear.
  28. 28.C. Lee, T. Yang, Z. Chen, Y. Su, and M. Lyu. Maat: Performance metric anomaly anticipation for cloud services with conditional diffusion. In 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE), pages 116–128. IEEE, 2023.
  29. 29.Y. Li, X. Zhang, S. He, Z. Chen, Y. Kang, J. Liu, L. Li, Y. Dang, F. Gao, Z. Xu, et al. An intelligent framework for timely, accurate, and comprehensive cloud incident detection. ACM SIGOPS Operating Systems Review, 56(1):1–7, 2022.
  30. 30.F. Lin, K. Muzumdar, N. P. Laptev, M.-V. Curelea, S. Lee, and S. Sankar. Fast dimensional analysis for root cause investigation in a large-scale service environment. Proceedings of the ACM on Measurement and Analysis of Computing Systems, 4(2):1–23, 2020.
  31. 31.J. Liu, J. Zhu, S. He, P. He, Z. Zheng, and M. R. Lyu. Logzip: Extracting hidden structures via iterative clustering for log compression. In 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE), pages 863–873. IEEE, 2019.
  32. 32.J. C. Mogul. Emergent (mis) behavior vs. complex software systems. ACM SIGOPS Operating Systems Review, 40(4):293–304, 2006.
  33. 33.C. M. Rosenberg and L. Moonen. Spectrum-based log diagnosis. In Proceedings of the 14th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM), pages 1–12, 2020.
  34. 34.D. Roy, X. Zhang, R. Bhave, C. Bansal, P. Las-Casas, R. Fonseca, and S. Rajmohan. Exploring llm-based agents for root cause analysis. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, FSE 2024, page 208–219, New York, NY, USA, 2024. Association for Computing Machinery.
  35. 35.V. Seshagiri, S. Balyan, V. Anand, K. Dhole, I. Sharma, A. Wildani, J. Cambronero, and A. Züfle. Chatting with logs: An exploratory study on finetuning llms for logql. arXiv preprint arXiv:2412.03612, 2024.
  36. 36.E. Shanks. Kubernetes - desired state and control loops. Accessed July, 2024 from https://theithollow.com/2019/09/16/kubernetes-desired-state-and-control-loops/, 2019.
  37. 37.G. Somashekar, K. Tandon, A. Kini, C.-C. Chang, P. Husak, R. Bhagwan, M. Das, A. Gandhi, and N. Natarajan. OPPerTune: Post-Deployment configuration tuning of services made easy. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), pages 1101–1120, Santa Clara, CA, Apr. 2024. USENIX Association.
  38. 38.P. Srinivas, F. Husain, A. Parayil, A. Choure, C. Bansal, and S. Rajmohan. Intelligent monitoring framework for cloud services: A data-driven approach. In Proceedings of the 46th International Conference on Software Engineering: Software Engineering in Practice, pages 381–391, 2024.
  39. 39.H. Wang, G. K. Tangirala, G. P. Naidu, C. Mayville, A. Roy, J. Sun, and R. B. Mandava. Anomaly detection for incident response at scale. arXiv preprint arXiv:2404.16887, 2024.
  40. 40.Z. Xie, Y. Zheng, L. Ottens, K. Zhang, C. Kozyrakis, and J. Mace. Cloud atlas: Efficient fault localization for clou systems using language models and causal insight. Accessed 11th July, 2024 from https://people.mpi-sws.org/~jcmace/papers/xie2024cloud.pdf, 2024.
  41. 41.G. Yu, P. Chen, Y. Li, H. Chen, X. Li, and Z. Zheng. Nezha: Interpretable fine-grained root causes analysis for microservices on multi-modal observability data. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pages 553–565, 2023.
  42. 42.D. Zhang, X. Zhang, C. Bansal, P. Las-Casas, R. Fonseca, and S. Rajmohan. Lm-pace: Confidence estimation by large language models for effective root causing of cloud incidents. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, pages 388–398, 2024.
  43. 43.X. Zhang, S. Ghosh, C. Bansal, R. Wang, M. Ma, Y. Kang, and S. Rajmohan. Automated root causing of cloud incidents using in-context learning with gpt-4. arXiv preprint arXiv:2401.13810, 2024.

Citation

MLA
Anand, V., et al. “Intent-based System Design and Operation”. arXiv, 2025, http://arxiv.org/abs/2502.05984v1.
APA
Anand, V., Li, Y., Kumbhare, A. G., Irvene, C., Bansal, C., Somashekar, G., Mace, J., Las-Casas, P., & Fonseca, R. (2025). Intent-based System Design and Operation. arXiv. http://arxiv.org/abs/2502.05984v1
Chicago
Anand, V., Y. Li, A. G. Kumbhare, et al. 2025. “Intent-based System Design and Operation”. arXiv. http://arxiv.org/abs/2502.05984v1.
Harvard
Anand, V. et al. (2025) “Intent-based System Design and Operation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.05984v1.
Vancouver
1. Anand V, Li Y, Kumbhare AG, Irvene C, Bansal C, Somashekar G, Mace J, Las-Casas P, Fonseca R (2025) Intent-based System Design and Operation. arXiv

BibTeX

@article{anand2025intent,
  title = {Intent-based System Design and Operation},
  author = {Anand, Vaastav and Li, Yichen and Kumbhare, Alok Gautam and Irvene, Celine and Bansal, Chetan and Somashekar, Gagan and Mace, Jonathan and Las-Casas, Pedro and Fonseca, Rodrigo},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.05984v1},
  eprint = {2502.05984}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/