To Copilot and Beyond: 22 AI Systems Developers Want Built
Rudrajit ChoudhuriChristian BirdCarmen BadeaAnita Sarma
Identifies 22 AI systems software engineers want built beyond code generation, using a survey of 860 developers to establish the principle of bounded delegation for designing tools that offload peripheral tasks without intruding on core engineering craft.
Software developers spend approximately one-tenth of their working hours actively writing code, yet the majority of commercial artificial intelligence tooling focuses heavily on code generation. This imbalance accelerates the creation of code while expanding downstream bottlenecks in review, testing, incident triage, and maintenance. As a result, engineering teams face growing technical debt, review fatigue, and poorly understood systems.
The article aims to identify the specific artificial intelligence systems that software developers actually want built across their broader workflow and to define the explicit behavioral boundaries and conditions required for these tools to be accepted in practice.
To investigate this, the study analyzed open-ended survey responses from 860 Microsoft software developers across global business units, roles, and geographies. The research team applied a reflexive qualitative analysis method using an independent multi-model council comprising three frontier language models from different providers (GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6), followed by human reconciliation and rigorous inter-rater reliability checks (achieving an average Krippendorff's alpha of 0.94).
The analysis identified 22 distinct AI systems across five engineering categories: Development, Design and Planning, Quality and Risk Management, Infrastructure and Operations, and Meta-Work. Across these systems, the primary findings show that developer demand is concentrated on the verification and maintenance side of development, with 50.1% of development respondents requesting tools to systematically address technical debt and 44.5% of quality respondents seeking change-aware test generation. Developers strongly rejected autonomous decision-making or direct production modifications, consistently asserting that tools should prepare drafts, compile evidence, and surface alternatives rather than approve changes, deploy code, or communicate directly with stakeholders.
These findings highlight an operating pattern termed bounded delegation: developers want artificial intelligence to absorb mechanical assembly and information retrieval tasks, but strictly preserve human agency and accountability over architectural choices and final evaluations. Unchecked code generation creates a right-shift problem by pushing verification burdens downstream. To prevent developer burnout and safeguard system integrity, organizations must shift quality signals earlier into the authoring process, embedding testing, security, and context mapping at the point of change.
For engineering and technology leaders, the article recommends prioritizing tooling that supports earlier defect detection, evidence gathering, and documentation synchronization over raw code-generation speed. Tool designers and organizations should enforce four mandatory guardrails across all AI systems: explicit authority scoping to halt tools before critical decisions, clear data provenance linking outputs to authoritative sources, proactive uncertainty signaling when confidence is low, and strict least-privilege access ensuring production environments remain read-only for AI.
Because the study reflects cross-sectional self-reported needs from engineers within a single large enterprise, caution is warranted when generalizing to smaller organizations, open-source projects, or differently regulated settings without further validation. Nonetheless, the high internal agreement and large sample size provide strong confidence that sustainable productivity gains depend on respecting developer agency and designing tools around bounded delegation.
- Paper: AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work, Rudrajit Choudhuri et al. (2025). This 2025 study establishes the same Microsoft developer survey’s task-preference and responsible-AI foundations, making its empirical framework essential context for the source’s later account of desired AI systems.
No sufficiently relevant recommendations were found.
