AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work
Rudrajit ChoudhuriCarmen BadeaChristian BirdJenna ButlerRobert DeLIneBrian Houck
Presents an empirical model derived from 860 developers that uses cognitive appraisal theory to identify where practitioners want generative AI assistance, where they resist automation, and how responsible AI priorities shift across software engineering tasks.
As generative artificial intelligence tools rapidly enter software engineering workflows, technology leaders face an ongoing challenge: tool capabilities are advancing quickly, yet organizations lack empirical clarity on where developers actually want assistance, where they prefer to retain manual control, and how to govern these tools responsibly. This lack of alignment risks investing in the wrong automation priorities and undermining meaningful developer work.
The main objective of the article is to establish an empirically validated framework showing how developers evaluate their daily tasks, how these evaluations predict their openness to and use of artificial intelligence, and which Responsible AI principles they prioritize across different types of software work.
To evaluate these dynamics, the researchers conducted a large-scale, mixed-methods survey in July 2025 across 860 software developers at Microsoft spanning global regions, roles, and experience levels. Grounded in psychological cognitive appraisal theory, the study gathered quantitative ratings across a comprehensive taxonomy of software engineering tasks, forced-choice trade-off rankings across eight Responsible AI principles, and qualitative rationales across thousands of open-ended responses. The data was analyzed using mixed-effects regression, hierarchical clustering, and reflexive thematic analysis.
The analysis produced several key findings. First, task appraisals strongly predict adoption: perceived task value, accountability, and cognitive demands each positively correlate with openness to and use of artificial intelligence, whereas professional identity exhibits a dual dynamic—reducing general openness to delegating work while increasing usage when tools complement personal craft. Second, tasks divide into distinct operational clusters: core development work (e.g., coding, testing, debugging) exhibits high demand for tool improvement, operational toil (e.g., continuous integration, environment maintenance) shows high demand for new automation, and interpersonal work (e.g., mentoring, stakeholder relations) meets strong resistance to automation. Third, Responsible AI priorities follow a clear hierarchy: across all tasks, baseline priorities focus heavily on Reliability & Safety (selected by 85% of respondents), Privacy & Security (77%), and Transparency (72%), whereas principles like Fairness (32%) and Inclusiveness (32%) become primary only in human-facing or design contexts. Finally, experience and individual traits matter: senior engineers and technophiles strongly prioritize steerability and control over direct execution.
These findings indicate that developers view artificial intelligence as a collaborator rather than an outright replacement. In high-stakes and system-critical tasks, automation errors carry high remediation costs, meaning developers insist on verifiability, provenance, and final sign-off rather than end-to-end automation. Conversely, for identity-defining and relational work, excessive automation risks deskilling engineers and damaging team culture.
For engineering leaders and tool designers, the article recommends prioritizing artificial intelligence development around cognitive augmentation rather than blunt replacement. Organizations should focus automation investments on reducing operational toil and assisting with boilerplate steps in core work, while keeping human oversight, suggest-only workflows, and reversible edits standard. Interpersonal activities like mentoring and direct communication should remain fundamentally human-led. Further research and internal piloting should focus on designing better observability logs, drift-prevention mechanisms, and adaptive autonomy controls.
The conclusions are drawn from a cross-sectional study within a single large enterprise (Microsoft), meaning findings capture self-reported perceptions and correlational patterns that may not fully generalize to open-source communities or smaller firms. Nonetheless, the substantial sample size and rigorous multi-method design provide high confidence in the identified appraisal drivers and design trade-offs.
- Paper: Generative AI, Stefan Feuerriegel et al. (2023). Its technical and organizational framework for generative AI clarifies the capabilities, probabilistic limits, and sociotechnical risks that the source maps onto developers’ daily software-engineering work.
- Paper: On the Opportunities and Risks of Foundation Models, Rishi Bommasani et al. (2021). Its account of foundation-model capabilities, homogenization, and inherited risks provides the conceptual basis for the source’s context-specific Responsible AI priorities.
- Paper: Large Language Models for Software Engineering: A Systematic Literature Review, Xinying Hou et al. (2023). Its systematic map of LLM applications across software-engineering tasks supplies the task landscape that the source refines through developer appraisals and adoption patterns.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Its taxonomy of LLM-agent architecture and evaluation prepares the reader to understand the source’s distinctions between AI assistance, autonomy, and human control across development work.
- Paper: Human-in-the-loop or AI-in-the-loop? Automate or Collaborate?, Sriraam Natarajan et al. (2025). Its distinction between human-in-the-loop and AI-in-the-loop systems establishes the control and authority concepts underlying the source’s analysis of responsible AI priorities.
- Paper: You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy, Rudrajit Choudhuri et al. (2026). It extends the source’s task-appraisal findings into a granular account of where developers accept, limit, or reject different levels of AI autonomy.
- Paper: From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction, Upol Ehsan et al. (2026). It continues the source’s concern with developer identity and human-facing work by examining how sustained AI use can erode expertise, judgment, and professional agency.
- Paper: The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior, Xiangzhe Xu et al. (2026). It applies the source’s call for contextual AI support to agent-tool design, showing how interface architecture changes exploration, consistency, and execution efficiency.
- Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). It operationalizes the source’s Responsible AI priorities by diagnosing unsafe multi-step agent behavior, tool misuse, prompt injection, and root causes of security failures.
- Paper: SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Yuhang Wang et al. (2026). It extends the source’s focus on fitting AI to developers’ workflows by reducing irrelevant repository context while preserving coding-agent performance and efficiency.
