You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy
Rudrajit ChoudhuriChristian BirdCarmen BadeaMarco GerosaAnita Sarma
Identifies the psychological and operational factors that govern where professional developers accept AI autonomy, showing that task identity and accountability restrict automated decision-making while high workload drives delegation.
As generative artificial intelligence tools rapidly expand across software engineering, organizations face critical choices about how much autonomy to grant machines versus how much control human practitioners must keep. If autonomy levels are set purely by what tools can technically do, organizations risk deskilling developers, obscuring accountability, and degrading long-term software quality. The article investigates where software developers draw boundaries around artificial intelligence autonomy, why they accept or resist delegating specific tasks, and what predicts their willingness to let tools act or decide on their behalf.
The article evaluates these questions using a mixed-methods study of 448 professional developers at Microsoft, analyzing 1,535 task-level responses across the software development lifecycle. The authors mapped open-ended responses onto a five-level autonomy framework—ranging from Level 1 (no artificial intelligence) to Level 5 (full automation)—using an ensemble of three distinct large language models validated by human inter-rater reliability. They then applied mixed-effects regression modeling to examine how developer traits (such as experience and risk tolerance) and cognitive task appraisals (value, identity, accountability, and workload demands) predict accepted autonomy levels and the crossing of two critical thresholds: the action boundary (Level 2 to Level 3, where the tool produces artifacts under human approval) and the decision-making boundary (Level 3 to Level 4, where the tool acts by default and human review is only optional).
The analysis reveals five primary findings. First, a vast majority of developers (74%) capped accepted autonomy at or below Level 3, welcoming artificial intelligence to produce work artifacts while firmly retaining decision-making authority. Second, accepted autonomy varied substantially by task type: developers favored higher autonomy for routine verification, testing, and operations, but restricted artificial intelligence to advisory roles (Level 2 or lower) in human-facing meta-work (such as mentoring and stakeholder communication) and design planning. Third, task identity was a strong negative predictor of autonomy, meaning developers actively retained control over work they found intrinsically fulfilling and central to their professional craft. Fourth, task accountability specifically lowered the odds of crossing the action boundary, as developers who felt personally responsible for outcomes refused to let machines produce artifacts unassisted. Fifth, while high task identity reduced the likelihood of letting artificial intelligence make decisions by default, high task demand increased it, showing that excessive workloads pressure developers to offload decision authority to machines.
These findings indicate that delegating authority to artificial intelligence is fundamentally a work-design challenge rather than a purely technical one. Pushing automation too far risks creating "rubber-stamp" reviews, degrading human oversight, and producing downstream defects that are costly to fix. The article conceptualizes task delegation as "Cascading Locks," where accountability acts as the first gate governing whether artificial intelligence may generate work, and identity acts as the second gate governing whether it may make decisions. When organizations reward sheer throughput without differentiating between routine toil and judgment-building challenges, they trigger organizational anti-patterns such as hollowed-out roles, homogenized problem-solving, and a severed developmental pipeline for junior staff.
To preserve meaningful and effective engineering work, leaders and tool designers should deliberately design workflows around human accountability and judgment rather than accepting default tool settings. Organizations should actively direct artificial intelligence toward repetitive toil while safeguarding the complex, judgment-intensive work that trains expertise. Review workflows should require substantive engagement before deployment rather than passive, optional vetoes, and entry-level tasks must be retained in part for practice to build institutional capability. Because developer experience and risk tolerance steadily increase over time, leaders should treat the article's appraisal framework as an ongoing diagnostic tool to periodically audit autonomy boundaries across teams.
These conclusions are bounded by a single organizational context at an artificial intelligence-forward enterprise, which may reflect higher-than-average baseline tool familiarity and risk tolerance compared to smaller companies or open-source ecosystems. Nonetheless, the high statistical consistency across regression models and robust qualitative agreement provide high confidence in the fundamental relationships identified: meaningful human engagement in automated workflows depends systematically on preserving accountability, protecting professional identity, and managing cognitive demands.
- Paper: Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy, Ben Shneiderman (2020). Establishes the foundational two-dimensional framework showing that high automation and high human control can coexist, directly informing how developers negotiate boundaries of AI autonomy.
- Paper: The SPACE of AI: Real-World Lessons on AI's Impact on Developers, Brian Houck et al. (2025). Provides essential baseline empirical data on multi-dimensional developer experience and real-world AI adoption across industry engineering teams.
- Paper: Human-in-the-loop or AI-in-the-loop? Automate or Collaborate?, Sriraam Natarajan et al. (2025). Clarifies the crucial conceptual distinction between human-in-the-loop automation and AI-in-the-loop collaboration regarding decision authority and control.
- Paper: Magentic-UI: Towards Human-in-the-loop Agentic Systems, Hussein Mozannar et al. (2025). Demonstrates practical human-in-the-loop interaction mechanisms and safety guardrails that developers rely on to oversee autonomous agent workflows.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). Details the agent-computer interfaces and execution feedback mechanisms that define how AI software engineering agents interact with code environments.
- Paper: VizCopilot: Fostering Appropriate Reliance on Enterprise Chatbots with Context Visualization, Sam Yu-Te Lee et al. (2025). Explores how visual oversight mechanisms and context engineering allow human workers to maintain appropriate reliance and accountability over AI tools.
- Paper: From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction, Upol Ehsan et al. (2026). Extends the study of worker autonomy by examining the long-term harms of AI delegation, specifically the erosion of expert skills, judgment, and professional identity.
- Paper: Intelligent AI Delegation, Nenad Tomašev et al. (2026). Builds upon developer delegation preferences by proposing a comprehensive architectural framework for transferring authority, responsibility, and accountability across agent systems.
- Paper: Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems, Jiacheng Liu et al. (2026). Analyzes the concrete system architectures, permission systems, and safety harnesses used in production-grade developer agents like Claude Code to balance autonomy with developer control.
- Paper: LLMs Corrupt Your Documents When You Delegate, Philippe Laban et al. (2026). Evaluates the tangible operational risks and silent document corruption that occur when knowledge workers delegate complex, multi-step editing tasks to autonomous language models.
- Paper: Code as Agent Harness, Xuying Ning et al. (2026). Surveys the design of code-level harnesses that provide programmatic oversight, verification, and grounded control over autonomous agent actions.
- Paper: AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery, Guiyao Tie et al. (2026). Generalizes the levels of task autonomy and oversight beyond software engineering to full workflow automation and accountability in scientific discovery.
