You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy

Rudrajit ChoudhuriChristian BirdCarmen BadeaMarco GerosaAnita Sarma

article2026arXiv2 citations

Identifies the psychological and operational factors that govern where professional developers accept AI autonomy, showing that task identity and accountability restrict automated decision-making while high workload drives delegation.

Listen

As generative artificial intelligence tools rapidly expand across software engineering, organizations face critical choices about how much autonomy to grant machines versus how much control human practitioners must keep. If autonomy levels are set purely by what tools can technically do, organizations risk deskilling developers, obscuring accountability, and degrading long-term software quality. The article investigates where software developers draw boundaries around artificial intelligence autonomy, why they accept or resist delegating specific tasks, and what predicts their willingness to let tools act or decide on their behalf.

The article evaluates these questions using a mixed-methods study of 448 professional developers at Microsoft, analyzing 1,535 task-level responses across the software development lifecycle. The authors mapped open-ended responses onto a five-level autonomy framework—ranging from Level 1 (no artificial intelligence) to Level 5 (full automation)—using an ensemble of three distinct large language models validated by human inter-rater reliability. They then applied mixed-effects regression modeling to examine how developer traits (such as experience and risk tolerance) and cognitive task appraisals (value, identity, accountability, and workload demands) predict accepted autonomy levels and the crossing of two critical thresholds: the action boundary (Level 2 to Level 3, where the tool produces artifacts under human approval) and the decision-making boundary (Level 3 to Level 4, where the tool acts by default and human review is only optional).

The analysis reveals five primary findings. First, a vast majority of developers (74%) capped accepted autonomy at or below Level 3, welcoming artificial intelligence to produce work artifacts while firmly retaining decision-making authority. Second, accepted autonomy varied substantially by task type: developers favored higher autonomy for routine verification, testing, and operations, but restricted artificial intelligence to advisory roles (Level 2 or lower) in human-facing meta-work (such as mentoring and stakeholder communication) and design planning. Third, task identity was a strong negative predictor of autonomy, meaning developers actively retained control over work they found intrinsically fulfilling and central to their professional craft. Fourth, task accountability specifically lowered the odds of crossing the action boundary, as developers who felt personally responsible for outcomes refused to let machines produce artifacts unassisted. Fifth, while high task identity reduced the likelihood of letting artificial intelligence make decisions by default, high task demand increased it, showing that excessive workloads pressure developers to offload decision authority to machines.

These findings indicate that delegating authority to artificial intelligence is fundamentally a work-design challenge rather than a purely technical one. Pushing automation too far risks creating "rubber-stamp" reviews, degrading human oversight, and producing downstream defects that are costly to fix. The article conceptualizes task delegation as "Cascading Locks," where accountability acts as the first gate governing whether artificial intelligence may generate work, and identity acts as the second gate governing whether it may make decisions. When organizations reward sheer throughput without differentiating between routine toil and judgment-building challenges, they trigger organizational anti-patterns such as hollowed-out roles, homogenized problem-solving, and a severed developmental pipeline for junior staff.

To preserve meaningful and effective engineering work, leaders and tool designers should deliberately design workflows around human accountability and judgment rather than accepting default tool settings. Organizations should actively direct artificial intelligence toward repetitive toil while safeguarding the complex, judgment-intensive work that trains expertise. Review workflows should require substantive engagement before deployment rather than passive, optional vetoes, and entry-level tasks must be retained in part for practice to build institutional capability. Because developer experience and risk tolerance steadily increase over time, leaders should treat the article's appraisal framework as an ongoing diagnostic tool to periodically audit autonomy boundaries across teams.

These conclusions are bounded by a single organizational context at an artificial intelligence-forward enterprise, which may reflect higher-than-average baseline tool familiarity and risk tolerance compared to smaller companies or open-source ecosystems. Nonetheless, the high statistical consistency across regression models and robust qualitative agreement provide high confidence in the fundamental relationships identified: meaningful human engagement in automated workflows depends systematically on preserving accountability, protecting professional identity, and managing cognitive demands.

arXiv: 2607.00533
Cover for You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy

Abstract

As AI takes on more software work, the line between human and AI effort is shifting. Where developers draw that line around AI autonomy bears on how we design tools and roles that preserve meaningful work. Drawing on cognitive appraisal theory, work design, and automation research, we conducted a mixed-methods study of 448 professional developers at Microsoft to investigate their accepted levels of AI autonomy across software engineering work. Most developers accepted AI producing work under their oversight, although accepted autonomy varied substantively across tasks and individuals. Acceptance was lowest for identity-defining, human-facing, and design-oriented work, and higher among developers with more AI experience and risk tolerance. Task accountability was associated with lower odds of allowing AI to act on developers' behalf, whereas task identity was associated with lower odds of granting AI decision-making autonomy. Task demands had the opposite effect, increasing willingness to delegate decision-making to AI. Our findings suggest that preferences for AI autonomy reflect how developers cognitively experience their work, highlighting important considerations for designing meaningful work.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Theory and Hypotheses
  • 4 Method
  • 4.1 Study Design
  • 4.1.1 Survey instrument design
  • 4.1.2 Data collection
  • 4.2 Data Analysis
  • 4.2.1 Autonomy level classification
  • 4.2.2 Mixed-methods analysis
  • 4.3 Threats to Validity
  • 5 Results
  • 5.1 RQ1: Accepted Autonomy Levels and their Predictors
  • 5.1.1 Autonomy Level Ceilings
  • 5.1.2 Predictors of accepted AI autonomy
  • 5.2 RQ2: Crossing the Action and Decision-Making Boundaries
  • 6 Cascading Locks of Meaningful Work Design
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Five-Level AI Autonomy Scale and Autonomy Transition Boundaries

    definition

    In software engineering human-AI collaboration, the division of authority and execution across tasks is formalized using a five-level ordinal autonomy scale adapted from Endsley and Kiris (1995), characterized by who decides and who acts:

    • Level 1 (None): Developer Decides and Acts; AI has no role.
    • Level 2 (Decision Support): Developer Decides and Acts; AI Suggests or flags possibilities while the developer creates the work artifact.
    • Level 3 (Consensual): AI Acts to produce work artifacts, but the developer Decides and Acts such that results take effect only after mandatory human sign-off/approval.
    • Level 4 (Monitored): AI Decides and Acts by default; the developer retains an optional Veto.
    • Level 5 (Full Automation): AI Decides and Acts autonomously without human intervention.

    Two critical transitions demarcate distinct shifts in human control:

    1. The Action Boundary (L2→L3L2 \rightarrow L3): AI transitions from advising to generating the concrete work artifact or executing the action under mandatory human approval.
    2. The Decision-Making Boundary (L3→L4L3 \rightarrow L4): Human sign-off transitions from mandatory to optional, delegating default decision-making authority to the AI.
  2. Knowl 2 — Predictors of Accepted AI Autonomy in Software Development Tasks

    data/table

    A Cumulative Link Mixed Model (CLMM) with a logit link estimated the effects of cognitive task appraisals, developer traits, and software development lifecycle (SDLC) task categories on the highest level of AI autonomy developers accept (L1L1 to L5L5). The model was fit on n=1,476n = 1{,}476 task-level responses from 448 professional developers, using a participant-level random intercept and treating boundary-ambiguous labels as interval-censored observations.

    Predictor β\beta 95% CI rr pp
    Task Appraisals
    Task Value 0.060.06 [−0.07,0.18][-0.07, 0.18] 0.030.03 0.3730.373
    Task Identity −0.13-0.13 [−0.24,−0.02][-0.24, -0.02] 0.070.07 0.017∗0.017^*
    Task Accountability −0.07-0.07 [−0.19,0.05][-0.19, 0.05] 0.040.04 0.2760.276
    Task Demand 0.050.05 [−0.06,0.17][-0.06, 0.17] 0.030.03 0.3730.373
    Developer Traits
    Software Engineering Experience −0.10-0.10 [−0.33,0.13][-0.33, 0.13] 0.060.06 0.3930.393
    AI Experience 0.260.26 [0.07,0.45][0.07, 0.45] 0.140.14 0.007∗∗0.007^{**}
    Risk Tolerance 0.330.33 [0.07,0.60][0.07, 0.60] 0.180.18 0.013∗0.013^*
    Technophilia 0.020.02 [−0.22,0.25][-0.22, 0.25] 0.010.01 0.8980.898
    SDLC Categories vs. Development (baseline) OR 95% CI rr pp
    Design Planning 0.420.42 [0.31,0.57][0.31, 0.57] 0.480.48 <.001∗∗∗<.001^{***}
    Quality Risk 1.671.67 [1.18,2.38][1.18, 2.38] 0.280.28 0.004∗∗0.004^{**}
    Infrastructure Ops 0.790.79 [0.58,1.19][0.58, 1.19] 0.130.13 0.0530.053
    Meta-work 0.330.33 [0.24,0.46][0.24, 0.46] 0.610.61 <.001∗∗∗<.001^{***}
    Rm2=0.07,  Rc2=0.49R^2_m = 0.07,\; R^2_c = 0.49; ∗p<.05,∗∗p<.01,∗∗∗p<.001{}^*p<.05, {}^{**}p<.01, {}^{***}p<.001 (Benjamini-Hochberg corrected).

    Among cognitive appraisals, only Task Identity significantly predicted overall accepted AI autonomy (standardized β=−0.13,p=0.017\beta = -0.13, p = 0.017), demonstrating that developers cede less autonomy on tasks that define their professional identity. Developer traits showing significant positive associations were AI Experience (β=0.26,p=0.007\beta = 0.26, p = 0.007) and Risk Tolerance (β=0.33,p=0.013\beta = 0.33, p = 0.013). Relative to core Development work, developers accepted significantly less autonomy for Meta-work (OR=0.33\text{OR} = 0.33) and Design & Planning (OR=0.42\text{OR} = 0.42), while accepting significantly more autonomy for Quality & Risk tasks (OR=1.67\text{OR} = 1.67).

  3. Knowl 3 — Predictors of Crossing the Action and Decision-Making Boundaries

    data/table

    Binary logistic Generalized Linear Mixed Models (GLMMs) were fit to evaluate the odds of developers permitting AI to cross the Action Boundary (L2→L3L2 \rightarrow L3, AI produces work artifacts under human sign-off) versus the Decision-Making Boundary (L3→L4L3 \rightarrow L4, AI acts and decides by default with optional human veto) across n=1,476n = 1{,}476 task responses from 448 developers.

    Predictor Action Boundary (L2→L3L2 \rightarrow L3) Decision-Making Boundary (L3→L4L3 \rightarrow L4)
    OR [95% CI] OR [95% CI]
    Task Appraisals
    Task Value 0.95  [0.79,1.13]0.95\; [0.79, 1.13] 1.05  [0.87,1.27]1.05\; [0.87, 1.27]
    Task Identity 0.89  [0.74,1.07]0.89\; [0.74, 1.07] 0.78∗  [0.64,0.95]0.78^*\; [0.64, 0.95]
    Task Accountability 0.82∗  [0.70,0.98]0.82^*\; [0.70, 0.98] 0.91  [0.76,1.09]0.91\; [0.76, 1.09]
    Task Demand 0.97  [0.83,1.14]0.97\; [0.83, 1.14] 1.20∗  [1.05,1.42]1.20^*\; [1.05, 1.42]
    Developer Traits
    Software Engineering Experience 0.92  [0.68,1.24]0.92\; [0.68, 1.24] 0.84  [0.62,1.14]0.84\; [0.62, 1.14]
    AI Experience 1.27  [1.00,1.61]1.27\; [1.00, 1.61] 1.29∗  [1.01,1.65]1.29^*\; [1.01, 1.65]
    Risk Tolerance 1.40  [1.00,1.96]1.40\; [1.00, 1.96] 1.46∗  [1.03,2.07]1.46^*\; [1.03, 2.07]
    Technophilia 0.91  [0.67,1.23]0.91\; [0.67, 1.23] 1.06  [0.78,1.45]1.06\; [0.78, 1.45]
    Model Fit Rm2=0.103,  Rc2=0.574R^2_m = 0.103,\; R^2_c = 0.574 Rm2=0.065,  Rc2=0.543R^2_m = 0.065,\; R^2_c = 0.543
    ∗p<.05,∗∗p<.01,∗∗∗p<.001{}^*p < .05, {}^{**}p < .01, {}^{***}p < .001.

    The two autonomy boundaries answer to different psychological appraisals:

    1. Action Boundary (L2→L3L2 \rightarrow L3): Task Accountability was the sole significant cognitive appraisal predictor (OR=0.82,p=0.027\text{OR} = 0.82, p = 0.027). High felt responsibility for an outcome lowers the odds of allowing AI to generate the artifact, keeping AI restricted to suggestion.
    2. Decision-Making Boundary (L3→L4L3 \rightarrow L4): Task Identity and Task Demand pull in opposite directions. Higher Task Identity lowers the odds of ceding decision authority (OR=0.78,p=0.010\text{OR} = 0.78, p = 0.010), whereas higher Task Demand increases the odds of delegating decision authority to AI (OR=1.20,p=0.030\text{OR} = 1.20, p = 0.030). Additionally, AI Experience (OR=1.29,p=0.044\text{OR} = 1.29, p = 0.044) and Risk Tolerance (OR=1.46,p=0.034\text{OR} = 1.46, p = 0.034) significantly raise the odds of crossing the decision-making boundary.
  4. Knowl 4 — Distribution and Ceilings of Developer-Accepted AI Autonomy Across SDLC Categories

    empirical result

    Across 1,535 task responses from 448 professional developers, the overall distribution of accepted AI autonomy levels is:

    • Level 1 (None): 12%
    • Level 2 (Decision Support): 36%
    • Level 3 (Consensual): 26%
    • Level 4 (Monitored): 10%
    • Level 5 (Full Automation): 16%

    Overall, the median accepted autonomy level is Level 3, with 74% of all responses setting an autonomy ceiling at or below Level 3 (retaining human decision-making and required approval) and only 26% permitting AI to decide and act without required pre-commit human approval (>L3> L3).

    Accepted autonomy varies across software development lifecycle (SDLC) categories:

    • System-Facing Tasks (Median: Level 3):
      • Development: 15% L1, 26% L2, 29% L3, 13% L4, 17% L5 (70% kept ≤L3\le L3, 30% ceded >L3> L3). Routine coding, debugging, and refactoring are accepted under supervision, but complex core logic is retained.
      • Quality & Risk: 4% L1, 33% L2, 37% L3, 10% L4, 16% L5 (74% kept ≤L3\le L3, 26% ceded >L3> L3). Routine verification, vulnerability scanning, and test generation receive high autonomy, whereas code review verdicts require post-review oversight.
      • Infrastructure & Ops: 12% L1, 25% L2, 32% L3, 15% L4, 16% L5 (69% kept ≤L3\le L3, 31% ceded >L3> L3). Setup, log monitoring, and alert triage accept high autonomy, while production releases remain guarded.
    • Human- and Design-Facing Tasks (Median: Level 2):
      • Design & Planning: 14% L1, 51% L2, 15% L3, 4% L4, 16% L5 (65% kept ≤L2\le L2, 35% ceded >L2> L2). Suggestions and draft plans are accepted, but architectural decisions and cross-team coordination are kept human.
      • Meta-Work: 12% L1, 53% L2, 16% L3, 9% L4, 10% L5 (65% kept ≤L2\le L2, 35% ceded >L2> L2). Mentoring, onboarding, and interpersonal client communication are retained at L1/L2, while documentation authoring is delegated further (L3).
  5. Knowl 5 — Cascading Locks Framework and Meaningful Work Design Anti-Patterns

    model/method

    The Cascading Locks framework models AI autonomy on a task as a vessel navigating a sequential flight of canal locks. The first gate is at the Action Boundary (L2→L3L2 \rightarrow L3), governed by Task Accountability (felt responsibility holds AI at suggestion until cleared). The second gate is at the Decision-Making Boundary (L3→L4L3 \rightarrow L4), governed by Task Identity (attachment to craft holds decision power). Developer AI experience and risk tolerance raise the overall water level to clear gates, while task demand introduces an additional inflow at the decision-making gate, pushing tasks toward default AI decision-making.

    When tool defaults or velocity pressures bypass these locks, ten distinct anti-patterns emerge across four design themes:

    Theme Anti-pattern Description
    Accountability Deferred answerability Review steps automated; responsibility offloaded until no one owns the result.
    retention Right-shifted quality Quality gates migrate to verification; errors surface late and cost more to fix.
    Demand Throughput stampede Incentives reward velocity; metrics conflate toil and challenge, shedding work
    differentiation that once defined self-worth and identity.
    Identity Shrinking fence Craft core defined by AI capability gaps, which shrinks with each release.
    evolution Hollow orchestrator Role title remains, but person cannot follow what they approve; veto is nominal.
    Expertise commoditization AI does underlying work; expertise that distinguished the role stops being scarce.
    Thought homogenization Developers defer to same AI framing; judgment converges, reducing innovation.
    Identity Severed pipeline Entry-level work automated for throughput; juniors lack scaffolding for hard problems.
    formation Cognitive debt accrual Juniors offload learning work to AI; skill gaps hide until AI fails.
    Mentoring-by-bot Onboarding automated; relationships needed to pass tacit knowledge never form.
  6. Knowl 6 — Rule-Based and Multi-LLM Ensemble Classification of Autonomy Levels

    algorithm

    To classify open-ended survey responses into ordinal autonomy levels, an automated rule-based algorithm operates over paired Want (desired AI involvement) and Resist (refused AI involvement) text fields:

    Input: Want field ww and Resist field rr for a participant-category cell
    Output: One or more rows ⟨task,source,label,rationale⟩\langle \text{task}, \text{source}, \text{label}, \text{rationale} \rangle
    Function classify(ss):
        if REFUSAL of all AI for the task, no carve-out then
            return L1
        else if REFUSAL + SUGGESTION in limited cases then
            return L1-L2
        else if AI contributes SUGGESTION; takes no ACTION or produces no ARTIFACT then
            return L2
        else if AI produces an ARTIFACT the human is framed as extending then
            return L2-L3
        else if APPROVAL required before AI proceeds with an ACTION then
            return L3
        else if AI ACTION language present + no APPROVAL/VETO anchor then
            return L3-L4
        else if AI ACTION by default; human keeps VETO then
            return L4
        else if AI ACTION with MONITORING-ONLY then
            return L4-L5
        else if AI ACTION with NO-HUMAN-CONTROL then
            return L5
    W←Tasks(w)W \leftarrow \text{Tasks}(w); R←Tasks(r)R \leftarrow \text{Tasks}(r)
    Pair WW and RR by same-task hyponym match; set source ∈{Want,Resist,Both}\in \{\text{Want}, \text{Resist}, \text{Both}\}
    for each paired task row do
        s←ws \leftarrow w if Want, rr if Resist, w∥rw \parallel r if Both
        label←classify(s)\text{label} \leftarrow \text{classify}(s)
        return row with label and rationale

    The classification execution employs an AI-council ensemble of three frontier LLMs from distinct model families (gpt-5.4, gemini-3-flash, claude-sonnet-4.6 in high-reasoning mode) applying the rules independently. Output labels are reconciled via majority voting (≥2/3 \ge 2/3). Inter-rater reliability against human expert re-coding of a random sample (n=67n = 67) yielded human-human agreement κ=0.94\kappa = 0.94, human-council agreement κ=0.95\kappa = 0.95 and 0.920.92, and overall three-rater Krippendorff's α=0.93\alpha = 0.93.

  7. Knowl 7 — Cognitive Appraisal Survey Protocol for AI Delegation Preferences

    experimental setup

    The empirical study surveyed N=448N = 448 professional software developers at Microsoft (yielding 1,535 task responses and 1,476 complete predictor sets) across 5 software development lifecycle (SDLC) categories comprising 19 distinct tasks:

    1. Development: Coding, Bug fixing, Performance optimization, Refactoring, AI integration.
    2. Design & Planning: System design, Requirements engineering, Project planning & management.
    3. Quality & Risk: Testing/QA, Code review, Security & compliance.
    4. Infrastructure & Ops: DevOps (CI/CD), Environment setup & maintenance, Infrastructure monitoring, Customer support.
    5. Meta-work: Learning, Research, Documentation, Stakeholder communication, Mentoring.

    Participants appraised each task in their chosen categories on 5-point Likert scales across four theoretical constructs:

    • Task Value: Perceived importance of task to project success and personal goals (Job Characteristics Model).
    • Task Identity: Task enjoyment and alignment with professional self-concept (Self-Determination Theory).
    • Task Accountability: Felt personal responsibility and answerability for task outcomes (Felt Accountability Scale).
    • Task Demands: Cognitive effort and workload imposed by the task (Job Demands-Resources Model).

    Participants also completed the Cognitive Style Facet Survey for Risk Tolerance and Technophilia, reported Software Engineering and AI Tool Experience, and provided open-ended responses specifying desired (Want) and resisted (Resist) AI roles over a 1–3 year horizon.

Coverage note — All major contributions, empirical models, qualitative findings, classification algorithms, and design frameworks from the paper are included; pilot evaluation details and general demographic breakdowns were omitted as auxiliary.

References

  1. 1.[n. d.]. Supplemental Package. https://zenodo.org/record/21060399.
  2. 2.Sadia Afroz, Zixuan Feng, Tyler Menezes, Katie Kimura, Bianca Trinkenreich, Igor Steinmacher, and Anita Sarma. 2026. The Fast and Spurious: Developer Productivity with GenAI. ACM International Conference on the Foundations of Software Engineering.
  3. 3.George A. Akerlof and Rachel E. Kranton. 2000. Economics and Identity. The Quarterly Journal of Economics 115, 3 (2000), 715–753.
  4. 4.Moaath Alshaikh, Tasneem Alshaher, Ricardo Vieira, Beatriz Santana, Clelio Xavier, Jose Amancio, Glauco Carneiro, Julio Leite, Savio Freire, and Manoel Mendonca. 2026. Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study. In Proceedings of the 1st International Workshop on Prompt Engineering for Software Engineering (PROMPT-SE 2026). Co-located with EASE 2026; arXiv:2605.07422.
  5. 5.Andrew Anderson, Jimena Noa Guevara, Fatima Moussaoui, Tianyi Li, Mihaela Vorvoreanu, and Margaret Burnett. 2022. Measuring User Experience Inclusivity in Human-AI Interaction via Five User Problem-Solving Styles. ACM Transactions on Interactive Intelligent Systems (2022).
  6. 6.Anthropic. 2025. Claude. https://claude.ai/.
  7. 7.Julian Ashwin, Aditya Chhabra, and Vijayendra Rao. 2026. Using large language models for qualitative analysis can introduce serious bias. Sociological Methods & Research 55, 3 (2026), 795–839.
  8. 8.David H Autor. 2015. Why are there still so many jobs? The history and future of workplace automation. Journal of economic perspectives 29, 3 (2015), 3–30.
  9. 9.Lisanne Bainbridge. 1983. Ironies of automation. In Analysis, design and evaluation of man–machine systems. Elsevier, 129–135.
  10. 10.Arnold B Bakker and Evangelia Demerouti. 2007. The job demands-resources model: State of the art. Journal of managerial psychology 22, 3 (2007), 309–328.
  11. 11.Victor R Basili, Forrest Shull, and Filippo Lanubile. 1999. Building knowledge through families of experiments. IEEE transactions on software engineering 25, 4 (1999), 456–473.
  12. 12.Christian Bird, Denae Ford, Thomas Zimmermann, Nicole Forsgren, Eirini Kalliamvakou, Travis Lowdermilk, and Idan Gazit. 2022. Taking Flight with Copilot: Early insights and opportunities of AI-powered pair-programming tools. Queue 20, 6 (2022), 35–57.
  13. 13.Christian Bird, Nachiappan Nagappan, Brendan Murphy, Harald Gall, and Premkumar Devanbu. 2011. Don’t touch my code! Examining the effects of ownership on software quality. In FSE. 4–14.
  14. 14.Barry Boehm, Victor R Basili, et al. 2005. Software defect reduction top 10 list. Foundations of empirical software engineering: the legacy of Victor R. Basili 426, 37 (2005), 426–431.
  15. 15.Virginia Braun and Victoria Clark. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101.
  16. 16.Virginia Braun and Victoria Clarke. 2022. Conceptual and design thinking for thematic analysis. Qualitative Psychology 9, 1 (2022), 3.
  17. 17.Frederick Brooks and H Kugler. 1987. No silver bullet. April.
  18. 18.Adam Brown, Sarah D’Angelo, Ambar Murillo, Ciera Jaspan, and Collin Green. 2024. Identifying the factors that influence trust in AI code completion. In Proceedings of the 1st ACM International Conference on AI-Powered Software. 1–9.
  19. 19.Jenna Butler, Jina Suh, Sankeerti Haniyur, and Constance Hadley. 2025. Dear Diary: A randomized controlled trial of Generative AI coding tools in the workplace. In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 319–329.
  20. 20.Tavis S Campbell, Jillian A Johnson, and Kristin A Zernicke. 2020. Cognitive appraisal. In Encyclopedia of behavioral medicine. Springer, 486–487.
  21. 21.Ruijia Cheng, Ruotong Wang, Thomas Zimmermann, and Denae Ford. 2023. ‘It would work for me too’: How Online Communities Shape Software Developers’ Trust in AI-Powered Code Generation Tools. ACM Transactions on Interactive Intelligent Systems (2023).
  22. 22.Rudrajit Choudhuri, Carmen Badea, Christian Bird, Jenna Butler, Robert DeLine, and Brian Houck. 2026. AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work. In Proceedings of the 48th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP).
  23. 23.Rudrajit Choudhuri, Christian Bird, Carmen Badea, and Anita Sarma. 2026. To Copilot and Beyond: 22 AI Systems Developers Want Built. arXiv preprint arXiv:2604.07830 (2026).
  24. 24.Rudrajit Choudhuri, Christopher A Sanchez, Margaret Burnett, and Anita Sarma. 2026. Thinking Less, Trusting More: GenAI’s Impacts on Students’ Cognitive Habits. SSRN (2026).
  25. 25.Rudrajit Choudhuri, Bianca Trinkenreich, Rahul Pandita, Eirini Kalliamvakou, Igor Steinmacher, Marco Gerosa, Christopher Sanchez, and Anita Sarma. 2025. What Guides Our Choices? Modeling Developers’ Trust and Behavioral Intentions Towards GenAI. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, 1691–1703.
  26. 26.Rudrajit Choudhuri, Bianca Trinkenreich, Rahul Pandita, Eirini Kalliamvakou, Igor Steinmacher, Marco Gerosa, Christopher Sanchez, and Anita Sarma. 2025. What Needs Attention? Prioritizing Drivers of Developers’ Trust and Adoption of Generative AI. arXiv preprint arXiv:2505.17418 (2025).
  27. 27.Rune Haubo Bojesen Christensen. 2023. ordinal—Regression Models for Ordinal Data. https://CRAN.R-project.org/package=ordinal R package version 2023.12-4.
  28. 28.Jacob Cohen. 2013. Statistical power analysis for the behavioral sciences. Routledge.
  29. 29.Kevin Crowston and Francesco Bolici. 2025. Deskilling and upskilling with AI systems. Information Research an international electronic journal 30, iConf (2025), 1009–1023.
  30. 30.Joost CF De Winter and Dimitra Dodou. 2014. Why the Fitts list has persisted throughout the history of function allocation. Cognition, Technology & Work 16, 1 (2014), 1–11.
  31. 31.Zackary Okun Dunivin. 2024. Scalable Qualitative Coding with LLMs: Chain-of-Thought Reasoning Matches Human Performance in Some Hermeneutic Tasks. arXiv preprint arXiv:2401.15170 (2024).
  32. 32.Upol Ehsan, Samir Passi, Koustuv Saha, Todd McNutt, Mark O. Riedl, and Sara Alcorn. 2026. From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms to Foster Dignified human-AI Interaction. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). ACM. doi:10.1145/3772318.3791081
  33. 33.Mica R. Endsley and Esin O. Kiris. 1995. The Out-of-the-Loop Performance Problem and Level of Control in Automation. Human Factors 37, 2 (1995), 381–394. doi:10.1518/001872095779064555
  34. 34.Zixuan Feng, Sadia Afroz, and Anita Sarma. 2025. From Gains to Strains: Modeling Developer Burnout with GenAI Adoption. arXiv preprint arXiv:2510.07435 (2025).
  35. 35.Zixuan Feng, Reed Milewicz, Emerson Murphy-Hill, Tyler Menezes, Alexander Serebrenik, Igor Steinmacher, and Anita Sarma. 2025. Charting Uncertain Waters: A Socio-Technical Roadmap for Sustaining Open Source Communities in the Age of GenAI. ACM Transactions on Software Engineering and Methodology (2025). doi:10.1145/3789210
  36. 36.Paul M Fitts. 1951. Human engineering for an effective air-navigation and traffic-control system. (1951).
  37. 37.Bent Flyvbjerg. 2006. Five misunderstandings about case-study research. Qualitative inquiry 12, 2 (2006), 219–245.
  38. 38.Yitzhak Fried and Gerald R Ferris. 1987. The validity of the job characteristics model: A review and meta-analysis. Personnel psychology 40, 2 (1987), 287–322.
  39. 39.Dwight D. Frink and Richard J. Klimoski. 1998. Toward a Theory of Accountability in Organizations and Human Resources Management. In Research in Personnel and Human Resources Management, Gerald R. Ferris (Ed.). Vol. 16. JAI Press, Greenwich, CT, 1–51.
  40. 40.Andrew Gelman and Jennifer Hill. 2007. Data analysis using regression and multilevel/hierarchical models. Cambridge university press.
  41. 41.Amir Ghorbani, Nathan Cassee, Derek Robinson, Adam Alami, Neil A Ernst, Alexander Serebrenik, and Andrzej Wąsowski. 2023. Autonomy is an acquired taste: Exploring developer preferences for GitHub bots. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1405–1417.
  42. 42.GitHub. 2024. Copilot. https://github.com/features/copilot.
  43. 43.Kilem L Gwet. 2014. Handbook of inter-rater reliability: The definitive guide to measuring the extent of agreement among raters. Advanced Analytics, LLC.
  44. 44.J Richard Hackman and Greg R Oldham. 1976. Motivation through the design of work: Test of a theory. Organizational behavior and human performance 16, 2 (1976), 250–279.
  45. 45.Joseph F Hair. 2009. Multivariate data analysis. (2009).
  46. 46.Angela T Hall, Dwight D Frink, and M Ronald Buckley. 2017. An accountability account: A review and synthesis of the theoretical and empirical research on felt accountability. Journal of Organizational Behavior 38, 2 (2017), 204–224.
  47. 47.Donald Hedeker and Robert D Gibbons. 1994. A random-effects ordinal regression model for multilevel analysis. Biometrics (1994), 933–944.
  48. 48.Angjelin Hila and Elliott Hauser. 2025. Assessing the Reliability of Large Language Models for Deductive Qualitative Coding: A Comparative Study of ChatGPT Interventions. arXiv preprint arXiv:2507.14384 (2025).
  49. 49.Stephen E Humphrey, Jennifer D Nahrgang, and Frederick P Morgeson. 2007. Integrating motivational, social, and contextual work design features: a meta-analytic summary and theoretical extension of the work design literature. Journal of applied psychology 92, 5 (2007), 1332.
  50. 50.Brittany Johnson, Christian Bird, Denae Ford, Nicole Forsgren, and Thomas Zimmermann. 2023. Make Your Tools Sparkle with Trust: The PICSE Framework for Trust in Software Tools. In 2023 IEEE/ACM 45th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 409–419.
  51. 51.William A Kahn. 1990. Psychological conditions of personal engagement and disengagement at work. Academy of management journal 33, 4 (1990), 692–724.
  52. 52.Mansi Khemka and Brian Houck. 2024. Toward Effective AI Support for Developers: A survey of desires and concerns. Commun. ACM 67, 11 (2024), 42–49.
  53. 53.Barbara A Kitchenham and Shari L Pfleeger. 2008. Personal opinion surveys. In Guide to advanced empirical software engineering. Springer, 63–92.
  54. 54.Richard Koestner, Natasha Lekes, Theodore A Powers, and Emanuel Chicoine. 2002. Attaining personal goals: self-concordance plus implementation intentions equals success. Journal of personality and social psychology 83, 1 (2002), 231.
  55. 55.Sukrit Kumar, Drishti Goel, Thomas Zimmermann, Brian Houck, Balasubramanyan Ashok, and Chetan Bansal. 2025. Time Warp: The Gap Between Developers’ Ideal vs Actual Workweeks in an AI-Driven Era. In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 12–22.
  56. 56.Stefano Lambiase, Gemma Catolino, Fabio Palomba, Filomena Ferrucci, and Daniel Russo. 2025. Exploring Individual Factors in the Adoption of LLMs for Specific Software Engineering Tasks. arXiv preprint arXiv:2504.02553 (2025).
  57. 57.Richard S Lazarus. 1991. Emotion and adaptation. Oxford University Press.
  58. 58.John D Lee and Katrina A See. 2004. Trust in automation: Designing for appropriate reliance. Human factors 46, 1 (2004), 50–80.
  59. 59.Jeffery A. LePine, Nathan P. Podsakoff, and Marcie A. LePine. 2005. A Meta-Analytic Test of the Challenge Stressor–Hindrance Stressor Framework: An Explanation for Inconsistent Relationships Among Stressors and Performance. Academy of Management Journal 48, 5 (2005), 764–775. doi:10.5465/amj.2005.18803921
  60. 60.Jennifer S Lerner and Philip E Tetlock. 1999. Accounting for the effects of accountability. Psychological bulletin 125, 2 (1999), 255.
  61. 61.Zijian Li, Luzhen Tang, Mengyu Xia, Xinyu Li, Naping Chen, Dragan Gašević, and Yizhou Fan. 2026. When LLMs Fall Short in Deductive Coding: Model Comparisons and Human-AI Collaboration Workflow Design. In Proceedings of the 16th International Conference on Learning Analytics and Knowledge (LAK ’26). arXiv:2512.21041.
  62. 62.Hannah Limerick, David Coyle, and James W Moore. 2014. The experience of agency in human-computer interactions: a review. Frontiers in human neuroscience 8 (2014), 643.
  63. 63.Marjolein Lips-Wiersma and Lani Morris. 2009. Discriminating between ‘meaningful work’and the ‘management of meaning’. Journal of business ethics 88, 3 (2009), 491–511.
  64. 64.Brian Lubars and Chenhao Tan. 2019. Ask not what AI can do, but what AI should do: Towards a framework of task delegability. Advances in neural information processing systems 32 (2019).
  65. 65.Russell A Matthews, Laura Pineault, and Yeong-Hyun Hong. 2022. Normalizing the use of single-item measures: Validation of the single-item compendium for organizational psychology. Journal of Business and Psychology 37, 4 (2022), 639–673.
  66. 66.Richard D. McKelvey and William Zavoina. 1975. A statistical model for the analysis of ordinal level dependent variables. Journal of Mathematical Sociology 4, 1 (1975), 103–120. doi:10.1080/0022250X.1975.9989847
  67. 67.John P Meyer and Natalie J Allen. 1991. A three-component conceptualization of organizational commitment. Human resource management review 1, 1 (1991), 61–89.
  68. 68.Courtney Miller, Rudrajit Choudhuri, Mara Ulloa, Sankeerti Haniyur, Robert DeLine, Margaret-Anne Storey, Emerson Murphy-Hill, Christian Bird, and Jenna L Butler. 2025. ‘Maybe We Need Some More Examples:’ Individual and Team Drivers of Developer GenAI Tool Use. arXiv preprint arXiv:2507.21280 (2025).
  69. 69.Shinichi Nakagawa and Holger Schielzeth. 2013. A general and simple method for obtaining R2 from generalized linear mixed-effects models. Methods in Ecology and Evolution 4, 2 (2013), 133–142. doi:10.1111/j.2041-210x.2012.00261.x
  70. 70.Raja Parasuraman and Dietrich H Manzey. 2010. Complacency and bias in human use of automation: An attentional integration. Human factors 52, 3 (2010), 381–410.
  71. 71.Raja Parasuraman, Thomas B Sheridan, and Christopher D Wickens. 2000. A model for types and levels of human interaction with automation. IEEE Transactions on systems, man, and cybernetics-Part A: Systems and Humans 30, 3 (2000), 286–297.
  72. 72.Guilherme Vaz Pereira, Victoria Jackson, Rafael Prikladnicki, André van der Hoek, Luciane Fortes, Carolina Araújo, André Coelho, Ligia Chelli, and Diego Ramos. 2025. Exploring GenAI in Software Development: Insights from a Case Study in a Large Brazilian Company. In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 330–341.
  73. 73.Qualtrics. 2026. Qualtrics Survey Platform. https://www.qualtrics.com. Accessed June 11, 2026.
  74. 74.Foyzur Rahman and Premkumar Devanbu. 2011. Ownership, experience and defects: a fine-grained study of authorship. In Proceedings of the 33rd international conference on software engineering. 491–500.
  75. 75.Ira J Roseman and Craig A Smith. 2001. Appraisal theory. Appraisal processes in emotion: Theory, methods, research (2001), 3–19.
  76. 76.Daniel Russo. 2024. Navigating the complexity of generative AI adoption in software engineering. ACM Transactions on Software Engineering and Methodology (2024).
  77. 77.Richard M Ryan and Edward L Deci. 2000. Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American psychologist 55, 1 (2000), 68.
  78. 78.Yijia Shao, Humishka Zope, Yucheng Jiang, Jiaxin Pei, David Nguyen, Erik Brynjolfsson, and Diyi Yang. 2025. Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the US Workforce. arXiv preprint arXiv:2506.06576 (2025).
  79. 79.Thomas B Sheridan and William L Verplank. 1978. Human and computer control of undersea teleoperators. (1978).
  80. 80.Klaas-Jan Stol and Brian Fitzgerald. 2018. The ABC of software engineering research. ACM TOSEM 27, 3 (2018).
  81. 81.Margaret-Anne Storey. 2026. From technical debt to cognitive and intent debt: Rethinking software health in the age of AI. arXiv preprint arXiv:2603.22106 (2026).
  82. 82.Margaret-Anne Storey, Thomas Zimmermann, Christian Bird, Jacek Czerwonka, Brendan Murphy, and Eirini Kalliamvakou. 2019. Towards a theory of software developer job satisfaction and perceived productivity. IEEE Transactions on Software Engineering 47, 10 (2019), 2125–2142.
  83. 83.Jobin Alexander Strunk, Leonardo Banh, Anika Nissen, Gero Strobel, and Stefan Smolnik. 2024. To delegate or not to delegate? Factors influencing human-agentic IS interaction. (2024).
  84. 84.Philip E Tetlock. 1983. Accountability and complexity of thought. Journal of personality and social psychology 45, 1 (1983), 74.
  85. 85.David Thissen, Lynne Steinberg, and Daniel Kuang. 2002. Quick and easy implementation of the Benjamini-Hochberg procedure for controlling the false positive rate in multiple comparisons. Journal of educational and behavioral statistics 27, 1 (2002), 77–83.
  86. 86.Bianca Trinkenreich, Fabio Santos, and Klaas-jan Stol. 2024. Predicting attrition among software professionals: Antecedents and consequences of burnout and engagement. ACM Transactions on Software Engineering and Methodology 33, 8 (2024), 1–45.
  87. 87.Mara Ulloa, Jenna L Butler, Sankeerti Haniyur, Courtney Miller, Barrett Amos, Advait Sarkar, and Margaret-Anne Storey. 2025. Product Manager Practices for Delegating Work to Generative AI:" Accountability must not be delegated to non-human actors". arXiv preprint arXiv:2510.02504 (2025).
  88. 88.Ruotong Wang, Ruijia Cheng, Denae Ford, and Thomas Zimmermann. 2024. Investigating and designing for trust in AI-powered code generation tools. In Proceedings of the 2024 ACM conference on fairness, accountability, and transparency. 1475–1493.
  89. 89.Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, and Ran Wang. 2026. Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias. arXiv preprint arXiv:2604.02923 (2026).
  90. 90.Minge Xie, Douglas G Simpson, and Raymond J Carroll. 2000. Random effects in censored ordinal regression: Latent structure and Bayesian approach. Biometrics 56, 2 (2000), 376–383.

Citation

MLA
Choudhuri, R., et al. “You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy”. arXiv, 2026, http://arxiv.org/abs/2607.00533v2.
APA
Choudhuri, R., Bird, C., Badea, C., Gerosa, M., & Sarma, A. (2026). You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy. arXiv. http://arxiv.org/abs/2607.00533v2
Chicago
Choudhuri, R., C. Bird, C. Badea, M. Gerosa, and A. Sarma. 2026. “You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy”. arXiv. http://arxiv.org/abs/2607.00533v2.
Harvard
Choudhuri, R. et al. (2026) “You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2607.00533v2.
Vancouver
1. Choudhuri R, Bird C, Badea C, Gerosa M, Sarma A (2026) You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy. arXiv

BibTeX

@article{choudhuri2026you,
  title = {You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy},
  author = {Choudhuri, Rudrajit and Bird, Christian and Badea, Carmen and Gerosa, Marco and Sarma, Anita},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2607.00533v2},
  eprint = {2607.00533}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/