AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work

Rudrajit ChoudhuriCarmen BadeaChristian BirdJenna ButlerRobert DeLIneBrian Houck

article2025ICSE-SEIP6 citationsFirst Runner Up, IEEE Software Best Paper Award

Presents an empirical model derived from 860 developers that uses cognitive appraisal theory to identify where practitioners want generative AI assistance, where they resist automation, and how responsible AI priorities shift across software engineering tasks.

Listen

As generative artificial intelligence tools rapidly enter software engineering workflows, technology leaders face an ongoing challenge: tool capabilities are advancing quickly, yet organizations lack empirical clarity on where developers actually want assistance, where they prefer to retain manual control, and how to govern these tools responsibly. This lack of alignment risks investing in the wrong automation priorities and undermining meaningful developer work.

The main objective of the article is to establish an empirically validated framework showing how developers evaluate their daily tasks, how these evaluations predict their openness to and use of artificial intelligence, and which Responsible AI principles they prioritize across different types of software work.

To evaluate these dynamics, the researchers conducted a large-scale, mixed-methods survey in July 2025 across 860 software developers at Microsoft spanning global regions, roles, and experience levels. Grounded in psychological cognitive appraisal theory, the study gathered quantitative ratings across a comprehensive taxonomy of software engineering tasks, forced-choice trade-off rankings across eight Responsible AI principles, and qualitative rationales across thousands of open-ended responses. The data was analyzed using mixed-effects regression, hierarchical clustering, and reflexive thematic analysis.

The analysis produced several key findings. First, task appraisals strongly predict adoption: perceived task value, accountability, and cognitive demands each positively correlate with openness to and use of artificial intelligence, whereas professional identity exhibits a dual dynamic—reducing general openness to delegating work while increasing usage when tools complement personal craft. Second, tasks divide into distinct operational clusters: core development work (e.g., coding, testing, debugging) exhibits high demand for tool improvement, operational toil (e.g., continuous integration, environment maintenance) shows high demand for new automation, and interpersonal work (e.g., mentoring, stakeholder relations) meets strong resistance to automation. Third, Responsible AI priorities follow a clear hierarchy: across all tasks, baseline priorities focus heavily on Reliability & Safety (selected by 85% of respondents), Privacy & Security (77%), and Transparency (72%), whereas principles like Fairness (32%) and Inclusiveness (32%) become primary only in human-facing or design contexts. Finally, experience and individual traits matter: senior engineers and technophiles strongly prioritize steerability and control over direct execution.

These findings indicate that developers view artificial intelligence as a collaborator rather than an outright replacement. In high-stakes and system-critical tasks, automation errors carry high remediation costs, meaning developers insist on verifiability, provenance, and final sign-off rather than end-to-end automation. Conversely, for identity-defining and relational work, excessive automation risks deskilling engineers and damaging team culture.

For engineering leaders and tool designers, the article recommends prioritizing artificial intelligence development around cognitive augmentation rather than blunt replacement. Organizations should focus automation investments on reducing operational toil and assisting with boilerplate steps in core work, while keeping human oversight, suggest-only workflows, and reversible edits standard. Interpersonal activities like mentoring and direct communication should remain fundamentally human-led. Further research and internal piloting should focus on designing better observability logs, drift-prevention mechanisms, and adaptive autonomy controls.

The conclusions are drawn from a cross-sectional study within a single large enterprise (Microsoft), meaning findings capture self-reported perceptions and correlational patterns that may not fully generalize to open-source communities or smaller firms. Nonetheless, the substantial sample size and rigorous multi-method design provide high confidence in the identified appraisal drivers and design trade-offs.

arXiv: 2510.00762
  • Paper: Generative AI, Stefan Feuerriegel et al. (2023). Its technical and organizational framework for generative AI clarifies the capabilities, probabilistic limits, and sociotechnical risks that the source maps onto developers’ daily software-engineering work.
  • Paper: On the Opportunities and Risks of Foundation Models, Rishi Bommasani et al. (2021). Its account of foundation-model capabilities, homogenization, and inherited risks provides the conceptual basis for the source’s context-specific Responsible AI priorities.
  • Paper: Large Language Models for Software Engineering: A Systematic Literature Review, Xinying Hou et al. (2023). Its systematic map of LLM applications across software-engineering tasks supplies the task landscape that the source refines through developer appraisals and adoption patterns.
  • Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). Its taxonomy of LLM-agent architecture and evaluation prepares the reader to understand the source’s distinctions between AI assistance, autonomy, and human control across development work.
  • Paper: Human-in-the-loop or AI-in-the-loop? Automate or Collaborate?, Sriraam Natarajan et al. (2025). Its distinction between human-in-the-loop and AI-in-the-loop systems establishes the control and authority concepts underlying the source’s analysis of responsible AI priorities.
Cover for AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work

Abstract

Generative AI is reshaping software work, yet we lack clear guidance on where developers most need support and how to design it responsibly. We report a large-scale, mixed-methods study of N=860 developers examining where, why, and how they seek or limit AI help across SE tasks. Using cognitive appraisal theory, we provide the first empirically validated mapping of developers' task appraisals to AI adoption patterns and Responsible AI (RAI) priorities. Appraisals predict AI openness and use, revealing distinct patterns: strong current use and demand for improvement in core work (e.g., coding, testing); high demand to reduce toil (e.g., documentation, operations); and clear limits for identity- and relationship-centric work (e.g., mentoring). RAI priorities vary by context: reliability and security for systems-facing tasks; transparency, alignment, and steerability to maintain control; and fairness and inclusiveness for human-facing work. Our results offer concrete, contextual guidance for delivering AI where it matters to developers and their work.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Appraisal Foundations & Hypotheses
  • 4 Method
  • 4.1 Study Design
  • 4.2 Data Collection
  • 4.3 Data Analysis
  • 5 Results
  • 5.1 RQ1a: How do appraisals shape AI adoption?
  • 5.2 RQ1b: Where and why do developers seek or limit AI support?
  • 5.2.1 Core work (C1)
  • 5.2.2 People & AI-Building (C2)
  • 5.2.3 Ops & Coordination (C3)
  • 5.3 RQ2: Which RAI design principles do developers prioritize in AI for SE tasks?
  • 5.3.1 How do priorities vary across task categories?
  • 5.3.2 How do priorities vary by experience/AI dispositions?
  • 6 Discussion
  • 6.1 Implications for practice
  • 6.2 Implications for research
  • 6.3 Limitations
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Mixed-Effects Regression of Cognitive Task Appraisals on AI Openness and Reported AI Usage

    empirical result

    To examine how cognitive task appraisals shape developer attitudes and behaviors toward generative AI tools, linear mixed-effects regressions were estimated on data from N=860N=860 software developers across 10,44910{,}449 task evaluations. The models included task appraisals (Value, Identity, Accountability, Demands) as fixed effects, controls for developer software engineering (SE) experience and AI experience, and random effects for within-participant and across-task type dependence. All Variance Inflation Factors (VIFs) were <2< 2, and pp-values were adjusted using the Benjamini–Hochberg False Discovery Rate (FDR) procedure.

    Factor Openness to AI support Reported AI usage
    β\beta dd β\beta dd
    Value (H1H_1) .12∗∗∗.12^{***} .16.16 .16∗∗∗.16^{***} .18.18
    Identity (H2H_2) −.09∗∗∗-.09^{***} −.15-.15 .15∗∗∗.15^{***} .20.20
    Accountability (H3H_3) .07∗∗∗.07^{***} .10.10 .18∗∗∗.18^{***} .21.21
    Demand (H4H_4) .12∗∗∗.12^{***} .18.18 .09∗∗∗.09^{***} .10.10
    SE Experience – – −.09∗∗∗-.09^{***} −.13-.13
    AI Experience .19∗∗∗.19^{***} .27.27 .41∗∗∗.41^{***} .46.46
    Rm2/Rc2R^2_m / R^2_c .25/.45.25 / .45 .25/.48.25 / .48
    Observations 10,44910{,}449 10,44910{,}449

    Note: ∗p<.05^{*}p < .05, ∗∗p<.01^{**}p < .01, ∗∗∗p<.001^{***}p < .001. Effect size dd thresholds: d<0.02d < 0.02 no effect, d∈[0.02,0.15)d \in [0.02, 0.15) small, d∈[0.15,0.35)d \in [0.15, 0.35) medium, d>0.35d > 0.35 large. Blank cells denote non-significant associations after FDR correction. Rm2R^2_m represents marginal variance explained by fixed effects; Rc2R^2_c represents conditional variance explained by fixed and random effects.

    Key empirical findings:

    • Perceived task value significantly increases openness (β=.12,p<.001,d=.16\beta = .12, p < .001, d = .16) and reported usage (β=.16,p<.001,d=.18\beta = .16, p < .001, d = .18).
    • Perceived identity alignment exhibits a dual effect: it reduces openness to AI support (β=−.09,p<.001,d=−.15\beta = -.09, p < .001, d = -.15) while increasing reported usage (β=.15,p<.001,d=.20\beta = .15, p < .001, d = .20), indicating that developers protect ownership of identity-defining tasks but strategically use AI to augment their craft.
    • Felt accountability positively predicts openness (β=.07,p<.001,d=.10\beta = .07, p < .001, d = .10) and usage (β=.18,p<.001,d=.21\beta = .18, p < .001, d = .21). High-accountability tasks drive developers to seek external validation while maintaining strict final oversight.
    • Task demands positively predict openness (β=.12,p<.001,d=.18\beta = .12, p < .001, d = .18) and usage (β=.09,p<.001,d=.10\beta = .09, p < .001, d = .10), reflecting the use of AI to offload cognitive burden.
    • Greater SE experience is associated with lower reported AI usage (β=−.09,p<.001,d=−.13\beta = -.09, p < .001, d = -.13), whereas prior AI experience strongly increases both openness (β=.19,p<.001,d=.27\beta = .19, p < .001, d = .27) and usage (β=.41,p<.001,d=.46\beta = .41, p < .001, d = .46).
    • Moderation analyses show that risk-tolerant developers seek more AI support for high-value (Δβ=.06,p=.035\Delta\beta = .06, p = .035) and high-demand tasks (Δβ=.09,p=.001\Delta\beta = .09, p = .001). Accountability predicts openness and use only among high-technophiles (β=0.07\beta = 0.07 and eta = 0.20, p<.001p < .001), but not among low-technophiles.
  2. Knowl 2 — Logistic GLMMs Predicting Developer Responsible AI Design Priorities Across SE Task Categories

    data/table

    To determine how developers prioritize Responsible AI (RAI) design principles across different software engineering task contexts, logistic Generalized Linear Mixed-effects Models (GLMMs) were estimated across N=860N=860 participants. In the survey elicitation, participants selected their top 5 out of 8 RAI principles for each task category under a forced-choice format. Each GLMM predicted the binary outcome of whether a given principle was selected (11) or not (00) as a function of task category, mean-centered developer SE experience, AI experience, risk tolerance, and technophilic motivations, including participant random intercepts.

    The baseline constant represents the odds of selecting the principle in development-heavy work for a developer with average experience and dispositions. Odds Ratios (OR\text{OR}) indicate relative multiplicative shifts in odds compared to the baseline.

    Factor Reliability Privacy Transp. Goal Steer. AI Fairness Inclus.
    Safety Security Maint. Account.
    Constant (Base Odds) 18.15∗∗∗18.15^{***} 8.19∗∗∗8.19^{***} 5.17∗∗∗5.17^{***} 4.68∗∗∗4.68^{***} 3.65∗∗∗3.65^{***} 3.03∗∗∗3.03^{***} 0.21∗∗∗0.21^{***} 0.20∗∗∗0.20^{***}
    Task Categories
    Design Planning 0.49∗∗0.49^{**} – – 1.45∗1.45^{*} – – 1.48∗1.48^{*} 1.61∗∗1.61^{**}
    Quality Risk Mgmt – 1.91∗∗1.91^{**} – 0.61∗∗0.61^{**} – – 1.07∗∗1.07^{**} –
    Infra Operations – 1.38∗1.38^{*} – – – – – –
    Meta-work – – – – – – 3.06∗∗∗3.06^{***} 2.49∗∗∗2.49^{***}
    Dispositions (Centered)
    SE Experience 1.15∗1.15^{*} – – – 1.21∗1.21^{*} – – –
    AI Experience – – 1.30∗1.30^{*} – 1.11∗1.11^{*} – – –
    Risk Tolerance – – – – 1.13∗∗1.13^{**} – – –
    Technophilic Motivations – – – 1.16∗∗1.16^{**} 1.28∗1.28^{*} – – –
    Rm2/Rc2R^2_m / R^2_c 0.06/0.320.06 / 0.32 0.03/0.310.03 / 0.31 0.05/0.290.05 / 0.29 0.07/0.310.07 / 0.31 0.04/0.370.04 / 0.37 0.04/0.430.04 / 0.43 0.07/0.440.07 / 0.44 0.03/0.430.03 / 0.43

    Note: ∗p<.05^{*}p < .05, ∗∗p<.01^{**}p < .01, ∗∗∗p<.001^{***}p < .001 after Benjamini–Hochberg False Discovery Rate (FDR) adjustment. Blank cells denote odds not significantly different from baseline (OR=1\text{OR} = 1). Rm2R^2_m and Rc2R^2_c report marginal and conditional pseudo-R2R^2 values.

    Key takeaways from the logistic regression models:

    • Base odds in development work prioritize operational safety and control: Reliability & Safety (base odds 18.1518.15, ≈95%\approx 95\% probability), Privacy & Security (base odds 8.198.19, ≈89%\approx 89\%), Transparency (base odds 5.175.17, ≈84%\approx 84\%), Goal Maintenance (base odds 4.684.68, ≈82%\approx 82\%), Steerability (base odds 3.653.65, ≈79%\approx 79\%), and AI Accountability (base odds 3.033.03, ≈75%\approx 75\%), whereas Fairness (base odds 0.210.21) and Inclusiveness (base odds 0.200.20) have low base odds in systems-facing coding.
    • Privacy & Security is significantly more prioritized in Quality & Risk Management (OR=1.91,p<.01\text{OR} = 1.91, p < .01) and Infrastructure & Operations (OR=1.38,p<.05\text{OR} = 1.38, p < .05).
    • Fairness and Inclusiveness are substantially elevated in human-facing meta-work (OR=3.06\text{OR} = 3.06 and 2.49,p<.0012.49, p < .001) and Design & Planning (OR=1.48\text{OR} = 1.48 and 1.611.61).
    • In Design & Planning, developers downweight Reliability & Safety (OR=0.49,p<.01\text{OR} = 0.49, p < .01) to encourage creative ideation while increasing demand for Goal Maintenance (OR=1.45,p<.05\text{OR} = 1.45, p < .05).
    • Steerability priority increases consistently with higher SE experience (OR=1.21\text{OR} = 1.21), AI experience (OR=1.11\text{OR} = 1.11), risk tolerance (OR=1.13\text{OR} = 1.13), and technophilia (OR=1.28\text{OR} = 1.28).
  3. Knowl 3 — Task Appraisal Profiles and Three-Cluster Categorization of Software Engineering Work

    theoretical result

    Software engineering tasks cluster into three distinct groups based on developers' cognitive appraisals across four theoretical drivers: Value, Identity alignment, Felt Accountability, and Task Demands.

    The task clustering methodology:

    1. For each task, top-2 agreement proportions (percentage of respondents selecting 4 or 5 on a 5-point Likert scale) were computed for Value, Identity, Accountability, and Demand.
    2. Proportions were standardized to z-scores (z=(x−xˉ)/sd(x)z = (x - \bar{x})/\text{sd}(x)) across tasks.
    3. Agglomerative hierarchical clustering with Ward linkage was performed on the Euclidean distance matrix of these standardized appraisal scores with precision-weighting (inverse-variance shrinkage) to account for task sample sizes, validated via stratified bootstrap (B=1000B=1000). Silhouette analysis confirmed k=3k=3 as the optimal number of clusters.

    The three emergent task clusters and their appraisal signatures:

    Task Name Value (%) Identity (%) Accountability (%) Demand (%)
    Cluster 1: Core Work
    Coding / Programming 98.0 97.0 94.9 77.1
    System Design 97.9 91.3 87.6 90.7
    Testing Quality Assurance 96.9 59.0 84.4 86.8
    Bug Fixing / Debugging 96.8 73.4 92.3 85.7
    Code Review / Pull Requests 96.7 76.6 86.1 74.0
    Requirements Engineering 95.8 66.5 74.7 88.8
    Security Compliance 95.5 51.6 82.3 82.5
    Research Brainstorming 90.9 88.9 75.5 86.8
    Performance Optimization 87.8 75.0 75.9 89.5
    Learning 83.9 93.5 72.6 81.0
    Cluster 2: People AI Building
    Mentoring Onboarding 68.8 70.0 62.2 67.1
    AI Integration 67.4 75.0 59.9 71.5
    Cluster 3: Ops Coordination
    DevOps (CI/CD) 93.4 50.8 74.6 74.2
    Infrastructure Monitoring 91.1 48.3 74.9 80.4
    Planning Management 90.3 51.9 67.3 79.3
    Refactoring Maintenance 86.4 57.4 80.7 76.3
    Env. Setup Maintenance 85.7 41.9 64.8 72.4
    Documentation 85.4 31.4 68.8 79.3
    Customer Support 83.8 40.6 67.5 84.3
    Stakeholder Communication 80.5 35.0 61.9 67.8

    Cluster descriptions:

    • Cluster 1 (Core Work): High value (83.9%83.9\%--98.0%98.0\%) and high demands (74.0%74.0\%--90.7%90.7\%); moderate-to-high accountability (72.6%72.6\%--94.9%94.9\%); moderate-to-strong identity alignment (51.6%51.6\%--97.0%97.0\%). These tasks form the technical essence of development, quality assurance, and technical skill acquisition.
    • Cluster 2 (People & AI Building): Moderate value (67.4%67.4\%--68.8%68.8\%), moderate demands (67.1%67.1\%--71.5%71.5\%), and moderate accountability (59.9%59.9\%--62.2%62.2\%); strong identity alignment (70.0%70.0\%--75.0%75.0\%). Tasks reflect intrinsically motivating, craft- or relationship-centered activities.
    • Cluster 3 (Ops & Coordination): Moderate-to-high value (80.5%80.5\%--93.4%93.4\%), demands (67.8%67.8\%--84.3%84.3\%), and accountability (61.9%61.9\%--80.7%80.7\%); distinctly low identity alignment (31.4%31.4\%--57.4%57.4\%). Tasks encompass infrastructure toil ("run-the-systems") and relational administrative overhead.
  4. Knowl 4 — Openness-to-Support vs. AI Usage Quadrant Mapping Framework for SE Tooling

    model/method

    To identify unmet needs and tool-adoption bottlenecks across software engineering tasks, tasks are mapped onto a two-dimensional plane comparing standardized Openness to AI Support (xx-axis, representing developer demand/need) against standardized Reported AI Usage (yy-axis).

    Coordinates are computed as task-level z-scores (z=(x−xˉ)/sd(x)z = (x - \bar{x})/\text{sd}(x)), and divided into four quadrants via a mean split (z=0z=0):

    1. Build Quadrant (High Need, Low Use; Bottom-Right: x>0,y<0x > 0, y < 0):
      • Characteristics: Developers express strong desire for AI assistance, but current adoption remains low due to tooling gaps, integration friction, or trust concerns.
      • Strategy: Develop new prototypes, lower workflow friction, and enhance grounding and safety.
      • Tasks located here: DevOps (CI/CD), Environment Setup & Maintenance, Infrastructure Monitoring, Security & Compliance, Documentation, Customer Support, Performance Optimization.
    2. Improve Quadrant (High Need, High Use; Top-Right: x>0,y>0x > 0, y > 0):
      • Characteristics: High developer demand aligned with high active adoption; core development areas where AI is already integrated into the daily workflow.
      • Strategy: Focus on precision, reliability, context-awareness, and output quality to achieve compounding productivity gains.
      • Tasks located here: Coding/Programming, Bug Fixing/Debugging, Testing & QA, Code Review/Pull Requests, Refactoring & Maintenance, Learning, Research & Brainstorming.
    3. Sustain Quadrant (Low Need, High Use; Top-Left: x<0,y>0x < 0, y > 0):
      • Characteristics: Tools are used out of convenience or default workflow features, but perceived necessity and openness are below average.
      • Strategy: Maintain existing capabilities without substantial new capital or engineering over-investment.
      • Tasks located here: AI Integration.
    4. De-prioritize Quadrant (Low Need, Low Use; Bottom-Left: x<0,y<0x < 0, y < 0):
      • Characteristics: Low reported usage and low openness; tasks involving high contextual ambiguity, deep strategic domain knowledge, or relational/interpersonal empathy.
      • Strategy: Avoid heavy automation investments; keep tools backstage or restricted to peripheral administrative assistance.
      • Tasks located here: System Design, Requirements Engineering, Mentoring & Onboarding, Stakeholder Communication, Planning & Management.
  5. Knowl 5 — Cognitive Appraisal Theoretical Framework for Developer AI Support and Delegation

    theoretical result

    Drawing on cognitive appraisal theory and work-design models, developer willingness to delegate tasks to generative AI and their actual AI tool usage are governed by four primary psychological appraisal drivers:

    1. Task Value (H1H_1): The perceived importance of a task to project success, organizational outcomes, or personal goals.
      • Theoretical Mechanism: Higher task value heightens developer focus and motivation. Developers seek AI to boost efficiency and accelerate tedious components of valuable tasks, while retaining ultimate decision control rather than replacing human agency.
      • Hypothesis: Higher task value increases developers' openness to AI support and increases AI usage (H1H_1).
    2. Task Identity Alignment (H2H_2): The extent to which a task aligns with a developer's professional self-concept, core expertise, and intrinsic motivation.
      • Theoretical Mechanism: Tasks defining professional identity foster craft ownership, generating resistance to full AI delegation to avoid deskilling. However, developers leverage AI as an intellectual scaffold to refine and expand their craft.
      • Hypothesis: Higher task identity reduces openness to AI support, but increases usage when AI complements expertise (H2H_2).
    3. Task Felt Accountability (H3H_3): The perceived social, organizational, or reputational blame and responsibility associated with a task's outcome.
      • Theoretical Mechanism: Under anticipated evaluation and high risk of failure, developers use AI as an informational safeguard and verification check ("second set of eyes"), but lower automation bias by demanding strict oversight and veto authority.
      • Hypothesis: Higher task accountability increases developers' openness to AI support and usage, accompanied by insistence on decision control (H3H_3).
    4. Task Demands (H4H_4): The contextual cognitive effort, complexity, and mental workload imposed by the task.
      • Theoretical Mechanism: High cognitive load strains coping resources, prompting developers to offload rote, boiler-plate, or repetitive cognitive steps to AI to preserve mental bandwidth for creative problem-solving.
      • Hypothesis: Higher task demands increase developers' openness to AI support and usage (H4H_4).
  6. Knowl 6 — Grounded Taxonomy of Software Engineering Tasks and Responsible AI Principles

    experimental setup

    To investigate task-level AI adoption and ethical requirements, a multi-source grounded taxonomy of software engineering (SE) tasks and an expanded Responsible AI (RAI) framework were synthesized:

    1. Taxonomy of SE Tasks (5 Categories, 18 Tasks):

      • Development: Coding/Programming, Bug Fixing/Debugging, Performance Optimization, Refactoring & Maintenance/Updates, AI Integration.
      • Design & Planning: System Design, Requirements Engineering, Project Planning & Management.
      • Quality & Risk Management: Testing & Quality Assurance, Code Review/Pull Requests, Security & Compliance.
      • Infrastructure & Operations: DevOps (CI/CD), Environment Setup & Maintenance, Infrastructure Monitoring, Customer Support.
      • Meta-work (Collaboration / Knowledge work): Documentation, Client/Stakeholder Communication, Mentoring & Onboarding, Learning, Research & Brainstorming.
    2. Responsible AI (RAI) Principles Set (8 Principles):

      • Reliability & Safety: System outputs are accurate, dependable, robust against errors, and safe for production environments.
      • Privacy & Security: Safeguarding confidential code, proprietary data, and sensitive organizational information from unauthorized leaks or extraction.
      • Transparency: Making the AI model's internal reasoning, data sources, transformations, and confidence levels inspectable and explainable to users.
      • Goal Maintenance: The AI's sustained alignment with evolving user intent, contextual constraints, and project goals across interactions and multi-step sessions without semantic drift.
      • Steerability: Preserving developer autonomy and agency by offering explicit controls to interrupt, correct, redirect, or reverse AI actions.
      • AI Accountability (Provenance): Maintaining traceable provenance and auditability of generated artifacts to identify where and why decisions or errors occurred.
      • Fairness: Ensuring unbiased, equitable assessments in evaluative workflows (e.g., automated code review, PR feedback, and defect attribution).
      • Inclusiveness: Designing AI tools and generated communications that are accessible, accommodate diverse developer audiences, and reflect varied skill levels and backgrounds.
    3. Theoretical Construct Instruments:

      • Value: Job Characteristics Model.
      • Identity: Self-Determination Theory.
      • Accountability: Felt Accountability Scale.
      • Demands: Job Demands-Resources Model.
      • Openness to AI Support: Levels of Automation Framework.
      • AI Usage: Technology Acceptance Model (UTAUT).
      • Individual Dispositions (Risk Tolerance, Technophilia): Cognitive Style Facet Survey.
  7. Knowl 7 — Qualitative Rationale for AI Delegation and Boundary Limits Across Task Clusters

    empirical result

    Qualitative analysis of 1,5281{,}528 developer explanations reveals distinct psychological rationales for delegating or restricting AI assistance across the three task clusters:

    1. Core Work (Cluster 1: Coding, Debugging, Testing, Review, Learning, Research):

      • Where/Why AI is sought: Offloading tedious boilerplate ("generate boilerplate code, build configurations, test cases... which I know how to write, but I don't want to write"), proactive defect and regression detection across edge cases, and serving as a personalized adaptive tutor for skill acquisition.
      • Where/Why AI is limited: Developers strictly reject full delegation of high-stakes decisions and final approvals ("I can't fully delegate the final code review to AI—my approval puts my name on it"). Overreliance is resisted to prevent deskilling ("intellectual offloading can result in errors that eventually no one understands") and avoid hallucinations or unmaintainable technical debt. In System Design, AI is rejected due to lack of multi-system context and a tendency to output conventional rather than innovative architectures.
    2. People & AI Building (Cluster 2: AI Integration, Mentoring & Onboarding):

      • Where/Why AI is limited: Mentoring is treated as an intrinsically interpersonal, relational process essential for organizational trust, culture, and mutual professional growth ("Mentoring teaches the mentor as well... humans need to do it to grow themselves"). AI integration is reserved for human craftsmanship and deterministic pipelines rather than stochastic models. AI is permitted only in peripheral, rote administrative onboarding steps.
    3. Ops & Coordination (Cluster 3: DevOps, Setup, Monitoring, Docs, Comms, Support):

      • Where/Why AI is sought: Automating grunt operational maintenance ("run-the-systems"), CI/CD pipeline triage, dependency tracking, log and telemetry analysis across systems, and drafting meeting summaries or release notes.
      • Where/Why AI is limited: Operational tools are gated behind strict human verification ("no auto-deploys, no direct live changes... without supervision") to avoid production outages and operational intuition decay. Relational tasks (stakeholder communication, customer support) remain human-led because empathy, tone calibration, trust-building, and high-level strategic trade-offs cannot be automated without risking reputational damage.
  8. Knowl 8 — Responsible AI Principles Prioritization Dynamics Across SE Task Contexts

    empirical result

    Across 2,4532{,}453 qualitative rationales and forced-choice selections (N=860N=860 developers selecting their top 5 out of 8 principles per category), developers establish a pragmatic hierarchy of Responsible AI requirements:

    Overall selection frequencies across all software engineering categories:

    • Reliability & Safety: 85%85\% of respondents
    • Privacy & Security: 77%77\%
    • Transparency: 72%72\%
    • Goal Maintenance: 68%68\%
    • AI Accountability: 67%67\%
    • Steerability: 67%67\%
    • Fairness: 32%32\%
    • Inclusiveness: 32%32\%

    Contextual prioritization dynamics:

    1. Systems-Facing Tasks (Development, Quality & Risk, Infrastructure & Operations):
      • Developers treat Reliability & Safety and Privacy & Security as non-negotiable hard prerequisites. AI errors impose severe remediation costs ("Incorrect AI may as well be throwing spaghetti at a wall—it's more work to fix it").
      • Transparency is required to inspect model logic and verify personal accountability. Goal Maintenance, Steerability, and AI Accountability are prioritized to combat context loss and semantic drift during complex sessions.
      • Fairness becomes salient specifically in Quality & Risk Management (OR=1.07,p<.01\text{OR} = 1.07, p < .01) to prevent biased pull request evaluations or unfair defect attribution.
    2. Human-Facing and Collaborative Tasks (Meta-work, Design & Planning):
      • Fairness and Inclusiveness jump significantly in priority (OR=3.06\text{OR} = 3.06 and 2.492.49 in Meta-work; OR=1.48\text{OR} = 1.48 and 1.611.61 in Design & Planning). Developers demand that public documentation and stakeholder communications cater to diverse user bases.
    3. Ideation vs. Execution Trade-off:
      • In Design & Planning, developers intentionally relax Reliability & Safety (OR=0.49,p<.01\text{OR} = 0.49, p < .01) while increasing Goal Maintenance (OR=1.45,p<.05\text{OR} = 1.45, p < .05). Developers tolerate imperfect or exploratory suggestions during early brainstorming as long as the AI adheres to overall project goals and scaffolds creative exploration.
  9. Knowl 9 — Mixed-Methods Survey Methodology and Sampling for Developer AI Studies

    experimental setup

    The empirical study employed a cross-sectional mixed-methods survey administered via Qualtrics to software developers at Microsoft worldwide.

    1. Sampling and Power Analysis:
      • An a priori statistical power analysis in G*Power for multiple linear regression with repeated measures targeted the detection of a small effect size (d=0.05d = 0.05) at α=0.05\alpha = 0.05 with statistical power (1−β)=0.95(1 - \beta) = 0.95, indicating a minimum requirement of N=245N = 245.
      • The survey was emailed to 8,0008{,}000 developers sampled uniformly at random across product divisions, roles, and global regions in July 2025.
    2. Response Filtering and Final Cohort:
      • Total responses received: 1,1931{,}193 (gross response rate of 14.86%14.86\%).
      • Exclusions: Incomplete responses (n=152n = 152), patterned responses such as straight-lining (n=59n = 59), attention check failures (n=98n = 98), and respondents with zero AI tool experience (n=24n = 24).
      • Valid final sample: N=860N = 860 developers across 6 continents (57.4%57.4\% North America, 73.8%73.8\% men).
    3. Survey Flow and Fatigue Controls:
      • Participants selected 2–3 task categories reflecting their daily work. Meta-work was presented only if exactly 2 categories were chosen, bounding the maximum blocks per respondent to 3.
      • Psychometric validity was preserved using single-item measures for task Value, Identity, Accountability, and Demands on 5-point Likert scales (with an explicit "I'm not sure / N.A." option treated as missing data).
      • Forced-choice top-5 selection was implemented for the 8 RAI principles to prevent ceiling effects, accompanied by standardized plain-language explanations.
    4. Qualitative and Validation Procedures:
      • Free-text responses (1,5281{,}528 on AI delegation/limits and 2,4532{,}453 on RAI priorities) were analyzed via reflexive thematic analysis with negotiated consensus.
      • Member checking was conducted with 6262 respondents out of 371371 opt-ins, confirming the interpretive findings.
  10. Knowl 10 — Threats to Validity and Boundary Conditions of Developer Task Appraisals

    limitation

    The findings on developer task appraisals, AI adoption, and Responsible AI priorities are subject to specific validity threats and boundary conditions:

    1. Construct Validity:
      • Measurements relied on self-reported single-item psychometric scales and perceived AI usage rather than objective telemetry logs of IDE interactions. While single-item measures reduce fatigue, they may capture less variance than multi-item batteries.
      • Survey responses are susceptible to social desirability and self-reporting bias, mitigated through pilot testing (n=11n=11 sandbox, n=50n=50 pilot), question randomization, attention checks, and screening for straight-lining.
    2. Internal Validity and Normativity:
      • The cross-sectional design establishes correlational associations rather than causal relationships. Self-selection bias could lead developers with stronger positive or negative stances on AI to complete the survey.
      • The prioritization of Responsible AI principles represents empirical developer preferences under forced trade-offs and must not be interpreted as normative ethical guidance or policy dismissing lower-ranked principles.
    3. External Validity and Generalizability:
      • The sample was drawn entirely from professional developers within a single large technology multinational (Microsoft). While spanning multiple geographic sites, product groups, and technical domains, the findings may not fully generalize to open-source software (OSS) communities, early-stage startups, or non-enterprise development environments with different tool access and accountability structures.

Coverage note — Detailed qualitative participant quotes and minor subgroup regression breakdowns for specific demographic cuts were synthesized into overarching cluster themes and main effect models.

References

  1. 1.[n. d.]. Supplemental Package. https://zenodo.org/record/17224961.
  2. 2.Daron Acemoglu and Pascual Restrepo. 2019. Automation and new tasks: How technology displaces and reinstates labor. Journal of economic perspectives 33, 2 (2019), 3–30.
  3. 3.Blake A Allan, Cassondra Batz-Barbarich, Haley M Sterling, and Louis Tay. 2019. Outcomes of meaningful work: A meta-analysis. Journal of management studies 56, 3 (2019), 500–528.
  4. 4.Duane F Alwin and Jon A Krosnick. 1991. The reliability of survey attitude measurement: The influence of question and respondent attributes. Sociological methods & research 20, 1 (1991), 139–181.
  5. 5.Catherine Bailey, Ruth Yeoman, Adrian Madden, Marc Thompson, and Gary Kerridge. 2019. A review of the empirical literature on meaningful work: Progress and research agenda. Human Resource Development Review 18, 1 (2019), 83–113.
  6. 6.Arnold B Bakker and Evangelia Demerouti. 2007. The job demands-resources model: State of the art. Journal of managerial psychology 22, 3 (2007), 309–328.
  7. 7.Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI conference on human factors in computing systems. 1–16.
  8. 8.Christian Bird, Denae Ford, Thomas Zimmermann, Nicole Forsgren, Eirini Kalliamvakou, Travis Lowdermilk, and Idan Gazit. 2022. Taking Flight with Copilot: Early insights and opportunities of AI-powered pair-programming tools. Queue 20, 6 (2022), 35–57.
  9. 9.Norman M Bradburn, Seymour Sudman, and Brian Wansink. 2004. Asking questions: the definitive guide to questionnaire design–for market research, political polls, and social and health questionnaires. John Wiley & Sons.
  10. 10.Virginia Braun and Victoria Clark. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101.
  11. 11.Virginia Braun and Victoria Clarke. 2022. Conceptual and design thinking for thematic analysis. Qualitative Psychology 9, 1 (2022), 3.
  12. 12.Norman E Breslow and David G Clayton. 1993. Approximate inference in generalized linear mixed models. Journal of the American statistical Association 88, 421 (1993), 9–25.
  13. 13.Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021. To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-computer Interaction 5, CSCW1 (2021), 1–21.
  14. 14.Margaret Burnett, Simone Stumpf, Jamie Macbeth, Stephann Makri, Laura Beckwith, Irwin Kwan, Anicia Peters, and William Jernigan. 2016. GenderMag: A method for evaluating software’s gender inclusiveness. Interacting with Computers 28, 6 (2016), 760–787.
  15. 15.Jenna Butler, Jina Suh, Sankeerti Haniyur, and Constance Hadley. 2025. Dear Diary: A randomized controlled trial of Generative AI coding tools in the workplace. In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 319–329.
  16. 16.Tavis S Campbell, Jillian A Johnson, and Kristin A Zernicke. 2020. Cognitive appraisal. In Encyclopedia of behavioral medicine. Springer, 486–487.
  17. 17.Thomas Nixon Carver. 1924. Elements of rural economics. Ginn.
  18. 18.Stephen Cave, Claire Craig, Kanta Dihal, Sarah Dillon, Jessica Montgomery, Beth Singler, and Lindsay Taylor. 2018. Portrayals and perceptions of AI and why they matter. (2018).
  19. 19.Rudrajit Choudhuri, Bianca Trinkenreich, Rahul Pandita, Eirini Kalliamvakou, Igor Steinmacher, Marco Gerosa, Christopher Sanchez, and Anita Sarma. 2025. What Guides Our Choices? Modeling Developers’ Trust and Behavioral Intentions Towards GenAI. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, 1691–1703.
  20. 20.Rudrajit Choudhuri, Bianca Trinkenreich, Rahul Pandita, Eirini Kalliamvakou, Igor Steinmacher, Marco Gerosa, Christopher Sanchez, and Anita Sarma. 2025. What Needs Attention? Prioritizing Drivers of Developers’ Trust and Adoption of Generative AI. arXiv preprint arXiv:2505.17418 (2025).
  21. 21.Kelly Churchill and IBM Corporation. 2021. Foundations of Trustworthy AI: Governed Data and AI, AI Ethics and an Open Diverse Ecosystem. https://www.ibm.com/think/insights/trustworthy-ai-foundations. Accessed August 11, 2025.
  22. 22.Jacob Cohen. 2013. Statistical power analysis for the behavioral sciences. Routledge.
  23. 23.Microsoft Corporation. 2025. Responsible AI Principles and Approach. https://www.microsoft.com/en-us/ai/principles-and-approach. Accessed: 2025-08-11.
  24. 24.John W Creswell and Cheryl N Poth. 2016. Qualitative inquiry and research design: Choosing among five approaches. Sage publications.
  25. 25.Kevin Crowston and Francesco Bolici. 2025. Deskilling and upskilling with AI systems. Information Research an international electronic journal 30, iConf (2025), 1009–1023.
  26. 26.Université de Montréal. 2017. The Montreal Declaration for a Responsible Development of Artificial Intelligence. https://perma.cc/8LPD-JN74. Accessed August 11, 2025.
  27. 27.Edward L Deci and Richard M Ryan. 2000. The" what" and" why" of goal pursuits: Human needs and the self-determination of behavior. Psychological inquiry 11, 4 (2000), 227–268.
  28. 28.Norman K Denzin and Yvonna S Lincoln. 2011. The Sage handbook of qualitative research. sage.
  29. 29.K Anders Ericsson et al. 2006. The influence of experience and deliberate practice on the development of superior expert performance. The Cambridge handbook of expertise and expert performance 38, 685-705 (2006), 2–2.
  30. 30.S European Commission et al. 2019. Ethics guidelines for trustworthy AI. Publications Office (2019).
  31. 31.Franz Faul, Edgar Erdfelder, Axel Buchner, and Albert-Georg Lang. 2009. Statistical power analyses using G* Power 3.1: Tests for correlation and regression analyses. Behav Res Methods 41, 4 (2009), 1149–1160.
  32. 32.Bent Flyvbjerg. 2006. Five misunderstandings about case-study research. Qualitative inquiry 12, 2 (2006), 219–245.
  33. 33.Susan Folkman, Richard S. Lazarus, Christine Dunkel-Schetter, Anita DeLongis, and Rand J. Gruen. 1986. Dynamics of a Stressful Encounter: Cognitive Appraisal, Coping, and Encounter Outcomes. Journal of Personality and Social Psychology 50, 5 (1986), 992–1003. doi:10.1037/0022-3514.50.5.992
  34. 34.Yitzhak Fried and Gerald R Ferris. 1987. The validity of the job characteristics model: A review and meta-analysis. Personnel psychology 40, 2 (1987), 287–322.
  35. 35.Andrew Gelman and Jennifer Hill. 2007. Data analysis using regression and multilevel/hierarchical models. Cambridge university press.
  36. 36.Google. 2023. Artificial Intelligence at Google: Our Principles. https://ai.google/responsibility/principles/. Accessed August 11, 2025.
  37. 37.Wolfgang L Grichting. 1994. The meaning of “I Don’t Know” in opinion surveys: Indifference versus ignorance. Aust Psychol 29, 1 (1994).
  38. 38.J Richard Hackman and Greg R Oldham. 1976. Motivation through the design of work: Test of a theory. Organizational behavior and human performance 16, 2 (1976), 250–279.
  39. 39.Joseph F Hair. 2009. Multivariate data analysis. (2009).
  40. 40.Angela T Hall, Dwight D Frink, and M Ronald Buckley. 2017. An accountability account: A review and synthesis of the theoretical and empirical research on felt accountability. Journal of Organizational Behavior 38, 2 (2017), 204–224.
  41. 41.Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024. Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology 33, 8 (2024), 1–79.
  42. 42.Stephen E Humphrey, Jennifer D Nahrgang, and Frederick P Morgeson. 2007. Integrating motivational, social, and contextual work design features: a meta-analytic summary and theoretical extension of the work design literature. Journal of applied psychology 92, 5 (2007), 1332.
  43. 43.Maurice Jakesch, Zana Buçinca, Saleema Amershi, and Alexandra Olteanu. 2022. How different groups prioritize ethical values for responsible AI. In proceedings of the 2022 ACM conference on fairness, accountability, and transparency. 310–323.
  44. 44.Brittany Johnson, Christian Bird, Denae Ford, Nicole Forsgren, and Thomas Zimmermann. 2023. Make Your Tools Sparkle with Trust: The PICSE Framework for Trust in Software Tools. In 2023 IEEE/ACM 45th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 409–419.
  45. 45.William A Kahn. 1990. Psychological conditions of personal engagement and disengagement at work. Academy of management journal 33, 4 (1990), 692–724.
  46. 46.Mansi Khemka and Brian Houck. 2024. Toward Effective AI Support for Developers: A survey of desires and concerns. Commun. ACM 67, 11 (2024), 42–49.
  47. 47.Barbara A Kitchenham and Shari L Pfleeger. 2008. Personal opinion surveys. In Guide to advanced empirical software engineering. Springer, 63–92.
  48. 48.Richard Koestner, Natasha Lekes, Theodore A Powers, and Emanuel Chicoine. 2002. Attaining personal goals: self-concordance plus implementation intentions equals success. Journal of personality and social psychology 83, 1 (2002), 231.
  49. 49.Sukrit Kumar, Drishti Goel, Thomas Zimmermann, Brian Houck, B Ashok, and Chetan Bansal. 2025. Time Warp: The Gap Between Developers’ Ideal vs Actual Workweeks in an AI-Driven Era. arXiv preprint arXiv:2502.15287 (2025).
  50. 50.Adam Kuper. 2004. The social science encyclopedia. Routledge.
  51. 51.Riadh Ladhari. 2010. Developing e-service quality scales: A literature review. Journal of retailing and consumer services 17, 6 (2010), 464–477.
  52. 52.Stefano Lambiase, Gemma Catolino, Fabio Palomba, Filomena Ferrucci, and Daniel Russo. 2025. Exploring Individual Factors in the Adoption of LLMs for Specific Software Engineering Tasks. arXiv preprint arXiv:2504.02553 (2025).
  53. 53.Richard S Lazarus. 1991. Emotion and adaptation. Oxford University Press.
  54. 54.Hao-Ping Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wilson. 2025. The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. In Proceedings of the 2025 CHI conference on human factors in computing systems. 1–22.
  55. 55.Jennifer S Lerner and Philip E Tetlock. 1999. Accounting for the effects of accountability. Psychological bulletin 125, 2 (1999), 255.
  56. 56.Marjolein Lips-Wiersma, Catherine Bailey, Adrian Madden, and Lani Morris. 2022. Why we don’t talk about meaning at work. MIT Sloan Management Review 63, 4 (2022), 33–38.
  57. 57.Marjolein Lips-Wiersma and Lani Morris. 2009. Discriminating between ‘meaningful work’and the ‘management of meaning’. Journal of business ethics 88, 3 (2009), 491–511.
  58. 58.Brian Lubars and Chenhao Tan. 2019. Ask not what AI can do, but what AI should do: Towards a framework of task delegability. Advances in neural information processing systems 32 (2019).
  59. 59.Aengus Lynch, Benjamin Wright, Ethan Perez, and Evan Hubinger. 2025. Agentic Misalignment: How LLMs could be insider threats. https://www.anthropic.com/research/agentic-misalignment. Accessed August 11, 2025.
  60. 60.Russell A Matthews, Laura Pineault, and Yeong-Hyun Hong. 2022. Normalizing the use of single-item measures: Validation of the single-item compendium for organizational psychology. Journal of Business and Psychology 37, 4 (2022), 639–673.
  61. 61.McKinsey & Company. 2024. Unleashing developer productivity with generative AI. https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/unleashing-developer-productivity-with-generative-ai Accessed: September 18, 2025.
  62. 62.André N Meyer, Earl T Barr, Christian Bird, and Thomas Zimmermann. 2019. Today was a good day: The daily life of software developers. IEEE Transactions on Software Engineering 47, 5 (2019), 863–880.
  63. 63.John P Meyer and Natalie J Allen. 1991. A three-component conceptualization of organizational commitment. Human resource management review 1, 1 (1991), 61–89.
  64. 64.Microsoft. 2025. The New Future of Work. https://www.microsoft.com/en-us/research/project/the-new-future-of-work/.
  65. 65.Albert W Musschenga. 2005. Empirical ethics, context-sensitivity, and contextualism. The Journal of medicine and philosophy 30, 5 (2005), 467–490.
  66. 66.Raja Parasuraman and Dietrich H Manzey. 2010. Complacency and bias in human use of automation: An attentional integration. Human factors 52, 3 (2010), 381–410.
  67. 67.Raja Parasuraman, Thomas B Sheridan, and Christopher D Wickens. 2000. A model for types and levels of human interaction with automation. IEEE Transactions on systems, man, and cybernetics-Part A: Systems and Humans 30, 3 (2000), 286–297.
  68. 68.Guilherme Vaz Pereira, Victoria Jackson, Rafael Prikladnicki, André van der Hoek, Luciane Fortes, Carolina Araújo, André Coelho, Ligia Chelli, and Diego Ramos. 2025. Exploring GenAI in Software Development: Insights from a Case Study in a Large Brazilian Company. In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 330–341.
  69. 69.Teade Punter, Marcus Ciolkowski, Bernd Freimut, and Isabel John. 2003. Conducting on-line surveys in software engineering. In 2003 International Symposium on Empirical Software Engineering, 2003. ISESE 2003. Proceedings. IEEE, 80–88.
  70. 70.Qualtrics. 2025. Qualtrics Survey Platform. https://www.qualtrics.com. Accessed August 11, 2025.
  71. 71.Ira J Roseman and Craig A Smith. 2001. Appraisal theory. Appraisal processes in emotion: Theory, methods, research (2001), 3–19.
  72. 72.Peter J Rousseeuw. 1987. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of computational and applied mathematics 20 (1987), 53–65.
  73. 73.Daniel Russo. 2024. Navigating the complexity of generative AI adoption in software engineering. ACM Transactions on Software Engineering and Methodology (2024).
  74. 74.Richard M Ryan and Edward L Deci. 2000. Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American psychologist 55, 1 (2000), 68.
  75. 75.Yijia Shao, Humishka Zope, Yucheng Jiang, Jiaxin Pei, David Nguyen, Erik Brynjolfsson, and Diyi Yang. 2025. Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the US Workforce. arXiv preprint arXiv:2506.06576 (2025).
  76. 76.Ben Shneiderman. 2020. Human-centered artificial intelligence: Reliable, safe & trustworthy. International Journal of Human–Computer Interaction 36, 6 (2020), 495–504.
  77. 77.Stack Overflow. 2024. Stack Overflow Developer Survey 2024. https://survey.stackoverflow.co/2024/ Accessed: September 18, 2025.
  78. 78.Klaas-Jan Stol and Brian Fitzgerald. 2018. The ABC of software engineering research. ACM TOSEM 27, 3 (2018).
  79. 79.Kevin M. Storer. 2024. How gen AI affects the value of development work. https://dora.dev/research/ai/value-of-development-work/ DORA (DevOps Research and Assessment). Accessed: September 18, 2025.
  80. 80.Margaret-Anne Storey, Thomas Zimmermann, Christian Bird, Jacek Czerwonka, Brendan Murphy, and Eirini Kalliamvakou. 2019. Towards a theory of software developer job satisfaction and perceived productivity. IEEE Transactions on Software Engineering 47, 10 (2019), 2125–2142.
  81. 81.John Sweller. 1988. Cognitive load during problem solving: Effects on learning. Cognitive science 12, 2 (1988), 257–285.
  82. 82.Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. 2024. The metacognitive demands and opportunities of generative AI. In CHI Conference on Human Factors in Computing Systems. 1–24.
  83. 83.DORA Research Team. 2025. DORA Research: 2025 Overview. https://dora.dev/research/2025/. Accessed August 12, 2025.
  84. 84.Philip E Tetlock. 1983. Accountability and complexity of thought. Journal of personality and social psychology 45, 1 (1983), 74.
  85. 85.David Thissen, Lynne Steinberg, and Daniel Kuang. 2002. Quick and easy implementation of the Benjamini-Hochberg procedure for controlling the false positive rate in multiple comparisons. Journal of educational and behavioral statistics 27, 1 (2002), 77–83.
  86. 86.Bianca Trinkenreich, Fabio Santos, and Klaas-jan Stol. 2024. Predicting attrition among software professionals: Antecedents and consequences of burnout and engagement. ACM Transactions on Software Engineering and Methodology 33, 8 (2024), 1–45.
  87. 87.Bianca Trinkenreich, Klaas-Jan Stol, Igor Steinmacher, Marco A Gerosa, Anita Sarma, Marcelo Lara, Michael Feathers, Nicholas Ross, and Kevin Bishop. 2023. A Model for Understanding and Reducing Developer Burnout. In 2023 IEEE/ACM 45th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 48–60.
  88. 88.Nash Unsworth, Thomas S Redick, Gregory J Spillers, and Gene A Brewer. 2012. Variation in working memory capacity and cognitive control: Goal maintenance and microadjustments of control. Quarterly Journal of Experimental Psychology 65, 2 (2012), 326–355.
  89. 89.Priyan Vaithilingam, Elena D’Angelo, and Arto Hellas. 2022. Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models. In Proceedings of the 22nd Koli Calling International Conference on Computing Education Research. 1–11. doi:10.1145/3564721.3564728
  90. 90.Viswanath Venkatesh, Michael G Morris, Gordon B Davis, and Fred D Davis. 2003. User acceptance of information technology: Toward a unified view. MIS quarterly (2003), 425–478.
  91. 91.Viswanath Venkatesh, James YL Thong, and Xin Xu. 2012. Consumer acceptance and use of information technology: extending the unified theory of acceptance and use of technology. MIS quarterly (2012), 157–178.
  92. 92.Joe H Ward Jr. 1963. Hierarchical grouping to optimize an objective function. Journal of the American statistical association 58, 301 (1963), 236–244.

Citation

MLA
Choudhuri, R., et al. “AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work”. arXiv, 2025, http://arxiv.org/abs/2510.00762v2.
APA
Choudhuri, R., Badea, C., Bird, C., Butler, J., DeLine, R., & Houck, B. (2025). AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work. arXiv. http://arxiv.org/abs/2510.00762v2
Chicago
Choudhuri, R., C. Badea, C. Bird, J. Butler, R. DeLine, and B. Houck. 2025. “AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work”. arXiv. http://arxiv.org/abs/2510.00762v2.
Harvard
Choudhuri, R. et al. (2025) “AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2510.00762v2.
Vancouver
1. Choudhuri R, Badea C, Bird C, Butler J, DeLine R, Houck B (2025) AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work. arXiv

BibTeX

@article{choudhuri2025where,
  title = {AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work},
  author = {Choudhuri, Rudrajit and Badea, Carmen and Bird, Christian and Butler, Jenna and DeLine, Rob and Houck, Brian},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2510.00762v2},
  eprint = {2510.00762}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/