To Copilot and Beyond: 22 AI Systems Developers Want Built

Rudrajit ChoudhuriChristian BirdCarmen BadeaAnita Sarma

article2026arXiv5 citations

Identifies 22 AI systems software engineers want built beyond code generation, using a survey of 860 developers to establish the principle of bounded delegation for designing tools that offload peripheral tasks without intruding on core engineering craft.

Listen

Software developers spend approximately one-tenth of their working hours actively writing code, yet the majority of commercial artificial intelligence tooling focuses heavily on code generation. This imbalance accelerates the creation of code while expanding downstream bottlenecks in review, testing, incident triage, and maintenance. As a result, engineering teams face growing technical debt, review fatigue, and poorly understood systems.

The article aims to identify the specific artificial intelligence systems that software developers actually want built across their broader workflow and to define the explicit behavioral boundaries and conditions required for these tools to be accepted in practice.

To investigate this, the study analyzed open-ended survey responses from 860 Microsoft software developers across global business units, roles, and geographies. The research team applied a reflexive qualitative analysis method using an independent multi-model council comprising three frontier language models from different providers (GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6), followed by human reconciliation and rigorous inter-rater reliability checks (achieving an average Krippendorff's alpha of 0.94).

The analysis identified 22 distinct AI systems across five engineering categories: Development, Design and Planning, Quality and Risk Management, Infrastructure and Operations, and Meta-Work. Across these systems, the primary findings show that developer demand is concentrated on the verification and maintenance side of development, with 50.1% of development respondents requesting tools to systematically address technical debt and 44.5% of quality respondents seeking change-aware test generation. Developers strongly rejected autonomous decision-making or direct production modifications, consistently asserting that tools should prepare drafts, compile evidence, and surface alternatives rather than approve changes, deploy code, or communicate directly with stakeholders.

These findings highlight an operating pattern termed bounded delegation: developers want artificial intelligence to absorb mechanical assembly and information retrieval tasks, but strictly preserve human agency and accountability over architectural choices and final evaluations. Unchecked code generation creates a right-shift problem by pushing verification burdens downstream. To prevent developer burnout and safeguard system integrity, organizations must shift quality signals earlier into the authoring process, embedding testing, security, and context mapping at the point of change.

For engineering and technology leaders, the article recommends prioritizing tooling that supports earlier defect detection, evidence gathering, and documentation synchronization over raw code-generation speed. Tool designers and organizations should enforce four mandatory guardrails across all AI systems: explicit authority scoping to halt tools before critical decisions, clear data provenance linking outputs to authoritative sources, proactive uncertainty signaling when confidence is low, and strict least-privilege access ensuring production environments remain read-only for AI.

Because the study reflects cross-sectional self-reported needs from engineers within a single large enterprise, caution is warranted when generalizing to smaller organizations, open-source projects, or differently regulated settings without further validation. Nonetheless, the high internal agreement and large sample size provide strong confidence that sustainable productivity gains depend on respecting developer agency and designing tools around bounded delegation.

arXiv: 2604.07830

No sufficiently relevant recommendations were found.

Cover for To Copilot and Beyond: 22 AI Systems Developers Want Built

Abstract

Developers spend roughly one-tenth of their workday writing code, yet most AI tooling targets that fraction. This paper asks what should be built for the rest. We surveyed 860 Microsoft developers to understand where they want AI support, and where they want it to stay out. Using a human-in-the-loop, multi-model council-based thematic analysis, we identify 22 AI systems that developers want built across five task categories. For each, we describe the problem it solves, what makes it hard to build, and the constraints developers place on its behavior. Our findings point to a growing right-shift burden in AI-assisted development: developers wanted systems that embed quality signals earlier in their workflow to keep pace with accelerating code generation, while enforcing explicit authority scoping, provenance, uncertainty signaling, and least-privilege access throughout. This tension reveals a pattern we call "bounded delegation": developers wanted AI to absorb the assembly work surrounding their craft, never the craft itself. That boundary tracks where they locate professional identity, suggesting that the value of AI tooling may lie as much in where and how precisely it stops as in what it does.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Source Survey
  • 3.1.1 Survey Design
  • 3.1.2 Data collection
  • 3.2 Data Analysis
  • 3.2.1 Stage 1: Independent theme discovery.
  • 3.2.2 Stage 2: Codebook reconciliation.
  • 3.2.3 Stage 3: Author review and codebook approval
  • 3.2.4 Stage 4: Systematic coding.
  • 3.2.5 Stage 5: Inter-rater reliability and consensus.
  • 3.3 Limitations
  • 4 22 Systems Developers Want Built
  • 4.1 Development (N=353)
  • 4.1.1 Scoped-PR builder for tech debt removal
  • 4.1.2 Embedded quality gate for code and tests
  • 4.1.3 Trace-to-diff root cause workbench
  • 4.1.4 Repository context graph for cross-file changes
  • 4.2 Design and Planning (N=223)
  • 4.2.1 Design-To-sprint workbench
  • 4.2.2 Architecture studio for requirements-to-design
  • 4.2.3 Design analyzer for trade-off and risk
  • 4.2.4 Decision context and provenance graph
  • 4.2.5 Full-loop design doc and diagram workspace
  • 4.3 Quality and Risk Management (N=155)
  • 4.3.1 Change-aware test generation and quality gates
  • 4.3.2 Context-aware pull request review assistant
  • 4.3.3 Pre-merge security advisor with patch suggestions
  • 4.3.4 Compliance evidence compiler and interpreter
  • 4.3.5 Change risk radar for proactive regression warning
  • 4.4 Infrastructure and Operations (N=101)
  • 4.4.1 Telemetry correlation assistant for alert tuning and incident triage
  • 4.4.2 CI/CD and infrastructure-as-code blueprint builder
  • 4.4.3 Maintenance backlog prioritizer
  • 4.4.4 Customer support triage assistant
  • 4.5 Meta-Work: Documentation, Knowledge, Collaboration (N=157)
  • 4.5.1 DocSync: continuous documentation synchronizer
  • 4.5.2 Contextualized developer ramp-up coach
  • 4.5.3 Stakeholder communication drafting workbench
  • 4.5.4 Interactive exploration board for tech discovery
  • 5 Discussion
  • 5.1 Implications for practice
  • 5.1.1 Bounded Delegation: Protecting the Parts Worth Doing.
  • 5.1.2 The Right-Shift Problem
  • 5.1.3 Cross-Cutting Guardrails
  • 5.2 Implications for research
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Bounded delegation shifts AI toward assembly and verification, not developer judgment

    empirical result

    Across the systems developers requested, a recurring boundary separated work they were willing to delegate from work they wanted to retain. Developers wanted AI to take on tedious, context-heavy assembly—such as gathering evidence, structuring plans, drafting artifacts, and surfacing risks—while preserving human ownership of consequential decisions, craft, and accountability. The authors call this pattern “bounded delegation.” It appeared even for tasks respondents thought AI could plausibly perform, so the boundary was not solely a response to current model limitations; the authors interpret it as also reflecting developers’ agency and professional identity.

    The requested systems also pointed to a right-shift burden in AI-assisted development: faster code generation increases the importance and volume of downstream verification. Developers therefore wanted quality signals closer to the point of authorship or change, including test-gap detection, risk analysis, and review support. Their preferred direction was to scale verification alongside generation, not to give AI final authority over the work.

  2. Knowl 2 — Four recurring guardrails define acceptable AI behavior

    empirical result

    Across the 22 requested systems, developers repeatedly placed four constraints on acceptable AI behavior: (1) explicit authority scoping—the system must stay within a declared task boundary and stop when the work requires human judgment; (2) provenance—outputs must expose the sources or evidence supporting them; (3) uncertainty signaling—the system must reveal missing evidence or low confidence rather than guess confidently; and (4) least-privilege access, including strict limits on access to sensitive data and production systems. The authors treat these recurring constraints as design requirements for bounded delegation: delegated work should remain visible, attributable, and interruptible.

  3. Knowl 3 — Development systems target technical debt, early quality feedback, debugging evidence, and repository context

    empirical result

    Among 353 developers who gave substantive open-ended responses in the Development category (816 completed that category block), four requested systems addressed recurring development work. The reported percentages are category response shares; responses could receive multiple codes, so the percentages are not additive.

    • Scoped-PR builder for tech debt removal (50.1%). Developers wanted repository-aware refactoring and maintenance, including migrations, renames, and dependency work, packaged as small, reviewable diffs that follow local conventions. The system should respect explicit scope, avoid unapproved structural changes, and stop when required domain or tribal knowledge is unavailable.
    • Embedded quality gate for code and tests (27.8%). Developers wanted diff-anchored feedback during authorship, combining deterministic checks and targeted model reasoning to find defects, standards violations, vulnerabilities, or missing tests. Suggestions should require explicit human approval; the system should ask for clarification rather than silently modify code or guess intent.
    • Trace-to-diff root cause workbench (19.5%). Developers wanted a tool that starts from an incident, alert, or stack trace and assembles correlated logs, traces, a likely regression window, historical analogues, and competing root-cause hypotheses. It could propose a narrowly scoped patch and regression test after verification, but must state uncertainty, use isolated reproduction by default, and apply no code without human review.
    • Repository context graph for cross-file changes (18.4%). Developers wanted a persistent map of source, tests, and code structure enriched with rationale from discussions, design records, and linked bugs, so the system can report change impact before drafting edits. It should expose missing context rather than invent it; all diffs require human review, and substantial or out-of-scope restructuring requires approval.
  4. Knowl 4 — Quality and risk systems should detect and explain issues without certifying outcomes

    empirical result

    Among 155 developers who gave substantive open-ended responses in Quality & Risk Management (401 completed that category block), five requested systems focused on earlier, change-specific quality and risk signals. Reported percentages are category response shares and are not additive because responses could receive multiple codes.

    • Change-aware test generation and quality gates (44.5%). Build an impact map from a diff, recover intended behavior from requirements and existing tests, generate repository-style tests, run them, and report changed behavior that remains untested. The system must not choose validation criteria, commit tests automatically, or certify its own output; it should provide auditable evidence and request input when intent is unclear.
    • Context-aware pull request review assistant (23.2%). Summarize a pull request’s intent, affected modules, dependencies, and downstream impact; attach repo-aware findings to specific files or lines with rationale, severity, and a summary of what was and was not checked. It must not replace human review, approve or reject a pull request, or change code without approval; it should make uncertainty explicit and stop rather than infer ambiguous intent.
    • Pre-merge security advisor with patch suggestions (21.9%). Triage security findings against the changed code, dependencies, and configurations; provide evidence and conservative remediation options for review. It must not auto-commit or merge fixes, act as the sole security authority, expose sensitive information, or conceal low confidence.
    • Compliance evidence compiler and interpreter (14.2%). Translate selected policies into applicability rules, gather required artifacts into a provenance-bearing evidence ledger, and identify missing evidence. It should not handle raw customer or user data, make compliance attestations without human approval, or change production systems.
    • Change risk radar for proactive regression warning (11.6%). Correlate deployment or configuration changes with service-specific baselines, telemetry deviations, and historical incidents, then explain risk drivers in a brief. It must remain read-only and auditable and must not make ship/no-ship decisions, accept risk, or autonomously escalate incidents.
  5. Knowl 5 — Design and planning systems should structure exploration while leaving decisions with people

    empirical result

    Among 223 developers who gave substantive open-ended responses in Design & Planning (548 completed that category block), five requested systems to reduce the overhead around design while preserving human decision authority. Reported percentages are category response shares and are not additive because responses could receive multiple codes.

    • Design-to-sprint workbench (30.5%). Turn an approved plan into hierarchical work items, dependencies, estimate ranges based on comparable historical tasks, and a draft sprint sequence. Human approval is required before tracker changes; the system should not autonomously replan, assign people, contact colleagues, or invent missing requirements.
    • Architecture studio for requirements-to-design (27.8%). Maintain explicit design state—goals, constraints, open questions, and exploration history—and generate materially different architecture options grounded in local records and infrastructure context. It should ask for missing information and track option rationale, but must not settle the final design or provide an unsteerable, context-free design.
    • Design analyzer for trade-off and risk (20.6%). Review components, data flows, trust boundaries, requirements, and non-functional requirements; identify coverage gaps and hidden dependencies, and generate evidence-linked trade-offs or what-if scenarios. It must not make the final design decision or offer unsupported critiques.
    • Decision context and provenance graph (19.7%). Connect time-stamped decisions, alternatives, owners, and rationale across artifacts, linking earlier decisions to downstream consequences and exposing source freshness. It must distinguish evidence from guesses, surface temporal provenance, respect sensitive-data boundaries, and leave decisions to people.
    • Full-loop design doc and diagram workbench (15.2%). Transform selected notes, transcripts, outlines, and code changes into editable design documents and diagrams, with source context and targeted regeneration of sections. It must not silently assume missing requirements or make final design decisions; outputs should remain steerable and verifiable.
  6. Knowl 6 — Meta-work systems address documentation drift, onboarding, communication, and early exploration

    empirical result

    Among 157 developers who gave substantive open-ended responses in Meta-Work (532 completed that category block), four requested systems to reduce coordination and information-assembly overhead. Reported percentages are category response shares and are not additive because responses could receive multiple codes.

    • DocSync: Continuous Documentation Synchronizer (45.9%). Detect which documentation may have drifted after code changes and propose targeted updates grounded in authoritative code, tests, or schemas. Updates must remain drafts until approved; the system should not fabricate facts or present inferred rationale as established fact, and should leave placeholders or request input when sources are missing.
    • Contextualized Developer Ramp-Up Coach (30.6%). Use a learner’s goals, skill level, and constraints to create a repository-specific code map, reading path, and hands-on exercises that adapt as the learner progresses. Guidance should be cited and uncertainty-aware, require human review, and support rather than replace mentoring and team integration.
    • Stakeholder Communication Draft Workbench (12.7%). Draft technical updates from selected artifacts according to user-specified audience, purpose, technical fluency, language, and formality while retaining important nuance. Messages must not be sent automatically; sensitive or relationship-critical communication should remain with the human, with strong privacy boundaries for external recipients.
    • Interactive Exploration Board for Tech Discovery (11.7%). Start from a research question and user-defined constraints and evaluation criteria, then present multiple options with assumptions and local context. Users should be able to reweight criteria, add options, or challenge the set before committing. The system should be pull-based, must not decide or treat its first option set as exhaustive, and must mark claims as uncertain for human validation.
  7. Knowl 7 — Infrastructure and operations systems assemble operational context but remain read-only

    empirical result

    Among 101 developers who gave substantive open-ended responses in Infrastructure & Operations (283 completed that category block), four requested systems to reduce toil in incidents, pipelines, maintenance, and support. Reported percentages are category response shares and are not additive because responses could receive multiple codes.

    • Telemetry correlation assistant for alert tuning and incident triage (40.6%). Assemble correlated telemetry, related-service signals, recent changes, and similar incidents into a structured brief; use historical data to suggest alert-threshold, deduplication, and coverage improvements. It should support human judgment, not perform production remediation, rollbacks, or security-policy changes without explicit oversight.
    • CI/CD and infrastructure-as-code blueprint builder (33.7%). Explain repository-specific delivery workflows, generate organization-compliant pipeline or infrastructure skeletons, and diagnose build or deployment failures from configuration and logs. Outputs must be reproducible and inspectable; the system must have no live write permission and must not deploy or approve changes automatically.
    • Maintenance backlog prioritizer (16.8%). Consolidate deprecation notices, security findings, runtime drift, platform notices, and cost anomalies into a deduplicated, evidence-backed queue prioritized by severity, impact, and effort. It should be read-only and must not autonomously patch, upgrade, or change production security configurations.
    • Customer support triage assistant (11.9%). Organize and classify incoming tickets, redact sensitive fields, use approved telemetry and knowledge sources to retrieve similar cases, and draft a triage card and customer-safe reply for human review. It must not send messages autonomously or become the forced primary customer interface; human escalation must remain immediate, and customer data and internal artifacts must be protected.
  8. Knowl 8 — Human-in-the-loop, multi-model analysis produced the system catalog

    model/method

    The open-ended survey responses were analyzed in separate opportunity and constraint tracks using a human-in-the-loop thematic-analysis pipeline. GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 independently proposed themes from responses in each task category, with supporting participant IDs. GPT-5.2 reconciled overlapping themes; a theme proposed by only one model was retained only when supported by at least three distinct respondents. Two researchers then reviewed the evidence and theme boundaries, negotiated revisions, and approved the codebooks before coding.

    All three models coded responses against the approved codebooks, with a rationale for each assignment. Each response was paired with the participant’s answer to the complementary opportunity or constraint question in the same category to clarify underspecified answers. Models could assign multiple approved codes but could not invent new substantive codes. Per-theme inter-rater reliability ranged from Krippendorff’s α=0.81\alpha=0.81 to 0.970.97 (mean 0.940.94); final assignments required agreement from at least two of the three models. Responses flagged with an ISSUE_* data-quality code by at least two models were excluded from analysis, and researchers spot-checked assignments. The researchers then synthesized related themes into the 22 system descriptions.

  9. Knowl 9 — The catalog comes from 860 Microsoft developers’ stated preferences

    experimental setup

    The study analyzed open-ended answers from an IRB-approved survey conducted in July 2025. The survey was emailed to 8,000 Microsoft developers; 1,193 responses were received, and 860 were retained after excluding incomplete responses, patterned responses, attention-check failures, and participants reporting no AI experience. Participants selected two or three software-engineering task categories and answered where they wanted AI support and what they did not want AI to handle. The analysis covered 2,580 response sets across Development (816), Design & Planning (548), Meta-Work (532), Quality & Risk Management (401), and Infrastructure & Operations (283). The smaller category sample sizes reported in the system catalog count respondents who wrote substantive open-ended responses. The resulting findings describe stated preferences in one large organization, not observed tool adoption or behavior.

  10. Knowl 10 — The findings are limited to stated needs in a cross-sectional, single-organization study

    limitation

    The survey measured developers’ stated preferences at one point in time; it does not establish that respondents would adopt or continue using the requested systems in practice. The Microsoft sample spans multiple roles, domains, and global sites, but it does not establish that the systems or their relative importance generalize to smaller organizations, open-source communities, or regulated industries. Self-selection may also mean respondents held stronger views about AI support than the broader developer population. Finally, model-assisted qualitative coding could identify plausible themes that do not fully reflect participant intent. The authors mitigated this risk through independent discovery by models from different providers, researcher review of supporting responses, approved-codebook constraints, and spot checks, but note that replication in other contexts remains necessary.

Coverage note — All 22 requested systems are included in the five category knowls; detailed participant quotations and secondary methodological reflections were omitted because they add illustration rather than distinct system requirements or findings.

References

  1. 1.[n. d.]. Supplemental Package. https://cabird.github.io/22-systems-devs-want/.
  2. 2.Sadia Afroz, Zixuan Feng, Katie Kimura, Bianca Trinkenreich, Igor Steinmacher, and Anita Sarma. 2025. Developer Productivity with GenAI. arXiv preprint arXiv:2510.24265 (2025).
  3. 3.Blake A Allan, Cassondra Batz-Barbarich, Haley M Sterling, and Louis Tay. 2019. Outcomes of meaningful work: A meta-analysis. Journal of management studies 56, 3 (2019), 500–528.
  4. 4.David Autor. 2022. The labor market impacts of technological change: From unbridled enthusiasm to qualified optimism to vast uncertainty. Technical Report. National Bureau of Economic Research.
  5. 5.Catherine Bailey, Ruth Yeoman, Adrian Madden, Marc Thompson, and Gary Kerridge. 2019. A review of the empirical literature on meaningful work: Progress and research agenda. Human Resource Development Review 18, 1 (2019), 83–113.
  6. 6.Sebastian Baltes, Marc Cheong, and Christoph Treude. 2026. " An Endless Stream of AI Slop": The Growing Burden of AI-Assisted Software Development. arXiv preprint arXiv:2603.27249 (2026).
  7. 7.Leonardo Banh, Florian Holldack, and Gero Strobel. 2025. Copiloting the future: How generative AI transforms Software Engineering. Information and Software Technology 183 (2025), 107751.
  8. 8.Christian Bird, Denae Ford, Thomas Zimmermann, Nicole Forsgren, Eirini Kalliamvakou, Travis Lowdermilk, and Idan Gazit. 2022. Taking Flight with Copilot: Early insights and opportunities of AI-powered pair-programming tools. Queue 20, 6 (2022), 35–57.
  9. 9.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2021).
  10. 10.Virginia Braun and Victoria Clark. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101.
  11. 11.Virginia Braun and Victoria Clarke. 2022. Conceptual and design thinking for thematic analysis. Qualitative Psychology 9, 1 (2022), 3.
  12. 12.Erik Brynjolfsson. 2022. The turing trap: The promise & peril of human-like artificial intelligence. Daedalus 151, 2 (2022), 272–287.
  13. 13.Jenna Butler, Jina Suh, Sankeerti Haniyur, and Constance Hadley. 2025. Dear Diary: A randomized controlled trial of Generative AI coding tools in the workplace. In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 319–329.
  14. 14.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021).
  15. 15.Ruijia Cheng, Ruotong Wang, Thomas Zimmermann, and Denae Ford. 2023. “It would work for me too”: How Online Communities Shape Software Developers’ Trust in AI-Powered Code Generation Tools. ACM Transactions on Interactive Intelligent Systems (2023).
  16. 16.Rudrajit Choudhuri, Carmen Badea, Christian Bird, Jenna Butler, Rob DeLine, and Brian Houck. 2025. AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work. arXiv preprint arXiv:2510.00762 (2025).
  17. 17.Rudrajit Choudhuri, Bianca Trinkenreich, Rahul Pandita, Eirini Kalliamvakou, Igor Steinmacher, Marco Gerosa, Christopher Sanchez, and Anita Sarma. 2025. What Guides Our Choices? Modeling Developers’ Trust and Behavioral Intentions Towards GenAI. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, 1691–1703.
  18. 18.Rudrajit Choudhuri, Bianca Trinkenreich, Rahul Pandita, Eirini Kalliamvakou, Igor Steinmacher, Marco Gerosa, Christopher Sanchez, and Anita Sarma. 2025. What Needs Attention? Prioritizing Drivers of Developers’ Trust and Adoption of Generative AI. arXiv preprint arXiv:2505.17418 (2025).
  19. 19.John W Creswell and Cheryl N Poth. 2016. Qualitative inquiry and research design: Choosing among five approaches. Sage publications.
  20. 20.Zixuan Feng, Sadia Afroz, and Anita Sarma. 2025. From Gains to Strains: Modeling Developer Burnout with GenAI Adoption. arXiv preprint arXiv:2510.07435 (2025).
  21. 21.Bent Flyvbjerg. 2006. Five misunderstandings about case-study research. Qualitative inquiry 12, 2 (2006), 219–245.
  22. 22.Natasa Gisev, J Simon Bell, and Timothy F Chen. 2013. Interrater agreement and interrater reliability: key concepts, approaches, and applications. Research in Social and Administrative Pharmacy 9, 3 (2013), 330–338.
  23. 23.Kilem L Gwet. 2014. Handbook of inter-rater reliability: The definitive guide to measuring the extent of agreement among raters. Advanced Analytics, LLC.
  24. 24.J Richard Hackman and Greg R Oldham. 1976. Motivation through the design of work: Test of a theory. Organizational behavior and human performance 16, 2 (1976), 250–279.
  25. 25.Brittany Johnson, Christian Bird, Denae Ford, Ebtesam Al Haque, Nicole Forsgren, and Thomas Zimmermann. [n. d.]. Facilitating Trust in AI-assisted Software Tools. ACM Transactions on Software Engineering and Methodology ([n. d.]).
  26. 26.Brittany Johnson, Christian Bird, Denae Ford, Nicole Forsgren, and Thomas Zimmermann. 2023. Make Your Tools Sparkle with Trust: The PICSE Framework for Trust in Software Tools. In 2023 IEEE/ACM 45th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 409–419.
  27. 27.Eirini Kalliamvakou. 2024. A developer’s second brain: Reducing complexity through partnership with AI.
  28. 28.Mansi Khemka and Brian Houck. 2024. Toward Effective AI Support for Developers: A survey of desires and concerns. Commun. ACM 67, 11 (2024), 42–49.
  29. 29.Sukrit Kumar, Drishti Goel, Thomas Zimmermann, Brian Houck, B Ashok, and Chetan Bansal. 2025. Time Warp: The Gap Between Developers’ Ideal vs Actual Workweeks in an AI-Driven Era. arXiv preprint arXiv:2502.15287 (2025).
  30. 30.Adam Kuper. 2004. The social science encyclopedia. Routledge.
  31. 31.Stefano Lambiase, Gemma Catolino, Fabio Palomba, Filomena Ferrucci, and Daniel Russo. 2025. Exploring Individual Factors in the Adoption of LLMs for Specific Software Engineering Tasks. arXiv preprint arXiv:2504.02553 (2025).
  32. 32.Richard S Lazarus. 1991. Emotion and adaptation. Oxford University Press.
  33. 33.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110 (2022).
  34. 34.Yue Liu, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, and David Lo. 2026. Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild. arXiv preprint arXiv:2603.28592 (2026).
  35. 35.André N Meyer, Earl T Barr, Christian Bird, and Thomas Zimmermann. 2019. Today was a good day: The daily life of software developers. IEEE Transactions on Software Engineering 47, 5 (2019), 863–880.
  36. 36.Courtney Miller, Rudrajit Choudhuri, Mara Ulloa, Sankeerti Haniyur, Robert DeLine, Margaret-Anne Storey, Emerson Murphy-Hill, Christian Bird, and Jenna L Butler. 2025. " Maybe We Need Some More Examples:" Individual and Team Drivers of Developer GenAI Tool Use. arXiv preprint arXiv:2507.21280 (2025).
  37. 37.Kate Niederhoffer, Gabriella Rosen Kellerman, Angela Lee, Alex Liebscher, Kristina Rapuano, and Jeffrey T Hancock. 2025. AI-generated “workslop” is destroying productivity. Harvard Business Review (2025).
  38. 38.Ike Obi, Jenna Butler, Sankeerti Haniyur, Brian Hassan, Margaret-Anne Storey, and Brendan Murphy. 2025. Identifying factors contributing to “bad days” for software developers: A mixed-methods study. In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 1–11.
  39. 39.Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri. 2022. Asleep at the keyboard? assessing the security of GitHub copilot’s code contributions. In 2022 IEEE Symposium on Security and Privacy (SP). IEEE, 754–768.
  40. 40.Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri. 2025. Asleep at the keyboard? assessing the security of github copilot’s code contributions. Commun. ACM 68, 2 (2025), 96–105.
  41. 41.Guilherme Vaz Pereira, Victoria Jackson, Rafael Prikladnicki, André van der Hoek, Luciane Fortes, Carolina Araújo, André Coelho, Ligia Chelli, and Diego Ramos. 2025. Exploring GenAI in Software Development: Insights from a Case Study in a Large Brazilian Company. In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 330–341.
  42. 42.Teade Punter, Marcus Ciolkowski, Bernd Freimut, and Isabel John. 2003. Conducting on-line surveys in software engineering. In 2003 International Symposium on Empirical Software Engineering, 2003. ISESE 2003. Proceedings. IEEE, 80–88.
  43. 43.Ira J Roseman and Craig A Smith. 2001. Appraisal theory. Appraisal processes in emotion: Theory, methods, research (2001), 3–19.
  44. 44.Daniel Russo. 2024. Navigating the complexity of generative AI adoption in software engineering. ACM Transactions on Software Engineering and Methodology (2024).
  45. 45.Hope Schroeder, Marianne Aubin Le Quéré, Casey Randazzo, David Mimno, and Sarita Schoenebeck. 2025. Large language models in qualitative research: uses, tensions, and intentions. In Proceedings of the 2025 chi conference on human factors in computing systems. 1–17.
  46. 46.Ben Shneiderman. 2020. Human-centered artificial intelligence: Reliable, safe & trustworthy. International Journal of Human–Computer Interaction 36, 6 (2020), 495–504.
  47. 47.Margaret-Anne Storey, Thomas Zimmermann, Christian Bird, Jacek Czerwonka, Brendan Murphy, and Eirini Kalliamvakou. 2019. Towards a theory of software developer job satisfaction and perceived productivity. IEEE Transactions on Software Engineering 47, 10 (2019), 2125–2142.
  48. 48.Robert H Tai, Lillian R Bentley, Xin Xia, Jason M Sitt, Sarah C Fankhauser, Ana M Chicas-Mosier, and Barnas G Monteith. 2024. An examination of the use of large language models to aid analysis of textual data. International Journal of Qualitative Methods 23 (2024), 16094069241231168.
  49. 49.Eric Lansdown Trist and Kenneth W Bamforth. 1951. Some social and psychological consequences of the longwall method of coal-getting: An examination of the psychological situation and defences of a work group in relation to the social structure and technological content of the work system. Human relations 4, 1 (1951), 3–38.
  50. 50.Ruotong Wang, Ruijia Cheng, Denae Ford, and Thomas Zimmermann. 2023. Investigating and designing for trust in AI-powered code generation tools. arXiv preprint arXiv:2305.11248 (2023).
  51. 51.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837.
  52. 52.Shuai Wu, Xue Li, Yanna Feng, Yufang Li, and Zhijun Wang. 2026. Council Mode: Mitigating Hallucination and Bias in LLMs via Multi-Agent Consensus. arXiv preprint arXiv:2604.02923 (2026).
  53. 53.Kazuma Yamasaki, Joseph Ayobami Joshua, Tasha Settewong, Mahmoud Alfadel, Kazumasa Shimari, and Kenichi Matsumoto. 2026. Who Writes the Docs in SE 3.0? Agent vs. Human Documentation Pull Requests. arXiv preprint arXiv:2601.20171 (2026).
  54. 54.Albert Ziegler, Eirini Kalliamvakou, X Alice Li, Andrew Rice, Devon Rifkin, Shawn Simister, Ganesh Sittampalam, and Edward Aftandilian. 2024. Measuring GitHub Copilot’s Impact on Productivity. Commun. ACM (2024).

Citation

MLA
Choudhuri, R., et al. “To Copilot and Beyond: 22 AI Systems Developers Want Built”. arXiv, 2026, http://arxiv.org/abs/2604.07830v1.
APA
Choudhuri, R., Bird, C., Badea, C., & Sarma, A. (2026). To Copilot and Beyond: 22 AI Systems Developers Want Built. arXiv. http://arxiv.org/abs/2604.07830v1
Chicago
Choudhuri, R., C. Bird, C. Badea, and A. Sarma. 2026. “To Copilot and Beyond: 22 AI Systems Developers Want Built”. arXiv. http://arxiv.org/abs/2604.07830v1.
Harvard
Choudhuri, R. et al. (2026) “To Copilot and Beyond: 22 AI Systems Developers Want Built”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.07830v1.
Vancouver
1. Choudhuri R, Bird C, Badea C, Sarma A (2026) To Copilot and Beyond: 22 AI Systems Developers Want Built. arXiv

BibTeX

@article{choudhuri2026copilot,
  title = {To Copilot and Beyond: 22 AI Systems Developers Want Built},
  author = {Choudhuri, Rudrajit and Bird, Christian and Badea, Carmen and Sarma, Anita},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.07830v1},
  eprint = {2604.07830}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/