Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

Aayush KumarAvik DuttaSumit GulwaniGustavo SoaresAdvait SarkarEmerson Murphy-Hill

article2026arXiv0 citations

Demonstrates through a user study that adding a plan mode to spreadsheet AI agents reduces iterative prompt refinements and improves perceived creativity support and collaboration without compromising task outcomes.

Listen

Autonomous artificial intelligence agents are increasingly capable of executing complex end-user tasks, leading to the rapid adoption of interactive planning features—often termed "Plan Modes"—in developer tools. While software engineers rely on planning to review code structures and maintain execution control, end users working in spreadsheet environments typically approach problem solving through flexible, emergent trial and error. The article investigates whether structuring spreadsheet agent interactions through an upfront planning mode provides meaningful value to end users, evaluating its effects on information exchange, task outcomes, computational cost, and user experience.

To evaluate this capability, the researchers developed a high-fidelity prototype that incorporates clarifying questions, editable persistent plan displays, and partial execution controls within a spreadsheet environment. They tested this prototype against an immediate-execution baseline ("Act Mode") via a controlled, within-subjects study involving 24 experienced spreadsheet users. Participants completed open-ended creation tasks (such as personal budgets and schedules) and subjective data analysis tasks (such as vacation planning and movie selection), while researchers recorded behavioral logs, requirements evolution, resulting workbook artifacts, and survey measures evaluating creativity and collaboration.

The study revealed four primary findings. First, interactive planning fundamentally shifted how users communicated requirements: Plan Mode users expressed 44% of their requirements in response to clarifying questions and only 14% via post-execution edits, whereas baseline users derived over 41% of requirements from iterative refinements after the agent modified the sheet. Second, this upfront clarification reduced execution refinements (1.4 turns versus 2.2 turns) and decreased large language model computation costs by approximately 35% (averaging 11.0k tokens in Plan Mode versus 17.0k tokens in Act Mode). Third, despite differing interaction paths, the generated spreadsheets showed nearly identical feature distributions, complexity, and content across both conditions. Fourth, participants reported a distinct preference for Plan Mode across dimensions of creativity support and human-agent collaboration—particularly for open-ended creation tasks and longer workflows—even though they rarely manipulated the interactive plan user interface directly.

These findings indicate that the primary value of a planning mode in end-user tools stems from improved interaction mechanics rather than superior final workbooks. Clarifying questions efficiently extract user intent and mitigate costly computational cycles from trial-and-error edits. Furthermore, the presence of visible plan artifacts provides users with a psychological sense of oversight and control ("control without action"), bridging the gap between user intent and agent execution without requiring extensive manual corrections.

For product teams and enterprise stakeholders, the article supports integrating clarifying questions and structured planning into generative spreadsheet tools. However, systems should avoid rigid, mandatory planning steps for all workflows. Because users who preferred rapid iteration or provided detailed initial prompts found planning overly restrictive, tools should implement mixed-initiative triggers that adapt based on user prompting style, task scope, and problem complexity.

Confidence in these findings is supported by rigorous qualitative coding and high inter-rater agreement. Nonetheless, leaders should interpret the results within the study's boundaries: the evaluation utilized a modest sample size (N=24), examined short-duration tasks (15–20 minutes), and evaluated users largely unaccustomed to spreadsheet agents. Further pilot testing across broader enterprise datasets and longer project timelines is advisable before making definitive architectural commitments.

arXiv: 2607.23670
Cover for Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

Abstract

Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to develop a plan before task execution. However, it remains unclear whether the benefits of this feature translate to end-user programming environments such as spreadsheets. Since spreadsheet programmers tend to work iteratively and care less about technical correctness, upfront planning may not fit into their workflows as easily. In this paper, we build a prototype of a Plan Mode for spreadsheet programming and evaluate it against a non-planning baseline through a within-subjects user study (N=24). We found that despite similar task outcomes with both tools, using Plan Mode led to a reduction in refinement and a better perception of the tool across dimensions of creativity support and human-machine collaboration. We discuss the implications of these results for the future design of Plan Modes, and for the broader role of human-AI planning in end-user programming.

Table of Contents

  • I Introduction
  • II Related Work
  • II-A Interactive Human-AI Planning
  • II-B AI tools for Spreadsheet Programming
  • III The Plan Mode Prototype
  • III-A Illustrative Example
  • III-B Agent Instructions
  • III-C Design Features
  • III-D Tool Implementation
  • IV Methodology
  • IV-A Participants
  • IV-B Tasks
  • IV-C Protocol
  • IV-D Data Collection and Analysis
  • IV-D1 Conversation Logs
  • IV-D2 Spreadsheets Created
  • IV-D3 Post-study questionnaire
  • IV-D4 Session Recordings
  • IV-D5 Limitations
  • V Results
  • V-A RQ1: Interaction Patterns
  • V-A1 Requirements Elicitation
  • V-A2 Refinement Patterns
  • V-A3 Using the UI
  • V-B RQ2: Differences in generated spreadsheets
  • V-B1 AnalysisTasks
  • V-B2 CreationTasks
  • V-C RQ3: Perceived experiences
  • VI Discussion
  • VI-A Design Implications for Plan Mode
  • VI-B Control Without Action as Design Value
  • VI-C Tasks, Interaction Styles, and Mixed-Initiative Entry
  • VII Conclusion
  • VIII Acknowledgements
  • References
  • -A Relation of Design Features to Existing Plan Modes
  • -B Study Materials
  • -B1 Screener Survey
  • -B2 Study Protocol
  • -B3 Task Descriptions
  • -C Feature Codebook
  • -D Modifications to Post-Study Questionnaire
  • -E Additional Figures

Knowls

  1. Knowl 1 — Dual-Agent Architecture for Spreadsheet Planning and Execution

    model/method

    To provide an upfront planning workflow without tool execution conflicts, the spreadsheet agent system separates planning and execution across two distinct agents sharing a common conversational history:

    1. Plan Agent: Interacts with the user during the planning phase. It has read-only access to the spreadsheet workbook context and possesses a specialized tool to render a structured plan in a dedicated user interface (UI) pane. It lacks tools to write to or edit the spreadsheet.
    2. Act Agent: Takes over once execution is triggered. It has read-only access to the generated plan and possesses tools to modify the spreadsheet and dynamically update step statuses in the UI (Pending →\to In Progress →\to Done). It lacks tools to modify or regenerate the plan.

    The Plan Agent operates via a 3-step conversational workflow:

    1. Context Understanding: Ingest the active workbook content and full dialogue history.
    2. Clarifying Questions: If user intent is ambiguous or multiple valid implementations exist, ask clarifying questions in natural language via chat. The agent is instructed never to repeat unanswered questions; instead, it makes explicit assumptions and surfaces them in the plan.
    3. Plan Generation: Synthesize user answers into a step-by-step plan constrained to at most 77 steps, with exactly one sentence per step, omitting low-level formula implementation details to focus on high-level functional requirements.
  2. Knowl 2 — Interactive UI Affordances for Spreadsheet Plan Modes

    model/method

    A Plan Mode for agentic spreadsheet programming incorporates five specialized interaction affordances designed to balance user agency with automated execution:

    • Manipulable Persistent Plan Artifact: A persistent plan rendered in a dedicated UI pane alongside the spreadsheet. Users can directly manipulate the artifact (adding, deleting, reordering, or editing individual steps in place) or request modifications through chat.
    • Partial Execution: A granular execution control (such as an "Execute Step 1–x1\text{--}x" trigger) enabling users to run intermediate subsets of steps. This allows users to verify progressive workbook states before committing to full execution.
    • Dynamic Step-Status Tracking: Real-time visual progress indicators for each step in the plan (Pending, In Progress, Done). The In Progress state informs users when temporary incomplete spreadsheet states arise during multi-command subtasks.
    • Read-Only Planning Constraints: Strict sandbox restrictions preventing the agent from modifying the underlying spreadsheet while in Plan Mode, ensuring safe requirements exploration.
    • Explicit Mode Switching: A persistent manual toggle allowing users to enter or exit Plan Mode, with execution occurring strictly upon an explicit user action.
  3. Knowl 3 — Experimental Protocol for Evaluating Spreadsheet Planning Modes

    experimental setup

    A within-subjects laboratory study (N=24N = 24) evaluated the effect of a Plan Mode prototype against a non-planning agent baseline ('Act Mode') for end-user spreadsheet programming.

    • Participants: 2424 adult spreadsheet users (1313 men, 1111 women) who use spreadsheets at least weekly, have prior experience with AI tools, and possess intermediate-to-expert spreadsheet skills (excluding formula beginners).
    • LLM Backbone: Both modes utilized Claude Sonnet 4.5.
    • Tasks: Each participant completed two 1515-minute tasks under balanced task-order and mode assignments within one of two categories:
      • Creation Tasks: Authoring workbooks from scratch (Personal Budget tracking or Personal Schedule management).
      • Analysis Tasks: Open-ended multi-attribute filtering/sorting on large datasets (Vacation planning or Movie selection).
    • Evaluation Metrics: Coded conversation logs (requirement origin, requirement iteration type, turn count, token usage), structural codebook analysis of generated spreadsheets (sheets, tables, charts, data columns used), and an adapted Mixed-Initiative Creativity Support Scale covering five creativity dimensions, four collaboration dimensions, and subjective outcome perceptions.
  4. Knowl 4 — Shift in Requirements Elicitation and Refinement Patterns

    data/table

    Comparing interaction transcripts between Plan Mode and Act Mode revealed that total unique requirements were virtually identical (6.56.5 vs. 6.46.4 new requirements per session), but Plan Mode shifted the elicitation of requirements from post-hoc refinements to upfront agent clarifying questions.

    Mode Proactive Clarification Plan Refine
    Plan Mode 3.1 4.1 0.7 1.3
    Act Mode 4.3 0.2 – 3.1
    Mode New Follow-up Reiteration Change
    Plan Mode 6.5 2.0 0.1 0.6
    Act Mode 6.4 0.7 0.1 0.4

    In Plan Mode, participants answered 8888 out of 9999 clarifying questions asked by the agent. Consequently, 44%44\% of all requirements and 37%37\% of brand-new requirements in Plan Mode originated from clarifying questions, compared to only 14%14\% and 10%10\% emerging as post-execution spreadsheet refinements. In contrast, Act Mode participants provided 41%41\% of all requirements and 35%35\% of new requirements via post-hoc refinements. Furthermore, Plan Mode elicited nearly three times as many requirement follow-ups (2.02.0 vs. 0.70.7), with 82%82\% of follow-ups arising directly from responses to clarifying questions.

  5. Knowl 5 — Token Efficiency and Refinement Turn Reduction via Upfront Planning

    empirical result

    While Plan Mode increased total conversational interaction time (10.410.4 vs. 8.88.8 minutes) and total conversation turns (5.35.3 vs. 3.63.6 turns on average) due to the upfront question-answering phase, it produced a marked reduction in post-execution spreadsheet refinement and overall computational cost:

    • Spreadsheet Refinements: Participants in Plan Mode required fewer spreadsheet refinement turns after the initial generation (1.41.4 vs. 2.22.2 turns on average).
    • Plan Modifications: Only 1010 of 2424 participants modified the agent's first proposed plan.
    • LLM Token Consumption: Mean token consumption (measuring generated reasoning, conversational messages, and spreadsheet tool call arguments) dropped from 17.0k17.0\text{k} tokens in Act Mode to 11.0k11.0\text{k} tokens in Plan Mode across all tasks, representing an approximate 35%35\% reduction in generated tokens.
  6. Knowl 6 — Equivalence of Generated Spreadsheet Artifacts and Features Across Modes

    empirical result

    Upfront planning did not alter the substantive feature diversity, personalization, or analytical breadth of the final spreadsheets:

    • Creation Tasks: The median number of coded semantic features was similar between Plan Mode and Act Mode (99 vs. 10.510.5 for the Budget task; 33 vs. 4.54.5 for the Schedule task). Although Act Mode generated more structural artifacts in the Budget task (median 2.52.5 sheets and 10.510.5 tables vs. 1.51.5 sheets and 4.54.5 tables in Plan Mode), the additional items consisted predominantly of meta-information, text instructions, and legends rather than additional data categories.
    • Analysis Tasks: The median number of unique dataset columns used for filtering or sorting was comparable (4.54.5 in Plan Mode vs. 5.05.0 in Act Mode; Vacation median 6.56.5 vs. 8.08.0, Movie median 4.04.0 vs. 3.73.7), and participants referenced identical column distributions across both modes.
  7. Knowl 7 — User Preference for Plan Mode in Creativity Support and Collaboration

    empirical result

    On an adapted Mixed-Initiative Creativity Support Scale, participants exhibited a clear overall preference for Plan Mode over Act Mode across all five evaluated creativity support dimensions (Enjoyment, Exploration, Expressiveness, Attention, and Worth) and three collaboration dimensions (Communication, Alignment, and Partnership).

    Act Mode was rated higher on the Agency dimension, which measures the extent to which the tool steers the user toward its own goals. On the attention scale, participants reported that Plan Mode required more attention directed toward the tool rather than the task, attributable to the additional UI elements. Despite these experiential preferences, ratings for perceived creative responsibility (1616 Plan vs. 1414 Act for user responsibility) and output satisfaction (2222 Plan vs. 2323 Act satisfied) showed no substantial difference between modes.

  8. Knowl 8 — Task and Problem-Solving Moderators of Planning Tool Preference

    empirical result

    User satisfaction and preference between Plan Mode and Act Mode were moderated by task type, task complexity, and individual interaction styles:

    1. Task Type: Plan Mode was strongly favored for Creation tasks across all dimensions (Enjoyment, Exploration, Expressiveness, Alignment, Partnership), whereas participants were neutral or favored Act Mode for open-ended data Analysis tasks.
    2. Task Length/Complexity: Participants favored Plan Mode more heavily on lengthier, complex tasks (Budget and Vacation, averaging 11.911.9 minutes and 4.94.9 turns) than on shorter tasks (Schedule and Movie, averaging 7.37.3 minutes and 4.04.0 turns).
    3. Upfront Prompting Style: Participants who naturally provided comprehensive upfront prompts (>3.7>3.7 initial requirements, n=9n=9) preferred Act Mode for its speed and direct execution.
    4. Iterative Refinement Style: Participants who relied on selective, depth-first iterative refinement across tasks (n=9n=9) preferred Act Mode's real-time, freestyle feedback loop over structured upfront planning.
  9. Knowl 9 — Control Without Action Effect in AI Steering Interfaces

    theoretical result

    The presence of visual and interactive steering affordances delivers substantial perceived user agency and trust even when those affordances are rarely actively exercised during execution.

    In the study, only 22 of 2424 participants directly edited the plan artifact via the UI, and only 66 of 2424 utilized the partial execution control. However, 1414 of 2424 participants explicitly highlighted the ability to view, reorder, edit, and partially execute plan steps as a primary driver of confidence and steerability (e.g., describing it as a visual "game plan"). This demonstrates that passive inspection and latent manual override capabilities help users bridge the "gulf of envisioning" in generative AI systems, generating design value through perceived control rather than direct operational action.

  10. Knowl 10 — Limitations of the Spreadsheet Plan Mode Evaluation

    limitation

    The evaluation of the spreadsheet Plan Mode prototype has several key limitations:

    1. Sample Size: The within-subjects study was conducted with N=24N = 24 participants, restricting the statistical generalizability of the quantitative findings.
    2. Tool Novelty: Almost all participants were using agentic spreadsheet tools for the first time, which may have influenced their interaction behavior and perception.
    3. Tutorial Disparity: The tutorial for Plan Mode covered custom UI features (the Plan pane and status updates), whereas the Act Mode tutorial covered only basic chat, introducing a potential interface familiarity confound.
    4. Time Constraints: Tasks were bounded by a 15–2015\text{--}20 minute time limit, which may have truncated longer-horizon refinement iterations and compressed outcome differences between modes.

Coverage note — None was omitted; all key architectural components, UI design features, empirical findings, quantitative tables, qualitative interaction style analyses, and stated limitations are fully represented.

References

  1. 1.G. Bansal, J. W. Vaughan, S. Amershi, E. Horvitz, A. Fourney, H. Mozannar, V. Dibia, and D. S. Weld, “Challenges in Human-Agent Communication,” 2024.
  2. 2.Anthropic, “Choose a Permission Mode – Claude Code Docs.” https://code.claude.com/docs/en/permission-modes, 2026. Accessed: 2026-05-01.
  3. 3.Microsoft, “Planning with Agents in VS Code.” https://code.visualstudio.com/docs/copilot/agents/planning, 2026. Accessed: 2026-05-01.
  4. 4.Cursor, “Plan Mode – Cursor Docs.” https://cursor.com/docs/agent/plan-mode, 2026. Accessed: 2026-05-01.
  5. 5.Cline, “Plan & Act Mode – Cline Documentation.” https://docs.cline.bot/core-workflows/plan-and-act, 2026. Accessed: 2026-05-01.
  6. 6.A. J. Ko, R. Abraham, L. Beckwith, A. Blackwell, M. Burnett, M. Erwig, C. Scaffidi, J. Lawrance, H. Lieberman, B. Myers, M. B. Rosson, G. Rothermel, M. Shaw, and S. Wiedenbeck, “The state of the art in end-user software engineering,” ACM Comput. Surv., vol. 43, Apr. 2011.
  7. 7.A. Sarkar, A. D. Gordon, C. Negreanu, C. Poelitz, S. S. Ragavan, and B. Zorn, “What is it like to program with artificial intelligence?,” 2022.
  8. 8.R. Pandita, C. Parnin, F. Hermans, and E. Murphy-Hill, “No half-measures: A study of manual and tool-assisted end-user programming tasks in Excel,” in 2018 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), pp. 95–103, 2018.
  9. 9.L. Reicherts, Z. T. Zhang, E. von Oswald, Y. Liu, Y. Rogers, and M. Hassib, “AI, Help Me Think—but for Myself: Assisting People in Complex Decision-Making by Providing Different Kinds of Cognitive Support,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, (New York, NY, USA), Association for Computing Machinery, 2025.
  10. 10.M. Kazemitabaar, J. Williams, I. Drosos, T. Grossman, A. Z. Henley, C. Negreanu, and A. Sarkar, “Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition,” in Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, UIST ’24, (New York, NY, USA), Association for Computing Machinery, 2024.
  11. 11.W.-H. Chen, W. Tong, A. Case, and T. Zhang, “Dango: A Mixed-Initiative Data Wrangling System using Large Language Model,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, (New York, NY, USA), Association for Computing Machinery, 2025.
  12. 12.I. Drosos, J. Williams, A. Sarkar, N. Wilson, S. Rintel, and P. Panda, “Dynamic Prompt Middleware: Contextual Prompt Refinement Controls for Comprehension Tasks,” in Proceedings of the 4th Annual Symposium on Human-Computer Interaction for Work, CHIWORK ’25, (New York, NY, USA), Association for Computing Machinery, 2025.
  13. 13.J. T. Liang, A. Kumar, Y. Bajpai, S. Gulwani, V. Le, C. Parnin, A. Radhakrishna, A. Tiwari, E. Murphy-Hill, and G. Soares, “TableTalk: Scaffolding Spreadsheet Development with a Language Agent,” ACM Trans. Comput.-Hum. Interact., vol. 32, Dec. 2025.
  14. 14.M. X. Liu, A. Sarkar, C. Negreanu, B. Zorn, J. Williams, N. Toronto, and A. D. Gordon, ““What It Wants Me To Say”: Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, (New York, NY, USA), Association for Computing Machinery, 2023.
  15. 15.J. Zamfirescu-Pereira, E. Jun, M. Terry, Q. Yang, and B. Hartmann, “Beyond Code Generation: LLM-supported Exploration of the Program Design Space,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, (New York, NY, USA), Association for Computing Machinery, 2025.
  16. 16.A. Kumar, Y. Bajpai, S. Gulwani, G. Soares, and E. Murphy-Hill, “Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild,” in 2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE), pp. 432–444, 2025.
  17. 17.R. Cheng, T. Barik, A. Leung, F. Hohman, and J. Nichols, “BISCUIT: Scaffolding LLM-Generated Code with Ephemeral UIs in Computational Notebooks,” in 2024 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), pp. 13–23, 2024.
  18. 18.A. Rahman, K. Niinuma, and A. Gupta, “DataSpeck: An AI-Driven Human-in-the-Loop System for Automating Transformations in Data Conversion Workflows,” in Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, CHI ’26, (New York, NY, USA), Association for Computing Machinery, 2026.
  19. 19.I. Drosos, A. Sarkar, X. Xu, C. Negreanu, S. Rintel, and L. Tankelevitch, ““It’s like a rubber duck that talks back”: Understanding Generative AI-Assisted Data Analysis Workflows through a Participatory Prompting Study,” in Proceedings of the 3rd Annual Meeting of the Symposium on Human-Computer Interaction for Work, CHIWORK ’24, (New York, NY, USA), Association for Computing Machinery, 2024.
  20. 20.G. He, G. Demartini, and U. Gadiraju, “Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, (New York, NY, USA), Association for Computing Machinery, 2025.
  21. 21.H. Mozannar, G. Bansal, C. Tan, A. Fourney, V. Dibia, J. Chen, J. Gerrits, T. Payne, M. K. Maldaner, M. Grunde-McLaughlin, et al., “Magentic-UI: Towards Human-in-the-loop Agentic Systems,” arXiv preprint arXiv:2507.22358, 2025.
  22. 22.K. J. K. Feng, K. Pu, M. Latzke, T. August, P. Siangliulue, J. Bragg, D. S. Weld, A. X. Zhang, and J. C. Chang, “Cocoa: Co-Planning and Co-Execution with AI Agents,” in Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, CHI ’26, (New York, NY, USA), Association for Computing Machinery, 2026.
  23. 23.T. Dong, H. Sampath, J. Y. Lee, S. Y. Shi, and A. Macvean, “From correctness to collaboration: Toward a human-centered framework for evaluating ai agent behavior in software engineering,” 2025.
  24. 24.R. Bairi, A. Sonwane, A. Kanade, V. D. C., A. Iyer, S. Parthasarathy, S. Rajamani, B. Ashok, and S. Shet, “Codeplan: Repository-level coding using llms and planning,” Proc. ACM Softw. Eng., vol. 1, July 2024.
  25. 25.S. Passi, “Agentic AI has a Human Oversight Problem,” Available at SSRN 5529058, 2025.
  26. 26.H. Li, J. Su, Y. Chen, Q. Li, and Z. Zhang, “SheetCopilot: Bringing Software Productivity to the Next Level through Large Language Models,” 2023.
  27. 27.Y. Chen, Y. Yuan, Z. Zhang, Y. Zheng, J. Liu, F. Ni, J. Hao, H. Mao, and F. Zhang, “SheetAgent: Towards A Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language Models,” 2025.
  28. 28.L. Yan, A. Head, K. Milne, V. Le, S. Gulwani, C. Parnin, and E. Murphy-Hill, “The Invisible Mentor: Inferring User Actions from Screen Recordings to Recommend Better Workflows,” in Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, CHI ’26, (New York, NY, USA), Association for Computing Machinery, 2026.
  29. 29.K. Ferdowsi, J. Williams, I. Drosos, A. D. Gordon, C. Negreanu, N. Polikarpova, A. Sarkar, and B. Zorn, “COLDECO: An End User Spreadsheet Inspection Tool for AI-Generated Code,” in 2023 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), pp. 82–91, 2023.
  30. 30.Microsoft, “Excel Copilot.” https://support.microsoft.com/en-US/excel/copilot/get-started-with-copilot-in-excel, 2026. Accessed: 2026-06-21.
  31. 31.Shortcut, “Shortcut AI.” https://shortcut.ai/, 2026. Accessed: 2026-06-21.
  32. 32.Anthropic, “Claude for Excel.” https://claude.com/claude-for-excel, 2026. Accessed: 2026-06-21.
  33. 33.S. Srinivasa Ragavan, A. Sarkar, and A. D. Gordon, “Spreadsheet Comprehension: Guesswork, Giving Up and Going Back to the Author,” in Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, (New York, NY, USA), Association for Computing Machinery, 2021.
  34. 34.A. Sarkar, “Will Code Remain a Relevant User Interface for End-User Programming with Generative AI Models?,” in Proceedings of the 2023 ACM SIGPLAN International Symposium on New Ideas, New Paradigms, and Reflections on Programming and Software, Onward! 2023, (New York, NY, USA), p. 153–167, Association for Computing Machinery, 2023.
  35. 35.M. Brachman, A. El-Ashry, C. Dugan, and W. Geyer, “How Knowledge Workers Use and Want to Use LLMs in an Enterprise Context,” in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA ’24, (New York, NY, USA), Association for Computing Machinery, 2024.
  36. 36.N. McDonald, S. Schoenebeck, and A. Forte, “Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice,” Proc. ACM Hum.-Comput. Interact., vol. 3, Nov. 2019.
  37. 37.J. R. Landis and G. G. Koch, “The measurement of observer agreement for categorical data,” Biometrics, vol. 33, no. 1, pp. 159–174, 1977.
  38. 38.T. Lawton, F. J. Ibarrola, D. Ventura, and K. Grace, “Drawing with Reframer: Emergence and Control in Co-Creative AI,” in Proceedings of the 28th International Conference on Intelligent User Interfaces, IUI ’23, (New York, NY, USA), p. 264–277, Association for Computing Machinery, 2023.
  39. 39.R. Parasuraman and D. Manzey, “Complacency and Bias in Human Use of Automation: An Attentional Integration,” Human factors, vol. 52, pp. 381–410, 06 2010.
  40. 40.V. Padmakumar and H. He, “Does Writing with Language Models Reduce Content Diversity?,” in The Twelfth International Conference on Learning Representations, 2024.
  41. 41.A. Sarkar, “AI Should Challenge, Not Obey,” Commun. ACM, vol. 67, p. 18–21, Sept. 2024.
  42. 42.L. Tankelevitch, E. L. Glassman, J. He, A. Kittur, M. Lee, S. Palani, A. Sarkar, G. Ramos, Y. Rogers, and H. Subramonyam, “Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop,” 2025.
  43. 43.E. Murphy-Hill, D. Y. Lee, G. C. Murphy, and J. Mcgrenere, “How Do Users Discover New Tools in Software Development and Beyond?,” Comput. Supported Coop. Work, vol. 24, p. 389–422, Oct. 2015.
  44. 44.V. Danry, P. Pataranutaporn, Y. Mao, and P. Maes, “Don’t Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI explanations,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, (New York, NY, USA), Association for Computing Machinery, 2023.
  45. 45.A. L. Cox, S. J. Gould, M. E. Cecchinato, I. Iacovides, and I. Renfree, “Design Frictions for Mindful Interactions: The Case for Microboundaries,” in Proceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems, CHI EA ’16, (New York, NY, USA), p. 1389–1397, Association for Computing Machinery, 2016.
  46. 46.Z. Buçinca, M. B. Malaya, and K. Z. Gajos, “To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making,” Proc. ACM Hum.-Comput. Interact., vol. 5, Apr. 2021.
  47. 47.A. Anderson, J. N. Guevara, F. Moussaoui, T. Li, M. Vorvoreanu, and M. Burnett, “Measuring User Experience Inclusivity in Human-AI Interaction via Five User Problem-Solving Styles,” ACM Trans. Interact. Intell. Syst., vol. 14, Sept. 2024.

Citation

MLA
Kumar, A., et al. “Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents”. arXiv, 2026, http://arxiv.org/abs/2607.23670v2.
APA
Kumar, A., Dutta, A., Gulwani, S., Soares, G., Sarkar, A., & Murphy-Hill, E. (2026). Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents. arXiv. http://arxiv.org/abs/2607.23670v2
Chicago
Kumar, A., A. Dutta, S. Gulwani, G. Soares, A. Sarkar, and E. Murphy-Hill. 2026. “Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents”. arXiv. http://arxiv.org/abs/2607.23670v2.
Harvard
Kumar, A. et al. (2026) “Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2607.23670v2.
Vancouver
1. Kumar A, Dutta A, Gulwani S, Soares G, Sarkar A, Murphy-Hill E (2026) Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents. arXiv

BibTeX

@article{kumar2026plans,
  title = {Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents},
  author = {Kumar, Aayush and Dutta, Avik and Gulwani, Sumit and Soares, Gustavo and Sarkar, Advait and Murphy-Hill, Emerson},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2607.23670v2},
  eprint = {2607.23670}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/