Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents
Aayush KumarAvik DuttaSumit GulwaniGustavo SoaresAdvait SarkarEmerson Murphy-Hill
Demonstrates through a user study that adding a plan mode to spreadsheet AI agents reduces iterative prompt refinements and improves perceived creativity support and collaboration without compromising task outcomes.
Autonomous artificial intelligence agents are increasingly capable of executing complex end-user tasks, leading to the rapid adoption of interactive planning features—often termed "Plan Modes"—in developer tools. While software engineers rely on planning to review code structures and maintain execution control, end users working in spreadsheet environments typically approach problem solving through flexible, emergent trial and error. The article investigates whether structuring spreadsheet agent interactions through an upfront planning mode provides meaningful value to end users, evaluating its effects on information exchange, task outcomes, computational cost, and user experience.
To evaluate this capability, the researchers developed a high-fidelity prototype that incorporates clarifying questions, editable persistent plan displays, and partial execution controls within a spreadsheet environment. They tested this prototype against an immediate-execution baseline ("Act Mode") via a controlled, within-subjects study involving 24 experienced spreadsheet users. Participants completed open-ended creation tasks (such as personal budgets and schedules) and subjective data analysis tasks (such as vacation planning and movie selection), while researchers recorded behavioral logs, requirements evolution, resulting workbook artifacts, and survey measures evaluating creativity and collaboration.
The study revealed four primary findings. First, interactive planning fundamentally shifted how users communicated requirements: Plan Mode users expressed 44% of their requirements in response to clarifying questions and only 14% via post-execution edits, whereas baseline users derived over 41% of requirements from iterative refinements after the agent modified the sheet. Second, this upfront clarification reduced execution refinements (1.4 turns versus 2.2 turns) and decreased large language model computation costs by approximately 35% (averaging 11.0k tokens in Plan Mode versus 17.0k tokens in Act Mode). Third, despite differing interaction paths, the generated spreadsheets showed nearly identical feature distributions, complexity, and content across both conditions. Fourth, participants reported a distinct preference for Plan Mode across dimensions of creativity support and human-agent collaboration—particularly for open-ended creation tasks and longer workflows—even though they rarely manipulated the interactive plan user interface directly.
These findings indicate that the primary value of a planning mode in end-user tools stems from improved interaction mechanics rather than superior final workbooks. Clarifying questions efficiently extract user intent and mitigate costly computational cycles from trial-and-error edits. Furthermore, the presence of visible plan artifacts provides users with a psychological sense of oversight and control ("control without action"), bridging the gap between user intent and agent execution without requiring extensive manual corrections.
For product teams and enterprise stakeholders, the article supports integrating clarifying questions and structured planning into generative spreadsheet tools. However, systems should avoid rigid, mandatory planning steps for all workflows. Because users who preferred rapid iteration or provided detailed initial prompts found planning overly restrictive, tools should implement mixed-initiative triggers that adapt based on user prompting style, task scope, and problem complexity.
Confidence in these findings is supported by rigorous qualitative coding and high inter-rater agreement. Nonetheless, leaders should interpret the results within the study's boundaries: the evaluation utilized a modest sample size (N=24), examined short-duration tasks (15–20 minutes), and evaluated users largely unaccustomed to spreadsheet agents. Further pilot testing across broader enterprise datasets and longer project timelines is advisable before making definitive architectural commitments.
- Paper: Magentic-UI: Towards Human-in-the-loop Agentic Systems, Hussein Mozannar et al. (2025). It provides foundational principles and empirical results on integrating human oversight and co-planning into autonomous agent workflows, establishing the interactive paradigms evaluated in the source.
- Paper: Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models, Lei Wang et al. (2023). It introduces the concept of upfront planning and step decomposition prior to execution in language models, which forms the underlying algorithmic basis of plan modes.
- Paper: VizCopilot: Fostering Appropriate Reliance on Enterprise Chatbots with Context Visualization, Sam Yu-Te Lee et al. (2025). It examines how interface transparency and contextual scaffolding foster appropriate human reliance on AI assistants, motivating the human-AI interaction challenges studied in spreadsheet planning.
- Paper: Repair Is Nearly Generation: Multilingual Program Repair with LLMs, Harshit Joshi et al. (2023). It explores automated program and formula repair mechanisms in spreadsheet environments, providing necessary context on how end users iterate on spreadsheet tasks.
- Paper: Interactive Debugging and Steering of Multi-Agent AI Systems, Will Epperson et al. (2025). It details how human-in-the-loop steering and interactive intervention affect user control and workflow refinement during complex agent execution.
- Paper: You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy, Rudrajit Choudhuri et al. (2026). It broadens the inquiry into human-AI collaboration by analyzing where and why practitioners set boundaries on AI autonomy versus planning and control across different task types.
- Paper: The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior, Xiangzhe Xu et al. (2026). It extends the study of agent interfaces by systematically analyzing how tool architecture and interaction abstractions shape agent behavior and user execution consistency.
- Paper: From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills, Zisu Huang et al. (2026). It investigates how structured plans and execution trajectories from domains like spreadsheet manipulation can be extracted and consumed as reusable agent skills.
- Paper: Covering Human Action Space for Computer Use: Data Synthesis and Benchmark, Miaosen Zhang et al. (2026). It scales human-AI spreadsheet interaction down to visual action spaces by establishing benchmarks for multi-point and continuous GUI execution across spreadsheet interfaces.
- Paper: AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation, Priyam Sahoo et al. (2026). It complements the behavioral evaluation of agent planning by assessing the internal trajectory quality and stage ordering of agent problem-solving beyond binary task outcomes.
