Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
Sheshera MysoreDebarati DasHancheng CaoBahar Sarrafzadeh
Identifies prototypical multi-turn collaboration behaviors in real-world LLM-assisted writing and connects them to specific user goals, providing an empirical foundation for designing and aligning more responsive AI writing assistants.
As conversational large language models are increasingly adopted across professional, academic, and creative domains, evaluating real-world human-AI interaction is essential. Current industry practice largely optimizes these systems using single-turn interactions and explicit ratings, such as thumbs-up or thumbs-down feedback. However, users in practice rarely provide explicit satisfaction ratings and instead engage in multi-turn dialogues to steer, refine, and co-construct text. To understand how people truly work with these tools, the article investigates the high-level collaboration behaviors that emerge during writing tasks and analyzes how these behaviors vary across different writing goals.
The article evaluates user collaboration patterns using large-scale conversation logs from two distinct platforms: Bing Copilot, covering 20.5 million sessions over seven months in 2024, and the publicly available WildChat dataset, covering 800,000 sessions over thirteen months in 2023–2024. After filtering for multi-turn English writing sessions, the analysis examined 250,000 Bing Copilot sessions and 68,000 WildChat sessions originating from users across more than 150 countries. The approach used automated language model classifiers to categorize user intents and follow-up prompts, statistical clustering via principal component analysis to identify core collaboration behaviors, and regression modeling paired with manual qualitative review to correlate specific writing goals with these behaviors.
The analysis revealed several key findings regarding user interaction. First, seven prototypical collaboration behaviors account for 80% to 85% of all variance in user follow-ups across both platforms, demonstrating consistent interaction patterns regardless of the underlying interface or model. Users most frequently collaborate by revising or restating their initial prompt (accounting for 14.5% to 14.6% of variance), asking follow-up questions, requesting additional output variations, or directly modifying generated text. In contrast, explicit positive or negative satisfaction signals occurred in only 1% to 5% of sessions. Furthermore, specific writing goals strongly correlate with distinct collaboration behaviors: users generating titles or marketing copy repeatedly request more outputs to brainstorm ideas; users creating long narratives or scripts use iterative follow-ups with varying specificity to stage generation; and users drafting professional documents or technical texts frequently ask questions to learn domain-specific norms and actively inject personal background information missing from the draft.
These findings demonstrate that users treat language models as active co-creators and learning partners rather than mere task-execution engines. This dynamic reveals a clear mismatch between actual user behavior and standard system alignment techniques that rely on single-turn, explicit feedback. For organizations deploying these tools, relying solely on explicit user ratings creates a blind spot regarding how effectively systems support complex workflows. In addition, when systems fail to proactively elicit private user context or adapt to domain conventions, users are forced to expend extra effort steering the conversation manually.
To bridge this gap, system developers and product teams should shift toward session-level alignment frameworks that learn from implicit, multi-turn conversational feedback rather than isolated prompt-response pairs. Systems should be designed to handle under-specified user feedback, proactively solicit missing personal context when drafting specialized communications, and offer diverse outputs for brainstorming tasks. Developers must also align models to support user learning and feedback-seeking while establishing safeguards to prevent cultural homogenization and privacy leaks during implicit learning.
Confidence in these findings is supported by the massive sample size, cross-platform consistency, and manual validation of the classification models. Nevertheless, the conclusions carry some limitations: the analysis was restricted to English-language sessions on desktop computers, did not explicitly model the sequential order of multi-turn steps within a conversation, and relied on qualitative sampling rather than direct user interviews. Decision-makers should account for these boundary conditions when applying the insights to non-English or mobile-first environments.
No sufficiently relevant recommendations were found.
- Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). Building on the source’s finding that users revise intent across turns, this study tests how models handle evolving requirements, corrections, and task switches.
