Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild

Sheshera MysoreDebarati DasHancheng CaoBahar Sarrafzadeh

article2025EMNLP35 citationsSAC Highlight Award

Identifies prototypical multi-turn collaboration behaviors in real-world LLM-assisted writing and connects them to specific user goals, providing an empirical foundation for designing and aligning more responsive AI writing assistants.

Listen

As conversational large language models are increasingly adopted across professional, academic, and creative domains, evaluating real-world human-AI interaction is essential. Current industry practice largely optimizes these systems using single-turn interactions and explicit ratings, such as thumbs-up or thumbs-down feedback. However, users in practice rarely provide explicit satisfaction ratings and instead engage in multi-turn dialogues to steer, refine, and co-construct text. To understand how people truly work with these tools, the article investigates the high-level collaboration behaviors that emerge during writing tasks and analyzes how these behaviors vary across different writing goals.

The article evaluates user collaboration patterns using large-scale conversation logs from two distinct platforms: Bing Copilot, covering 20.5 million sessions over seven months in 2024, and the publicly available WildChat dataset, covering 800,000 sessions over thirteen months in 2023–2024. After filtering for multi-turn English writing sessions, the analysis examined 250,000 Bing Copilot sessions and 68,000 WildChat sessions originating from users across more than 150 countries. The approach used automated language model classifiers to categorize user intents and follow-up prompts, statistical clustering via principal component analysis to identify core collaboration behaviors, and regression modeling paired with manual qualitative review to correlate specific writing goals with these behaviors.

The analysis revealed several key findings regarding user interaction. First, seven prototypical collaboration behaviors account for 80% to 85% of all variance in user follow-ups across both platforms, demonstrating consistent interaction patterns regardless of the underlying interface or model. Users most frequently collaborate by revising or restating their initial prompt (accounting for 14.5% to 14.6% of variance), asking follow-up questions, requesting additional output variations, or directly modifying generated text. In contrast, explicit positive or negative satisfaction signals occurred in only 1% to 5% of sessions. Furthermore, specific writing goals strongly correlate with distinct collaboration behaviors: users generating titles or marketing copy repeatedly request more outputs to brainstorm ideas; users creating long narratives or scripts use iterative follow-ups with varying specificity to stage generation; and users drafting professional documents or technical texts frequently ask questions to learn domain-specific norms and actively inject personal background information missing from the draft.

These findings demonstrate that users treat language models as active co-creators and learning partners rather than mere task-execution engines. This dynamic reveals a clear mismatch between actual user behavior and standard system alignment techniques that rely on single-turn, explicit feedback. For organizations deploying these tools, relying solely on explicit user ratings creates a blind spot regarding how effectively systems support complex workflows. In addition, when systems fail to proactively elicit private user context or adapt to domain conventions, users are forced to expend extra effort steering the conversation manually.

To bridge this gap, system developers and product teams should shift toward session-level alignment frameworks that learn from implicit, multi-turn conversational feedback rather than isolated prompt-response pairs. Systems should be designed to handle under-specified user feedback, proactively solicit missing personal context when drafting specialized communications, and offer diverse outputs for brainstorming tasks. Developers must also align models to support user learning and feedback-seeking while establishing safeguards to prevent cultural homogenization and privacy leaks during implicit learning.

Confidence in these findings is supported by the massive sample size, cross-platform consistency, and manual validation of the classification models. Nevertheless, the conclusions carry some limitations: the analysis was restricted to English-language sessions on desktop computers, did not explicitly model the sequential order of multi-turn steps within a conversation, and relied on qualitative sampling rather than direct user interviews. Decision-makers should account for these boundary conditions when applying the insights to non-English or mobile-first environments.

No sufficiently relevant recommendations were found.

  • Paper: LLMs Get Lost in Evolving User Intent, Jihoon Tack et al. (2026). Building on the source’s finding that users revise intent across turns, this study tests how models handle evolving requirements, corrections, and task switches.
Cover for Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild

Abstract

As large language models (LLMs) are used in complex writing workflows, users engage in multi-turn interactions to steer generations to better fit their needs. Rather than passively accepting output, users actively refine, explore, and co-construct text. We conduct a large-scale analysis of this collaborative behavior for users engaged in writing tasks in the wild with two popular AI assistants, Bing Copilot and WildChat. Our analysis goes beyond simple task classification or satisfaction estimation common in prior work and instead characterizes how users interact with LLMs through the course of a session. We identify prototypical behaviors in how users interact with LLMs in prompts following their original request. We refer to these as Prototypical Human-AI Collaboration Behaviors (PATHs) and find that a small group of PATHs explain a majority of the variation seen in user-LLM interaction. These PATHs span users revising intents, exploring texts, posing questions, adjusting style or injecting new content. Next, we find statistically significant correlations between specific writing intents and PATHs, revealing how users' intents shape their collaboration behaviors. We conclude by discussing the implications of our findings on LLM alignment.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Analysis Setup
  • 4 Log Analysis with Paths
  • 5 Results – Exploring Paths
  • 6 Results – Correlating Intents and Paths
  • 6.1 Requesting more outputs to brainstorm or stage long generations
  • 6.2 Asking follow-up questions to learn or stage long generations
  • 6.3 Adding to generations when they lacked content known only to the user
  • 7 Discussion
  • 8 Limitations
  • References
  • A Extended Related Work
  • B Analysis Setup - Details
  • B.1 Dataset Description and Preliminary Filtering
  • B.2 Identifying Writing Sessions
  • B.3 Validating Task Classifiers
  • C Log Analysis with Paths - Details
  • C.1 Identifying Paths
  • C.2 Correlating intents and Paths
  • C.3 Validating Follow-up and Intent Classifiers
  • D Extended Results
  • D.1 Changing style to align generations with readers
  • D.2 Dataset dependent trends in revising and elaborating on requests

Knowls

  1. Knowl 1 — Seven prototypical collaboration behaviors explain most follow-up variation

    empirical result

    Across the Bing Copilot and WildChat writing-session datasets, seven principal components of user follow-up behavior explained 80–85% of the variation in follow-up types. Each component explained a relatively small 8–14%, and most were dominated by one follow-up type, suggesting that distinct collaboration behaviors usually occur separately rather than as frequent combinations. Restating an original request was the largest component in both datasets, explaining 14.5–14.6% of variance; requesting additional outputs was the second largest. The components were broadly similar across the two deployments, although the components associated with questions and generation changes differed: questions co-occurred with style changes in WildChat, while the fourth component was associated with adding content in Bing Copilot and changing style in WildChat.

  2. Knowl 2 — Two large in-the-wild writing corpora

    experimental setup

    The analysis used English writing sessions from Bing Copilot and WildChat. Bing Copilot writing data (BCPWr) contained 250,000 sessions, sampled from 500,000 sessions classified as writing among 2.8 million English sessions with at least two user utterances; the source was a daily random sample of 20.5 million Bing Copilot sessions collected from April through October 2024. WildChat writing data (WCWr) contained 68,000 sessions classified as writing among 160,000 English sessions with at least two user utterances, drawn from the public WildChat-1M collection gathered from April 2023 through May 2024. BCPWr represented about 202,000 users in 219 countries, and WCWr about 22,000 users in 166 countries. Writing included generating or rewriting communicative, creative, and technical text, as well as summarization. Sessions averaged 2.1 follow-ups in BCPWr and 1.9 in WCWr. Bing Copilot interactions were powered by GPT-4; WildChat interactions used GPT-4 or GPT-3.5-Turbo.

  3. Knowl 3 — PATH discovery and intent-correlation analysis

    model/method

    The analysis first used GPT-4o classifiers to label user utterances as an original request or one of eleven follow-up types: restating a request, elaborating on it, requesting answers, requesting more outputs, changing style, adding content, removing content, courtesy response, positive response, negative response, or undefined response. A separate multi-label GPT-4o classifier assigned writing intents to original requests, including improving text; generating messages, professional documents, summaries, technical text, essays, online posts, catchy text, biographies, stories, scripts, characters, poems, songs, and jokes; getting references; asking about writing; and undefined requests. For each session, follow-up-type counts were normalized by the number of user utterances and weighted by inverse type frequency across the dataset. Principal component analysis (PCA) on this session-by-type matrix produced components interpreted as Prototypical Human-AI Collaboration Behaviors (PATHs); the components accounting for 80–85% of variance were retained. To test associations with writing intents, the researchers used one-hot intent indicators as predictors in separate logistic regressions for each PATH. PATH membership was defined by retaining the 15–20% of sessions with the highest component scores; follow-up-type presence was also used directly as a regression outcome when a corresponding PATH was not shared across datasets. Classifiers were manually evaluated: follow-up-label accuracy was 79.09% for BCPWr and 84.55% for WCWr; intent-label accuracy was 81.58% and 78.70%, respectively. Cohen’s kappa values for these judgments ranged from 0.74 to 0.82. Writing-session classification accuracy was 85.71% for Bing Copilot and 91.07% for WildChat.

  4. Knowl 4 — Common follow-ups and the rarity of explicit feedback

    empirical result

    Across both writing corpora, follow-ups commonly involved revising or elaborating on the original request, asking questions about a generation, requesting additional outputs, or modifying generated text; each of these broad behaviors appeared in 18–30% of sessions. Explicitly positive or negative reactions and courtesy responses were much less common, appearing in only 1–5% of sessions. Thus, users’ observable collaboration behavior was more often conveyed through requests that continued or changed the work than through direct satisfaction statements.

  5. Knowl 5 — Requesting more outputs supports brainstorming and staged creative generation

    empirical result

    The PATH characterized by requests for additional outputs was positively associated with the intent to generate catchy text, such as product names, titles, or email subjects. Qualitative analysis indicated that users often sought multiple alternatives to brainstorm more creative or attention-grabbing wording. Weaker positive associations appeared for script and fictional-character generation. In those sessions, users sometimes explored alternatives but also used successive requests to build long creative texts, with follow-ups ranging from a general request for more output to more specific direction. These observations document both alternative-generation behavior and multi-turn construction of creative narratives in real-world writing sessions.

  6. Knowl 6 — Questions serve knowledge seeking, norm learning, and narrative direction

    empirical result

    When users asked questions about a model’s writing output, the associated writing intents included professional-document, summary, technical-text, and fictional-character generation. The purposes varied by intent. Questions during technical writing and summarization often sought domain knowledge or clarification grounded in the task context. Questions about professional documents, including cover letters, could seek advice about professional norms and expectations or feedback on the draft. In fictional-character and story-generation sessions, questions often referred back to the model’s preceding output and helped users direct the development of a longer narrative. The authors also found that questions in fictional-character sessions were especially similar to the preceding generation, consistent with questions being grounded in the generated narrative.

  7. Knowl 7 — Users add information missing from generated drafts

    empirical result

    The behavior of adding content to a model’s generation was associated with requests for professional documents, interpersonal messages, and creative writing such as stories, scripts, and characters. Qualitative examination found that users commonly supplied information likely to be known only to them: personal experience or skills for professional documents, personal stories for messages, and plot details for fiction. The finding identifies situations in which a draft may lack user-specific information; the authors suggest personalization or proactive requests for missing details as possible design directions, rather than demonstrating that either approach improves outcomes.

  8. Knowl 8 — Style changes target readers and communicative expectations

    empirical result

    Requests to change a generation’s style were associated with improving existing text and generating messages, summaries, and online posts. Users described in the logs changed wording or format to better fit a communication’s assumed norms, a business context, a particular group or individual, or readers’ likely preferences. In summaries, style changes could target readability. Users also sometimes requested greater concision. This indicates that style modification in these writing sessions often served reader- and context-sensitive goals, not merely a generic preference for a different tone.

  9. Knowl 9 — Creative writing follow-ups become more specific over a session

    empirical result

    In multi-turn creative writing, users could progress from requesting more outputs, to asking questions about the generation, to adding concrete content. The analysis found increasing similarity between these successive follow-up types and the next model generation, which the authors interpret as increasing specificity in user feedback. Examples in fictional-character writing show this progression from a broad request for more material toward questions and plot additions that more directly guide the output. The pattern motivates alignment methods that can respond usefully even when early session feedback is underspecified.

  10. Knowl 10 — Implications for session-level alignment and user learning

    theoretical result

    The observed follow-ups suggest that alignment for writing should account for interaction across a session, rather than relying only on explicit single-turn preference judgments. Many users supplied implicit feedback by revising requests, changing style, asking questions, or adding content, while explicit positive and negative feedback was rare. The study also distinguishes task-oriented behavior, such as changing or adding to a draft, from exploratory behavior, such as requesting alternatives or asking questions. The authors propose learning from realistic multi-turn feedback, handling underspecified goals, and supporting exploration and user learning as promising alignment goals. These are implications for future system design, not experimentally tested alignment interventions.

  11. Knowl 11 — Scope and modeling limitations

    limitation

    The study covered English writing sessions only, so it may miss collaboration behaviors specific to other languages. Its PCA representation is linear and does not model the order of follow-ups; richer sequence-aware methods could reveal patterns that this approach cannot capture. The intent–behavior interpretations also relied on lightweight qualitative review of conversations rather than interviews or controlled studies with the users themselves. Finally, using implicit interactions for alignment could expose systems to adversarial feedback or create privacy risks when follow-ups contain personal information; the authors identify safety and privacy protections as necessary for such work.

Coverage note — The dataset-dependent intent associations for request-revision and request-elaboration behaviors are not detailed because the paper reports no positive, statistically significant intent correlations replicated across both corpora; raw classifier prompts and example conversations are omitted because they support, but do not add substantially to, the core findings.

References

  1. 1.Dhruv Agarwal, Mor Naaman, and Aditya Vashistha. 2025. Ai suggestions homogenize writing toward western styles and diminish cultural nuances. Preprint, arXiv:2409.11360.
  2. 2.Marwah Alaofi, Luke Gallagher, Dana Mckay, Lauren L. Saling, Mark Sanderson, Falk Scholer, Damiano Spina, and Ryen W. White. 2022. Where do queries come from? In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’22, page 2850–2862, New York, NY, USA. Association for Computing Machinery.
  3. 3.Aparna Ananthasubramaniam, Hong Chen, Jason Yan, Kenan Alkiek, Jiaxin Pei, Agrima Seth, Lavinia Dunagan, Minje Choi, Benjamin Litterer, and David Jurgens. 2023. Exploring linguistic style matching in online communities: The role of social context and conversation dynamics. In Proceedings of the First Workshop on Social Influence in Conversations (SICon 2023), pages 64–74, Toronto, Canada. Association for Computational Linguistics.
  4. 4.Yoshua Bengio, Aaron Courville, and Pascal Vincent. 2013. Representation learning: A review and new perspectives. IEEE Trans. Pattern Anal. Mach. Intell., 35(8):1798–1828.
  5. 5.Advait Bhat, Saaket Agashe, Parth Oberoi, Niharika Mohile, Ravi Jangir, and Anirudha Joshi. 2023. Interacting with next-phrase suggestions: How suggestion systems aid and influence the cognitive processes of writing. In Proceedings of the 28th International Conference on Intelligent User Interfaces, IUI ’23, page 436–452, New York, NY, USA. Association for Computing Machinery.
  6. 6.Ananya Bhattacharjee, Jina Suh, Mahsa Ershadi, Shamsi T. Iqbal, Andrew D. Wilson, and Javier Hernandez. 2024. Understanding communication preferences of information workers in engagement with text-based conversational agents. Preprint, arXiv:2410.20468.
  7. 7.Param Biyani, Yasharth Bajpai, Arjun Radhakrishna, Gustavo Soares, and Sumit Gulwani. 2024. Rubicon: Rubric-based evaluation of domain-specific human ai conversations. In Proceedings of the 1st ACM International Conference on AI-Powered Software, AIware 2024, page 161–169, New York, NY, USA. Association for Computing Machinery.
  8. 8.Faeze Brahman, Alexandru Petrusca, and Snigdha Chaturvedi. 2020. Cue me in: Content-inducing approaches to interactive story generation. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, pages 588–597, Suzhou, China. Association for Computational Linguistics.
  9. 9.Natalie Grace Brigham, Chongjiu Gao, Tadayoshi Kohno, Franziska Roesner, and Niloofar Mireshghallah. 2024. Developing story: Case studies of generative ai’s use in journalism. Preprint, arXiv:2406.13706.
  10. 10.Daniel Buschek. 2024. Collage is the new writing: Exploring the fragmentation of text and user interfaces in ai tools. In Proceedings of the 2024 ACM Designing Interactive Systems Conference, DIS ’24, page 2719–2737, New York, NY, USA. Association for Computing Machinery.
  11. 11.Tuhin Chakrabarty, Vishakh Padmakumar, Faeze Brahman, and Smaranda Muresan. 2024. Creativity support in the age of large language models: An empirical study involving emerging writers. Preprint, arXiv:2309.12570.
  12. 12.Eric Chamoun, Michael Schlichtkrull, and Andreas Vlachos. 2024. Automated focused feedback generation for scientific writing assistance. In Findings of the Association for Computational Linguistics: ACL 2024, pages 9742–9763, Bangkok, Thailand. Association for Computational Linguistics.
  13. 13.Shreyas Chaudhari, Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande, and Bruno Castro da Silva. 2024. Rlhf deciphered: A critical analysis of reinforcement learning from human feedback for llms. arXiv preprint arXiv:2404.08555.
  14. 14.Jia Chen, Jiaxin Mao, Yiqun Liu, Fan Zhang, Min Zhang, and Shaoping Ma. 2021. Towards a better understanding of query reformulation behavior in web search. In Proceedings of the Web Conference 2021, WWW ’21, page 743–755, New York, NY, USA. Association for Computing Machinery.
  15. 15.Katherine M. Collins, Albert Q. Jiang, Simon Frieder, Lionel Wong, Miri Zilka, Umang Bhatt, Thomas Lukasiewicz, Yuhuai Wu, Joshua B. Tenenbaum, William Hart, Timothy Gowers, Wenda Li, Adrian Weller, and Mateja Jamnik. 2024a. Evaluating language models for mathematics through interactions. Proceedings of the National Academy of Sciences, 121(24):e2318124121.
  16. 16.Katherine M Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E Zhang, Tan Zhi-Xuan, Mark Ho, Vikash Mansinghka, and 1 others. 2024b. Building machines that learn and think with people. Nature human behaviour, 8(10):1851–1863.
  17. 17.Yang Deng, Lizi Liao, Wenqiang Lei, Grace Hui Yang, Wai Lam, and Tat-Seng Chua. 2025. Proactive conversational ai: A comprehensive survey of advancements and opportunities. ACM Trans. Inf. Syst., 43(3).
  18. 18.Paramveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, and Lionel Peter Robert. 2024. Shaping human-ai collaboration: Varied scaffolding levels in co-writing with language models. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY, USA. Association for Computing Machinery.
  19. 19.Shachar Don-Yehiya, Leshem Choshen, and Omri Abend. 2024. Naturally Occurring Feedback is Common, Extractable and Useful. Preprint, arXiv:2407.10944.
  20. 20.Fiona Draxler, Anna Werner, Florian Lehmann, Matthias Hoppe, Albrecht Schmidt, Daniel Buschek, and Robin Welsch. 2024. The ai ghostwriter effect: When users do not perceive ownership of ai-generated text but self-declare as authors. ACM Trans. Comput.-Hum. Interact., 31(2).
  21. 21.Ian Drosos, Advait Sarkar, Neil Toronto, and 1 others. 2025. " it makes you think": Provocations help restore critical thinking to ai-assisted knowledge work. arXiv preprint arXiv:2501.17247.
  22. 22.Susan Dumais, Robin Jeffries, Daniel M. Russell, Diane Tang, and Jaime Teevan. 2014. Understanding User Behavior Through Log Data and Analysis, pages 349–372. Springer New York, New York, NY.
  23. 23.Nathan Eagle and Alex Sandy Pentland. 2009. Eigenbehaviors: Identifying structure in routine. Behavioral ecology and sociobiology, 63:1057–1066.
  24. 24.Jonas Frich, Lindsay MacDonald Vermeulen, Christian Remy, Michael Mose Biskjaer, and Peter Dalsgaard. 2019. Mapping the landscape of creativity support tools in hci. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, CHI ’19, page 1–18, New York, NY, USA. Association for Computing Machinery.
  25. 25.Yue Fu, Sami Foell, Xuhai Xu, and Alexis Hiniker. 2024. From text to self: Users’ perception of aimc tools on interpersonal communication and self. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY, USA. Association for Computing Machinery.
  26. 26.Ge Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro, and Dipendra Misra. 2024. Aligning LLM agents by learning latent preference from user edits. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
  27. 27.Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2024. A survey of confidence estimation and calibration in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 6577–6595, Mexico City, Mexico. Association for Computational Linguistics.
  28. 28.Katy Ilonka Gero, Tao Long, and Lydia B Chilton. 2023. Social dynamics of ai support in creative writing. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA. Association for Computing Machinery.
  29. 29.Derek Greene, Derek O’Callaghan, and Pádraig Cunningham. 2014. How many topics? stability analysis for topic models. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2014, Nancy, France, September 15-19, 2014. Proceedings, Part I 14, pages 498–513. Springer.
  30. 30.Damodar N Gujarati. 2021. Essentials of econometrics. Sage Publications.
  31. 31.Jeffrey T Hancock, Mor Naaman, and Karen Levy. 2020. Ai-mediated communication: Definition, research agenda, and ethical considerations. Journal of Computer-Mediated Communication, 25(1):89–100.
  32. 32.Kunal Handa, Alex Tamkin, Miles McCain, Saffron Huang, Esin Durmus, Sarah Heck, Jared Mueller, Jerry Hong, Stuart Ritchie, Tim Belonax, and 1 others. 2025. Which economic tasks are performed with ai? evidence from millions of claude conversations. arXiv preprint arXiv:2503.04761.
  33. 33.Bairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas, Shiyu Chang, and Yang Zhang. 2024. Decomposing uncertainty for large language models through input clarification ensembling. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 19023–19042. PMLR.
  34. 34.Xi Yu Huang, Krishnapriya Vishnubhotla, and Frank Rudzicz. 2024. The gpt-writingprompts dataset: A comparative analysis of character portrayal in short stories. Preprint, arXiv:2406.16767.
  35. 35.Angel Hsing-Chi Hwang, Q. Vera Liao, Su Lin Blodgett, Alexandra Olteanu, and Adam Trischler. 2024. "It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models. Preprint, arXiv:2411.13032.
  36. 36.Bernard J. Jansen and Amanda Spink. 2006. How are we searching the world wide web? a comparison of nine search engine transaction logs. Information Processing & Management, 42(1):248–263. Formal Methods for Information Retrieval.
  37. 37.Maryam Kamvar and Shumeet Baluja. 2006. A large scale study of wireless search behavior: Google mobile search. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’06, page 701–709, New York, NY, USA. Association for Computing Machinery.
  38. 38.Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, and Scott A. Hale. 2024. The prism alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models. Preprint, arXiv:2404.16019.
  39. 39.Katarzyna Kobalczyk, Nicolas Astorga, Tennison Liu, and Mihaela van der Schaar. 2025. Active task disambiguation with llms. arXiv preprint arXiv:2502.04485.
  40. 40.Mina Lee, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A. Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C. Guo, Md Naimul Hoque, Yewon Kim, Simon Knight, Seyed Parsa Neshaei, and 17 others. 2024. A design space for intelligent and interactive writing assistants. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY, USA. Association for Computing Machinery.
  41. 41.Mina Lee, Percy Liang, and Qian Yang. 2022. Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22, New York, NY, USA. Association for Computing Machinery.
  42. 42.Ge Li, Danai Vachtsevanou, Jérémy Lemée, Simon Mayer, and Jannis Strecker. 2024. Reader-aware writing assistance through reader profiles. In Proceedings of the 35th ACM Conference on Hypertext and Social Media, HT ’24, page 344–350, New York, NY, USA. Association for Computing Machinery.
  43. 43.Shuyue Stella Li, Jimin Mun, Faeze Brahman, Jonathan S Ilgen, Yulia Tsvetkov, and Maarten Sap. 2025. Aligning llms to ask good questions a case study in clinical reasoning. arXiv preprint arXiv:2502.14860.
  44. 44.Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel A. McFarland, and James Y. Zou. 2024a. Monitoring ai-modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews. In ICML.
  45. 45.Weixin Liang, Yaohui Zhang, Mihai Codreanu, Jiayu Wang, Hancheng Cao, and James Zou. 2025. The widespread adoption of large language model-assisted writing across society. Preprint, arXiv:2502.09747.
  46. 46.Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Yi Ding, Xinyu Yang, Kailas Vodrahalli, Siyu He, Daniel Scott Smith, Yian Yin, Daniel A. McFarland, and James Zou. 2024b. Can large language models provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI, 1(8):AIoa2400196.
  47. 47.Ying-Chun Lin, Jennifer Neville, Jack Stokes, Longqi Yang, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Xiaofeng Xu, Deepak Gupta, Sujay Kumar Jauhar, Xia Song, Georg Buscher, Saurabh Tiwary, Brent Hecht, and Jaime Teevan. 2024. Interpretable user satisfaction estimation for conversational systems with large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11100–11115, Bangkok, Thailand. Association for Computational Linguistics.
  48. 48.Lucie Charlotte Magister, Katherine Metcalf, Yizhe Zhang, and Maartje ter Hoeve. 2024. On the way to llm personalization: Learning to remember user conversations. arXiv preprint arXiv:2411.13405.
  49. 49.Moushumi Mahato, Avinash Kumar, Kartikey Singh, Javaid Nabi, Debojyoti Saha, and Krishna Singh. 2024. Exploring user dissatisfaction: Taxonomy of implicit negative feedback in virtual assistants. In Proceedings of the 21st International Conference on Natural Language Processing (ICON), pages 230–242, AU-KBC Research Centre, Chennai, India. NLP Association of India (NLPAI).
  50. 50.Chaitanya Malaviya, Joseph Chee Chang, Dan Roth, Mohit Iyyer, Mark Yatskar, and Kyle Lo. 2024. Contextualized evaluations: Taking the guesswork out of language model evaluations. arXiv preprint arXiv:2411.07237.
  51. 51.Yusuf Mehdi. 2023. Reinventing search with a new ai-powered microsoft bing and edge, your copilot for the web. Accessed: 3 March 2025.
  52. 52.Sheshera Mysore, Zhuoran Lu, Mengting Wan, Longqi Yang, Bahareh Sarrafzadeh, Steve Menezes, Tina Baghaee, Emmanuel Barajas Gonzalez, Jennifer Neville, and Tara Safavi. 2024. Pearl: Personalizing large language model writing assistants with generation-calibrated retrievers. In Proceedings of the 1st Workshop on Customizable NLP: Progress and Challenges in Customizing NLP for a Domain, Application, Group, or Individual (CustomNLP4U), pages 198–219, Miami, Florida, USA. Association for Computational Linguistics.
  53. 53.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  54. 54.Maria-Teresa De Rosa Palmini, Laura Wagner, and Eva Cetinic. 2024. Civiverse: A dataset for analyzing user engagement with open-source text-to-image models. Preprint, arXiv:2408.15261.
  55. 55.Shramay Palta, Nirupama Chandrasekaran, Rachel Rudinger, and Scott Counts. 2025. Speaking the right language: The impact of expertise alignment in user-ai interactions. Preprint, arXiv:2502.18685.
  56. 56.Chau Minh Pham, Simeng Sun, and Mohit Iyyer. 2024. Suri: Multi-constraint instruction following in long-form text generation. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 1722–1753, Miami, Florida, USA. Association for Computational Linguistics.
  57. 57.Martin J Pickering and Simon Garrod. 2013. An integrated theory of language production and comprehension. Behavioral and brain sciences, 36(4):329–347.
  58. 58.Chen Qu, Liu Yang, W. Bruce Croft, Johanne R. Trippas, Yongfeng Zhang, and Minghui Qiu. 2018. Analyzing and characterizing user intent in information-seeking conversations. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR ’18, page 989–992, New York, NY, USA. Association for Computing Machinery.
  59. 59.Jonathan Reades, Francesco Calabrese, and Carlo Ratti. 2009. Eigenplaces: analysing cities using the space–time structure of the mobile phone network. Environment and planning B: Planning and design, 36(5):824–836.
  60. 60.Vasile Rus and Panayiota Kendeou. 2025. Are llms actually good for learning? AI & SOCIETY, pages 1–2.
  61. 61.Rupak Sarkar, Bahareh Sarrafzadeh, Nirupama Chandrasekaran, Nagu Rangan, Philip Resnik, Longqi Yang, and Sujay Kumar Jauhar. 2025. Conversational user-ai intervention: A study on prompt rewriting for improved llm response generation. Preprint, arXiv:2503.16789.
  62. 62.Bahareh Sarrafzadeh, Sujay Kumar Jauhar, Michael Gamon, Edward Lank, and Ryen W. White. 2021. Characterizing stage-aware writing assistance for collaborative document authoring. Proc. ACM Hum.-Comput. Interact., 4(CSCW3).
  63. 63.Zejiang Shen, Tal August, Pao Siangliulue, Kyle Lo, Jonathan Bragg, Jeff Hammerbacher, Doug Downey, Joseph Chee Chang, and David Sontag. 2023. Beyond summarization: Designing ai support for real-world expository writing tasks. arXiv preprint arXiv:2304.02623.
  64. 64.Taiwei Shi, Zhuoer Wang, Longqi Yang, Ying-Chun Lin, Zexue He, Mengting Wan, Pei Zhou, Sujay Jauhar, Sihao Chen, Shan Xia, and 1 others. 2024. Wildfeedback: Aligning llms with in-situ user interactions and feedback. arXiv preprint arXiv:2408.15549.
  65. 65.Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett. 2024. A long way to go: Investigating length correlations in RLHF. In First Conference on Language Modeling.
  66. 66.Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell L. Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. 2024. Position: A roadmap to pluralistic alignment. In ICML.
  67. 67.Siddharth Suri, Scott Counts, Leijie Wang, Chacha Chen, Mengting Wan, Tara Safavi, Jennifer Neville, Chirag Shah, Ryen W. White, Reid Andersen, Georg Buscher, Sathish Manivannan, Nagu Rangan, and Longqi Yang. 2024. The use of generative search engines for knowledge work and complex tasks. Preprint, arXiv:2404.04268.
  68. 68.Alex Tamkin, Miles McCain, Kunal Handa, Esin Durmus, Liane Lovitt, Ankur Rathi, Saffron Huang, Alfred Mountfield, Jerry Hong, Stuart Ritchie, Michael Stern, Brian Clarke, Landon Goldberg, Theodore R. Sumers, Jared Mueller, William McEachen, Wes Mitchell, Shan Carter, Jack Clark, and 2 others. 2024. Clio: Privacy-preserving insights into real-world ai use. Preprint, arXiv:2412.13678.
  69. 69.Trang Tran and Mari Ostendorf. 2016. Characterizing the language of online communities and its relation to community reception. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1030–1035, Austin, Texas. Association for Computational Linguistics.
  70. 70.Johanne R. Trippas, Sara Fahad Dawood Al Lawati, Joel Mackenzie, and Luke Gallagher. 2024. What do users really ask large language models? an initial log analysis of google bard interactions in the wild. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’24, page 2703–2707, New York, NY, USA. Association for Computing Machinery.
  71. 71.Kailas Vodrahalli and James Zou. 2024. ArtWhisperer: A dataset for characterizing human-AI interactions in artistic creations. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 49627–49654. PMLR.
  72. 72.Qian Wan, Siying Hu, Yu Zhang, Piaohong Wang, Bo Wen, and Zhicong Lu. 2024. "it felt like having a second mind": Investigating human-ai co-creativity in prewriting with large language models. Proc. ACM Hum.-Comput. Interact., 8(CSCW1).
  73. 73.Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2024. Text embeddings by weakly-supervised contrastive pre-training. Preprint, arXiv:2212.03533.
  74. 74.Sterling Williams-Ceci, Maurice Jakesch, Advait Bhat, Kowe Kadoma, Lior Zalmanson, and Mor Naaman. 2024. Bias in ai autocomplete suggestions leads to attitude shift on societal issues. PsyArXiv.
  75. 75.Shujin Wu, Yi R. Fung, Cheng Qian, Jeonghwan Kim, Dilek Hakkani-Tur, and Heng Ji. 2025. Aligning LLMs with individual preferences via interaction. In Proceedings of the 31st International Conference on Computational Linguistics, pages 7648–7662, Abu Dhabi, UAE. Association for Computational Linguistics.
  76. 76.Rui Xin, Niloofar Mireshghallah, Shuyue Stella Li, Michael Duan, Hyunwoo Kim, Yejin Choi, Yulia Tsvetkov, Sewoong Oh, and Pang Wei Koh. 2024. A false sense of privacy: Evaluating textual data sanitization beyond surface-level privacy leakage. In Neurips Safe Generative AI Workshop 2024.
  77. 77.Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Li, Nora Bradford, Branda Sun, Tran Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, and Mark Warschauer. 2022. Fantastic questions and where to find them: FairytaleQA – an authentic dataset for narrative comprehension. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 447–460, Dublin, Ireland. Association for Computational Linguistics.
  78. 78.Lili Yao, Nanyun Peng, Ralph Weischedel, Kevin Knight, Dongyan Zhao, and Rui Yan. 2019. Plan-and-write: Towards better automatic storytelling. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01):7378–7385.
  79. 79.Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024a. Wildchat: 1m chatGPT interaction logs in the wild. In The Twelfth International Conference on Learning Representations.
  80. 80.Wenting Zhao, Alexander M Rush, and Tanya Goyal. 2024b. Challenges in trustworthy human evaluation of chatbots. arXiv preprint arXiv:2412.04363.
  81. 81.Zixin Zhao, Damien Masson, Young-Ho Kim, Gerald Penn, and Fanny Chevalier. 2025. Making the write connections: Linking writing support tools with writer’s needs. arXiv preprint arXiv:2502.13320.
  82. 82.Shengqi Zhu, Jeffrey M. Rzeszotarski, and David Mimno. 2025. Data paradigms in the era of llms: On the opportunities and challenges of qualitative data in the wild. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA ’25, New York, NY, USA. Association for Computing Machinery.

Citation

MLA
Mysore, S., et al. “Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 16819–46, https://doi.org/10.18653/v1/2025.emnlp-main.852.
APA
Mysore, S., Das, D., Cao, H., & Sarrafzadeh, B. (2025). Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 16819–16846. https://doi.org/10.18653/v1/2025.emnlp-main.852
Chicago
Mysore, S., D. Das, H. Cao, and B. Sarrafzadeh. 2025. “Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild”. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 16819–46. https://doi.org/10.18653/v1/2025.emnlp-main.852.
Harvard
Mysore, S. et al. (2025) “Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild”, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 16819–16846. Available at: https://doi.org/10.18653/v1/2025.emnlp-main.852.
Vancouver
1. Mysore S, Das D, Cao H, Sarrafzadeh B (2025) Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 16819–16846

BibTeX

@inproceedings{mysore-etal-2025-prototypical,
    title = "Prototypical Human-{AI} Collaboration Behaviors from {LLM}-Assisted Writing in the Wild",
    author = "Mysore, Sheshera  and
      Das, Debarati  and
      Cao, Hancheng  and
      Sarrafzadeh, Bahareh",
    editor = "Christodoulopoulos, Christos  and
      Chakraborty, Tanmoy  and
      Rose, Carolyn  and
      Peng, Violet",
    booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2025",
    address = "Suzhou, China",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.emnlp-main.852/",
    doi = "10.18653/v1/2025.emnlp-main.852",
    pages = "16819--16846",
    ISBN = "979-8-89176-332-6"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/