The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions

Siru OuyangShuohang WangYang LiuMing ZhongYizhu JiaoDan IterReid PryzantChenguang ZhuHeng JiJiawei Han

article2023EMNLP64 citations

Reveals a critical misalignment between academic NLP benchmarks and real-world needs by analyzing over 94,000 user-GPT interactions, identifying frequently requested yet neglected tasks such as planning, designing, and advising.

Listen

The rapid advancement and consumer adoption of large language models have fundamentally altered how people interact with artificial intelligence. However, standard academic benchmarks in natural language processing were largely developed before this widespread adoption, creating a potential mismatch between academic evaluation and real-world utility. Understanding whether current artificial intelligence models actually meet the everyday requirements of real users is essential for guiding future system development, evaluation standards, and deployment strategies.

The article investigates the divergence between traditional natural language processing research benchmarks and genuine user needs by analyzing real-world interactions with large language models. Its main objective is to identify emerging, overlooked task categories in daily usage and provide a clear roadmap for aligning artificial intelligence capabilities with human demands.

To conduct this evaluation, the researchers developed an automated annotation framework powered by GPT-4 using chain-of-thought prompting, demonstration sampling, and post-processing clustering. They analyzed 94,145 real-world user-model interactions from the public ShareGPT repository and compared them against 2,911 standard datasets from the Huggingface repository. A human assessment of 100 randomly sampled conversations confirmed high annotation reliability, achieving strong correctness scores and near-perfect inter-rater agreement.

The analysis revealed several critical findings regarding user behavior and benchmark misalignment. First, traditional benchmarks are overwhelmingly dominated by question answering and text classification, which make up more than two-thirds of established datasets, whereas real-world usage consists almost entirely of open-ended, free-form text generation. Second, established benchmarks draw more than 80% of their text from Wikipedia and news sources, while user queries reflect diverse, everyday domains spanning education, business, technology, and personal life. Third, while coding assistance (about 20%) and writing assistance (about 21%) represent the two largest query categories, a significant long-tail of overlooked tasks accounts for roughly 25% to 40% of user requests. These overlooked tasks include textual analysis (7.3%), evaluation based on custom rubrics (4.0%), open discussion and debate (3.8%), advice seeking (3.0%), travel and schedule planning (2.7%), and creative design (2.5%). Finally, error analysis indicates that state-of-the-art models exhibit substantial failure rates on these complex tasks—ranging from 40% to 65% for GPT-4 and 55% to 80% for GPT-3.5—frequently struggling with constraint satisfaction, spatial reasoning, emotional empathy, and hallucinated information.

These findings indicate that existing evaluation frameworks do not adequately capture the operational risks and functional demands of real-world deployment. In production environments, relying on models for complex planning, subjective evaluation, and personalized guidance introduces notable risks regarding safety, factual accuracy, and brand trust. The persistent failure of models to handle multi-constraint planning or demonstrate genuine empathy suggests that organizations cannot rely solely on standard benchmarks to predict real-world performance.

The article outlines clear recommendations for artificial intelligence developers and decision-makers. Research and development priorities must shift beyond basic classification and extraction toward complex reasoning, multi-turn interactivity, multimodal integration, dynamic world knowledge, and emotional perception. Dataset curation must also evolve to incorporate multi-format, realistic data rather than relying primarily on news articles and encyclopedia text. Furthermore, developers must carefully balance deep personalization with fairness to prevent biased outcomes across diverse user groups.

These conclusions carry high confidence regarding the general divergence between user requests and standard benchmarks, supported by systematic, human-verified categorization of a large dataset. Nevertheless, leaders should exercise appropriate caution: the findings rely on public, user-uploaded data that may overrepresent technically inclined users, and the primary analysis utilized an automated language model for large-scale annotation. Broader evaluations across enterprise-specific interaction logs are advised before implementing targeted downstream interventions.

Cover for The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions

Abstract

Recent progress in Large Language Models (LLMs) has produced models that exhibit remarkable performance across a variety of NLP tasks. However, it remains unclear whether the existing focus of NLP research accurately captures the genuine requirements of human users. This paper provides a comprehensive analysis of the divergence between current NLP research and the needs of real-world NLP applications via a large-scale collection of user-GPT conversations. We analyze a large-scale collection of real user queries to GPT. We compare these queries against existing NLP benchmark tasks and identify a significant gap between the tasks that users frequently request from LLMs and the tasks that are commonly studied in academic research. For example, we find that tasks such as “design” and “planning” are prevalent in user interactions but are largely neglected or different from traditional NLP benchmarks. We investigate these overlooked tasks, dissect the practical challenges they pose, and provide insights toward a roadmap to make LLMs better aligned with user needs.

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 2.1 ShareGPT
  • 2.2 Self-demonstrated annotation
  • 2.3 Human Evaluation
  • 2.4 Post-processing for analysis
  • 3 Overall Investigation
  • 3.1 Domain and Task Distribution
  • 3.2 Distribution Difference with Conventional Datasets
  • 4 Shifted and Overlooked Tasks
  • 4.1 Task of Providing Advice
  • 4.2 Task of Designing
  • 4.3 Task of Planning
  • 4.4 Task of Discussion
  • 4.5 Task of Analysis
  • 4.6 Task of Evaluation
  • 5 Emerging Trends and Challenges
  • 5.1 What trends are reflected in ShareGPT user queries?
  • 5.2 What challenges are proposed by trending and future tasks
  • 6 Conclusion and Future Works
  • Limitations
  • Ethics Statement
  • Acknowledgement
  • References
  • A Details in Post-processing
  • B Human Evaluation Interface
  • C Popular programming languages seen in Section 3
  • D Common topics and LLM performance for tasks listed in Section 4
  • E Concrete examples mentioned in Section 5

Knowls

  1. Knowl 1 — Real user requests differ sharply from benchmark task distributions

    empirical result

    The study annotated 94,145 split user–GPT conversations from ShareGPT and compared them with 2,911 licensed, English-language, non-multimodal, downloadable Hugging Face NLP datasets. ShareGPT requests were predominantly free-form generation or generation under user-specified constraints, whereas question answering and text classification together accounted for more than two-thirds of the benchmark dataset collection. The comparison therefore indicates a mismatch between common benchmark task formats and the types of assistance people request from an LLM in conversation.

  2. Knowl 2 — Six substantial ShareGPT task categories extend beyond conventional task settings

    empirical result

    The analysis identified six task categories that are common in everyday user requests but are not well captured by conventional NLP task definitions. Their reported approximate shares of ShareGPT are: advice, 3% (tailored guidance for a person’s circumstances); design, 2.5% (constructing specified objects, concepts, or activities); planning, 2.7% (sequencing steps toward an objective); discussion, 3.8% (exchanging views on a topic); analysis, 7.3% (examining a target whose scope or analytic aspects may be unspecified); and evaluation, 4% (judging an object against user-provided or open-ended criteria). The resulting requests tend to be open-ended and personalized: advice often asks for broad everyday guidance, design can concern conceptual or creative outputs, and planning can require coordinated real-world constraints. Discussion is dynamic, analysis may lack predefined targets, and evaluation can involve varied input formats and criteria rather than a fixed task metric.

  3. Knowl 3 — Everyday LLM use implies a cross-task capability roadmap

    model/method

    The paper’s proposed roadmap for serving ShareGPT-style requests spans several capabilities. Advice calls for emotion perception and personalization; design for multimodal input and interactive refinement; planning for stronger reasoning and world knowledge; discussion for personalized, interactive dialogue; analysis for multimodal handling and better reasoning; and evaluation for fairness and personalization. Across these tasks, the authors also observe requests tied to both personal and professional daily life and a broader range of user groups, including people of different ages, professions, and cultural backgrounds. They argue that models need to handle static facts and changing information, while balancing tailored assistance with fairness and attention to bias and misinformation.

  4. Knowl 4 — User and benchmark datasets differ in domain representation and source material

    empirical result

    In the ShareGPT domain distribution, technology accounts for about a quarter of queries (26.12% in the distribution chart); education, business, and language are also prominent and together contribute roughly another quarter. Technology is prominent in the Hugging Face dataset collection as well, but political, legal, personal-life, and film domains make up notable portions there. To estimate benchmark-domain coverage, the authors classified 10 randomly selected samples from each of the 2,911 datasets. They also used dataset-card metadata to estimate benchmark data sources and found that Wikipedia and news together account for over 80% of sources. The authors interpret this contrast as motivation to curate data for domains common in user requests and to include source material beyond Wikipedia and news.

  5. Knowl 5 — Coding and writing assistance are prevalent but include broader needs than standard benchmarks

    empirical result

    The paper reports coding assistance as 19.9% and writing assistance as 21.3% of ShareGPT requests. Coding requests include familiar settings such as code generation and debugging, but also higher-level program understanding, including code simplification and suggestions about design patterns. Writing help extends beyond grammatical or stylistic correction and story generation: article drafting and editing account for up to 5.1% of writing-assistance requests, and email editing for 2.6%. Users also request creative work in varied formats and help with procedural or how-to writing, indicating demand for explanatory and pedagogical support as well as generated text.

  6. Knowl 6 — GPT-4 self-demonstration pipeline annotates queries with domains, summaries, and task types

    algorithm

    The authors used GPT-4 to annotate each ShareGPT user query with a domain or topic, a one-sentence query summary, and one or more fine-grained task types. The procedure first prompted GPT-4 in a manual chain-of-thought format to identify the domain, summarize the query, and then infer task types. It next used in-context demonstrations sampled from an initial pool of 20 examples covering varied domains and task types; each query received three demonstrations. To expand that pool while limiting uncontrolled proliferation of free-form labels, the authors tracked task-type frequencies and added a sample containing a task type when its frequency exceeded 5% of the current samples. For annotation runs, three samples were concatenated per GPT-4 request and the temperature was 0.4. The annotation produced 13,783 task-type labels and 8,392 domain labels across the ShareGPT samples.

  7. Knowl 7 — Post-processing groups free-form labels into analyzable clusters

    model/method

    Because GPT-4 produced free-form domain and task phrases, the authors post-processed them to support distributional analysis. They first used frequency-based heuristics and an external thesaurus to combine synonyms, such as advice, tips, and suggestions. They then used similarity among the generated query-summary sentences to find nearest neighbors and union samples judged to belong to the same cluster. For low-frequency phrases not handled by the preceding steps, they assigned the phrase to the nearest existing cluster as an approximation, then manually removed unrelated results. This process converts varied surface labels into clusters suitable for comparing task and domain frequencies.

  8. Knowl 8 — Human assessment found high completeness and correctness for GPT-4 task labels

    empirical result

    Three graduate-student assessors with NLP and machine-learning research experience independently evaluated task-type annotations for 100 randomly selected ShareGPT samples. Completeness was scored 0 or 1, and correctness was scored 0, 1, or 2. Mean completeness was 0.95 with Fleiss’ κ = 0.96; mean correctness was 1.76 with κ = 0.83. None of the 100 samples received an incorrect correctness rating. These results support the quality of the generated task labels on this sample, while the evaluation directly assessed task types rather than validating every domain label or every annotation in the corpus.

  9. Knowl 9 — Manual case review reports substantial failure rates on the six highlighted tasks

    data/table

    The authors randomly selected 20 ShareGPT samples for each of the six highlighted task types and manually assessed GPT-3.5-turbo and GPT-4 responses. They report the following values as failure rates; the table is reproduced as reported. The task columns are advice, planning, design, discussion, analysis, and evaluation, respectively.

    Model Advice Planning Design Discussion Analysis Evaluation
    GPT-3.5-turbo 0.55 0.80 0.70 0.65 0.70 0.75
    GPT-4 0.40 0.60 0.65 0.45 0.50 0.45

    The review thus documents failures across every task category for both models in this small case-based sample. The accompanying examples identify constraint violations and factual errors in travel planning, infeasible or overlapping course schedules, limited visual-design ability, weak emotional attunement in advice, contradictions during philosophical discussion, and hallucinations in literary analysis.

  10. Knowl 10 — The study’s conclusions are bounded by dataset coverage and annotation costs

    limitation

    The authors identify three limits on generalizing their findings. First, ShareGPT and the selected Hugging Face datasets cannot represent the full breadth of real-world user needs or NLP benchmarks, and both collections change over time. Second, GPT-4 may produce inaccurate domain or task annotations, which can affect subsequent clustering and analysis despite the human assessment. Third, reproducing the annotation requires applying GPT-4 at large scale and was resource-intensive and time-consuming; the authors report that annotating the 94,000-plus ShareGPT samples took around 10 days under their setup.

Coverage note — The paper’s detailed per-topic examples and individual conversation transcripts are omitted because they illustrate, rather than add separate findings beyond, the task categories, capability gaps, and case-review results captured above.

References

  1. 1.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  2. 2.Robert H Bonczek, Clyde W Holsapple, and Andrew B Whinston. 1979. Computer-based support of organizational decision making. Decision Sciences, 10(2):268–291.
  3. 3.Francesco Borrelli, Dharmashankar Subramanian, Arvind U Raghunathan, and Lorenz T Biegler. 2006. Milp and nlp techniques for centralized trajectory planning of multiple unmanned air vehicles. In 2006 American Control Conference, pages 6–pp. IEEE.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Erik Cambria, Björn Schuller, Yunqing Xia, and Catherine Havasi. 2013. New avenues in opinion mining and sentiment analysis. IEEE Intelligent systems, 28(2):15–21.
  6. 6.Kathleen Carley. 1994. Extracting culture through textual analysis. Poetics, 22(4):291–312.
  7. 7.Jiuhai Chen, Lichang Chen, Heng Huang, and Tianyi Zhou. 2023. When do you need chain-of-thought prompting for chatgpt? arXiv preprint arXiv:2304.03262.
  8. 8.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  9. 9.Meng Chen, Ruixue Liu, Lei Shen, Shaozu Yuan, Jingyan Zhou, Youzheng Wu, Xiaodong He, and Bowen Zhou. 2020. The JDDC corpus: A large-scale multi-turn chinese dialogue dataset for e-commerce customer service. In Proceedings of The 12th Language Resources and Evaluation Conference, LREC 2020, Marseille, France, May 11-16, 2020, pages 459–466. European Language Resources Association.
  10. 10.Xiaoping Chen, Jiehui Jiang, Jianmin Ji, Guoqiang Jin, and Feng Wang. 2009. Integrating nlp with reasoning about actions for autonomous agents communicating with humans. In 2009 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology, volume 2, pages 137–140. IEEE.
  11. 11.Xinyun Chen, Chang Liu, and Dawn Song. 2018. Tree-to-tree neural networks for program translation. Advances in neural information processing systems, 31.
  12. 12.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing 3.54 with 90%* chatgpt quality.
  13. 13.Nancy Chinchor. 1992. MUC-4 evaluation metrics. In Fourth Message Uunderstanding Conference (MUC-4): Proceedings of a Conference Held in McLean, Virginia, June 16-18, 1992.
  14. 14.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  15. 15.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics.
  17. 17.William B. Dolan and Chris Brockett. 2005. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005).
  18. 18.Deepak Chowdary Edara, Lakshmi Prasanna Vanukuri, Venkatramaphanikumar Sistla, and Venkata Krishna Kishore Kolli. 2023. Sentiment analysis and text categorization of cancer medical records with lstm. Journal of Ambient Intelligence and Humanized Computing, 14(5):5309–5325.
  19. 19.CRN Estevam, MJ Rider, E Amorim, and JRS Mantovani. 2010. Reactive power dispatch and planning using a non-linear branch-and-bound algorithm. IET generation, transmission & distribution, 4(8):963–973.
  20. 20.Norman Fairclough. 1992. Discourse and text: Linguistic and intertextual analysis within discourse analysis. Discourse & society, 3(2):193–217.
  21. 21.Norman Fairclough. 2003. Analysing discourse: Textual analysis for social research. Psychology Press.
  22. 22.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  23. 23.Ingrid E Fisher, Margaret R Garnsey, and Mark E Hughes. 2016. Natural language processing in accounting, auditing and finance: A synthesis of the literature with a roadmap for future research. Intelligent Systems in Accounting, Finance and Management, 23(3):157–214.
  24. 24.Joseph L. Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological Bulletin, 76:378–382.
  25. 25.Robert W Floyd. 1963. Syntactic analysis and operator precedence. Journal of the ACM (JACM), 10(3):316–333.
  26. 26.Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. Gptscore: Evaluate as you desire. arXiv preprint arXiv:2302.04166.
  27. 27.Claude Goldenberg. 1992. Instructional conversations: Promoting comprehension through discussion. The Reading Teacher, 46(4):316–326.
  28. 28.Barbara J Grosz and Candace L Sidner. 1988. Plans for discourse. Technical report, BBN LABS INC CAMBRIDGE MA.
  29. 29.Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. The false promise of imitating proprietary llms. arXiv preprint arXiv:2305.15717.
  30. 30.Mohammad Kasra Habib. 2019. On the automated entity-relationship and schema design by natural language processing. Int. J. Eng. Sci, 8(11):42–48.
  31. 31.Zellig S Harris and Zellig S Harris. 1970. Discourse analysis. Springer.
  32. 32.Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al. 2021a. Measuring coding challenge competence with apps. arXiv preprint arXiv:2105.09938.
  33. 33.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021b. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR).
  34. 34.Carl Hewitt. 1969. Planner: A language for proving theorems in robots. In Proceedings of the 1st International Joint Conference on Artificial Intelligence, IJCAI’69, page 295–301, San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.
  35. 35.Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019. Codesearchnet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436.
  36. 36.Yizhu Jiao, Ming Zhong, Sha Li, Ruining Zhao, Siru Ouyang, Heng Ji, and Jiawei Han. 2023. Instruct and extract: Instruction tuning for on-demand information extraction. In the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics.
  37. 37.James M Keller, Michael R Gray, and James A Givens. 1985. A fuzzy k-nearest neighbor algorithm. IEEE transactions on systems, man, and cybernetics, (4):580–585.
  38. 38.Dan Klein and Christopher D Manning. 2003. Accurate unlexicalized parsing. In Proceedings of the 41st annual meeting of the association for computational linguistics, pages 423–430.
  39. 39.Thomas K Landauer, Peter W Foltz, and Darrell Laham. 1998. An introduction to latent semantic analysis. Discourse processes, 25(2-3):259–284.
  40. 40.Shu-Hsien Liao. 2005. Expert system methodologies and applications—a decade review from 1995 to 2004. Expert systems with applications, 28(1):93–103.
  41. 41.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics.
  42. 42.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. Gpteval: Nlg evaluation using 3.54 with better human alignment. arXiv preprint arXiv:2303.16634.
  43. 43.Tim Loughran and Bill McDonald. 2020. Textual analysis in finance. Annual Review of Financial Economics, 12:357–375.
  44. 44.Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2023. Chatgpt as a factual inconsistency evaluator for abstractive text summarization. arXiv preprint arXiv:2303.15621.
  45. 45.William C Mann and Sandra A Thompson. 1988. Rhetorical structure theory: Toward a functional theory of text organization. Text-interdisciplinary Journal for the Study of Discourse, 8(3):243–281.
  46. 46.Tomás Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In 1st International Conference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Workshop Track Proceedings.
  47. 47.Shachar Mirkin, Michal Jacovi, Tamar Lavee, Hong-Kwang Kuo, Samuel Thomas, Leslie Sager, Lili Kotlerman, Elad Venezian, and Noam Slonim. 2018. A recorded debating dataset. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  48. 48.Tetsuya Nasukawa and Jeonghee Yi. 2003. Sentiment analysis: Capturing favorability using natural language processing. In Proceedings of the 2nd international conference on Knowledge capture, pages 70–77.
  49. 49.Hwee Tou Ng, Siew Mei Wu, Ted Briscoe, Christian Hadiwinoto, Raymond Hendy Susanto, and Christopher Bryant. 2014. The CoNLL-2014 shared task on grammatical error correction. In Proceedings of the Eighteenth Conference on Computational Natural Language Learning: Shared Task, pages 1–14, Baltimore, Maryland. Association for Computational Linguistics.
  50. 50.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. Codegen: An open large language model for code with multi-turn program synthesis. ICLR.
  51. 51.OpenAI. 2023. 3.54 technical report. ArXiv, abs/2303.08774.
  52. 52.Siru Ouyang, Zhuosheng Zhang, and Hai Zhao. 2021. Dialogue graph modeling for conversational machine reading. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pages 3158–3169. Association for Computational Linguistics.
  53. 53.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  54. 54.Niels Pinkwart, Vincent Aleven, Kevin Ashley, and Collin Lynch. 2006. Toward legal argument instruction with graph grammars and collaborative filtering techniques. In Intelligent Tutoring Systems: 8th International Conference, ITS 2006, Jhongli, Taiwan, June 26-30, 2006. Proceedings 8, pages 227–236. Springer.
  55. 55.Marilyn Radford. 2003. Practice papers personal financial services in a digital age. Journal of Consumer Behaviour: An International Research Review, 2(3):287–295.
  56. 56.Marzieh Saeidi, Max Bartolo, Patrick Lewis, Sameer Singh, Tim Rocktäschel, Mike Sheldon, Guillaume Bouchard, and Sebastian Riedel. 2018. Interpretation of natural language rules in conversational machine reading. EMNLP.
  57. 57.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761.
  58. 58.A Carlisle Scott, William J Clancey, Randall Davis, and Edward H Shortliffe. 1977. Explanation capabilities of production-based consultation systems. Technical report, Stanford Univ CA Dept of Computer Science.
  59. 59.Tianxiao Shen, Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2017. Style transfer from non-parallel text by cross-alignment. Advances in neural information processing systems, 30.
  60. 60.Edward H Shortliffe, Stanton G Axline, Bruce G Buchanan, Thomas C Merigan, and Stanley N Cohen. 1973. An artificial intelligence program to advise physicians regarding antimicrobial therapy. Computers and Biomedical Research, 6(6):544–560.
  61. 61.Abhijeet R Sontakke and Amit Pimpalkar. 2014. A rule based graphical user interface to relational database using nlp. interaction, 5(6).
  62. 62.Michael Stubbs. 1996. Text and corpus analysis: Computer-assisted studies of language and culture. Blackwell Oxford.
  63. 63.Tian-Xiang Sun, Xiang-Yang Liu, Xi-Peng Qiu, and Xuan-Jing Huang. 2022. Paradigm shift in natural language processing. Machine Intelligence Research, 19(3):169–183.
  64. 64.Gerald Sussman and Terry Winograd. 1970. Micro-planner reference manual. Technical report, USA.
  65. 65.Jeffrey Svajlenko, Judith F. Islam, Iman Keivanloo, Chanchal K. Roy, and Mohammad Mamun Mia. 2014. Towards a big data curated benchmark of inter-project code clones. In 2014 IEEE International Conference on Software Maintenance and Evolution, pages 476–480.
  66. 66.Jiwei Tan, Xiaojun Wan, and Jianguo Xiao. 2017. Abstractive document summarization with a graph-based attentional neural model. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1171–1181.
  67. 67.Steven Tay et al. 2023. Sharegpt.
  68. 68.Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk. 2019. An empirical study on learning bug-fixing patches in the wild via neural machine translation. ACM Transactions on Software Engineering and Methodology (TOSEM), 28(4):1–29.
  69. 69.Karthik Valmeekam, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati. 2022. Large language models still can’t plan (a benchmark for llms on planning and reasoning about change). arXiv preprint arXiv:2206.10498.
  70. 70.Karthik Valmeekam, Sarath Sreedharan, Matthew Marquez, Alberto Olmo, and Subbarao Kambhampati. 2023. On the planning abilities of large language models (a critical investigation with a proposed benchmark). arXiv preprint arXiv:2302.06706.
  71. 71.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  72. 72.Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, and Huan Sun. 2022a. Towards understanding chain-of-thought prompting: An empirical study of what matters. arXiv preprint arXiv:2212.10001.
  73. 73.Qingyun Wang, Manling Li, Hou Pong Chan, Lifu Huang, Julia Hockenmaier, Chowdhary Girish, and Heng Ji. 2023. Multimedia generative script learning for task planning. In Proc. The 61st Annual Meeting of the Association for Computational Linguistics (ACL2023) Findings.
  74. 74.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022b. Self-instruct: Aligning language model with self generated instructions.
  75. 75.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS.
  76. 76.Yaqi Xie, Chen Yu, Tongyao Zhu, Jinbin Bai, Ze Gong, and Harold Soh. 2023. Translating natural language to planning goals with large-language models. arXiv preprint arXiv:2302.05128.
  77. 77.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244.
  78. 78.Da Yin, Xiao Liu, Fan Yin, Ming Zhong, Hritik Bansal, Jiawei Han, and Kai-Wei Chang. 2023. Dynosaur: A dynamic growth paradigm for instruction-tuning data curation. arXiv preprint arXiv:2305.14327.
  79. 79.Hainan Zhang, Yanyan Lan, Liang Pang, Jiafeng Guo, and Xueqi Cheng. 2019. Recosa: Detecting the relevant contexts with self-attention for multi-turn dialogue generation. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 3721–3730. Association for Computational Linguistics.
  80. 80.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  81. 81.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric Xing, et al. 2023. Lmsys-chat-1m: A large-scale real-world llm conversation dataset. arXiv preprint arXiv:2309.11998.
  82. 82.Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022. Towards a unified multi-dimensional evaluator for text generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 2023–2038. Association for Computational Linguistics.

Citation

MLA
Ouyang, S., et al. “The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 2375–93, https://doi.org/10.18653/v1/2023.emnlp-main.146.
APA
Ouyang, S., Wang, S., Liu, Y., Zhong, M., Jiao, Y., Iter, D., Pryzant, R., Zhu, C., Ji, H., & Han, J. (2023). The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2375–2393. https://doi.org/10.18653/v1/2023.emnlp-main.146
Chicago
Ouyang, S., S. Wang, Y. Liu, et al. 2023. “The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2375–93. https://doi.org/10.18653/v1/2023.emnlp-main.146.
Harvard
Ouyang, S. et al. (2023) “The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 2375–2393. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.146.
Vancouver
1. Ouyang S, Wang S, Liu Y, Zhong M, Jiao Y, Iter D, Pryzant R, Zhu C, Ji H, Han J (2023) The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 2375–2393

BibTeX

@inproceedings{ouyang-etal-2023-shifted,
    title = "The Shifted and The Overlooked: A Task-oriented Investigation of User-{GPT} Interactions",
    author = "Ouyang, Siru  and
      Wang, Shuohang  and
      Liu, Yang  and
      Zhong, Ming  and
      Jiao, Yizhu  and
      Iter, Dan  and
      Pryzant, Reid  and
      Zhu, Chenguang  and
      Ji, Heng  and
      Han, Jiawei",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.146/",
    doi = "10.18653/v1/2023.emnlp-main.146",
    pages = "2375--2393"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/