Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review

Rock Yuren PangHope SchroederKynnedy Simone SmithSolon BarocasZiang XiaoEmily TsengDanielle Bragg

article2025International Conference on Human Factors in Computing Systems81 citationsBest Paper Award

Establishes a comprehensive taxonomy of large language model adoption across 153 CHI papers, identifying their roles as research tools and simulated participants while providing practical questions to address prevalent validity and reproducibility concerns.

Listen

Large language models (LLMs) are rapidly altering computing research, reshaping both user-facing systems and internal research workflows. Because computing disciplines increasingly incorporate human feedback into artificial intelligence development, human-computer interaction (HCI) plays a crucial role in evaluating these technologies. However, there has been limited systematic understanding of how LLMs are being integrated across the discipline and what methodological challenges arise from their adoption. The article addresses this gap by evaluating the landscape of LLM research at the field's flagship venue, the ACM CHI Conference on Human Factors in Computing Systems (CHI), to determine how these tools are applied and whether their use meets scientific standards of rigor.

The article's main objective is to evaluate how LLM-related scholarship has expanded across application domains, contribution types, system roles, and self-reported limitations. To assess this, the authors conducted a systematic literature review of 153 generative LLM-related papers published in CHI proceedings from 2020 through 2024. Using an adapted PRISMA framework, the research team qualitatively coded the corpus across multiple dimensions using iterative human codebook development, achieving high interrater reliability across all categorized codes.

The investigation produced five central findings. First, LLM-focused publications experienced exponential growth, rising from 2 papers (0.26% of all conference papers) in 2020 to 115 papers (10.88%) in 2024. Second, researchers applied LLMs across 10 distinct domains, dominated by Communication and Writing (22.88%), Augmenting Capabilities (16.99%), and Education (14.38%). Third, contributions were heavily skewed toward empirical evaluations (98.70%) and artifact system building (61.44%), with few theoretical (5.23%), methodological (10.46%), or dataset (4.00%) advances. Fourth, the article mapped five primary roles for LLMs: as system engines (62.74%), subjects in user perception studies (23.53%), objects of study (9.80%), research tools (9.15%), and simulated human participants (7.19%). Fifth, 90.85% of papers acknowledged research validity limitations, and 84.98% relied on closed, proprietary models from the GPT family. Furthermore, 40.41% of prompt-based studies failed to disclose their exact prompts.

These findings indicate substantial risks for research reproducibility, validity, and scientific progress. Reliance on closed, proprietary application programming interfaces (APIs) means underlying models are non-deterministic, frequently updated without disclosure, and opaque regarding training data. The widespread omission of prompts further harms replicability. Additionally, using LLMs as simulated research participants creates serious validity and ethical issues, as models struggle to capture genuine human experiences and cannot replace human consent. While LLMs lower the barrier to rapidly prototyping software artifacts, the literature frequently mentions vague, unspecified model errors rather than conducting precise failure analyses, and only 22.88% of studies explicitly address broader societal consequences such as economic displacement or misinformation.

To ensure scientific integrity and responsible adoption, the article provides actionable guidance structured around critical reflection. Decision-makers and researchers must explicitly justify whether an LLM is necessary over simpler baselines, evaluate open versus closed models based on reproducibility requirements, and comprehensively disclose model versions, parameters, and full prompt templates. When LLMs serve as research tools or simulated users, rigorous human validation must be integrated rather than relying uncritically on synthetic outputs. Future research should prioritize underrepresented areas such as theoretical modeling, standardized benchmarking, and structured ethical impact assessments.

The conclusions are drawn from a comprehensive, rigorous qualitative review of peer-reviewed papers from the primary international venue in HCI. Readers should note that the scope was intentionally restricted to CHI conference proceedings and primarily focused on prompt-based generative models. As a result, the findings provide a highly confident diagnostic of cutting-edge human-centered computing research, though further analysis is warranted to examine specialized sub-disciplines and advanced technical configurations such as parameter-efficient fine-tuning and multi-agent systems.

arXiv: 2501.12557
Cover for Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review

Abstract

Large language models (LLMs) have been positioned to revolutionize HCI, by reshaping not only the interfaces, design patterns, and sociotechnical systems that we study, but also the research practices we use. To-date, however, there has been little understanding of LLMs' uptake in HCI. We address this gap via a systematic literature review of 153 CHI papers from 2020-24 that engage with LLMs. We taxonomize: (1) domains where LLMs are applied; (2) roles of LLMs in HCI projects; (3) contribution types; and (4) acknowledged limitations and risks. We find LLM work in 10 diverse domains, primarily via empirical and artifact contributions. Authors use LLMs in five distinct roles, including as research tools or simulated users. Still, authors often raise validity and reproducibility concerns, and overwhelmingly study closed models. We outline opportunities to improve HCI research with and on LLMs, and provide guiding questions for researchers to consider the validity and appropriateness of LLM-related work.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Literature Reviews in HCI
  • 2.2 Literature Reviews of LLM Papers
  • 2.3 How LLMs Can and Should Change Research
  • 3 Methods
  • 3.1 Data
  • 3.2 Analysis
  • 3.3 Research Positionality
  • 4 Results
  • 4.1 Application Domains
  • 4.2 Contribution Types
  • 4.3 LLM Roles
  • 4.4 Limitations
  • 4.4.1 LLM Performance (42.48%, N=65)
  • 4.4.2 Resource Limitation (28.76%, N=44)
  • 4.4.3 Research Validity(90.85%, N=139)
  • 4.4.4 Consequences (22.88%, N=35)
  • 5 Discussion
  • 5.1 Revealed Growth Opportunities for HCI
  • 5.1.1 Beyond language-based applications
  • 5.1.2 Beyond empirical and artifact contributions
  • 5.1.3 How LLMs impact prototyping standards
  • 5.2 Challenges: Validity, Reproducibility, and Consequences
  • 5.2.1 Proprietary LLMs raise reproducibility concerns
  • 5.2.2 LLM properties introduce additional research validity concerns
  • 5.2.3 Consequences, Risks, and Broader Impacts
  • 5.3 Guiding questions for HCI researchers using LLMs
  • 5.4 Limitations
  • References

Knowls

  1. Knowl 1 — Corpus Selection and Systematic Coding Protocol for CHI LLM Literature Review

    experimental setup

    To systematically analyze the uptake and impact of large language models (LLMs) in human-computer interaction (HCI), a PRISMA-adapted review methodology was conducted across full papers published in the ACM Conference on Human Factors in Computing Systems (CHI) proceedings from 2020 through 2024.

    The initial corpus of 4,077 full papers was filtered by matching titles and abstracts against a target keyword list: "language model", "llm", "foundation model", "foundational model", "GPT", "ChatGPT", "Claude", "Gemini", and "Falcon". This yielded 152 candidate papers. A stratified validation check on a random sample of 200 excluded papers identified 1 false negative (a 0.5% error rate), bringing the final analyzed corpus to N=153N=153 papers.

    Coding was executed using an iterative codebook development process over four rounds with randomly selected batches of 10 papers. Disagreements were resolved by consensus, achieving high interrater reliability assessed via Krippendorff's alpha:

    • Contribution Types: α=0.866\alpha = 0.866
    • Application Domains: α=0.849\alpha = 0.849
    • LLM Roles: α=0.773\alpha = 0.773
    • Limitations & Risks: α=0.887\alpha = 0.887

    The remaining papers were subsequently partitioned and coded independently by three researchers who developed the codebook, with ongoing consensus meetings.

  2. Knowl 2 — Taxonomy and Distribution of LLM Application Domains in HCI Research

    data/table

    Across the corpus of N=153N=153 LLM-related CHI papers published between 2020 and 2024, LLM applications spanned 10 distinct domains. Papers could be tagged with multiple domain codes.

    Application Domain Paper Count (NN) Percentage (%)
    Communication Writing 35 22.88%
    Augmenting Capabilities 26 16.99%
    Education 22 14.38%
    Responsible Computing 19 12.42%
    Programming 17 11.11%
    Reliability Validity of LLMs 16 10.46%
    Well-being Health 14 9.15%
    Design 13 8.50%
    Accessibility Aging 12 7.84%
    Creativity 9 5.88%
    • Communication & Writing (22.88%): Investigates text generation, editing, AI-mediated communication (AIMC), email composition, storytelling, and implicit bias in writing assistance.
    • Augmenting Capabilities (16.99%): Covers tools that enhance workplace and academic productivity, sensemaking over document collections, and physical-digital bridges (e.g., video conferencing and mixed reality).
    • Education (14.38%): Examines learner interactions, domain-specific tutoring (e.g., math, programming, vocabulary), and teacher-support tools.
    • Responsible Computing (12.42%): Addresses algorithmic fairness, representational harms, privacy preservation, misinformation, and tools fostering responsible AI prototyping.
    • Programming (11.11%): Focuses on code synthesis, code explanation, no-code/end-user programming interfaces, and prompt engineering tools.
    • Reliability & Validity of LLMs (10.46%): Evaluates model output correctness, benchmark pipelines, hallucination highlighting, and prompt structuring frameworks.
    • Well-being & Health (9.15%): Studies clinical decision support for healthcare providers and self-tracking, journaling, or cognitive restructuring tools for patients.
    • Design (8.50%): Investigates ideation support, UI prototyping, and visual design tools tailored for UX/UI designers.
    • Accessibility & Aging (7.84%): Addresses tools for blind and low-vision users, autistic individuals, augmentative and alternative communication (AAC), situational impairments, and older adults.
    • Creativity (5.88%): Scrutinizes computational creativity metrics and multi-model creative prototyping environments.
  3. Knowl 3 — Distribution of Research Contribution Types in LLM-Focused HCI Scholarship

    empirical result

    Classification of N=153N=153 LLM-related CHI papers (2020–2024) according to Wobbrock and Kientz's HCI contribution taxonomy demonstrates a strong concentration in empirical and artifact contributions, with limited representation across theoretical, methodological, dataset, and opinion types:

    • Empirical Contributions: 98.70% (N=151N=151). Almost all papers evaluated user perceptions, system performance, or human-LLM interactions via user studies or empirical data analyses.
    • Artifact Contributions: 61.44% (N=94N=94). More than six in ten papers contributed a novel system, interface, prototype, or framework. This rate is over 2.5 times the historical baseline rate for artifact contributions at CHI overall (24.50%), indicating that LLMs significantly lower technical barriers to prototyping research systems.
    • Methodological Contributions: 10.46% (N=16N=16). Includes LLM-augmented qualitative coding pipelines, UX evaluation frameworks, and synthetic data generation techniques for research methods.
    • Theoretical Contributions: 5.23% (N=8N=8). Encompasses conceptual interaction frameworks, design spaces for writing assistants, and cognitive modeling for prompting.
    • Dataset Contributions: 4.00% (N=6N=6). Dedicated benchmark datasets representing authentic user interactions remain rare.
    • Survey Contributions: 0.65% (N=1N=1).
    • Opinion Contributions: 0.00% (N=0N=0).
  4. Knowl 4 — Taxonomy and Prevalence of LLM Roles Across the HCI Research Lifecycle

    model/method

    LLMs occupy five distinct roles within HCI research project pipelines (coded across N=153N=153 papers, allowing multiple roles per paper):

    1. LLMs as System Engines (62.74%, N=96N=96): LLMs serve as core algorithmic components inside interactive systems, prototypes, or workflows. Common functions include generating content (text, code, dialogue, synthetic UI components) or processing and extracting structured insights from unstructured data (summarization, user intent inference).
    2. Users' Perceptions of LLMs (23.53%, N=36N=36): Studies evaluating how human populations perceive, adopt, or interact with public or standalone LLM tools (e.g., commercial chatbots) outside the context of evaluating a newly contributed research artifact.
    3. LLMs as Objects of Study (9.80%, N=15N=15): Direct scientific investigations into the underlying properties, internal mechanisms, training data compositions, bias distributions, ecological validity, and failure modes (such as hallucinations) of LLMs themselves.
    4. LLMs as Research Tools (9.15%, N=15N=15): LLMs deployed to automate or assist tasks traditionally executed by researchers, including qualitative coding, literature synthesis, data annotation, and synthetic research dataset generation.
    5. LLMs as Participants and Users (7.19%, N=11N=11): LLMs prompted to simulate human behaviors, act as personas, generate synthetic user research responses, or provide automated usability and heuristic feedback in place of human participants.
  5. Knowl 5 — Prevalence of Proprietary LLMs and Prompt Transparency Deficits at CHI

    empirical result

    Analysis of the N=153N=153 CHI LLM papers reveals widespread dependencies on closed-source model ecosystems alongside significant gaps in experimental disclosure:

    • Dominance of Closed GPT Models: 84.98% (N=130N=130) of all papers used or evaluated proprietary models from the OpenAI GPT family. Specifically, N=61N=61 papers utilized GPT-4, N=41N=41 used GPT-3.5, and N=26N=26 used GPT-3.
    • Lack of Model Justification: The majority of papers did not provide explicit methodological justification for their specific model choices, despite documented behavioral discrepancies between base, instruction-tuned, and chat-tuned model variants.
    • Prompt Disclosure Deficit: Of the 146 papers that conducted studies prompting LLMs, 40.41% (N=59N=59) failed to disclose the exact prompts or prompt templates used, either within the main text or in supplementary materials.

    These practices undermine scientific reproducibility in HCI because proprietary APIs are subject to unannounced backend updates, non-deterministic drift, and undocumented system prompt modifications.

  6. Knowl 6 — Taxonomy and Distribution of Acknowledged Limitations and Risks in CHI LLM Research

    data/table

    Analysis of limitation disclosures across N=153N=153 CHI papers identified four top-level categories and 22 sub-codes. Dedicated limitations sections were present in 94.77% (N=145N=145) of papers, while dedicated ethics or broader impact statements appeared in only 14.38% (N=22N=22).

    Limitation / Risk Category Paper Count (NN) Percentage (%)
    Research Validity 139 90.85%
    Validity Across Users Contexts – Prevalent
    Validity Across Models Prompts – Prevalent
    LLM Performance 65 42.48%
    Unspecified Errors and Biases 26 16.99%
    LLM Bias Toward Different Groups 17 11.11%
    Limited Data Coverage in Training Data 15 9.80%
    Hallucination 13 8.50%
    Non-deterministic Response 12 7.84%
    Resource Limitations 44 28.76%
    Lack of Evaluation Standards / Metrics 26 16.99%
    Computational Cost (e.g., GPU/Token Limits) 14 9.15%
    Financial Cost (e.g., API/Subscription Fees) 5 3.27%
    Consequences / Risks to Society 35 22.88%
    Economic Harms (e.g., Job Displacement) 17 11.11%
    Representational Harms 9 5.88%
    Misinformation Harms 4 2.61%
    Malicious Use 3 1.96%
    Hate Speech 3 1.96%
    Environmental Harms 1 0.65%

    The most frequent performance limitation code, Unspecified Errors and Biases (16.99%), reflects a widespread tendency of authors to acknowledge general AI error-making without characterizing the exact failure mechanics or behavioral boundaries of the underlying black-box models.

  7. Knowl 7 — Methodological Validation Practices for LLMs Used as Research Tools and Simulated Participants

    empirical result

    In the emerging methodologies where LLMs perform researcher tasks or simulate study subjects, authors almost universally provided textual justifications and empirical validation:

    • LLMs as Research Tools (N=15N=15): 100% (N=15N=15) of papers provided textual justification for model capabilities and task suitability supported by NLP literature citations. 93.33% (N=14N=14) conducted explicit experimental validations (e.g., comparative user studies against non-LLM baselines, manual verification of system-generated codes, or expert evaluation). 14 of 15 papers relied on human validation, while 1 used computational metrics.
    • LLMs as Participants & Users (N=11N=11): 90.91% (N=10N=10) provided both textual justification and experimental validation across user studies, computational evaluations, and human rating comparisons. The remaining 1 paper was an explicit critical audit demonstrating the empirical invalidity and risks of using LLMs to simulate human conversational partners.

    Despite high validation rates, studies utilizing LLM personas and simulated participants frequently face validity boundaries regarding the model's inability to represent authentic marginalized identities, lack of user consent, and failure to capture lived human experiences.

  8. Knowl 8 — Guiding Reflective Framework for LLM Integration Across HCI Research Lifecycles

    model/method

    A five-part reflective framework designed to guide HCI researchers in evaluating the validity, appropriateness, and societal impacts of incorporating LLMs into study designs:

    1. General Role Appropriateness (G1G_1): Identify the exact stage and role where the LLM enters the pipeline. Explicitly evaluate whether an LLM is necessary, or if simpler models, deterministic heuristics, or human involvement provide higher validity or lower resource overhead.
    2. Model Selection Justification (G2G_2): Weigh trade-offs between closed/proprietary APIs and open-weight models regarding reproducibility, participant data privacy consent, and system modularity.
    3. Disclosure of Models and Prompts (G3G_3): Fully document exact model checkpoint strings (e.g., gpt-4o-2024-08-06 rather than generic gpt-4o) and provide full prompt templates and parameter configurations.
    4. Role-Specific Limitation Scrutiny (G4G_4 / SS):
      • As System Engines: Assess whether lower-fidelity techniques (e.g., Wizard of Oz) are more scientifically appropriate than commercial LLM APIs; determine how prompt variations alter user study validity.
      • As Research Tools: Evaluate the risk of the "automation trap" where human validation effort exceeds the efficiency gained by LLM deployment; assess how labeling noise impacts empirical claims.
      • As Participants/Users: Account for identity misrepresentation, consent deprivation, and the fundamental gap between language modeling and embodied human behavior/opinion dynamics.
      • As Objects of Study: Clearly bound claims to specific model versions versus generalizable LLM properties.
      • For User Perception Studies: Control for user preconceptions, marketing hype, and sample representativeness.
    5. Societal Consequences Reflection (G5G_5): Explicitly evaluate environmental energy costs, participant privacy/data governance, and downstream labor or economic impacts.
  9. Knowl 9 — Limitations of the CHI Systematic Literature Review Methodology

    limitation

    The findings and taxonomy of this systematic literature review are subject to three primary methodological constraints:

    1. Corpus Scope and Sampling Filtering: The dataset is restricted to full papers published in the main CHI proceedings (2020–2024), excluding extended abstracts, doctoral consortium submissions, and specialized SIGCHI conferences (e.g., CSCW, UIST, DIS, ASSETS). Title and abstract keyword filtering can omit papers that use LLMs only in the method section without mentioning them in the abstract (observed as a 0.5% false-negative rate in validation sampling).
    2. Interface Modality Focus: The review predominantly captures studies that interact with LLMs via natural language prompting interfaces. Newer interaction and architectural paradigms—such as parameter-efficient fine-tuning (PEFT), dense LLM embedding spaces, and autonomous multi-agent systems—were less represented in the 2020–2024 CHI corpus.
    3. Qualitative Human Coding Constraints: Automated LLM coding of papers was tested in preliminary phases (using gpt-4o-2024-05-13) but was discarded because extracted themes were too generic without iterative human prompt engineering and codebook validation. The resulting manual, multi-coder qualitative process constrained the review's scale to CHI proceedings.

Coverage note — Individual qualitative case studies and specific system descriptions from surveyed papers were synthesized into the taxonomy, distribution tables, and methodological validation knowls rather than extracted as individual paper-level summaries.

References

  1. 1.William Agnew, A. Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R. McKee. 2024. The Illusion of Artificial Inclusion. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 286, 12 pages. https://doi.org/10.1145/3613904.3642703
  2. 2.Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. 2019. Guidelines for Human-AI Interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI '19). Association for Computing Machinery, New York, NY, USA, 1–13. https://doi.org/10.1145/3290605.3300233
  3. 3.Ian Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg, and Elena L. Glassman. 2024. ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 304, 18 pages. https://doi.org/10.1145/3613904.3642016
  4. 4.Marianne Aubin Le Quéré, Hope Schroeder, Casey Randazzo, Jie Gao, Ziv Epstein, Simon Tangi Perrault, David Mimno, Louise Barkhuus, and Hanlin Li. 2024. LLMs as Research Tools: Applications and Evaluations in HCI Data Work. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI EA '24). Association for Computing Machinery, New York, NY, USA, Article 479, 7 pages. https://doi.org/10.1145/3613905.3636301
  5. 5.Lisanne Bainbridge. 1983. Ironies of automation. Automatica 19, 6 (1983), 775–779. https://doi.org/10.1016/0005-1098(83)90046-8
  6. 6.Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017. The problem with bias: Allocative versus representational harms in machine learning. In 9th Annual conference of the special interest group for computing, information and society. New York, NY, 1.
  7. 7.Nikki M Barrington, Nithin Gupta, Basel Musmar, David Doyle, Nicholas Panico, Nikhil Godbole, Taylor Reardon, and Randy S D’Amico. 2023. A bibliometric analysis of the rise of ChatGPT in medical research. Medical Sciences 11, 3 (2023), 61.
  8. 8.Christoph Bartneck and Jun Hu. 2009. Scientometric analysis of the CHI proceedings. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Boston, MA, USA) (CHI '09). Association for Computing Machinery, New York, NY, USA, 699–708. https://doi.org/10.1145/1518701.1518810
  9. 9.Yasmine Belghith, Atefeh Mahdavi Goloujeh, Brian Magerko, Duri Long, Tom Mcklin, and Jessica Roberts. 2024. Testing, Socializing, Exploring: Characterizing Middle Schoolers’ Approaches to and Conceptions of ChatGPT. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 276, 17 pages. https://doi.org/10.1145/3613904.3642332
  10. 10.Victoria Bellotti, Maribeth Back, W. Keith Edwards, Rebecca E. Grinter, Austin Henderson, and Cristina Lopes. 2002. Making sense of sensing systems: five questions for designers and researchers. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Minneapolis, Minnesota, USA) (CHI '02). Association for Computing Machinery, New York, NY, USA, 415–422. https://doi.org/10.1145/503376.503450
  11. 11.Karim Benharrak, Tim Zindulka, Florian Lehmann, Hendrik Heuer, and Daniel Buschek. 2024. Writer-Defined AI Personas for On-Demand Feedback Generation. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 1049, 18 pages. https://doi.org/10.1145/3613904.3642406
  12. 12.Abeba Birhane, Atoosa Kasirzadeh, David Leslie, and Sandra Wachter. 2023. Science in the age of large language models. Nature Reviews Physics 5, 5 (2023), 277–280.
  13. 13.Danielle Bragg, Abraham Glasser, Fyodor Minakov, Naomi Caselli, and William Thies. 2022. Exploring collection of sign language videos through crowdsourcing. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2 (2022), 1–24.
  14. 14.Daniel Buschek, Martin Zürn, and Malin Eiband. 2021. The Impact of Multiple Parallel Phrase Suggestions on Email Input and Composition Behaviour of Native and Non-Native English Writers. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI '21). Association for Computing Machinery, New York, NY, USA, Article 732, 13 pages. https://doi.org/10.1145/3411764.3445372
  15. 15.Kelly Caine. 2016. Local Standards for Sample Size at CHI. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (San Jose, California, USA) (CHI '16). Association for Computing Machinery, New York, NY, USA, 981–992. https://doi.org/10.1145/2858036.2858498
  16. 16.Hancheng Cao, Yujie Lu, Yuting Deng, Daniel Mcfarland, and Michael S. Bernstein. 2023. Breaking Out of the Ivory Tower: A Large-scale Analysis of Patent Citations to HCI Research. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 760, 24 pages. https://doi.org/10.1145/3544548.3581108
  17. 17.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023. Quantifying Memorization Across Neural Language Models. arXiv:2202.07646 [cs.LG] https://arxiv.org/abs/2202.07646
  18. 18.Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu. 2024. Art or Artifice? Large Language Models and the False Promise of Creativity. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 30, 34 pages. https://doi.org/10.1145/3613904.3642731
  19. 19.Minsuk Chang, John Joon Young Chung, Katy Ilonka Gero, Ting-Hao Kenneth Huang, Dongyeop Kang, Vipul Raheja, Sarah Sterman, and Thiemo Wambsganss. 2024. Dark Sides: Envisioning, Understanding, and Preventing Harmful Effects of Writing Assistants - The Third Workshop on Intelligent and Interactive Writing Assistants. , 6 pages. https://doi.org/10.1145/3613905.3636312
  20. 20.Chaoran Chen, Weijun Li, Wenxin Song, Yanfang Ye, Yaxing Yao, and Toby Jia-Jun Li. 2024. An Empathy-Based Sandbox Approach to Bridge the Privacy Gap among Attitudes, Goals, Knowledge, and Behaviors. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 234, 28 pages. https://doi.org/10.1145/3613904.3642363
  21. 21.John Chen, Xi Lu, Yuzhou Du, Michael Rejtig, Ruth Bagley, Mike Horn, and Uri Wilensky. 2024. Learning Agent-based Modeling with LLM Companions: Experiences of Novices and Experts Using ChatGPT & NetLogo Chat. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 141, 18 pages. https://doi.org/10.1145/3613904.3642377
  22. 22.Alan Y. Cheng, Meng Guo, Melissa Ran, Arpit Ranasaria, Arjun Sharma, Anthony Xie, Khuyen N. Le, Bala Vinaithirthan, Shihe (Tracy) Luan, David Thomas Henry Wright, Andrea Cuadra, Roy Pea, and James A. Landay. 2024. Scientific and Fantastical: Creating Immersive, Culturally Relevant Learning Experiences with Augmented Reality and Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 275, 23 pages. https://doi.org/10.1145/3613904.3642041
  23. 23.Furui Cheng, Vilém Zouhar, Simran Arora, Mrinmaya Sachan, Hendrik Strobelt, and Mennatallah El-Assady. 2024. RELIC: Investigating Large Language Model Responses using Self-Consistency. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 647, 18 pages. https://doi.org/10.1145/3613904.3641904
  24. 24.Madiha Zahrah Choksi, Marianne Aubin Le Quéré, Travis Lloyd, Ruojia Tao, James Grimmelmann, and Mor Naaman. 2024. Under the (neighbor)hood: Hyperlocal Surveillance on Nextdoor. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 771, 22 pages. https://doi.org/10.1145/3613904.3641967
  25. 25.John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. 2022. TaleBrush: Sketching Stories with Generative Pretrained Language Models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI '22). Association for Computing Machinery, New York, NY, USA, Article 209, 19 pages. https://doi.org/10.1145/3491102.3501819
  26. 26.Andy Cockburn, Carl Gutwin, and Alan Dix. 2018. HARK No More: On the Preregistration of CHI Experiments. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI '18). Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3173574.3173715
  27. 27.Lucas Colusso, Ridley Jones, Sean A. Munson, and Gary Hsieh. 2019. A Translational Science Model for HCI. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI '19). Association for Computing Machinery, New York, NY, USA, 1–13. https://doi.org/10.1145/3290605.3300231
  28. 28.Michael Correll. 2019. Ethical Dimensions of Visualization Research. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI '19). Association for Computing Machinery, New York, NY, USA, 1–13. https://doi.org/10.1145/3290605.3300418
  29. 29.Andrea Cuadra, Maria Wang, Lynn Andrea Stein, Malte F. Jung, Nicola Dell, Deborah Estrin, and James A. Landay. 2024. The Illusion of Empathy? Notes on Displays of Emotion in Human-Computer Interaction. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 446, 18 pages. https://doi.org/10.1145/3613904.3642336
  30. 30.Nils Dahlbäck, Arne Jönsson, and Lars Ahrenberg. 1993. Wizard of Oz studies: why and how. In Proceedings of the 1st International Conference on Intelligent User Interfaces (Orlando, Florida, USA) (IUI '93). Association for Computing Machinery, New York, NY, USA, 193–200. https://doi.org/10.1145/169891.169968
  31. 31.Hai Dang, Sven Goller, Florian Lehmann, and Daniel Buschek. 2023. Choice Over Control: How Users Write with Large Language Models using Diegetic and Non-Diegetic Prompting. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 408, 17 pages. https://doi.org/10.1145/3544548.3580969
  32. 32.Fernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski-Fahey, Judith Amores Fernandez, and Jaron Lanier. 2024. LLMR: Real-time Prompting of Interactive Worlds using Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 600, 22 pages. https://doi.org/10.1145/3613904.3642579
  33. 33.Nicola Dell and Neha Kumar. 2016. The Ins and Outs of HCI for Development. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (San Jose, California, USA) (CHI '16). Association for Computing Machinery, New York, NY, USA, 2220–2232. https://doi.org/10.1145/2858036.2858081
  34. 34.Xiaohan Ding, Buse Carik, Uma Sushmitha Gunturi, Valerie Reyna, and Eugenia Ha Rim Rho. 2024. Leveraging Prompt-Based Large Language Models: Predicting Pandemic Health Decisions and Outcomes Through Social Media Language. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 443, 20 pages. https://doi.org/10.1145/3613904.3642117
  35. 35.Kimberly Do*, Rock Yuren Pang*, Jiachen Jiang, and Katharina Reinecke. 2023. “That’s important, but...”: How Computer Science Researchers Anticipate Unintended Consequences of Their Research Innovations. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 602, 16 pages. https://doi.org/10.1145/3544548.3581347
  36. 36.Peitong Duan, Jeremy Warner, Yang Li, and Bjoern Hartmann. 2024. Generating Automatic Feedback on UI Mockups with Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 6, 20 pages. https://doi.org/10.1145/3613904.3642782
  37. 37.Lizhou Fan, Lingyao Li, Zihui Ma, Sanggyu Lee, Huizi Yu, and Libby Hemphill. 2024. A Bibliometric Review of Large Language Models Research from 2017 to 2023. ACM Trans. Intell. Syst. Technol. (may 2024). https://doi.org/10.1145/3664930 Just Accepted.
  38. 38.Li Feng, Ryan Yen, Yuzhe You, Mingming Fan, Jian Zhao, and Zhicong Lu. 2024. Coprompt: Supporting prompt sharing and referring in collaborative natural language programming. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–21.
  39. 39.Sidong Feng, Suyu Ma, Han Wang, David Kong, and Chunyang Chen. 2024. MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–14.
  40. 40.Raymond Fok, Nedim Lipka, Tong Sun, and Alexa F Siu. 2024. Marco: Supporting Business Document Workflows via Collection-Centric Information Foraging with Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–20.
  41. 41.Liye Fu, Benjamin Newman, Maurice Jakesch, and Sarah Kreps. 2023. Comparing sentence-level suggestions to message-level suggestions in AI-mediated communication. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–13.
  42. 42.Yue Fu, Sami Foell, Xuhai Xu, and Alexis Hiniker. 2024. From Text to Self: Users’ Perception of AIMC Tools on Interpersonal Communication and Self. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 977, 17 pages. https://doi.org/10.1145/3613904.3641955
  43. 43.Takao Fujii, Katie Seaborn, and Madeleine Steeds. 2024. Silver-tongued and sundry: Exploring intersectional pronouns with chatgpt. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–14.
  44. 44.Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 (2023).
  45. 45.Katy Ilonka Gero, Tao Long, and Lydia B Chilton. 2023. Social Dynamics of AI Support in Creative Writing. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 245, 15 pages. https://doi.org/10.1145/3544548.3580782
  46. 46.Andreas Göldi, Thiemo Wambsganss, Seyed Parsa Neshaei, and Roman Rietsche. 2024. Intelligent Support Engages Writers Through Relevant Cognitive Processes. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–12.
  47. 47.Alice Good and Arunasalam Sambhanthan. 2014. Accessing web based health care and resources for mental health: interface design considerations for people experiencing mental illness. In Design, User Experience, and Usability. User Experience Design for Everyday Life Applications and Services: Third International Conference, DUXU 2014, Held as Part of HCI International 2014, Heraklion, Crete, Greece, June 22-27, 2014, Proceedings, Part III 3. Springer, 25–33.
  48. 48.Mary L Gray and Siddharth Suri. 2019. Ghost work: How to stop Silicon Valley from building a new global underclass. Eamon Dolan Books.
  49. 49.Jonathan Grudin. 2009. AI and HCI: Two fields divided by a common focus. AI magazine 30, 4 (2009), 48–48.
  50. 50.Ken Gu, Ruoxi Shang, Tim Althoff, Chenglong Wang, and Steven M Drucker. 2024. How Do Analysts Understand and Verify AI-Assisted Data Analyses?. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–22.
  51. 51.Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V Chawla, Olaf Wiest, and Xiangliang Zhang. 2024. Large language model based multi-agents: A survey of progress and challenges. arXiv preprint arXiv:2402.01680 (2024).
  52. 52.Justin Anthony Haegele and Samuel Hodge. 2016. Disability discourse: Overview and critiques of the medical and social models. Quest 68, 2 (2016), 193–206.
  53. 53.Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating large language models in generating synthetic hci research data: a case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–19.
  54. 54.Ariel Han, Xiaofei Zhou, Zhenyao Cai, Shenshen Han, Richard Ko, Seth Corrigan, and Kylie A Peppler. 2024. Teachers, Parents, and Students’ perspectives on Integrating Generative AI into Elementary Literacy Education. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–17.
  55. 55.Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang. 2024. Parameter-efficient fine-tuning for large models: A comprehensive survey. arXiv preprint arXiv:2403.14608 (2024).
  56. 56.Jeffrey T Hancock, Mor Naaman, and Karen Levy. 2020. AI-Mediated Communication: Definition, Research Agenda, and Ethical Considerations. Journal of Computer-Mediated Communication 25, 1 (01 2020), 89–100. https://doi.org/10.1093/jcmc/zmz022 arXiv:https://academic.oup.com/jcmc/article-pdf/25/1/89/32961176/zmz022.pdf
  57. 57.Yuexing Hao, Zeyu Liu, Robert N. Riter, and Saleh Kalantari. 2024. Advancing Patient-Centered Shared Decision-Making with AI Systems for Older Adult Cancer Patients. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 437, 20 pages. https://doi.org/10.1145/3613904.3642353
  58. 58.Zeyu He, Chieh-Yang Huang, Chien-Kuang Cornelia Ding, Shaurya Rohatgi, and Ting-Hao Kenneth Huang. 2024. If in a Crowdsourced Data Annotation Pipeline, a GPT-4. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 1040, 25 pages. https://doi.org/10.1145/3613904.3642834
  59. 59.Brent Hecht, Lauren Wilcox, Jeffrey P. Bigham, Johannes Schöning, Ehsan Hoque, Jason Ernst, Yonatan Bisk, Luigi De Russis, Lana Yarosh, Bushra Anjum, Danish Contractor, and Cathy Wu. 2021. It’s Time to Do Something: Mitigating the Negative Impacts of Computing Through a Change to the Peer Review Process. arXiv:2112.09544 [cs.CY] https://arxiv.org/abs/2112.09544
  60. 60.Michael A. Hedderich, Natalie N. Bazarova, Wenting Zou, Ryun Shim, Xinda Ma, and Qian Yang. 2024. A Piece of Theatre: Investigating How Teachers Design LLM Chatbots to Assist Adolescent Cyberbullying Education. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 668, 17 pages. https://doi.org/10.1145/3613904.3642379
  61. 61.Hendrik Heuer and Daniel Buschek. 2021. Methods for the Design and Evaluation of HCI+NLP Systems. In Proceedings of the First Workshop on Bridging Human–Computer Interaction and Natural Language Processing, Su Lin Blodgett, Michael Madaio, Brendan O’Connor, Hanna Wallach, and Qian Yang (Eds.). Association for Computational Linguistics, Online, 28–33. https://aclanthology.org/2021.hcinlp-1.5
  62. 62.Md Naimul Hoque, Tasfia Mashiat, Bhavya Ghai, Cecilia D. Shelton, Fanny Chevalier, Kari Kraus, and Niklas Elmqvist. 2024. The HaLLMark Effect: Supporting Provenance and Transparent Use of Large Language Models in Writing with Interactive Visualization. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 1045, 15 pages. https://doi.org/10.1145/3613904.3641895
  63. 63.Yihan Hou, Manling Yang, Hao Cui, Lei Wang, Jie Xu, and Wei Zeng. 2024. C2Ideas: Supporting Creative Interior Color Design Ideation with a Large Language Model. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 172, 18 pages. https://doi.org/10.1145/3613904.3642224
  64. 64.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685 (2021).
  65. 65.Forrest Huang, Gang Li, Tao Li, and Yang Li. 2024. Automatic Macro Mining from Interaction Traces at Scale. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–16.
  66. 66.Rong Huang, Haichuan Lin, Chuanzhang Chen, Kang Zhang, and Wei Zeng. 2024. PlantoGraphy: Incorporating iterative design process into generative artificial intelligence for landscape rendering. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–19.
  67. 67.Jessica Hullman. 2024. Status update on Twitter. https://x.com/JessicaHullman/status/1791645453223608422 Accessed: 2024-07-15.
  68. 68.Gautier Izacard and Édouard Grave. 2021. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 874–880.
  69. 69.Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naaman. 2023. Co-writing with opinionated language models affects users’ views. In Proceedings of the 2023 CHI conference on human factors in computing systems. 1–15.
  70. 70.JiWoong Jang, Sanika Moharana, Patrick Carrington, and Andrew Begel. 2024. “It’s the only thing I can trust”: Envisioning Large Language Model Use by Autistic Workers for Communication Assistance. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–18.
  71. 71.Ellen Jiang, Edwin Toh, Alejandra Molina, Kristen Olson, Claire Kayacik, Aaron Donsbach, Carrie J Cai, and Michael Terry. 2022. Discovering the Syntax and Strategies of Natural Language Programming with Generative Language Models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI '22). Association for Computing Machinery, New York, NY, USA, Article 386, 19 pages. https://doi.org/10.1145/3491102.3501870
  72. 72.Hyoungwook Jin, Seonghee Lee, Hyungyu Shin, and Juho Kim. 2024. Teach AI How to Code: Using Large Language Models as Teachable Agents for Programming Education. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–28.
  73. 73.Eunkyung Jo, Daniel A. Epstein, Hyunhoon Jung, and Young-Ho Kim. 2023. Understanding the Benefits and Challenges of Deploying Conversational AI Leveraging Large Language Models for Public Health Intervention. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 18, 16 pages. https://doi.org/10.1145/3544548.3581503
  74. 74.Mirabelle Jones, Christina Neumayer, and Irina Shklovski. 2023. Embodying the Algorithm: Exploring Relationships with Large Language Models Through Artistic Performance. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 654, 24 pages. https://doi.org/10.1145/3544548.3580885
  75. 75.Samia Kabir, David N. Udo-Imeh, Bonan Kou, and Tianyi Zhang. 2024. Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 935, 17 pages. https://doi.org/10.1145/3613904.3642596
  76. 76.Shivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi Li, and Hong Shen. 2024. "I’m categorizing LLM as a productivity tool": Examining ethics of LLM use in HCI research practices. arXiv:2403.19876 [cs.HC] https://arxiv.org/abs/2403.19876
  77. 77.Jin K. Kim, Michael Chua, Mandy Rickard, and Armando Lorenzo. 2023. ChatGPT and large language model (LLM) chatbots: The current state of acceptability and a proposal for guidelines on utilization in academic medicine. Journal of Pediatric Urology 19, 5 (2023), 598–604. https://doi.org/10.1016/j.jpurol.2023.05.018
  78. 78.Taewan Kim, Seolyeong Bae, Hyun Ah Kim, Su-woo Lee, Hwajung Hong, Chanmo Yang, and Young-Ho Kim. 2024. MindfulDiary: Harnessing Large Language Model to Support Psychiatric Patients’ Journaling. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–20.
  79. 79.Taewan Kim, Donghoon Shin, Young-Ho Kim, and Hwajung Hong. 2024. DiaryMate: Understanding User Perceptions and Experience in Human-AI Collaboration for Personal Journaling. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–15.
  80. 80.Tae Soo Kim, DaEun Choi, Yoonseo Choi, and Juho Kim. 2022. Stylette: Styling the Web with Natural Language. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI '22). Association for Computing Machinery, New York, NY, USA, Article 5, 17 pages. https://doi.org/10.1145/3491102.3501931
  81. 81.Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2024. EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 306, 21 pages. https://doi.org/10.1145/3613904.3642216
  82. 82.Hyung-Kwon Ko, Hyeon Jeon, Gwanmo Park, Dae Hyun Kim, Nam Wook Kim, Juho Kim, and Jinwook Seo. 2024. Natural language dataset generation framework for visualizations powered by large language models. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–22.
  83. 83.Charlotte Kobiella, Yarhy Said Flores López, Franz Waltenberger, Fiona Draxler, and Albrecht Schmidt. 2024. " If the Machine Is As Good As Me, Then What Use Am I?"–How the Use of ChatGPT Changes Young Professionals’ Perception of Productivity and Accomplishment. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–16.
  84. 84.Emily Kuang, Minghao Li, Mingming Fan, and Kristen Shinohara. 2024. Enhancing UX Evaluation Through Collaboration with Conversational AI Assistants: Effects of Proactive Dialogue and Timing. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 3, 16 pages. https://doi.org/10.1145/3613904.3642168
  85. 85.Tzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, and Haiyi Zhu. 2024. Wikibench: Community-driven data curation for ai evaluation on wikipedia. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–24.
  86. 86.Michelle S. Lam, Janice Teoh, James A. Landay, Jeffrey Heer, and Michael S. Bernstein. 2024. Concept Induction: Analyzing Unstructured Text with High-Level Concepts Using LLooM. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 766, 28 pages. https://doi.org/10.1145/3613904.3642830
  87. 87.Lane Lawley and Christopher Maclellan. 2024. VAL: Interactive Task Learning with GPT Dialog Parsing. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 5, 18 pages. https://doi.org/10.1145/3613904.3641915
  88. 88.Jungeun Lee, Suwon Yoon, Kyoosik Lee, Eunae Jeong, Jae-Eun Cho, Wonjeong Park, Dongsun Yim, and Inseok Hwang. 2024. Open Sesame? Open Salami! Personalizing Vocabulary Assessment-Intervention for Children via Pervasive Profiling and Bespoke Storybook Generation. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–32.
  89. 89.Mina Lee, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A. Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C. Guo, Md Naimul Hoque, Yewon Kim, Simon Knight, Seyed Parsa Neshaei, Antonette Shibani, Disha Shrivastava, Lila Shroff, Agnia Sergeyuk, Jessi Stark, Sarah Sterman, Sitong Wang, Antoine Bosselut, Daniel Buschek, Joseph Chee Chang, Sherol Chen, Max Kreminski, Joonsuk Park, Roy Pea, Eugenia Ha Rim Rho, Zejiang Shen, and Pao Siangliulue. 2024. A Design Space for Intelligent and Interactive Writing Assistants. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 1054, 35 pages. https://doi.org/10.1145/3613904.3642697
  90. 90.Mina Lee, Percy Liang, and Qian Yang. 2022. Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities. In Proceedings of the 2022 CHI conference on human factors in computing systems. 1–19.
  91. 91.Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, et al. 2022. Evaluating human-language model interaction. arXiv preprint arXiv:2212.09746 (2022).
  92. 92.Yoonjoo Lee, Hyeonsu B Kang, Matt Latzke, Juho Kim, Jonathan Bragg, Joseph Chee Chang, and Pao Siangliulue. 2024. PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Papers. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–19.
  93. 93.Florian Leiser, Sven Eckhardt, Valentin Leuthe, Merlin Knaeble, Alexander Maedche, Gerhard Schwabe, and Ali Sunyaev. 2024. Hill: A hallucination identifier for large language models. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–13.
  94. 94.Jiahao Nick Li, Yan Xu, Tovi Grossman, Stephanie Santosa, and Michelle Li. 2024. OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–22.
  95. 95.Zhuoyan Li, Chen Liang, Jing Peng, and Ming Yin. 2024. The Value, Benefits, and Concerns of Generative AI-Powered Assistance in Writing. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 1048, 25 pages. https://doi.org/10.1145/3613904.3642625
  96. 96.Zekun Li, Baolin Peng, Pengcheng He, Michel Galley, Jianfeng Gao, and Xifeng Yan. 2023. Guiding Large Language Models via Directional Stimulus Prompting. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 62630–62656. https://proceedings.neurips.cc/paper_files/paper/2023/file/c5601d99ed028448f29d1dae2e4a926d-Paper-Conference.pdf
  97. 97.Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D Manning, and James Y. Zou. 2024. Mapping the Increasing Use of LLMs in Scientific Papers. arXiv:2404.01268 [cs.CL] https://arxiv.org/abs/2404.01268
  98. 98.Q. Vera Liao, Hariharan Subramonyam, Jennifer Wang, and Jennifer Wortman Vaughan. 2023. Designerly Understanding: Information Needs for Model Transparency to Support Design Ideation for AI-Powered User Experience. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 9, 21 pages. https://doi.org/10.1145/3544548.3580652
  99. 99.Q. Vera Liao and Ziang Xiao. 2023. Rethinking Model Evaluation as Narrowing the Socio-Technical Gap. arXiv:2306.03100 [cs.HC] https://arxiv.org/abs/2306.03100
  100. 100.David Chuan-En Lin and Nikolas Martelaro. 2024. Jigsaw: Supporting Designers to Prototype Multimodal Applications by Chaining AI Foundation Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 4, 15 pages. https://doi.org/10.1145/3613904.3641920
  101. 101.Susan Lin, Jeremy Warner, J.D. Zamfirescu-Pereira, Matthew G Lee, Sauhard Jain, Shanqing Cai, Piyawat Lertvittayakumjorn, Michael Xuelin Huang, Shumin Zhai, Bjoern Hartmann, and Can Liu. 2024. Rambler: Supporting Writing With Speech via LLM-Assisted Gist Manipulation. In Proceedings of the CHI Conference on Human Factors in Computing Systems, Vol. 22. ACM, 1–19. https://doi.org/10.1145/3613904.3642217
  102. 102.Yupeng Lin and Zhonggen Yu. 2024. A bibliometric analysis of artificial intelligence chatbots in educational contexts. Interactive Technology and Smart Education 21, 2 (2024), 189–213.
  103. 103.Sebastian Linxen, Christian Sturm, Florian Brühlmann, Vincent Cassau, Klaus Opwis, and Katharina Reinecke. 2021. How WEIRD is CHI?. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI '21). Association for Computing Machinery, New York, NY, USA, Article 143, 14 pages. https://doi.org/10.1145/3411764.3445488
  104. 104.Michael Xieyang Liu, Advait Sarkar, Carina Negreanu, Benjamin Zorn, Jack Williams, Neil Toronto, and Andrew D Gordon. 2023. “What it wants me to say”: Bridging the abstraction gap between end-user programmers and code-generating large language models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–31.
  105. 105.Xingyu" Bruce" Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang" Anthony" Chen, and Ruofei Du. 2023. Visual captions: augmenting verbal communication with on-the-fly visuals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–20.
  106. 106.Xingyu Bruce Liu, Jiahao Nick Li, David Kim, Xiang’Anthony’ Chen, and Ruofei Du. 2024. Human I/O: Towards a Unified Approach to Detecting Situational Impairments. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–18.
  107. 107.Yong Liu, Jorge Goncalves, Denzil Ferreira, Bei Xiao, Simo Hosio, and Vassilis Kostakos. 2014. CHI 1994-2013: mapping two decades of intellectual progress through co-word analysis. In Proceedings of the SIGCHI conference on human factors in computing systems. 3553–3562.
  108. 108.Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. 2024. The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv:2408.06292 [cs.AI] https://arxiv.org/abs/2408.06292
  109. 109.Sasha Luccioni, Yacine Jernite, and Emma Strubell. 2024. Power hungry processing: Watts driving the cost of AI deployment?. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. 85–99.
  110. 110.Zilin Ma, Yiyang Mei, Yinru Long, Zhaoyuan Su, and Krzysztof Z Gajos. 2024. Evaluating the Experience of LGBTQ+ People Using Large Language Model Based Chatbots for Mental Health Support. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–15.
  111. 111.Kelly Mack, Emma McDonnell, Dhruv Jain, Lucy Lu Wang, Jon E. Froehlich, and Leah Findlater. 2021. What do we mean by “accessibility research”? A literature survey of accessibility papers in CHI and ASSETS from 1994 to 2019. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–18.
  112. 112.I Scott MacKenzie. 2024. Human-computer interaction: An empirical research perspective. (2024).
  113. 113.Damien Masson, Sylvain Malacria, Géry Casiez, and Daniel Vogel. 2024. Directgpt: A direct manipulation interface to interact with large language models. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–16.
  114. 114.David Maulsby, Saul Greenberg, and Richard Mander. 1993. Prototyping an intelligent agent through Wizard of Oz. In Proceedings of the INTERACT’93 and CHI’93 conference on Human factors in computing systems. 277–284.
  115. 115.Piotr Mirowski, Kory W Mathewson, Jaylen Pittman, and Richard Evans. 2023. Co-writing screenplays and theatre scripts with language models: Evaluation by industry professionals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–34.
  116. 116.David Moher, Alessandro Liberati, Jennifer Tetzlaff, Douglas G Altman, Prisma Group, et al. 2010. Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement. International journal of surgery 8, 5 (2010), 336–341.
  117. 117.Rajiv Movva, Sidhika Balachandar, Kenny Peng, Gabriel Agostini, Nikhil Garg, and Emma Pierson. 2024. Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 1223–1243.
  118. 118.Priyanka Nanayakkara, Jessica Hullman, and Nicholas Diakopoulos. 2021. Unpacking the expressed consequences of AI research in broader impact statements. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society. 795–806.
  119. 119.Sydney Nguyen, Hannah McLean Babe, Yangtian Zi, Arjun Guha, Carolyn Jane Anderson, and Molly Q Feldman. 2024. How Beginning Programmers and Code LLMs (Mis) read Each Other. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–26.
  120. 120.Rajvardhan Oak and Zubair Shafiq. 2024. Understanding Underground Incentivized Review Services. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–18.
  121. 121.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35 (2022), 27730–27744.
  122. 122.Shuyin Ouyang, Jie M Zhang, Mark Harman, and Meng Wang. 2023. LLM is Like a Box of Chocolates: the Non-determinism of ChatGPT in Code Generation. arXiv preprint arXiv:2308.02828 (2023).
  123. 123.Matthew J Page, Joanne E McKenzie, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoffmann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, et al. 2021. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. bmj 372 (2021).
  124. 124.Alexis Palmer, Noah A Smith, and Arthur Spirling. 2024. Using proprietary language models in academic research requires explicit justification. Nature Computational Science 4, 1 (2024), 2–3.
  125. 125.Rock Yuren Pang, Sebastin Santy, René Just, and Katharina Reinecke. 2024. BLIP: Facilitating the Exploration of Undesirable Consequences of Digital Technologies. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 290, 18 pages. https://doi.org/10.1145/3613904.3642054
  126. 126.Hyanghee Park and Daehwan Ahn. 2024. The Promise and Peril of ChatGPT in Higher Education: Opportunities, Challenges, and Design Implications. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 271, 21 pages. https://doi.org/10.1145/3613904.3642785
  127. 127.Savvas Petridis, Nicholas Diakopoulos, Kevin Crowston, Mark Hansen, Keren Henderson, Stan Jastrzebski, Jeffrey V Nickerson, and Lydia B Chilton. 2023. Anglekindling: Supporting journalistic angle ideation with large language models. In Proceedings of the 2023 CHI conference on human factors in computing systems. 1–16.
  128. 128.Alina Petukhova, Joao P Matos-Carvalho, and Nuno Fachada. 2024. Text clustering with LLM embeddings. arXiv preprint arXiv:2403.15112 (2024).
  129. 129.Heila Precel, Allison McDonald, Brent Hecht, and Nicholas Vincent. 2024. A Canary in the AI Coal Mine: American Jews May Be Disproportionately Harmed by Intellectual Property Dispossession in Large Language Model Training. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–17.
  130. 130.Mirjana Prpa, Giovanni Maria Troiano, Matthew Wood, and Yvonne Coady. 2024. Challenges and Opportunities of LLM-Based Synthetic Personae and Data in HCI. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI EA '24). Association for Computing Machinery, New York, NY, USA, Article 461, 5 pages. https://doi.org/10.1145/3613905.3636293
  131. 131.John Pruitt and Jonathan Grudin. 2003. Personas: practice and theory. In Proceedings of the 2003 conference on Designing for user experiences. 1–15.
  132. 132.Katharina Reinecke and Krzysztof Z. Gajos. 2015. LabintheWild: Conducting Large-Scale Online Experiments With Uncompensated Samples. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing (Vancouver, BC, Canada) (CSCW '15). Association for Computing Machinery, New York, NY, USA, 1364–1378. https://doi.org/10.1145/2675133.2675246
  133. 133.Katharina Reinecke, Tom Yeh, Luke Miratrix, Rahmatri Mardiko, Yuechen Zhao, Jenny Liu, and Krzysztof Z. Gajos. 2013. Predicting users’ first impressions of website aesthetics with a quantification of perceived visual complexity and colorfulness. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Paris, France) (CHI '13). Association for Computing Machinery, New York, NY, USA, 2049–2058. https://doi.org/10.1145/2470654.2481281
  134. 134.Anna Rogers, Niranjan Balasubramanian, Leon Derczynski, Jesse Dodge, Alexander Koller, Sasha Luccioni, Maarten Sap, Roy Schwartz, Noah A. Smith, and Emma Strubell. 2023. Closed AI Models Make Bad Baselines. https://hackingsemantics.xyz/2023/closed-baselines/
  135. 135.Kavous Salehzadeh Niksirat, Lahari Goswami, Pooja S. B. Rao, James Tyler, Alessandro Silacci, Sadiq Aliyu, Annika Aebli, Chat Wacharamanotham, and Mauro Cherubini. 2023. Changes in Research Ethics, Openness, and Transparency in Empirical Studies between CHI 2017 and CHI 2022. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 505, 23 pages. https://doi.org/10.1145/3544548.3580848
  136. 136.Joni Salminen, Chang Liu, Wenjing Pian, Jianxing Chi, Essi Häyhänen, and Bernard J Jansen. 2024. Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona Descriptions. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 510, 20 pages. https://doi.org/10.1145/3613904.3642036
  137. 137.Mark A Schmuckler. 2001. What is ecological validity? A dimensional analysis. Infancy 2, 4 (2001), 419–436.
  138. 138.Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2023. Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting. arXiv preprint arXiv:2310.11324 (2023).
  139. 139.Orit Shaer, Angelora Cooper, Osnat Mokryn, Andrew L Kun, and Hagit Ben Shoshan. 2024. AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–17.
  140. 140.Omar Shaikh, Valentino Emil Chai, Michele Gelfand, Diyi Yang, and Michael S Bernstein. 2024. Rehearsal: Simulating conflict to teach conflict resolution. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–20.
  141. 141.Ashish Sharma, Kevin Rushton, Inna Wanyin Lin, Theresa Nguyen, and Tim Althoff. 2024. Facilitating Self-Guided Mental Health Interventions Through Human-Language Model Interaction: A Case Study of Cognitive Restructuring. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 700, 29 pages. https://doi.org/10.1145/3613904.3642761
  142. 142.Nikhil Sharma, Q Vera Liao, and Ziang Xiao. 2024. Generative Echo Chamber? Effect of LLM-Powered Search Systems on Diverse Information Seeking. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–17.
  143. 143.Hong Shen, Tianshi Li, Toby Jia-Jun Li, Joon Sung Park, and Diyi Yang. 2023. Shaping the emerging norms of using large language models in social computing research. In Companion Publication of the 2023 Conference on Computer Supported Cooperative Work and Social Computing. 569–571.
  144. 144.Donghoon Shin, Lucy Lu Wang, and Gary Hsieh. 2024. From Paper to Card: Transforming Design Implications with Generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–15.
  145. 145.Jessie J Smith, Saleema Amershi, Solon Barocas, Hanna Wallach, and Jennifer Wortman Vaughan. 2022. Real ml: Recognizing, exploring, and articulating limitations of machine learning research. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 587–597.
  146. 146.Evropi Stefanidi, Marit Bentvelzen, Paweł W. Woźniak, Thomas Kosch, Mikołaj P. Woźniak, Thomas Mildner, Stefan Schneegass, Heiko Müller, and Jasmin Niess. 2023. Literature Reviews in HCI: A Review of Reviews. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 509, 24 pages. https://doi.org/10.1145/3544548.3581332
  147. 147.Konstantin R. Strömel, Stanislas Henry, Tim Johansson, Jasmin Niess, and Paweł W. Woźniak. 2024. Narrating Fitness: Leveraging Large Language Models for Reflective Fitness Tracker Data Interpretation. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 646, 16 pages. https://doi.org/10.1145/3613904.3642032
  148. 148.Hari Subramonyam, Roy Pea, Christopher Pondoc, Maneesh Agrawala, and Colleen Seifert. 2024. Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMs. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 1039, 19 pages. https://doi.org/10.1145/3613904.3642754
  149. 149.Jiao Sun, Tongshuang Wu, Yue Jiang, Ronil Awalegaonkar, Xi Victoria Lin, and Diyi Yang. 2022. Pretty princess vs. successful leader: Gender roles in greeting card messages. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–15.
  150. 150.Harini Suresh, Emily Tseng, Meg Young, Mary Gray, Emma Pierson, and Karen Levy. 2024. Participation in the age of foundation models. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (Rio de Janeiro, Brazil) (FAccT '24). Association for Computing Machinery, New York, NY, USA, 1609–1621. https://doi.org/10.1145/3630106.3658992
  151. 151.Maryam Taeb, Amanda Swearngin, Eldon Schoop, Ruijia Cheng, Yue Jiang, and Jeffrey Nichols. 2024. AXNav: Replaying Accessibility Tests from Natural Language. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 962, 16 pages. https://doi.org/10.1145/3613904.3642777
  152. 152.Mei Tan and Hari Subramonyam. 2024. More than Model Documentation: Uncovering Teachers’ Bespoke Information Needs for Informed Classroom Integration of ChatGPT. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 269, 19 pages. https://doi.org/10.1145/3613904.3642592
  153. 153.Yilin Tang, Liuqing Chen, Ziyu Chen, Wenkai Chen, Yu Cai, Yao Du, Fan Yang, and Lingyun Sun. 2024. EmoEden: Applying Generative Artificial Intelligence to Emotional Learning for Children with High-Function Autism. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 1001, 20 pages. https://doi.org/10.1145/3613904.3642899
  154. 154.Thitaree Tanprasert, Sidney S Fels, Luanne Sinnamon, and Dongwook Yoon. 2024. Debate Chatbots to Facilitate Critical Thinking on YouTube: Social Identity and Conversational Style Make A Difference. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 805, 24 pages. https://doi.org/10.1145/3613904.3642513
  155. 155.Nina Tran, Richard E Ladner, and Danielle Bragg. 2023. US Deaf Community Perspectives on Automatic Sign Language Translation. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility. 1–7.
  156. 156.Stephanie Valencia, Richard Cave, Krystal Kallarackal, Katie Seaver, Michael Terry, and Shaun K Kane. 2023. “The less I type, the better”: How AI Language Models can Enhance or Impede Communication for AAC Users. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–14.
  157. 157.Angelina Wang, Jamie Morgenstern, and John P Dickerson. 2024. Large language models cannot replace human participants because they cannot portray identity groups. arXiv preprint arXiv:2402.01908 (2024).
  158. 158.Bryan Wang, Gang Li, and Yang Li. 2023. Enabling Conversational Interaction with Mobile UI using Large Language Models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 432, 17 pages. https://doi.org/10.1145/3544548.3580895
  159. 159.Jiyao Wang, Haolong Hu, Zuyuan Wang, Song Yan, Youyu Sheng, and Dengbo He. 2024. Evaluating Large Language Models on Academic Literature Understanding and Review: An Empirical Study among Early-stage Scholars. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 12, 18 pages. https://doi.org/10.1145/3613904.3641917
  160. 160.Sitong Wang, Savvas Petridis, Taeahn Kwon, Xiaojuan Ma, and Lydia B Chilton. 2023. PopBlends: Strategies for conceptual blending with large language models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–19.
  161. 161.Xinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra, and Zhengjie Miao. 2024. Human-LLM Collaborative Annotation Through Effective Verification of LLM Labels. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 303, 21 pages. https://doi.org/10.1145/3613904.3641960
  162. 162.Zijie J Wang, Chinmay Kulkarni, Lauren Wilcox, Michael Terry, and Michael Madaio. 2024. Farsight: Fostering Responsible AI Awareness During AI Application Prototyping. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–40.
  163. 163.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652 (2021).
  164. 164.Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al. 2022. Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 214–229.
  165. 165.Joel Wester, Tim Schrills, Henning Pohl, and Niels van Berkel. 2024. “As an AI language model, I cannot”: Investigating LLM Denials of User Requests. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI '24). Association for Computing Machinery, New York, NY, USA, Article 979, 14 pages. https://doi.org/10.1145/3613904.3642135
  166. 166.Jules White, Sam Hays, Quchen Fu, Jesse Spencer-Smith, and Douglas C. Schmidt. 2024. ChatGPT Prompt Patterns for Improving Code Quality, Refactoring, Requirements Elicitation, and Software Design. Springer Nature Switzerland, Cham, 71–108. https://doi.org/10.1007/978-3-031-55642-5_4
  167. 167.Jacob O. Wobbrock and Julie A. Kientz. 2016. Research contributions in human-computer interaction. Interactions 23, 3 (apr 2016), 38–44. https://doi.org/10.1145/2907069
  168. 168.Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI '22). Association for Computing Machinery, New York, NY, USA, Article 385, 22 pages. https://doi.org/10.1145/3491102.3517582
  169. 169.Dirk U Wulff, Zak Hussain, and Rui Mata. 2024. The Behavioral and Social Sciences Need Open LLMs. https://doi.org/10.31219/osf.io/ybvzs
  170. 170.Wei Xiang, Hanfei Zhu, Suqi Lou, Xinli Chen, Zhenghua Pan, Yuping Jin, Shi Chen, and Lingyun Sun. 2024. SimUser: Generating Usability Feedback by Simulating Various Users Interacting with Mobile Applications. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–17.
  171. 171.Ziang Xiao, Wesley Hanwen Deng, Michelle S. Lam, Motahhare Eslami, Juho Kim, Mina Lee, and Q. Vera Liao. 2024. Human-Centered Evaluation and Auditing of Language Models. In Extended Abstracts of the 2024 CHI Conference on Human Factors in Computing Systems (CHI EA '24). Association for Computing Machinery, New York, NY, USA, Article 476, 6 pages. https://doi.org/10.1145/3613905.3636302
  172. 172.Anna Xygkou, Chee Siang Ang, Panote Siriaraya, Jonasz Piotr Kopecki, Alexandra Covaci, Eiman Kanjo, and Wan-Jou She. 2024. MindTalker: Navigating the Complexities of AI-Enhanced Social Engagement for People with Early-Stage Dementia. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–15.
  173. 173.Litao Yan, Alyssa Hwang, Zhiyuan Wu, and Andrew Head. 2024. Ivie: Lightweight anchored explanations of just-generated code. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–15.
  174. 174.Qian Yang, Nikola Banovic, and John Zimmerman. 2018. Mapping Machine Learning Advances from HCI Research to Reveal Starting Places for Design Innovation. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI '18). Association for Computing Machinery, New York, NY, USA, 1–11. https://doi.org/10.1145/3173574.3173704
  175. 175.Qian Yang, Justin Cranshaw, Saleema Amershi, Shamsi T Iqbal, and Jaime Teevan. 2019. Sketching nlp: A case study of exploring the right things to design with language intelligence. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–12.
  176. 176.Qian Yang, Yuexing Hao, Kexin Quan, Stephen Yang, Yiran Zhao, Volodymyr Kuleshov, and Fei Wang. 2023. Harnessing Biomedical Literature to Calibrate Clinicians’ Trust in AI Decision Support Systems. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 14, 14 pages. https://doi.org/10.1145/3544548.3581393
  177. 177.Qian Yang, Aaron Steinfeld, Carolyn Rosé, and John Zimmerman. 2020. Re-examining whether, why, and how human-AI interaction is uniquely difficult to design. In Proceedings of the 2020 chi conference on human factors in computing systems. 1–13.
  178. 178.Nur Yildirim, Hannah Richardson, Maria Teodora Wetscherek, Junaid Bajwa, Joseph Jacob, Mark Ames Pinnock, Stephen Harris, Daniel Coelho De Castro, Shruthi Bannur, Stephanie Hyland, et al. 2024. Multimodal healthcare AI: identifying and designing clinically relevant vision-language applications for radiology. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–22.
  179. 179.Chao Zhang, Xuechen Liu, Katherine Ziska, Soobin Jeon, Chi-Lin Yu, and Ying Xu. 2024. Mathemyths: leveraging large language models to teach mathematical language through Child-AI co-creative storytelling. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–23.
  180. 180.Lotus Zhang, Abigale Stangl, Tanusree Sharma, Yu-Yun Tseng, Inan Xu, Danna Gurari, Yang Wang, and Leah Findlater. 2024. Designing Accessible Obfuscation Support for Blind Individuals’ Visual Privacy Management. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–19.
  181. 181.Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, et al. 2023. Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792 (2023).
  182. 182.Zhiping Zhang, Michelle Jia, Hao-Ping Lee, Bingsheng Yao, Sauvik Das, Ada Lerner, Dakuo Wang, and Tianshi Li. 2024. “It’s a Fair Game”, or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–26.
  183. 183.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023).
  184. 184.Jiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G Parker, and Munmun De Choudhury. 2023. Synthetic lies: Understanding ai-generated misinformation and evaluating algorithmic and human solutions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–20.
  185. 185.Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024. Can large language models transform computational social science? Computational Linguistics 50, 1 (2024), 237–291.
  186. 186.Wazeer Deen Zulfikar, Samantha Chan, and Pattie Maes. 2024. Memoro: Using Large Language Models to Realize a Concise Interface for Real-Time Memory Augmentation. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–18.

Citation

MLA
Pang, R. Y., et al. “Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI Through a Systematic Literature Review”. arXiv, 2025, http://arxiv.org/abs/2501.12557v1.
APA
Pang, R. Y., Schroeder, H., Smith, K. S., Barocas, S., Xiao, Z., Tseng, E., & Bragg, D. (2025). Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review. arXiv. http://arxiv.org/abs/2501.12557v1
Chicago
Pang, R. Y., H. Schroeder, K. S. Smith, et al. 2025. “Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI Through a Systematic Literature Review”. arXiv. http://arxiv.org/abs/2501.12557v1.
Harvard
Pang, R.Y. et al. (2025) “Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.12557v1.
Vancouver
1. Pang RY, Schroeder H, Smith KS, Barocas S, Xiao Z, Tseng E, Bragg D (2025) Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review. arXiv

BibTeX

@article{pang2025understanding,
  title = {Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review},
  author = {Pang, Rock Yuren and Schroeder, Hope and Smith, Kynnedy Simone and Barocas, Solon and Xiao, Ziang and Tseng, Emily and Bragg, Danielle},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.12557v1},
  eprint = {2501.12557}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/