Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
Rock Yuren PangHope SchroederKynnedy Simone SmithSolon BarocasZiang XiaoEmily TsengDanielle Bragg
Establishes a comprehensive taxonomy of large language model adoption across 153 CHI papers, identifying their roles as research tools and simulated participants while providing practical questions to address prevalent validity and reproducibility concerns.
Large language models (LLMs) are rapidly altering computing research, reshaping both user-facing systems and internal research workflows. Because computing disciplines increasingly incorporate human feedback into artificial intelligence development, human-computer interaction (HCI) plays a crucial role in evaluating these technologies. However, there has been limited systematic understanding of how LLMs are being integrated across the discipline and what methodological challenges arise from their adoption. The article addresses this gap by evaluating the landscape of LLM research at the field's flagship venue, the ACM CHI Conference on Human Factors in Computing Systems (CHI), to determine how these tools are applied and whether their use meets scientific standards of rigor.
The article's main objective is to evaluate how LLM-related scholarship has expanded across application domains, contribution types, system roles, and self-reported limitations. To assess this, the authors conducted a systematic literature review of 153 generative LLM-related papers published in CHI proceedings from 2020 through 2024. Using an adapted PRISMA framework, the research team qualitatively coded the corpus across multiple dimensions using iterative human codebook development, achieving high interrater reliability across all categorized codes.
The investigation produced five central findings. First, LLM-focused publications experienced exponential growth, rising from 2 papers (0.26% of all conference papers) in 2020 to 115 papers (10.88%) in 2024. Second, researchers applied LLMs across 10 distinct domains, dominated by Communication and Writing (22.88%), Augmenting Capabilities (16.99%), and Education (14.38%). Third, contributions were heavily skewed toward empirical evaluations (98.70%) and artifact system building (61.44%), with few theoretical (5.23%), methodological (10.46%), or dataset (4.00%) advances. Fourth, the article mapped five primary roles for LLMs: as system engines (62.74%), subjects in user perception studies (23.53%), objects of study (9.80%), research tools (9.15%), and simulated human participants (7.19%). Fifth, 90.85% of papers acknowledged research validity limitations, and 84.98% relied on closed, proprietary models from the GPT family. Furthermore, 40.41% of prompt-based studies failed to disclose their exact prompts.
These findings indicate substantial risks for research reproducibility, validity, and scientific progress. Reliance on closed, proprietary application programming interfaces (APIs) means underlying models are non-deterministic, frequently updated without disclosure, and opaque regarding training data. The widespread omission of prompts further harms replicability. Additionally, using LLMs as simulated research participants creates serious validity and ethical issues, as models struggle to capture genuine human experiences and cannot replace human consent. While LLMs lower the barrier to rapidly prototyping software artifacts, the literature frequently mentions vague, unspecified model errors rather than conducting precise failure analyses, and only 22.88% of studies explicitly address broader societal consequences such as economic displacement or misinformation.
To ensure scientific integrity and responsible adoption, the article provides actionable guidance structured around critical reflection. Decision-makers and researchers must explicitly justify whether an LLM is necessary over simpler baselines, evaluate open versus closed models based on reproducibility requirements, and comprehensively disclose model versions, parameters, and full prompt templates. When LLMs serve as research tools or simulated users, rigorous human validation must be integrated rather than relying uncritically on synthetic outputs. Future research should prioritize underrepresented areas such as theoretical modeling, standardized benchmarking, and structured ethical impact assessments.
The conclusions are drawn from a comprehensive, rigorous qualitative review of peer-reviewed papers from the primary international venue in HCI. Readers should note that the scope was intentionally restricted to CHI conference proceedings and primarily focused on prompt-based generative models. As a result, the findings provide a highly confident diagnostic of cutting-edge human-centered computing research, though further analysis is warranted to examine specialized sub-disciplines and advanced technical configurations such as parameter-efficient fine-tuning and multi-agent systems.
- Paper: Ethical and social risks of harm from Language Models, Laura Weidinger et al. (2022). Provides a comprehensive taxonomy of ethical and social harms in language models, framing the specific risks and limitations analyzed in HCI literature.
- Paper: The Rise and Potential of Large Language Model Based Agents: A Survey, Zhiheng Xi et al. (2023). Surveys the conceptual foundations, interactive architectures, and operational roles of LLM-based autonomous agents, which directly underpin how CHI researchers deploy LLMs as interactive tools and simulated users.
- Paper: A Survey on Evaluation of Large Language Models, Yu-Chu Chang et al. (2023). Outlines foundational evaluation paradigms and benchmarks for large language models, providing necessary background for understanding the validity and reproducibility challenges identified in CHI studies.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). Establishes a structured framework for evaluating model trustworthiness and safety, clarifying the technical vulnerabilities and evaluation metrics relevant to human-computer interaction research.
- Paper: The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values, Hannah Kirk et al. (2023). Examines how human feedback and subjective values are integrated into LLMs, contextualizing how HCI researchers conceptualize human preferences and alignment.
- Paper: A Comprehensive Overview of Large Language Models, Humza Naveed et al. (2023). Provides a core overview of foundational LLM architectures, alignment methods, and tool-augmented capabilities that form the technical basis of HCI artifact contributions.
- Paper: Large Language Model-Brained GUI Agents: A Survey, Chaoyun Zhang et al. (2025). Directly extends the study of LLMs in HCI by providing an in-depth survey of LLM-powered graphical user interface agents that interact with software systems like human users.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). Deepens the source paper's discussion on LLM evaluation roles and validity concerns by systematically analyzing the methodologies and biases of using LLMs as automated judges.
- Paper: TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents, Bofei Zhang et al. (2026). Applies the concepts of user-simulating GUI agents by introducing multimodal trajectory learning for generalized graphical user interface interaction.
