Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration

Leixian ShenYifang WangHuamin QuXing XieHaotian Li

article2026International Conference on Human Factors in Computing Systems8 citationsHonorable Mention Award

Develops the Interaction-Augmented Instruction framework, an entity-relation model that combines text prompts with graphical user interface interactions into twelve composable paradigms to guide the systematic design of generative AI systems.

Listen

Free-form natural language prompts have become the primary method for interacting with generative artificial intelligence, yet text alone is frequently too ambiguous and coarse to communicate precise, fine-grained, or referential user intent. While combining text prompts with graphical user interface interactions such as clicking, brushing, and dragging offers a compelling solution, prior efforts have largely produced fragmented demonstrations without a shared theoretical foundation. The article addresses this gap by establishing the Interaction-Augmented Instruction model, a formal framework designed to systematically describe, differentiate, and generate interactive generative artificial intelligence interfaces.

To develop and validate the model, the authors followed an iterative, deductive modeling approach to identify a minimal and expressive set of entities and relations. They then conducted a structured qualitative analysis across a curated corpus of 66 representative interactive systems, mapping each system's workflow to directed paradigm graphs based on the model. This systematic review evaluated how diverse systems combine linguistic instructions with interface interactions across different stages of execution.

The analysis yielded several key findings. First, all examined workflows can be formally represented using six core entities: Human, Interaction, Text Prompt, Augmented Instruction, Generative AI, and Artifact. A critical finding is that treating Augmented Instruction as an explicit entity is necessary to capture how structured, non-linguistic constraints merge with text before model execution. Second, the article identified twelve recurring, composable atomic interaction paradigms categorized along two main dimensions: interaction timing (occurring before or after model invocation) and resource availability (prompt-only versus artifact-grounded). Third, through four diverse usage scenarios spanning data analysis, creative arts, and multi-agent systems, the authors demonstrated that the model effectively guides the extension, refinement, and generation of new interaction paradigms.

These findings indicate that user intent is best supported when interface designs strategically match task clarity and context. Pre-invocation and artifact-grounded paradigms reduce ambiguity and improve precision for deterministic tasks like editing and coding, whereas post-invocation paradigms provide essential scaffolding for open-ended exploration. Transitioning from pure prompt engineering to structured interaction-augmented instruction enhances system controllability, provenance tracking, and referential fidelity, potentially shortening trial-and-error cycles and reducing user cognitive burden.

Decision-makers and interface designers should use the article's framework as a structured design blueprint. Teams should evaluate whether user goals are known upfront or exploratory to select appropriate interaction timing, and leverage existing artifacts to anchor prompts whenever possible. Rather than treating interaction paradigms as rigid templates, developers should chain and remix these atomic patterns to support complex workflows. Organizations should also expand their evaluation metrics beyond artifact output quality to include controllability, convergence cycles, and referential accuracy.

The findings are bounded by certain limitations: the model currently focuses on single-user, single-agent interactions, the empirical corpus is limited to 66 tools, and validation relies on qualitative case synthesis rather than formal user testing. While confidence in the model's descriptive and structuring power is high, future empirical evaluations and deployments will be valuable to validate user adoption and performance in production environments.

No sufficiently relevant recommendations were found.

Cover for Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration

Abstract

Text prompt is the most common way for human-generative AI (GenAI) communication. Though convenient, it is challenging to convey fine-grained and referential intent. One promising solution is to combine text prompts with precise GUI interactions, like brushing and clicking. However, there lacks a formal model to capture synergistic designs between prompts and interactions, hindering their comparison and innovation. To fill this gap, via an iterative and deductive process, we develop the Interaction-Augmented Instruction (IAI) model, a compact entity-relation graph formalizing how the combination of interactions and text prompts enhances human-GenAI communication. With the model, we distill twelve recurring and composable atomic interaction paradigms from prior tools, verifying our model's capability to facilitate systematic design characterization and comparison. Four usage scenarios further demonstrate the model's utility in applying, refining, and innovating these paradigms. These results illustrate the IAI model's descriptive, discriminative, and generative power for shaping future GenAI systems.

Citation

MLA
Shen, L., et al. “Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration”. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 2026, pp. 1–1, https://doi.org/10.1145/3772318.3790505.
APA
Shen, L., Wang, Y., Qu, H., Xie, X., & Li, H. (2026). Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 1–21. https://doi.org/10.1145/3772318.3790505
Chicago
Shen, L., Y. Wang, H. Qu, X. Xie, and H. Li. 2026. “Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration”. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 1–21. https://doi.org/10.1145/3772318.3790505.
Harvard
Shen, L. et al. (2026) “Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration”, Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. ACM, pp. 1–21. Available at: https://doi.org/10.1145/3772318.3790505.
Vancouver
1. Shen L, Wang Y, Qu H, Xie X, Li H (2026) Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration. In: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. ACM, pp 1–21

BibTeX

@inproceedings{Shen_2026, series={CHI ’26}, title={Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration}, url={http://dx.doi.org/10.1145/3772318.3790505}, DOI={10.1145/3772318.3790505}, booktitle={Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems}, publisher={ACM}, author={Shen, Leixian and Wang, Yifang and Qu, Huamin and Xie, Xing and Li, Haotian}, year={2026}, month=Apr, pages={1–21}, collection={CHI ’26} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/