Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI Collaboration
Leixian ShenYifang WangHuamin QuXing XieHaotian Li
Develops the Interaction-Augmented Instruction framework, an entity-relation model that combines text prompts with graphical user interface interactions into twelve composable paradigms to guide the systematic design of generative AI systems.
Free-form natural language prompts have become the primary method for interacting with generative artificial intelligence, yet text alone is frequently too ambiguous and coarse to communicate precise, fine-grained, or referential user intent. While combining text prompts with graphical user interface interactions such as clicking, brushing, and dragging offers a compelling solution, prior efforts have largely produced fragmented demonstrations without a shared theoretical foundation. The article addresses this gap by establishing the Interaction-Augmented Instruction model, a formal framework designed to systematically describe, differentiate, and generate interactive generative artificial intelligence interfaces.
To develop and validate the model, the authors followed an iterative, deductive modeling approach to identify a minimal and expressive set of entities and relations. They then conducted a structured qualitative analysis across a curated corpus of 66 representative interactive systems, mapping each system's workflow to directed paradigm graphs based on the model. This systematic review evaluated how diverse systems combine linguistic instructions with interface interactions across different stages of execution.
The analysis yielded several key findings. First, all examined workflows can be formally represented using six core entities: Human, Interaction, Text Prompt, Augmented Instruction, Generative AI, and Artifact. A critical finding is that treating Augmented Instruction as an explicit entity is necessary to capture how structured, non-linguistic constraints merge with text before model execution. Second, the article identified twelve recurring, composable atomic interaction paradigms categorized along two main dimensions: interaction timing (occurring before or after model invocation) and resource availability (prompt-only versus artifact-grounded). Third, through four diverse usage scenarios spanning data analysis, creative arts, and multi-agent systems, the authors demonstrated that the model effectively guides the extension, refinement, and generation of new interaction paradigms.
These findings indicate that user intent is best supported when interface designs strategically match task clarity and context. Pre-invocation and artifact-grounded paradigms reduce ambiguity and improve precision for deterministic tasks like editing and coding, whereas post-invocation paradigms provide essential scaffolding for open-ended exploration. Transitioning from pure prompt engineering to structured interaction-augmented instruction enhances system controllability, provenance tracking, and referential fidelity, potentially shortening trial-and-error cycles and reducing user cognitive burden.
Decision-makers and interface designers should use the article's framework as a structured design blueprint. Teams should evaluate whether user goals are known upfront or exploratory to select appropriate interaction timing, and leverage existing artifacts to anchor prompts whenever possible. Rather than treating interaction paradigms as rigid templates, developers should chain and remix these atomic patterns to support complex workflows. Organizations should also expand their evaluation metrics beyond artifact output quality to include controllability, convergence cycles, and referential accuracy.
The findings are bounded by certain limitations: the model currently focuses on single-user, single-agent interactions, the empirical corpus is limited to 66 tools, and validation relies on qualitative case synthesis rather than formal user testing. While confidence in the model's descriptive and structuring power is high, future empirical evaluations and deployments will be valuable to validate user adoption and performance in production environments.
- Paper: AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools, Nathalie Henry Riche et al. (2025). Its prompt-as-graphical-instrument paradigm provides a direct precursor for understanding how the source formalizes combinations of text prompts and GUI interaction.
- Paper: Intent Tagging: Exploring Micro-Prompting Interactions for Supporting Granular Human-GenAI Co-Creation Workflows, Frederic Gmeiner et al. (2025). Its granular visual tagging approach grounds the source’s treatment of composable interaction patterns for expressing and refining human intent.
- Paper: PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs, Soroush Nasiriany et al. (2024). Its iterative visual prompting method offers a concrete example of pairing model instructions with precise on-screen interaction, clarifying the design space the source systematizes.
No sufficiently relevant recommendations were found.
