AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools
Nathalie Henry RicheAnna OffenwangerFrederic GmeinerDavid BrownHugo RomatMichel PahudNicolai MarquardtKori InkpenKen Hinckley
Introduces an interaction framework that turns text prompts into reusable, direct-manipulation graphical tools, allowing designers to resolve ambiguous intents and dynamically generate custom controls for generative AI systems.
Mainstream interfaces for generative artificial intelligence rely heavily on linear, text-based chat prompts. This conversational paradigm creates significant friction for creative workflows, as users often struggle to articulate nuanced concepts in words, explore divergent ideas, iteratively steer model outputs, and resolve ambiguous creative goals.
The article demonstrates and evaluates an interaction paradigm called AI-Instruments, which embodies text prompts as reusable graphical interface objects. Its objective is to show how grounding artificial intelligence in instrumental interaction enables non-linear exploration, finer direct manipulation, and clearer intent formulation across generative workflows.
The authors developed a web-based prototype implementing four exemplar instruments—Fragments, Transformative Lenses, Generative Containers, and Fillable Brushes—integrated with large language and image diffusion models. They evaluated this interaction model through a qualitative study involving twelve experienced generative artificial intelligence users who performed content generation and styling tasks.
The evaluation revealed several key findings. First, reifying prompts into graphical objects allowed users to shift attention from crafting text to directly manipulating visual outcomes, enabling precise spatial control and scope selection. Second, surfacing multi-faceted prompt structures (reflection-in-intent) and multi-output variations (reflection-in-response) significantly reduced the cognitive burden of intent formulation and prompt disambiguation. Third, grounding instruments in existing visual examples or previous instruments enabled users to extract and transfer complex styles without requiring specialized descriptive vocabulary. Finally, participants found the instruments vastly superior for non-linear iterative workflows and localized image steering compared to linear prompting.
These findings indicate that moving beyond text prompts to direct manipulation instruments can reduce time lost to trial-and-error prompting and streamline creative production pipelines. Although graphical controls risk increasing interface clutter and computational overhead, combining them with generative meta-instruments provides structured abstraction and modular tool creation without hard-coding software functions.
Organizations developing generative artificial intelligence tools should adopt hybrid user interfaces that pair direct manipulation instruments with on-demand access to underlying text prompts, while incorporating robust versioning and history-tracking mechanisms. Future research should expand beyond image synthesis to evaluate AI-instruments across heterogeneous artifacts, including structured documents, code, and slide presentations.
Confidence in the conceptual framework is high, as all twelve participants validated the core principles. However, since the study relied on a small sample focused primarily on 2D image tasks with early-stage prototype probes, readers should view performance in broader multi-modal and textual domains as an area requiring further empirical testing.
- Paper: InstructPix2Pix: Learning to Follow Image Editing Instructions, Tim Brooks et al. (2023). Its instruction-based image-editing system establishes the text-to-image editing context that AI-Instruments reworks into direct, iterative controls.
- Paper: Imagic: Text-Based Real Image Editing with Diffusion Models, Bahjat Kawar et al. (2022). Imagic shows how text prompts can steer real-image edits, providing a useful foundation for understanding AI-Instruments’ focus on refining intent through interactive controls.
- Paper: StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery, Or Patashnik et al. (2021). StyleCLIP demonstrates semantic image manipulation through language, grounding the visual-generation task that AI-Instruments makes more directly steerable.
- Paper: Optimizing Prompts for Text-to-Image Generation, Yaru Hao et al. (2023). PROMPTIST’s automated prompt refinement clarifies the prompt-optimization approach that AI-Instruments contrasts with reusable, manipulable interface objects.
- Paper: DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection, Yuying Tang et al. (2026). DuoDrama extends reflection-centered AI interaction into screenplay refinement, coordinating perspectives to help writers iteratively evaluate and reshape creative work.
- Paper: InfoAlign: A Human-AI Co-Creation System for Storytelling with Infographics, Jielin Feng et al. (2026). InfoAlign carries AI-supported, editable creative workflows into infographic storytelling, where users selectively steer generated structure and visual composition.
