AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools

Nathalie Henry RicheAnna OffenwangerFrederic GmeinerDavid BrownHugo RomatMichel PahudNicolai MarquardtKori InkpenKen Hinckley

article2025International Conference on Human Factors in Computing Systems41 citationsHonorable Mention, CHI 2025

Introduces an interaction framework that turns text prompts into reusable, direct-manipulation graphical tools, allowing designers to resolve ambiguous intents and dynamically generate custom controls for generative AI systems.

Listen

Mainstream interfaces for generative artificial intelligence rely heavily on linear, text-based chat prompts. This conversational paradigm creates significant friction for creative workflows, as users often struggle to articulate nuanced concepts in words, explore divergent ideas, iteratively steer model outputs, and resolve ambiguous creative goals.

The article demonstrates and evaluates an interaction paradigm called AI-Instruments, which embodies text prompts as reusable graphical interface objects. Its objective is to show how grounding artificial intelligence in instrumental interaction enables non-linear exploration, finer direct manipulation, and clearer intent formulation across generative workflows.

The authors developed a web-based prototype implementing four exemplar instruments—Fragments, Transformative Lenses, Generative Containers, and Fillable Brushes—integrated with large language and image diffusion models. They evaluated this interaction model through a qualitative study involving twelve experienced generative artificial intelligence users who performed content generation and styling tasks.

The evaluation revealed several key findings. First, reifying prompts into graphical objects allowed users to shift attention from crafting text to directly manipulating visual outcomes, enabling precise spatial control and scope selection. Second, surfacing multi-faceted prompt structures (reflection-in-intent) and multi-output variations (reflection-in-response) significantly reduced the cognitive burden of intent formulation and prompt disambiguation. Third, grounding instruments in existing visual examples or previous instruments enabled users to extract and transfer complex styles without requiring specialized descriptive vocabulary. Finally, participants found the instruments vastly superior for non-linear iterative workflows and localized image steering compared to linear prompting.

These findings indicate that moving beyond text prompts to direct manipulation instruments can reduce time lost to trial-and-error prompting and streamline creative production pipelines. Although graphical controls risk increasing interface clutter and computational overhead, combining them with generative meta-instruments provides structured abstraction and modular tool creation without hard-coding software functions.

Organizations developing generative artificial intelligence tools should adopt hybrid user interfaces that pair direct manipulation instruments with on-demand access to underlying text prompts, while incorporating robust versioning and history-tracking mechanisms. Future research should expand beyond image synthesis to evaluate AI-instruments across heterogeneous artifacts, including structured documents, code, and slide presentations.

Confidence in the conceptual framework is high, as all twelve participants validated the core principles. However, since the study relied on a small sample focused primarily on 2D image tasks with early-stage prototype probes, readers should view performance in broader multi-modal and textual domains as an area requiring further empirical testing.

Cover for AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools

Abstract

Chat-based prompts respond with verbose linear-sequential texts, making it difficult to explore and refine ambiguous intents, back up and reinterpret, or shift directions in creative AI-assisted design work. AI-Instruments instead embody "prompts" as interface objects via three key principles: (1) Reification of user-intent as reusable direct-manipulation instruments; (2) Reflection of multiple interpretations of ambiguous user-intents (Reflection-in-intent) as well as the range of AI-model responses (Reflection-in-response) to inform design "moves" towards a desired result; and (3) Grounding to instantiate an instrument from an example, result, or extrapolation directly from another instrument. Further, AI-Instruments leverage LLM's to suggest, vary, and refine new instruments, enabling a system that goes beyond hard-coded functionality by generating its own instrumental controls from content. We demonstrate four technology probes, applied to image generation, and qualitative insights from twelve participants, showing how AI-Instruments address challenges of intent formulation, steering via direct manipulation, and non-linear iterative workflows to reflect and resolve ambiguous intents.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Human-AI Interaction
  • 2.2 Interaction Models in the Era of AI
  • 2.3 Content Generation and Creativity Support
  • 3 Example Walkthrough
  • 4 Instrumental Interaction with AI
  • 4.1 Reification of User Intent
  • 4.2 Reflection
  • 4.3 Grounding
  • 5 Examples of AI-Instruments
  • 5.1 Fragments
  • 5.2 Transformative Lenses
  • 5.3 Generative Containers
  • 5.4 Fillable Brushes
  • 5.5 Generated Instruments and Meta-Instruments
  • 6 Implementation
  • 7 Study
  • 7.1 Procedure
  • 7.2 Participants
  • 7.3 Material and Analysis
  • 7.4 Insights on Model Principles
  • 7.5 Challenges Addressed by AI-instruments
  • 8 Discussion and Future Work
  • 8.1 Revisiting and Extending the Instrumental Interaction Model
  • 8.2 Beyond Images, Applying AI-Instruments to Other Forms of Content
  • 9 Conclusion
  • References

Knowls

  1. Knowl 1 — AI-Instruments extend instrumental interaction for generative AI

    model/method

    An AI-instrument is an AI-powered mediator between a user and the content the user acts on. The model turns user intent and generated results into reusable interface objects that can be manipulated directly, supporting workflows that branch and recombine rather than proceeding only through sequential prompt-and-response exchanges. It is organized around three principles: reification of user intent, reflection on possible intents and model responses, and grounding instruments in examples or other instruments. In relation to classic instrumental interaction, it broadens reification from predefined commands to natural-language intent, reframes polymorphism as reflection over multiple interpretations, and extends reuse into grounding and the creation of instruments from other instruments.

  2. Knowl 2 — Reification makes intent directly manipulable across scope and abstraction

    definition

    Reification of user intent turns a user’s input and AI-generated output into graphical objects that can be reused and directly manipulated, rather than leaving interaction as a sequence of prompts that must be rewritten to make changes. Users can compose interactions in phrases, reuse earlier outputs as inputs, and use direct-manipulation gestures such as selections to specify the scope of an instruction. The intent represented by an instrument can be high-level or detailed; other instruments can help users move between these levels, for example by turning an abstract idea into specific editable attributes or combining specific attributes into a broader direction.

  3. Knowl 3 — Reflection exposes alternatives in both intent and AI responses

    definition

    Reflection is an AI-instrument capability that helps users steer generation by making alternatives visible in two places. Reflection-in-intent surfaces multiple facets or formulations of what the user may mean, helping the user clarify, revise, or add to their intent. Reflection-in-response presents multiple possible outputs or interpretations of an input so the user can assess the range of results and choose a next design action. An instrument can produce these alternatives by varying an input, changing model parameters, or asking the model to interpret different contextual information.

  4. Knowl 4 — Grounding instantiates instruments from examples and other instruments

    definition

    Grounding creates or informs an instrument using selected content, an example of a desired result, or another instrument. For visual content, AI segmentation can let users refer to elements of an example without naming them, or extract a property such as style for reuse on other content. Grounding can also produce an instrument from another instrument, allowing users to explore related controls or reuse an existing intent in a different context. This makes content and prior work sources of instrumental controls, not merely outputs.

  5. Knowl 5 — Fragments reify prompts as composable attribute cards

    model/method

    Fragments use a language model to decompose a text or image-generation prompt into interactive attribute cards, each represented as a pair [type,value][\text{type},\text{value}]—for example, a dimension such as style paired with a value such as illustration. The cards expose the conceptual dimensions attributed to the prompt, supporting reflection-in-intent. Users reveal fragments from content, add or remove cards by dragging them onto or away from content, and regenerate content after a change. They can request value variations for a fragment or additional fragment types; the interface organizes types in rows and value variations beneath the relevant type. Fragments can also be transferred between content objects to reuse selected aspects of one generation in another. To avoid overwhelming users with choices, fragments are surfaced on demand rather than all at once.

  6. Knowl 6 — Transformative Lenses generate content through spatial layering

    model/method

    A Transformative Lens is a movable, resizable interface object coupled to a generative prompt. Placing a lens over image content generates an image that synthesizes the prompt and the content beneath it; placing content over a lens can likewise recombine the two. Users can layer lenses and images to complete a partial image, compose multiple pieces of content, or regenerate an image in a different style. A blank lens containing a prompt can generate a backdrop around existing content. The probe regenerates after a two-second idle period following movement or resizing. Because generation is nondeterministic, removing a layer and recreating it does not necessarily restore the earlier result, unlike a reversible undo operation.

  7. Knowl 7 — Generative Containers present four parallel alternatives for exploration

    model/method

    A Generative Container associates a prompt with a fixed 2×22\times2 grid of four generated image variations. Users can edit the prompt or ground the container by dragging in example content or another instrument, such as a Fragment; the container then generates a new set of alternatives. Showing several responses together supports reflection-in-response, while generating variations of a Fragment can help users turn an abstract intention into more concrete editing options. The probe accepts one grounding example per container. Users can create and chain multiple containers, transferring results from one to another to explore successive changes and build a more refined direction.

  8. Knowl 8 — Fillable Brushes apply grounded prompts to selected image regions

    model/method

    A Fillable Brush is a persistent, pen-like instrument that applies an encapsulated prompt to a selected region of an existing image. Users can type or edit the prompt, or ground an empty brush in example content so that the system extracts descriptive content or style language for reuse. Applying the brush to an image scopes the requested change to the painted area; approximate strokes are converted into object masks using the Segment Anything Model, so precise tracing is not required. The probe then uses the mask for localized image generation. Brushes can be reused on multiple images, combined by drag-and-drop, and applied repeatedly to emphasize an effect.

  9. Knowl 9 — Meta-instruments organize and generate reusable instrument collections

    model/method

    Meta-instruments operate on instruments as well as on content. For example, a Generative Container can operate on Fragments to turn a vague attribute into more concrete alternatives. Since unrestricted generation of instruments could clutter the workspace, the authors introduce Palettes as meta-instruments for storing or generating collections of content and instruments. A Palette supports retrieval and reuse of controls, can group different instruments for a task, and can use collected work or examples as generative seeds. Such collections can also give users starting material when beginning a creative task would otherwise require a blank canvas.

  10. Knowl 10 — Qualitative study found perceived benefits alongside interaction and probe limitations

    empirical result

    Twelve participants who used generative AI weekly completed a 60-minute qualitative study involving image-generation and editing tasks. They first used a chat-based prompting probe with the same image-generation model, then tried the four AI-instrument probes after demonstrations. The tasks included generating images and changing their style; participants also considered which interaction technique best suited combining, splitting, iterating, editing by example, and expanding content. The researchers coded 156 participant statements. These findings describe participants’ perceptions of the probes, not a controlled measure of performance or evidence that the approach is superior in every task.

    All 12 participants identified reification, reflection-in-intent, reflection-in-response, and grounding as advantages over prompting. Participants made 38 comments about selection scope, emphasizing that brushes could target image regions and lenses could position or combine image elements. Reflection-in-intent drew 31 comments, with 24 associated with Fragments; reflection-in-response drew 31 comments, with 26 associated with Containers. Grounding drew 45 comments and was particularly valued for transferring visual properties that participants found difficult to describe in words. Participants also described easier exploration and iteration through parallel alternatives and reusable controls.

    The study surfaced qualifications as well as benefits. Four participants noted drawbacks of GUI interaction relative to chat, including accessibility, extra interactions, unexpected behavior, or the perception that prompting could be faster for a desired result. Participants also raised concerns that tightly grounded instruments could hinder larger divergent changes, and that prior edits could be lost when Fragments or Lenses regenerated content. Of the 48 comments describing aspects participants did not value, 28 concerned probe-specific limitations, including usability and regeneration timing. The results therefore support the perceived usefulness of the principles while leaving open how well they generalize beyond the image-focused probes and study tasks.

Coverage note — Detailed implementation-stack and model-configuration specifics are omitted because they support the prototypes rather than define the interaction contribution; exploration of text and heterogeneous artifacts is discussed as future work, not established by the image-focused study.

References

  1. 1.Adobe. 2024. Firefly. https://www.adobe.com/products/firefly
  2. 2.Shm Garanganao Almeda, J.D. Zamfirescu-Pereira, Kyu Won Kim, Pradeep Mani Rathnam, and Bjoern Hartmann. 2024. Prompting for Discovery: Flexible Sense-Making for AI Art-Making with Dreamsheets. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–17. https://doi.org/10.1145/3613904.3642858
  3. 3.Caroline Appert, Michel Beaudouin-Lafon, and Wendy E Mackay. 2005. Context matters: Evaluating interaction techniques with the CIS model. In People and Computers XVIII—Design for Life: Proceedings of HCI 2004. Springer, Springer, New York, NY, USA, 279–295.
  4. 4.Yavar Bathaee. 2017. The artificial intelligence black box and the failure of intent and causation. Harv. JL & Tech. 31 (2017), 889.
  5. 5.Michel Beaudouin-Lafon. 2000. Instrumental interaction: an interaction model for designing post-WIMP user interfaces. In Proceedings of the SIGCHI conference on Human factors in computing systems. Association for Computing Machinery, New York, NY, USA, 446–453.
  6. 6.Michel Beaudouin-Lafon, Susanne Bødker, and Wendy E Mackay. 2021. Generative theories of interaction. ACM Transactions on Computer-Human Interaction (TOCHI) 28, 6 (2021), 1–54.
  7. 7.Michel Beaudouin-Lafon and Wendy E Mackay. 2000. Reification, polymorphism and reuse: three principles for designing visual interfaces. In Proceedings of the working conference on Advanced visual interfaces. Association for Computing Machinery, New York, NY, USA, 102–109.
  8. 8.Eric A. Bier, Maureen C. Stone, Ken Pier, Ken Fishkin, Thomas Baudel, Matt Conway, William Buxton, and Tony DeRose. 1994. Toolglass and magic lenses: the see-through interface. In Conference Companion on Human Factors in Computing Systems (Boston, Massachusetts, USA) (CHI ’94). Association for Computing Machinery, New York, NY, USA, 445–446. https://doi.org/10.1145/259963.260447
  9. 9.Aras Bozkurt. 2024. Tell Me Your Prompts and I Will Make Them True: The Alchemy of Prompt Engineering and Generative AI. Open Praxis 16, 2 (April 2024), 111–118. https://doi.org/10.55982/openpraxis.16.2.661
  10. 10.Stephen Brade, Bryan Wang, Mauricio Sousa, Sageev Oore, and Tovi Grossman. 2023. Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language Models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23). Association for Computing Machinery, New York, NY, USA, 1–14. https://doi.org/10.1145/3586183.3606725
  11. 11.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712 (2023).
  12. 12.Bill Buxton. 2007. Sketching User Experiences: Getting the Design Right and the Right Design. Morgan Kaufmann, Burlington. https://doi.org/10.1016/B978-0-12-374037-3.X5043-3
  13. 13.William Buxton. 1983. Lexical and pragmatic considerations of input structures. SIGGRAPH Comput. Graph. 17, 1 (jan 1983), 31–37. https://doi.org/10.1145/988584.988586
  14. 14.William Buxton. 1995. Chunking and phrasing and the design of human-computer dialogues. In Readings in human–computer interaction. Elsevier, 494–499.
  15. 15.Tracy Diane Cassidy. 2008. Mood boards: Current practice in learning and teaching strategies and students’ understanding of the process. International journal of fashion design 1, 1 (2008), 43–54.
  16. 16.DaEun Choi, Sumin Hong, Jeongeon Park, John Joon Young Chung, and Juho Kim. 2024. CreativeConnect: Supporting Reference Recombination for Graphic Design Ideation with Generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–25. https://doi.org/10.1145/3613904.3642794
  17. 17.John Joon Young Chung and Eytan Adar. 2023. PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23). Association for Computing Machinery, New York, NY, USA, 1–17. https://doi.org/10.1145/3586183.3606777
  18. 18.John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. 2022. TaleBrush: Sketching Stories with Generative Pretrained Language Models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22). Association for Computing Machinery, New York, NY, USA, 1–19. https://doi.org/10.1145/3491102.3501819
  19. 19.Gregory Thomas Croisdale, John Joon Young Chung, Emily Huang, Gage Birchmeier, Xu Wang, and Anhong Guo. 2023. DeckFlow: A Card Game Interface for Exploring Generative Model Flows. In Adjunct Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23 Adjunct). Association for Computing Machinery, New York, NY, USA, 1–3. https://doi.org/10.1145/3586182.3615821
  20. 20.Nicholas Davis, Chih-Pin Hsiao, Yanna Popova, and Brian Magerko. 2015. An enactive model of creativity for computational collaboration and co-creation. Creativity in the digital age (2015), 109–133.
  21. 21.Umer Farooq, John M Carroll, and Craig H Ganoe. 2005. Supporting creativity in distributed scientific communities. In Proceedings of the 2005 ACM International Conference on Supporting Group Work. Association for Computing Machinery, New York, NY, USA, 217–226.
  22. 22.OpenJS Foundation. 2024. Node.js. https://nodejs.org/en
  23. 23.Charles Freeman, Sara Marcketti, and Elena Karpova. 2017. Creativity of images: using digital consensual assessment to evaluate mood boards. Fashion and Textiles 4 (2017), 1–15.
  24. 24.Katy Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K. Kummerfeld, and Elena L. Glassman. 2024. Supporting Sensemaking of Large Language Model Outputs at Scale. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 838, 21 pages. https://doi.org/10.1145/3613904.3642139
  25. 25.James J Gibson. 1977. The theory of affordances. Hilldale, USA 1, 2 (1977), 67–82.
  26. 26.Joshua Hailpern, Erik Hinterbichler, Caryn Leppert, Damon Cook, and Brian P. Bailey. 2007. TEAM STORM: demonstrating an interaction model for working with multiple ideas during creative group work. In Proceedings of the 6th ACM SIGCHI Conference on Creativity & Cognition (Washington, DC, USA) (C&C ’07). Association for Computing Machinery, New York, NY, USA, 193–202. https://doi.org/10.1145/1254960.1254987
  27. 27.Christine A. Halverson. 2002. Activity Theory and Distributed Cognition: Or What Does CSCW Need to DO with Theories? Computer Supported Cooperative Work (CSCW) 11, 1 (March 2002), 243–267. https://doi.org/10.1023/A:1015298005381
  28. 28.Björn Hartmann, Scott R. Klemmer, Michael Bernstein, Leith Abdulla, Brandon Burr, Avi Robinson-Mosher, and Jennifer Gee. 2006. Reflective physical prototyping through integrated design, test, and analysis. In Proceedings of the 19th Annual ACM Symposium on User Interface Software and Technology (Montreux, Switzerland) (UIST ’06). Association for Computing Machinery, New York, NY, USA, 299–308. https://doi.org/10.1145/1166253.1166300
  29. 29.Rorik Henrikson, Bruno De Araujo, Fanny Chevalier, Karan Singh, and Ravin Balakrishnan. 2016. Storeoboard: Sketching Stereoscopic Storyboards. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (San Jose, California, USA) (CHI ’16). Association for Computing Machinery, New York, NY, USA, 4587–4598. https://doi.org/10.1145/2858036.2858079
  30. 30.Robert R Hoffman, Gary Klein, and Shane T Mueller. 2018. Explaining explanation for “explainable AI”. In Proceedings of the human factors and ergonomics society annual meeting, Vol. 62. SAGE Publications Sage CA: Los Angeles, CA, SAGE Publications, Los Angeles, CA, USA, 197–201.
  31. 31.Kasper Hornbæk and Antti Oulasvirta. 2017. What is interaction?. In Proceedings of the 2017 CHI conference on human factors in computing systems. Association for Computing Machinery, New York, NY, USA, 5040–5052.
  32. 32.Eric Horvitz. 1999. Principles of mixed-initiative user interfaces. In Proceedings of the SIGCHI conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 159–166.
  33. 33.Edwin L Hutchins, James D Hollan, and Donald A Norman. 1985. Direct manipulation interfaces. Human–computer interaction 1, 4 (1985), 311–338.
  34. 34.Hilary Hutchinson, Wendy Mackay, Bo Westerlund, Benjamin B. Bederson, Allison Druin, Catherine Plaisant, Michel Beaudouin-Lafon, Stéphane Conversy, Helen Evans, Heiko Hansen, Nicolas Roussel, and Björn Eiderbäck. 2003. Technology probes: inspiring design for and with families. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Ft. Lauderdale, Florida, USA) (CHI ’03). Association for Computing Machinery, New York, NY, USA, 17–24. https://doi.org/10.1145/642611.642616
  35. 35.Takeo Igarashi and John F. Hughes. 2007. A suggestive interface for 3D drawing. In ACM SIGGRAPH 2007 Courses (San Diego, California) (SIGGRAPH ’07). Association for Computing Machinery, New York, NY, USA, 20–es. https://doi.org/10.1145/1281500.1281531
  36. 36.Robert JK Jacob, Audrey Girouard, Leanne M Hirshfield, Michael S Horn, Orit Shaer, Erin Treacy Solovey, and Jamie Zigelbaum. 2008. Reality-based interaction: a framework for post-WIMP interfaces. In Proceedings of the SIGCHI conference on Human factors in computing systems. Association for Computing Machinery, New York, NY, USA, 201–210.
  37. 37.Peiling Jiang, Jude Rayan, Steven P. Dow, and Haijun Xia. 2023. Graphologue: Exploring Large Language Model Responses with Interactive Diagrams. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23). Association for Computing Machinery, New York, NY, USA, 1–20. https://doi.org/10.1145/3586183.3606737
  38. 38.Juriy Zaytsev, Stefan Kienzle, and Andrea Bogazzi. 2024. Fabric.js. https://github.com/fabricjs/fabric.js
  39. 39.Tae Soo Kim, Yoonjoo Lee, Minsuk Chang, and Juho Kim. 2023. Cells, generators, and lenses: Design framework for object-oriented interaction with large language models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. Association for Computing Machinery, New York, NY, USA, 1–18.
  40. 40.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. 2023. Segment Anything. https://arxiv.org/abs/2304.02643v1
  41. 41.David Kirsh. 1995. The intelligent use of space. Artif. Intell. 73, 1–2 (feb 1995), 31–68. https://doi.org/10.1016/0004-3702(94)00017-U
  42. 42.Amy J Ko, Brad A Myers, and Htet Htet Aung. 2004. Six learning barriers in end-user programming systems. In 2004 IEEE Symposium on Visual Languages-Human Centric Computing. IEEE, IEEE, New York, NY, USA, 199–206.
  43. 43.Jingyi Li, Eric Rawn, Jacob Ritchie, Jasper Tran O’Leary, and Sean Follmer. 2023. Beyond the Artifact: Power as a Lens for Creativity Support Tools. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. Association for Computing Machinery, New York, NY, USA, 1–15.
  44. 44.Vivian Liu. 2023. Beyond Text-to-Image: Multimodal Prompts to Explore Generative AI. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems (CHI EA ’23). Association for Computing Machinery, New York, NY, USA, 1–6. https://doi.org/10.1145/3544549.3577043
  45. 45.Vivian Liu and Lydia B Chilton. 2022. Design Guidelines for Prompt Engineering Text-to-Image Generative Models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22). Association for Computing Machinery, New York, NY, USA, 1–23. https://doi.org/10.1145/3491102.3501825
  46. 46.Wendy E Mackay. 2002. Which interaction technique works when? Floating palettes, marking menus and toolglasses support different task strategies. In Proceedings of the Working Conference on Advanced Visual Interfaces. Association for Computing Machinery, New York, NY, USA, 203–208.
  47. 47.Wendy E Mackay and Anne-Laure Fayard. 1997. HCI, natural science and design: a framework for triangulation across disciplines. In Proceedings of the 2nd conference on Designing interactive systems: processes, practices, methods, and techniques. Association for Computing Machinery, New York, NY, USA, 223–234.
  48. 48.Atefeh Mahdavi Goloujeh, Anne Sullivan, and Brian Magerko. 2024. Is It AI or Is It Me? Understanding Users’ Prompt Journey with Text-to-Image Generative AI Tools. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–13. https://doi.org/10.1145/3613904.3642861
  49. 49.J. Marks, B. Andalman, P. A. Beardsley, W. Freeman, S. Gibson, J. Hodgins, T. Kang, B. Mirtich, H. Pfister, W. Ruml, K. Ryall, J. Seims, and S. Shieber. 1997. Design galleries: a general approach to setting parameters for computer graphics and animation. In Proceedings of the 24th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’97). ACM Press/Addison-Wesley Publishing Co., USA, 389–400. https://doi.org/10.1145/258734.258887
  50. 50.Damien Masson, Sylvain Malacria, Géry Casiez, and Daniel Vogel. 2024. DirectGPT: A Direct Manipulation Interface to Interact with Large Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–16. https://doi.org/10.1145/3613904.3642462
  51. 51.Donald A Norman. 1986. Cognitive engineering. User centered system design 31, 61 (1986), 2.
  52. 52.Changhoon Oh, Jungwoo Song, Jinhan Choi, Seonghyeon Kim, Sungwoo Lee, and Bongwon Suh. 2018. I Lead, You Help but Only with Enough Details: Understanding User Experience of Co-Creation with Artificial Intelligence. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–13. https://doi.org/10.1145/3173574.3174223
  53. 53.OpenAI. 2024. GPT-4o System Card. https://openai.com/index/gpt-4o-system-card/
  54. 54.Xiaohan Peng, Janin Koch, and Wendy E. Mackay. 2024. DesignPrompt: Using Multimodal Interaction for Design Exploration with Generative AI. In Proceedings of the 2024 ACM Designing Interactive Systems Conference (Copenhagen, Denmark) (DIS ’24). Association for Computing Machinery, New York, NY, USA, 804–818. https://doi.org/10.1145/3643834.3661588
  55. 55.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988 (2022).
  56. 56.Yvonne Rogers. 2004. New Theoretical Approaches for Human-Computer Interaction. Annual Review of Information Science and Technology (ARIST) 38 (2004), 87–143. ERIC Number: EJ678114.
  57. 57.Yvonne Rogers. 2012. HCI Theory: Classical, Modern, and Contemporary (1 ed.). Morgan & Claypool Publishers.
  58. 58.Hugo Romat, Nicolai Marquardt, Ken Hinckley, and Nathalie Henry Riche. 2022. Style Blink: exploring digital inking of structured information via handcrafted styling as a first-class object. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 1–14.
  59. 59.Karl Toby Rosenberg, Rubaiat Habib Kazi, Li-Yi Wei, Haijun Xia, and Ken Perlin. 2024. DrawTalking: Towards Building Interactive Worlds by Sketching and Speaking. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. ACM, Honolulu HI USA, 1–8. https://doi.org/10.1145/3613905.3651089
  60. 60.Runway. 2024. Motion Brushes. https://youtu.be/zQ3fQt8swEI
  61. 61.D. A. Schön. 1992. Designing as reflective conversation with the materials of a design situation. Know.-Based Syst. 5, 1 (mar 1992), 3–14. https://doi.org/10.1016/0950-7051(92)90020-G
  62. 62.Ben Shneiderman. 2007. Creativity support tools: accelerating discovery and innovation. Commun. ACM 50, 12 (Dec. 2007), 20–32. https://doi.org/10.1145/1323688.1323689
  63. 63.StabilityAI. 2024. CompVis/stable-diffusion · Hugging Face. https://huggingface.co/CompVis/stable-diffusion
  64. 64.Hari Subramonyam, Roy Pea, Christopher Pondoc, Maneesh Agrawala, and Colleen Seifert. 2024. Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMs. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–19. https://doi.org/10.1145/3613904.3642754
  65. 65.Sangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li, and Haijun Xia. 2024. Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–26. https://doi.org/10.1145/3613904.3642400
  66. 66.Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language Models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. ACM, San Francisco CA USA, 1–18. https://doi.org/10.1145/3586183.3606756
  67. 67.Michael Terry and Elizabeth D. Mynatt. 2002. Recognizing creative needs in user interface design. In Proceedings of the 4th Conference on Creativity & Cognition (Loughborough, UK) (C&C ’02). Association for Computing Machinery, New York, NY, USA, 38–44. https://doi.org/10.1145/581710.581718
  68. 68.Michael Terry and Elizabeth D. Mynatt. 2002. Side views: persistent, on-demand previews for open-ended tasks. In Proceedings of the 15th Annual ACM Symposium on User Interface Software and Technology (Paris, France) (UIST ’02). Association for Computing Machinery, New York, NY, USA, 71–80. https://doi.org/10.1145/571985.571996
  69. 69.Maryam Tohidi, William Buxton, Ronald Baecker, and Abigail Sellen. 2006. Getting the right design and the design right. In Proceedings of the SIGCHI conference on Human Factors in computing systems. 1243–1252.
  70. 70.Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. 2022. Phenaki: Variable length video generation from open domain textual descriptions. In International Conference on Learning Representations.
  71. 71.Yael Vinker, Andrey Voynov, Daniel Cohen-Or, and Ariel Shamir. 2023. Concept Decomposition for Visual Exploration and Inspiration. ACM Trans. Graph. 42, 6, Article 241 (Dec. 2023), 13 pages. https://doi.org/10.1145/3618315
  72. 72.Zhijie Wang, Yuheng Huang, Da Song, Lei Ma, and Tianyi Zhang. 2024. PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–21. https://doi.org/10.1145/3613904.3642803
  73. 73.Xingjiao Wu, Luwei Xiao, Yixuan Sun, Junhang Zhang, Tianlong Ma, and Liang He. 2022. A survey of human-in-the-loop for machine learning. Future Generation Computer Systems 135 (2022), 364–381.
  74. 74.Haijun Xia, Bruno Araujo, Tovi Grossman, and Daniel Wigdor. 2016. Object-oriented drawing. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 4610–4621.
  75. 75.Haijun Xia, Ken Hinckley, Michel Pahud, Xiao Tu, and Bill Buxton. 2017. WritLarge: Ink Unleashed by Unified Scope, Action, & Zoom. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (Denver, Colorado, USA) (CHI ’17). Association for Computing Machinery, New York, NY, USA, 3227–3240. https://doi.org/10.1145/3025453.3025664
  76. 76.Zihan Yan, Chunxu Yang, Qihao Liang, and Xiang ’Anthony’ Chen. 2023. XCreation: A Graph-based Crossmodal Generative Creativity Support Tool. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23). Association for Computing Machinery, New York, NY, USA, Article 48, 15 pages. https://doi.org/10.1145/3586183.3606826
  77. 77.Ryan Yen and Jian Zhao. 2024. Memolet: Reifying the Reuse of User-AI Conversational Memories. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. Association for Computing Machinery, New York, NY, USA, 1–22.
  78. 78.J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. 2023. Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Machinery, New York, NY, USA, 1–21. https://doi.org/10.1145/3544548.3581388
  79. 79.Chao Zhang, Cheng Yao, Jiayi Wu, Weijia Lin, Lijuan Liu, Ge Yan, and Fangtian Ying. 2022. StoryDrawer: A Child–AI Collaborative Drawing System to Support Children’s Creative Visual Storytelling. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22). Association for Computing Machinery, New York, NY, USA, 1–15. https://doi.org/10.1145/3491102.3501914
  80. 80.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding Conditional Control to Text-to-Image Diffusion Models. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, Paris, France, 3813–3824. https://doi.org/10.1109/ICCV51070.2023.00355

Citation

MLA
Riche, N., et al. “AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools”. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025, pp. 1–8, https://doi.org/10.1145/3706598.3714259.
APA
Riche, N., Offenwanger, A., Gmeiner, F., Brown, D., Romat, H., Pahud, M., Marquardt, N., Inkpen, K., & Hinckley, K. (2025). AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–18. https://doi.org/10.1145/3706598.3714259
Chicago
Riche, N., A. Offenwanger, F. Gmeiner, et al. 2025. “AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools”. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–18. https://doi.org/10.1145/3706598.3714259.
Harvard
Riche, N. et al. (2025) “AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools”, Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, pp. 1–18. Available at: https://doi.org/10.1145/3706598.3714259.
Vancouver
1. Riche N, Offenwanger A, Gmeiner F, Brown D, Romat H, Pahud M, Marquardt N, Inkpen K, Hinckley K (2025) AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools. In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, pp 1–18

BibTeX

@inproceedings{Riche_2025, series={CHI ’25}, title={AI-Instruments: Embodying Prompts as Instruments to Abstract & Reflect Graphical Interface Commands as General-Purpose Tools}, url={http://dx.doi.org/10.1145/3706598.3714259}, DOI={10.1145/3706598.3714259}, booktitle={Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems}, publisher={ACM}, author={Riche, Nathalie and Offenwanger, Anna and Gmeiner, Frederic and Brown, David and Romat, Hugo and Pahud, Michel and Marquardt, Nicolai and Inkpen, Kori and Hinckley, Ken}, year={2025}, month=Apr, pages={1–18}, collection={CHI ’25} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/