Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI

Yanwei HuangWesley DengSijia XiaoMotahhare EslamiJason I. HongArpit NarechaniaAdam Perer

article2025International Conference on Human Factors in Computing Systems1 citations

Presents Vipera, an interactive auditing interface that integrates scene-graph visual cues with large language model suggestions to help human auditors systematically discover, organize, and evaluate problematic text-to-image generations.

Listen

Text-to-image generative artificial intelligence systems are increasingly deployed across creative and enterprise domains, yet they often generate biased, offensive, or inaccurate visual content. Auditing these models systematically is difficult because images contain vast semantic details that cannot be captured by rigid checklists, and human auditors frequently rely on ad-hoc intuition, missing critical systemic flaws.

The article demonstrates and evaluates Vipera, an interactive auditing interface that blends visual guidance—using interactive, statistics-augmented scene graphs—with large language model (LLM) suggestions to help auditors systematically discover, evaluate, and document vulnerabilities in text-to-image models.

The researchers formulated design goals through a formative study with five experienced auditors and implemented the Vipera system, which automatically summarizes image semantics into hierarchical tree structures and provides automated labeling and prompt recommendations. They evaluated the approach through a controlled user study involving 24 participants (20 general and 4 expert auditors) performing structured auditing tasks on the Stable Diffusion XL model across four system variants representing different combinations of visual and AI-driven guidance.

The study yielded several key findings: First, the full Vipera system achieved a statistically significant improvement in self-rated auditing performance compared to the baseline interface. Second, AI-driven guidance served as the primary driver of mental workload reduction, leading to a substantial drop in mental demand for participants. Third, while visual guidance through scene graphs helped users decompose image scenes and explore more diverse criteria, it added cognitive overhead and temporal demand when used in isolation. Fourth, combining both modalities led to the highest exploration rates, with AI generating up to roughly 78% to 87% of audit criteria, although the number of bookmarked findings per prompt decreased as users prioritized broad exploration over manual documentation.

These results indicate that visual structures and conversational AI guidance are highly complementary for risk governance and quality assurance. Visual aids ground the analysis in concrete data, while automated language models help prioritize focus and lower the cognitive barrier to exploring complex image spaces. Organizations deploying generative AI can use blended human-AI auditing to surface safety, bias, and compliance risks more efficiently than manual testing alone.

Organizations should consider integrating hybrid visual-AI auditing workflows into model validation pipelines, red-teaming processes, and developer environments like computational notebooks. Systems should also provide clear transparency and confidence indicators, as the study revealed that participants occasionally accepted incorrect AI labels without verification. Further technical development is needed to personalize guidance based on auditor behavior and to build aggregation tools that combine multi-auditor reports into cohesive compliance records.

The findings are subject to several limitations, including a participant sample comprised mostly of student auditors rather than dedicated compliance officers, a controlled laboratory setting with 15-minute audit tasks, and potential fatigue effects from the experimental sequence. Nevertheless, the study provides strong, credible evidence that blending visual analytics with generative language guidance creates a scalable, structured approach to generative AI auditing.

No sufficiently relevant recommendations were found.

Cover for Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI

Abstract

Despite their increasing capabilities, text-to-image generative AI systems are known to produce biased, offensive, and otherwise problematic outputs. While recent advancements have supported testing and auditing of generative AI, existing auditing methods still face challenges in supporting effectively explore the vast space of AI-generated outputs in a structured way. To address this gap, we conducted formative studies with five AI auditors and synthesized five design goals for supporting systematic AI audits. Based on these insights, we developed Vipera, an interactive auditing interface that employs multiple visual cues including a scene graph to facilitate image sensemaking and inspire auditors to explore and hierarchically organize the auditing criteria. Additionally, Vipera leverages LLM-powered suggestions to facilitate exploration of unexplored auditing directions. Through a controlled experiment with 24 participants experienced in AI auditing, we demonstrate Vipera's effectiveness in helping auditors navigate large AI output spaces and organize their analyses while engaging with diverse criteria.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Auditing generative AI (at scale)
  • 2.2 Visual analytics for sensemaking image collection
  • 3 Design Study & Goals
  • 3.1 Formative study
  • 3.1.1 Study setup
  • 3.1.2 Findings
  • 3.2 Design goals
  • 4 The Vipera System
  • 4.1 Interface Design
  • 4.1.1 Input view
  • 4.1.2 Analysis view
  • 4.1.3 Note view
  • 4.2 User workflow and technical pipeline
  • 4.3 Usage Scenario
  • 5 User Study
  • 5.1 Study Design and Methodology
  • 6 Results
  • 6.1 Mixed-method Analysis of Questionnaires and Audit Logs
  • 6.1.1 Questionnaires
  • 6.1.2 Auditing logs
  • 6.1.3 Auditing reports
  • 6.2 Qualitative Findings from Interviews
  • 6.2.1 The scene graph encourages systematic auditing while bringing additional cognitive demand and pressure.
  • 6.2.2 AI-powered auditing suggestions, though inspiring and effort-saving, needed to be personalized and upskilled.
  • 6.2.3 Effective auditing relies on drawing on diverse sources of insight and integrating them to perform a comprehensive assessment.
  • 6.2.4 The auditing approaches varied among participants, with both breadth-oriented and depth-oriented patterns in prompts and criteria.
  • 7 Limitations
  • 8 Discussion
  • 8.1 Opportunities from involving visual guidance in algorithm auditing
  • 8.2 Leveraging AI for inspirations and guidance in AI auditing
  • 8.3 Integrating different types of guidance for effective human-AI collaboration
  • 8.4 Incorporating intelligent auditing system into real world
  • 8.5 Future work
  • 9 Conclusion
  • References

Knowls

  1. Knowl 1 — Vipera Interactive Interface and Architectural Components

    model/method

    Vipera (Visual Intelligence-Promoted End User Auditing) is an interactive visual analytics system designed for systematic, multi-faceted auditing of generative text-to-image (T2I) models. The interface comprises three interconnected primary views:

    1. Input View: Contains a text field for entering prompts, a numerical selector for the batch size of images to generate, and a dropdown menu to select the target T2I model.

    2. Analysis View: Supports iterative exploration and hypothesis testing through four coordinated subviews:

      • Prompts View: Displays all user-initiated prompts, assigning a distinct categorical color to each prompt to maintain visual provenance across subsequent views.
      • Images View: Displays image thumbnails generated across prompts, bordered by their corresponding prompt color. Users can hover over an image to view its attributed labels and highlight corresponding nodes in the scene graph, right-click to manually correct inaccurate labels, or click to enlarge and bookmark images.
      • Scene Graph View: A tree-structured hierarchy organizing image semantics into visual objects (non-leaf nodes) and specific auditing criteria (leaf attribute nodes). Attribute nodes embed interactive stacked bar charts showing label distributions broken down by prompt color.
      • Audit Suggestions View: Presents AI-driven prompt modifications and image-difference-based criterion suggestions.
    3. Note View: Direct report authoring panel consisting of a General Notes editor (equipped with large language model auto-completion triggered by the Tab key) and an Evidence tab collecting bookmarked charts and images with user-written qualitative annotations.

  2. Knowl 2 — Automated Tree-Structured Scene Graph Construction and Dynamic Labeling Pipeline

    algorithm

    Vipera generates a semantic visual overview of generated images and dynamically evaluates user-defined attribute criteria using omni-modal large language models (LLMs). The end-to-end pipeline operates as follows:

    Input: User prompt PP, generation count NN, T2I model MM, extraction model LLMextractLLM_{extract} (Gemini 2.5 Flash), labeling model LLMlabelLLM_{label} (GPT-5-mini)
    Output: Interactive scene graph tree TT with embedded distribution charts
    1. Generate image set I=M(P,N)I = M(P, N)
    2. Sample image subset IsampleextfromIextsuchthat∣Isample∣=extmin(4,∣I∣)I_{sample} ext{ from } I ext{ such that } |I_{sample}| = ext{min}(4, |I|)
    3. Initialize empty scene graph list Glist=[]G_{list} = []
    4. for each image imgextinIsampleimg ext{ in } I_{sample} do
    5. gimgext=LLMextract(img)g_{img} ext{ = } LLM_{extract}(img) constrained to root branches "foreground" and "background"
    6. Append gimgexttoGlistg_{img} ext{ to } G_{list}
    7. end for
    8. Tmergedext=Merge(Glist)T_{merged} ext{ = Merge}(G_{list}) by merging identical object nodes and leaf nodes
    9. Text=Prune(Tmerged)T ext{ = Prune}(T_{merged}) by randomly retaining at most 5 leaf nodes (targeting approximately 7 total nodes)
    10. Render TT in an interactive tree layout
    11. upon user adding attribute node uu under parent object vv with candidate set C={c1,…,cm}C = \{c_1, \dots, c_m\} and scope S⊆IS \subseteq I:
    12. Propagate scope SS upward to all ancestors of uu in TT
    13. Extract partial graph schema path(root,u)path(root, u)
    14. for each image img∈Simg \in S do
    15. labelimg = LLMlabel(img,path(root,u),C)label_{img} \text{ = } LLM_{label}(img, path(root, u), C)
    16. end for
    17. D = AggregateFrequencies({labelimg∣img∈S}) grouped by prompt colorD \text{ = AggregateFrequencies}(\{label_{img} \mid img \in S\}) \text{ grouped by prompt color}
    18. Embed stacked bar chart of distribution DD within attribute node uu
  3. Knowl 3 — LLM-Driven Auditing Guidance via Audit Analysis Support and Prompt Suggestion

    model/method

    Vipera incorporates two complementary LLM-driven guidance modules to guide auditors toward unexplored testing directions:

    1. Audit Analysis Support (Criteria Generation): Focuses on surfacing "unknown unknowns" by identifying subtle visual discrepancies. The module selects pairs of generated images showing noticeable variations and prompts an LLM to recommend new auditing attribute criteria. To prevent intent drift, users can provide custom keywords or select from LLM-generated keywords (e.g., composition, style, coherence, diversity). Suggestions are filtered and displayed only when the LLM's self-reported confidence score exceeds a predefined threshold across multiple iterative evaluation passes.

    2. Prompt Suggestion (Divergent Prompting and Inheritance): Proposes new prompts by substituting specific tokens or phrases in existing prompts (e.g., replacing "doctor" with "nurse" in "a cinematic photo of a doctor"). When an auditor accepts a prompt suggestion, Vipera generates the new image batch, creates a corresponding sibling node in the scene graph (e.g., a "nurse" node alongside "doctor"), and automatically duplicates all existing descendant attribute criteria (e.g., "gender", "stethoscope") into the new branch. This enables immediate multi-prompt comparative auditing across identical evaluation axes.

  4. Knowl 4 — Controlled Experimental Setup for Evaluating Blended Auditing Guidance

    experimental setup

    A controlled laboratory user study evaluated the effectiveness of Vipera's scene graph and AI suggestions in auditing text-to-image models:

    • Participants: 24 participants (14 male, 10 female; age range 20–29, mean 23.9), comprising 20 student auditors (4 undergraduate, 7 Master's, 9 Ph.D.; self-rated AI auditing familiarity M=3.04/5,SD=1.16M=3.04/5, SD=1.16; T2I usage frequency M=2.92/5,SD=0.929M=2.92/5, SD=0.929) and 4 expert auditors from industry and academia.
    • Target Model and Tasks: Auditing Stable Diffusion XL starting from one of four seed prompts: (a) "A couple on their wedding day", (b) "A family having a picnic in the park", (c) "Worldwide athletes in the Olympic Games", or (d) "An award-winning chef preparing a gourmet meal".
    • System Configurations:
      • System A (Baseline): No scene graph and no AI guidance; equipped with a flat Criteria View allowing manual criteria creation and bar chart rendering.
      • System B: Vipera with scene graph visual guidance, but without AI guidance modules.
      • System C: Vipera with AI guidance modules and flat Criteria View, but without the scene graph.
      • System D (Full Vipera): Integrated scene graph and AI guidance modules.
    • Study Protocol: A mixed-design progressive evaluation split participants into Group I (n=12n=12, using A →\rightarrow B →\rightarrow D) and Group II (n=12n=12, using A →\rightarrow C →\rightarrow D). Each system session lasted 15 minutes with think-aloud protocols, NASA-TLX workload questionnaires, 7-point Likert ratings, system interaction logs, and auditing report collection.
  5. Knowl 5 — Workload and Performance Effects of Visual and LLM Guidance Modalities

    empirical result

    Statistical analysis of NASA-TLX ratings across 24 participants revealed distinct impacts of visual and AI-driven guidance:

    • Overall Performance: Full Vipera (System D) produced a statistically significant improvement in self-rated Performance compared to baseline System A (t(23)=2.685,p=0.0066t(23) = 2.685, p = 0.0066, Holm–Bonferroni corrected p′=0.0397p' = 0.0397).
    • AI Guidance Workload Reduction (Group II): Introducing AI auditing support alone (System C vs. System A) caused a statistically significant reduction in Mental Demand (t(11)=−5.063,p=0.0002,p′=0.0022t(11) = -5.063, p = 0.0002, p' = 0.0022). Non-significant reduction trends were observed for Effort (p=0.0086,p′=0.0515p = 0.0086, p' = 0.0515), Physical Demand (p=0.0380,p′=0.0912p = 0.0380, p' = 0.0912), and Temporal Demand (p=0.0310,p′=0.0931p = 0.0310, p' = 0.0931).
    • Visual Guidance Cognitive Trade-Off (Group I): Adding the scene graph alone (System B vs. System A) trended toward improved Performance (t(11)=3.023,p=0.0058,p′=0.0696t(11) = 3.023, p = 0.0058, p' = 0.0696), but slightly increased median Mental Demand due to cognitive and temporal overhead in navigating dense graph hierarchies. Subsequent addition of AI support (System D vs. System B: t(11)=1.876,p=0.0437,p′=0.1748t(11) = 1.876, p = 0.0437, p' = 0.1748) alleviated this overhead.
    • Component Usability: On 7-point Likert scales, participants rated note-taking features highest (M=3.71,SD=1.083M=3.71, SD=1.083), followed by AI auditing support (M=3.67,SD=0.868M=3.67, SD=0.868), the scene graph (M=3.50,SD=0.885M=3.50, SD=0.885), and automatic image labeling (M=3.375,SD=0.970M=3.375, SD=0.970).
  6. Knowl 6 — Auditing Exploration Behaviors and Authoring Agency Under Blended Guidance

    empirical result

    System interaction log analysis across Systems A, B, C, and D demonstrated distinct behavioral shifts:

    • Criteria and Prompt Generation: Participants created more auditing criteria across all guided systems (B, C, D) than in baseline System A, with the highest count in full Vipera (System D). Prompts and image generation volume increased primarily when AI guidance was present (Systems C and D).
    • Evidence Bookmarking: The number of bookmarked charts and images per user decreased consistently in Systems C and D, indicating a shift from passively documenting existing views to actively generating and exploring new criteria and prompts.
    • Authoring Agency Disparity: In Systems C and D, AI generated 87.2%87.2\% and 78.4%78.4\% of used criteria, respectively, whereas AI generated only 61.8%61.8\% and 50.8%50.8\% of used prompts. This demonstrates that auditors prefer greater manual control over prompt formulation than over criteria specification.
    • Prompt Consistency: Across all sessions, the average pairwise BERT cosine similarity among user-generated prompts was 0.9040.904 (scale 0–10\text{--}1), showing that users deliberately executed controlled prompt variations to isolate specific visual factors.
    • Output Efficiency: The average number of bookmarked insights per prompt decreased from 1.21.2 in System A and 1.21.2 in System B to 0.90.9 in System C and 0.80.8 in System D as exploration scope widened.
  7. Knowl 7 — Thematic Classification of Text-to-Image Auditing Insights

    data/table

    Thematic analysis of the qualitative insights authored by study participants categorized auditing findings into four primary dimensions and nine subcategories:

    Category Subcategory Example Quotes (Participant ID)
    Quality Exquisiteness "The details are slightly lacking" (P08); "The picture looks beautiful" (P16); "Mosaics appearing" (P19)
    Authenticity "The person looks like a dummy" (P02); "The dishes are floating" (P12); "People's faces look distorted" (P05)
    Style "Prefer to see more realistic images but the majority are artistic" (P21)
    Diversity "The characters are highly homogeneous" (P19); "Racial bias" (P02); ÄI may discriminate unwealthy people" (P13)
    Prompt Interpretation Alignment with the prompt "The content of the picture aligns with the 'worldwide' theme" (P15); "No 'athletes' in the image" (P03)
    Alignment with common sense "There are six fingers in one hand" (P08); "People's mood is wrong; they should be happy on the wedding day" (P05)
    Alignment with user understanding Ëven if I added the word 'beautiful' to the prompt, the images were still ugly" (P02); Ït misunderstood me" (P03)
    Reasoning for LLM Behavior Prompt comparison "The pictures are better when more details are included in the prompt" (P06)
    Express the tendency of the LLM "The chance of errors increases when more elements exist" (P16); "LLM didn't understand subjective keywords well" (P02)
    Explain for the LLM behavior Ï probably want a more diverse set of colors. Could be attributed to most of the training images having this color" (P21)
    Critic on the System - ÄI gives me more options... But the recommended prompt may have little novelty" (P14); "The classification of images is problematic" (P06)

    The taxonomy demonstrates how multimodal guidance enables auditors to progress from pixel-level artifact detection (exquisiteness and authenticity) to semantic prompt alignment, demographic distribution analysis, and hypotheses regarding training data distribution.

  8. Knowl 8 — Breadth-Oriented vs. Depth-Oriented Auditing Behavioral Strategies

    empirical result

    Auditors exhibited two distinct operational strategies when navigating the auditing space of text-to-image models:

    1. Breadth in Criteria / Depth in Prompts: The majority of participants (including 9 participants who averaged fewer than 2 prompts per system) focused on deeply decomposing image batches for a single prompt across many criteria. They exhausted a diverse hierarchy of attributes on existing images before moving to new prompts.
    2. Depth in Criteria / Breadth in Prompts (Iterative Prompt Engineering): Other participants focused on prompt exploration, iteratively perturbing prompt phrases to observe how model fidelity, artifact frequency, or composition shifted across successive prompts rather than evaluating images across multiple static criteria.
    3. Balanced Hybrid Navigation: A subset of participants actively regulated their workflow by monitoring their depth along a specific branch in the scene graph and intentionally switching to broad prompt exploration when perceiving diminishing returns.
  9. Knowl 9 — Auditor Passivity and Trust in Automated LLM Labeling Failures

    empirical result

    During user interactions with automated vision-language labeling in Vipera:

    • Although participants were explicitly tasked with auditing the generative text-to-image model, only approximately 50%50\% of participants who noticed labeling errors made by the omni-modal LLM attempted to manually correct labels or request relabeling with modified attribute descriptions.
    • The remaining half of the participants accepted incorrect LLM-generated categorical labels without intervention, subsequently incorporating inaccurate visual statistics and charts directly into their final audit reports as evidence.
    • Multiple participants mistakenly critiqued the scene graph and the automated AI labeler in their auditing reports rather than isolating and evaluating the underlying generative image model.
  10. Knowl 10 — Stated Limitations of Vipera and its Empirical Evaluation

    limitation

    The authors identified several limitations in the system and experimental evaluation:

    1. Ecological Validity and Session Duration: Controlled 15-minute laboratory tasks with fixed durations may not capture the open-ended, longitudinal nature of professional red-teaming or organizational AI compliance audits.
    2. Sample Demographics: The evaluation cohort consisted mostly of student auditors (20 out of 24), whose strategies and domain expertise may differ from professional industry compliance teams.
    3. Insight Proxy Metric: Using the quantity of bookmarked items as a quantitative proxy for insight discovery does not measure the qualitative depth, novelty, or validity of the discovered model vulnerabilities.
    4. Order and Learning Effects: The progressive non-counterbalanced presentation order (A →\rightarrow B/C →\rightarrow D), adopted to prevent contrast bias against simpler baselines, could introduce fatigue or learning effects.
    5. Automated Vision-Language Errors: Errors in the omni-modal LLMs used for scene graph generation and image classification can mislead users and degrade report accuracy.
    6. Single-Model Scope: The system was evaluated on a single generative model (Stable Diffusion XL) without multi-model parallel reconciliation.

Coverage note — None was omitted beyond standard formative interview quotes, preliminary pilot testing remarks, and general related work background.

References

  1. 1.Shehzad Afzal, Sohaib Ghani, Mohamad Mazen Hittawe, Sheikh Faisal Rashid, Omar M Knio, Markus Hadwiger, and Ibrahim Hoteit. 2023. Visualization and Visual Analytics Approaches for Image and Video Datasets: A Survey. ACM Transactions on Interactive Intelligent Systems 13, 1 (2023), 1–41.
  2. 2.Shm Garanganao Almeda, JD Zamfirescu-Pereira, Kyu Won Kim, Pradeep Mani Rathnam, and Bjoern Hartmann. 2024. Prompting for Discovery: Flexible Sense-Making for AI Art-Making with Dreamsheets. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–17. doi:10.1145/3613904.3642858
  3. 3.Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Zico Kolter, Matt Fredrikson, et al. 2025. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. International Conference on Learning Representations (2025).
  4. 4.Ian Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg, and Elena L. Glassman. 2024. ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing. In Proceedings of the CHI Conference on Human Factors in Computing Systems. doi:10.1145/3613904.3642016
  5. 5.Joshua Asplund, Motahhare Eslami, Hari Sundaram, Christian Sandvig, and Karrie Karahalios. 2020. Auditing race and gender discrimination in online housing markets. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 14. 24–35.
  6. 6.Agathe Balayn, Mireia Yurrita, Jie Yang, and Ujwal Gadiraju. 2023. “Fairness Toolkits, A Checkbox Culture?” On the Factors that Fragment Developer Practices in Handling Algorithmic Harms. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society. 482–495.
  7. 7.Abeba Birhane, Ryan Steed, Victor Ojewale, Briana Vecchione, and Inioluwa Deborah Raji. 2024. AI auditing: The broken bus on the road to AI accountability. In IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). IEEE, 612–643.
  8. 8.Edyta Bogucka, Marios Constantinides, Sanja Šćepanović, and Daniele Quercia. 2024. Co-designing an AI impact assessment report template with AI practitioners and AI compliance experts. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, Vol. 7. 168–180.
  9. 9.Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psychology. Qualitative research in psychology 3, 2 (2006), 77–101.
  10. 10.Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency. PMLR, 77–91.
  11. 11.Ángel Alexander Cabrera, Abraham J Druck, Jason I Hong, and Adam Perer. 2021. Discovering and Validating AI Errors With Crowdsourced Failure Reports. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–22.
  12. 12.Ángel Alexander Cabrera, Abraham J. Druck, Jason I. Hong, and Adam Perer. 2021. Discovering and Validating AI Errors With Crowdsourced Failure Reports. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (Oct. 2021). doi:10.1145/3479569
  13. 13.Ángel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein, Ameet Talwalkar, Jason I Hong, and Adam Perer. 2023. Zeno: An interactive framework for behavioral evaluation of machine learning. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–14. doi:10.1145/3544548.3581268
  14. 14.Wang Claire, Wesley Hanwen Deng, Jason Hong, Ken Holstein, and Motahhare Eslami. 2024. Designing a Crowdsourcing Pipeline to Verify Reports from User AI Audits. Work in Progress of the AAAI Conference on Human Computation and Crowdsourcing (2024).
  15. 15.Dazhen Deng, Chuhan Zhang, Huawei Zheng, Yuwen Pu, Shouling Ji, and Yingcai Wu. 2025. AdversaFlow: Visual Red Teaming for Large Language Models with Multi-Level Adversarial Flow . IEEE Transactions on Visualization and Computer Graphics 31, 01 (Jan. 2025), 492–502. doi:10.1109/TVCG.2024.3456150
  16. 16.Wesley Hanwen Deng, Wang Claire, Howard Ziyu Han, Jason I. Hong, Kenneth Holstein, and Motahhare Eslami. 2025. WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI. Proceedings of the ACM on Human-Computer Interaction 9, 7 (Oct. 2025). doi:10.1145/3757702
  17. 17.Wesley Hanwen Deng, Boyuan Guo, Alicia Devrio, Hong Shen, Motahhare Eslami, and Kenneth Holstein. 2023. Understanding Practices, Challenges, and Opportunities for User-Engaged Algorithm Auditing in Industry Practice. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–18. doi:10.1145/3544548.3581026
  18. 18.Wesley Hanwen Deng, Manish Nagireddy, Michelle Seng Ah Lee, Jatinder Singh, Zhiwei Steven Wu, Kenneth Holstein, and Haiyi Zhu. 2022. Exploring how machine learning practitioners (try to) use fairness toolkits. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency. 473–484.
  19. 19.Wesley Hanwen Deng, Nur Yildirim, Monica Chang, Motahhare Eslami, Kenneth Holstein, and Michael Madaio. 2023. Investigating practices and opportunities for cross-functional collaboration around AI fairness in industry practice. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency. 705–716.
  20. 20.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 4171–4186.
  21. 21.Nathalie Diberardino, Clair Baleshta, and Luke Stark. 2024. Algorithmic Harms and Algorithmic Wrongs. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency. 1725–1732. doi:10.1145/3630106.3659001
  22. 22.Moreno D’Incà, Elia Peruzzo, Massimiliano Mancini, Dejia Xu, Vidit Goel, Xingqian Xu, Zhangyang Wang, Humphrey Shi, and Nicu Sebe. 2024. Openbias: Open-set bias detection in text-to-image generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12225–12235.
  23. 23.Steven P Dow, Alana Glassco, Jonathan Kass, Melissa Schwarz, Daniel L Schwartz, and Scott R Klemmer. 2010. Parallel prototyping leads to better design results, more divergence, and increased self-efficacy. ACM Transactions on Computer-Human Interaction (TOCHI) 17, 4 (2010), 1–24.
  24. 24.Klaus Eckelt, Kiran Gadhave, Alexander Lex, and Marc Streit. 2025. Loops: Leveraging Provenance and Visualization to Support Exploratory Data Analysis in Notebooks. IEEE Transactions on Visualization and Computer Graphics 31, 1 (2025), 1213–1223. doi:10.1109/TVCG.2024.3456186
  25. 25.Fei Fang, Miao Yi, Hui Feng, Shenghong Hu, and Chunxia Xiao. 2017. Narrative Collage of Image Collections by Scene Graph Recombination. IEEE Transactions on Visualization and Computer Graphics 24, 9 (2017), 2559–2572. doi:10.1109/TVCG.2017.2759265
  26. 26.Lin Gao, Jing Lu, Zekai Shao, Ziyue Lin, Shengbin Yue, Chiokit Leong, Yi Sun, Rory James Zauner, Zhongyu Wei, and Siming Chen. 2025. Fine-Tuned Large Language Model for Visualization System: A Study on Self-Regulated Learning in Education. IEEE Transactions on Visualization and Computer Graphics 31, 1 (2025), 514–524. doi:10.1109/TVCG.2024.3456145
  27. 27.Katy Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K. Kummerfeld, and Elena L. Glassman. 2024. Supporting Sensemaking of Large Language Model Outputs at Scale. In Proceedings of the CHI Conference on Human Factors in Computing Systems. doi:10.1145/3613904.3642139
  28. 28.Yuhan Guo, Hanning Shao, Can Liu, Kai Xu, and Xiaoru Yuan. 2024. PrompTHis: Visualizing the Process and Influence of Prompt Editing during Text-to-Image Creation. IEEE Transactions on Visualization and Computer Graphics (2024). doi:10.1109/TVCG.2024.3408255
  29. 29.Aniko Hannak, Gary Soeller, David Lazer, Alan Mislove, and Christo Wilson. 2014. Measuring price discrimination and steering on e-commerce web sites. In Proceedings of the 2014 Conference on Internet Measurement Conference. 305–318.
  30. 30.Jeffrey Heer. 2019. Agency plus automation: Designing artificial intelligence into interactive systems. Proceedings of the National Academy of Sciences 116, 6 (2019), 1844–1850.
  31. 31.Florian Heimerl, Steffen Lohmann, Simon Lange, and Thomas Ertl. 2014. Word Cloud Explorer: Text analytics based on word clouds. In Proceedings of the 2014 47th Hawaii International Conference on System Sciences. IEEE, 1833–1842.
  32. 32.Fred Hohman, Kanit Wongsuphasawat, Mary Beth Kery, and Kayur Patel. 2020. Understanding and Visualizing Data Iteration in Machine Learning. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–13. doi:10.1145/3313831.3376177
  33. 33.Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach. 2019. Improving fairness in machine learning systems: What do industry practitioners need?. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–16. doi:10.1145/3290605.3300830
  34. 34.Peiling Jiang, Jude Rayan, Steven P Dow, and Haijun Xia. 2023. Graphologue: Exploring Large Language Model Responses with Interactive Diagrams. In Proceedings of the Annual ACM Symposium on User Interface Software and Technology. 1–20. doi:10.1145/3586183.3606737
  35. 35.Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei. 2015. Image Retrieval Using Scene Graphs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3668–3678.
  36. 36.Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan. 2020. Interpreting Interpretability: Understanding Data Scientists’ Use of Interpretability Tools for Machine Learning. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–14. doi:10.1145/3313831.3376219
  37. 37.Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021. Dynabench: Rethinking Benchmarking in NLP. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4110–4124. doi:10.18653/v1/2021.naacl-main.324
  38. 38.Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2024. EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria. In Proceedings of the CHI Conference on Human Factors in Computing Systems. doi:10.1145/3613904.3642216
  39. 39.Apoorva Nalini Pradeep Kumar, Justus Bogner, Markus Funke, and Patricia Lago. 2024. Balancing Progress and Responsibility: A Synthesis of Sustainability Trade-Offs of AI-Based Systems. In IEEE 21st International Conference on Software Architecture Companion (ICSA-C). 207–214.
  40. 40.Michelle S. Lam, Mitchell L. Gordon, Danaë Metaxa, Jeffrey T. Hancock, James A. Landay, and Michael S. Bernstein. 2022. End-User Audits: A System Empowering Communities to Lead Large-Scale Investigations of Harmful Algorithmic Behavior. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2 (Nov. 2022). doi:10.1145/3555625
  41. 41.Clayton Lewis and Robert Mack. 1982. Learning to use a text processing system: Evidence from “thinking aloud” protocols. In Proceedings of the Conference on Human Factors in Computing Systems. 387–392. doi:10.1145/800049.801817
  42. 42.Guanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang, and Tianwei Zhang. 2024. ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users. In Advances in Neural Information Processing Systems, Vol. 37. 91184–91219. doi:10.52202/079017-2894
  43. 43.Xinfeng Li, Yuchen Yang, Jiangyi Deng, Chen Yan, Yanjiao Chen, Xiaoyu Ji, and Wenyuan Xu. 2024. SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models. In Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS).
  44. 44.Yiran Li, Junpeng Wang, Prince Aboagye, Chin-Chia Michael Yeh, Yan Zheng, Liang Wang, Wei Zhang, and Kwan-Liu Ma. 2024. Visual Analytics for Efficient Image Exploration and User-Guided Image Captioning. IEEE Transactions on Visualization and Computer Graphics (2024). doi:10.1109/TVCG.2024.3388514
  45. 45.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved Baselines with Visual Instruction Tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 26286–26296. doi:10.1109/CVPR52733.2024.02484
  46. 46.David C. Logan. 2009. Known knowns, known unknowns, unknown unknowns and the propagation of scientific enquiry. Journal of Experimental Botany 60, 3 (03 2009), 712–714. doi:10.1093/jxb/erp043
  47. 47.Kelly Avery Mack, Rida Qadri, Remi Denton, Shaun K Kane, and Cynthia L Bennett. 2024. “They only care to show us the wheelchair”: disability representation in text-to-image AI models. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–23. doi:10.1145/3613904.3642166
  48. 48.Michael A Madaio, Luke Stark, Jennifer Wortman Vaughan, and Hanna Wallach. 2020. Co-designing checklists to understand organizational challenges and opportunities around fairness in AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–14. doi:10.1145/3313831.3376445
  49. 49.Matheus Kunzler Maldaner, Wesley Hanwen Deng, Jason Hong, Ken Holstein, and Motahhare Eslami. 2024. MIRAGE: Multi-model Interface for Reviewing and Auditing Generative Text-to-Image AI. Demo of the AAAI Conference on Human Computation and Crowdsourcing (2024).
  50. 50.Danaë Metaxa, Joon Sung Park, Ronald E Robertson, Karrie Karahalios, Christo Wilson, Jeff Hancock, Christian Sandvig, et al. 2021. Auditing algorithms: Understanding algorithmic systems from the outside in. Foundations and Trends® in Human–Computer Interaction 14, 4 (2021), 272–344.
  51. 51.Katelyn Morrison, Arpit Mathur, Aidan Bradshaw, Tom Wartmann, Steven Lundi, Afrooz Zandifar, Weichang Dai, Kayhan Batmanghelich, Motahhare Eslami, and Adam Perer. 2025. A Human-Centered Approach to Identifying Promises, Risks, & Challenges of Text-to-Image Generative AI in Radiology. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, Vol. 8. 1758–1770.
  52. 52.Ranjita Naik and Besmira Nushi. 2023. Social Biases through the Text-to-Image Generation Lens. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society. 786–808. doi:10.1145/3600211.3604711
  53. 53.Safiya Umoja Noble. 2018. Algorithms of oppression: How search engines reinforce racism. NYU Press.
  54. 54.Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane, and Inioluwa Deborah Raji. 2025. Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling. In Proceedings of the CHI Conference on Human Factors in Computing Systems. doi:10.1145/3706598.3713301
  55. 55.Hadas Orgad, Bahjat Kawar, and Yonatan Belinkov. 2023. Editing implicit assumptions in text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 7053–7061.
  56. 56.Qian Pan, Zahra Ashktorab, Michael Desmond, Martín Santillán Cooper, James Johnson, Rahul Nair, Elizabeth Daly, and Werner Geyer. 2024. Human-Centered Design Recommendations for LLM-as-a-Judge. In Proceedings of the 1st Human-Centered Large Language Modeling Workshop. 16–29. doi:10.18653/v1/2024.hucllm-1.2
  57. 57.Xingjia Pan, Fan Tang, Weiming Dong, Chongyang Ma, Yiping Meng, Feiyue Huang, Tong-Yee Lee, and Changsheng Xu. 2019. Content-Based Visual Summarization for Image Collections. IEEE Transactions on Visualization and Computer Graphics 27, 4 (2019), 2298–2312. doi:10.1109/TVCG.2019.2948611
  58. 58.Samir Passi and Solon Barocas. 2019. Problem formulation and fairness. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency. 39–48.
  59. 59.Rune Pettersson. 1993. Visual information. Educational Technology.
  60. 60.Marcelo OR Prates, Pedro H Avelar, and Luís C Lamb. 2020. Assessing gender bias in machine translation: a case study with google translate. Neural Computing and Applications 32, 10 (2020), 6363–6381.
  61. 61.Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. 2024. Visual adversarial examples jailbreak aligned large language models. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence. AAAI Press. doi:10.1609/aaai.v38i19.30150
  62. 62.Xin Qian, Ryan A. Rossi, Fan Du, Sungchul Kim, Eunyee Koh, Sana Malik, Tak Yeon Lee, and Joel Chan. 2021. Learning to Recommend Visualizations from Data. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery & Data Mining. 1359–1369. doi:10.1145/3447548.3467224
  63. 63.Bogdana Rakova, Jingying Yang, Henriette Cramer, and Rumman Chowdhury. 2021. Where responsible AI meets reality: Practitioner perspectives on enablers for shifting organizational practices. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1 (2021), 1–23.
  64. 64.Charvi Rastogi, Marco Tulio Ribeiro, Nicholas King, Harsha Nori, and Saleema Amershi. 2023. Supporting Human-AI Collaboration in Auditing LLMs with LLMs. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society. 913–926. doi:10.1145/3600211.3604712
  65. 65.Charvi Rastogi, Marco Tulio Ribeiro, Nicholas King, Harsha Nori, and Saleema Amershi. 2023. Supporting Human-AI Collaboration in Auditing LLMs with LLMs. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society. 913–926. doi:10.1145/3600211.3604712
  66. 66.Mark Ryan, Eleni Christodoulou, Josephina Antoniou, and Kalypso Iordanou. 2024. An AI ethics ‘David and Goliath’: value conflicts between large tech companies and their employees. AI & SOCIETY 39, 2 (2024), 557–572.
  67. 67.Christian Sandvig, Kevin Hamilton, Karrie Karahalios, and Cedric Langbort. 2014. Auditing algorithms: Research methods for detecting discrimination on internet platforms. Data and Discrimination: Converting Critical Concerns into Productive Inquiry (2014).
  68. 68.Johanna Schmidt, M Eduard Gröller, and Stefan Bruckner. 2013. VAICo: Visual analysis for image comparison. IEEE Transactions on Visualization and Computer Graphics 19, 12 (2013), 2090–2099. doi:10.1109/TVCG.2013.213
  69. 69.Patrick Schramowski, Manuel Brack, Björn Deiseroth, and Kristian Kersting. 2023. Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22522–22531.
  70. 70.Shreya Shankar, J.D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya Parameswaran, and Ian Arawjo. 2024. Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences. In Proceedings of the Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3654777.3676450
  71. 71.Renee Shelby, Shalaleh Rismani, and Negar Rostamzadeh. 2024. Generative AI in Creative Practice: ML-Artist Folk Theories of T2I Use, Harm, and Harm-Reduction. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–17. doi:10.1145/3613904.3642461
  72. 72.Hong Shen, Alicia DeVos, Motahhare Eslami, and Kenneth Holstein. 2021. Everyday Algorithm Auditing: Understanding the Power of Everyday Users in Surfacing Harmful Algorithmic Behaviors. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (Oct. 2021). doi:10.1145/3479577
  73. 73.Hong Shen, Leijie Wang, Wesley H Deng, Ciell Brusse, Ronald Velgersdijk, and Haiyi Zhu. 2022. The model card authoring toolkit: Toward community-centered, deliberation-driven AI design. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency. 440–451.
  74. 74.Leixian Shen, Haotian Li, Yifang Wang, Xing Xie, and Huamin Qu. 2025. Prompting generative AI with interaction-augmented instructions. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–9.
  75. 75.Jaemarie Solyst, Cindy Peng, Wesley Hanwen Deng, Praneetha Pratapa, Amy Ogan, Jessica Hammer, Jason Hong, and Motahhare Eslami. 2025. Investigating Youth AI Auditing. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency. 2098–2111.
  76. 76.Hendrik Strobelt, Daniela Oelke, Christian Rohrdantz, Andreas Stoffel, Daniel A Keim, and Oliver Deussen. 2009. Document Cards: A Top Trumps Visualization for Documents. IEEE Transactions on Visualization and Computer Graphics 15, 6 (2009), 1145–1152. doi:10.1109/TVCG.2009.139
  77. 77.Sangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li, and Haijun Xia. 2024. Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation. In Proceedings of the CHI Conference on Human Factors in Computing Systems. doi:10.1145/3613904.3642400
  78. 78.Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language Models. In Proceedings of the Annual ACM Symposium on User Interface Software and Technology. 1–18. doi:10.1145/3586183.3606756
  79. 79.Latanya Sweeney. 2013. Discrimination in online ad delivery. Queue 11, 3 (2013), 10–29.
  80. 80.April Yi Wang, Will Epperson, Robert A DeLine, and Steven M Drucker. 2022. Diff in the Loop: Supporting Data Comparison in Exploratory Data Analysis. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–10. doi:10.1145/3491102.3502123
  81. 81.Qiaosi Wang, Michael Madaio, Shaun Kane, Shivani Kapania, Michael Terry, and Lauren Wilcox. 2023. Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI Challenges. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–16. doi:10.1145/3544548.3581278
  82. 82.Zijie J. Wang, Chinmay Kulkarni, Lauren Wilcox, Michael Terry, and Michael Madaio. 2024. Farsight: Fostering Responsible AI Awareness During AI Application Prototyping. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–40. doi:10.1145/3613904.3642335
  83. 83.Luoxuan Weng, Xingbo Wang, Junyu Lu, Yingchaojie Feng, Yihan Liu, Haozhe Feng, Danqing Huang, and Wei Chen. 2025. InsightLens: Augmenting LLM-Powered Data Analysis with Interactive Insight Management and Navigation. IEEE Transactions on Visualization and Computer Graphics (2025). doi:10.1109/TVCG.2025.3567131
  84. 84.David Gray Widder, Laura Dabbish, James D Herbsleb, and Nikolas Martelaro. 2024. Power and Play: Investigating"License to Critique"in Teams’ AI Ethics Discussions. Proceedings of the ACM on Human-Computer Interaction 8, CSCW2 (2024), 1–23.
  85. 85.Xiao Xie, Xiwen Cai, Junpei Zhou, Nan Cao, and Yingcai Wu. 2018. A Semantic-Based Method for Visualizing Large Image Collections. IEEE Transactions on Visualization and Computer Graphics 25, 7 (2018), 2362–2377. doi:10.1109/TVCG.2018.2835485
  86. 86.Ka-Ping Yee, Kirsten Swearingen, Kevin Li, and Marti Hearst. 2003. Faceted Metadata for Image Search and Browsing. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 401–408. doi:10.1145/642611.642681
  87. 87.Nur Yildirim, Mahima Pushkarna, Nitesh Goyal, Martin Wattenberg, and Fernanda Viégas. 2023. Investigating How Practitioners Use Human-AI Guidelines: A Case Study on the People + AI Guidebook. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–13. doi:10.1145/3544548.3580900
  88. 88.Jan Zahálka, Marcel Worring, and Jarke J Van Wijk. 2020. II-20: Intelligent and pragmatic analytic categorization of image collections. IEEE Transactions on Visualization and Computer Graphics 27, 2 (2020), 422–431. doi:10.1109/TVCG.2020.3030383
  89. 89.Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2024. SafetyBench: Evaluating the Safety of Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 15537–15553. doi:10.18653/v1/2024.acl-long.830
  90. 90.Yuheng Zhao, Xueli Shu, Liwen Fan, Lin Gao, Yu Zhang, and Siming Chen. 2025. ProactiveVA: Proactive Visual Analytics with LLM-Based UI Agent. IEEE VIS (2025).
  91. 91.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2024. Judging LLM-as-a-judge with MT-bench and Chatbot Arena. In Proceedings of the 37th International Conference on Neural Information Processing Systems.

Citation

MLA
Huang, Y., et al. “Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI”. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 2026, pp. 1–8, https://doi.org/10.1145/3772318.3791942.
APA
Huang, Y., Hanwen Deng, W., Xiao, S., Eslami, M., Hong, J. I., Narechania, A., & Perer, A. (2026). Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 1–18. https://doi.org/10.1145/3772318.3791942
Chicago
Huang, Y., W. Hanwen Deng, S. Xiao, et al. 2026. “Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI”. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 1–18. https://doi.org/10.1145/3772318.3791942.
Harvard
Huang, Y. et al. (2026) “Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI”, Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. ACM, pp. 1–18. Available at: https://doi.org/10.1145/3772318.3791942.
Vancouver
1. Huang Y, Hanwen Deng W, Xiao S, Eslami M, Hong JI, Narechania A, Perer A (2026) Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI. In: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. ACM, pp 1–18

BibTeX

@inproceedings{Huang_2026, series={CHI ’26}, title={Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI}, url={http://dx.doi.org/10.1145/3772318.3791942}, DOI={10.1145/3772318.3791942}, booktitle={Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems}, publisher={ACM}, author={Huang, Yanwei and Hanwen Deng, Wesley and Xiao, Sijia and Eslami, Motahhare and Hong, Jason I. and Narechania, Arpit and Perer, Adam}, year={2026}, month=Apr, pages={1–18}, collection={CHI ’26} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/