Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AI
Yanwei HuangWesley DengSijia XiaoMotahhare EslamiJason I. HongArpit NarechaniaAdam Perer
Presents Vipera, an interactive auditing interface that integrates scene-graph visual cues with large language model suggestions to help human auditors systematically discover, organize, and evaluate problematic text-to-image generations.
Text-to-image generative artificial intelligence systems are increasingly deployed across creative and enterprise domains, yet they often generate biased, offensive, or inaccurate visual content. Auditing these models systematically is difficult because images contain vast semantic details that cannot be captured by rigid checklists, and human auditors frequently rely on ad-hoc intuition, missing critical systemic flaws.
The article demonstrates and evaluates Vipera, an interactive auditing interface that blends visual guidance—using interactive, statistics-augmented scene graphs—with large language model (LLM) suggestions to help auditors systematically discover, evaluate, and document vulnerabilities in text-to-image models.
The researchers formulated design goals through a formative study with five experienced auditors and implemented the Vipera system, which automatically summarizes image semantics into hierarchical tree structures and provides automated labeling and prompt recommendations. They evaluated the approach through a controlled user study involving 24 participants (20 general and 4 expert auditors) performing structured auditing tasks on the Stable Diffusion XL model across four system variants representing different combinations of visual and AI-driven guidance.
The study yielded several key findings: First, the full Vipera system achieved a statistically significant improvement in self-rated auditing performance compared to the baseline interface. Second, AI-driven guidance served as the primary driver of mental workload reduction, leading to a substantial drop in mental demand for participants. Third, while visual guidance through scene graphs helped users decompose image scenes and explore more diverse criteria, it added cognitive overhead and temporal demand when used in isolation. Fourth, combining both modalities led to the highest exploration rates, with AI generating up to roughly 78% to 87% of audit criteria, although the number of bookmarked findings per prompt decreased as users prioritized broad exploration over manual documentation.
These results indicate that visual structures and conversational AI guidance are highly complementary for risk governance and quality assurance. Visual aids ground the analysis in concrete data, while automated language models help prioritize focus and lower the cognitive barrier to exploring complex image spaces. Organizations deploying generative AI can use blended human-AI auditing to surface safety, bias, and compliance risks more efficiently than manual testing alone.
Organizations should consider integrating hybrid visual-AI auditing workflows into model validation pipelines, red-teaming processes, and developer environments like computational notebooks. Systems should also provide clear transparency and confidence indicators, as the study revealed that participants occasionally accepted incorrect AI labels without verification. Further technical development is needed to personalize guidance based on auditor behavior and to build aggregation tools that combine multi-auditor reports into cohesive compliance records.
The findings are subject to several limitations, including a participant sample comprised mostly of student auditors rather than dedicated compliance officers, a controlled laboratory setting with 15-minute audit tasks, and potential fatigue effects from the experimental sequence. Nevertheless, the study provides strong, credible evidence that blending visual analytics with generative language guidance creates a scalable, structured approach to generative AI auditing.
- Paper: OpenBias: Open-Set Bias Detection in Text-to-Image Generative Models, Moreno D'Incà et al. (2024). OpenBias establishes the foundational paradigm of combining large language models and vision-question answering for open-set bias discovery in text-to-image models, directly motivating Vipera's interactive visual and LLM auditing guidance.
- Paper: DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models, Zijie J. Wang et al. (2023). DiffusionDB provides essential empirical background on real-world text-to-image prompt structures, generation failures, and user interactions that auditing tools like Vipera aim to systematically evaluate.
- Paper: Discovering and Mitigating Visual Biases Through Keyword Explanation, Younghyun Kim et al. (2024). Understanding how visual biases can be automatically surfaced and articulated into textual explanations provides key groundwork for structuring auditing criteria in generative media.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). VIEScore introduces explainable, multimodal criteria-based evaluation for conditional image generation, which informs the structured criteria organization implemented in Vipera.
- Paper: Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation, Mayu Otani et al. (2023). This work analyzes the pitfalls of standard automated metrics and the necessity of rigorous human evaluation frameworks when auditing text-to-image synthesis.
- Paper: What the DAAM: Interpreting Stable Diffusion Using Cross Attention, Raphael Tang et al. (2023). DAAM demonstrates how visual attribution and cross-attention maps reveal semantic failures in diffusion models, offering foundational visual sensemaking techniques used in generative AI audits.
No sufficiently relevant recommendations were found.
