OpenBias: Open-Set Bias Detection in Text-to-Image Generative Models
Moreno D'IncàElia PeruzzoMassimiliano ManciniDejia XuVidit GoelXingqian XuZhangyang WangHumphrey ShiNicu Sebe
Proposes OpenBias, a framework combining large language models and vision question answering to automatically discover and quantify unexpected open-set biases in text-to-image generative models without relying on predefined attribute lists.
As text-to-image generative models are increasingly deployed into real-world consumer and commercial applications, ensuring their fairness and safety has become a major operational concern. These models often inherit and perpetuate systemic biases present in their training data. Traditional auditing techniques rely on predefined, closed sets of sensitive attributes—such as adult gender, age, and race—leaving a wide range of domain-specific, contextual, or previously unconsidered biases entirely undetected.
The article introduces and evaluates OpenBias, an automated framework designed to discover, assess, and quantify biases in text-to-image models in an open-set manner, without relying on predefined lists of target concepts or specialized manual training datasets.
OpenBias operates via a three-stage automated pipeline that treats the generative model as a black box. First, a large language model analyzes a dataset of text prompts to propose potential biases, associated categories, and visual inspection questions, while filtering out concepts explicitly mentioned in the original prompts. Second, the target generative model synthesizes images conditioned on these prompts across multiple random seeds. Third, a vision question answering model inspects the generated images to answer the bias-related questions, enabling the quantification of bias severity via normalized entropy scores. The framework was evaluated across text prompts from the COCO and Flickr30k datasets and tested on several widely used text-to-image systems, specifically Stable Diffusion 1.5, 2, and XL.
The findings confirm that OpenBias closely matches established closed-set benchmarks and human judgment while uncovering previously unstudied biases. When evaluated on standard demographic categories, the selected vision question answering model (LLaVA-1.5-13B) showed high agreement with FairFace classifiers across real and synthetic images. In a user study spanning 2,200 human evaluations, OpenBias demonstrated a low Absolute Mean Error of 0.15 against human-rated bias intensity and agreed with human judgment on the primary bias direction in 67% of cases. Crucially, the system exposed extensive novel biases beyond traditional categories, including strong defaults for commercial brands (such as exclusively generating Apple logos for laptops), specific animal breeds, object colors, and socioeconomic stereotyping (such as pairing non-white children with impoverished settings). Context-aware evaluations revealed that newer models like Stable Diffusion XL exhibit subtle bias amplification compared to earlier iterations.
These results demonstrate that bias in generative models extends far beyond traditional demographic categories and often depends heavily on surrounding contextual text. Organizations deploying generative AI risk perpetuating brand favoritism, harmful social stereotypes, and unrepresentative visual outputs if auditing is limited to traditional closed-set metrics. Because OpenBias operates modularly and externally without requiring internal model modifications, it offers an effective mechanism to expand current compliance, safety, and bias mitigation protocols.
Practitioners and decision-makers should integrate open-set auditing pipelines into existing model evaluation workflows before deploying text-to-image systems to the public. However, stakeholders should recognize that OpenBias relies on foundation models (LLaMA-2 and LLaVA-1.5) that may carry their own inherent biases or perceptual limitations. Additionally, the systematic role of prompt context requires deeper formal exploration. Overall, the methodology provides high confidence for identifying emerging generative risks and serves as a modular foundation that can readily incorporate more advanced evaluation models as they emerge.
- Paper: SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis, Dustin Podell et al. (2024). Introduces the architecture and training advances of Stable Diffusion XL, one of the primary generative systems audited for subtle bias amplification in OpenBias.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). Establishes the methodology of using automated language models to red-team and probe target AI systems, which OpenBias adapts to multimodal bias discovery.
- Paper: Taxonomizing and Measuring Representational Harms: A Look at Image Tagging, Jared Katzman et al. (2023). Develops a foundational taxonomy for categorizing and quantifying representational harms in visual AI systems that informs the conceptual scope of open-set bias evaluation.
- Paper: Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification, Joy Buolamwini et al. (2018). Demonstrates the benchmark closed-set demographic auditing methodology against which OpenBias compares and expands open-set detection.
- Paper: A Survey on Bias and Fairness in Machine Learning, Ninareh Mehrabi et al. (2019). Provides a comprehensive survey of the definitions, sources, and propagation of algorithmic bias across machine learning pipelines.
- Paper: Discovering and Mitigating Visual Biases Through Keyword Explanation, Younghyun Kim et al. (2024). Extends the discovery of open-ended visual biases by extracting and validating interpretable text keywords to guide active bias mitigation in vision models.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). Applies multimodal large language models to construct explainable, metric-driven evaluations for conditional image synthesis and generation quality.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). Broadens the evaluation of foundation model fairness and safety into a unified benchmark covering multiple trustworthiness dimensions across modern AI systems.
- Paper: FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts, Yichen Gong et al. (2025). Explores multimodal safety vulnerabilities in large vision-language models, the class of models relied upon as visual evaluators within automated auditing pipelines.
