French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English
Aurélie NévéolYoann DupontJulien BezançonKarën Fort
Presents a culturally adapted French extension of the CrowS-Pairs dataset and revised English pairs to effectively measure and compare social biases across languages in masked language models.
Artificial intelligence language models increasingly power critical daily communication and technology tools, yet they risk learning and amplifying harmful social stereotypes. Most research assessing these biases focuses strictly on the English language and United States cultural contexts. Because social biases and linguistic structures do not transfer identically across cultures, global organizations and developers need reliable benchmarks tailored to non-English languages.
The article establishes a bilingual challenge benchmark to measure social bias in French language models while correcting flaws in existing evaluation datasets. Specifically, it assesses whether leading French and multilingual language models systematically prefer stereotypical statements over non-stereotypical alternatives across ten demographic categories.
To accomplish this, the authors translated and culturally adapted 1,467 sentence pairs from the United States-focused CrowS-Pairs dataset and crowd-sourced 210 original French stereotype pairs using a citizen-science platform. Professional linguists and translators refined the data into minimal sentence pairs differing only by the targeted demographic term, correcting structural flaws present in earlier datasets. They then evaluated bias by presenting these paired sentences to three popular French language models (CamemBERT, FlauBERT, and FrALBERT) alongside a multilingual baseline (multilingual BERT) to measure probability preferences.
The evaluation revealed that French language models exhibit significant, measurable social bias. All tested models consistently favored stereotypical sentences over neutral counterparts, scoring well above the neutral benchmark threshold of 50. Stereotypical preferences were most pronounced in categories such as religion (reaching up to 72.2 in FrALBERT and 69.6 in CamemBERT), socioeconomic status, and nationality. Overall bias scores were slightly lower in French models (ranging from roughly 51 to 59) than in English models (which scored up to 65.1), and multilingual BERT demonstrated the lowest overall bias. Additionally, the study found that translation alone is insufficient for cross-lingual benchmarks: 17 original American sentences were untranslatable due to distinct cultural concepts, and French stereotypes relied heavily on explicit group naming rather than individual first names.
These findings demonstrate that organizations deploying AI models in multilingual settings cannot assume English debiasing transfers abroad, nor can they rely on simple machine-translated benchmarks. Deploying biased models creates compliance, reputational, and ethical risks in customer-facing and decision-support systems. Pretraining corpus size and model architecture substantially affect bias levels, indicating that engineering choices directly shape how stereotypes are encoded.
Organizations developing or deploying French-language AI systems should adopt tailored, culturally native bias evaluation suites rather than simple translated datasets. Teams extending bias benchmarks to additional languages should employ expert human adaptation to manage grammar rules like gender agreement and establish standardized localization strategies. Furthermore, future benchmark expansions should incorporate formal frameworks, such as Social Bias Frames, to better represent nuanced cultural contexts.
The study's primary limitations include its reliance on volunteer crowdsourcing, which yielded fewer anti-stereotypical examples, and its structural focus on masked language models rather than newer generative models. Nonetheless, the findings provide a robust, statistically validated baseline confirming the presence of systematic demographic bias in French language representations.
- Paper: StereoSet: Measuring stereotypical bias in pretrained language models, Moin Nadeem et al. (2020). StereoSet establishes a prior benchmark for measuring stereotypes in pretrained language models, clarifying the evaluation tradition that French CrowS-Pairs adapts to a new language and culture.
No sufficiently relevant recommendations were found.
