French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English

Aurélie NévéolYoann DupontJulien BezançonKarën Fort

article2022ACL60 citations

Presents a culturally adapted French extension of the CrowS-Pairs dataset and revised English pairs to effectively measure and compare social biases across languages in masked language models.

Listen

Artificial intelligence language models increasingly power critical daily communication and technology tools, yet they risk learning and amplifying harmful social stereotypes. Most research assessing these biases focuses strictly on the English language and United States cultural contexts. Because social biases and linguistic structures do not transfer identically across cultures, global organizations and developers need reliable benchmarks tailored to non-English languages.

The article establishes a bilingual challenge benchmark to measure social bias in French language models while correcting flaws in existing evaluation datasets. Specifically, it assesses whether leading French and multilingual language models systematically prefer stereotypical statements over non-stereotypical alternatives across ten demographic categories.

To accomplish this, the authors translated and culturally adapted 1,467 sentence pairs from the United States-focused CrowS-Pairs dataset and crowd-sourced 210 original French stereotype pairs using a citizen-science platform. Professional linguists and translators refined the data into minimal sentence pairs differing only by the targeted demographic term, correcting structural flaws present in earlier datasets. They then evaluated bias by presenting these paired sentences to three popular French language models (CamemBERT, FlauBERT, and FrALBERT) alongside a multilingual baseline (multilingual BERT) to measure probability preferences.

The evaluation revealed that French language models exhibit significant, measurable social bias. All tested models consistently favored stereotypical sentences over neutral counterparts, scoring well above the neutral benchmark threshold of 50. Stereotypical preferences were most pronounced in categories such as religion (reaching up to 72.2 in FrALBERT and 69.6 in CamemBERT), socioeconomic status, and nationality. Overall bias scores were slightly lower in French models (ranging from roughly 51 to 59) than in English models (which scored up to 65.1), and multilingual BERT demonstrated the lowest overall bias. Additionally, the study found that translation alone is insufficient for cross-lingual benchmarks: 17 original American sentences were untranslatable due to distinct cultural concepts, and French stereotypes relied heavily on explicit group naming rather than individual first names.

These findings demonstrate that organizations deploying AI models in multilingual settings cannot assume English debiasing transfers abroad, nor can they rely on simple machine-translated benchmarks. Deploying biased models creates compliance, reputational, and ethical risks in customer-facing and decision-support systems. Pretraining corpus size and model architecture substantially affect bias levels, indicating that engineering choices directly shape how stereotypes are encoded.

Organizations developing or deploying French-language AI systems should adopt tailored, culturally native bias evaluation suites rather than simple translated datasets. Teams extending bias benchmarks to additional languages should employ expert human adaptation to manage grammar rules like gender agreement and establish standardized localization strategies. Furthermore, future benchmark expansions should incorporate formal frameworks, such as Social Bias Frames, to better represent nuanced cultural contexts.

The study's primary limitations include its reliance on volunteer crowdsourcing, which yielded fewer anti-stereotypical examples, and its structural focus on masked language models rather than newer generative models. Nonetheless, the findings provide a robust, statistically validated baseline confirming the presence of systematic demographic bias in French language representations.

Névéol et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English

Table of Contents

  • 1 Introduction
  • 2 Corpus development
  • 3 Measuring Bias in masked language models for English and French
  • 4 Corpus analysis
  • 4.1 Comments on the translation process
  • 4.2 Comparison to CrowS-Pairs
  • 4.3 Recommendations for further extension to other languages.
  • 4.4 Expression of bias in corpus
  • 5 Related work
  • 6 Conclusion
  • 7 Ethical Considerations and limitations of this study
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Sample of newly collected stereotypes and their translation into English.
  • A.2 Data Statement

Knowls

  1. Knowl 1 — French CrowS-Pairs corpus composition

    definition

    The enriched CrowS-Pairs challenge corpus contains 1,677 French sentence pairs: 1,467 translated and adapted from the original English CrowS-Pairs material and 210 newly collected French contributions. Each pair contrasts a more stereotypical statement about a socially disadvantaged group with a less stereotypical counterpart concerning an advantaged group. The dataset covers ethnicity/color, gender identity or expression, socioeconomic status/occupation, nationality, religion, age, sexual orientation, physical appearance, and disability; an additional other category is used for new examples that do not fit those nine categories. The resource also includes English translations of the new French material and revised English versions of original pairs.

  2. Knowl 2 — Translation and repair of CrowS-Pairs examples

    model/method

    To produce French examples, the authors divided the 1,508 original CrowS-Pairs into 16 batches of 90 pairs and one batch of 68. A French-speaking author translated a selected sentence from each pair, and another author reviewed and validated the translation. The team adapted names and cultural references where needed, and flagged examples that were culturally untranslatable. They also repaired three source-pair problems in both the French translations and a revised English corpus: non-minimal pairs, where unrelated wording also changed; double switches, where an extra change altered the sentence meaning; and bias mismatches, where the pair actually instantiated different bias categories. The reported number of pairs affected by these repairs was 22 non-minimal pairs, 64 double switches, and 64 bias mismatches. Across all translation and adaptation types, 670 pairs were affected: names 361, origin 97, country/location 22, religion 7, sport 6, food 6, other 21, US-culture adaptation 24, and untranslatable examples 17, in addition to the three repair categories.

  3. Knowl 3 — Collection and curation of native French stereotypes

    model/method

    The authors used the LanguageARC citizen-science platform to collect French stereotypes relevant to French social and cultural contexts. Contributors could submit a stereotyped French statement and select among the nine CrowS-Pairs bias categories plus other; separate tasks asked contributors to judge translation fluency or classify the bias category of a translated example. From August 1 to October 1, 2021, 26 users submitted 229 raw statements. The authors merged strict duplicates, manually checked near-duplicates and categories, split statements containing multiple stereotypes, and removed examples for which they could not identify a bias or construct a minimal pair. They retained 210 contributions and translated them into English, with feedback from a native English speaker. Nationality and gender were the most common categories in the published distribution of new contributions: 64 (30.2%) and 60 (28.3%), respectively. That distribution reports 212 statements in total, whereas the curation narrative reports 210 retained contributions.

  4. Knowl 4 — Validation results reveal uneven translation and category agreement

    empirical result

    In the fluency task, contributors assessed 426 translated sentences: 336 (79%) were validated as fluent, and 90 received correction suggestions that the authors used to revise the translations. Bias-category annotation was less consistent: Krippendorff’s alpha was 0.41. Of the assessed translations, 1,310 (50%) received the same category as the original CrowS-Pairs sentence, while 481 (19%) were assigned multiple categories that included the original category. The remaining annotations were classified as irrelevant to any category (18%), relevant to other (2%), or relevant to a different category from the original (11%). The authors’ manual review suggested that divergent judgments often reflected culturally specific examples or ambiguity about whether a sentence instantiated a stereotype on its own.

  5. Knowl 5 — Masked-language-model evaluation protocol

    experimental setup

    The authors evaluated masked language models by comparing their preference for the more stereotypical sentence in each pair. A metric score of 50 represents no preference; a score above 50 indicates greater preference for stereotypical sentences. They also report stereo and anti-stereo scores, which adjust the comparison according to the bias orientation of the target group. The French evaluation used the base versions of CamemBERT, FlauBERT, and FrALBERT, plus multilingual BERT (mBERT). English evaluations included BERT, RoBERTa, and mBERT. To improve comparability, the English evaluation used the revised CrowS-Pairs corpus and excluded examples judged untranslatable or strongly tied to US culture; it also included the English translations of the newly collected French examples. Experiments ran on a single GPU. In a reproduction check on the original English corpus, the reported metric scores were 60.5 for BERT, 65.4 for RoBERTa, and 60.5 for ALBERT.

  6. Knowl 6 — Overall evaluation finds stereotypical preferences in most models

    empirical result

    On the enriched corpus, the overall metric scores for French were 59.3 for CamemBERT, 53.7 for FlauBERT, 55.9 for FrALBERT, and 50.9 for mBERT. For English they were 52.9 for mBERT, 61.3 for BERT, and 65.1 for RoBERTa. All scores except French mBERT’s were significantly above 50 in the reported t-tests (p<0.05p<0.05), indicating preference for stereotypical sentences under this evaluation. The model-score differences were significant for English; for French, differences between FrALBERT and FlauBERT and between FlauBERT and mBERT were not significant. The authors characterize English-model bias as generally higher than French or multilingual-model bias, while noting that the English scores changed little when they evaluated the revised and filtered corpus instead of the original.

  7. Knowl 7 — Bias scores vary substantially across categories and models

    data/table

    The category-level metric scores show that the overall preference for stereotypical sentences conceals substantial variation. Scores below 50 indicate preference in the opposite direction; 50 indicates no preference. In the sequences below, model order is French CamemBERT, FlauBERT, FrALBERT, and mBERT, followed by English mBERT, BERT, and RoBERTa. Each category also lists its number of pairs and percentage of the 1,677-pair evaluation set.

    • Ethnicity/color, 460 pairs (27.4%): 58.6, 51.4, 56.7, 47.3; 54.4, 59.3, 62.9.
    • Gender, 321 pairs (19.1%): 54.8, 51.7, 47.7, 48.0; 46.2, 58.4, 58.4.
    • Socioeconomic status, 196 pairs (11.7%): 64.3, 54.1, 58.2, 56.1; 52.4, 57.1, 67.2.
    • Nationality, 253 pairs (15.1%): 60.1, 53.0, 60.5, 53.4; 50.9, 60.6, 64.8.
    • Religion, 115 pairs (6.9%): 69.6, 63.5, 72.2, 51.3; 56.8, 71.2, 71.2.
    • Age, 90 pairs (5.4%): 61.1, 58.9, 38.9, 54.4; 50.5, 53.9, 71.4.
    • Sexual orientation, 91 pairs (5.4%): 50.5, 47.2, 81.3, 55.0; 65.6, 65.6, 65.6.
    • Physical appearance, 72 pairs (4.3%): 58.3, 51.4, 40.3, 51.4; 59.7, 66.7, 76.4.
    • Disability, 66 pairs (3.9%): 63.6, 65.2, 42.4, 54.5; 50.8, 61.5, 69.2.
    • Other, 13 pairs (0.8%): 53.9, 61.5, 53.9, 46.1; 27.3, 72.7, 63.6.

    The strongest category-level scores differ by model: FrALBERT scores 81.3 for sexual orientation and 72.2 for religion, while English RoBERTa scores 76.4 for physical appearance and 71.4 for age. Several model-category scores fall below 50, demonstrating that aggregate scores do not imply a uniform preference across bias types.

  8. Knowl 8 — Native and translated portions produce different model scores

    empirical result

    The comparison between native and translated examples indicates that measured preferences depend on which portion of the corpus is evaluated. For French, native examples (210 pairs) yielded scores of 56.1 for CamemBERT, 47.2 for FlauBERT, 54.3 for FrALBERT, and 57.1 for mBERT; translated examples (1,467 pairs) yielded 59.9, 54.4, 55.6, and 50.2, respectively. For English, the native portion (1,508 pairs) yielded 60.9 for BERT, 65.2 for RoBERTa, and 53.0 for mBERT; the translated portion (210 pairs) yielded 53.8, 62.9, and 50.0. These are reported score comparisons, not claims that the native-versus-translated differences are statistically significant.

  9. Knowl 9 — Cultural adaptation and native collection address different gaps

    model/method

    The translation work identified examples whose stereotypes were culturally specific to the United States, culturally weaker in France, or difficult to translate without changing the bias category. The authors marked 24 pairs as US-culture examples and 17 as untranslatable. They also found cases where an ethnicity term paired with a nationality term in English would mix bias categories in French, so they adapted the pair to preserve a single category. Conversely, newly collected French examples captured local references, regional stereotypes, and idiomatic phrasing that translated examples did not necessarily contain. On this basis, the authors recommend direct human translation rather than editing machine translation, creative but minimal-pair-preserving phrasing to accommodate grammatical differences such as French gender agreement, consistent policies for adapting names and locations, and native data collection alongside translation. They also suggest that formal descriptions of the social frames in examples could support more precise cross-cultural comparison.

  10. Knowl 10 — Study scope and dataset-use limitations

    limitation

    The dataset broadens cultural coverage beyond the United States but represents only French and US contexts. Its new examples came from unpaid volunteers recruited through channels accessible to the French research community, so the participant pool may not represent the wider population; the authors note that requiring a platform account may also have discouraged participation. The new collection contains only one statement consistent with an anti-stereotype, which the authors suggest may reflect how the collection task was phrased. The benchmark is designed primarily for masked language models, a subset of language models, although the authors propose comparing sentence-pair perplexities as a possible adaptation for causal or generative models. They caution that exposing models to the benchmark during training would undermine its use for bias assessment.

Coverage note — Detailed examples of individual sentence translations and model-architecture descriptions are omitted because they illustrate the corpus and experiments but do not add separate methods or findings beyond the knowls above.

References

  1. 1.Marion Bartl, Malvina Nissim, and Albert Gatt. 2020. Unmasking contextual stereotypes: Measuring and mitigating BERT’s gender bias. In Proceedings of the Second Workshop on Gender Bias in Natural Language Processing, pages 1–16, Barcelona, Spain (Online). Association for Computational Linguistics.
  2. 2.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? . In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623, New York, NY, USA. Association for Computing Machinery.
  3. 3.Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online. Association for Computational Linguistics.
  4. 4.Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, Online. Association for Computational Linguistics.
  5. 5.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186.
  6. 6.Oralie Cattan, Christophe Servan, and Sophie Rosset. 2021. On the usability of transformers-based models for a french question-answering task. In Recent Advances in Natural Language Processing (RANLP).
  7. 7.Jon Chamberlain, Karën Fort, Udo Kruschwitz, Mathieu Lafourcade, and Massimo Poesio. 2013. Using games to create language resources: Successes and limitations of the approach. In Iryna Gurevych and Jungi Kim, editors, The People’s Web Meets NLP, Theory and Applications of Natural Language Processing, pages 3–44. Springer Berlin Heidelberg.
  8. 8.K. Bretonnel Cohen, Jingbo Xia, Pierre Zweigenbaum, Tiffany Callahan, Orin Hargraves, Foster Goss, Nancy Ide, Aurélie Névéol, Cyril Grouin, and Lawrence E. Hunter. 2018. Three Dimensions of Reproducibility in Natural Language Processing. In Proceedings of LREC, page 156–165.
  9. 9.Hal Daumé III. 2016. Language bias and black sheep.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.James Fiumara, Christopher Cieri, Jonathan Wright, and Mark Liberman. 2020. LanguageARC: Developing language resources through citizen linguistics. In Proceedings of the LREC 2020 Workshop on “Citizen Linguistics in Language Resource Development”, pages 1–6, Marseille, France. European Language Resources Association.
  12. 12.Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sanchez, Mugdha Pandya, and Adam Lopez. 2021. Intrinsic bias metrics do not correlate with application bias. In Proceedings of ACL 2021.
  13. 13.Dirk Hovy and Shannon L. Spruit. 2016. The social impact of natural language processing. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 591–598, Berlin, Germany. Association for Computational Linguistics.
  14. 14.Ann Irvine, John Morgan, Marine Carpuat, Hal Daumé III, and Dragos Munteanu. 2013. Measuring machine translation errors in new domains. Transactions of the Association for Computational Linguistics, 1:429–440.
  15. 15.Mascha Kurpicz-Briki. 2020. Cultural differences in bias? origin and gender bias in pre-trained german and french word embeddings. In Proceedings of the 5th Swiss Text Analytics Conference (SwissText) & 16th Conference on Natural Language Processing (KONVENS), volume 2624, Zurich, Switzerland (held online due to COVID19 pandemic). CEUR Workshop proceedings.
  16. 16.Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. Albert: A lite bert for self-supervised learning of language representations. In International Conference on Learning Representations.
  17. 17.Hang Le, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoit Crabbé, Laurent Besacier, and Didier Schwab. 2020. FlauBERT: Unsupervised language model pre-training for French. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 2479–2490, Marseille, France. European Language Resources Association.
  18. 18.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, M. Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. ArXiv, abs/1907.11692.
  19. 19.Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020. CamemBERT: a tasty French language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7203–7219, Online. Association for Computational Linguistics.
  20. 20.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1953–1967, Online. Association for Computational Linguistics.
  21. 21.Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019. Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures. In 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7), Cardiff, United Kingdom. Leibniz-Institut für Deutsche Sprache.
  22. 22.Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2021. Social bias frames: Reasoning about social and power implications of language. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, Online. Association for Computational Linguistics.
  23. 23.Jean-Paul Vinay and Jean Darbelnet. 1958. Stylistique comparée du français et de l’anglais [Texte imprimé] : méthode de traduction / J.P. Vinay, J. Darbelnet. Bibliothèque de stylistique comparée. Didier, Paris.
  24. 24.Jieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang, and Ahmed Hassan Awadallah. 2020. Gender bias in multilingual embeddings and cross-lingual transfer. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2896–2907, Online. Association for Computational Linguistics.

Citation

MLA
Neveol, A., et al. “French CrowS-Pairs: Extending a Challenge Dataset for Measuring Social Bias in Masked Language Models to a Language Other Than English”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 8521–31, https://doi.org/10.18653/v1/2022.acl-long.583.
APA
Neveol, A., Dupont, Y., Bezançon, J., & Fort, K. (2022). French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8521–8531. https://doi.org/10.18653/v1/2022.acl-long.583
Chicago
Neveol, A., Y. Dupont, J. Bezançon, and K. Fort. 2022. “French CrowS-Pairs: Extending a Challenge Dataset for Measuring Social Bias in Masked Language Models to a Language Other Than English”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8521–31. https://doi.org/10.18653/v1/2022.acl-long.583.
Harvard
Neveol, A. et al. (2022) “French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8521–8531. Available at: https://doi.org/10.18653/v1/2022.acl-long.583.
Vancouver
1. Neveol A, Dupont Y, Bezançon J, Fort K (2022) French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8521–8531

BibTeX

@inproceedings{neveol-etal-2022-french,
    title = "{F}rench {C}row{S}-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than {E}nglish",
    author = {N{\'e}v{\'e}ol, Aur{\'e}lie  and
      Dupont, Yoann  and
      Bezan{\c{c}}on, Julien  and
      Fort, Kar{\"e}n},
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.583/",
    doi = "10.18653/v1/2022.acl-long.583",
    pages = "8521--8531"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/