BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages
Shamsuddeen Hassan MuhammadNedjma OusidhoumIdris AbdulmuminJan Philip WahleTerry RuasMeriem BeloucifChristine de KockNirmal SurangeDaniela TeodorescuIbrahim Said Ahmad
Presents BRIGHTER, a benchmark of nearly 100,000 human-annotated examples spanning 28 typologically diverse, predominantly low-resource languages, providing multi-label emotion and intensity annotations to expose and address cross-lingual performance gaps in large language models.
Automated emotion recognition is an essential component across digital health, dialog systems, computational social science, and customer service. However, research and available tools remain heavily skewed toward a few high-resource Western languages, frequently relying on translated datasets that fail to capture cultural nuance. In response to this gap, the article introduces BRIGHTER, a benchmark collection of multi-labeled emotion datasets across 28 diverse languages, primarily covering under-resourced languages from Africa, Asia, Eastern Europe, and Latin America.
The main objective of the article is to establish a high-quality human-annotated benchmark for text-based emotion recognition and intensity scoring, and to evaluate how effectively state-of-the-art multilingual language models and large language models detect perceived emotions across different languages and domains.
The authors curated nearly 100,000 text instances from social media, literature, news, and speeches, engaging fluent native speakers to annotate perceived emotions—joy, sadness, anger, fear, surprise, disgust, and neutral—alongside four levels of intensity. The evaluation benchmarked both smaller multilingual models and large language models across monolingual, few-shot, and cross-lingual transfer settings.
The findings show that emotion recognition remains a challenging task for automated systems, with significant performance disparities between high-resource and under-resourced languages. Large language models struggle broadly with perceived emotion detection, achieving overall average macro F1 scores below 50% across the 28 languages. Low-resource languages exhibited the lowest performance, with some African languages scoring below 25%. Cross-lingual transfer within language families yielded mixed results, failing almost entirely on severely under-resourced languages such as Emakhuwa and Yoruba. Furthermore, large language models performed substantially better on low-resource texts when prompted in English rather than the native language, and their outputs remained highly sensitive to slight variations in prompt wording.
These results demonstrate that standard foundation models cannot yet be reliably deployed for emotion analysis across global languages without substantial adaptation. Deploying automated affective tools in cross-cultural settings risks severe misclassification, misinterpretation of user sentiment, and uneven user experiences across linguistic demographics.
For organizations building global natural language applications, the article demonstrates that prompting models in English offers higher zero- and few-shot reliability for low-resource languages, though direct fine-tuning on native data remains preferable. System architects should avoid relying on out-of-the-box large language models for emotion detection and should instead leverage targeted community-annotated datasets. Decision-makers must note that these models evaluate perceived emotions from short texts rather than true internal emotional states. Systems trained on these datasets should not be deployed in high-stakes or critical decision-making contexts without appropriate expert human oversight.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). Provides foundational methodology for benchmarking generative LLMs across diverse language families and highlights how prompt language affects low-resource performance.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). Establishes the taxonomy and theoretical framing for resource disparities and linguistic exclusion across global languages in NLP research.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). Demonstrates techniques and baseline benchmarks for extending language models and NLP evaluation to hundreds of under-resourced languages.
- Paper: XNLI: Evaluating Cross-lingual Sentence Representations, Alexis Conneau et al. (2018). Introduces foundational standards and transfer methodologies for evaluating cross-lingual sentence understanding across diverse language typologies.
- Paper: CROWDSOURCING A WORD–EMOTION ASSOCIATION LEXICON, Saif M. Mohammad et al. (2013). Pioneers crowdsourced human annotation frameworks for categorical emotion associations in textual data.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). Builds on empirical observations of cross-lingual model failures by introducing an internal metric to quantify representation disparities between high- and low-resource languages in LLMs.
- Paper: MMTEB: Massive Multilingual Text Embedding Benchmark, Kenneth C. Enevoldsen et al. (2025). Expands massively multilingual benchmarking beyond emotion classification to broad multilingual text embeddings across over 250 languages.
