BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages

Shamsuddeen Hassan MuhammadNedjma OusidhoumIdris AbdulmuminJan Philip WahleTerry RuasMeriem BeloucifChristine de KockNirmal SurangeDaniela TeodorescuIbrahim Said Ahmad

article2025ACL102 citations

Presents BRIGHTER, a benchmark of nearly 100,000 human-annotated examples spanning 28 typologically diverse, predominantly low-resource languages, providing multi-label emotion and intensity annotations to expose and address cross-lingual performance gaps in large language models.

Listen

Automated emotion recognition is an essential component across digital health, dialog systems, computational social science, and customer service. However, research and available tools remain heavily skewed toward a few high-resource Western languages, frequently relying on translated datasets that fail to capture cultural nuance. In response to this gap, the article introduces BRIGHTER, a benchmark collection of multi-labeled emotion datasets across 28 diverse languages, primarily covering under-resourced languages from Africa, Asia, Eastern Europe, and Latin America.

The main objective of the article is to establish a high-quality human-annotated benchmark for text-based emotion recognition and intensity scoring, and to evaluate how effectively state-of-the-art multilingual language models and large language models detect perceived emotions across different languages and domains.

The authors curated nearly 100,000 text instances from social media, literature, news, and speeches, engaging fluent native speakers to annotate perceived emotions—joy, sadness, anger, fear, surprise, disgust, and neutral—alongside four levels of intensity. The evaluation benchmarked both smaller multilingual models and large language models across monolingual, few-shot, and cross-lingual transfer settings.

The findings show that emotion recognition remains a challenging task for automated systems, with significant performance disparities between high-resource and under-resourced languages. Large language models struggle broadly with perceived emotion detection, achieving overall average macro F1 scores below 50% across the 28 languages. Low-resource languages exhibited the lowest performance, with some African languages scoring below 25%. Cross-lingual transfer within language families yielded mixed results, failing almost entirely on severely under-resourced languages such as Emakhuwa and Yoruba. Furthermore, large language models performed substantially better on low-resource texts when prompted in English rather than the native language, and their outputs remained highly sensitive to slight variations in prompt wording.

These results demonstrate that standard foundation models cannot yet be reliably deployed for emotion analysis across global languages without substantial adaptation. Deploying automated affective tools in cross-cultural settings risks severe misclassification, misinterpretation of user sentiment, and uneven user experiences across linguistic demographics.

For organizations building global natural language applications, the article demonstrates that prompting models in English offers higher zero- and few-shot reliability for low-resource languages, though direct fine-tuning on native data remains preferable. System architects should avoid relying on out-of-the-box large language models for emotion detection and should instead leverage targeted community-annotated datasets. Decision-makers must note that these models evaluate perceived emotions from short texts rather than true internal emotional states. Systems trained on these datasets should not be deployed in high-stakes or critical decision-making contexts without appropriate expert human oversight.

arXiv: 2502.11926
Cover for BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages

Abstract

People worldwide use language in subtle and complex ways to express emotions. Although emotion recognition—an umbrella term for several NLP tasks—impacts various applications within NLP and beyond, most work in this area has focused on high-resource languages. This has led to significant disparities in research efforts and proposed solutions, particularly for under-resourced languages, which often lack high-quality annotated datasets. In this paper, we present BRIGHTER—a collection of multi-labeled, emotion-annotated datasets in 28 different languages and across several domains. BRIGHTER primarily covers low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers. We highlight the challenges related to the data collection and annotation processes, and then report experimental results for monolingual and crosslingual multi-label emotion identification, as well as emotion intensity recognition. We analyse the variability in performance across languages and text domains, both with and without the use of LLMs, and show that the BRIGHTER datasets represent a meaningful step towards addressing the gap in text-based emotion recognition.

Table of Contents

  • 1 Introduction
  • 2 The BRIGHTER Dataset Collection
  • 2.1 Data Sources
  • 2.2 Pre-processing and Quality Control
  • 2.3 Annotating BRIGHTER
  • 2.4 Annotators' Reliability
  • 2.5 Determining the Final Labels
  • 2.6 Final Data Statistics
  • 3 Experiments
  • 3.1 Setup
  • 3.2 Experimental Results
  • 4 Analysis
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Ethical Considerations
  • Acknowledgments
  • References
  • A PLMs and LLMs Used
  • A.1 PLMs
  • A.2 LLMs
  • B Data sources
  • C Annotation
  • C.1 Annotation Guidelines and Definitions
  • C.1.1 Emotion Intensity
  • C.2 Pilot Annotation
  • D SCHMP Calculation
  • E Experimental Settings

Knowls

  1. Knowl 1 — BRIGHTER multilingual emotion dataset collection

    data/table

    BRIGHTER is a collection of human-annotated, perceived-emotion datasets covering 28 languages from seven language families: Afrikaans, Algerian Arabic, Moroccan Arabic, Chinese, German, English, Latin American Spanish, Hausa, Hindi, Igbo, Indonesian, Javanese, Kinyarwanda, Marathi, Nigerian Pidgin, Brazilian Portuguese, Mozambican Portuguese, Romanian, Russian, Sundanese, Swahili, Swedish, Tatar, Ukrainian, Emakhuwa, isiXhosa, Yoruba, and isiZulu. The paper describes the collection as containing nearly 100,000 text instances, drawn from speeches, literature, news, social media, personal narratives, reviews, and newly created text. The language inventory reports 3–122 annotators per language and 2–30 annotators per sampled instance; individual train, development, and test sets contain approximately 1,000–5,350, 108–940, and 745–4,509 instances, respectively. Twenty-four languages have training data, while Indonesian, Javanese, isiXhosa, and isiZulu are test-only datasets.

    The sources were adapted to local availability: speeches for Afrikaans; a translated and post-processed Algerian Arabic novel for Algerian Arabic; social media and news for many African languages; Weibo for Chinese; Reddit for German and English; YouTube comments for Spanish, Indonesian, Javanese, and Sundanese; Twitter for Russian and some African languages; and manually created or machine-assisted text for Hindi and Marathi. Before annotation, the authors removed duplicates, invisible characters, garbled encodings, and malformed emoticons, anonymised texts, and excluded excessive expletives and dehumanising content.

    The datasets, annotation guidelines, and individual non-aggregated annotations were publicly released so that disagreement between annotators remains available for research.

  2. Knowl 2 — Multi-label perceived-emotion and intensity annotation scheme

    definition

    Each BRIGHTER instance is annotated for perceived emotion: the emotions that most people infer the speaker may have felt from a short text, rather than the speaker’s objectively verifiable emotional state. Annotators select every applicable category from anger, sadness, fear, disgust, joy, and surprise. An instance is neutral when none of the six categories is selected, so a text can have multiple emotion labels or a neutral label.

    For every selected emotion, annotators assign an intensity from 0 to 3: 0 means no emotion, 1 slight or low emotion, 2 moderate emotion, and 3 high emotion. Emotion labels are available for all 28 languages, whereas intensity labels are provided only for the 10 datasets in which most instances were annotated by at least five people. Fluent speakers performed the annotations; English was annotated through Amazon Mechanical Turk, Russian, Ukrainian, and Tatar through Toloka, and other languages through directly recruited speakers using Label Studio or Potato.

  3. Knowl 3 — Aggregation of annotator labels

    model/method

    BRIGHTER aggregates each emotion independently from the annotators’ intensity ratings. Let Ai∈{0,1,2,3}A_i \in \{0,1,2,3\} be the rating from annotator ii, let NN be the number of annotators for an instance, and let T=0.5T=0.5. An emotion is included in the final multi-label annotation only when at least two annotators assign it a nonzero intensity and the mean rating exceeds TT:

    Lemotion=1iff∑i=1N1[Ai∈{1,2,3}]≥2 and 1N∑i=1NAi>T.L_{\mathrm{emotion}} = 1 \quad\text{iff}\quad \sum_{i=1}^{N}\mathbf{1}[A_i\in\{1,2,3\}]\geq 2 \ \text{and}\ \frac{1}{N}\sum_{i=1}^{N}A_i>T.

    Otherwise, the emotion is absent. For an emotion retained in the final label set, the annotators’ intensity ratings are averaged and converted to the reported integer intensity level. The paper’s formal interval mapping assigns levels 0, 1, and 2 to mean scores in [0,1)[0,1), [1,2)[1,2), and [2,3)[2,3), respectively, and level 3 when the mean equals 3. This aggregation deliberately retains multi-label co-occurrence while reducing the effect of isolated annotations.

  4. Knowl 4 — Split-Half Class Match Percentage for annotation reliability

    model/method

    The authors evaluate annotation reliability with Split-Half Class Match Percentage (SHCMP), an adaptation of split-half reliability for discrete emotion-intensity classes. For each of 1,000 random repetitions, annotations are divided into bins or halves; each item receives a class from each bin, and the proportion of items whose resulting classes match is recorded. The final SHCMP is the average percentage over repetitions. In the appendix’s binned-score formulation, scores lie in [−3,3][-3,3], the bin width is b=6/Kb=6/K for KK bins, and an item counts as consistent when its two bin indices c1c_1 and c2c_2 satisfy ∣c1−c2∣<1|c_1-c_2|<1.

    Higher SHCMP means that repeated annotation subsets produce similar class labels. The BRIGHTER heatmap reports generally high reliability: every dataset exceeds 60% when annotations are divided into two bins, although reliability decreases for several datasets as the number of bins increases from 2 to 10. The authors therefore characterize the aggregated annotations as reliable while retaining the original annotator-level labels because disagreement is itself informative.

  5. Knowl 5 — Baseline evaluation protocol for multilingual emotion recognition

    experimental setup

    The paper evaluates BRIGHTER on multi-label emotion classification and emotion-intensity prediction. Multi-label classification is measured with macro-F1; intensity prediction is measured with Pearson correlation. Five large language models (LLMs)—Qwen2.5-72B, Dolly-v2-12B, Llama-3.3-70B, Mixtral-8x7B, and DeepSeek-R1-70B—receive eight few-shot examples and chain-of-thought instructions, and only the first generated answer is scored. Multilingual language models (MLMs) are evaluated with LaBSE, RemBERT, XLM-R, multilingual BERT, and mDeBERTa.

    For cross-lingual classification, an MLM is trained on all available languages in a language family except the held-out target language, then tested on that target language without using target-language training data. For singleton families, the authors use specified surrogate training groups: Russian and Ukrainian for Tatar, Swahili and Yoruba for Nigerian Pidgin, and Russian for Chinese. Intensity models are trained and evaluated on the 10 languages with intensity-labeled training data. Both MLMs and LLM fine-tuning use two epochs and learning rate 10−510^{-5}. LLM decoding uses temperature 0 and top-k=1k=1, except the pass@kk ablation, which uses temperature 0.7.

  6. Knowl 6 — Multi-label classification remains difficult across BRIGHTER languages

    empirical result

    The baseline results show that perceived-emotion classification is difficult even for large models and high-resource languages. In eight-shot classification across all 28 test languages, the average macro-F1 scores were 49.71 for Qwen2.5-72B, 49.21 for DeepSeek-R1-70B, 47.12 for Llama-3.3-70B, 43.56 for Mixtral-8x7B, and 26.88 for Dolly-v2-12B. Qwen2.5-72B had the best average score, while DeepSeek-R1-70B produced the highest individual result reported for Yoruba, 27.44; Hindi and Marathi also obtained comparatively high scores, partly because roughly 80% of Hindi and 70% of Marathi test instances were single-labeled, making the task easier than genuinely multi-label prediction.

    The cross-lingual MLM averages were lower: LaBSE 40.50, RemBERT 33.63, XLM-R 30.61, mDeBERTa 32.38, and multilingual BERT 24.16 macro-F1. Performance varied sharply by language. For example, the best cross-lingual scores included 77.24 for Marathi with RemBERT, 69.96 for Hindi with XLM-R, 65.21 for Romanian with XLM-R, and 60.66 for Tatar with LaBSE, whereas Emakhuwa and Yoruba generally remained below 13. The results indicate that multilingual models can behave unreliably on languages absent from their pretraining data.

  7. Knowl 7 — Cross-lingual transfer depends on language family and pretraining coverage

    empirical result

    Cross-lingual transfer is not uniformly improved by adding more languages. Training on related languages sometimes outperforms few-shot prompting; Swedish, for example, benefits when RemBERT is fine-tuned on other Germanic languages. However, Niger-Congo languages—especially Emakhuwa—benefit least from cross-lingual transfer, with RemBERT performing particularly poorly, reflecting the severe data scarcity of these languages.

    Model behavior also follows pretraining coverage. XLM-R performs strongly for German, Chinese, Hindi, and Brazilian Portuguese but struggles for Swedish and Mozambican Portuguese. mDeBERTa is more consistent across languages overall, but performs poorly for Igbo, Emakhuwa, and Yoruba, which are not represented in the CC-100 corpus used for its relevant pretraining coverage. The authors conclude that multilingual models transfer more effectively to languages encountered during pretraining and may produce near-random or unreliable predictions for languages not represented there.

  8. Knowl 8 — Emotion-intensity prediction favors reasoning-capable LLMs on several low-resource languages

    data/table

    For the 10 languages with intensity annotations, the reported Pearson correlations were:

    Language LaBSE RemBERT XLM-R mBERT mDeBERTa Qwen2.5 Dolly Llama-3.3 Mixtral DeepSeek-R1
    arq 1.42 1.64 0.89 1.10 0.47 29.54 3.80 36.29 31.05 36.37
    chn 23.37 40.53 36.92 21.96 23.25 46.17 8.11 51.86 46.52 48.57
    deu 28.93 56.21 38.30 17.35 18.14 43.30 7.43 53.46 47.60 54.78
    eng 35.34 64.15 37.36 25.74 8.85 55.99 13.35 44.14 55.26 48.08
    esp 56.89 72.59 55.72 27.94 29.18 51.11 10.49 51.64 55.54 60.74
    hau 26.13 27.03 24.68 2.79 0.00 27.00 6.43 39.16 25.84 38.85
    ptbr 20.62 29.74 18.24 8.36 1.32 38.20 9.02 40.90 39.17 46.72
    ron 35.57 55.66 37.77 21.99 4.63 55.48 12.62 45.87 57.07 57.69
    rus 68.43 87.66 68.96 37.63 5.03 58.25 13.96 57.56 56.01 62.28
    ukr 13.75 39.94 36.16 4.32 3.51 37.74 6.04 36.99 38.74 43.54
    Average 30.54 46.61 35.25 16.35 9.97 43.03 8.74 45.78 43.97 48.88

    DeepSeek-R1-70B has the highest average correlation, 48.88, and leads on Chinese, German, Spanish, Brazilian Portuguese, and Ukrainian. RemBERT is strongest on German, English, Spanish, and Russian, illustrating that MLMs remain competitive for high-resource languages. LLMs show especially large gains on primarily spoken or low-resource languages: DeepSeek-R1-70B reaches 36.37 for Algerian Arabic, compared with 1.64 for RemBERT, an improvement of more than 34 correlation points.

  9. Knowl 9 — LLM predictions are sensitive to prompt wording, shot count, and sampling

    empirical result

    Ablations on the English test set show that LLM emotion predictions are not invariant to equivalent prompt formulations. Across three paraphrased prompts, macro-F1 varied for every model: DeepSeek-R1-70B scored 58.95, 57.98, and 58.29; Dolly-v2-12B 64.64, 59.35, and 61.15; Llama-3.3-70B 63.17, 66.34, and 62.13; Mixtral-8x7B 44.51, 42.26, and 36.53; and Qwen2.5-72B 55.09, 59.79, and 57.17.

    Increasing the number of few-shot examples generally improves performance. The gains tend to plateau at four shots for all tested models except Qwen2.5-72B, suggesting that four to eight examples are usually sufficient for stable English results. Allowing multiple sampled generations through pass@kk further increases the chance of retrieving a correct answer; DeepSeek-R1-70B exceeds 90 macro-F1 at k=8k=8, while the ranking of models at k=8k=8 remains consistent with the ranking at k=1k=1.

  10. Knowl 10 — English prompting substantially improves low-resource-language LLM performance

    empirical result

    When the same LLMs are prompted in English versus the target language, English prompts generally yield higher multi-label emotion-classification scores. The improvement is especially pronounced for low-resource languages such as Hausa, Marathi, and Emakhuwa, where Dolly-v2-12B and Llama-3.3-70B perform poorly when prompted in the target language. The principal exception is Algerian Arabic: Qwen2.5-72B performs better when prompted in Modern Standard Arabic than when prompted in English. This result indicates that multilingual LLM performance depends not only on the input language but also on the language used to specify the task.

  11. Knowl 11 — Limits on interpretation and use of BRIGHTER

    limitation

    BRIGHTER does not claim to identify speakers’ true emotions. The labels represent annotators’ perceptions from short text snippets, and perceptions can differ because of cultural background, social group, personal experience, missing context, and the absence of tone or other non-textual cues. The collection is not claimed to represent all language use or all possible emotions in the 28 languages.

    Some low-resource datasets rely on limited domains or small source pools, so they are not suitable for tasks requiring large amounts of representative language data. Text sources and annotators may introduce social and cultural biases, and inappropriate content may not have been completely removed. Systems trained on BRIGHTER may be unreliable for individual instances and sensitive to domain shift; the authors prohibit unapproved commercial or high-risk state use and warn against critical decisions about individuals, including health-related decisions, without expert oversight.

Coverage note — The appendix’s full monolingual result table, detailed emotion-frequency plots, complete split-distribution table, and example prompt templates were not made separate knowls because they support the main dataset, evaluation, and ablation contributions already captured above.

References

  1. 1.Felermino Ali, Henrique Lopes Cardoso, and Rui Sousa Silva. 2024. Building resources for emakhuwa: machine translation and news classification benchmarks. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing.
  2. 2.Magda B Arnold. 1960. Emotion and personality. vol. i. psychological aspects.
  3. 3.Etienne Barnard, Marelie H Davel, Charl van Heerden, Febe De Wet, and Jaco Badenhorst. 2014. The nchlt speech corpus of the south african languages. Workshop Spoken Language Technologies for Under-resourced Languages (SLTU).
  4. 4.L.F. Barrett. 2017. How Emotions are Made: The Secret Life of the Brain. Expert Thinking Series. Macmillan.
  5. 5.Lisa Feldman Barrett. 2016. The theory of constructed emotion: an active inference account of interoception and categorization. Social Cognitive and Affective Neuroscience, 12(1):1–23.
  6. 6.Federico Bianchi, Debora Nozza, and Dirk Hovy. 2021. FEEL-IT: Emotion and sentiment classification for the Italian language. In Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pages 76–83, Online. Association for Computational Linguistics.
  7. 7.Federico Bianchi, Debora Nozza, and Dirk Hovy. 2022. Xlm-emo: Multilingual emotion prediction in social media text. In Proceedings of the 12th Workshop on Computational Approaches to Subjectivity, Sentiment & Social Media Analysis, pages 195–203.
  8. 8.Kateryna Bobrovnyk. 2019. Automated building and analysis of ukrainian twitter corpus for toxic text detection. In COLINS 2019. Volume II: Workshop.
  9. 9.Ankush Chatterjee, Kedhar Nath Narahari, Meghana Joshi, and Puneet Agrawal. 2019. SemEval-2019 task 3: EmoContext contextual emotion detection in text. In Proceedings of the 13th International Workshop on Semantic Evaluation, pages 39–48, Minneapolis, Minnesota, USA. Association for Computational Linguistics.
  10. 10.Alexandra Ciobotaru, Mihai Vlad Constantinescu, Liviu P. Dinu, and Stefan Dumitrescu. 2022. RED v2: Enhancing RED dataset for multi-label emotion detection. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 1392–1399, Marseille, France. European Language Resources Association.
  11. 11.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  12. 12.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm.
  13. 13.DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, and 181 others. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. Preprint, arXiv:2501.12948.
  14. 14.Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions: A dataset of fine-grained emotions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4040–4054, Online. Association for Computational Linguistics.
  15. 15.Paul Ekman. 1992. Are there basic emotions?
  16. 16.PC Ellsworth. 2013. Appraisal theory: old and new questions. emot. rev. 5, 125–131.
  17. 17.Nico H Frijda. 1986. The emotions. Studies in Emotion and Social Interaction.
  18. 18.Shreya Havaldar, Bhumika Singhal, Sunny Rai, Langchen Liu, Sharath Chandra Guntuku, and Lyle Ungar. 2023. Multilingual language models are not multicultural: A case study in emotion. In Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Media Analysis, pages 202–214, Toronto, Canada. Association for Computational Linguistics.
  19. 19.Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022. Challenges and strategies in cross-cultural NLP. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6997–7013, Dublin, Ireland. Association for Computational Linguistics.
  20. 20.MD Asif Iqbal, Avishek Das, Omar Sharif, Mohammed Moshiul Hoque, and Iqbal H Sarker. 2022. Bemoc: A corpus for identifying emotion in bengali texts. SN Computer Science, 3(2):135.
  21. 21.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, and 1 others. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  22. 22.Svetlana Kiritchenko and Saif M. Mohammad. 2016. Capturing reliable fine-grained sentiment associations by crowdsourcing and best–worst scaling. In Proceedings of The 15th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL), San Diego, California.
  23. 23.Irina Krylova, Boris Orekhov, Ekaterina Stepanova, and Lyudmila Zaydelman. 2016. Languages of russia: Using social networks to collect texts. Information Retrieval: 9th Russian Summer School, RuSSIR 2015, Saint Petersburg, Russia, August 24-28, 2015, Revised Selected Papers 9, pages 179–185.
  24. 24.Shivani Kumar, Anubhav Shrimal, Md Shad Akhtar, and Tanmoy Chakraborty. 2022. Discovering emotion and reasoning its flip in multi-party conversations using masked memory network and transformer. Knowledge-Based Systems, 240:108112.
  25. 25.Richard S Lazarus. 1991. Emotion and adaptation, volume 557. Oxford University Press.
  26. 26.Youssef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov, Philip Torr, Kenneth Church, and Mohamed Elhoseiny. 2024. No culture left behind: ArtELingo-28, a benchmark of WikiArt with captions in 28 languages. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 20939–20962, Miami, Florida, USA. Association for Computational Linguistics.
  27. 27.Saif Mohammad. 2023. Best practices in the creation and use of emotion lexicons. In Findings of the Association for Computational Linguistics: EACL 2023, pages 1825–1836, Dubrovnik, Croatia. Association for Computational Linguistics.
  28. 28.Saif Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018a. Semeval-2018 task 1: Affect in tweets. In Proceedings of the 12th international workshop on semantic evaluation, pages 1–17.
  29. 29.Saif Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018b. SemEval-2018 task 1: Affect in tweets. In Proceedings of the 12th International Workshop on Semantic Evaluation, pages 1–17, New Orleans, Louisiana. Association for Computational Linguistics.
  30. 30.Saif Mohammad and Svetlana Kiritchenko. 2018. Understanding emotions: A dataset of tweets to study interactions between affect categories. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  31. 31.Saif M. Mohammad. 2016. 9 - sentiment analysis: Detecting valence, emotions, and other affectual states from text. In Herbert L. Meiselman, editor, Emotion Measurement, pages 201–237. Woodhead Publishing.
  32. 32.Saif M. Mohammad. 2022. Ethics sheet for automatic emotion recognition and sentiment analysis. Preprint, arXiv:2109.08256.
  33. 33.Saif M. Mohammad. 2024. WorryWords: Norms of anxiety association for over 44k English words. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 16261–16278, Miami, Florida, USA. Association for Computational Linguistics.
  34. 34.Agnes Moors, Phoebe C Ellsworth, Klaus R Scherer, and Nico H Frijda. 2013. Appraisal theories of emotion: State of the art and future development. Emotion review, 5(2):119–124.
  35. 35.Shamsuddeen Muhammad, Idris Abdulmumin, Abinew Ayele, Nedjma Ousidhoum, David Adelani, Seid Yimam, Ibrahim Ahmad, Meriem Beloucif, Saif Mohammad, Sebastian Ruder, Oumaima Hourrane, Alipio Jorge, Pavel Brazdil, Felermino Ali, Davis David, Salomey Osei, Bello Shehu-Bello, Falalu Lawan, Tajuddeen Gwadabe, and 8 others. 2023a. AfriSenti: A Twitter sentiment analysis benchmark for African languages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13968–13981, Singapore. Association for Computational Linguistics.
  36. 36.Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Seid Muhie Yimam, David Ifeoluwa Adelani, Ibrahim Sa’id Ahmad, Nedjma Ousidhoum, Abinew Ayele, Saif M. Mohammad, Meriem Beloucif, and Sebastian Ruder. 2023b. SemEval-2023 task 12: Sentiment analysis for african languages (AfriSenti-SemEval). In Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023). Association for Computational Linguistics.
  37. 37.Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Seid Muhie Yimam, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine De Kock, Tadesse Destaw Belay, Ibrahim Said Ahmad, Nirmal Surange, Daniela Teodorescu, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino Ali, Vladimir Araujo, Abinew Ali Ayele, Oana Ignat, Alexander Panchenko, and 2 others. 2025. Semeval-2025 task 11: Bridging the gap in text-based emotion detection. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria. Association for Computational Linguistics.
  38. 38.Emily Öhman, Marc Pàmies, Kaisla Kajava, and Jörg Tiedemann. 2020. Xed: A multilingual dataset for sentiment analysis and emotion detection. arXiv preprint arXiv:2011.01612.
  39. 39.Andrew Ortony, Gerald L Clore, and Allan Collins. 2022. The cognitive structure of emotions. Cambridge university press.
  40. 40.Jessica Ouyang and Kathleen McKeown. 2015. Modeling reportable events as turning points in narrative. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 2149–2158.
  41. 41.Jiaxin Pei, Aparna Ananthasubramaniam, Xingyao Wang, Naitian Zhou, Apostolos Dedeloudis, Jackson Sargent, and David Jurgens. 2022. Potato: The portable text annotation tool. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations.
  42. 42.Barbara Plank. 2022. The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671–10682.
  43. 43.Flor Miriam Plaza-del Arco, Alba A. Cercas Curry, Amanda Cercas Curry, and Dirk Hovy. 2024. Emotion analysis in NLP: Trends, gaps and roadmap for future directions. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 5696–5710, Torino, Italia. ELRA and ICCL.
  44. 44.Robert Plutchik. 1980. Chapter 1 - a general psycho-evolutionary theory of emotion. In Robert Plutchik and Henry Kellerman, editors, Theories of Emotion, pages 3–33. Academic Press.
  45. 45.Ira J Roseman. 2013. Appraisal in the emotion system: Coherence in strategies for coping. Emotion Review, 5(2):141–149.
  46. 46.Alieh Hajizadeh Saffar, Tiffany Katharine Mann, and Bahadorreza Ofoghi. 2023. Textual emotion detection in health: Advances and applications. Journal of Biomedical Informatics, 137:104258.
  47. 47.Mei Silviana Saputri, Rahmad Mahendra, and Mirna Adriani. 2018. Emotion classification on indonesian twitter dataset. In 2018 International Conference on Asian Language Processing (IALP), pages 90–95.
  48. 48.Klaus R Scherer. 2009. The dynamic architecture of emotion: Evidence for the component process model. Cognition and emotion, 23(7):1307–1351.
  49. 49.Armin Seyeditabari, Narges Tabari, and Wlodek Zadrozny. 2018. Emotion detection in text: a review. arXiv preprint arXiv:1806.00674.
  50. 50.Språkbanken Text. 2024. Svensk absabank.
  51. 51.Carlo Strapparava and Rada Mihalcea. 2007. SemEval-2007 task 14: Affective text. In Proceedings of the Fourth International Workshop on Semantic Evaluations (SemEval-2007), pages 70–74, Prague, Czech Republic. Association for Computational Linguistics.
  52. 52.Maxim Tkachenko, Mikhail Malyuk, Andrey Holmanyuk, and Nikolai Liubimov. 2020-2025. Label Studio: Data labeling software. Open source software available from https://github.com/HumanSignal/label-studio.
  53. 53.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, and 1 others. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  54. 54.Krishnapriya Vishnubhotla, Daniela Teodorescu, Mallory J Feldman, Kristen Lindquist, and Saif M. Mohammad. 2024. Emotion granularity from text: An aggregate-level indicator of mental health. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 19168–19185, Miami, Florida, USA. Association for Computational Linguistics.
  55. 55.Janyce Wiebe, Theresa Wilson, and Claire Cardie. 2005. Annotating expressions of opinions and emotions in language. Language resources and evaluation, 39:165–210.
  56. 56.An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, and 1 others. 2024. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115.
  57. 57.Yuan Zhuang, Tianyu Jiang, and Ellen Riloff. 2024. My heart skipped a beat! recognizing expressions of embodied emotion in natural language. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 3525–3537.

Citation

MLA
Muhammad, S. H., et al. “BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 8895–916, https://doi.org/10.18653/v1/2025.acl-long.436.
APA
Muhammad, S. H., Ousidhoum, N., Abdulmumin, I., Wahle, J. P., Ruas, T., Beloucif, M., de Kock, C., Surange, N., Teodorescu, D., Ahmad, I. S., Adelani, D. I., Aji, A. F., Ali, F. D. M. A., Alimova, I., Araujo, V., Babakov, N., Baes, N., Bucur, A.-M., Bukula, A., … Mohammad, S. M. (2025). BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8895–8916. https://doi.org/10.18653/v1/2025.acl-long.436
Chicago
Muhammad, S. H., N. Ousidhoum, I. Abdulmumin, et al. 2025. “BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8895–8916. https://doi.org/10.18653/v1/2025.acl-long.436.
Harvard
Muhammad, S.H. et al. (2025) “BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages”, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8895–8916. Available at: https://doi.org/10.18653/v1/2025.acl-long.436.
Vancouver
1. Muhammad SH, Ousidhoum N, Abdulmumin I, et al (2025) BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8895–8916

BibTeX

@inproceedings{Muhammad_2025, title={BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages}, url={http://dx.doi.org/10.18653/v1/2025.acl-long.436}, DOI={10.18653/v1/2025.acl-long.436}, booktitle={Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, publisher={Association for Computational Linguistics}, author={Muhammad, Shamsuddeen Hassan and Ousidhoum, Nedjma and Abdulmumin, Idris and Wahle, Jan Philip and Ruas, Terry and Beloucif, Meriem and de Kock, Christine and Surange, Nirmal and Teodorescu, Daniela and Ahmad, Ibrahim Said and Adelani, David Ifeoluwa and Aji, Alham Fikri and Ali, Felermino D. M. A. and Alimova, Ilseyar and Araujo, Vladimir and Babakov, Nikolay and Baes, Naomi and Bucur, Ana-Maria and Bukula, Andiswa and Cao, Guanqun and Tufiño, Rodrigo and Chevi, Rendi and Chukwuneke, Chiamaka Ijeoma and Ciobotaru, Alexandra and Dementieva, Daryna and Gadanya, Murja Sani and Geislinger, Robert and Gipp, Bela and Hourrane, Oumaima and Ignat, Oana and Lawan, Falalu Ibrahim and Mabuya, Rooweither and Mahendra, Rahmad and Marivate, Vukosi and Panchenko, Alexander and Piper, Andrew and Ferreira, Charles Henrique Porto and Protasov, Vitaly and Rutunda, Samuel and Shrivastava, Manish and Udrea, Aura Cristina and Wanzare, Lilian Diana Awuor and Wu, Sophie and Wunderlich, Florian Valentin and Zhafran, Hanif Muhammad and Zhang, Tianhui and Zhou, Yi and Mohammad, Saif M.}, year={2025}, pages={8895–8916} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/