EmoBench: Evaluating the Emotional Intelligence of Large Language Models

Sahand SabourSiyang LiuZheyuan ZhangJune M. LiuJinfeng ZhouAlvionna S. SunaryoTatia M. C. LeeRada MihalceaMinlie Huang

article2024ACL82 citations

Proposes EmoBench, a bilingual evaluation benchmark grounded in psychological theories that tests large language models on both emotional understanding and application, exposing a significant performance gap between leading models and human capabilities.

Listen

Large language models are increasingly deployed in sensitive, interpersonal domains such as customer service, education, and mental health support. However, existing benchmarks primarily evaluate emotional intelligence through simple pattern recognition and text extraction, ignoring complex reasoning, emotion regulation, and practical application. This gap creates operational and safety risks when organizations rely on automated systems to navigate nuanced human emotions.

The article introduces and validates a new evaluation benchmark, called EmoBench, to assess the emotional intelligence of large language models. The main objective is to measure how effectively these models understand emotional states and causes, and how well they apply that understanding to resolve emotional dilemmas.

The authors designed 400 hand-crafted, multiple-choice questions in both English and Chinese, rooted in established psychological frameworks. The evaluation is split across two core dimensions: emotional understanding (200 questions requiring models to identify both the emotion and its underlying cause across complex scenarios) and emotional application (200 questions testing the selection of optimal responses or actions in personal and social dilemmas). The study evaluated leading proprietary and open-source models across multiple prompt variations and compared their scores against a validated human baseline of 48 emotionally mature adults.

The findings reveal that current artificial intelligence systems exhibit a substantial emotional intelligence deficit compared to humans. First, the highest-performing model, GPT-4, achieved an overall accuracy of roughly 54% to 60% in emotional understanding and 74% to 76% in emotional application, remaining consistently below the average human performance and well below high-performing humans. Second, all evaluated models struggled severely with emotional understanding compared to application, particularly failing in perspective-taking tasks where smaller models often performed worse than a simple frequency baseline. Third, larger model scale strongly correlated with improved performance, but implementing step-by-step reasoning prompts provided minimal benefit and actively degraded the accuracy of smaller models under 14 billion parameters. Finally, models frequently defaulted to generic rules or surface-level patterns, misinterpreting implicit cues and ignoring situational context.

These results indicate that automated systems cannot yet reliably infer complex emotional states or manage delicate human interactions. Deploying current models in high-stakes environments without human oversight introduces significant risks of miscommunication, inappropriate advice, and damaged trust. The evidence also disproves the common assumption that standard reasoning techniques, such as step-by-step prompting, naturally enhance social and emotional judgment.

Organizations deploying artificial intelligence in user-facing and supportive roles should avoid full automation and maintain humans in the loop. Decision-makers should implement specialized evaluation frameworks that test implicit reasoning rather than relying on standard benchmarks. Before adopting language models for sensitive emotional workflows, teams should conduct controlled pilot programs and develop reasoning methods specifically tailored to social context.

The assessment is limited by its focus on text-only scenarios, a scope restricted to English and Chinese, and a relatively small sample size of 400 questions designed around shared cultural commonalities. While confidence in the benchmark's internal consistency and objective scoring is high—supported by strong human annotator agreement—stakeholders should exercise caution and not extrapolate these text-based findings to real-time multimodal or voice interactions.

Abstract

Recent advances in Large Language Models (LLMs) have highlighted the need for robust, comprehensive, and challenging benchmarks. Yet, research on evaluating their Emotional Intelligence (EI) is considerably limited. Existing benchmarks have two major shortcomings: first, they mainly focus on emotion recognition, neglecting essential EI capabilities such as emotion regulation and thought facilitation through emotion understanding; second, they are primarily constructed from existing datasets, which include frequent patterns, explicit information, and annotation errors, leading to unreliable evaluation. We propose EmoBench, a benchmark that draws upon established psychological theories and proposes a comprehensive definition for machine EI, including Emotional Understanding and Emotional Application. EmoBench includes a set of 400 hand-crafted questions in English and Chinese, which are meticulously designed to require thorough reasoning and understanding. Our findings reveal a considerable gap between the EI of existing LLMs and the average human, highlighting a promising direction for future research. Our code and data are publicly available at https://github.com/Sahandfer/EmoBench.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Definition of Emotional Intelligence
  • 2.2 Measures of Emotional Intelligence
  • 3 EMOBENCH
  • 3.1 Emotional Understanding
  • 3.2 Emotional Application
  • 4 Experiments
  • 4.1 Task Formulation
  • 4.2 Baselines
  • 4.3 Implementation Details
  • 5 Results and Findings
  • 6 Comparison with Human Performance
  • 7 Error Analysis
  • 8 Conclusion and Future Work
  • 9 Limitations
  • Ethical Considerations
  • Acknowledgements
  • References
  • A Scenario Taxonomy
  • A.1 Complex Emotions
  • A.2 Personal Beliefs and Experiences
  • A.3 Emotional Cues
  • A.4 Perspective Taking
  • B Experiment Prompts
  • C Emotion Distribution
  • D Human Evaluation
  • Emotional Understanding Guideline
  • Emotional Application Guideline

Knowls

  1. Knowl 1 — Emotional intelligence is evaluated as understanding plus application

    definition

    EMOBENCH defines machine emotional intelligence (EI) through two abilities. Emotional Understanding (EU) is the ability to perceive, identify, and monitor emotions, including reasoning about the causes of emotions and the perspectives of people in a situation. Emotional Application (EA) is the ability to use an understanding of emotional states to facilitate thought, manage emotions, and select an effective action or response in an emotional dilemma. The benchmark treats these as complementary dimensions: recognizing an emotion is not by itself evidence that a system can apply that understanding.

  2. Knowl 2 — EU scenarios test reasoning beyond stereotyped emotion patterns

    model/method

    EMOBENCH's Emotional Understanding scenarios are organized into four categories: complex emotions (emotion transitions, mixtures of emotions, and unexpected outcomes); personal beliefs and experiences (cultural values, sentimental value, and persona, including traits or prior experiences); emotional cues (visual and vocal cues); and perspective-taking (false beliefs, faux pas, and strange stories). Scenarios are designed so that common event–emotion associations are not sufficient: the system may need to infer unstated causes, track changing or mixed feelings, or distinguish what different people know and value. Many scenarios involve multiple people whose perspectives lead to different emotions about the same event.

  3. Knowl 3 — EA assesses whether emotional understanding guides a dilemma response

    model/method

    In Emotional Application, a system receives a scenario involving either a personal relationship (such as family, friends, or romantic partners) or a social relationship (such as coworkers, teachers, or a boss). The problem concerns either the subject answering the question (self) or another person (others); interpersonal conflicts are included under problems concerning others. The system chooses the most effective action (what to do) or response (what to say). The task is intended to assess whether the answer fits the people’s mental states and relationship, rather than merely selecting a generally plausible solution.

  4. Knowl 4 — The benchmark contains 400 hand-authored questions in two languages

    data/table

    EMOBENCH contains 400 multiple-choice questions (MCQs), available in English and Chinese: 200 assess Emotional Understanding and 200 assess Emotional Application. The EU questions were built from 121 manually crafted scenarios involving one to three people; the scenarios yield questions about individuals’ emotions and their causes. The 200 EA questions cover four relationship–problem combinations—personal-self, personal-others, social-self, and social-others—with 25 items in each combination. The authors used GPT-4-generated examples as inspiration for topic diversity, but report that they manually wrote the benchmark scenarios because generated examples tended to state emotions or causes explicitly and require little reasoning. Items were translated between English and Chinese and reviewed for content and quality.

  5. Knowl 5 — EU uses a scalable emotion-label ontology

    definition

    EMOBENCH bases its emotion labels on Plutchik’s wheel, using eight basic emotions—sadness, anger, joy, fear, anticipation, trust, disgust, and surprise—with intensity distinctions and combinations of basic emotions. The benchmark also incorporates labels from prior emotion taxonomies. Its mixed-emotion labels include guilt, pride, excitement or hopeful optimism, love/caring/gratitude, amusement/delight, disapproval/disappointment, sentimentality, jealousy, pessimism, remorse, hopelessness, embarrassment, nervousness, and curiosity. When a person is not experiencing an emotion, the labels distinguish being unbothered or indifferent from being oblivious, depending on the situation.

  6. Knowl 6 — EA answers are scored through distributed annotator preferences

    model/method

    Because several responses to an Emotional Application dilemma can be plausible, annotators assigned preference scores across the available choices rather than treating every alternative as wholly implausible. Each annotator distributed four units of 0.25, totaling 1, among the choices; for example, a preference for option A over a plausible option B could be recorded as 0.75 for A and 0.25 for B. Scores were averaged across annotators, and the choice with the highest resulting score defined the benchmark answer. Four workers participated in judging each item. Fleiss’ kappa was 0.852, which the authors report as excellent agreement.

  7. Knowl 7 — Model evaluation averages repeated, choice-order-varied MCQ runs

    experimental setup

    For EU, models answer an emotion question and then a corresponding cause question; for EA, they choose an action or response. The study compares zero-shot prompts with task instructions (Base) against prompts requesting chain-of-thought reasoning (CoT). Each model is prompted five times per MCQ, and the most frequent answer is selected. To test sensitivity to answer ordering, the choices are randomly reordered three times in addition to the original order; accuracy is averaged over the four choice orders. The evaluated systems include GPT-4, GPT-3.5, ChatGLM 3, Baichuan 2, Llama 2, Qwen, and Yi models, as well as random-choice and most-frequent-choice baselines.

  8. Knowl 8 — GPT-4 leads the benchmark but remains below human performance

    empirical result

    GPT-4 achieved the strongest overall LLM results reported for both tasks. Its Base accuracies were 59.75% on English EU, 54.12% on Chinese EU, 75.50% on English EA, and 73.75% on Chinese EA. With CoT prompting, its corresponding accuracies were 58.25%, 51.75%, 75.88%, and 73.50%. GPT-3.5 scored 33.12% and 26.38% on English and Chinese EU, and 61.38% and 55.75% on English and Chinese EA in the Base setting. In the human evaluation, participants outperformed the LLMs on both tasks; GPT-4 came closer to average human performance on EA but did not surpass the average human result. The human comparison used 48 participants, with 30 randomly selected benchmark questions per language–task group.

  9. Knowl 9 — EU is harder than EA, and CoT gains are inconsistent

    empirical result

    Across the evaluated systems, Emotional Understanding was considerably more difficult than Emotional Application. The authors note that EU requires correctly identifying both an emotion and its cause, and its scenarios more often depend on implications or unusual outcomes; they also caution that the EA scenarios remained more susceptible to familiar response patterns. CoT prompting generally produced little or no improvement and sometimes reduced smaller models’ accuracy, particularly for models below 14B parameters. Performance generally rose with model size. English results were often slightly higher than Chinese results, with Yi and ChatGLM-6B as exceptions; the authors do not claim that this difference is caused by training-data composition because they lack access to the models’ training data.

  10. Knowl 10 — Observed errors expose shallow perspective and relationship reasoning

    empirical result

    The authors’ qualitative analysis found that CoT outputs often evaluated the answer choices themselves instead of tracing events and changes in the people’s mental states. EU errors included unsupported assumptions about what a person knew, relying on a typical emotion association despite a scenario-specific exception, and failing to reason from a character’s perspective. In EA, models often favored broadly sensible responses without accounting for the relationship: for example, taking accountability may suit criticism, while gentle humor may be more appropriate for a friend’s teasing. CoT could also lead to off-topic discussion or refusal to select an answer; these behaviors were less common among larger models, particularly those above 50B parameters.

  11. Knowl 11 — Benchmark scope and evaluation have stated limitations

    limitation

    The authors identify several limits on what EMOBENCH establishes. Manual creation and supervision made the dataset small relative to benchmarks in other tasks, and it covers only English and Chinese and text rather than multimodal cues such as tone of voice or facial expression. GPT-4-generated examples informed topic selection, which could have introduced bias. The measured model performance depends on prompts that the authors do not claim are optimal; only CoT was tested as a reasoning intervention. EI and emotional responses are not fully objective, and the requirement that annotators agree on scenarios and answers restricted the topics and relationships represented. The benchmark also does not examine fine-grained personal traits that may alter emotional reactions.

Coverage note — The full per-model score matrices for every EU subcategory and EA relationship–problem combination are omitted; the benchmark’s principal aggregate comparisons and reported performance patterns are retained.

References

  1. 1.Mostafa M. Amin, Rui Mao, Erik Cambria, and Björn W. Schuller. 2023. A wide evaluation of chatgpt on affective computing tasks.
  2. 2.Neal M Ashkanasy and Catherine S Daus. 2005. Rumors of the death of emotional intelligence in organizational behavior are vastly exaggerated. Journal of Organizational Behavior, 26(4):441–452.
  3. 3.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  4. 4.Reuven Bar-On. 1997. BarOn emotional quotient inventory, volume 40. Multi-health systems.
  5. 5.Simon Baron-Cohen, Alan M Leslie, and Uta Frith. 1985. Does the autistic child have a “theory of mind”? Cognition, 21(1):37–46.
  6. 6.Simon Baron-Cohen, Michelle O’riordan, Valerie Stone, Rosie Jones, and Kate Plaisted. 1999. Recognition of faux pas by normally developing children and children with asperger syndrome or high-functioning autism. Journal of autism and developmental disorders, 29:407–418.
  7. 7.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  8. 8.Jeffrey M Conte. 2005. A review and critique of emotional intelligence measures. Journal of organizational behavior, 26(4):433–440.
  9. 9.Marzia Del Prete. 2021. Emotional artificial intelligence: detecting and managing customer emotions in automated customer service.
  10. 10.Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. Goemotions: A dataset of fine-grained emotions. arXiv preprint arXiv:2005.00547.
  11. 11.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335.
  12. 12.Murray J Dyck, Kara Ferguson, and Ian M Shochet. 2001. Do autism spectrum disorders differ from each other and from non-spectrum disorders on emotion recognition tests? European child & adolescent psychiatry, 10:105–116.
  13. 13.Paul Ekman. 1984. Expression and the nature of emotion. Approaches to emotion, 3(19):344.
  14. 14.Lisa Fan, Matthias Scheutz, Monika Lohani, Marissa McCoy, and Charlene Stokes. 2017. Do we need emotionally intelligent artificial agents? first results of human perceptions of emotional intelligence in humans compared to robots. In Intelligent Virtual Agents: 17th International Conference, IVA 2017, Stockholm, Sweden, August 27-30, 2017, Proceedings 17, pages 129–141. Springer.
  15. 15.Fiona J Ferguson and Elizabeth J Austin. 2010. Associations of trait and ability emotional intelligence with performance on theory of mind tasks in an adult sample. Personality and individual differences, 49(5):414–418.
  16. 16.Joseph L Fleiss and Jacob Cohen. 1973. The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability. Educational and psychological measurement, 33(3):613–619.
  17. 17.Deepanway Ghosal, Siqi Shen, Navonil Majumder, Rada Mihalcea, and Soujanya Poria. 2022. CICERO: A dataset for contextualized commonsense inference in dialogues. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5010–5028, Dublin, Ireland. Association for Computational Linguistics.
  18. 18.Daniel Goleman. 1996. Emotional intelligence. why it can matter more than iq. Learning, 24(6):49–50.
  19. 19.Francesca GE Happé. 1994. An advanced test of theory of mind: Understanding of story characters’ thoughts and feelings by able autistic, mentally handicapped, and normal children and adults. Journal of autism and Developmental disorders, 24(2):129–154.
  20. 20.Yinghui He, Yufan Wu, Yilin Jia, Rada Mihalcea, Yulong Chen, and Naihao Deng. 2023. Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models. arXiv preprint arXiv:2310.16755.
  21. 21.Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael R Lyu. 2023. Who is chatgpt? benchmarking llms’ psychological portrayal using psychobench. arXiv preprint arXiv:2310.01386.
  22. 22.Mete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva, and Antoine Bosselut. 2023. CRoW: Benchmarking commonsense reasoning in real-world tasks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9785–9821, Singapore. Association for Computational Linguistics.
  23. 23.Mirjana Ivanovic, Miloš Radovanović, Zoran Budimac, Dejan Mitrovic, Vladimir Kurbalija, Weihui Dai, and Weidong Zhao. 2014. Emotional intelligence and agents: Survey and possible applications. In Proceedings of the 4th International Conference on Web Intelligence, Mining and Semantics (WIMS14), pages 1–7.
  24. 24.Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, et al. 2021. Can machines learn morality? the delphi experiment. arXiv preprint arXiv:2110.07574.
  25. 25.Mridula C Jobson. 2020. Emotional maturity among adolescents and its importance. Indian Journal of Mental Health, 7(1):35–41.
  26. 26.Elke Kalbe, Marius Schlegel, Alexander T. Sack, Dennis A. Nowak, Manuel Dafotakis, Christopher Bangard, Matthias Brand, Simone Shamay-Tsoory, Oezguer A. Onur, and Josef Kessler. 2010. Dissociating cognitive from affective theory of mind: A tms study. Cortex, 46(6):769–780.
  27. 27.Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Bras, Gunhee Kim, Yejin Choi, and Maarten Sap. 2023. FANToM: A benchmark for stress-testing machine theory of mind in interactions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14397–14413, Singapore. Association for Computational Linguistics.
  28. 28.Cheng Li, Jindong Wang, Yixuan Zhang, Kaijie Zhu, Wenxin Hou, Jianxun Lian, Fang Luo, Qiang Yang, and Xing Xie. 2023. Large language models understand and can be enhanced by emotional stimuli. arXiv preprint arXiv:2305.15068.
  29. 29.Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. arXiv preprint arXiv:1710.03957.
  30. 30.Siyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li, Zhou Yu, Yong Jiang, and Minlie Huang. 2021. Towards emotional support dialog systems. arXiv preprint arXiv:2106.01144.
  31. 31.Paulo N Lopes, Marc A Brackett, John B Nezlek, Astrid Schütz, Ina Sellin, and Peter Salovey. 2004. Emotional intelligence and social interaction. Personality and social psychology bulletin, 30(8):1018–1034.
  32. 32.Xiaomeng Ma, Lingyu Gao, and Qihui Xu. 2023. Tom-challenges: A principle-guided dataset and diverse evaluation tasks for exploring theory of mind. arXiv preprint arXiv:2305.15068.
  33. 33.Carolyn MacCann and Richard D Roberts. 2008. New paradigms for assessing emotional intelligence: theory and data. Emotion, 8(4):540.
  34. 34.John D Mayer, David R Caruso, and Peter Salovey. 1999. Emotional intelligence meets traditional standards for an intelligence. Intelligence, 27(4):267–298.
  35. 35.John D Mayer, Peter Salovey, and David R Caruso. 2007. Mayer-salovey-caruso emotional intelligence test.
  36. 36.Daniela Mier, Stefanie Lis, Kerstin Neuthe, Carina Sauer, Christine Esslinger, Bernd Gallhofer, and Peter Kirsch. 2010. The involvement of emotion recognition in affective theory of mind. Psychophysiology.
  37. 37.Peter J O’Connor, Andrew Hill, Maria Kaya, and Brett Martin. 2019. The measurement of emotional intelligence: A critical review of the literature and recommendations for researchers and practitioners. Frontiers in psychology, 10:1116.
  38. 38.OpenAI. 2023. Gpt-4 technical report.
  39. 39.Samuel J Paech. 2023. Eq-bench: An emotional intelligence benchmark for large language models. arXiv preprint arXiv:2312.06281.
  40. 40.Rosalind W Picard. 2008. Toward machines with emotional intelligence.
  41. 41.Rosalind W. Picard, Elias Vyzas, and Jennifer Healey. 2001. Toward machine emotional intelligence: Analysis of affective physiological state. IEEE transactions on pattern analysis and machine intelligence, 23(10):1175–1191.
  42. 42.Robert Plutchik. 1982. A psychoevolutionary theory of emotions.
  43. 43.Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. MELD: A multimodal multi-party dataset for emotion recognition in conversations. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 527–536, Florence, Italy. Association for Computational Linguistics.
  44. 44.Soujanya Poria, Navonil Majumder, Devamanyu Hazarika, Deepanway Ghosal, Rishabh Bhardwaj, Samson Yu Bai Jian, Pengfei Hong, Romila Ghosh, Abhinaba Roy, Niyati Chhaya, et al. 2021. Recognizing emotion cause in conversations. Cognitive Computation, 13:1317–1332.
  45. 45.Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2018. Towards empathetic open-domain conversation models: A new benchmark and dataset. arXiv preprint arXiv:1811.00207.
  46. 46.Byron Reeves and Clifford Nass. 1996. The media equation: How people treat computers, television, and new media like real people. Cambridge, UK, 10(10).
  47. 47.Susan E Rivers, Isaac J Handley-Miner, John D Mayer, and David R Caruso. 2020. Emotional intelligence.
  48. 48.Sahand Sabour, Wen Zhang, Xiyao Xiao, Yuwei Zhang, Yinhe Zheng, Jiaxin Wen, Jialu Zhao, and Minlie Huang. 2022. Chatbots for mental health support: Exploring the impact of emohaa on reducing mental distress in china. arXiv preprint arXiv:2209.10183.
  49. 49.Peter Salovey and John D Mayer. 1990. Emotional intelligence. Imagination, cognition and personality, 9(3):185–211.
  50. 50.Maarten Sap, Ronan Le Bras, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A Smith, and Yejin Choi. 2019. Atomic: An atlas of machine commonsense for if-then reasoning. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 3027–3035.
  51. 51.Dagmar Schuller and Björn W Schuller. 2018. The age of artificial emotional intelligence. Computer, 51(9):38–46.
  52. 52.Nicola S Schutte, John M Malouff, Chad Bobik, Tracie D Coston, Cyndy Greeson, Christina Jedlicka, Emily Rhodes, and Greta Wendorf. 2001. Emotional intelligence and interpersonal relations. Journal of social psychology, 141(4):523–536.
  53. 53.Nicola S Schutte, John M Malouff, Maureen Simunek, Jamie McKenley, and Sharon Hollander. 2002. Characteristic emotional intelligence and emotional well-being. Cognition & Emotion, 16(6):769–785.
  54. 54.Candace L Sidner. 2016. Engagement, emotions, and relationships: on building intelligent agents. In Emotions, Technology, Design, and Learning, pages 273–294. Elsevier.
  55. 55.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  56. 56.Tomer Ullman. 2023. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399.
  57. 57.Xuena Wang, Xueting Li, Zi Yin, Yue Wu, and Jia Liu. 2023. Emotional intelligence of large language models. Journal of Pacific Rim Psychology, 17:18344909231213958.
  58. 58.Lynn Waterhouse. 2006. Multiple intelligences, the mozart effect, and emotional intelligence: A critical review. Educational psychologist, 41(4):207.
  59. 59.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  60. 60.Yige Xu, Zhiwei Zeng, and Zhiqi Shen. 2023. Efficient cross-task prompt tuning for few-shot conversational emotion recognition. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 11654–11666.
  61. 61.Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, et al. 2023a. Baichuan 2: Open large-scale language models. arXiv preprint arXiv:2309.10305.
  62. 62.Kailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie, and Sophia Ananiadou. 2023b. On the evaluations of chatgpt and emotion-enhanced prompting for mental health analysis. arXiv preprint arXiv:2304.03347.
  63. 63.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. 2022. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414.
  64. 64.Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2023. Safety-bench: Evaluating the safety of large language models with multiple choice questions. arXiv preprint arXiv:2309.07045.
  65. 65.Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2023. Large language models are not robust multiple choice selectors.
  66. 66.Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023. Agieval: A human-centric benchmark for evaluating foundation models. arXiv preprint arXiv:2304.06364.
  67. 67.Aojun Zhou, Ke Wang, Zimu Lu, Weikang Shi, Sichun Luo, Zipeng Qin, Shaoqing Lu, Anya Jia, Linqi Song, Mingjie Zhan, et al. 2023a. Solving challenging math word problems using gpt-4 code interpreter with code-based self-verification. arXiv preprint arXiv:2308.07921.
  68. 68.Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, et al. 2023b. Sotopia: Interactive evaluation for social intelligence in language agents. arXiv preprint arXiv:2310.11667.

Citation

MLA
Sabour, S., et al. “EmoBench: Evaluating the Emotional Intelligence of Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 5986–6004, https://doi.org/10.18653/v1/2024.acl-long.326.
APA
Sabour, S., Liu, S., Zhang, Z., Liu, J., Zhou, J., Sunaryo, A., Lee, T., Mihalcea, R., & Huang, M. (2024). EmoBench: Evaluating the Emotional Intelligence of Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5986–6004. https://doi.org/10.18653/v1/2024.acl-long.326
Chicago
Sabour, S., S. Liu, Z. Zhang, et al. 2024. “EmoBench: Evaluating the Emotional Intelligence of Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5986–6004. https://doi.org/10.18653/v1/2024.acl-long.326.
Harvard
Sabour, S. et al. (2024) “EmoBench: Evaluating the Emotional Intelligence of Large Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5986–6004. Available at: https://doi.org/10.18653/v1/2024.acl-long.326.
Vancouver
1. Sabour S, Liu S, Zhang Z, Liu J, Zhou J, Sunaryo A, Lee T, Mihalcea R, Huang M (2024) EmoBench: Evaluating the Emotional Intelligence of Large Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5986–6004

BibTeX

@inproceedings{sabour-etal-2024-emobench,
    title = "{E}mo{B}ench: Evaluating the Emotional Intelligence of Large Language Models",
    author = "Sabour, Sahand  and
      Liu, Siyang  and
      Zhang, Zheyuan  and
      Liu, June  and
      Zhou, Jinfeng  and
      Sunaryo, Alvionna  and
      Lee, Tatia  and
      Mihalcea, Rada  and
      Huang, Minlie",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.326/",
    doi = "10.18653/v1/2024.acl-long.326",
    pages = "5986--6004"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/