EmoBench: Evaluating the Emotional Intelligence of Large Language Models
Sahand SabourSiyang LiuZheyuan ZhangJune M. LiuJinfeng ZhouAlvionna S. SunaryoTatia M. C. LeeRada MihalceaMinlie Huang
Proposes EmoBench, a bilingual evaluation benchmark grounded in psychological theories that tests large language models on both emotional understanding and application, exposing a significant performance gap between leading models and human capabilities.
Large language models are increasingly deployed in sensitive, interpersonal domains such as customer service, education, and mental health support. However, existing benchmarks primarily evaluate emotional intelligence through simple pattern recognition and text extraction, ignoring complex reasoning, emotion regulation, and practical application. This gap creates operational and safety risks when organizations rely on automated systems to navigate nuanced human emotions.
The article introduces and validates a new evaluation benchmark, called EmoBench, to assess the emotional intelligence of large language models. The main objective is to measure how effectively these models understand emotional states and causes, and how well they apply that understanding to resolve emotional dilemmas.
The authors designed 400 hand-crafted, multiple-choice questions in both English and Chinese, rooted in established psychological frameworks. The evaluation is split across two core dimensions: emotional understanding (200 questions requiring models to identify both the emotion and its underlying cause across complex scenarios) and emotional application (200 questions testing the selection of optimal responses or actions in personal and social dilemmas). The study evaluated leading proprietary and open-source models across multiple prompt variations and compared their scores against a validated human baseline of 48 emotionally mature adults.
The findings reveal that current artificial intelligence systems exhibit a substantial emotional intelligence deficit compared to humans. First, the highest-performing model, GPT-4, achieved an overall accuracy of roughly 54% to 60% in emotional understanding and 74% to 76% in emotional application, remaining consistently below the average human performance and well below high-performing humans. Second, all evaluated models struggled severely with emotional understanding compared to application, particularly failing in perspective-taking tasks where smaller models often performed worse than a simple frequency baseline. Third, larger model scale strongly correlated with improved performance, but implementing step-by-step reasoning prompts provided minimal benefit and actively degraded the accuracy of smaller models under 14 billion parameters. Finally, models frequently defaulted to generic rules or surface-level patterns, misinterpreting implicit cues and ignoring situational context.
These results indicate that automated systems cannot yet reliably infer complex emotional states or manage delicate human interactions. Deploying current models in high-stakes environments without human oversight introduces significant risks of miscommunication, inappropriate advice, and damaged trust. The evidence also disproves the common assumption that standard reasoning techniques, such as step-by-step prompting, naturally enhance social and emotional judgment.
Organizations deploying artificial intelligence in user-facing and supportive roles should avoid full automation and maintain humans in the loop. Decision-makers should implement specialized evaluation frameworks that test implicit reasoning rather than relying on standard benchmarks. Before adopting language models for sensitive emotional workflows, teams should conduct controlled pilot programs and develop reasoning methods specifically tailored to social context.
The assessment is limited by its focus on text-only scenarios, a scope restricted to English and Chinese, and a relatively small sample size of 400 questions designed around shared cultural commonalities. While confidence in the benchmark's internal consistency and objective scoring is high—supported by strong human annotator agreement—stakeholders should exercise caution and not extrapolate these text-based findings to real-time multimodal or voice interactions.
No sufficiently relevant recommendations were found.
No sufficiently relevant recommendations were found.
