SENSE-7: Taxonomy and Dataset for Measuring User Perceptions of Empathy in Sustained Human-AI Conversations
Jina SuhLindy LeErfan ShayeganiGonzalo RamosJudith AmoresDesmond C. OngMary CzerwinskiJavier Hernandez
Introduces the SENSE-7 taxonomy and dataset of real-world human-AI conversations with turn-by-turn user annotations, providing a practical foundation for measuring and modeling context-dependent empathy in conversational agents.
As artificial intelligence conversational agents are increasingly deployed in workplace and personal settings, digital empathy has become a central focus for user engagement and satisfaction. However, conventional engineering approaches often attempt to simulate internal, humanlike emotional states or rely on third-party crowdworker ratings, ignoring the subjective, dynamic, and relational nature of how users actually perceive empathy during real-world interactions.
The article aims to introduce a human-centered, seven-dimension behavioral taxonomy of digital empathy and evaluate how real users perceive these empathic behaviors across sustained, multi-turn human-AI conversations.
To conduct this evaluation, the researchers developed the SENSE-7 framework and gathered 695 naturalistic conversations from 109 information workers interacting with four distinct large language model configurations across informational, personal, and work-related tasks. Participants provided per-turn and post-conversation empathy ratings alongside psychometric profiles, contextual goals, and mood indicators. The authors then performed mixed-effects statistical modeling, qualitative feedback analyses, and an automated classification experiment on a fully anonymized 672-conversation subset using an advanced language model.
The analysis yielded several key findings: First, empathy perception is highly fragile; even a single poor turn substantially reduced overall conversation empathy ratings and significantly harmed user engagement, task success, and future adoption intent. Second, user expectations and individual traits heavily dictate desired empathy; participants tackling personal or work issues desired significantly higher empathy, whereas individuals with high cognitive reappraisal skills desired less empathy when handling personal or work problems. Third, users prioritized cognitive understanding and response appropriateness over superficial affective displays, frequently expressing frustration with unsolicited advice or generic bulleted lists. Finally, baseline automated classification using an advanced language model achieved an accuracy of 48.7% and exact-or-within-one-level accuracy of 96.0% (Spearman rank correlation of 0.369), demonstrating the feasibility of automated empathy measurement while highlighting remaining technical challenges.
These findings indicate that empathy in conversational systems should be treated as an interactional, context-sensitive behavior rather than a simulated internal state. For enterprise deployments, uncalibrated or generic emotional expressions risk appearing superficial, invalidating user feelings, and eroding trust. Conversely, systems that appropriately adapt their communication style to user context and task type can enhance engagement and user satisfaction.
Organizations developing or integrating conversational agents should implement dynamic empathy calibration that adjusts responses based on task context and explicit user preferences, allowing users to opt out of empathic styling when only factual assistance is required. AI architectures should incorporate conversational repair mechanisms to detect and recover from misaligned turns, while expanding memory mechanisms to preserve relational continuity across multi-turn exchanges.
These conclusions should be interpreted in light of certain limitations: the participant sample was restricted to information workers at a single technology company with generally positive attitudes toward AI, and findings were derived exclusively from text-based models. While confidence in the seven-dimension behavioral taxonomy and empirical findings is high within this operational setting, cautious validation is recommended before generalizing to non-technical demographics or high-stakes contexts such as specialized mental healthcare.
- Paper: Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation, Dongjin Kang et al. (2024). This study analyzes how large language models navigate multi-stage emotional support conversations and suffer from strategy bias, providing key background for evaluating observable empathic behaviors.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). This paper establishes evaluation frameworks for long-term conversational memory in AI agents, contextualizing how continuity failures disrupt sustained human-AI relationships.
- Paper: The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values, Hannah Kirk et al. (2023). This survey examines the challenges of learning from subjective human preferences and values, grounding SENSE-7's approach to user-perceived and individualized empathy.
- Paper: MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues, Ge Bai et al. (2024). This work establishes a fine-grained evaluation taxonomy for multi-turn conversational competencies in large language models, setting the methodological baseline for multi-turn dialogue assessment.
- Paper: How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation, Chia-Wei Liu et al. (2016). This foundational paper demonstrates the discordance between automated metrics and human judgments in dialogue systems, motivating human-grounded evaluation datasets.
- Paper: Personalizing Dialogue Agents: I have a dog, do you have pets too?, Saizheng Zhang et al. (2018). This landmark research introduces persona conditioning to maintain conversational consistency and user engagement over interactive dialogues.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). This paper investigates how reasoning models internally simulate conversational and socio-emotional roles, offering a mechanistic perspective on generating the adaptive empathic behaviors highlighted by SENSE-7.
- Paper: Metacognition in LLMs: Foundations, Progress, and Opportunities, Gabrielle Kaili-May Liu et al. (2026). This comprehensive volume surveys self-monitoring and cognitive regulation in language models, providing frameworks to build conversational agents that reliably adapt to user expectations.
- Paper: Agentic Reasoning for Large Language Models, Tianxin Wei et al. (2026). This work details architectures for memory, feedback, and interactive reasoning in autonomous agents, providing the engineering principles necessary to implement sustained, context-sensitive agentic behaviors.
