Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation

Dongjin KangSunghwan KimTaeyoon KwonSeungjun MoonHyunsouk ChoYoungjae YuDongha LeeJinyoung Yeo

article2024ACL65 citationsOutstanding Paper Award

Reveals that large language models struggle in emotional support conversations due to an inherent bias toward specific conversational strategies and establishes methods to reduce this bias using external assistance.

Abstract

Emotional Support Conversation (ESC) is a task aimed at alleviating individuals' emotional distress through daily conversation. Given its inherent complexity and non-intuitive nature, ESConv dataset incorporates support strategies to facilitate the generation of appropriate responses. Recently, despite the remarkable conversational ability of large language models (LLMs), previous studies have suggested that they often struggle with providing useful emotional support. Hence, this work initially analyzes the results of LLMs on ESConv, revealing challenges in selecting the correct strategy and a notable preference for a specific strategy. Motivated by these, we explore the impact of the inherent preference in LLMs on providing emotional support, and consequently, we observe that exhibiting high preference for specific strategies hinders effective emotional support, aggravating its robustness in predicting the appropriate strategy. Moreover, we conduct a methodological study to offer insights into the necessary approaches for LLMs to serve as proficient emotional supporters. Our findings emphasize that (1) low preference for specific strategies hinders the progress of emotional support, (2) external assistance helps reduce preference bias, and (3) existing LLMs alone cannot become good emotional supporters. These insights suggest promising avenues for future research to enhance the emotional intelligence of LLMs.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries & Related Work
  • 2.1 Emotional Support Conversation
  • 2.2 Incorporating Strategies into ESC Systems
  • 2.3 Emotional Support from LLMs
  • 3 Evaluation Setup
  • 3.1 Task and Focus
  • 3.2 Evaluation Set
  • 3.3 Metrics
  • 4 Proficiency and Preference of LLMs on Strategy
  • 4.1 Models & Implementation Details
  • 4.2 RQ1: Does the preference affect providing emotional support?
  • 5 Methodological Study: Mitigating Preference Bias
  • 5.1 Methods
  • 5.2 RQ2: How to mitigate the preference bias on LLMs?
  • 5.3 RQ3: Does improving preference bias help to become a better emotional supporter?
  • 6 Discussion and Conclusions
  • Limitations
  • Ethical Considerations
  • Acknowledgements
  • References
  • A Details of Preliminary Studies
  • A.1 Analysis of LLMs on ESC
  • A.2 Importance of Strategy
  • B ESConv Dataset
  • B.1 Definitions of Stages
  • B.2 Definitions of Strategies
  • C Experiments Details
  • C.1 Evaluation Sets
  • C.2 Preference Metric
  • C.3 Models
  • C.4 Prompts Details
  • C.5 Methods Details
  • D Implementation Details
  • E Details on Human Evaluation
  • E.1 Human Evaluation Criteria
  • E.2 Implementations of Human Evaluation
  • F Additional Analysis
  • F.1 LLMs' Proficiency for Each Strategy
  • F.2 Relation between Proficiency and Preference
  • F.3 Preference for Strategies by the Number of Examples.
  • F.4 Supervised Fine-tuning on ESC Task
  • G Case Study
  • G.1 Responses of LLMs by Stages
  • G.2 Comparison between Self-Contact and External-Contact
  • G.3 Misalignment between Strategy and Response

Knowls

  1. Knowl 1 — Strategy-Centric Two-Stage Formulation of Emotional Support Response Generation

    model/method

    Emotional support conversation (ESC) aims to alleviate a seeker's distress and guide them through personal challenges. In a strategy-centric framework, generating a supporter response is decomposed into two distinct probabilistic steps:

    1. Support Strategy Prediction: Given the seeker's background information II (e.g., emotion type, problem category, and situation summary) and the dialogue context history CC, the model predicts a conversational support strategy SS: S∼Pθ(⋅∣I,C)S \sim P_\theta(\cdot \mid I, C) where SS is chosen from a predefined taxonomy of 8 strategies: Question, Restatement or Paraphrasing, Reflection of Feelings, Self-disclosure, Affirmation and Reassurance, Providing Suggestions, Information, and Others.

    2. Strategy-Constrained Response Generation: Given II, CC, and the predicted strategy SS, the model generates the final conversational response RR: R∼Pθ(⋅∣I,C,S)R \sim P_\theta(\cdot \mid I, C, S)

    When augmented with an external strategy planner model θ′\theta', strategy selection is decoupled from the response generator θ\theta such that S^∼Pθ′(⋅∣I,C)\hat{S} \sim P_{\theta'}(\cdot \mid I, C) and R∼Pθ(⋅∣I,C,S^)R \sim P_\theta(\cdot \mid I, C, \hat{S}).

  2. Knowl 2 — Strategy Selection Metrics: Proficiency, Preference, and Preference Bias

    definition

    To analyze how language models select conversational support strategies across an NN-class strategy set (N=8N = 8), three quantitative metrics are defined:

    • Proficiency (qiq_i and QQ): The strategy-level proficiency qiq_i is the F1 score for predicting strategy ii. The overall proficiency across all strategies is the macro F1 score: Q=1N∑i=1NqiQ = \frac{1}{N} \sum_{i=1}^N q_i Stage-specific proficiency is measured using the weighted F1 score evaluated exclusively on samples belonging to that conversation stage.

    • Strategy Preference (pip_i): Quantifies how strongly a model tends to select strategy ii relative to ground-truth strategy jj. Estimated via Bradley-Terry modeling, the preference parameters pi>0p_i > 0 are normalized so that their sum equals NN: ∑i=1Npi=N\sum_{i=1}^N p_i = N Under this normalization, the average preference is pˉ=1\bar{p} = 1. A value pi>1p_i > 1 indicates an elevated preference for strategy ii, whereas pi<1p_i < 1 indicates dispreference.

    • Preference Bias (BB): The standard deviation of the strategy preferences pip_i across all NN strategies: B=∑i=1N(pi−pˉ)2NB = \sqrt{\frac{\sum_{i=1}^N (p_i - \bar{p})^2}{N}} A higher BB indicates severe disparity where a model excessively over-selects certain strategies while under-selecting others.

  3. Knowl 3 — Iterative Bradley-Terry Algorithm for Estimating Strategy Preferences

    algorithm

    Strategy preferences pip_i for each strategy i∈{1,…,N}i \in \{1, \dots, N\} are computed from the pairwise confusion matrix W=[wij]W = [w_{ij}], where wijw_{ij} is the frequency with which the model predicted strategy ii when the ground-truth strategy was jj. Under the Bradley-Terry model, the pairwise preference probability is P(i>j)=pipi+pjP(i > j) = \frac{p_i}{p_i + p_j}. The estimation procedure runs as follows:

    Input: Confusion matrix W∈RN×NW \in \mathbb{R}^{N \times N}, number of strategies N=8N=8, iterations K=20K=20
    Output: Normalized strategy preference vector p∈RNp \in \mathbb{R}^N
    Initialize pi←1p_i \leftarrow 1 for all i∈{1,…,N}i \in \{1, \dots, N\}
    for k=1k = 1 to KK do
        for i=1i = 1 to NN do
            pi′←∑j≠iwijpjpi+pj∑j≠iwjipi+pjp'_i \leftarrow \frac{\sum_{j \neq i} \frac{w_{ij} p_j}{p_i + p_j}}{\sum_{j \neq i} \frac{w_{ji}}{p_i + p_j}}
        end for
        G←(∏j=1Npj′)1/NG \leftarrow \left( \prod_{j=1}^N p'_j \right)^{1/N}
        for i=1i = 1 to NN do
            pi←pi′Gp_i \leftarrow \frac{p'_i}{G}
        end for
    end for
    S←∑i=1NpiS \leftarrow \sum_{i=1}^N p_i
    for i=1i = 1 to NN do
        pi←pi⋅NSp_i \leftarrow \frac{p_i \cdot N}{S}
    end for
    return pp
  4. Knowl 4 — Stage-Partitioned Benchmark Construction on ESConv

    experimental setup

    To evaluate model performance across conversational phases grounded in Hill's Helping Skills Theory, the 1,300 dialogues of the ESConv dataset are partitioned into three distinct stage-specific test sets (D1,D2,D3D_1, D_2, D_3) without dialogue overlap:

    • Exploration (D1D_1): 549 samples from 433 dialogues (average 9.95 turns, 16.27 words per utterance), focused on identifying and clarifying the seeker's underlying issues.
    • Comforting (D2D_2): 524 samples from 434 dialogues (average 10.04 turns, 16.81 words per utterance), focused on expressing empathy, validation, and emotional reassurance.
    • Action (D3D_3): 816 samples from 433 dialogues (average 10.66 turns, 18.92 words per utterance), focused on facilitating problem-solving, goal-setting, and actionable suggestions.

    Samples are created by slicing dialogues into 5–15 turns. The stage label for each target utterance is assigned based on the majority stage of surrounding strategies within a window size of 4. Instances whose assigned stage conflicts with the test set partition are removed, and slices are constrained such that the proportion of the uninformative strategy Others does not exceed 5%.

  5. Knowl 5 — Strategy Preference Skew and Stage-Specific Performance Disparities in LLMs

    empirical result

    Standard LLMs exhibit severe strategy preference biases that directly correlate with stage-specific failure modes:

    • Skewed Selection: GPT-4 and ChatGPT allocate the majority of their predictions to Affirmation and Reassurance (ChatGPT: 64.0% selection ratio, preference pi=4.49p_i = 4.49; GPT-4: 60.0% ratio, pi=4.26p_i = 4.26), while severely neglecting exploration strategies such as Question (ChatGPT: 1.4%, pi=0.12p_i = 0.12; GPT-4: 1.4%, pi=0.11p_i = 0.11) and Restatement or Paraphrasing (ChatGPT: 2.2%, pi=0.27p_i = 0.27; GPT-4: 0.0%, pi=0.00p_i = 0.00).

    • Stage Performance Gap: Because GPT-4 and ChatGPT strongly disprefer exploration strategies, they perform poorly on Exploration test set D1D_1 (weighted F1: GPT-4 2-shot 14.61, ChatGPT 2-shot 15.16) compared to Comforting D2D_2 (GPT-4 22.55, ChatGPT 19.07) and Action D3D_3 (GPT-4 24.68, ChatGPT 20.10).

    • Correlation between Preference and Proficiency: Across diverse LLMs, Pearson correlation between strategy preference pip_i and strategy proficiency qiq_i is strongly positive: Mistral-7B (0.943), Vicuna-13B (0.935), Tulu-70B (0.899), GPT-4 (0.820), LLaMA2-70B (0.772), ChatGPT (0.752), Solar-10.7B (0.747), and LLaMA2-7B (0.600).

  6. Knowl 6 — Self-Contact versus External-Contact Approaches for Mitigating Preference Bias

    empirical result

    Testing adaptation methods on ChatGPT and LLaMA2-70B shows that self-contained prompting exacerbates bias, whereas external knowledge and assistance reduce bias:

    • Self-Contact Degradation: Relying strictly on the model's internal reasoning—Direct-Refine, Self-Refine, and Emotional-CoT—worsens preference bias BB and reduces macro F1 proficiency QQ. On zero-shot ChatGPT (baseline Q=13.50,B=1.38Q = 13.50, B = 1.38), Direct-Refine yields Q=13.40,B=1.60Q = 13.40, B = 1.60; Self-Refine yields Q=12.37,B=1.53Q = 12.37, B = 1.53; and Emotional-CoT yields Q=9.55,B=1.56Q = 9.55, B = 1.56. Iterative refinement further amplifies the selection of initially preferred strategies (pi>1p_i > 1) and reduces dispreferred strategies (pi<1p_i < 1).

    • External-Contact Mitigation: Providing external knowledge or specialized modular assistance alleviates preference bias and improves proficiency. For ChatGPT, adding COMET commonsense reduces BB to 0.95 (Q=12.78Q = 12.78), expanding few-shot examples to 4-shot reduces BB to 0.82 (Q=16.91Q = 16.91), and decoupling strategy selection to a fine-tuned Strategy Planner reduces BB to 0.36 while raising QQ to 21.09.

    • LLaMA2-70B Behavior: For 2-shot LLaMA2-70B (baseline Q=14.55,B=0.47Q = 14.55, B = 0.47), Self-Refine degrades metrics to Q=13.15,B=0.55Q = 13.15, B = 0.55, whereas the fine-tuned Strategy Planner achieves Q=21.09,B=0.36Q = 21.09, B = 0.36 and weighted F1 scores of 22.59 on D1D_1, 21.85 on D2D_2, and 23.77 on D3D_3.

  7. Knowl 7 — Architectural Comparison of Fine-Tuned Strategy Planners

    model/method

    A Strategy Planner is trained to predict the next conversational support strategy S^∈{1,…,8}\hat{S} \in \{1, \dots, 8\} conditioned on dialogue context CC and background II. Autoregressive LLM backbones trained with QLoRA (4-bit quantization, rank r=64r = 64, α=16\alpha = 16, learning rate 5×10−55 \times 10^{-5} over 5 epochs) outperform encoder classification models in mitigating preference bias:

    Base Model Q↑Q \uparrow B↓B \downarrow D1D_1 F1 D2D_2 F1 D3D_3 F1
    BERT 18.02 0.50 18.17 22.68 19.25
    RoBERTa 21.01 0.60 21.34 24.18 22.99
    Mistral-7B 21.89 0.45 22.61 23.57 24.59
    LLaMA2-7B 21.10 0.36 22.59 21.85 23.77

    While encoder backbones (BERT, RoBERTa) achieve comparable overall F1, their higher preference bias (B=0.50–0.60B = 0.50\text{--}0.60) produces larger performance gaps between Exploration (D1D_1) and Comforting (D2D_2). LLaMA2-7B achieves the lowest preference bias (B=0.36B = 0.36) and the most uniform cross-stage F1 scores.

  8. Knowl 8 — Multidimensional Psychological Human Evaluation Framework for Emotional Support

    definition

    To assess emotional support responses beyond standard reference-based n-gram metrics, a four-dimensional evaluation framework was developed in collaboration with psychologists:

    1. Seeker's Satisfaction (Sat.): An overarching construct comprising three fine-grained sub-criteria evaluated from the seeker's perspective:

      • Acceptance: Whether the seeker can receive and accept the response without emotional resistance, irritation, or discomfort.
      • Effectiveness: Whether the response actively helps mitigate negative emotional intensity and shift the seeker's mindset toward a constructive direction.
      • Sensitivity: Whether the response accurately perceives, validates, and adapts to the seeker's holistic state (mood, immediate needs, emotional resources, and life context).
    2. Strategy Alignment:

      • Alignment: Whether the linguistic structure and conversational intent of the generated response faithfully adhere to the target or planned support strategy.

    Evaluation is conducted using both pairwise comparative selection (Win / Tie / Lose) and absolute 5-point Likert scales with specific behavioral rubrics.

  9. Knowl 9 — Human Evaluation: Impact of Preference Bias on Emotional Support Quality and Failure Rate

    empirical result

    Human evaluation on 100 diverse multi-turn dialogue samples demonstrates that mitigating strategy preference bias directly reduces harmful/low-quality responses (defined as Likert score <3< 3 on Seeker's Satisfaction):

    • Failure Rate Reductions:

      • Vanilla ChatGPT: 16.7% failure rate (83.3% acceptable ≥3\ge 3).
      • ChatGPT + Direct-Refine (aggravated bias): 21.2% failure rate (78.8% acceptable).
      • ChatGPT + Self-Refine (aggravated bias): 17.4% failure rate (82.6% acceptable).
      • ChatGPT + Strategy Planner (mitigated bias): 8.0% failure rate (92.0% acceptable).
      • ChatGPT + Ground-Truth Oracle Strategy: 3.8% failure rate (96.2% acceptable).
    • Pairwise Win Rates against Vanilla ChatGPT:

      • Self-Refine achieves a 50.5% win rate on Satisfaction (44.1% Effectiveness, 55.9% Sensitivity).
      • COMET augmentation achieves a 52.1% win rate on Satisfaction (42.7% Effectiveness, 58.3% Sensitivity).
      • Example Expansion (4-shot) achieves a 57.2% win rate on Satisfaction (48.5% Effectiveness, 62.6% Sensitivity).
      • Strategy Planner achieves a 61.1% win rate on Satisfaction (70.8% Acceptance, 54.2% Effectiveness, 58.3% Sensitivity).
  10. Knowl 10 — Limitations of Strategy Prediction and Generation in LLM Emotional Supporters

    limitation

    LLM-based emotional support conversation systems face several inherent operational limitations:

    1. Residual Failure Under Oracle Strategies: Even when supplied with perfect ground-truth support strategies, LLMs produce unhelpful or distress-intensifying responses in 3.8% of cases, demonstrating that accurate strategy selection alone does not guarantee high-quality generation.
    2. Knowledge-Strategy Misalignment: When provided with an external strategy that conflicts with the LLM's internal next-token preference, models can experience knowledge conflict and generate responses misaligned with the intended strategy (e.g., producing reassurance instead of providing requested factual information).
    3. Broad Catch-All Category Ambiguity: The Others strategy category absorbs diverse conversational utterances, obscuring finer-grained model preferences when few-shot examples scale up (n>8n > 8).
    4. Prompt Formatting Artifacts: Open-source foundation models frequently fail to output valid strategy names in zero-shot settings, requiring 2-shot exemplars that may artificially improve measured baseline proficiency over pure zero-shot capabilities.

Coverage note — None was omitted; all primary contributions—including strategy-centric formulation, preference bias metrics and estimation algorithm, stage benchmark construction, empirical findings on self/external-contact approaches, strategy planner ablations, psychological human evaluation criteria, and empirical failure analyses—are captured in the knowls.

References

  1. 1.Thomas Allport, Pettigrew, Kerstin Hammann, and S Salzborn. 1954. Gordon willard allport: The nature of prejudice. Samuel Salzborn (Hg.): Klassiker der Sozialwissenschaften, 100:193–197.
  2. 2.Satanjeev Banerjee and Alon Lavie. 2005. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In IEEvaluation@ACL.
  3. 3.Ralph Allan Bradley and Milton E Terry. 1952. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345.
  4. 4.Brant R Burleson. 2003. Emotional support skill. In Handbook of Communication and Social Interaction Skills, page 551. Psychology Press.
  5. 5.Hyungjoo Chae, Yongho Song, Kai Ong, Taeyoon Kwon, Minjin Kim, Youngjae Yu, Dongha Lee, Dongyeop Kang, and Jinyoung Yeo. 2023. Dialogue chain-of-thought distillation for commonsense-aware conversational agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5606–5632, Singapore. Association for Computational Linguistics.
  6. 6.Maximillian Chen, Xiao Yu, Weiyan Shi, Urvi Awasthi, and Zhou Yu. 2023a. Controllable mixed-initiative dialogue generation through prompting. In Annual Meeting of the Association for Computational Linguistics.
  7. 7.Yirong Chen, Xiaofen Xing, Jingkai Lin, Huimin Zheng, Zhenyu Wang, Qi Liu, and Xiangmin Xu. 2023b. Soulchat: Improving llms’ empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations.
  8. 8.Jiale Cheng, Sahand Sabour, Hao Sun, Zhuang Chen, and Minlie Huang. 2023. Pal: Persona-augmented emotional support conversation generation. In ACL.
  9. 9.Yi Cheng, Wenge Liu, Wenjie Li, Jiashuo Wang, Ruihui Zhao, Bang Liu, Xiaodan Liang, and Yefeng Zheng. 2022. Improving multi-turn emotional support dialogue generation with lookahead strategy planning. In Conference on Empirical Methods in Natural Language Processing.
  10. 10.Neo Christopher Chung, George Dyer, and Lennart Brocki. 2023. Challenges of large language models for mental health counseling. arXiv preprint arXiv:2311.13857.
  11. 11.Yang Deng, Wenxuan Zhang, Yifei Yuan, and Wai Lam. 2023. Knowledge-enhanced mixed-initiative dialogue system for emotional support conversations. In ACL.
  12. 12.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT.
  14. 14.Mahshid Eshghie and Mojtaba Eshghie. 2023. Chatgpt as a therapist assistant: A suitability study. arXiv preprint arXiv:2304.09873.
  15. 15.Faiza Farhat. 2023. Chatgpt as a complementary mental health resource: a boon or a bane. Annals of Biomedical Engineering, pages 1–4.
  16. 16.Luke Friedman, Sameer Ahuja, David Allen, Zhenning Tan, Hakim Sidahmed, Changbo Long, Jun Xie, Gabriel Schubiner, Ajay Patel, Harsh Lara, Brian Chu, Zexi Chen, and Manoj Tiwari. 2023. Leveraging large language models in conversational recommender systems.
  17. 17.Jun Gao, Wei Bi, Ruifeng Xu, and Shuming Shi. 2022a. Ream♯: An enhancement approach to reference-based evaluation metrics for open-domain dialog generation.
  18. 18.Silin Gao, Jena D. Hwang, Saya Kanno, Hiromi Wakaki, Yuki Mitsufuji, and Antoine Bosselut. 2022b. ComFact: A benchmark for linking contextual commonsense knowledge. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 1656–1675, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  19. 19.Jennifer C Greene. 2003. Handbook of Communication and Social Interaction Skills. Psychology Press.
  20. 20.Catherine A Heaney and Barbara A Israel. 2008. Social networks and social support. 4:189–210.
  21. 21.Clara E Hill. 2009. Helping Skills: Facilitating, Exploration, Insight, and Action. American Psychological Association.
  22. 22.Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi. 2020. Comet-atomic 2020: On symbolic and neural commonsense knowledge graphs. In AAAI Conference on Artificial Intelligence.
  23. 23.Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi. 2023. Camels in a changing climate: Enhancing lm adaptation with tulu 2.
  24. 24.Shaoxiong Ji, Tianlin Zhang, Kailai Yang, Sophia Ananiadou, and Erik Cambria. 2023. Rethinking large language models in mental health applications.
  25. 25.Mengzhao Jia, Qianglong Chen, Liqiang Jing, Dawei Fu, and Renyu Li. 2023. Knowledge-enhanced memory model for emotional support conversation. arXiv preprint arXiv:2310.07700.
  26. 26.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b.
  27. 27.Dahyun Kim, Chanjun Park, Sanghoon Kim, Wonsung Lee, Wonho Song, Yunsu Kim, Hyeonwoo Kim, Yungi Kim, Hyeonju Lee, Jihoo Kim, Changbae Ahn, Seonghoon Yang, Sukyung Lee, Hyunbyung Park, Gyoungjin Gim, Mikyoung Cha, Hwalsuk Lee, and Sunghun Kim. 2023. Solar 10.7b: Scaling large language models with simple yet effective depth upscaling.
  28. 28.Catherine Penny Hinson Langford, Juanita Bowsher, Joseph P Maloney, and Patricia P Lillis. 1997. Social support: A conceptual analysis. Journal of Advanced Nursing, 25(1):95–100.
  29. 29.Gibbeum Lee, Volker Hartmann, Jongho Park, Dimitris Papailiopoulos, and Kangwook Lee. 2023. Prompted LLMs as chatbot modules for long open-domain conversation. In Findings of the Association for Computational Linguistics: ACL 2023, pages 4536–4554, Toronto, Canada. Association for Computational Linguistics.
  30. 30.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and William B. Dolan. 2016. A diversity-promoting objective function for neural conversation models. In NAACL.
  31. 31.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Annual Meeting of the Association for Computational Linguistics.
  32. 32.Siyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li, Zhou Yu, Yong Jiang, and Minlie Huang. 2021. Towards emotional support dialog systems. In ACL.
  33. 33.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  34. 34.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Sean Welleck, Bodhisattwa Prasad Majumder, Shashank Gupta, Amir Yazdanbakhsh, and Peter Clark. 2023. Self-refine: Iterative refinement with self-feedback. ArXiv, abs/2303.17651.
  35. 35.Shikib Mehri and Maxine Eskenazi. 2020. Usr: An unsupervised and reference free evaluation metric for dialog generation.
  36. 36.M. E. J. Newman. 2023. Efficient computation of rankings from pairwise comparisons. Journal of Machine Learning Research, 24(238):1–25.
  37. 37.OpenAI. 2023a. Chatgpt. https://openai.com/blog/chatgpt.
  38. 38.OpenAI. 2023b. Gpt-4 technical report.
  39. 39.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Annual Meeting of the Association for Computational Linguistics.
  40. 40.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model.
  41. 41.Inhwa Song, Sachin R. Pendse, Neha Kumar, and Munmun De Choudhury. 2024. The typing cure: Experiences with large language model chatbots for mental health support.
  42. 42.Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Cantón Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. ArXiv, abs/2307.09288.
  43. 43.Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2014. Cider: Consensus-based image description evaluation. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4566–4575.
  44. 44.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  45. 45.Ernst Zermelo. 1929. Die berechnung der turnierergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung. Mathematische Zeitschrift, 29(1):436–460.
  46. 46.Weixiang Zhao, Yanyan Zhao, Xin Lu, Shilong Wang, Yanpeng Tong, and Bing Qin. 2023a. Is chatgpt equipped with emotional dialogue capabilities? arXiv preprint arXiv:2304.09582.
  47. 47.Weixiang Zhao, Yanyan Zhao, Shilong Wang, and Bing Qin. 2023b. Transesc: Smoothing emotional support conversation via turn-level state transition. In Annual Meeting of the Association for Computational Linguistics.
  48. 48.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023a. Judging llm-as-a-judge with mt-bench and chatbot arena.
  49. 49.Zhonghua Zheng, Lizi Liao, Yang Deng, and Liqiang Nie. 2023b. Building emotional support chatbots in the era of llms. ArXiv, abs/2308.11584.

Citation

MLA
Kang, D., et al. “Can Large Language Models Be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 15232–61, https://doi.org/10.18653/v1/2024.acl-long.813.
APA
Kang, D., Kim, S. M., Kwon, T., Moon, S., Cho, H., Yu, Y., Lee, D., & Yeo, J. (2024). Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15232–15261. https://doi.org/10.18653/v1/2024.acl-long.813
Chicago
Kang, D., S. M. Kim, T. Kwon, et al. 2024. “Can Large Language Models Be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15232–61. https://doi.org/10.18653/v1/2024.acl-long.813.
Harvard
Kang, D. et al. (2024) “Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 15232–15261. Available at: https://doi.org/10.18653/v1/2024.acl-long.813.
Vancouver
1. Kang D, Kim SM, Kwon T, Moon S, Cho H, Yu Y, Lee D, Yeo J (2024) Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 15232–15261

BibTeX

@inproceedings{kang-etal-2024-large,
    title = "Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation",
    author = "Kang, Dongjin  and
      Kim, Sunghwan  and
      Kwon, Taeyoon  and
      Moon, Seungjun  and
      Cho, Hyunsouk  and
      Yu, Youngjae  and
      Lee, Dongha  and
      Yeo, Jinyoung",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.813/",
    doi = "10.18653/v1/2024.acl-long.813",
    pages = "15232--15261"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/