HealMe: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy

Mengxi XiaoQianqian XieZiyan KuangZhicheng LiuKailai YangMin PengWeiguang HanJimin Huang

article2024ACL75 citations

Presents HealMe, a conversational model that moves beyond simple sentence rewriting to guide clients through a structured, multi-step cognitive reframing process evaluated on real-world psychological metrics.

Listen

The growing demand for mental health support is often hindered by resource shortages, high therapist skill variability, and client hesitation rooted in shame or distrust. While artificial intelligence and large language models offer scalable potential, early attempts in cognitive reframing—a core cognitive-behavioral therapy technique—largely treated the process as simple text rewriting. This approach frequently failed to foster genuine client self-discovery and struggled to provide sustained empathy and purposeful clinical guidance.

To resolve these challenges, the article evaluates HealMe, a specialized conversational artificial intelligence model designed to guide clients through structured cognitive reframing rather than imposing top-down advice. The overarching objective is to demonstrate that a systematically trained model can empower clients to alter negative thinking patterns autonomously while maintaining high conversational quality and emotional resonance.

Researchers built a multi-turn dialogue dataset based on 1,000 thinking trap scenarios, prompting an advanced language model to simulate both client and therapist roles across a three-step therapeutic framework: distinguishing facts from thoughts, brainstorming alternative perspectives, and formulating empathetic, actionable conclusions. This dataset was used to fine-tune an open-source 7-billion parameter chat model over three training epochs. Evaluation encompassed two psychologist raters assessing 300 test cases for empathy, logical coherence, and guidance on a 0-to-3 scale, followed by a preliminary real-world pilot with human participants measured via the Positive and Negative Affect Schedule.

The findings confirm that HealMe substantially outperforms baseline models in therapeutic conversational efficacy. In simulated dialogues, HealMe achieved the highest ratings across all dimensions, scoring 2.500 in empathy, 2.650 in logical coherence, and 2.275 in guidance, resulting in an overall score of 2.125 compared to 1.750 for the baseline chat model and 1.675 for a bilingual comparator. In human trials, the model supported an average 41% reduction in negative emotion scores—markedly outperforming the 10% baseline drift in the control group—with several extreme negative emotional traits dropping from maximum severity to minimal levels.

These results indicate that structured guidance and prompt design can bridge the gap between superficial chat responses and authentic psychological intervention. By empowering users to generate their own alternative viewpoints, such models mitigate the risk of preachiness and support self-efficacy, making them strong candidates for text-based mental health support tools.

Organizations exploring automated mental health tools should consider adopting structured, multi-turn cognitive frameworks over single-turn rewriting models. However, broad deployment should follow expanded clinical trials with larger cohorts and more diverse psychological scenarios to address multi-issue complexities beyond rigid three-round dialogue constraints.

Confidence in the model's core conversational superiority is solid, backed by expert blind ratings and established psychological scales. Nevertheless, readers should interpret the human client results cautiously given the preliminary nature of the six-person pilot and the scope limits of the fixed dialogue structure.

No sufficiently relevant recommendations were found.

Cover for HealMe: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy

Abstract

Large Language Models (LLMs) can play a vital role in psychotherapy by adeptly handling the crucial task of cognitive reframing and overcoming challenges such as shame, distrust, therapist skill variability, and resource scarcity. Previous LLMs in cognitive reframing mainly converted negative emotions to positive ones, but these approaches have limited efficacy, often not promoting clients’ self-discovery of alternative perspectives. In this paper, we unveil the Helping and Empowering through Adaptive Language in Mental Enhancement (HealMe) model. This novel cognitive reframing therapy method effectively addresses deep-rooted negative thoughts and fosters rational, balanced perspectives. Diverging from traditional LLM methods, HealMe employs empathetic dialogue based on psychotherapeutic frameworks. It systematically guides clients through distinguishing circumstances from feelings, brainstorming alternative viewpoints, and developing empathetic, actionable suggestions. Moreover, we adopt the first comprehensive and expertly crafted psychological evaluation metrics, specifically designed to rigorously assess the performance of cognitive reframing, in both AI-simulated dialogues and real-world therapeutic conversations. Experimental results show that our model outperforms others in terms of empathy, guidance, and logical coherence, demonstrating its effectiveness and potential positive impact on psychotherapy.

Table of Contents

  • 1 Introduction
  • 2 Problem Definition and Goals
  • 3 Dataset Construction
  • 3.1 Step 1: Separating Emotions from Facts
  • 3.2 Step 2: Brainstorming
  • 3.3 Step 3: Empathetic Response
  • 4 Dataset Evaluation
  • 4.1 Evaluation of the AI Client
  • 4.2 Evaluation of the AI Therapist
  • 5 Training
  • 6 Experiments
  • 6.1 Experimental Settings
  • 6.2 Testing Therapy Models with an AI Client
  • 6.3 Analysis of Experimental Results
  • 6.4 Case Study
  • 6.5 A Supplementary Test with Real Person Clients
  • 7 Related Work
  • 8 Conclusion
  • 9 Limitations
  • 10 Ethical Considerations
  • 11 Acknowledgement
  • References
  • A Evaluation Metrics in AI-to-AI Conversations
  • B Case Study in AI-to-AI conversations
  • B.1 Case Study in Guidance
  • B.2 A Low-scoring Case of HealMe
  • C Evaluation Details in AI-to-Human Conversations
  • C.1 Pre-test Evaluation
  • C.2 Evaluation Metrics in Real Conversations
  • C.2.1 Testing Procedure
  • C.3 Clients PANAS Score changes
  • C.4 Client's feedback on baseline models

Knowls

  1. Knowl 1 — HealMe’s three-stage cognitive reframing dialogue

    model/method

    HealMe implements cognitive reframing as a guided dialogue intended to help clients discover alternative interpretations rather than receive a therapist-written replacement thought. First, the therapist acknowledges the client’s feelings and helps distinguish the situation from the client’s thoughts about it, checking that the distinction is accurate. Second, the therapist guides the client to brainstorm other interpretations or responses that fit the same situation and take the client’s thinking trap into account; the aim is to open up viewpoints, not to force a positive or perfect reframing. Third, the therapist acknowledges the client’s effort, responds empathetically to the specific situation, and combines any useful reframed thought with persuasion and actionable suggestions. If the client cannot generate alternatives, the therapist should still respond to the client’s difficulty with empathy.

  2. Knowl 2 — A psychological rubric for evaluating AI therapists

    definition

    The paper evaluates AI-therapist replies on empathy, logical coherence, and guidance, each scored from 0 to 3, and assigns an overall score from 0 to 3. For empathy, 0 means disregarding the client’s content and feelings; 1 means reflecting content without recognizing emotion; 2 means reflecting both content and emotion; and 3 means integrating the client’s signals into an effective response. For logical coherence, 0 indicates loss of focus or severe logical errors; 1 indicates weak coherence, insufficient use of client evidence, or unclear expression; 2 indicates a generally clear, consistent, evidence-based response with at most minor issues; and 3 indicates rigorous reasoning from ample evidence and clear premises, without contradictions. For guidance, 0 indicates nonspecific or impractical suggestions without clear goals or implementation; 1 indicates basic but limited guidance; 2 indicates targeted, detailed, feasible recommendations; and 3 indicates highly targeted, feasible guidance that accounts for real-world factors and the client’s future development. The overall categories are: 0 for poor performance with empathy and logical coherence at 1 or below; 1 for empathy and coherence of at least 2 but guidance of 1 or below; 2 for empathy and coherence of 3 with guidance of 2; and 3 when all three component scores are 3.

  3. Knowl 3 — Construction and quality control of the dialogue dataset

    experimental setup

    The training corpus begins with 1,000 manually selected, well-composed pairs of a thinking trap and a client thought from an existing cognitive-reframing dataset. ChatGPT (gpt-3.5-turbo-0125) simulates both the client and therapist to expand each case into a three-round dialogue. Human intervention and manual inspection are used at each round to keep the simulated client in role. Client outputs are checked for clear expression of the situation and emotions, role adherence, and compliance with therapist instructions; each criterion is binary, and outputs are revised until they satisfy all three. To include clients who cannot readily brainstorm alternatives, 20 training examples are selected for prompts instructing the simulated client to respond negatively and challenge the therapist. The resulting split contains 900 three-round training cases and 100 three-round validation cases. Testing uses 300 three-round cases from a separate collection of anonymized real-life scenarios.

  4. Knowl 4 — Fine-tuning recipe for HealMe

    model/method

    HealMe is produced by fine-tuning LLaMA2-7b-chat on the constructed therapist-side dialogues and prompts. Training runs for 3 epochs using AdamW, with a maximum learning rate of 3×10−43\times10^{-4} and a warm-up ratio of 1%. The best checkpoint is selected using the designated validation set. The reported training run took 2 hours, 12 minutes, and 44 seconds on four Nvidia GeForce RTX 3090 GPUs, each with 24 GB of memory.

  5. Knowl 5 — HealMe outperforms two baselines in AI-client dialogues

    empirical result

    In three-round conversations with a simulated ChatGPT client, two psychologists scored HealMe and two baseline models on empathy, logical coherence, guidance, and overall performance, using the paper’s 0–3 scale. The reported mean scores were: ChatGLM3-6b, 2.150 empathy, 2.075 logical coherence, 2.000 guidance, and 1.675 overall; LLaMA2-7b-chat, 2.325, 1.900, 1.925, and 1.750, respectively; and HealMe, 2.500, 2.650, 2.275, and 2.125, respectively. HealMe scored highest in all four categories in this evaluation.

  6. Knowl 6 — Exploratory evaluation with real-person clients

    experimental setup

    The real-person study involved six volunteers who interacted with the three models, with two clients assigned to each model, and two additional volunteers serving as a no-model control group. Participants were selected for short-term negative experiences lasting more than a day but less than a week, and the study describes them as having mild temperaments and similar ages, education, and life situations. Model assignments were anonymous. Model conversations were limited to three rounds; control participants stayed in the therapy room for 30 minutes between completing the two questionnaires. Participants completed a 20-item Positive and Negative Affect Schedule (PANAS) before and after the interaction or control interval. The questionnaire covers 10 positive and 10 negative affect items, each rated from 1 to 5. The authors describe this small-scale study as exploratory and state that its purpose was to gather feedback and examine possible emotional changes, not to compare numerical changes between models.

  7. Knowl 7 — Negative-affect changes in the small real-person study

    empirical result

    For each study group, the paper reports average negative-affect fluctuation, defined as the absolute change in negative-affect scores divided by the group’s total pre-intervention negative-affect score. The reported values were 44% for ChatGLM3-6b, 21% for LLaMA2-7b-chat, 41% for HealMe, and 10% for the control group. The standard deviations of negative-affect change scores were 0.77, 0.95, 1.23, and 0.55, respectively. The authors report that some HealMe clients showed reduced negative emotions and increased determination; they give examples of two clients describing helpful shifts in perspective. Because the study included only two clients per model and reports these descriptive quantities, the results are exploratory rather than a robust comparative estimate of treatment effects.

  8. Knowl 8 — Expert and GPT-4 assessment of the training dialogue quality

    data/table

    Two psychological experts manually scored 70 randomly sampled training dialogues, and GPT-4 was prompted with the expert-scored examples and rubric to score the training set. Mean scores on the 0–3 scale were as follows: manual assessment—empathy 2.255, logical coherence 2.613, guidance 1.985, overall 1.916; GPT-4 assessment—empathy 3, logical coherence 3, guidance 2.456, overall 2.460. The paper also reports average absolute score differences (Avg. Diff) and standard deviations (Std. Dev) for evaluator comparisons. Expert 1 versus expert 2: empathy 0.75 and 0.71, logical coherence 0.65 and 0.80, guidance 0.88 and 0.71, overall 0.96 and 0.75. GPT-4 versus expert 1: 0.23 and 0.51, 0.32 and 0.63, 0.61 and 0.62, 0.57 and 0.58, respectively. GPT-4 versus expert 2: 0.84 and 0.77, 0.62 and 0.76, 0.80 and 0.69, 0.80 and 0.69, respectively. The authors characterize the training dialogues as having strong empathy and logic, while noting that moderate guidance can be appropriate for clients who primarily want to express themselves.

  9. Knowl 9 — AI-client comparison protocol

    experimental setup

    The AI-to-AI comparison tests HealMe against ChatGLM3-6b and its base model, LLaMA2-7b-chat, on 300 cases from a public, anonymized set of real-life scenarios reviewed by psychology experts. ChatGPT simulates the client, and each model follows the same three-round client and therapist prompts used to construct the training dialogues. Models use their default parameter settings and receive a consistent prompt instructing them first to identify cognitive errors and then generate analysis text. Two psychologists score anonymized conversations presented in random order; the final scores are the average of their ratings.

  10. Knowl 10 — Scope limitations of the structured three-round interaction

    limitation

    HealMe’s three-round, step-by-step dialogue structure can limit flexibility when a client presents multiple concerns: the therapist may address only some of them, leaving feelings unresolved when the conversation ends. The prompts that support specific guidance also constrain adaptation to the client’s changing needs. The authors identify longer, more flexible dialogues and a broader range of psychotherapeutic strategies as directions for future work.

Coverage note — Detailed case-study transcripts and item-level PANAS records are omitted because they illustrate individual interactions and are too granular to add to the main method and aggregate findings captured here.

References

  1. 1.American Psychological Association et al. 2017. Ethical principles of psychologists and code of conduct (2002, amended effective june 1, 2010, and january 1, 2017).
  2. 2.John W Ayers, Adam Poliak, Mark Dredze, Eric C Leas, Zechariah Zhu, Jessica B Kelley, Dennis J Faix, Aaron M Goodman, Christopher A Longhurst, Michael Hogarth, et al. 2023. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA internal medicine.
  3. 3.Arthur C Bohart and Leslie S Greenberg. 1997. Empathy and psychotherapy: An introductory overview.
  4. 4.Ruth C Brown, Michael A Southam-Gerow, Bryce D McLeod, Emily B Wheat, Carrie B Tully, Steven P Reise, Philip C Kendall, and John R Weisz. 2018. The global therapist competence scale for youth psychosocial treatment: Development and initial validation. Journal of Clinical Psychology, 74(4):649–664.
  5. 5.David D Burns and Susan Nolen-Hoeksema. 1992. Therapeutic empathy and recovery from depression in cognitive-behavioral therapy: a structural equation model. Journal of consulting and clinical psychology, 60(3):441.
  6. 6.Linda L Carli. 1999. Cognitive reconstruction, hindsight, and reactions to victims and perpetrators. Personality and Social Psychology Bulletin, 25(8):966–979.
  7. 7.Siyuan Chen, Mengyue Wu, Kenny Q Zhu, Kunyao Lan, Zhiling Zhang, and Lyuchun Cui. 2023a. Llm-empowered chatbots for psychiatrist and patient simulation: Application and evaluation. arXiv preprint arXiv:2305.13614.
  8. 8.Zhiyu Chen, Yujie Lu, and William Yang Wang. 2023b. Empowering psychotherapy with large language models: Cognitive distortion detection through diagnosis of thought prompting. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 4295–4304. Association for Computational Linguistics.
  9. 9.John R Crawford and Julie D Henry. 2004. The positive and negative affect schedule (panas): Construct validity, measurement properties and normative data in a large non-clinical sample. British journal of clinical psychology, 43(3):245–265.
  10. 10.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. Glm: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 320–335.
  11. 11.David JA Edwards. 1989. Cognitive restructuring through guided imagery: Lessons from gestalt therapy. Comprehensive handbook of cognitive therapy, pages 283–297.
  12. 12.Robert Elliott, Arthur C Bohart, Jeanne C Watson, and David Murphy. 2018. Therapist empathy and client outcome: An updated meta-analysis. Psychotherapy, 55(4):399.
  13. 13.Stefan G Hofmann, David JA Dozois, Winfried Ed Rief, and Jasper AJ Smits. 2014. The Wiley handbook of cognitive behavioral therapy, Vols. 1-3. Wiley Blackwell.
  14. 14.Jen-tse Huang, Man Ho Lam, Eric John Li, Shujie Ren, Wenxuan Wang, Wenxiang Jiao, Zhaopeng Tu, and Michael R Lyu. 2023. Emotionally numb or empathetic? evaluating how llms feel using emotionbench. arXiv preprint arXiv:2308.03656.
  15. 15.Ian Andrew James, Rachel Morse, and Alan Howarth. 2010. The science and art of asking questions in cognitive therapy. Behavioural and Cognitive Psychotherapy, 38(1):83–93.
  16. 16.C Johnco, VM Wuthrich, and RM Rapee. 2014. The influence of cognitive flexibility on treatment outcome and cognitive restructuring skill acquisition during cognitive behavioural treatment for anxiety and depression in older adults: Results of a pilot study. Behaviour research and therapy, 57:55–64.
  17. 17.Andreas Larsson, Nic Hooper, Lisa A Osborne, Paul Bennett, and Louise McHugh. 2016. Using brief cognitive restructuring and cognitive defusion techniques to cope with negative thoughts. Behavior modification, 40(3):452–482.
  18. 18.Deborah Roth Ledley, Brian P Marx, and Richard G Heimberg. 2011. Making cognitive-behavioral therapy work: Clinical process for new practitioners. Guilford Press.
  19. 19.Ilya Loshchilov and Frank Hutter. 2018. Decoupled weight decay regularization. In International Conference on Learning Representations.
  20. 20.Mounica Maddela, Megan Ung, Jing Xu, Andrea Madotto, Heather Foran, and Y-Lan Boureau. 2023. Training models to generate, recognize, and reframe unhelpful thoughts. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 13641–13660. Association for Computational Linguistics.
  21. 21.Bryce D McLeod, Michael A Southam-Gerow, Adriana Rodríguez, Alexis M Quinoy, Cassidy C Arnold, Philip C Kendall, and John R Weisz. 2018. Development and initial psychometrics for a therapist competence instrument for cbt for youth anxiety. Journal of Clinical Child & Adolescent Psychology, 47(1):47–60.
  22. 22.Kate Muse, Freda McManus, Sarah Rakovshik, and Richard Thwaites. 2017. Development and psychometric evaluation of the assessment of core cbt skills (accs): An observation-based tool for assessing cognitive behavioral therapy competence. Psychological assessment, 29(5):542.
  23. 23.James P Robson Jr and Meredith Troutman-Jordan. 2014. A concept analysis of cognitive reframing. Journal of Theory Construction & Testing, 18(2).
  24. 24.Tulika Saha, Vaibhav Gakhreja, Anindya Sundar Das, Souhitya Chakraborty, and Sriparna Saha. 2022. Towards motivational and empathetic response generation in online mental health support. In Proceedings of the 45th international ACM SIGIR conference on research and development in information retrieval, pages 2650–2656.
  25. 25.Ashish Sharma, Inna W Lin, Adam S Miner, David C Atkins, and Tim Althoff. 2023a. Human–ai collaboration enables more empathic conversations in text-based peer-to-peer mental health support. Nature Machine Intelligence, 5(1):46–57.
  26. 26.Ashish Sharma, Kevin Rushton, Inna E. Lin, David Wadden, Khendra G. Lucas, Adam S. Miner, Theresa Nguyen, and Tim Althoff. 2023b. Cognitive reframing of negative thoughts through human-language model interaction. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 9977–10000. Association for Computational Linguistics.
  27. 27.Amy E Sickel, Jason D Seacat, and Nina A Nabors. 2014. Mental health stigma update: A review of consequences. Advances in Mental Health, 12(3):202–215.
  28. 28.Vera Sorin, Danna Brin, Yiftach Barash, Eli Konen, Alexander Charney, Girish Nadkarni, and Eyal Klang. 2023. Large language models (llms) and empathy-a systematic review. medRxiv, pages 2023–08.
  29. 29.Elizabeth Stade, Shannon Wiltsey Stirman, Lyle H Ungar, Cody L Boland, H Andrew Schwartz, David Bryce Yaden, João Sedoc, Robert DeRubeis, Robb Willer, et al. 2023. Large language models could change the future of behavioral healthcare: A proposal for responsible development and evaluation.
  30. 30.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  31. 31.Myrra Vernooij-Dassen, Irena Draskovic, Jenny McCleery, and Murna Downs. 2011. Cognitive reframing for carers of people with dementia. Cochrane database of systematic reviews, (11).
  32. 32.Caleb Ziems, Minzhi Li, Anthony Zhang, and Diyi Yang. 2022. Inducing positive perspectives with text reframing. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3682–3700.

Citation

MLA
Xiao, M., et al. “HealMe: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1707–25, https://doi.org/10.18653/v1/2024.acl-long.93.
APA
Xiao, M., Xie, Q., Kuang, Z., Liu, Z., Yang, K., Peng, M., Han, W., & Huang, J. (2024). HealMe: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1707–1725. https://doi.org/10.18653/v1/2024.acl-long.93
Chicago
Xiao, M., Q. Xie, Z. Kuang, et al. 2024. “HealMe: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1707–25. https://doi.org/10.18653/v1/2024.acl-long.93.
Harvard
Xiao, M. et al. (2024) “HealMe: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1707–1725. Available at: https://doi.org/10.18653/v1/2024.acl-long.93.
Vancouver
1. Xiao M, Xie Q, Kuang Z, Liu Z, Yang K, Peng M, Han W, Huang J (2024) HealMe: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1707–1725

BibTeX

@inproceedings{xiao-etal-2024-healme,
    title = "{H}eal{M}e: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy",
    author = "Xiao, Mengxi  and
      Xie, Qianqian  and
      Kuang, Ziyan  and
      Liu, Zhicheng  and
      Yang, Kailai  and
      Peng, Min  and
      Han, Weiguang  and
      Huang, Jimin",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.93/",
    doi = "10.18653/v1/2024.acl-long.93",
    pages = "1707--1725"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/