Cognitive Reframing of Negative Thoughts through Human-Language Model Interaction
Ashish SharmaKevin RushtonInna E. LinDavid WaddenKhendra G. LucasAdam S. MinerTheresa NguyenTim Althoff
Demonstrates how language models can assist in cognitive therapy by generating and controlling linguistic attributes of reframed thoughts, backed by a randomized field study of over 2,000 participants showing that users favor empathic and specific reframes over overly positive ones.
Cognitive reframing is an established psychological technique that helps individuals overcome distressing negative thoughts by replacing them with constructive, alternative perspectives. While highly effective, widespread access to this intervention is severely limited by clinician shortages, high financial costs, and social stigma. The article addresses whether language models can assist individuals by automatically generating relatable, helpful, and memorable reframed thoughts in real time.
The main objective of the article is to define a measurable framework of reframing attributes, develop an artificial intelligence method to generate and control these reframes, and evaluate which linguistic characteristics make reframed thoughts most effective for people experiencing negative thoughts.
To achieve this, the authors collaborated with clinical psychologists to establish a framework of seven linguistic attributes: addressing thinking traps, rationality, positivity, empathy, actionability, specificity, and readability. They created an expert-annotated dataset of 600 situations, thoughts, and reframes sourced from mental health practitioners. Using this data, they designed a retrieval-enhanced language model that generates reframes and modulates their specific attributes. They subsequently validated the approach through automated metrics, clinical expert ratings, and a month-long, randomized field study involving 2,067 consented participants on Mental Health America, a major national mental health platform.
The field study and expert evaluations revealed several key findings regarding what users value in cognitive reframing. Highly empathic reframes were preferred 55.7% more often than low-empathy reframes, and highly specific reframes were preferred 43.1% more often than generic ones. In contrast, reframes with high positivity were preferred 22.7% less often than those with lower positivity, indicating that overly optimistic messaging can alienate individuals facing emotional distress. Additionally, reframes grounded in rationality were 10.8% more relatable, while reframes that directly addressed thinking traps, offered actionable guidance, or provided situational specificity scored significantly higher in helpfulness and long-term memorability. In technical benchmarks, the retrieval-enhanced approach outperformed standard baseline models in both linguistic overlap metrics and practitioner ratings for helpfulness and relatability.
These findings suggest that artificial intelligence can provide valuable, scalable scaffolding for in-the-moment cognitive support if properly guided. Crucially, the results show that simply maximizing positive sentiment is counterproductive; effective reframing relies instead on empathy, logical soundness, and actionable steps. System safety proved high, with only 0.56% of suggestions flagged by users, none of which involved harmful or unsafe content.
Organizations developing digital mental health tools should design generative systems that prioritize empathy, specificity, and cognitive distortion resolution over superficial optimism. Before deploying these systems widely or across clinical workflows, stakeholders must conduct further research across diverse demographic groups, evaluate non-English languages, and measure longitudinal psychological outcomes beyond single-session interactions. High confidence in these findings is supported by the randomized field trial with over 2,000 real-world users, though caution is warranted regarding long-term clinical efficacy and generalizability outside English-speaking web audiences.
No sufficiently relevant recommendations were found.
No sufficiently relevant recommendations were found.
