top of page

Supportive AI Can Still Reinforce Distress

  • Receptiviti Labs
  • Jul 29
  • 2 min read

Updated: Aug 4

Research suggests that evaluating supportive AI requires measuring more than empathy or perceived helpfulness.


Researchers at the University of Pennsylvania, Stony Brook University, and the University of Texas at Dallas (Chi et al., arXiv:2605.21569) found that LLM responses that appear supportive can also contain behaviors that reinforce a user’s distress.


In the paper "When Support Escalates Distress: Regulation and Escalation in LLM Responses to Venting and Advice-Seeking," researchers analyzed 178,800 Reddit posts from 14,040 users, comparing venting with advice-seeking. They found that venting language contained more absolutist language and emotional intensity, while advice-seeking language was more reflective and solution-oriented.


Using 9,000 GPT-5.3 responses, the authors evaluated two independent dimensions: Regulation (behaviors that help stabilize emotion) and Escalation (behaviors that reinforce or amplify distress). The dimensions were weakly correlated (r = .13); responses to venting became more regulating while also becoming more escalatory. The pattern resembles co-rumination, in which emotional validation co-occurs with processes that reinforce negative emotional states.


A friend persona produced the highest escalation, while a therapist persona reduced escalation while maintaining regulation, with no significant difference in how helpful raters found the responses.


Two expert psychologists, annotating 20 default-persona responses, agreed strongly with the framework's escalation scores (κ = 0.81). Lay raters showed only modest agreement (κ = 0.22-0.26). Warmth is easy to recognize, but emotional reinforcement is much harder.


Chiu et al.'s BOLT framework (arXiv:2401.00820) reached a compatible conclusion through a different methodology, finding that current LLMs often resemble lower-quality therapy on several behavioral dimensions.


The implications extend beyond mental health: The paper argues that evaluating supportive AI requires measuring the interaction between the user and the model, not just the model’s responses in isolation. It shows that the same model responds differently to different help-seeking styles, increasing both regulation and escalation in response to venting. Those interactional dynamics are not captured by empathy, helpfulness, or refusal-based safety metrics alone. As AI systems become longer-term conversational partners, understanding their effects on people becomes an increasingly important part of evaluating them.


If you're working on measuring human outcomes or interaction state in AI systems, we'd be interested in hearing about your work.

 
 

Subscribe to Field Notes

bottom of page