top of page

AI Can Cause Harm: The Case for Psycholinguistic AI Safety

  • Receptiviti Labs
  • Jul 17
  • 3 min read

In a recent Nature Correspondence, Neuroscience Researcher Ziv Ben-Zion argues that emotionally responsive AI has crossed a line because they build dependency, reinforce harmful beliefs, and manufacture a sense of connection that feels real to the person using them. He makes the case that the voluntary safeguards deployed by AI companies are no longer sufficient, and he proposes four mandatory guardrails.


  1. Continuous disclosure that the system is not human, in both language and interface, with reminders during emotionally intense exchanges that it is not a therapist.


  2. Flagging language in the user's prompts that signals psychological distress — hopelessness, withdrawal, agitation — and pausing to offer crisis resources or escalate to human support.


  3. Clear conversational boundaries, so systems avoid simulating romantic intimacy or engaging on suicide, death, and metaphysics.


  4. Involvement of clinicians, ethicists, and human–AI interaction specialists, with regular audits for unsafe behavior.



Guardrails 1, 3, and 4 are policy decisions that can be implemented by introducing a disclosure rule, limiting the model from discussing certain topics, and by convening an advisory board.


The second guardrail Ben-Zion proposes is directionally correct, but insufficient as described. What's missing is psycholinguistic AI safety: a way to identify distress in how a user's language changes, not just the content of one message.


AI safety inherited much of its evaluation framework from social media moderation, where harms were primarily associated with individual pieces of content. But in conversational AI, the harm signals are not as overt. Many of the most significant risks don't appear in individual prompts or in the language that safety researchers often assume they do. Recent research into the relationship between language and self-harm shows that the earliest and most important signals can't be identified using prompt-by-prompt flagging.


In a npj Mental Health Research study, Charlotte Entwistle, Katie Hoemann, Sophie Nightingale, and Receptiviti's Ryan Boyd analyzed the language of 992 people who identified as having borderline personality disorder, a group with high rates of self-harm. They studied which language patterns track self-harm, and how that language changes in the weeks around suicidality and non-suicidal self-injury (NSSI) events.


Two findings are key for AI-user distress detection.


  1. Self-harm has steady linguistic correlates, and most of them are words no keyword filter would flag, specifically increased use of self-focused language, more negative emotion, sadness, anger, swearing, and more use of absolutist language including words like "always" and "never." In other words, a user who is manifesting suicidal thoughts may never express them overtly, so detection requires a far more nuanced approach.


  1. The markers move on a schedule:


Markers of suicidality: Anxiety language climbed sharply two to three weeks before suicidality and stayed high into the final week, while sadness and swearing rose in that last week, and social language dropped just before.


Markers of NSSI events: Sadness rose two weeks prior, then dropped significantly in the days right before the event, which the authors read as a stretch of numbness or dissociation, a known precursor to self-injury.


The schedule is critical because distress follows an arc that starts before any individual message would be flagged. Seeing that arc means measuring these markers turn after turn, not scanning one message at a time. That's the case for evaluating AI safety at the level of a user's trajectory, not just the individual prompt.


AI user safety must be evaluated at the level of user trajectories.

Because of this, AI user safety must not only be evaluated at the individual message level, but also at the level of user trajectories.


In practice this means tracking a set of validated language markers as the conversation progresses, comparing them against the person's own earlier messages, and surfacing a rising trend as a probability. When the trend crosses a threshold, the system should do what Ben-Zion describes - pause, offer resources, or hand off to a clinician.


Counting language features costs are negligible, as is any additional latency. And the same measurement layer feeds Ben-Zion's fourth guardrail, since you can't audit for interaction patterns you can't measure.


It also stops well short of surveillance because signals can be computed for one decision and dropped, or aggregated without profiling.


Ben-Zion is right that these systems need to catch distress before it becomes a crisis. Catching it means building the part that's currently missing: a way to measure how a person is changing, not just what they typed. It's a research and engineering problem, but the tools to solve it already exist.



This piece draws on research using the LIWC framework, for which Receptiviti and Receptiviti Labs hold exclusive commercial rights. Ryan Boyd, a co-author of the study discussed, works closely with Receptiviti. The suicidality and NSSI findings are research results, not clinical guidance.


References:


Ben-Zion, Z. (2026). AI can cause harm: safeguards must catch up. Nature. https://www.nature.com/articles/d41586-026-02109-z


Entwistle, C., Hoemann, K., Nightingale, S. J., & Boyd, R. L. (2025). Psychosocial dynamics of suicidality and nonsuicidal self-injury: a digital linguistic perspective. npj Mental Health Research, 4:28.

 
 

Subscribe to Field Notes

bottom of page