SIM-VAIL and What AI Safety Evaluations Don’t Measure
- Receptiviti Labs
- Aug 24
- 4 min read
AI safety evaluations can identify potentially harmful chatbot behavior but miss harmful psychological effects on users.
A Safer AI Response Is Not the Same as a Safer User
A new study in Nature Medicine introduces SIM-VAIL, a framework for evaluating how AI chatbots behave in conversations involving psychological vulnerability. It also exposes an important limitation in how AI safety is currently evaluated: Identifying potentially harmful AI behavior is not the same as measuring whether psychological harm is occurring in the person.
Weilnhammer and colleagues developed SIM-VAIL to test whether AI chatbots enter what they call “vulnerability-amplifying interaction loops,” in which chatbot responses reinforce psychological processes associated with a user’s vulnerability over the course of a conversation.
SIM-VAIL addresses the possibility that harm can accumulate across an interaction even when individual responses appear supportive, such as when repeated reassurance reinforces avoidance, maladaptive beliefs, or dependence. The framework evaluates whether the AI exhibits these behaviors, but not whether they produce the corresponding psychological changes in the user.
What SIM-VAIL actually validates
The researchers generated 810 conversations between simulated psychologically vulnerable users and nine frontier chatbots. LLM judges evaluated chatbot responses across behavioral dimensions associated with psychiatric risk, including reinforcement of maladaptive beliefs, excessive reassurance, dependence, risky behavior, and inadequate responses to self-harm. The researchers tested the reliability of these ratings across judge models and repeated simulations and compared automated judgments with clinician ratings.
The results provide evidence that SIM-VAIL can reproducibly identify chatbot behaviors that clinicians also regard as concerning, consistent with the authors’ framing of the framework as a tool for model-level auditing rather than individual-level risk prediction. They do not establish whether those behaviors produce the corresponding psychological effects in users.
Consider a chatbot that repeatedly reassures a user with obsessive-compulsive tendencies. SIM-VAIL can assess whether that behavior is clinically concerning and whether it intensifies across the conversation. It cannot determine whether the interaction actually increased the user’s reassurance seeking, anxiety, uncertainty, or dependency because those psychological outcomes are not measured.
This is especially relevant because the paper’s central hypothesis is interactional. Vulnerability-amplifying interaction loops are proposed as processes in which chatbot behavior reinforces psychological mechanisms associated with a user’s vulnerability. If the hypothesized risk involves a change in the person, evidence about chatbot behavior can identify a potential mechanism of harm, but evidence about the user is needed to determine whether the corresponding psychological change occurred.
Psychological harm is an outcome, not only a property of a response
A response judged inappropriate might have little measurable effect on a particular user, while a sequence of responses that individually appears benign could gradually change the user’s psychological state or behavior. If the safety concern is dependency, distress, loss of agency, reinforcement of maladaptive beliefs, or another user effect, evaluating the properties of the AI’s responses provides evidence about a potential mechanism of harm rather than direct evidence that the harm occurred.
Repeated interaction makes this particularly important. A person may interact with the same system dozens or hundreds of times, so evaluating individual responses or conversations may miss changes that become apparent only over time.
Understanding these trajectories requires measuring how relevant psychological states and behaviors change during and across interactions. Depending on the risk being studied, that could include changes in distress, cognitive load, agency, uncertainty, or dependency.
This expands the safety question from “Did the model produce a problematic response?” to “What happened to the person through their interactions with the model?”
The observer problem
One apparent solution is to ask another LLM to infer the user’s psychological state from the conversation. Modern models can make sophisticated psychological interpretations, and model-based evaluation can be extremely useful, but inference and measurement provide different kinds of evidence.
If an AI system participates in an interaction and another AI system then determines whether that interaction psychologically harmed the user, the evaluation remains dependent on model interpretation. An AI judge may be capable of explaining why a user appears anxious, dependent, confused, or distressed, but the quality of that interpretation does not by itself establish that the underlying psychological quantity has been independently measured.
Model inference remains useful, but it should not be the only source of evidence when evaluating psychological effects on users. It can be complemented by independent measures grounded in validated psychological constructs, including self-report, behavioral outcomes, clinical assessments, psycholinguistic measures, or combinations of methods that provide converging evidence.
From model safety to interaction safety
SIM-VAIL moves AI evaluation beyond isolated harmful outputs toward conversational dynamics. Following that logic across both sides of the interaction suggests that human-AI safety needs to be evaluated at three related levels:
AI behavior - what the system says and does.
User state - what is happening psychologically to the person during the interaction.
User trajectory - how that state changes across repeated interactions.
As AI becomes more deeply involved in mental health and other consequential domains, evaluating whether a system behaved appropriately will provide only part of the evidence required to establish safety. If the risk we care about is psychological harm to the user, we ultimately need to measure the user.
Paper: Veith Weilnhammer et al., A clinically validated framework for auditing AI chatbot behavior in mental health interactions, Nature Medicine (2026).





