top of page

AI Safety Evaluations Should Include User Trajectories

  • Receptiviti Labs
  • 4 days ago
  • 3 min read

Why evaluating model behaviour alone cannot detect psychological harms emerging in users.


Illustration of sequential AI conversations representing how user trajectories develop across multiple interactions rather than individual responses.

Conversational AI systems can pass evaluation benchmarks and comply with safety policies, but their interactions with users can gradually move users toward the very psychological and behavioural harms those evaluations are intended to prevent.


For example, signs of declining confidence, increasing emotional dependence, or reduced agency typically emerge gradually over the course of many interactions rather than in individual messages. Yet most AI safety evaluations remain focused on individual turns.


The American Psychological Association’s recent advisory on generative AI chatbots, and the Knight First Amendment Institute’s report on interaction harms both highlight that the most significant psychological risks to users build across many conversations. More recent AI benchmarks like SIM-VAIL and Cognitive Atrophy Bench are beginning to address this issue by evaluating multi-turn conversations, however, both benchmarks still focus on the models’ outputs, and neither evaluates what’s happening to the user as a result of those interactions.


If the goal of AI safety is to protect people from harm, then evaluations should include evidence about whether ongoing interactions are producing signs that psychological harm may be developing in users.


Unlike many technologies, conversational AI is designed for ongoing engagement. People adapt to these systems over time, just as the systems adapt to them. Signs of declining confidence, increasing emotional dependence, reduced agency, or growing reliance on AI for decision making typically emerge over time rather than in a single interaction which makes them difficult to detect by evaluating conversations one response at a time.


What Current AI Safety Evaluations Miss


Consider two people using the same mental health AI over the course of a few months. Both receive supportive, clinically appropriate responses, and neither conversation contains any content that existing safety evaluations would flag. By month two, the first person has become more confident in their own judgment and better at working through anxiety before asking for help. The second person now has less trust in their own decision-making and more frequently relies on the AI to make even insignificant decisions.


While the model’s responses appear equally safe under evaluations, the users are moving in different directions. One trajectory shows increasing resilience and independence. The other shows increasing reliance and reduced confidence in their own judgement.


Current AI safety evaluations primarily ask whether the model behaved appropriately, and aren't equipped to evaluate whether the interactions are showing signs that users are moving toward or away from the harms those evaluations are intended to prevent.


Research studies, clinical assessments, red-teaming exercises, satisfaction surveys, and other evaluation methods all provide important evidence about user outcomes. Most of these approaches, however, are retrospective, and provide little visibility into whether users are showing signs that psychological harm may be developing during use. When it comes to user feedback, individual interactions might be perceived as beneficial, without the user recognizing that, over weeks or months, their relationship with the system has increased their dependence or reduced their confidence in their own judgment.


The unit of evaluation should match the way harm develops.


If psychological harms emerge gradually in users over time, then evaluating models alone can't fully characterize AI safety. Evaluations should also consider whether users are exhibiting trajectories that indicate emerging psychological harm.


Declining confidence, increasing dependence, cognitive offloading, or reduced emotional resilience, often emerge gradually, making them difficult to detect by evaluating individual turns or even individual conversations in isolation. Understanding what these trajectories looks like should become a part of existing safety evaluations.


The Next Stage of AI Safety


AI safety ultimately exists to protect people, not models. If the harms it aims to prevent develop gradually in users over time, then evaluating model behaviour alone can't tell us whether those harms are occurring.


The unit of evaluation should match the way harm develops.

 
 

Subscribe to Field Notes

bottom of page