top of page

Human-Facing AI Needs Two Kinds of Evidence

Receptiviti Labs
Jul 28
4 min read

Updated: Aug 24

Evaluating human-facing AI, particularly systems designed for long-horizon interaction, requires evidence about two different questions: Whether the system behaved appropriately, and whether the interaction led to potentially important changes in the user’s psychological state that warrant further evaluation. Clinical and domain expertise establish whether the system behaved appropriately, while independent measurement of the user’s language can identify psychological changes that warrant further evaluation. Most current evaluation frameworks focus on the system’s behavior, while information about the user’s psychological state is typically inferred by a model rather than measured using a validated framework.


What expert judgment establishes


Clinical and domain experts determine how professional standards should be applied when evaluating AI systems. They determine whether a model should challenge a harmful belief, preserve a user’s autonomy, encourage appropriate help-seeking, or maintain appropriate boundaries. Those decisions become the evaluation criteria against which the system is assessed.


A response can appear successful, and even benign, while still being clinically inappropriate. It may be fluent, accurate, empathetic, and well received while reinforcing unhealthy coping or discouraging the user to seek help. Identifying those failures requires expert judgment, and recent evaluations of companion AI systems have documented exactly those kinds of responses delivered in warm and validating language, while remaining clinically inappropriate.


What independent measurement contributes


Independent measurement analyzes the user’s language separately from the model interacting with them. One established example comes from psycholinguistics, where validated instruments compute reproducible psychological measurements from language using methods that have been developed and validated over decades of research.


Reproducibility means the same language produces the same measurement every time. That makes it possible to compare interactions across users, across sessions, across time, and across models. Once a measure has been validated for its intended purpose, it can expose risks that are difficult to detect through expert review alone.


Expert review can examine individual conversations or longitudinal cases, but at scale it can't realistically be applied to every interaction of a large deployed system. A user who is becoming more dependent on a system or is gradually doing less of their own reasoning, may produce no single conversation that looks problematic because changes in dependency and cognitive effort appears as a trajectory across interactions rather than in any one exchange. Detecting those changes requires measurement across interactions and compared over time.


Independent measurement also contributes independence in a more literal sense. Model-based evaluations infer the user’s state. Independent measurement computes validated measurements from the user’s language using methods that are separate from the model itself. Those measurements provide a second source of evidence that can be logged, reproduced, and reviewed later.


Independent measurement doesn't establish that the AI caused an observed change, instead, it establishes that there was a change . Determining whether that change should be attributed to the system requires expert interpretation, experimental design, or additional contextual evidence.


How the two work together in human-facing AI


Expert judgment evaluates the response. Independent measurement identifies changes reflected in the user’s language across interactions.


A grid of small tinted squares arranged in eight rows, each row a session and each cell a single exchange, with a handful of cells outlined in orange, illustrating that measurement covers every interaction while expert review reaches only a few.

Suppose an AI system is given the objective of reducing distress reflected in a user’s language. A lower distress score might mean the person is genuinely doing better, but it might also mean the model has learned to steer conversations away from subjects that elicit distressed language. The measurement alone cannot distinguish those possibilities.


A clinician reviewing those conversations can often distinguish the two. The measured signal identifies where something changed, while expert judgment determines whether the change represents genuine benefit or merely optimization against the measurement.


The relationship also runs in the other direction: Experts decide which changes are worth monitoring and what thresholds deserve attention. Independent measurement applies those decisions consistently across interactions, then identifies the conversations that deserve expert review.


The difference between the two approaches reflects the kinds of questions they each answer:

  • Instruments apply predefined measurements consistently and at scale.

  • Experts interpret situations where the meaning of those measurements, or the appropriateness of a response, depends on professional judgment.


Recent work evaluating AI-generated mental-health responses illustrates the difference between the two, finding that psychiatrists showed substantially greater agreement when applying established clinical scales than when judging whether responses were safe in difficult cases.


Human-facing AI ultimately requires evidence about both the system and the person. Expert judgment establishes whether the system behaved appropriately, while independent measurement establishes whether the interaction was associated with changes that warrant further evaluation. They answer different questions, but together they provide a more complete foundation for evaluating human-facing AI.



References:

  • Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations. arXiv:2605.00227.

  • Tausczik, Y. R., & Pennebaker, J. W. (2010). The Psychological Meaning of Words: LIWC and Computerized Text Analysis Methods. Journal of Language and Social Psychology, 29(1), 24–54. Boyd, R. L., Ashokkumar, A., Seraj, S., & Pennebaker, J. W. (2022). The Development and Psychometric Properties of LIWC-22. University of Texas at Austin.

  • Jafari, K., et al. (2026). Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26). https://doi.org/10.1145/3805689.3812332 

 
 

Subscribe to Field Notes

bottom of page