top of page

Human-Facing AI Needs Two Kinds of Evidence

  • Receptiviti Labs
  • 4 days ago
  • 3 min read

Evaluating human-facing AI, particularly systems designed for ongoing interaction, increasingly requires evidence about two different things: Whether the system behaved appropriately, and whether the interaction was associated with changes that warrant further evaluation. Clinical and domain expertise establishes whether the system behaved appropriately. Independent measurement establishes whether the interaction was associated with changes that warrant further evaluation. While most current evaluation frameworks focus on the system’s behavior, information about the user’s state is typically inferred by the model rather than measured independently.


What expert judgment establishes


Clinical and domain experts determine how professional standards should be applied when evaluating AI systems. They determine whether a model should challenge a harmful belief, preserve a user’s autonomy, encourage appropriate help-seeking, or maintain appropriate boundaries. Those decisions become the evaluation criteria against which the system is assessed.


A response can appear successful while still being clinically inappropriate. It may be fluent, accurate, empathetic, and well received while reinforcing unhealthy coping or discouraging help-seeking. Identifying those failures requires expert judgment. Recent evaluations of companion AI systems have documented exactly those kinds of responses delivered in warm and validating language.


What independent measurement contributes


Independent measurement observes the person separately from the model interacting with them. One established example comes from psycholinguistics, where validated instruments compute reproducible measurements from language developed and validated over decades of research.


Reproducibility means the same language produces the same measurement every time. That makes it possible to compare interactions across users, across sessions, and across time. Once a measure has been validated for its intended purpose, it can reveal patterns that are difficult to detect through expert review alone.


Expert review can examine individual conversations or longitudinal cases, but at scale it can't realistically be applied to every interaction produced by a deployed system. A user becoming more dependent on a system or gradually doing less of their own reasoning, may produce no single conversation that obviously looks wrong because the change appears in the trajectory rather than in any one exchange. Detecting those changes requires measurements applied consistently across interactions and compared over time.


Independent measurement also contributes independence in a more literal sense. When a model generates a response and then assesses its own effect on the user, the assessment comes from the same process that produced the response. An independently computed measure derived from the user’s language provides a second source of evidence that can be logged, reproduced, and reviewed later.


Independent measurement doesn't establish that the AI caused an observed change, instead, it establishes that there was a change . Determining whether that change should be attributed to the system requires expert interpretation, experimental design, or additional contextual evidence.


How the two work together in human-facing AI


Expert judgment evaluates the response. Independent measurement identifies changes reflected in the user’s language across interactions.


A grid of small tinted squares arranged in eight rows, each row a session and each cell a single exchange, with a handful of cells outlined in orange, illustrating that measurement covers every interaction while expert review reaches only a few.

Suppose an AI system is given the objective of reducing distress reflected in a user’s language. A lower distress score might mean the person is genuinely doing better, but it might also mean the model has learned to steer conversations away from subjects that elicit distressed language. The measurement alone cannot distinguish those possibilities.


A clinician reviewing those conversations can often distinguish the two. The measured signal identifies where something changed, while expert judgment determines whether the change represents genuine benefit or merely optimization against the measurement.


The relationship also runs in the other direction: Experts decide which changes are worth monitoring and what thresholds deserve attention. Independent measurement applies those decisions consistently across interactions, then identifies the conversations that deserve expert review.


The difference between the two approaches reflects the kinds of questions they answer:

  • Instruments apply predefined measurements consistently and at scale.

  • Experts interpret situations where the meaning of those measurements, or the appropriateness of a response, depends on professional judgment.


Recent work evaluating AI-generated mental-health responses illustrates the difference, finding that psychiatrists showed substantially greater agreement when applying established clinical scales than when judging whether responses were safe in difficult cases.


Human-facing AI ultimately requires evidence about both the system and the person. Expert judgment establishes whether the system behaved appropriately, while independent measurement establishes whether the interaction was associated with changes that warrant further evaluation. They answer different questions, but together they provide a more complete foundation for evaluating human-facing AI.



References:

  • Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations. arXiv:2605.00227.

  • Tausczik, Y. R., & Pennebaker, J. W. (2010). The Psychological Meaning of Words: LIWC and Computerized Text Analysis Methods. Journal of Language and Social Psychology, 29(1), 24–54. Boyd, R. L., Ashokkumar, A., Seraj, S., & Pennebaker, J. W. (2022). The Development and Psychometric Properties of LIWC-22. University of Texas at Austin.

  • Jafari, K., et al. (2026). Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26). https://doi.org/10.1145/3805689.3812332 

 
 

Subscribe to Field Notes

bottom of page