top of page

EVALUATION & OBSERVABILITY

The human-state eval dimension AI stacks are missing.

Derived from users’ language. Independent of the model being evaluated.

The human-state eval dimension AI stacks are missing.

Eval stacks instrument the model extensively - traces, spans, latency, accuracy, hallucination rate, faithfulness. Receptiviti adds a dimension conventional eval stacks rarely capture: what is happening on the human side of the interaction.

THE GAP

AI evaluations measure what models do, but rarely measure how interactions affect the people using them.

Faithfulness, hallucination rate, relevance, toxicity, and drift are predominantly model-side measures. Receptiviti adds structured measurement of user-side characteristics and change, including cognitive load, analytical thinking, anxiety, affect, rapport, and related signals.

The gap.png

INTEGRATION

One API call. Drops into existing eval pipelines.

Not a replacement for LangSmith, Arize, Braintrust, W&B, or Langfuse — an addition. Receptiviti scores sit alongside your existing metrics in the same pipeline, the same traces, the same dashboards.

The human-state eval dimension AI stacks are missing.
On-prem containerized deployment

On-prem containerized deployment

Available for privacy-sensitive environments. No data leaves your infrastructure. Contact us to discuss deployment options.

WHY NOT LLM-AS-JUDGE

Deterministic. Reproducible.
Not prompt-dependent.

LLM-as-judge scores can vary with prompt phrasing, model choice, model version, and sampling conditions. That makes them useful for many evaluation tasks, but less suitable when the requirement is a fixed, independently reproducible psychological measurement.

​

Receptiviti scores are explicitly computed using validated psycholinguistic methods. Under a fixed scoring version, the same input produces the same score, independent of the model being evaluated.

LLM-AS-JUDGE

Prompt-sensitive, shifts across model versions, not independently citable,  circular when evaluating the model producing the judgment.

RECEPTIVITI

Version-stable, same input - same score, 34,000+ citations, independent of the model being evaluated.

SESSION-LEVEL & LONGITUDINAL

Single-turn evals miss what emerges across a conversation.

A model can pass output-level evals while user-side measures worsen across a conversation - rising cognitive load, declining analytical engagement, increasing distress, or growing patterns of deference. Trajectory-level measurement makes those changes observable.

ONLINE EVALS

Score production traces in real time

Receptiviti scores append to live traces alongside your existing online eval metrics - latency, accuracy, safety flags.

OFFLINE EVALS

Score golden datasets & regression suites

Add interaction-state scoring to your existing golden datasets. Track human-side regression across model versions.

SESSION-LEVEL

Trajectory across the full conversation

Cognitive load trend, emotional trajectory, dependency signals - measured across turns, not just at the final output.

CI GATES

Block deploys on human-side regression

Add interaction-state thresholds to your CI pipeline - catch model updates that improve output quality but worsen the human experience.

RESEARCH

We publish. We contribute.

Receptiviti's team and academic advisors publish peer-reviewed research on the psychological dimensions of language and human behavior. Receptiviti also conducts its own research and experiments — contributing to the questions AI evaluation teams are actively working on.

PUBLISHED: npj Artificial Intelligence · 2026

PsychAdapter: adapting LLMs to reflect traits, personality, and mental health

Vu, Boyd, Eichstaedt et al., 2026. A lightweight LLM architectural modification that generates text reliably reflecting Big Five personality traits (87.3% accuracy) and mental health variables (96.7% accuracy).

Co-authored by Ryan Boyd (UT Dallas), academic advisor to Receptiviti.

​

Read more →

SCOPED

Interaction state as a missing dimension in AI evaluation
The case for human-state signals as a first-class eval criterion alongside accuracy, helpfulness, and harmlessness.

​

​Research partnerships →

PUBLISHED: PNAS Nexus, 2024

Large language models display human-like social desirability biases in personality surveys

Salecha, Ireland et al.,2024. Large language models display human-like response biases when they infer they are being evaluated, with effects up to 1.20 human SD across GPT-4, Claude 3, Llama 3, and PaLM-2. Co-authored by Molly Ireland (Receptiviti).

​​

Read more →

INTERNAL STUDY

Providing measured user state as context to educational AI

GPT Study Mode: Providing psycholinguistic interaction-state signals produced a +4% improvement in overall educational effectiveness across 25 blinded evaluations. Largest gains in reasoning & scaffolding (+6.3%) and cognitive-load management (+5.9%).

Add the human-state dimension to your eval stack.

API access · Python and R packages · On-prem deployment available

bottom of page