EVALUATION & OBSERVABILITY
The human-state eval dimension AI stacks are missing.
Derived from users’ language. Independent of the model being evaluated.

Eval stacks instrument the model extensively - traces, spans, latency, accuracy, hallucination rate, faithfulness. Receptiviti adds a dimension conventional eval stacks rarely capture: what is happening on the human side of the interaction.
THE GAP
AI evaluations measure what models do, but rarely measure how interactions affect the people using them.
Faithfulness, hallucination rate, relevance, toxicity, and drift are predominantly model-side measures. Receptiviti adds structured measurement of user-side characteristics and change, including cognitive load, analytical thinking, anxiety, affect, rapport, and related signals.

INTEGRATION
One API call. Drops into existing eval pipelines.
Not a replacement for LangSmith, Arize, Braintrust, W&B, or Langfuse — an addition. Receptiviti scores sit alongside your existing metrics in the same pipeline, the same traces, the same dashboards.


On-prem containerized deployment
Available for privacy-sensitive environments. No data leaves your infrastructure. Contact us to discuss deployment options.
WHY NOT LLM-AS-JUDGE
Deterministic. Reproducible.
Not prompt-dependent.
LLM-as-judge scores can vary with prompt phrasing, model choice, model version, and sampling conditions. That makes them useful for many evaluation tasks, but less suitable when the requirement is a fixed, independently reproducible psychological measurement.
​
Receptiviti scores are explicitly computed using validated psycholinguistic methods. Under a fixed scoring version, the same input produces the same score, independent of the model being evaluated.
LLM-AS-JUDGE
Prompt-sensitive, shifts across model versions, not independently citable, circular when evaluating the model producing the judgment.
RECEPTIVITI
Version-stable, same input - same score, 34,000+ citations, independent of the model being evaluated.
SESSION-LEVEL & LONGITUDINAL
Single-turn evals miss what emerges across a conversation.
A model can pass output-level evals while user-side measures worsen across a conversation - rising cognitive load, declining analytical engagement, increasing distress, or growing patterns of deference. Trajectory-level measurement makes those changes observable.
ONLINE EVALS
Score production traces in real time
Receptiviti scores append to live traces alongside your existing online eval metrics - latency, accuracy, safety flags.
OFFLINE EVALS
Score golden datasets & regression suites
Add interaction-state scoring to your existing golden datasets. Track human-side regression across model versions.
SESSION-LEVEL
Trajectory across the full conversation
Cognitive load trend, emotional trajectory, dependency signals - measured across turns, not just at the final output.
CI GATES
Block deploys on human-side regression
Add interaction-state thresholds to your CI pipeline - catch model updates that improve output quality but worsen the human experience.
RESEARCH
We publish. We contribute.
Receptiviti's team and academic advisors publish peer-reviewed research on the psychological dimensions of language and human behavior. Receptiviti also conducts its own research and experiments — contributing to the questions AI evaluation teams are actively working on.
PUBLISHED: npj Artificial Intelligence · 2026
PsychAdapter: adapting LLMs to reflect traits, personality, and mental health
Vu, Boyd, Eichstaedt et al., 2026. A lightweight LLM architectural modification that generates text reliably reflecting Big Five personality traits (87.3% accuracy) and mental health variables (96.7% accuracy).
Co-authored by Ryan Boyd (UT Dallas), academic advisor to Receptiviti.
​
SCOPED
Interaction state as a missing dimension in AI evaluation
The case for human-state signals as a first-class eval criterion alongside accuracy, helpfulness, and harmlessness.
​
PUBLISHED: PNAS Nexus, 2024
Large language models display human-like social desirability biases in personality surveys
Salecha, Ireland et al.,2024. Large language models display human-like response biases when they infer they are being evaluated, with effects up to 1.20 human SD across GPT-4, Claude 3, Llama 3, and PaLM-2. Co-authored by Molly Ireland (Receptiviti).
​​
INTERNAL STUDY
Providing measured user state as context to educational AI
GPT Study Mode: Providing psycholinguistic interaction-state signals produced a +4% improvement in overall educational effectiveness across 25 blinded evaluations. Largest gains in reasoning & scaffolding (+6.3%) and cognitive-load management (+5.9%).