EVALUATION & OBSERVABILITY
The human-state eval dimension AI stacks are missing.
Derived from language.
Independent of the model being evaluated.

Eval stacks instrument the model completely — traces, spans, latency, accuracy, hallucination rate, faithfulness. Receptiviti adds the one dimension none of it covers: how the interaction affects the user.
Drops into your existing pipeline · Python and R packages available
THE GAP
Every platform measures whether the output is good. None measure how it impacts the user.
Faithfulness, hallucination rate, relevance, toxicity, drift — all model-side. The interaction itself — whether it built understanding or eroded it, whether it raised anxiety or reduced it, whether rapport is accumulating or collapsing — has no instrumentation at all.

INTEGRATION
One API call. Drops into existing eval pipelines.
Not a replacement for LangSmith, Arize, Braintrust, W&B, or Langfuse — an addition. Receptiviti scores sit alongside your existing metrics in the same pipeline, the same traces, the same dashboards.


On-prem containerized deployment
Available for privacy-sensitive environments. No data leaves your infrastructure. Contact us to discuss deployment options.
WHY NOT LLM-AS-JUDGE
Version-stable. Citable. Not prompt-dependent.
LLM-as-judge is known to be inconsistent across model versions and prompt phrasings. The same conversation scores differently on a different day. It can't be cited in a research paper or an accountability conversation.
Receptiviti scores are derived from validated psycholinguistic frameworks — explicitly computed, consistent across runs, traceable to 34,000+ peer-reviewed citations. The same input yields the same score.
LLM-AS-JUDGE
Prompt-sensitive, shifts across model versions, not independently citable, circular when evaluating the model producing the judgment.
RECEPTIVITI
Version-stable, same input - same score, 34,000+ citations, independent of the model being evaluated.
SESSION-LEVEL & LONGITUDINAL
Single-turn evals miss what emerges across a conversation.
A model can pass every single-turn eval and still be systematically raising cognitive load, eroding confidence, or building dependency over the course of a session. Trajectory-level measurement catches what per-turn scoring cannot.
ONLINE EVALS
Score production traces in real time
Receptiviti scores append to live traces alongside your existing online eval metrics — latency, accuracy, safety flags.
OFFLINE EVALS
Score golden datasets & regression suites
Add interaction-state scoring to your existing golden datasets. Track human-side regression across model versions.
SESSION-LEVEL
Trajectory across the full conversation
Cognitive load trend, emotional trajectory, dependency signals — measured across turns, not just at the final output.
CI GATES
Block deploys on human-side regression
Add interaction-state thresholds to your CI pipeline — catch model updates that improve output quality but worsen the human experience.
RESEARCH
We publish. We contribute.
Receptiviti's team and academic advisors publish peer-reviewed research on the psychological dimensions of language and human behavior. Receptiviti also conducts its own research and experiments — contributing to the questions AI evaluation teams are actively working on.
PUBLISHED: npj Artificial Intelligence · 2026
PsychAdapter: adapting LLMs to reflect traits, personality, and mental health
Vu, Boyd, Eichstaedt et al., 2026. A lightweight LLM architectural modification that generates text reliably reflecting Big Five personality traits (87.3% accuracy) and mental health variables (96.7% accuracy).
Co-authored by Ryan Boyd (UT Dallas), academic advisor to Receptiviti.
SCOPED
Interaction state as a missing dimension in AI evaluation
The case for human-state signals as a first-class eval criterion alongside accuracy, helpfulness, and harmlessness.
PUBLISHED: PNAS Nexus, 2024
Large language models display human-like social desirability biases in personality surveys
Salecha, Ireland et al.,2024. Large language models display human-like response biases when they infer they are being evaluated, with effects up to 1.20 human SD across GPT-4, Claude 3, Llama 3, and PaLM-2. Co-authored by Molly Ireland (Receptiviti).
INTERNAL STUDY
Conditioning educational AI on measured user state
GPT Study Mode: conditioning on psycholinguistic interaction-state signals produced a +4% improvement in overall educational effectiveness across 25 blinded evaluations. Largest gains in reasoning & scaffolding (+6.3%) and cognitive-load management (+5.9%).