top of page

EVALUATION & OBSERVABILITY

The human-state eval dimension AI stacks are missing.

Derived from language.
Independent of the model being evaluated.

The human-state eval dimension AI stacks are missing.

Eval stacks instrument the model completely — traces, spans, latency, accuracy, hallucination rate, faithfulness. Receptiviti adds the one dimension none of it covers: how the interaction affects the user.

Drops into your existing pipeline · Python and R packages available

THE GAP

Every platform measures whether the output is good. None measure how it impacts the user.

Faithfulness, hallucination rate, relevance, toxicity, drift — all model-side. The interaction itself — whether it built understanding or eroded it, whether it raised anxiety or reduced it, whether rapport is accumulating or collapsing — has no instrumentation at all.

Every platform measures whether the output is good. None measure how it impacts the user.

INTEGRATION

One API call. Drops into existing eval pipelines.

Not a replacement for LangSmith, Arize, Braintrust, W&B, or Langfuse — an addition. Receptiviti scores sit alongside your existing metrics in the same pipeline, the same traces, the same dashboards.

The human-state eval dimension AI stacks are missing.
On-prem containerized deployment

On-prem containerized deployment

Available for privacy-sensitive environments. No data leaves your infrastructure. Contact us to discuss deployment options.

WHY NOT LLM-AS-JUDGE

Version-stable. Citable. Not prompt-dependent.

LLM-as-judge is known to be inconsistent across model versions and prompt phrasings. The same conversation scores differently on a different day. It can't be cited in a research paper or an accountability conversation.

Receptiviti scores are derived from validated psycholinguistic frameworks — explicitly computed, consistent across runs, traceable to 34,000+ peer-reviewed citations. The same input yields the same score.

LLM-AS-JUDGE

Prompt-sensitive, shifts across model versions, not independently citable,  circular when evaluating the model producing the judgment.

RECEPTIVITI

Version-stable, same input - same score, 34,000+ citations, independent of the model being evaluated.

SESSION-LEVEL & LONGITUDINAL

Single-turn evals miss what emerges across a conversation.

A model can pass every single-turn eval and still be systematically raising cognitive load, eroding confidence, or building dependency over the course of a session. Trajectory-level measurement catches what per-turn scoring cannot.

ONLINE EVALS

Score production traces in real time

Receptiviti scores append to live traces alongside your existing online eval metrics — latency, accuracy, safety flags.

OFFLINE EVALS

Score golden datasets & regression suites

Add interaction-state scoring to your existing golden datasets. Track human-side regression across model versions.

SESSION-LEVEL

Trajectory across the full conversation

Cognitive load trend, emotional trajectory, dependency signals — measured across turns, not just at the final output.

CI GATES

Block deploys on human-side regression

Add interaction-state thresholds to your CI pipeline — catch model updates that improve output quality but worsen the human experience.

RESEARCH

We publish. We contribute.

Receptiviti's team and academic advisors publish peer-reviewed research on the psychological dimensions of language and human behavior. Receptiviti also conducts its own research and experiments — contributing to the questions AI evaluation teams are actively working on.

PUBLISHED: npj Artificial Intelligence · 2026

PsychAdapter: adapting LLMs to reflect traits, personality, and mental health

Vu, Boyd, Eichstaedt et al., 2026. A lightweight LLM architectural modification that generates text reliably reflecting Big Five personality traits (87.3% accuracy) and mental health variables (96.7% accuracy).

Co-authored by Ryan Boyd (UT Dallas), academic advisor to Receptiviti.

Read the paper →

SCOPED

Interaction state as a missing dimension in AI evaluation
The case for human-state signals as a first-class eval criterion alongside accuracy, helpfulness, and harmlessness.

​Research partnerships →

PUBLISHED: PNAS Nexus, 2024

Large language models display human-like social desirability biases in personality surveys

Salecha, Ireland et al.,2024. Large language models display human-like response biases when they infer they are being evaluated, with effects up to 1.20 human SD across GPT-4, Claude 3, Llama 3, and PaLM-2. Co-authored by Molly Ireland (Receptiviti).

Read the paper →

INTERNAL STUDY

Conditioning educational AI on measured user state

GPT Study Mode: conditioning on psycholinguistic interaction-state signals produced a +4% improvement in overall educational effectiveness across 25 blinded evaluations. Largest gains in reasoning & scaffolding (+6.3%) and cognitive-load management (+5.9%).

Add the human-state dimension to your eval stack.

API access · Python and R packages · On-prem deployment available

bottom of page