top of page

ALIGNMENT & SAFETY

Validated, independent measurement of how AI models impact users.

Validated, independent measurement of how AI models impact users

Derived from language.
Independent of the model being evaluated.

Alignment research requires evidence the model can't provide about itself. A model that benefits from user engagement cannot be the reliable source of evidence about whether that engagement is harmful — and without accurate signal about how the interaction is affecting the user, it can't respond appropriately either. Receptiviti measures the psychological dimension of every interaction — independently, longitudinally, and appended directly to existing logs and traces.

Production API · Containerized on-prem deployment available.

HOW IT WORKS

Pass conversation text. Receive structured psychological variables. Append to your existing logs.

No changes to your data infrastructure. Receptiviti scores sit alongside your existing traces, conversation logs, and eval outputs — adding the human-state dimension to data you are already collecting.

Conversation turn

User or model language - submitted as plain text

API CALL

Receptiviti

200+ psychological variables <65ms

APPEND

Your logs & traces

Existing eval data + human-state scores

Receptiviti measures psychological signals from language. It does not diagnose, classify individuals, or replace clinical judgment. Scores are structured variables for research and evaluation purposes.

On-prem deployment available

For privacy-sensitive research contexts, Receptiviti is available as a containerized on-prem deployment. Conversation text never leaves your infrastructure.

Contact us to discuss deployment options.

THE CORE PROBLEM

AI models cannot independently measure their own effect on users.

Current evaluation frameworks measure what the model outputs. They can't measure what those outputs do to users over time. That evidence has to come from outside the model being evaluated.

01

Circular evaluation

Asking the model to assess its own psychological impact on users is not independent measurement. Sycophancy, over-reliance, and dependency emerge precisely because users trust the model's signals — including signals about whether the interaction is going well.

02

Socioaffective alignment is a non-stationary target

The human-AI relationship changes user preferences and perceptions through mutual influence. Static benchmarks and single-session evals cannot capture what emerges longitudinally — cognitive offloading, dependency, erosion of independent reasoning — because these patterns only become visible across time.

03

Wellbeing claims require evidence

Safety and wellbeing commitments require evidence that survives external scrutiny. Inference-based signals are not auditable, not citable, and not independent. Structured, peer-reviewed measurement is.

WHAT RECEPTIVITI MEASURES

The dimensions alignment research needs — externally validated, derived from language.

200+ psychological dimensions scored from language, independent of model behavior. The same input yields the same score across model versions, prompt variations, and evaluation contexts.

COGNITIVE

Over-reliance & cognitive offloading

Analytical thinking scores track whether users are reasoning independently or deferring. Cognitive load measures whether interactions are adding friction. Trajectory shows whether reasoning engagement declines over sustained use.

AFFECTIVE

Emotional state & wellbeing trajectory
Emotions, distress signals, and cognitive load scored from language — not from the model's inference about how the user feels. Longitudinal wellbeing trajectory across interactions, not just single-session snapshots.

RELATIONAL

Dependency & attachment signals
Rapport and affiliation scores surface when relational dynamics are forming. Over-reliance signals are detectable as quantified variables before they become behavioural patterns.

SOCIAL

Agency & autonomy preservation

Clout, dominance, and deference signals indicate whether users maintain autonomy or defer to the model — critical for evaluating whether AI is supporting or subtly influencing user agency and autonomy.

WHY MEASUREMENT, NOT INFERENCE

Psychological state needs to be measured independently. Not inferred.

We need a science of AI safety that studies real human-AI interactions in natural contexts and treats the psychological and behavioural responses of users as key objects of inquiry.
Kirk et al. — npj Mental Health Research, 2025 · Oxford, Google DeepMind, UK AI Security Institute

It is important for model developers to consider the socioaffective alignment of their models, taking into account how models influence users' psychological states and social environments.

OpenAI — Investigating Affective Use and Emotional Well-being on ChatGPT, 2025

Three approaches exist for understanding what AI models do to users over time. They differ on key properties that matter to safety and alignment research. Unlike retrospective surveys, Receptiviti enables scoring of interactions in real time — turn by turn, across sessions. Unlike LLM inference, the scores are independent of the model being evaluated.

LLM inference


Varies with prompt phrasing and model version. Circular — the model being evaluated assesses its own impact. Not independently citable or auditable.

Retrospective surveys


After-the-fact, low-frequency, recall-biased. Cannot capture what happened turn by turn during the interaction.

Receptiviti measurement

Real-time, turn-level, and longitudinal. Independent of model behavior. Every dimension traceable to 34,000+ peer-reviewed citations. The same input yields the same score.

INTEGRATION

Into real-time adaptation, eval harnesses, red-teaming workflows, and research pipelines.

Receptiviti adds the human-state dimension to your existing evaluation infrastructure.

REAL-TIME ADAPTATION

Score each conversation turn as it happens. Feed interaction-state signals directly into your model's context or system prompt - so the model can adjust its behavior based on what's actually happening in the interaction, not what it infers.

EVAL HARNESSES

Score conversation logs alongside accuracy, faithfulness, and safety benchmarks. Human-state scores appended to existing eval outputs with a single API call per conversation turn or bulk-scored across entire log datasets.

RED-TEAMING

Track how interaction state evolves across adversarial sequences. Detect when scenarios produce distress, dependency, or cognitive offloading signals — beyond whether the model refused.

LONGITUDINAL STUDIES

Session-level and cross-session measurement of wellbeing trajectory, over-reliance, and autonomy preservation. The longitudinal dimension that single-turn evals cannot capture.

RESEARCH DATASETS

Annotate conversation datasets with validated psychological variables — richer supervision signal and wellbeing-aware benchmarking. On-prem deployment available for privacy-sensitive research.

RESEARCH

We publish. We contribute.

Receptiviti's team and academic advisors publish peer-reviewed research on the psychological dimensions of language and human behavior. Receptiviti also conducts its own research and experiments, contributing to the questions AI safety, alignment, and evaluation teams are actively working on.

PUBLISHED: PNAS Nexus, 2024

Large language models display human-like social desirability biases in personality surveys

Salecha, Ireland et al.,2024. Large language models display human-like response biases when they infer they are being evaluated, with effects up to 1.20 human SD across GPT-4, Claude 3, Llama 3, and PaLM-2. Co-authored by Molly Ireland (Receptiviti).

Read the paper →

PUBLISHED: npj Mental Health Research, 2025

Psychosocial dynamics of suicidality and nonsuicidal self-injury: a digital linguistic perspective

Entwistle, Hoemann, Nightingale & Boyd, 2025. Large-scale naturalistic study of the language dynamics surrounding suicidality and self-injury in 992 individuals with borderline personality disorder (66,786 posts). Co-authored by Ryan Boyd (UT Dallas), academic advisor to Receptiviti.

Read the paper →

PUBLISHED: npj Artificial Intelligence · 2026

PsychAdapter: adapting LLMs to reflect traits, personality, and mental health
Vu, Boyd, Eichstaedt et al., 2026. A lightweight LLM architectural modification that generates text reliably reflecting Big Five personality traits (87.3% accuracy) and mental health variables (96.7% accuracy). Co-authored by Ryan Boyd (UT Dallas), academic partner to Receptiviti.

Read the paper →

PUBLISHED: Perspectives on Psychological Science, 2026

Artificial intelligence and the psychology of human connection

Boyd & Markowitz, 2026. Introduces the MIRA model - a theoretical framework for when and how AI functions as a relational entity in human ecosystems. Language is the primary modality through which that relationship operates. Co-authored by Ryan Boyd (UT Dallas), academic partner to Receptiviti.

Read the paper →

PREPRINT: arXiv, May 2026, under review

When support escalates distress: regulation and escalation in LLM responses to venting and advice-seeking

Chi, Ganesan, Boyd, Ungar & Guntuku, 2026. Across 178,800 Reddit posts, LLM responses to venting simultaneously regulate and escalate distress — and the escalation is invisible to standard safety evaluations. Therapist personas reduce escalation; friend personas increase both. Measured using LIWC-22. Co-authored by Ryan Boyd (UT Dallas), academic advisor to Receptiviti.

​Read the preprint →

 SCOPED 

Interaction state as a missing dimension in AI evaluation
The case for human-state signals as a first-class eval criterion alongside accuracy, helpfulness, and harmlessness.

​Research partnerships →

Real-time, independent measurement of how your model impacts users.

Research partnerships · API integration · On-prem deployment

bottom of page