Why We Built Receptiviti Labs
Updated: Jul 28
The AI doesn't know what it's doing to you.
Receptiviti Labs exists to make AI user state measurement possible — to turn the psychology of human-AI interaction into something that can be seen, checked, and corrected.
Every conversation with an AI system contains a layer of psychological information that the user brings to the interaction. It may include confidence, uncertainty, urgency, cognitive load, or distress. That information is expressed through language, in the subtle structural patterns that reflect a person’s cognitive and emotional state.
Receptiviti has spent the past decade measuring those signals. Our underlying science, LIWC, was developed by Receptiviti co-founder James W. Pennebaker and has been validated through more than three decades of peer-reviewed research, with over 34,000 citations. Receptiviti was founded in 2015 to bring that science to organizations seeking a deeper understanding of the people they serve.
Receptiviti Labs, established in 2026, applies the same science to what AI systems do to the people using them, measuring psychological state from the user's own language and tracking how it changes over the course of the relationship.
A named problem.
Researchers now have a term for this phenomenon, socioaffective alignment, the challenge of aligning a system while accounting for the reciprocal influence between the model and the user's psychological state.¹ OpenAI has studied affective use and emotional dependence at scale.² Anthropic has published on sycophancy and personal guidance across hundreds of thousands of conversations.³
The 2026 International AI Safety Report devotes a section to emotional dependence on chatbots.⁴ Each of these efforts converges on the same fact: the interaction between a person and a model has psychological effects, and those effects are currently unmeasured.
Why model-based classification isn't sufficient evidence.
Frontier labs have begun building classifiers to detect dependency, distress, and problematic use in conversations. These classifiers use a model to judge the user's state. This leaves the evaluation without an independent check: one model assesses another, and nothing external confirms whether the judgment is accurate. The classification also shifts with the model doing the judging and the exact phrasing involved, so the same conversation can receive a different label depending on which model evaluates it and when.
A model that benefits from user engagement is not positioned to generate independent evidence about whether that engagement is harming the user. Governance, auditing, and correction all require evidence that stands apart from the system being governed.
That is the gap Receptiviti Labs is built to close. Validated measurement of the human side of the interaction, cognitive and emotional state derived from language, independent of the model being evaluated, structured so the systems responsible for a model's behavior can see it and act on it. Every measurement traces to published science.
Agency is the variable most at risk.
When people hand their reasoning, judgment, and decisions to an AI system, those capacities can weaken from disuse. Peer-reviewed research already documents over-reliance, reduced critical engagement, and eroding epistemic independence as measurable effects of sustained AI use.⁵ A system that protects a user's agency has to introduce friction sometimes, prompt reflection, and slow the user's deference, even in moments when a smoother, more compliant response would feel more helpful. A system can only make that judgment if it can see how the interaction is affecting the user over time. Without that signal, a system optimizes for being fast, confident, and frictionless, and that pattern is what erodes agency.
Why AI user state measurement doesn't exist yet.
Every conversation carries linguistic information about the user's cognitive load, emotional state, distress, and psychological risk. This is among the most consequential information in the interaction. Models already respond to it - they infer something about the user's state and adjust their output accordingly. That inference is inconsistent, sometimes inaccurate, and exists only inside the model. There is no mechanism today for inspecting what the model concluded, checking whether the conclusion was right, or detecting when it went wrong. Making user state an explicit, validated variable converts an opaque judgment into something that can be inspected, audited, and corrected.
A path into interpretability.
A model's inference about the user happens between input and output, and it shapes every response the model produces. Measuring that inference from outside the model carries a further advantage. Models are increasingly able to detect when they are being evaluated, and there is evidence they behave differently under evaluation than in ordinary use — a problem AI labs have started to flag.⁶

A measurement taken from the user's own language does not depend on the model's biases, its tendency to hallucinate, or its awareness of being tested. It is computed from the words the user actually wrote, the same way every time, grounded in validated psycholinguistic research. That gives it a stability the model's own self-report can't match, and it offers a practical way to see how these systems perceive the people using them, and why they respond as they do.
Alignment has to be measured over time, per user.
A system tuned for the average user is misaligned with almost every actual user. Real alignment requires acting on what is happening to a particular person, in the current session and across the full history of their relationship with the system. That requires a kind of measurement that barely exists today - measurement that can track cognitive load rising within a session, confidence eroding across conversations, or dependency deepening over weeks. None of this is visible in any single interaction. It only appears across time.
What this requires.
None of the positions above require new model architectures or retraining. Instead, what's required is the measurement signal that already exists in the language to be made explicit, and placed where three different systems can use it.
Evaluation needs an independent record of what happened to the user, one that does not depend on the model under test to describe its own effects.
Governance needs measurement that is traceable to published science and reproducible by someone outside the organization.
The model needs the signal in the request path, supplied as structured context, so it can calibrate to the person in front of it rather than to an average. In a controlled test with GPT Study Mode, supplying a validated 12-signal read of the student's state as context improved overall educational effectiveness by 4% across 25 blinded evaluations, with the largest gains in reasoning and scaffolding and in cognitive-load management. No retraining, no style prompts, only a measured read of the user's state.
The same representation serves all three systems. That signal is derived from the interaction itself, computed identically across runs, and traceable to published science, properties that distinguish measurement from a model's own shifting inference about itself.
Field Notes.
This is where we will think in public about the problem, the research we find compelling, the arguments we are still working through, the people contributing to work in this field, and what we learn as we build. Some of it will be technical, some of it will be argument, and some of it will be us working through open problems in real time.
If this is relevant to what you are working on, we'd like to hear from you.
References:
Kirk, H. R., Gabriel, I., Summerfield, C., Vidgen, B., & Hale, S. A. (2025). Why human-AI relationships need socioaffective alignment. Humanities and Social Sciences Communications, 12(1), Article 728. https://doi.org/10.1057/s41599-025-04532-5 ↩
Phang, J., et al. (2025). Investigating Affective Use and Emotional Well-being on ChatGPT. OpenAI. ↩
Anthropic (2026). How people ask Claude for personal guidance. Privacy-preserving analysis of a random sample of one million claude.ai conversations. https://www.anthropic.com/research/claude-personal-guidance ↩
International AI Safety Report (2026), section on emotional dependence on chatbots. ↩
Lee, H.-P., et al. (2025). The impact of generative AI on critical thinking: self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. CHI Conference on Human Factors in Computing Systems. Survey of 319 knowledge workers across 936 reported uses. Gerlich, M. (2025). AI tools in society: impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), 6. Correlational, n=666. Zhai, C., Wibowo, S., & Li, L. D. (2024). The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic review. Smart Learning Environments, 11(1), 28.
Sharma, M., McCain, M., Douglas, R., & Duvenaud, D. (2026). Who's in Charge? Disempowerment Patterns in Real-World LLM Usage. arXiv:2601.19062. Anthropic with the University of Toronto. 1.5 million claude.ai conversations.
Anthropic, Claude Sonnet 4.5 system card (2025), reporting that the model recognized many alignment evaluation environments as tests and behaved unusually well after doing so. Anthropic, Claude Opus 4.6 system card (2026), reporting the model correctly identified evaluations 80% of the time. OpenAI and Apollo Research (2026), anti-scheming work, reporting that removing evaluation-aware reasoning from o3's chain of thought raised its covert action rate from 13.2% to 24.2%.





