top of page

AI Can't Tell If It's Helping Your Judgment or Replacing It

  • Receptiviti Labs
  • Jul 10
  • 3 min read

Updated: Aug 4

Thinking Machines built its first year around a clear claim: AI should extend human judgment, not replace it.¹


While most AI development treats human involvement as a temporary limitation, something to be engineered away as capability improves, Thinking Machines is arguing the opposite: that human participation is the point, not the bottleneck.


We agree with the premise. This post is about what it takes to actually deliver on it.


AI Over-Reliance: Extending judgment and eroding it look the same from inside the model.


A conversation can go well and still leave someone worse off. A helpful answer today can make it a little harder to work through the same problem alone tomorrow, and harder again after that. Peer-reviewed research documents that sustained AI use is associated with over-reliance, reduced critical engagement, and eroding epistemic independence.² The effect never appears inside a single message; it accumulates across many messages.


That accumulation is what makes it hard to catch. Someone who feels well-supported this week may have quietly stopped forming their own judgments by month two, and the model answering each message has no way to notice the change, because each message on its own looks fine. The change exists in the trajectory, where the model never looks.

A person who feels well-supported by an AI model this week can be a person who has quietly stopped forming their own judgments by month two, and the model answering each individual message has no way to notice, because the change never shows up inside any one message. It only shows up in the pattern across many.

This is the gap in Thinking Machine's "extend, don't replace" as a design principle. The intention is right, but the system executing it has no way to check whether it's actually succeeding.


A model can't audit its own effect on the user.


If a system is meant to extend judgment rather than substitute for it, something has to be able to tell whether that's actually happening — whether a given person is getting sharper or more dependent, more confident or more deferential, across the sessions that make up their relationship with the tool. That evaluation can't come from the model itself. A model that benefits from engagement, that is rewarded for being helpful in the moment, is not positioned to generate independent evidence about whether it's helping or quietly training the person to stop thinking for themselves. The judgment needs a source of evidence outside the system being judged.


The evidence is already in the language.


People's cognitive and emotional state is carried in the structural features of how they write — word choice, sentence construction, the small linguistic markers that shift with confidence, doubt, and cognitive load. This has been measurable for three decades, through peer-reviewed psycholinguistic research.


It's a signal that can be measured in every conversation already, and us traceable to published science. That independence is what makes it so useful for AI: it doesn't inherit the biases of the system it's evaluating, because it isn't produced by that system.


What "extend, not replace" requires in practice.


If Thinking Machine's goal is a system that strengthens a person's judgment over time, that system needs to see the person's cognitive and emotional trajectory across sessions, not just whether the last response was helpful, but whether the pattern across weeks looks beneficial (like growing capability) or detrimental (like growing dependence). Most eval stacks don't evaluate this today, and the model being evaluated can't evaluate it either.


This is the layer we're building at Receptiviti Labs: validated measurement of what's happening to the person across the relationship, independent of the model. A model's in-conversation judgment is only useful if the model or eval team can tell whether it's working. This layer supplies that check: evidence for whether the model's judgment is actually extending the person's capability, or if it just seems helpful in the moment.


Building toward a future where AI extends human capability requires a way to check, session by session, whether the intention is being realized.


If you're working on that problem, we'd like to hear from you.



  1. Thinking Machines Lab, "The Future Worth Building Is Human" (2026).

  2. Gerlich, M. (2025). "AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking." Societies 15(1), 6. doi:10.3390/soc15010006.


 
 

Subscribe to Field Notes

bottom of page