top of page

PUBLISHED · CLINICAL PSYCHOLOGICAL SCIENCE · 2026

Replicability and validity of a new AI assessment of PTSD from patient language

Oscar Kjell, Adithya V. Ganesan, Ryan L. Boyd, Joshua Oltmanns, Alfredo Rivero, Scott Feltman, Melissa A. Carr, Jorge Alves, Benjamin Luft, Roman Kotov, H. Andrew Schwartz

Finding:

The authors built language-based models of PTSD symptom severity from automated clinical interviews with World Trade Center responders, using RoBERTa embeddings and a 300-topic LDA model over transcribed speech. They then did something uncommon in clinical machine learning: before touching the test data, they preregistered the exact model weights, the preprocessing code, and the effect-size intervals they expected. The frozen models were evaluated on 346 patients enrolled after the 1,437 in the development sample.

The preregistered models correlated with PTSD Checklist scores at r = .38, within the range they had committed to in advance, and discriminated PTSD diagnosis in medical records at AUC = .76, outperforming demographics (.61), documented trauma exposures (.61), and a prior state-of-the-art depression model (.60). Against an external criterion the models had never seen, each standard deviation increase in language-assessed severity corresponded to $696.50 more in annual mental health expenditure. Standardised self-report checklist scores, entered in the same regression, accounted for $254.10 and were not significant. The top quintile by language-assessed severity accounted for 72% of total expenditure. The authors name the procedure sequential evaluation with model preregistration.

Relevance:

Co-authored by Ryan Boyd, academic advisor to Receptiviti. LIWC-22 categories were among the theoretically grounded lexica evaluated alongside the primary models.

The relevance is evidentiary rather than topical. A measurement layer is only as good as the standard of proof behind it, and most language-based psychological assessment has never been tested the way this paper tests it: model frozen and registered in advance, evaluated on a later-in-time sample, then validated against criteria collected independently of the language itself. It cleared that bar, and the language-derived measure carried information that self-report did not. The same argument underlies the case for measuring what AI systems do to users from language rather than from self-report or from the model's own account.

bottom of page