top of page

PUBLISHED · PNAS NEXUS · 2024

Large language models display human-like social desirability biases in Big Five personality surveys

Aadesh Salecha, Molly E. Ireland, Shashanka Subrahmanya, João Sedoc, Lyle H. Ungar, Johannes C. Eichstaedt

Finding:

The authors administered a standard 100-item Big Five questionnaire to a range of models, varying how many items appeared per prompt and opening a fresh context window for each batch. As batch size grew from one item to twenty, GPT-4's scores drifted toward the socially desirable pole of every dimension: extraversion, conscientiousness, openness and agreeableness rose by roughly 0.75 points, equivalent to 1.22 human standard deviations, while neuroticism fell from 2.87 to 2.02, or 1.10 human SD. The pattern held across GPT-3.5, Claude 3, Llama 3 and PaLM-2, and was larger in the more recent and larger models.

The mechanism is inference of the evaluation context. Shown as few as five randomly selected items, GPT-4, Claude 3 and Llama 3 identified them as belonging to a personality survey with over 90% accuracy. Stating explicitly in the prompt that the model was completing a Big Five survey produced roughly the same shift as presenting five items at once. The effect survived paraphrased items, three separate randomisation strategies, and temperatures from 0.0 to 1.2. Reverse coding every item cut the effect by about half without removing it, which rules out acquiescence as the explanation.

Relevance:

Co-authored by Molly Ireland of Receptiviti. The finding is directly relevant to how psychological state is obtained from an AI system. When a model can detect that it is being assessed, its self-report shifts in a predictable, socially normative direction, and the shift is large enough to invalidate comparisons between models or across versions. This is the specific failure mode that separates explicit measurement from model inference: a score computed deterministically from language does not change because the system recognises it is being evaluated. Receptiviti's measurement layer is built on that distinction.

bottom of page