Back to feed
arXiv cs.AI·

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior

Signal
75
Hype
15
In three linesarXiv study comparing self-reports (Big 5 vs Theory of Planned Behavior) with actual behavior of 11 frontier LLMs across 4 tasks. TPB reaches human-level coherence within shared conversation but fails across separate conversations. Broad personality frameworks poorly predict real behavior.
Read source
Your take?
EvalsAI safetyAlignment

Summary generated by Claude — human-verified