When Roleplaying, Do Models Believe What They Say?
Signal
78
Hype
15
In three linesStudy distinguishing what language models say from what they internally believe. Using linear truth probes on Claude, Qwen, and Llama role-playing historical personas, authors show persona adoption changes outputs more than internal truth representations. Contrasts with Emergent Misalignment where false claims shift toward true belief space.Read source
Your take?
Summary generated by Claude — human-verified