Back to feed
arXiv cs.AI·

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

Signal
78
Hype
15
In three linesStudy evaluating 4 lie detectors across 31 models (2B-1T parameters). Detectors (CoT judge, logprob classifier, activation probes, DYL) perform well on prompted lying but fail on trained model organisms with verified beliefs. Only CoT judge maintains 0.82 balanced accuracy.
Read source
Your take?
EvalsReasoningAlignmentAI safety

Summary generated by Claude — human-verified