"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms
Signal
78
Hype
15
In three linesStudy evaluating 4 lie detectors across 31 models (2B-1T parameters). Detectors (CoT judge, logprob classifier, activation probes, DYL) perform well on prompted lying but fail on trained model organisms with verified beliefs. Only CoT judge maintains 0.82 balanced accuracy.Read source
Your take?
Summary generated by Claude — human-verified