Back to feed
arXiv cs.LG·

Self-CTRL: Self-Consistency Training with Reinforcement Learning

Signal
78
Hype
25
In three linesSelf-CTRL optimizes consistency between language models' self-explanations and behavior via reinforcement learning. On probabilistic reasoning tasks, the method improves R² correlation from 0.24 to 0.64. In constitutional AI, it increases refusal prediction from 36% to 92% and reduces HarmBench failure rate from 15.0% to 0.5%.
Read source
Your take?
Reinforcement learningAlignmentAI safetyEvals

Summary generated by Claude — human-verified