Back to feed
arXiv cs.CL·

When to Think Deeply: Inhibitory Deliberation for LLM Reasoning

Signal
72
Hype
18
In three linesIDPR is a framework that dynamically decides when to invoke slow reasoning in LLMs. An inhibition controller analyzes the fast response (confidence, logit margin, cost) and suppresses or validates before slow reasoning. On 5000 math examples, IDPR uses slow reasoning on 8.20% of cases and improves accuracy from 47.90% to 48.92%.
Read source
Your take?
ReasoningEvalsReinforcement learning

Summary generated by Claude — human-verified