Back to feed
arXiv cs.AI·

FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games

Signal
78
Hype
15
In three linesFALSIFYBENCH evaluates inductive reasoning in 12 LLMs through rule discovery games inspired by the Wason 2-4-6 task. Reasoning models outperform instruction-tuned models, but none approach optimal performance. Success depends primarily on capacity for negative testing and hypothesis falsification.
Read source
Your take?
BenchmarksReasoningEvals

Summary generated by Claude — human-verified