LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?")
Signal
78
Hype
25
In three linesLEVANTE-bench benchmarks vision-language models against children aged 5-12 (N=1547) on cognitive tasks across 3 countries. Larger models show better overall alignment, but smaller models match younger children's error patterns better. VLMs struggle on matrix reasoning and mental rotation tasks.Read source
Your take?
Summary generated by Claude — human-verified