Back to feed
arXiv cs.LG·

LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?")

Signal
78
Hype
25
In three linesLEVANTE-bench benchmarks vision-language models against children aged 5-12 (N=1547) on cognitive tasks across 3 countries. Larger models show better overall alignment, but smaller models match younger children's error patterns better. VLMs struggle on matrix reasoning and mental rotation tasks.
Read source
Your take?
VisionBenchmarksEvals

Summary generated by Claude — human-verified