Back to feed
arXiv cs.CL·

Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark

Signal
78
Hype
15
In three linesarXiv study measuring VLM reliance on textual priors over visual content. 540-image benchmark with 4 question variants per image. 11 models tested: all degrade on hardest variant, open-source models drop furthest. No-image ablation reduces open models to 1–9% performance. GRPO post-training improves image-dependence across variants.
Read source
Your take?
VisionEvalsBenchmarksAlignment

Summary generated by Claude — human-verified