Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark
Signal
78
Hype
15
In three linesarXiv study measuring VLM reliance on textual priors over visual content. 540-image benchmark with 4 question variants per image. 11 models tested: all degrade on hardest variant, open-source models drop furthest. No-image ablation reduces open models to 1–9% performance. GRPO post-training improves image-dependence across variants.Read source
Your take?
Summary generated by Claude — human-verified