Back to feed
arXiv cs.AI·

VISTA: View-Consistent Self-Verified Training for GUI Grounding

Signal
78
Hype
15
In three linesVISTA proposes a GRPO-based fine-tuning method for GUI grounding. It generates multiple views of the same screen (crops preserving target elements) to create more robust comparison groups. On ScreenSpot-Pro, it improves Qwen3-VL 4B/8B/30B from 55.5/52.7/53.7 to 63.4/65.8/67.0.
Read source
Your take?
Reinforcement learningVisionBenchmarksQwen

Summary generated by Claude — human-verified