Back to feed
Reddit r/LocalLLaMA·

Vista 9B/4B from inclusionAI

Signal
72
Hype
25
In three linesinclusionAI releases VISTA-9B and VISTA-4B, vision-language models built on Qwen 3.5 backbones for GUI grounding. Trained with VISTA (View-Consistent Self-Verified Training), they map screenshots + natural-language instructions to normalized click coordinates 0-1000. Uses GRPO with target-preserving views and self-verified cross-view anchoring.
Read source
Your take?
QwenVisionAI AgentsOpen source

Summary generated by Claude — human-verified