Learning GUI Grounding with Spatial Reasoning from Visual Feedback

By Yu Zhao, Wei Ning Chen, Huseyin A. Inan, Samuel Kessler, Lu Wang, Lukas Wutschitz, Fangkai Yang, Chaoyun Zhang, Pasquale Minervini, Saravan Rajmohan, and Robert Sim, July 6, 2026

In ICML 2026

GUI grounding is usually posed as one-shot coordinate prediction, but vision-language models often struggle to map instructions to precise locations in high-resolution, complex interfaces. This work instead frames grounding as an interactive search: the model moves a rendered cursor, observes where it landed, and progressively approaches the requested element.

GUI-Cursor is trained with multi-step online reinforcement learning and a dense trajectory-based reward. At each step, the model identifies the target, reasons about its spatial relation to the cursor, and conditions the next movement on the interaction history.

Experiments on GUI grounding and agentic tasks show improvements over strong baselines using the same base models and less training data. The learned policy also adapts its search length to difficulty and transfers its spatial reasoning more effectively to out-of-distribution domains.

Paper: https://arxiv.org/abs/2509.21552

Stay ahead with research-backed solutions

From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.

Book a Consultation