acceptodds
Under review as a conference paper at ICLR 2027

DragForge: Procedural Evaluation of Drag for GUI Agents

Abstract

Drag actions combine object selection, a cursor trajectory, and release in a valid state. Endpoint grounding alone cannot evaluate interactions whose outcome depends on the intervening path. We introduce DragForge, a procedural benchmark with 18 task families and 60 configurations spanning placement, path-constrained manipulation, shape production, selection, and dynamic control. Policies observe screenshots and execute bounded waypoint sequences; evaluators measure the resulting task state. Privileged references solve all 180 instances used in the policy evaluation. In paired experiments on 117 development and 117 evaluation instances, replacing successful paths with straight segments while preserving their endpoints retains all 45 placement successes but only one of 72 successes in path-dependent families, in both splits. We evaluate five screenshot-conditioned vision-language policies on the same 180 instances under one protocol: four open-weight models and Claude Opus 5.5. Family-balanced task quality ranges from 6.8 to 70.1 on a 0–100 scale. Among the four open-weight models, quality rises descriptively with size, from 6.8 for a dense 27B model that fails even placement to 60.8 for a 2.8T-parameter model, although size is not isolated from training. The strongest policy, Claude Opus 5.5, solves every placement family but scores 78.4 on geometry and 32.0 on dynamic control. Recorded static prefixes improve with early calls and saturate as episodes terminate under native limits. In track, maze and curve tracing, replayed failures of the stronger policies usually press and release at the correct locations but violate the path. DragForge supplies executable tasks, state-based scoring, and controlled trajectory interventions for studying drag behavior beyond endpoint prediction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.