acceptodds
Under review as a conference paper at ICLR 2027

Robot-Centric Scaffolding Advances the Frontier of VLM-Based Manipulation

Abstract

In 2024, as foundation models became able to handle real software tasks, coding scaffolds emerged and greatly accelerated AI coding. General-purpose vision-language models (VLMs) can now turn natural-language goals into useful manipulation without task-specific policies. Existing coding scaffolds complete some robot tasks but fail systematically, because they rely on three assumptions that hold only in digital environments: that actions can be checked afterward and undone, that targets are named exactly by symbols, and that habits formed on digital tasks carry over. We therefore propose Qoboter, a robot-centric scaffold that redesigns its loop, tools, and skills accordingly. Its loop plans task-level subgoals but commits only to the next local action, chosen from the latest observation. Its five core tools let the model select target pixels, then convert them into 3D points, check reachability before execution, and batch frequently co-occurring operations. Its skills guide the model to interpret occlusion, contact, and control error and to judge progress by observed evidence. On four simulators and a real robot, we weigh success against token and turn cost. With model weights fixed, Qoboter raises success on every frontier model at a similar token cost, from 75.8% to 90.0% on GPT-6-Astra and from 43.3% to 55.0% on Opus 5. The same configuration transfers unchanged to three open-weight models, raising their success by 5.0 to 8.4 points. These results suggest that scaffolding complements model scaling as a lever for reliable VLM-based manipulation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.