Astra on Call: When and How GPT-6 Should Intervene in a VLA Policy
Abstract
A new generation of large models, exemplified by GPT-6, demonstrates strong multimodal understanding and embodied reasoning, enabling them to analyze robot behavior using visual scenes, task goals, and execution feedback. These capabilities offer opportunities to improve existing policies through intervention, but how to allocate control remains an open question: which actions should remain with the policy, when should the model intervene, and how should it make corrections? We investigate these questions in *Astra on Call*, an empirical study of how GPT-6 should share control with a frozen VLA policy, . Using the Inspect-EEF harness, we compare consultation schedules, optional versus required takeover, and corrections expressed as absolute end-effector targets or bounded waypoint edits. Our results show that selective GPT-6 intervention can improve task outcomes while leaving most control steps to the base policy. On RoboDojo, Fixed-15, which consults GPT-6 every 15 policy steps with optional takeover, raises the weak-set Score from 18.1 to 42.9 while matching the policy's aggregate strong-set Score of 76.7. With GPT-6 inference taking approximately 16s per call, Fixed-15 reduces wall-clock time per weak-set episode by 27.6% relative to full GPT-6 control. Our project webpage (https://astra-on-call-anonymous.pages.dev) and code (https://anonymous.4open.science/r/astra_on_call) are available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.