acceptodds
Under review as a conference paper at ICLR 2027

Correcting Strategy Selection in Generalist Robot Policies

Abstract

Large-scale robot datasets record multiple strategies, or ways of completing the same task: a microwave can be opened by pulling its handle or by pressing its release button, and a lid can be taken off by lifting its knob or by pushing its rim. A given object, however, often affords only one of them. We find that generalist policies trained on such data can carry out an unsupported strategy, one the object does not afford, such as reaching for a handle on a microwave that does not have one. The training data demonstrate the requested strategy about as often as or more often than the strategy the policy carries out, and stating the requested strategy in the instruction does not fix the failure. We therefore ask what drives the selection, and identify two features of the input that matter: the color of the object and the object nouns in the instruction. Recoloring the object to the most common color of the objects on which the training data demonstrate the requested strategy, and deleting object nouns that draw the policy's attention away from the part this strategy acts on, both shift the policy toward the requested strategy. We turn these findings into a method that decides for each input whether and how to edit it, without retraining the policy. Across three simulated tasks, the edits raise the average success of -DROID from 8% to 57%. Same edits also improve two other DROID-trained policies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.