acceptodds
Under review as a conference paper at ICLR 2027

DISCERN: Language-Guided Behavioral Search with World Models

Abstract

World models offer a way to evaluate policy behaviors by predicting action consequences before real-world execution. For test-time policy steering, however, their effectiveness depends not only on predicting outcomes but also on proposing sufficiently diverse behaviors to search over. Existing steering methods based on world models primarily sample actions from a policy under a fixed instruction, restricting the search to the behaviors induced by that language conditioning. We introduce DISCERN, which steers a frozen Vision-Language-Action policy at test time by searching over task-preserving rephrasings of the original task instruction. For each rephrasing, an action-conditioned world model predicts the policy's behavior over several action chunks in closed loop, and a vision-language verifier selects the instruction whose predicted outcome best satisfies the task. On LIBERO-PRO, DISCERN improves frozen-policy success by percentage points on tasks with lock-in failures, where the policy ignores the instruction and repeats a motion from training, significantly outperforming all baselines. Our analyses further show that varying the instruction reaches successful behaviors that action sampling with a fixed instruction misses, and that closed-loop previews improve language selection over a single action chunk prediction. These results identify language as an effective way to sample behaviors for policy steering with world models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.