acceptodds
Under review as a conference paper at ICLR 2027

When to Warn, Not What to Do: Calibrated Language Warnings for Frozen World-Action Models

Abstract

World-action models (WAMs) combine future visual prediction and action generation for language-conditioned robot control. Collision safety requires anticipating unsafe contact and guiding responses that preserve task progress, yet collision records do not specify corrective actions. Learning recovery or language-feedback policies requires additional supervision or interaction. We propose risk-triggered language steering, which learns when to warn and guides a frozen WAM with a fixed prompt without corrective-action supervision. A classifier trained on collision-derived -step warning labels reads risk from predicted-video and action hidden states. Unsafe-class split conformal calibration sets a threshold targeting unsafe recall under the unmodified policy. Alerts trigger regeneration of the pending chunk with the warning appended to the task instruction. Experiments with LingBot-VA on seen tasks across four SafeLIBERO simulation suites show 89.3–99.0% AUROC and mean unsafe recall meeting nominal targets. Online, mean collision avoidance rate rises from 9.4% to 18.8%, versus 3.8% with the same warning always active. Mean task success is 30.0%, versus 29.4% without intervention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.