When to Warn, Not What to Do: Calibrated Language Warnings for Frozen World-Action Models
Abstract
World-action models (WAMs) combine future visual prediction and action generation for language-conditioned robot control. Collision safety requires anticipating unsafe contact and guiding responses that preserve task progress, yet collision records do not specify corrective actions. Learning recovery or language-feedback policies requires additional supervision or interaction. We propose risk-triggered language steering, which learns when to warn and guides a frozen WAM with a fixed prompt without corrective-action supervision. A classifier trained on collision-derived -step warning labels reads risk from predicted-video and action hidden states. Unsafe-class split conformal calibration sets a threshold targeting unsafe recall under the unmodified policy. Alerts trigger regeneration of the pending chunk with the warning appended to the task instruction. Experiments with LingBot-VA on seen tasks across four SafeLIBERO simulation suites show 89.3–99.0% AUROC and mean unsafe recall meeting nominal targets. Online, mean collision avoidance rate rises from 9.4% to 18.8%, versus 3.8% with the same warning always active. Mean task success is 30.0%, versus 29.4% without intervention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.