acceptodds
Under review as a conference paper at ICLR 2027

Utility-Guided World Feedback for Visuotactile Manipulation

Abstract

World–action models jointly predict future observations and actions, but predictive accuracy alone does not determine which visual and tactile forecasts are useful for control. This gap is particularly important in contact-rich manipulation, where the utility of predicted evidence varies across sensors and time. We propose Utility-Guided World Feedback (UWF), a framework that learns to assess and selectively incorporate predicted evidence into action generation. UWF estimates the marginal utility of sensor–time evidence groups against a shared supervised action objective, converts utility-weighted evidence into action-chunk corrections, and regulates their adoption through sequence-level risk estimates. A shared action reference supports consistent evidence assessment, while joint world–action denoising updates feedback as predictions evolve. The framework learns from demonstrations without additional reward annotations or simulator rollouts for utility supervision. Across eight UniVTAC simulation tasks, UWF achieves a mean success rate of 90.1%. On eight real-robot tasks evaluated against four baselines with 20 trials per task and method, UWF achieves 91.3% mean success, exceeding the strongest baseline by 11.9 percentage points and obtaining six task wins and two ties. These results support utility-guided selection of predicted visuotactile evidence as an effective approach to contact-rich manipulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.