acceptodds
Under review as a conference paper at ICLR 2027

PIVLA: Physical-Interaction-aware Vision-Language-Action Policy for Force-Sensitive Manipulation

Abstract

Vision-Language-Action (VLA) models produce semantically and geometrically correct motions, yet the same motion can crush or drop an object depending on grasp force. Recent force-aware VLAs treat force as an input or control target but do not represent what interaction a given context calls for. We present our method, a VLA that learns an intermediate physical interaction representation: a dedicated token fusing vision, language, and contact observations, supervised by per-group safe-force distributions and recurring interaction prototypes, and used to directly guide a flow-matching action expert. We validate on two real-robot tasks transferring i) damage-prone produce and ii) appearance-identical cartons with disjoint safe grasp intervals, where the required grasp can only be resolved through contact. Our method outperforms force-agnostic and equal-input force-aware baselines, and ablations confirm the contribution of each component. Code and data will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.