acceptodds
Under review as a conference paper at ICLR 2027

Grounding Force: Image-Aligned Contact and Force Prediction for Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models inherit semantic understanding from pretrained vision-language models and map observations and instructions directly to robot actions. The success of an action, however, also depends on the physical interaction, where the robot makes contact with an object and how it applies force. A VLA trained only to imitate actions never represents this physical interaction explicitly. Prior work trains the VLA to predict targets that describe the scene rather than the interaction, or measures the interaction with force or tactile sensors, which most robots and datasets lack. We present Grounding Force (GF), which trains a VLA to predict, alongside its next action, the contact and force that action will produce, as a contact map and force maps on the wrist-camera image. This supervision is used only in training, so the policy needs no force or tactile sensor at test time. Across four backbones (, , GR00T N1.6, and OpenVLA-OFT) on RoboCasa and RoboTwin 2.0, Grounding Force consistently improves the base VLAs, by up to 7.3% in success rate, and helps more than predicting future images, depth, or semantics. With pseudo-labels from a labeler trained only in simulation, it also improves by 13.0% on both real single-arm and bimanual robots. These results indicate that predicting the physical interaction in the image frame, without any sensor at test time, is an effective supervision signal for VLAs. Code and videos: https://groundingforce.github.io.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.