FE-VLA: Force Encoder Vision-Language-Action Policies for Compliant Manipulation
Abstract
Vision-Language-Action (VLA) models have driven remarkable progress in general-purpose robotic manipulation, yet they remain largely kinematics-centric and struggle in contact-rich scenarios where excessive force causes catastrophic failure, such as crushing force-sensitive objects. We advance VLA capabilities by introducing a framework for force-conditioned manipulation that treats force awareness as both a sensory input and a perceptual output. Our approach integrates a lightweight fingertip force sensor into the VLA prefix via a force encoder, allowing the model to condition its actions on real-time force feedback. An auxiliary force-prediction head taps directly into the VLM's perceptual trunk to predict the maximum grasping force, enabling the model to anticipate an object's physical limits from vision and force information before contact is established. A reactive safety layer then intercepts motor commands and freezes the gripper when live sensor readings exceed the predicted threshold. Evaluated across five force-sensitive everyday objects, our approach reduces peak applied force by approximately 50%, while maintaining comparable task success
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.