CAFE-VLA: Controllable Avoidance with Field Encoding for Compact Vision-Language-Action Models
Abstract
We present CAFE-VLA (Controllable Avoidance with Field Encoding), a controllable obstacle-avoidance framework for compact vision-language-action (VLA) models. The CAFE-VLA policy is conditioned on two complementary inputs: an avoidance-strength scalar and an end-effector-centric 3D geometry field. The avoidance-strength scalar specifies the requested degree of avoidance, while the 3D geometry field represents local obstacle distances and clearance directions at a fixed set of probes. We inject the scalar condition into the action expert via Feature-wise Linear Modulation (FiLM), while the 3D geometry field conditions the policy through cross-attention to field tokens and visual prompting encoding obstacle distances. This structure-aware design only increases the parameter count by less than 5%. A sweep over the avoidance-strength scalar confirms monotonic control over obstacle clearance (Spearman ρ = 0.953), while classifier-free guidance enables stable extrapolation to a strength 50% above the training maximum. On risk-stratified SafeLIBERO, CAFE-VLA achieves the highest task success rate at every risk level, with an average success rate of 74.32%, compared with 70.34% for the strongest safety-aware baseline. Meanwhile, CAFE-VLA retains a task success rate of 82.54% on the original LIBERO benchmark, the highest among the compared safety-aware methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.