JEPA-VA: JEPA Is a Good Visual Action Verifier for Generalizable Robot Policies
Abstract
Vision-language-action (VLA) policies are commonly trained to align predicted actions with demonstrations, while world-action models (WAMs) jointly predict future video and action; neither explicitly checks whether an action produces its intended visual consequence. We study this complementary training signal: does a policy's proposed action lead to the expected visual change? We introduce \method, a JEPA visual verifier that evaluates proposed action chunks through their predicted feature changes relative to demonstrated transitions. Conditioned on the current observation and an action chunk, the verifier predicts the feature change at the end of the chunk. During policy training, a change-weighted cosine objective aligns the predicted change direction with the demonstrated one, sending a consequence-consistency gradient through the verifier's action input back to the policy. The verifier is not needed at deployment and adds no inference cost. Verifier-guided policy learning raises LIBERO-Plus success from 62.68% to 71.73% and clean-to-randomized RoboTwin success from 12.00% to 17.80%, suggesting that JEPA-VA helps policies learn more generalizable robot manipulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.