acceptodds
Under review as a conference paper at ICLR 2027

JEPA-VA: JEPA Is a Good Visual Action Verifier for Generalizable Robot Policies

Abstract

Vision-language-action (VLA) policies are commonly trained to align predicted actions with demonstrations, while world-action models (WAMs) jointly predict future video and action; neither explicitly checks whether an action produces its intended visual consequence. We study this complementary training signal: does a policy's proposed action lead to the expected visual change? We introduce \method, a JEPA visual verifier that evaluates proposed action chunks through their predicted feature changes relative to demonstrated transitions. Conditioned on the current observation and an action chunk, the verifier predicts the feature change at the end of the chunk. During policy training, a change-weighted cosine objective aligns the predicted change direction with the demonstrated one, sending a consequence-consistency gradient through the verifier's action input back to the policy. The verifier is not needed at deployment and adds no inference cost. Verifier-guided policy learning raises LIBERO-Plus success from 62.68% to 71.73% and clean-to-randomized RoboTwin success from 12.00% to 17.80%, suggesting that JEPA-VA helps policies learn more generalizable robot manipulation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.