acceptodds
Under review as a conference paper at ICLR 2027

Self-Supervised VLA Preference Optimisation via Temporally Shifted Guidance

Abstract

Vision-language-action (VLA) models learn diverse robotic behaviors from demonstrations, but improving them beyond imitation often requires costly sources of feedback. Preference optimization offers a direct and effective post-training strategy, yet obtaining useful action comparisons can still depend on human intervention, additional interactions with environments, or external evaluators. We introduce temporally shifted classifier-free guidance (TS-CFG) for flow-matching action policies, using their learned action structure to construct hard negatives from demonstrations. The central idea is to steer a demonstration's reconstruction through the model's conditional predictions, targeting hard negatives with subtle deviations from demonstrated behavior. TS-CFG adapts classifier-free guidance by combining predictions under the original context and a nearby temporal context from the same episode. Temporal proximity provides related conditioning, intended to avoid mixing markedly different action distributions while introducing a structured perturbation to generation. Treating each demonstration as preferred over its temporally shifted alternative provides a preference-learning signal using only existing demonstrations and the pretrained policy. This construction supports direct preference post-training without new environment interaction, human preference annotation, or auxiliary world-model or reward evaluation. The additional computation is confined to training, leaving the policy's deployment procedure unchanged. Extensive experiments on RoboTwin and RoboCasa show improvements in task success over the starting policies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.