acceptodds
Under review as a conference paper at ICLR 2027

BranchPO: Demonstration-Anchored Branching Preference Optimization for VLA Post-Training

Abstract

Pretrained vision–language–action policies can remain unreliable on multi-stage deployment tasks, while collecting expert corrections at learner-visited states is expensive and from-start autonomous rollouts can undersample later phases. We present BranchPO, a post-training framework that reuses one successful demonstration as a phase-indexed interaction scaffold. Semantic roots provide access to distinct task phases; autonomous sibling continuations are collected from each root and compared with cached expert segments. For flow-matching policies, BranchPO optimizes a root-balanced, reference-relative energy preference objective with explicit expert attraction, requiring neither a learned critic nor additional expert action labels. An idealized analysis characterizes root-local outcome contrast, heterogeneous phase corrections, and the comparison weights retained by the implemented loss. Across RoboCasa and real-robot manipulation tasks, BranchPO improves the SFT initialization on every evaluated task and generally outperforms AWR and trajectory-level DPO. Scaling from-start whole-trajectory preferences does not recover the benefit of rooted continuations, while a positive-only adaptation control explains only a small part of the gain. These results show that demonstration structure can guide both autonomous data collection and local credit assignment for effective VLA post-training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.