CoVE-PPO: Change-of-Variables Exact-Density PPO for Flow Policies
Abstract
Flow models transport probability density through deterministic maps, creating expressive multi-modal distributions attractive for robotic manipulation. Computing the standard policy gradient requires differentiating through action log-density, which is available through change of variables. Prior methods avoid the exact density citing Jacobian-trace estimation noise, cost, or numerical instability. We revisit these objections, and show that change of variables is tractable and effective. Sharing trace estimation probes across the ratio makes the estimator's error scale with the Jacobian difference rather than the Jacobian, and as few as 1 probe matches the full trace's learning performance. Much of the instability comes from a stiff forward solve, and can be fixed by placing integration nodes within the solver's stability region. We introduce CoVE-PPO, which optimizes the change-of-variables density of the executed action, with no surrogate or augmented variable: in a general form for any velocity field, and in EDC, a parameterization whose divergence is a bounded forward network output giving stable solves and free divergence. Across state- and vision-based robot manipulation, we show our method matches or exceeds on-policy baselines at matched or higher sample efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.