acceptodds
Under review as a conference paper at ICLR 2027

ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies

Abstract

Behavior-cloned diffusion policies are expressive but remain vulnerable to covariate shift: small deviations from demonstrated states can compound into task failure. Existing methods address this either by expanding the training distribution through expert corrections or synthetic augmentation, or by steering a frozen policy at test time with guidance from a learned model. The former can be expensive or assumption-dependent, while the latter discards the corrected trajectories after execution. We introduce **ReGuide**, a self-improving framework that treats guided rollouts as reusable data. ReGuide first uses **Phase-Conditioned Guidance (PCG)** to generate corrective rollouts: it constructs phase-specific latent targets, applies guidance only in the drifted-but-recoverable regime, and guides through the estimated clean action. Successful guided rollouts are then absorbed back into the policy through **ReGuide-FT**, which fine-tunes the current checkpoint, or **ReGuide-FS**, which retrains from scratch on the augmented dataset; the two can also be composed and iterated. On Robomimic Can, Square, Transport, and Tool Hang, absorbing success-filtered rollouts lifts base-policy success by – from – demonstrations. At matched training-data volume, rollouts collected under PCG yield a better policy than unguided rollouts: by across three training seeds on Can, and at every rollout count on Square, so guidance improves the quality of self-collected data rather than merely its quantity. PCG's role is to generate training data rather than to improve a single rollout: at test time it matches LPB and DynaGuide on the three tasks where no guidance method separates from the base policy, and beats both on the one where any does. On LIBERO, fine-tuning a pretrained multi-task checkpoint on guided rollouts improves it on the three tasks it handles worst. Code is available at [anonymous.4open.science/r/reguide-anon-27B4](https://anonymous.4open.science/r/reguide-anon-27B4/README.md).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.