Learning Flow Policies from Unary Feedback for Real-World Robot Manipulation
Abstract
In real-world robot manipulation, flow policies can perform complex tasks after supervised fine-tuning (SFT), yet still make recurring errors during deployment. Expanding the expert demonstration dataset to correct these errors can be costly. Deployment failures provide another source of information for policy refinement. For flow policies, however, direct optimization through action generation can produce unstable gradients, and action log-probabilities are costly to evaluate. We propose Flow-KTO, an offline post-training method for flow policies. It learns from success and failure feedback on recorded execution clips. Flow-KTO accounts for asynchronous inference latency to align execution feedback with the corresponding action trajectories. The method replaces direct computation of policy log-likelihood ratios with a surrogate, avoiding backpropagation through action generation. A policy drift estimate modulates the strength of feedback updates at each training step. Across four real-robot tasks, Flow-KTO reduces the average target-mistake recurrence rate by 33.3% relative to the strongest baseline, while improving average success rate by 13.3% and increasing task throughput by 17.7% on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.