Beyond Positive Bias: Customizing and Correcting Generative Robot Policies via Negative Guidance
Abstract
Diffusion policies can represent multimodal robot manipulation behaviors, but deployment may introduce constraints that were not present during policy training. Extending negative guidance to robot control is nontrivial: guided action chunks are repeatedly executed in a closed loop, where excessive repulsion can compound over time, reduce task completion, or collapse the remaining valid behaviors. We extend Dynamic Negative Guidance to diffusion robot policies and develop a multi-target formulation for simultaneously avoiding multiple undesirable outcome modes. We instantiate the framework through two complementary applications. First, language-conditioned negative guidance provides a semantic interface for specifying single or multiple outcomes to avoid. Second, when undesirable behavior is identified through deployment feedback rather than language, an auxiliary diffusion policy trained on a small set of failure rollouts provides the negative guidance signal while the base-policy parameters remain fixed. We evaluate the framework across object picking, target placement, planar pushing, and precise charger insertion, including a real-robot manipulation experiment. As an additional generalization study, we further evaluate language-conditioned negative guidance with flow-matching policies on LIBERO. Results show that dynamic negative guidance suppresses specified undesirable outcomes while better preserving task completion and alternative valid behaviors than static negative guidance, and that failure-policy guidance improves task success across the evaluated deployment-time environment changes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.