Flow-based Policy Adaptation without Policy Updates
Abstract
Foundation models and other policies trained on large and diverse datasets can endow robots with semantic and task knowledge that would be difficult to acquire from task-specific demonstrations alone. Despite their broad scene understand- ing and reasoning capabilities, however, these models often lack the environment- and embodiment-specific knowledge required for precise and reliable execution. We propose GLOVES, a flow-based framework that complements an existing agent with a lightweight, goal-agnostic expert action prior that refines the action chunks proposed by the agent without updating its parameters. The prior captures observation-conditioned expert behavior, while the agent’s proposal supplies task- and context-specific intent that the correction aims to preserve. The learned flow also provides a natural in-distribution scoring mechanism through reverse flow evaluation. We use this signal as an intervention gate: actions that appear consis- tent with the expert distribution are passed through unchanged, while anomalous or out-of-distribution (OOD) actions are corrected. In this way, GLOVES only provides assistance when necessary. By learning local goal-agnostic expert ac- tion patterns and using them to adapt agent-proposed actions during execution, GLOVES provides a lightweight shared-control module for robust action adap- tation across tasks and environments. Across controlled OOD-recovery studies, simulation benchmarks, and real-robot experiments, GLOVES consistently im- proves the success of vision-language-action models (VLAs) and Code-as-Policy agents without retraining the underlying policy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.