acceptodds
Under review as a conference paper at ICLR 2027

Soft-Value Flow Matching for Offline Reinforcement Learning

Abstract

Flow-matching policies represent multimodal actions, but improving them with a learned critic in offline reinforcement learning faces two difficulties: backpropagating through multi-step sampling is costly and unstable, and critic queries on actor-generated actions can expose the policy to out-of-distribution value errors. We introduce Soft-Value Flow Matching (SVFM), which trains a multi-step actor with the gradient of a conditional soft value for KL-regularized policy improvement. Starting from noised data actions, SVFM samples stochastic continuations of a fixed behavior process and evaluates the critic at their endpoints, keeping the queries used for guidance independent of the current actor. A log-mean-exp of these values provides a Monte Carlo target for a soft-value network, whose gradient guides actor updates without trajectory backpropagation. We further propose an adaptive temperature that scales with the spread of critic values across continuations. SVFM achieves higher average success rates than the compared baselines on OGBench, LIBERO, and LIBERO-Plus. It also improves real-world manipulation with both task-specific flow policies and a pretrained VLA model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.