Bellman-Aware Conservative Intervention: Recursive Actor-Flow Selection for Offline Reinforcement Learning
Abstract
Flow-based offline reinforcement learning can combine an expressive flow policy with an efficient one-step actor, yielding two candidate actions at each state. Selecting between them is not purely a local ranking problem because the value of a current action can depend on future actor–flow selections. We introduce Bellman-Aware Conservative Intervention (BACI), an offline actor-flow selection algorithm that keeps both action generators fixed and learns a dedicated decision critic. Its Bellman target recursively accounts for future actor–flow selections. At deployment, a flow proposal can replace the actor action when its estimated advantage exceeds a replacement cost, with accepted replacements executed with probability . Under explicit assumptions, we establish exact policy improvement within the restricted actor–flow policy class, identify conditions under which recursive selection captures gains missed by actor-policy evaluation, and characterize the effect of value-ranking errors under approximate deployment. Experiments with frozen actor–flow pairs trained by Guided Flow Policy (GFP) and Flow Actor-Critic (FAC) show average improvements over the corresponding pretrained actors on OGBench and D4RL. Fitted Q Evaluation (FQE) and original-critic controls show that recursive value learning provides additional benefits beyond direct ranking, with its advantage over FQE varying across benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.