Data-Anchored Dynamic Routing for One-Step Offline Reinforcement Learning
Abstract
One-step actors keep inference cheap in offline reinforcement learning, but still have to improve under a critic without moving too far from the data. In teacher–student methods such as FQL, the same student output is asked to do both jobs: move toward higher Q and stay near its paired teacher endpoint. When these directions disagree, the update is a compromise on that sample, even if another output could represent the target. We propose DROL, which replaces this teacher–student pairing with dynamic routing and trains the actor directly from offline data. For each state–action pair, the actor samples candidates and updates only the nearest one with behavior cloning and critic guidance. As the actor changes, another candidate can take over the reconstruction target, allowing different latent regions to specialize without fixing their assignments in advance. Multiple candidates are needed only during training; inference uses a single actor pass. Our local analysis explains this assignment flexibility, and diagnostics track it in trained policies. On OGBench and D4RL, DROL is competitive with one-step baselines, with multi-candidate training improving performance across manipulation and puzzle task families.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.