acceptodds
Under review as a conference paper at ICLR 2027

TollGate: Safe Reset-Free Reinforcement Learning via Feasibility and Budget Gating

Abstract

Reinforcement learning typically relies on resets between attempts, which are readily available in simulation but costly in the real world. Prior work pairs a task-directed forward policy with a reset policy that learns to return the agent to the initial state, using estimated recoverability to determine handoffs. However, both task execution and recovery can incur safety costs, and within each attempt they draw on a single, shared safety budget. Recoverability therefore does not imply affordability: a state can remain recoverable while the remaining budget no longer covers the cost of recovery. We introduce TollGate, a budget-aware handoff framework that combines constrained learning of both policies with a dual-gate switching mechanism. An outcome-supervised feasibility gate predicts whether recovery will succeed, and a budget gate checks whether the predicted recovery cost would exhaust the remaining budget; either gate can trigger a handoff to the reset policy. On a grid world and on continuous-control tasks, TollGate trains both policies without demonstrations or a separately designed reset reward. TollGate learns the tasks, replaces the vast majority of manual resets, and drives safety costs below their thresholds. Together, these capabilities bring reinforcement learning a step closer to the real world, where neither manual resets nor unsafe exploration comes cheap.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.