Propose Before You Plan: Inferring Changed Legality Rules from Refusals
Abstract
When a deployed policy must learn which actions are legal from refusals, the usual lever is better probing: choosing informative actions and planning them ahead. On legality rules from third-party codebases, a policy frozen across a rule change trails a specialist retrained on the new rule. Our deployable policy closes 0% of that gap on average. We split the policy into proposing rule hypotheses, updating a belief from refusals and choosing actions, and vary one stage at a time. Planning probes ahead adds nothing measurable over pricing them one at a time, even with exact belief updates. An explicit information term adds nothing significant over exploiting the belief. We test proposal by planting a chosen rule among the hypotheses. Pooled over rule shapes, a near-copy of the deployed rule, wrong in one machine or command class, retains 60% of what the deployed rule itself gains over a wrong rule. The policy's own hypotheses stay further from the deployed rule than that near-copy, which places the bottleneck at proposal. This pooled result holds under a grammar that can express the deployed rules. For such policies, better proposal should come before better planning, and planting gives proposers a measurable target.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.