When Is One Reading Enough? Decision-Relevant Ambiguity in Constrained Reinforcement Learning
Abstract
Safety rules for AI agents, online services, and industrial equipment are usually written in natural language, yet a learned controller can obey a rule only once it is a precise cost that training can measure. One sentence, such as "drain a server when load stays high," often admits several plausible executable readings, each a different constraint. We cast this as constrained reinforcement learning (RL) with candidate cost functions and ask whether the controller can train under one reading or must obey them all. Obeying all is safe but can give up performance, while picking one can yield a controller another reading forbids. We propose ARROW (**A**mbiguity **R**elevance for **R**einforcement Learning **O**ver **W**ritten Rules), a test that runs before training and searches for a controller nearly optimal under one reading yet violating another. If none exists, training under that reading is safe for every candidate reading. Otherwise ARROW keeps the competing readings or asks the operator targeted questions, re-testing after each answer. The test is exact when the dynamics are known, gives a high-confidence certificate from a fixed log, and works with any constrained learner. On agent-service tasks, training under the reading ARROW certifies raises return by a per-learner median of 15.0 to 37.3 percentage points of the best return that obeys every reading, across six learners. On service-monitoring rules, four questions settle what to enforce for 99% of ambiguous rules, usually before the intended reading is known. On 182 real ventilator, robotics, and cyber-physical requirements, 87.4% keep several distinct readings after redundant ones are removed, so the question is routine.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.