Safe Executions Can Hide Policy Failure: A Proposal-Level Audit of Hybrid Control in Synthetic Interviews
Abstract
Safe executions can conceal unsafe proposals or a policy that avoids useful actions. We audit these failure modes in a synthetic structured-interview benchmark using joint proposal, execution, intervention and utility traces. A common hybrid-action interface supports TD3-style and PPO controllers with memory, reliability conditioning, nuisance consistency and proposal-cost optimization. Across 60 fitted policies and 23,520 evaluation episodes, both full controllers avoid follow-ups in the nominal panel, while a transparent rule achieves higher utility. Component-stack effects reverse direction between algorithms. An auxiliary control isolates exploration-induced constraint pressure: training exploration raises an always-proceed controller's mean discounted proposal cost to 0.696, above its budget of 0.5, despite zero deployed proposal cost. Released checkpoints, turn traces and executable audits connect every outcome table to its source, with explicit denominators and paired subject-level uncertainty. The findings motivate evaluating executed safety jointly with proposal safety and useful activity, and distinguishing exploratory behavior costs from target-policy costs. The evidence concerns a controlled synthetic setting; real-audio transfer and clinical effectiveness remain untested.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.