acceptodds
Under review as a conference paper at ICLR 2027

Safe Executions Can Hide Policy Failure: A Proposal-Level Audit of Hybrid Control in Synthetic Interviews

Abstract

Safe executions can conceal unsafe proposals or a policy that avoids useful actions. We audit these failure modes in a synthetic structured-interview benchmark using joint proposal, execution, intervention and utility traces. A common hybrid-action interface supports TD3-style and PPO controllers with memory, reliability conditioning, nuisance consistency and proposal-cost optimization. Across 60 fitted policies and 23,520 evaluation episodes, both full controllers avoid follow-ups in the nominal panel, while a transparent rule achieves higher utility. Component-stack effects reverse direction between algorithms. An auxiliary control isolates exploration-induced constraint pressure: training exploration raises an always-proceed controller's mean discounted proposal cost to 0.696, above its budget of 0.5, despite zero deployed proposal cost. Released checkpoints, turn traces and executable audits connect every outcome table to its source, with explicit denominators and paired subject-level uncertainty. The findings motivate evaluating executed safety jointly with proposal safety and useful activity, and distinguishing exploratory behavior costs from target-policy costs. The evidence concerns a controlled synthetic setting; real-audio transfer and clinical effectiveness remain untested.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.