Knowing When Not to Act: Latent No-Action Region Recovery Hidden in Neural Control
Abstract
Linear action costs induce structured inaction regions in which the optimal control is exactly zero. Neural policies can nevertheless achieve near-saturated objective values while blurring this regime structure into near-zero actions. Post-hoc magnitude thresholding can enforce zeros, but its induced geometry is heuristic, need not match the true switching region, and may degrade performance by suppressing beneficial interventions. We instead recover the latent action/inaction regime directly from a frozen neural controller, without learning a global value function, explicitly parameterizing a free boundary, or introducing a heuristic threshold on action magnitude. Local continuation gradients yield marginal-value ratios, whose known action-cost thresholds determine the regime labels; a Yosida map then realizes the recovered hold regime as exact zero while preserving the switching structure. Repeated same-time mapping further converts the recovered regime into boundary-directed deployment. Across linear-action-cost benchmarks, including high-dimensional dynamic portfolio choice, we show that structurally meaningful inaction cannot generally be inferred from raw action magnitude alone, and that recovering the correct switching regime can improve both interpretability and control performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.