acceptodds
Under review as a conference paper at ICLR 2027

RC-DQN: A RISK-AWARE DUAL-CRITIC CONFORMAL DEEP Q-NETWORK FOR DEXMEDETOMIDINE DOSE TITRATION IN POSTOPERATIVE ICU PATIENTS

Abstract

Dexmedetomidine dose titration in postoperative intensive care unit patients is a sequential decision-making problem requiring a balance among prespecified treatment objectives, physiological tolerance, and action support. Applying offline reinforcement learning to this setting is complicated by the ordinal dose space, the scalarization of treatment utility and physiological cost, and extrapolation to actions that are poorly supported by retrospective data. We developed a Risk-Aware Dual-Critic Conformal Deep Q-Network (RC-DQN) using retrospective trajectories from two critical-care databases. Patient records were organized into six-hour decision windows, and dexmedetomidine exposure was represented by 16 ordered actions. RC-DQN combines an ordinal behavior-policy model, separate utility and physiological-cost-score critics, risk-modulated conformal calibration, an absolute cost-score shield, and rule-based fallback. The primary offline evaluation used terminal-outcome fitted Q evaluation (FQE) to estimate the mean initial-state value for the alive-without-delirium terminal outcome. An independent FQE evaluated returns under the prespecified composite reward, while action matching, ablation analysis, and bidirectional cross-database sensitivity analysis provided complementary assessments. Among the evaluated policies, RC-DQN produced the numerically highest within-database point estimates for both the alive-without-delirium endpoint and composite-reward return across both databases. Its endpoint estimates were higher than the corresponding observed test-cohort proportions, while action-matching analysis showed partial agreement with documented clinical decisions. Cross-database evaluation produced asymmetric results when source-trained policies and evaluators were applied to target-database states, indicating sensitivity to database-specific distributions. Overall, RC-DQN provides a structured approach for retrospectively learning and evaluating ordinal medication-titration policies under behavioral-support and physiological-cost constraints.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.