Recurrent Clinical Safety Queues Reverse Multimodal Model Selection
Abstract
At every decision time, a recurrent clinical safety queue spends capacity on new cases and material updates but not on unresolved carryovers; deferral, dismissal, cooldown, and re-entry then determine the next admissible set. Completed-record sorting and next-admissible selection consequently need not commute. On 9,714 patient-episodes and 1,074 adjudicated events from three held-out trials, completed-record discrimination selects gated fusion (AUROC/AUPRC/PPV@top-5%: 0.90/0.47/65.5), while recurrent review selects ChronoGuard. Under the five-selection per-instant boundary, ChronoGuard recognizes events 1.1 days earlier than gated fusion (95% CI, 0.6–1.6), raises severe-event sensitivity by 0.06 (0.02–0.10) and selected-alert precision by 0.07 (0.03–0.11), and removes 0.31 false alerts per monitor-week (0.17–0.45). AdmitLock makes this reversal testable by locking the evidence filtration, scorer-specific admission, recurrent transition, capacity clock, and endpoint registry, then reporting both Pareto orders, selected-episode overlap, and paired queue effects. Only 12–16 of each model's 25 highest-ranked completed-record episodes occur among its first 25 recurrent selections. Against operational key risk indicators, ChronoGuard improves recognition by 3.5 days, severe sensitivity by 0.14, precision by 0.19, and false-alert burden by 0.95 per monitor-week; day-28 restricted mean recognition time falls from 11.1 to 7.4 days. An independently governed four-trial portfolio reproduces both point-estimate leaders after refitting every learned method, and a 228-monitor-week shadow evaluation preserves ChronoGuard's operating profile under live arrivals. Model selection for scarce clinical review should therefore target the stateful sequence that reaches experts, not a ranking formed after records close.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.