When to Wait, When to Intervene: Separating Hazard-Related Delays from Computational Congestion
Abstract
A delayed safety check may signal a hazardous situation or simply reflect computational overload. Confusing these causes can distort risk estimates and lead to poorly timed interventions. We propose a congestion-aware sequential decision framework for deciding whether to wait or intervene. A partially observable Markov decision process jointly represents uncertainty about task risk and computational load, combining elapsed checker response times, available verdicts, and lightweight workload measurements. The model captures dependence between checks affected by a common bottleneck. We compare fixed deadlines, independent-delay inference, telemetry-aware threshold rules, and joint belief-based decisions at matched information and resource budgets. Experiments combine measured checker runtimes under controlled computational load with replay tests that preserve tasks, verdicts, and each checker’s risk-conditional response-time distribution while varying cross-checker dependence. Outcomes include risk calibration, unsafe approvals, unnecessary interventions, and waiting time on held-out tasks and load conditions. The study tests when separating hazard-related delays from computational congestion improves decisions between waiting and intervention, and when simpler stopping rules suffice.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.