Risk Bounds for Physical AI: Decision Ambiguity and Model Incompleteness
Abstract
Physical AI safety depends not only on what a system perceives, but on what its action will do and whether a response can still work in time. From classical decision, identification, and testing results we compose one risk bound for a single irreversible decision that separately budgets the loss of continuing while the consequence model is valid, of continuing after it has silently failed, and of responding, including responses that fail or come too late. A hand-constructed finite example computes every term of the bound, and a synthetic model instantiates its monitoring term. A MetaDrive study, which does not estimate the bound, asks whether detect-and-brake monitors pass a preset intervention screen in intact driving and how responses change repaired and new failures. A hidden fault that scales braking commands to 70% produces 21 failures in 120 drives selected for nominal success. None of twelve detect-and-brake candidates frozen before the audit passes the screen: the least-triggering ones intervene in 3 of 24 intact scenes where at most 2 were allowed; with 24 scenes this fails the screen but cannot show that the true rate exceeds 10%. Offline re-scoring with a post-hoc predictor whose one-step error is 37.7% lower detects more faults, but its best variants still cross the threshold in 3 of 24 intact scenes. In post-hoc same-scene comparisons, compensation with the true fault gain, started at the known onset in fault drives only, repairs 15 of 21 baseline failures and introduces two; started by an onset-aligned trigger in every drive, it leaves 16 fault failures and causes 14 intact-drive failures. Three braking responses on that trigger each introduce more failures than they repair in point estimates (paired intervals include or touch zero). Together, the analysis and experiments motivate treating information sufficiency, consequence validity, and response execution as separate assurance obligations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.