acceptodds
Under review as a conference paper at ICLR 2027

Limits of Statistical Assurance for Embodied AI under Adaptive Adversaries

Abstract

Embodied AI increasingly places language models, vision-language-action models, and learned controllers in physical systems, and their deployment rests on finite statistical assessment, presuming that inspected behavior persists once inspection ends. We show that a statistically valid certificate can become an attack signal for an adversary with continuing access. An adversary that preserves the unattacked system's joint observation distribution until observable, irrevocable certification and attacks afterward attains expected post-certification events over trials, where is the event probability under attack and the unattacked certification time. The bound holds for adaptive and privately randomized evaluators, regardless of policy architecture, and depends on the full certification-time distribution, so equal pass rates can conceal different exposure. For the fixed likelihood-ratio test, a pathwise coupling proves switching at certification exactly optimal under memoryless interventions that raise event probability. Requiring a fixed positive certification probability for unattacked systems forces linear worst-case expected excess over unmonitored deployment, even when certification is deferred to the assessment deadline. In simulated RoCo multi-robot collaboration, a switching controller records commanded separation at or below a calibrated threshold in 34 of 36 post-certification episodes. In preregistered randomized AgentDojo trials, attacks initiated after certification cause 1.98 additional security violations per four tasks (95% interval ). A language-model controller infers whether certification has occurred from noisy signals of evaluation status. Informative signals raise observed verified violations from 1.41 to 2.07 per trajectory, a 47% increase over uninformative signals. Fresh preregistered trials independently confirm the increase, yielding 0.93 additional violations per trajectory with no missing deployment outcomes. With certificate revocation, a mixture likelihood-ratio test bounds cumulative excess violations relative to a prespecified reference rate at square-root order, up to logarithmic factors, for every adaptive adversary. This bound holds on every sample path. Monitoring with a single fixed alternative can permit linear expected excess even under stationary attacks. Against the same controller given the monitor's exact rule and current statistic, live mixture monitoring reduces verified violations from 12.0 to 3.8 per 24-task deployment, a 68% reduction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.