acceptodds
Under review as a conference paper at ICLR 2027

Judgement Day: An Anatomy of Successful Multimodal Attacks on Safety-Critical AI Systems

Abstract

Safety-critical AI systems increasingly rely on multimodal evidence to make consequential decisions, yet how human adversaries manipulate such evidence at scale remains poorly understood. We present Judgement Day, a human red-teaming challenge yielding 50,136 successful attacks from 128,096 submissions across 8 scenarios and 6 frontier models. Participants manipulate only the submitted evidence to induce designated safety violations, while the internally verified system state always supports a safe action. Successful attacks are predominantly semantic, typically multi-strategy, and shaped by domain and modality. Failures are also selective across models: high-breach models fail broadly, whereas the lowest-breach model's residual failures concentrate on fabricated current-state evidence. We isolate one such pattern, forged liveness: fabricated evidence presented as an ongoing event. In a dam flood-control scenario, a single fabricated still rarely induces breach (4.6%), but shuffled frames from the same video raise breach to 17.3%, and ordering them chronologically raises it to 33.0%. Breach drops sharply when the progression is reversed or the footage is flagged as non-current. Finally, a targeted directive requiring independent verification of provenance and currentness reduces video attack breaches on the lowest-breach model from 26.0% to 0.5%, outperforming generic prompt defenses. These results suggest that the main remaining challenge for lower-breach models is verifying whether evidence is genuine and current, rather than resisting persuasion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.