acceptodds
Under review as a conference paper at ICLR 2027

Gatekeeper Arena: Diagnosing Perception and Memory for Multimodal Policy Decisions

Abstract

Policy decisions can fail when relevant state is absent from the current encounter. We present Gatekeeper Arena, an offline testbed with 100 fictional entrants, 222 Unreal Engine clips, a disclosed nine-role policy, and a three-stage lifecycle: profile registration, unscored familiarization, and frozen-memory evaluation. Models return a fixed record; event-level balanced accuracy becomes the 0–100 Arena Score, accompanied by decision coverage. We evaluate the current- evidence setting with three open-weight VLMs across all five text/visual conditions, and ten hosted families join them on a shared text-plus-eight-frame leaderboard (13 models; 2,925 scheduled decisions). All cluster near chance (49.78– 54.01) and are strongly deny-biased. FlyWire-v783 recurrence retains delayed synthetic input, yet its Arena Score is 45.25 versus 50.53 for a degree-preserving rewire. Additional frames and generic recurrence therefore do not supply the profile and encounter state required by a longitudinal gatekeeper.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.