acceptodds
Under review as a conference paper at ICLR 2027

WMAttack: Automated Attack Search for Adversarial Evaluation of World-Model Agents

Abstract

Assessing adversarial robustness in world-model agents requires costly, noisy closed-loop rollouts. We cast this task as a search for attack configurations against a fixed victim, subject to observation perturbation limits and a finite online budget. We present WMAttack, a closed-loop framework with two modules: Representation-Guided Attack Retrieval (RGAR) and Self-Correcting Attack Search (SCAS). RGAR matches internal and behavioral descriptors to stored attack records, retrieving configurations that give search an informed starting point. SCAS uses feedback from the current task to guide the construction, ranking, and selection of feasible candidates. An Auditor summarizes return degradation, policy-output disagreement, runtime, and return variability into structured advice. Fixed rules use this advice to build and rank candidates, and an LLM proposer selects from the resulting shortlist. A scout–confirm schedule screens candidates, confirms promising choices, and freezes the selected configuration for independent testing. Test results never feed back into search. On 26 Atari and 20 DeepMind Control tasks with DreamerV3, full WMAttack exceeds the highest-mean baseline, Successive Halving, by 0.138 and 0.133 in mean independent-test normalized return degradation, respectively, under matched online episode budgets; only WMAttack uses historical records. Controlled ablations support both modules' value, while further evaluations demonstrate applicability to three other world-model architectures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.