acceptodds
Under review as a conference paper at ICLR 2027

Evolving Gaming Harnesses via Knowledge–Probing Co-Evolution

Abstract

Long-horizon gaming requires agents to diagnose stalled execution, recover from failures, and reuse experience across tasks. Existing execution harnesses typically rely on fixed intervention rules, limiting their ability to adapt to visually ambiguous failures. We introduce Knowledge–Probing Co-Evolution Harness (), a framework that improves the execution harness surrounding a frozen vision-language-action model through interaction. separates fast action execution from a slow research loop that diagnoses failures, proposes harness modifications, and verifies them through isolated paired sandbox evaluations. Its central mechanism is knowledge–probing co-evolution, which turns outcome-driven harness search into evidence-guided evolution: transferable recovery knowledge provides intervention hypotheses, active probes test when they apply, and validated outcomes improve future knowledge and probe selection. We instantiate with a gaming policy trained through supervised fine-tuning and reinforcement learning, and keep its parameters frozen throughout harness adaptation. Experiments on 149 Minecraft tasks spanning mining, combat, and crafting show that knowledge-guided harness evolution substantially improves task coverage over both the bare actor and a manually specified fixed harness, while the complete framework further improves over knowledge-only evolution. Frozen evaluation on unseen tasks further shows that the evolved harness transfers beyond tasks encountered during adaptation. These results support external harness evolution as a means of converting grounded interaction feedback into transferable execution capability without further updating the underlying actor. The project page is available at https://gaming-harness.github.io/adaptive-gaming-harnesses/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.