When the Display Is Wrong: Concept-Dependent Compensation in World-Model Agents
Abstract
A world-model agent may retain task information when its observations become misleading, but whether that information sustains control requires a causal test. In Craftax-Classic, we falsify inventory counts without changing actual resources. We compare four -IRIS agents, whose policies receive frames, with ten DreamerV3 agents, whose policies receive recurrent and stochastic world-model states; all were trained with visible inventory. A false zero stone count reduces stone-related achievement rates from to in -IRIS, but only from to in DreamerV3. This preserved performance motivates a within-agent test of recurrent support. Removing a readout-defined recurrent stone subspace has little effect with an accurate display but substantially increases impairment when the display is false. Replacing that component with one from a recorded donor state largely restores performance when the donor's resource status matches the agent's actual inventory; wrong-status donors do not. The correct-donor advantage grows from percentage points with an accurate display to with a false display, an increase of points (Holm-corrected ). Wood interventions are individually costly, and newly trained probes still decode resource availability after projection, limiting generalisation across resources and ruling out complete information erasure. However, correct- and wrong-status patches differ both in the magnitude of the recurrent-state change they induce and in their contextual similarity to the receiver, so restoration cannot be attributed solely to resource information. The results support partial recurrent compensation for misleading sensory input in the tested DreamerV3 agents, while the wood comparison shows that this capacity cannot be assumed for every resource represented by the same world model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.