Breaking the Historical Action Shortcut: Mitigating Visual Modality Collapse in Gaming VLA Models via Historical Action Noise Injection
Abstract
Interleaved multimodal inputs—streams of multi-frame images alternating with historical action tokens—have become a primary data format for Vision-Language-Action (VLA) models used in Embodied AI and gaming agents. Under Supervised Fine-Tuning (SFT), we identify a failure mode of this paradigm: because expert gameplay produces action sequences with strong temporal correlation and low conditional entropy, the autoregressive transformer tends to route gradient updates through the low-entropy historical pathway, which can starve the high-dimensional visual encoder—a phenomenon we term visual modality collapse. Unlike the classical copycat/inertia shortcut in unimodal behavioral cloning, this failure is (i) conditional—it persists even after perfect marginal class balancing because within any trajectory, and (ii) cross-modal—both modalities are read by the same causal transformer, so the shortcut manifests as gradient starvation of the visual pathway rather than a data-imbalance artifact. We provide a first-principles account of this collapse under standard simplifying assumptions from information-bottleneck theory and gradient dynamics, and introduce two directly measurable diagnostics—Modality Information Gain (IG) and KL Divergence Shift—that isolate which modality actually drives predictions. Building on this analysis, we propose Historical Action Noise Injection (HANI): with probability , historical action tokens are replaced by uniformly sampled tokens from the discrete action vocabulary, keeping the ground-truth target intact. This discrete, semantically mismatched signal is harder to recover from than continuous noise, which drives up the historical pathway's prediction error and routes larger corrective gradients back onto the visual encoder. On the extended MiniWorld WallGap benchmark with Qwen3-VL-8B-Instruct as backbone, HANI outperforms Standard SFT, NEFTune, and Modality Dropout on task success and modality diagnostics (: ; : ). Additional experiments on three visually richer 3D environments—ViZDoom (Deadly Corridor), Minecraft (MineRL Navigate), and RoboTHOR (ObjectNav)—suggest that the shortcut is not tied to a low-texture testbed, and that HANI's relative improvement is largest in the environment (Minecraft) whose expert data most closely matches the low-entropy regime our analysis targets. HANI is a single-stage, one-line data-loader intervention with zero inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.