Starve the MLP, Feed the Attention: Resolving Vision-Knowledge Conflicts in VideoLLMs
Abstract
Modern video models have significantly advanced the field of video understanding. However, as these models generate responses, they overly depend on intrinsic textual priors rather than actual visual evidence, fundamentally restricting their performance. This over-reliance, called vision-knowledge conflicts, is especially evident in videos depicting counterfactual physics, where models hallucinate standard physics rules rather than processing the actual counterfactual visual input. Our mechanistic analyses reveal that as Transformer layers deepen, these language priors progressively dominate the residual stream, effectively overshadowing the visual tokens. To resolve this issue, we propose Vertex (Visual Evidence Restoration and Textual EXclusion), a zero-shot, inference-time intervention framework. Vertex first isolates physical motion to generate binary kinematic masks, which are directly utilized to construct distinct latent geometric subspaces. During inference, the framework actively rebalances the model by strengthening its focus on actual visual evidence while filtering out parametric textual biases. Evaluations on counterfactual physics benchmarks show that Vertex reduces vision-knowledge conflicts, outperforming standard baselines and providing a more robust framework for grounded video understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.