acceptodds
Under review as a conference paper at ICLR 2027

Starve the MLP, Feed the Attention: Resolving Vision-Knowledge Conflicts in VideoLLMs

Abstract

Modern video models have significantly advanced the field of video understanding. However, as these models generate responses, they overly depend on intrinsic textual priors rather than actual visual evidence, fundamentally restricting their performance. This over-reliance, called vision-knowledge conflicts, is especially evident in videos depicting counterfactual physics, where models hallucinate standard physics rules rather than processing the actual counterfactual visual input. Our mechanistic analyses reveal that as Transformer layers deepen, these language priors progressively dominate the residual stream, effectively overshadowing the visual tokens. To resolve this issue, we propose Vertex (Visual Evidence Restoration and Textual EXclusion), a zero-shot, inference-time intervention framework. Vertex first isolates physical motion to generate binary kinematic masks, which are directly utilized to construct distinct latent geometric subspaces. During inference, the framework actively rebalances the model by strengthening its focus on actual visual evidence while filtering out parametric textual biases. Evaluations on counterfactual physics benchmarks show that Vertex reduces vision-knowledge conflicts, outperforming standard baselines and providing a more robust framework for grounded video understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.