acceptodds
Under review as a conference paper at ICLR 2027

Editing What the Video Contributes: Visual Intervention Meets Fisher Geometry

Abstract

Successful editing in Video-LLM can mask Visual Blindness: the correction may rely on text-side shortcuts rather than videos. We study this ambiguity through a structure-preserving Visual-Null intervention, revealing that substantial editing gains persist without informative video. We translate this behavioral contrast into Fisher geometry to isolate editing directions beyond the blind response. The resulting Fisher-orthogonal residual defines video-grounded editability, capturing both separation from the Visual-Null response and editing target controllability. This formulation yields VIGOR (Visual-Intervention Guided Orthogonal Residualization), a closed-form, minimum-information update that improves the video-conditioned objective while preserving the Visual-Null objective to first order. We prove that greater video-grounded editability lowers the minimum functional cost of a prescribed target improvement, informing principled layer selection within the locate-then-edit paradigm. Across four Video-LLMs on VMEB, VIGOR achieves the highest average editing score among evaluated methods, combining near-perfect reliability with near-zero Visual-Null gains. It also largely preserves general video understanding and representation structure after 100 sequential edits. Our code is provided in the supplementary materials; all data will be released soon on Github.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.