acceptodds
Under review as a conference paper at ICLR 2027

EviHarness: Evidence-Grounded Learnable Harness for Long Video Understanding

Abstract

Long-video question answering must locate sparse, answer-critical evidence under a limited observation budget. Existing selective-acquisition methods allocate observations to find relevant intervals, but compressing them can discard details needed to answer; timestamp-based memory points to a location without identifying which unsupported claim needs further evidence. We address this gap by linking compact claims to executable source pointers and using verification diagnoses to guide targeted re-observation. We introduce EviHarness, a learnable harness with raw-addressable evidence records, a frozen verifier that diagnoses missing or conflicting support, and a budget-aware controller that searches, revisits sources, or answers. We train only the controller: supervised fine-tuning (SFT) learns from valid, correct, low-cost trajectories, and GRPO optimizes answer correctness and acquisition cost while observation and verification components remain fixed. With Qwen2.5-VL-7B under shared inference caps, SFT+RL improves over prompting on Video-MME, LVBench, and MLVU; on LVBench, independently audited source support is 77.1% versus 61.4% for timestamp-only memory.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.