acceptodds
Under review as a conference paper at ICLR 2027

BEPA: Agent-Guided Evidence Policy Adaptation for Video Question Answering

Abstract

Under a limited frame budget, question-relevant observations may remain compatible with multiple answers, while shared scene content obscures discriminative cues. We propose BEPA (Agent-Guided Evidence Policy Adaptation) to organize evidence around inter-option distinguishability and signal–background separability with a frozen video answerer. Before answering, CORE (Contrastive Option-guided Refinement of Evidence) suppresses shared option responses and refines a few BOLT fill frames. In an independent post-answer setting, a text agent turns prediction disagreement into visual queries for one additional observation. On 4,703 VUDG questions across 11 domains with VideoLLaMA3-7B, exploratory evaluation yields 75.82% domain-averaged accuracy (D-Avg) for CORE, versus 75.16% for prompt-matched BOLT, using one answer and at most 32 frames. Independent agent-guided acquisition reaches 77.26% D-Avg with the original prompts at 2.52 video answers per question on average.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.