acceptodds
Under review as a conference paper at ICLR 2027

Calling Is Not Looking: Entropy-Resolved Evidence Acquisition for Video Agents

Abstract

MLLM-based video agents call tools to retrieve missing visual evidence, but calling a tool does not guarantee that the returned evidence is actually or properly used. We propose Entropy-Resolved Evidence Acquisition (ERA), which tracks the agent's top high-entropy tokens, where it is most uncertain how to continue, and locates the commitment point at which this uncertainty resolves. The position of this point relative to the tool calls reveals two failure modes of current recipes: agents trained with tool-calling rewards commit to an answer before the tool returns and overlook the evidence, while agents cold-started under LLM-as-a-Judge supervision answer after any return and overrely on it, even when the return is uninformative. ERA turns this diagnosis into a turn-level reward that scores where the agent commits and contrasts its reasoning under an uninformative and a decisive return. On tool-required questions from four video benchmarks, we swap the first return for an uninformative one and assert that it is exactly what the agent requested. Supervised fine-tuning on counterfactual re-seek demonstrations teaches the agent to re-seek, but this behavior collapses under such assertions, whereas RL with the ERA reward keeps re-seek tied to what the returned frames show, lowers the hallucination rate, and preserves accuracy under pressure. The reward also improves an existing agent, VideoZoomer. These results suggest that where an agent's high-entropy tokens fall relative to its tool calls reveals whether it uses the evidence it acquires.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.