ECR: Evidence-Certified Revision for Self-Correcting Long-Video Agents
Abstract
Long-video agents increasingly answer questions by actively searching for task-relevant evidence. Once an answer has been formed, a separate decision arises: when does new evidence justify revising it? Evidence that supports an alternative answer does not necessarily refute the current one. We introduce ECR (Evidence-Certified Revision), a training-free framework for post-answer revision. ECR keeps the current answer as an anchor and treats a new answer as a challenger. It separately evaluates challenger support, anchor refutation, and evidence admissibility, records these judgments in a structured certificate, and applies a fixed reason-conditioned policy to retain the anchor, switch to the challenger, or escalate selected residual conflicts to blind pairwise arbitration that does not reveal anchor identity. On Video-MME Long, the full evaluation pipeline improves a frozen AVP agent from 52.11% to 62.33% accuracy while preserving 94.9% of initially correct answers. On a 655-question exact-replay cohort with frozen base executions and proposal outputs, ECR reaches 65.50% accuracy with 99 fixes and 13 breaks, compared with 65.19%, 141 fixes, and 57 breaks under unguarded proposal adoption. Under matched evidence access and nearly matched inference budgets, ECR-K2 reaches 55.29% accuracy versus 52.16% for a compute-matched symmetric selector, with more fixes and fewer breaks. Across five heterogeneous base agents on a shared 128-question cohort, ECR yields positive point estimates in all five settings. These results support treating the decision to revise as distinct from the process that produces complementary evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.