acceptodds
Under review as a conference paper at ICLR 2027

VidDeduce: Enhancing Long-form Video Understanding via Deductive Verification

Abstract

Long-form video understanding poses challenges in clue localization and logical reasoning. While agents enhance reasoning via adaptive region selection and evidence gathering, current paradigms either introduce query-agnostic redundancy via global text preprocessing, or neglect global context through localized iterative retrieval. These limitations lead agents to form misleading intermediate hypotheses, trapping them in unviable reasoning trajectories. To enable evidence-guided search refinement, we propose VidDeduce, a training-free multi-agent framework comprising a cognitive planner, visual observer, and a verifier based on Natural Language Inference (NLI). The planner integrates clue-relevant captions with prior verification feedback to update accumulated evidence and unresolved gaps, and formulates subsequent search strategies. Translating these strategies into actionable perception, the observer distills relevant captions and adaptively performs visual inspection to acquire observations. The verifier assesses new observations against the accumulated evidence and unresolved gaps via NLI, providing explicit feedback to guide continued investigation of useful clues or revise unproductive search directions, ensuring every exploratory step possesses a traceable logical rationale. Extensive experiments across four video understanding benchmarks demonstrate the effectiveness of VidDeduce, which outperforms strong agentic baselines including DVD and LensWalk on Video-MME, LongVideoBench, and EgoSchema.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.