acceptodds
Under review as a conference paper at ICLR 2027

HieraSeek: Learning to Search Visual Hierarchies for Long-Video Reasoning

Abstract

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in video understanding, yet long-video reasoning remains challenging because task-relevant evidence is sparse and easily obscured by redundant visual content. Existing methods typically compress long videos into fixed representations or iteratively search over flat temporal sequences, which can limit effective evidence discovery. We propose **HieraSeek**, a hierarchical framework that constructs a reusable multi-granularity visual representation directly within the ViT encoder through progressive frame pruning and temporal merging, and exposes this representation as the search space for multi-round reasoning. This shared latent structure tightly couples representation construction with evidence acquisition, avoiding a separate retrieval or memory pipeline during reasoning. A learned policy iteratively refines or relocates its temporal focus, while reusing intermediate visual states to reduce repeated computation. We further optimize the search policy with an evidence-seeking reinforcement learning objective. Across six long-video understanding and reasoning benchmarks, HieraSeek achieves the best performance on all reported metrics, outperforming strong multi-round baselines under the same visual budget while maintaining competitive inference efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.