acceptodds
Under review as a conference paper at ICLR 2027

Joint Latent Reasoning and Observation for Agentic Video Understanding

Abstract

Agentic video understanding requires models to acquire useful visual evidence through repeated tool interactions. Distilling this capability into smaller models requires transferring not only answer prediction, but also the teacher's evidence-acquisition strategy. Conventional distillation relies on textual chain-of-thought and verbalized observations, incurring autoregressive generation costs while potentially omitting visual details needed for subsequent decisions. We introduce Latent Agentic Reasoning for Video Analysis (LARVA), a framework that jointly learns latent reasoning and latent visual observations. A reasoner selects acquisition actions and produces answers, while an observer communicates selected visual evidence through compact continuous representations. We train LARVA on teacher trajectories generated from VideoMarathon and evaluate it on VideoHolmes, Video-MME, and LVBench. Our evaluations show significant advantages over directly distilled CoT models and recent agentic video-understanding methods. Trajectory analysis suggests that LARVA more closely follows the teacher's progression from broad inspection to targeted acquisition than a CoT-and-caption student. These observations motivate investigating joint latent distillation as a means of transferring effective evidence-acquisition strategies for video understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.