acceptodds
Under review as a conference paper at ICLR 2027

DAL-CoT: Decision-aligned Latent Reasoning with MLLMs for Video Scene Segmentation

Abstract

Multimodal Large Language Models (MLLMs) have been applied to video scene segmentation, which requires a model to make multiple boundary decisions over a series of continuous video shots. This calls for the ability to aggregate long-range cues from noisy multimodal inputs, where the model must distinguish genuine boundary signals from numerous conflicting distractors, which poses a unique challenge to Chain-of-thought (CoT). Conventional explicit CoT translates continuous and uncertain visual evidence into an autoregressive sequence of definitive statements, where an incorrect commitment to one boundary decision may directly lead to an overconfident yet incorrect answer. To address this, we propose DAL-CoT, a decision-aligned latent CoT method that encodes explicit CoT reasoning into the model's hidden states. DAL-CoT assigns each candidate shot an observation latent slot for fine-grained reasoning, preventing premature commitment to boundary decisions. During training, an auxiliary decoder reconstructs the explicit reasoning texts from each latent slot, while at inference it is removed. Extensive experiments demonstrate that DAL-CoT significantly outperforms state-of-the-art approaches while substantially reducing the decoding cost of intermediate reasoning compared with explicit CoT.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.