acceptodds
Under review as a conference paper at ICLR 2027

Calibrating Perception and Temporal Trajectories for Test-Time Latent Alignment in Multimodal LLMs

Abstract

While existing MLLMs enhance logical reasoning, they still suffer from severe confidence miscalibration. Moreover, current explicit and implicit reasoning mechanisms are constrained by discrete tokens and autoregressive mapping, which easily induces "premature semantic collapse" and "tunnel vision," failing to meet the holistic and hierarchical demands of visual perception. To address these challenges, we propose CHARM, a progressive three-stage architecture: First, Metacognitive Visual Learning (MVL) uses reinforcement learning to jointly optimize accuracy, calibration, and sensitivity, thereby aligning perception with confidence. Second, Chronological Trajectory Refinement (CTR) introduces temporal perturbation penalties to reinforce explicit temporal reasoning and keyframe grounding. Finally, Multi-Horizon Latent Alignment (MHLA) operates during test-time inference, leveraging pre-calibrated uncertainty to dynamically mix future multi-step semantic superposition distributions with hard labels, optimizing latent reasoning to decode the final explicit textual answer. Built on Qwen2.5-VL-7B, our method achieves 80.5% (OE) and 87.3% (MC) on Math-Vista, as well as 72.8% (STEM) and 80.6% (HASS) on MMMU. When scaled to Qwen3-VL-8B-Instruct, it reaches 69.7 ([email protected]) on the ActivityNet benchmark and sets new state-of-the-art records on QVHighlights and TVGBench with 90.3 and 65.1 ([email protected]).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.