acceptodds
Under review as a conference paper at ICLR 2027

Lost in Adaptation: Layer-Selective Recovery of Temporal Reasoning in Video-Language Models

Abstract

Multimodal adaptation can weaken temporal reasoning (TR) in video-language models (VLMs) even when the visual evidence needed for an answer is correctly perceived. This weakened competence, however, remains partly recoverable from the model's paired pre-adaptation language checkpoint. We introduce MERIT, a gradient-free framework that selectively restores this competence across the language backbone while protecting temporal perception (TP). It assigns each self-attention layer a VLM-dominant or LLM-dominant interpolation and uses the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) to search the recipe space under a perception-constrained objective. Across three VLM families and five video benchmarks, MERIT improves reasoning-aligned performance in all 15 model–benchmark evaluations while preserving TP across all three Video-MME settings. Recipes selected using 232 diagnostic examples transfer to four unseen benchmarks, with relative gains up to 27.8%, and outperform full-backbone, uniform-attention, and alpha-tuned random-layer merging. Interventional masking shows that MERIT-selected layers are disproportionately important for TR, while frame-level attribution reveals increased reliance on temporally distributed, causally relevant evidence. These results show that adaptation-induced reasoning degradation can be repaired post hoc through selective reuse of pre-adaptation competence, without learning new parameters or adding inference overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.