acceptodds
Under review as a conference paper at ICLR 2027

Latent Depth Recursion Reveals Post-Training Dynamics in Multimodal LLMs

Abstract

Recent work has shown that frozen large language models (LLMs) can correct their own predictions by re-executing a block of their own layers. Prior methods exploit this recursive structure at inference time, either by searching for the right block to repeat or by training a router to select one. All of these evaluations target pretrained or instruction-tuned LLMs. Every deployed model in practice, however, undergoes a longer post-training pipeline that adds multimodal adaptation and reinforcement learning (RL). In this work, we investigate whether the recursive structure that these methods rely on survives this pipeline. To this end, we reframe the recursive capability as a measurable budget: the fraction of instances that the default forward pass gets wrong but that some recursion route gets right. We compute this budget by exhaustively evaluating every route in a small, closed recursion family. Tracking it across the training stages of several LLM and multimodal LLM (MLLM) families, we find that the budget is large in pretrained models and shrinks consistently as post-training proceeds. Both supervised fine-tuning (SFT) and RL absorb the recursive structure into the default forward pass, and their updates align with where recursion was previously helpful. As a result, little budget remains for test-time routing to exploit, which explains why existing methods fail on fully post-trained models. Instead of recovering recursion at test time, we propose recursion-aware post-training, which deliberately spends the remaining budget by adding a fixed per-category recursion route to the training forward pass. At the same parameter count, our scheme outperforms standard GSPO post-training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.