Why Do Efficient Architectures Learn Expensive Reasoning?
Abstract
Hybrid architectures combining linear and softmax attention increasingly undergo reinforcement learning from verifiable rewards (RLVR), yet architecture selection still relies on pre-training metrics. We ask whether long, explicit state-tracking trajectories selected by correctness-only RL reflect architectural requirements or training-induced strategies. We study parameter-matched 126M and 325M attention, linear-attention, and hybrid models under matched pre-training, supervised fine-tuning, and RL on verifiable working-memory tasks. Several hybrids adopt per-step state summaries, increasing output length 2.7× and inference time 2.8–3.1×. Three controlled analyses separate selection from necessity. First, independently trained terse policies match or outperform long policies for several architecture–task pairs, but lose 5–16 percentage points for other hybrids. Second, corrupting summaries reduces accuracy by 36–61 percentage points, showing that dependence within a learned policy does not establish necessity across policies. Third, although summary trajectories are less accurate overall throughout training, they gain an advantage within the mixed-outcome groups that drive policy updates, reinforcing their selection. Stratifying advantages within each (prompt, style) removes this between-style term and prevents switching in all nine runs across two scales and two tasks, with accuracy within one percentage point of the summary-forbidden control. Experiments varying pre-training seeds and task difficulty support the same mechanism: bases with similar pre-training loss differ in state fidelity, and switching depends on terse-policy failures and the conditional advantage of summaries within the training band. These RL-selected trajectories reverse the hybrids’ deployment-cost advantage, raising inference time from 0.6× to 1.6–2.0× that of full attention. Our findings show that efficiency depends on the interaction between architecture and post-training strategy selection, which pre-training metrics alone fail to capture. We release generators, checkpoints, per-item records, and an online diagnostic that predicts switching.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.