Block-Level Recursion: Adaptive Test-Time Routing in Large Language Models
Abstract
Test-time routing improves frozen large language models (LLMs) by taking non-linear paths through their layers, without modifying weights or generating extra tokens. Existing approaches define route spaces that grow exponentially with depth, making them costly to search and hard to learn from. We therefore introduce (), a restricted route family that repeats a single contiguous block of transformer layers once. This reduces the number of routes from exponential to quadratic in the number of layers, making exhaustive per-instance evaluation tractable and the oracle upper bound directly measurable. Despite this restriction, yields large accuracy gains while keeping route selection practical. Across six model families and ten reasoning benchmarks, the optimal block varies across models, tasks, and individual inputs, with per-instance oracle gains of overall and up to on individual tasks. also supports two practical policies: a single train-selected block () that requires no router or per-input overhead, and a learned global router () trained from dense per-instance rewards over all routes. already improves accuracy substantially, while improves further by selecting routes per input. With a frozen Qwen2.5-0.5B backbone, achieves higher accuracy than the unrouted Qwen2.5-7B model at lower FLOPs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.