TA-IBAR: Task-Adaptive Inference Block Attention Residuals
Abstract
Attention residuals (AttnRes) improve information integration across network depth, but Full AttnRes incurs increasing memory and communication costs as historical representations accumulate. Block AttnRes reduces these costs by compressing consecutive layers into block-level sources, yet uses a fixed depth partition across tasks. We propose Task-Adaptive Inference Block Attention Residuals (TA-IBAR), which derives a task-specific contiguous partition from residual representations using centred Linear CKA and structurally constrained dynamic programming. The partition remains fixed during post-training, while the backbone and lightweight task-specific routing parameters are jointly optimised. At inference time, a lightweight probe selects the corresponding task-adapted configuration. Experiments span training-from-scratch and post-training settings, model sizes from 220M to 14B, and network depths from 28 to 40. Across Qwen3 models from 1.7B to 14B, TA-IBAR achieves the lowest validation loss in all eight Math and Multi-hop settings, including reductions from 0.4810 to 0.4629 on 14B Math and from 9.7513 to 8.9869 on 14B Multi-hop over fixed Block AttnRes. Across five additional model families, TA-IBAR outperforms fixed Block AttnRes in all 10 model–task settings and Full AttnRes in 8 of 10 while retaining the memory efficiency of blockwise residual access.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.