ReBound: Readout-Aware Boundary Pruning for Efficient Mamba Prefill
Abstract
Mamba language models make autoregressive decoding efficient, but prefill still processes every prompt token through every layer. We propose \methodname, a training-free readout-aware boundary pruning method that runs the selected Mamba layer at full length and then removes low-impact boundary representations before downstream layers. Its score combines channel-mean , selective-update energy, and proximity to the final autoregressive readout, while a calibration-only KL sweep selects the earliest behaviorally safe pruning layer. At 10% FLOP reduction on Mamba-1 2.8B, \methodname reaches 68.0% average lm-evaluation-harness accuracy, only 0.62 points below baseline, and raises WikiText-103 continuation perplexity by only 0.61%. A complete seven-way score ablation shows that the full score is best, and calibration generalizes across architectures, with sharp safe-layer transitions at L44 for Mamba-1 2.8B and L52 for Mamba-2 2.7B. On five LongBench tasks, \methodname preserves long-context quality substantially better than -only DTP on both architectures. Measured hardware performance realizes these savings: at L44 with 50% token keep, latency improves by and throughput by , close to the FLOP-based ceiling. These results show that Mamba token compression should preserve the final readout rather than rely on local token importance alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.