ReCompose: Diagnosing and Improving Conversion to Sparse Transformer-SSM Hybrids
Abstract
Converting a pretrained Transformer into a sparse hybrid replaces most attention modules with state-space models (SSMs), retaining only a few attention layers. We present **ReCompose**, a diagnosis-guided approach to recovering capabilities lost during conversion. We examine how replacement supervision, attention-layer selection, and joint training affect capability loss and recovery within existing conversion pipelines. The repair trains SSM replacements and their output projections to match complete teacher-layer outputs. It then jointly trains consecutive student layers to match teacher block outputs before whole-model distillation. The local and blockwise stages share a design choice: match the output passed onward while allowing the trained components to adapt within that computation. Controlled replacement experiments support the combined local repair through improved downstream accuracy. In a same-start Qwen comparison at matched post-assembly data exposure, blockwise-then-global training yields 0.6–1.0 percentage points higher MMLU accuracy, while direct global distillation yields higher Lambada accuracy at the tested endpoints. MMLU-selected layouts retain higher MMLU accuracy and weaker retrieval than retrieval-selected layouts. Subsequent training partly recovers lost capabilities with the selected attention layers held fixed. On Qwen and Llama, ReCompose recovers performance across knowledge, commonsense, and text-completion benchmarks while retaining only four attention layers. These results inform how to supervise replacements, select retained layers, and allocate training to recover the intended capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.