acceptodds
Under review as a conference paper at ICLR 2027

ReCompose: Diagnosing and Improving Conversion to Sparse Transformer-SSM Hybrids

Abstract

Converting a pretrained Transformer into a sparse hybrid replaces most attention modules with state-space models (SSMs), retaining only a few attention layers. We present **ReCompose**, a diagnosis-guided approach to recovering capabilities lost during conversion. We examine how replacement supervision, attention-layer selection, and joint training affect capability loss and recovery within existing conversion pipelines. The repair trains SSM replacements and their output projections to match complete teacher-layer outputs. It then jointly trains consecutive student layers to match teacher block outputs before whole-model distillation. The local and blockwise stages share a design choice: match the output passed onward while allowing the trained components to adapt within that computation. Controlled replacement experiments support the combined local repair through improved downstream accuracy. In a same-start Qwen comparison at matched post-assembly data exposure, blockwise-then-global training yields 0.6–1.0 percentage points higher MMLU accuracy, while direct global distillation yields higher Lambada accuracy at the tested endpoints. MMLU-selected layouts retain higher MMLU accuracy and weaker retrieval than retrieval-selected layouts. Subsequent training partly recovers lost capabilities with the selected attention layers held fixed. On Qwen and Llama, ReCompose recovers performance across knowledge, commonsense, and text-completion benchmarks while retaining only four attention layers. These results inform how to supervise replacements, select retained layers, and allocate training to recover the intended capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.