acceptodds
Under review as a conference paper at ICLR 2027

D-model and D-Module: Identifying and Eliminating Model-, Task-, and Request-Level Transformer Redundancy

Abstract

Which computation in a trained Transformer can be removed while improving quality and reducing inference cost? We study this question under a fixed local compensation budget. Dynamic H couples D-model, which predicts the gains of a trained replacement, with D-Module, a causal rank-96 bridge trained through the frozen model’s suffix. The same mechanism supports model-, task-, and request-level reuse; only request-level recognition incurs per-request cost. Twenty of 24 fixed crossings across three Qwen backbones improve observed mean task CE, accuracy, and model-path time. A matched control improves quality without removing layers but adds latency, separating adaptation’s quality benefit from removal’s execution benefit. The same training and selection procedure produces joint-positive actions on Qwen3, SmolLM2, and Pythia. Automatic task selection and request routing also retain joint gains. Shared-checkpoint model reuse improves mixture CE and time at unchanged accuracy, although task effects oppose each other. These findings establish useful compensated removals and procedural transfer across model families. They also distinguish bridge recoverability from selector quality: the learned selector finds usable actions, but composite-utility superiority over geometric selection is not established.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.