acceptodds
Under review as a conference paper at ICLR 2027

ResponseScope: Measuring Finite-Response Distortion in Compressed Models

Abstract

Model compression is commonly evaluated using clean-task performance and static teacher–student agreement, but these criteria do not characterize responses to finite, task-valid input changes. We study three related questions: how to measure this behavioral drift, predict compression vulnerability before full recovery and final utility evaluation, and preserve response behavior at a fixed model footprint. ResponseScope (RS) is a teacher-normalized finite-response distortion measure. The evaluation spans six datasets, nine teacher–student architecture pairs, four compression families, three storage budgets, 540 recipe means, and 2,700 independently stored seed checkpoints. Combining teacher layer fragility with operator-local distortion and applicable initialization compatibility yields task-equal leave-onedataset-out R2 = .687 and MAE 1.800 percentage points for predicting utility damage prospectively. After compression, adding measured RS to a scale- and endpoint-matched baseline raises R2 from .587 to .778 and reduces selection regret from 1.260 to .688 points. At footprint ratio .25, IF-Reg reduces mean held-out RS from .440 to .257 and improves the task-specific utility composite by 3.017 points relative to augmented logit distillation. A baseline matched for endpoints, teacher scale, and margin explains part of this gain. These results connect the measurement, prediction, and preservation of interventional fidelity

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.