DistillScope: Benchmarking Flow-Map Distillation for Language Generation
Abstract
Distillation promises fast language generation, but a recipe's apparent advantage can depend on how quality and computation are measured. We introduce , a common training and evaluation framework for flow-map language-model distillation that examines how recipe comparisons change across quality criteria, training budgets, and inference settings. Semigroup distillation achieves lower external perplexity than Lagrangian distillation, while their ordering reverses under MAUVE. Denoising readout and length-extrapolation generation also favor different recipes. Accounting for measured training time reverses a recipe advantage observed at equal updates. Additional inference computation allows the same trained student to pass degeneration checks that it fails under tighter inference budgets. Initialization, objective ablations, and cross-environment comparisons delineate the scope of these findings. An exposure-matched sampling control and supporting analysis of composition and categorical readout clarify the distinctions between training allocation, map fidelity, and decoded quality. makes these dependencies explicit, supporting recipe choices tied to a stated quality target and computational constraint.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.