LoadCalib: Cognitive-Load-Calibrated Chain-of-Thought Synthesis for Small Language Models
Abstract
Chain-of-thought (CoT) distillation from a large teacher into a small student is typically performed with a uniform verbosity: every problem receives the same style of reasoning trace, regardless of its difficulty or of the student's capacity. We show this one-size-fits-all supervision is measurably suboptimal and propose LoadCalib, which calibrates the cognitive load of teacher CoT to (i) the difficulty of each problem and (ii) the measured capacity of the target student. LoadCalib measures a capacity matrix (the student's accuracy on difficulty bin when distilled with verbosity tier ) on a held-out split, derives a per-difficulty policy with a sufficiency rule (the cheapest tier within of the best), and refines it with a short closed loop rewarded by the student's downstream accuracy. On GSM8K with Qwen2.5 students of 0.5B/1.5B/3B parameters, the oracle-difficulty variant (LoadCalib-oracle) reaches the best uniform tier's accuracy at 3B while using 16% fewer reasoning tokens ( points, not significant; the advantage over the shortest tier is larger and significant, points, paired bootstrap ), and trades 1.1–1.5 accuracy points for 10–12% token savings at 0.5B/1.5B (the oracle variant 1.1–1.7 points for 10–14%). Reverse-calibration and random-tier controls confirm the gains come from calibrated assignment rather than verbosity variance, the variant ordering survives a second training mix and seed changes, and only mid-capacity students show per-difficulty tier differentiation: larger students prefer shorter CoT on easy problems. The deployable learned-difficulty variant (LoadCalib) matches the oracle at 0.5B/1.5B and trails it by 1.2 points at 3B.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.