Megrez: Reconstructing Intermediate Supervision for Ternary Looped Transformers
Abstract
Ternary looped transformers reuse the same quantized weights across recurrent calls, coupling the approximation of intermediate teacher states. We present Megrez, a calibration framework that adapts the supervision targets to the current student. We reconstruct teacher targets to minimize alignment error while requiring the teacher's unchanged native readout to return its original output on those targets. Each token's normalized teacher states are jointly transformed across recurrent calls by a single orthogonal map. The constrained alignment has an exact low-dimensional solution. Megrez alternates target reconstruction with shared ternary weight updates, accommodating gated and terminal readouts without changing the inference architecture. Across Ouro-1.4B, Parcae-1.3B and LoopFormer, reconstruction reduces squared target error by 20.2–55.5% on average per window at fixed student weights. On Ouro-1.4B, the complete framework lowers C4 perplexity from 45.44 to 29.54 and raises HellaSwag accuracy from 39.83% to 44.92% compared with the restricted TernaryLLM adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.