acceptodds
Under review as a conference paper at ICLR 2027

BudgetCMT: Joint Allocation of Teacher Compute and Student Supervision Budgets in Consistency Mid-Training

Abstract

Consistency mid-training (CMT) trains a few-step student on intermediate states of online teacher ODE trajectories. Generating these trajectories incurs teacher computation, while retaining student activations incurs memory cost. We introduce BudgetCMT, a joint allocation framework with three controls: teacher discretization , trajectory refresh stride , and student supervision budget . Uniform sampling without replacement preserves the dense gradient in conditional expectation at a fixed parameter state and teacher pool, with covariance given by the finite-population correction. Complementary two-step reuse covers the pool once per cycle while halving teacher evaluations. On CIFAR-10, reducing supervision from 128 to 64 cells lowers reported interval peak GPU allocation by 41.8% in a three-seed 10k-update study; quality intervals allow both improvement and degradation. In a separate three-seed 30k-update continuation, reuse reduces loop time by a median of 30.6%. Its 2-NFE FID is 1.40% higher at equal updates and 1.42% lower at a pre-specified approximate equal-time endpoint. A single paired run from a second warm start shows the same direction, with a smaller equal-time gain. These results establish measured memory–time–quality trade-offs within CMT; the tested downstream fine-tuning recipe does not preserve checkpoint quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.