acceptodds
Under review as a conference paper at ICLR 2027

CCRT: A Single-State Optimizer with Retained Gradient History and Bounded Nonlinear Updates

Abstract

Modern optimizers often retain multiple full-shape statistics, making optimizer state a substantial cost at scale. Current Carry Root Transport (CCRT) instead persists one BF16 gradient-history tensor per parameter, reinjects the fresh gradient, and applies a bounded root-rational readout without a second moment. Across three observed 134.8M training trajectories—two fresh paired runs and one earlier independent run—CCRT shows the same terminal validation-CE and validation-AUC ordering against calibrated AdamW and PowerStep. In a complementary single-seed 3B-scale reproduction, CCRT reaches validation perplexity 127.15, versus 246.22 for AdamW and 378.64 for PowerStep, while using 75% and 50% less persistent optimizer state, respectively. CCRT also completes a 1K-step single-GPU run at 6.98B parameters on an 80GB A800; validation CE decreases across all five recorded checkpoints, with a 13.00 GiB BF16 optimizer state. We also characterize the ideal full-precision update mathematically. Exact scalar quadratic dynamics exhibit a structured nonzero fixed-step two-cycle, while a deterministic smooth-nonconvex analysis proves a natural-stationarity residual of order and conventional first-order stationarity complexity. The scalar cycle matches the step-size exponent of the general residual bound. Together, the empirical and theoretical results show that a low-state nonlinear optimizer can combine substantial state savings with effective large-model optimization and an explicit connection between fixed-step dynamics and general nonconvex stationarity complexity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.