Adaptive Objectives Have Clocks: Transactional Gradient Accumulation
Abstract
Can a training checkpoint preserve its algorithm when memory constraints require a different gradient-accumulation factor? For stateful adaptive objectives, invoking the controller per microbatch changes its observation, state, and possibly random-event sequence. We identify this undeclared coupling in pinned PhysicsNeMo-Sym and PaddleScience integrations and specify an observe–prepare–realize–commit contract for an explicit optimizer-step reference. We characterize gradient agreement, establish reference equivalence under stated assumptions, and distinguish controller transactions from statistic reconstruction in a controlled 2×2 study. Paired PaddleScience experiments isolate action-connected clock effects; MultiMNIST checkpoint switches demonstrate prediction changes beyond ordinary accumulation controls. In a five-seed NYUv2 continuation study, Transactional Replay reduces prediction-distance AUC in every seed and task, with median paired reductions of 38–46%, while substantial finite-precision deviation remains. Retained graphs, replay, gradient bases, and an explicitly lagged alternative expose the cost of preserving current-step semantics. Trajectory recovery does not imply better task performance: Replay worsens NYUv2 depth absolute error in four of five seeds. Controller cadence is therefore part of an algorithm’s specification, not merely its memory configuration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.