Additivity and Selection Value in Depth-Indexed Model Updates
Abstract
Selecting multi-step losses for a shared dynamics model requires predicting the value of training on a set of objectives. Does this prediction need interactions between losses? We study the question through finite-update probes that fit additive and quadratic set predictors to a fixed teacher-relative utility. For fixed-length SGD blocks, we derive orders , , and for raw singleton, pair, and triple effects at step size . The leading singleton term is initial-gradient alignment. Experiments on two continuous-control tasks with feedforward and recurrent dynamics exhibit this hierarchy. Ridge fitting nevertheless assigns substantial coefficients to pair features: removing the penalty shifts most larger-set prediction comparisons in favor of the quadratic model but changes no choice within the three-loss family. Additive and quadratic choices agree at 178 of 180 common states. In an optimizer stress test, ten-step Adam blocks from fresh optimizer state produce larger higher-order effects, and both predictors select negative-gain updates despite positive mean oracle gain. Closed-loop controls distinguish useful composition from useful interaction modelling. Selected triples improve lower-tail return over All-five updates at matched total loss weight, but increasing All-five's total weight reverses the ordering; continued inexpensive training also outperforms quadratic selection under the recorded cost budgets. These results identify three requirements for interaction modelling to improve training: nonadditive update effects, decision-relevant prediction, and benefits that justify acquiring the labels. no choice within the three-loss family. Additive and quadratic choices agree at 178 of 180 common states. This agreement is not universal: Ten-step Adam blocks from fresh optimizer state produce larger higher-order effects, and both predictors select negative-gain updates despite positive mean oracle gain. Closed-loop controls distinguish useful composition from useful interaction modelling. Selected triples improve lower-tail return over All-five updates at matched total loss weight, but increasing All-five's total weight reverses the ordering; continued inexpensive training also outperforms quadratic selection under the recorded cost budgets. These results identify three requirements for interaction modelling to improve training: nonadditive update effects, decision-relevant prediction, and benefits that justify acquiring the labels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.