Scaling Width and Scaling Compute Teach Different Algorithms
Abstract
Adding parameters and adding optimisation steps are both called scaling, and they push a network toward different algorithms. We show this on two tasks that share no structure, each solvable either by a cheap shortcut or by a general rule, over a grid of width and step budget at fixed batch size and seeds per cell. Compute favours the rule: averaged over widths, the fraction of runs that acquire it climbs monotonically with the budget in every width column of both tasks, from to on one and from to on the other. Width favours the shortcut once past a small optimum, falling between and on the first task and past its peak on the second. So “more scale” is not one direction, and a study that grows a single compound budget averages two effects of opposite sign. The outcome distribution turns out to depend on the rule rather than on the resources: for parity, an all-or-nothing computation, of networks land on one of exactly two solutions, each hitting its predicted accuracy ( and ) to within with nothing fitted, and the smooth averaged curve is then the mixing fraction between them, matched by to an RMS of , on which no individual model lies. For associative recall, which admits partial credit, the same measurement gives graded outcomes, with of runs in the interior. Discreteness therefore requires a rule that cannot be half-learned; the opposition between width and compute does not.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.