Scaling Laws for Quantized Neural Computation: Geometry, Synthesis, and Execution
Abstract
Classical empirical neural scaling laws commonly treat scaling exponents as fitted constants, limiting mechanistic prediction before costly resource sweeps are executed. We develop a first-principles theory of quantized sequential neural computation that derives joint scaling surfaces directly from independently measurable structural profiles. First, we prove that marginal depth and library-size scaling rates cannot identify their joint law because depth dynamically reuses a finite operation library as a decoder. We resolve this bottleneck by introducing the task-quotiented shared-library reachable width. For a canonical spectral residual class, we solve the finite-resource minimax problem exactly, obtaining , where is depth, is library cardinality, and is the task-visible operation spectrum. This formula shows that depth resolves a clipped Cesàro profile rather than a primitive power law. Monotone regular variation yields a sharp summability transition in which nonsummable spectra retain spectrum-dependent rates, summable spectra enter a universal synthesis regime, and critical spectra are governed by an integrated slowly varying profile. Routed nonlinear residual adapters realize these phases, while task norm and loss geometry transform the observed exponent. Furthermore, execution semantics supplies a second structural law for arithmetic transport. Under fixed-horizon residual refinement, exact Green-operator identities show that full-state write-back has arithmetic amplification via backward-propagator mass, whereas first-order error feedback remains via adjoint semivariation. On a frozen pretrained ResNet-18 representation, a task-visible spectrum estimated from 5,000 calibration examples predicts all 49 registered depth–library errors on untouched test data with 1.02% median and 2.49% worst relative error without fitting a held-out exponent. Holding the ideal route fixed while altering only arithmetic semantics shifts the measured depth slope from 0.90 under write-back to 0.00 under error feedback, isolating a bit-width law of . Consequently, neural scaling exponents emerge not as primitive constants, but as identifiable, predictable summaries of operation geometry, finite synthesis, task geometry, and execution-error transport.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.