The State Between Stages: Identifying Sufficient Statistics from Second Derivatives
Abstract
A deterministic differentiable program often lets one block of its inputs interact with the rest only through a few numbers, its interface, whose dimension is the state a staged solver must carry. Representation learning estimates such a summary from data; we compute it from second derivatives, even where no variable holds it. The blocks' cross-Hessian factors through the interface, so its rank at any set of second-block values bounds from below the dimension of every linear interface and every smooth one with a full-rank Jacobian. Under analyticity, constant rank, a full-rank summary map, and values drawn from a density, the rank almost surely equals the minimal dimension of an analytic interface, which can lie below that of any linear one. On a control rollout it reads 2 where the loop carries 4 numbers, recovering the zero-effort miss (where the mass would land if it coasted), which no line computes. Solving in stages through it reaches the joint optimum. On a published robot-arm benchmark it reads 2 against 3 carried numbers. Given only the objective, it recovers the interface of 200 of 200 random objectives and of 6 Feynman formulas, and two ratio interfaces as payoff differences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.