The Intrinsic Geometry of a Task
Abstract
Understanding why deep learning models generalize on a task requires explaining how their internal representations encode information relevant to that task. Although empirical studies have revealed regularities in representation geometry, a theory connecting generalization to this geometry remains incomplete. We study this connection using the Fisher–Rao geometry induced by cross-entropy: a task already has a geometry, and that geometry is the reference a representation must match. A representation is sufficient if it recovers the task's conditional laws, and canonical if it is also minimal. Sufficiency implies isometry to this geometry; local deviations from a sufficient predictor determine excess loss. Pulling the same geometry back to parameter space yields local loss bounds and a criterion for pruning. Experiments confirm the theorems in the regime they assume and show that pullback residual length predicts pointwise excess loss. We then operationalize this geometry through pruning experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.