Same Information, Different Generalization: Representation Sensitivity in Neural Network Regression
Abstract
Invertible feature transformations preserve predictive information, yet can affect finite-sample learning performance. We study this representation sensitivity in controlled synthetic tabular regression. Our paired experiments compare transformations with the same normalized population error under optimal affine approximation, thereby matching the total amount of induced nonlinearity while varying its structure. Despite this matching, neural networks can exhibit different learning curves, and the ordering of representation penalties can reverse depending on whether the transformed coordinates define the inputs or the target relationship. In contrast, XGBoost is nearly unaffected by the transformations. We characterize these differences by decomposing nonlinear target energy across polynomial degrees using projections adapted to the observed input distribution. Our results show that neural-network representation sensitivity depends on the target's nonlinear structure in the observed coordinates, the input distribution, and the learner.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.