When Good Brain Predictors Make Bad Brain Controllers: A Regression Theory of Prediction–Control Dissociation
Abstract
Deep neural networks are increasingly used both as encoding models of visual cortex and as controllers that synthesize images to drive neural responses toward desired states. Recent in vivo experiments reveal a striking dissociation: models with nearly identical held-out accuracy can generate qualitatively different stimuli and differ sharply in control, and these differences track input-gradient structure better than predictive accuracy. We develop a geometric and analytical account of this phenomenon in a linear teacher–student model. The key is that control creates a self-generated distribution shift. Natural-image generalization, self-accentuation, and peer review evaluate the same fitted predictor under three different input distributions; correspondingly, their equal-error sets in weight space are ellipsoids, spheres, and hyperplanes. Using deterministic equivalents for ridge regression, we show that prediction-optimal and control-optimal regularization satisfy different spectral conditions, so cross-validation can leave large control error when responses are noisy and inputs are anisotropic. Extending the analysis to fixed linear feature maps, we construct a prediction-preserving gauge family with identical generalization and cross-validated regularization but arbitrarily different control. Control is most fragile when feature directions that survive regularization backpropagate into low-variance, high-spatial-frequency image directions. The resulting gradient-geometry factor extends to nonlinear networks as a Jacobian-based diagnostic and predicts across-model variation in in vivo control better than held-out prediction error. Thus, predictive accuracy alone cannot certify control; the model’s gradient geometry or generated stimuli must also be evaluated.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.