Specific Algorithmic Interpretability of Neural Networks: A Case Study on Textures
Abstract
We develop a principled framework for constructing neural networks whose specific parameter realizations admit an explicit algorithmic interpretation. Existing algorithm-inspired architectures can explain the computational structure of a network, yet after standard training the learned parameters need not retain a clear relation to the motivating algorithm. We address this gap as follows. First, we model each data point as a sample of a class-dependent stochastic process and assume that statistics of this process can be estimated from a single sample and these statistics are sufficient to distinguish the classes . We then construct a neural network whose initial parameters exactly implement an algorithm for estimating these statistics, making the network fully interpretable. To account for mismatch between the idealized model and real data, we fine-tune this network while controlling its deviation from the algorithmic initialization. The trained network hence roughly retains the interpretation of the initial network. A PAC-Bayesian analysis yields a uniform generalization bound whose complexity term scales with the fine-tuning radius, providing a statistical motivation for our approach. We instantiate the framework for texture classification using the scattering transform to estimate the discriminative statistics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.