SPAN-INR: SPATIAL ASSIGNMENT NETWORKS FOR DEEP SINUSOIDAL IMPLICIT REPRESENTATIONS
Abstract
Real signals are detailed in some places and smooth in others, so a sinusoidal implicit neural representation (INR) has to produce different harmonics at different locations. Existing sinusoidal INRs, however, spend their parameters on a separate matrix in every layer, which decides which channels can combine (*reach*), and do little to control what is produced where (*assignment*): the scale and offset of the argument entering each sine, which by the Jacobi–Anger expansion decide how high the harmonic orders go and the balance between odd and even harmonics, are the same at every location or are adjusted only by coarse masks on groups of features. In this paper, we first show that in our setting reach is cheap and assignment is what has to be learned. Analytically, which harmonic interactions are available depends on the layer matrix only through its connectivity, and the scale of each of its rows can be taken over by a gain; empirically, a single dense mixing matrix shared by all stages and frozen at its random initialization loses only 0.14 dB against a fully learned one, whereas restricting it to its diagonal loses 5.81 dB, and a model with an independent matrix per stage loses 3.48 dB at the same budget. Based on this observation, we propose a new sinusoidal INR, called Spatial Assignment Network (SPAN-INR), which spends its parameters on assignment rather than on reach: most of them go to coordinate controllers, one per stage, that predict a gain and a shift for every channel at every location, while the single shared mixing matrix that provides the reach is small by comparison. Since a new stage needs only a new controller and no new matrix, the same budget buys more stages, and every stage is one more place to assign harmonics locally. Experiments show that SPAN-INR achieves 38.90 dB PSNR on image fitting with 126K trainable parameters and the best results among the compared methods on super-resolution, denoising, 3D occupancy and novel-view synthesis. Analyses of the fitted models further confirm that the gain tracks the local spectral centroid in the later stages, the shift tracks the odd/even balance, and the empirical neural tangent kernel is 9.0 times narrower at edges than in smooth regions, against 4.8 times for SASNet.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.