acceptodds
Under review as a conference paper at ICLR 2027

Neural Networks Spontaneously Dissociate Visual from Informational Persistence Using Iconic Memory

Abstract

Tied classically to decaying persistence in early visual cortex, is a high-capacity sensory trace that briefly outlasts its stimulus. Yet, fading inside the cortex is not confined to any single region, nor does any reverberant activity that counteracts such fading cease uniformly. The ventral stream sustains retrograde insignia intractably, and the recurrence that re-amplifies this decaying activity exists at nearly every level of the visual hierarchy, rendering behavior hard to isolate: Appended to neural networks where such stage-wise differentiation already exists, would a fading percept let us determine its functional use? Confer benefits beyond what a mere expansion in parameter space warrants? To find out, a proper test should heed all stages equally, letting the network determine its own specialized regions, . Consequently, we place the same lightweight following each stage of quaternate-depth convolutional and attention-based classifiers—governed by a learnable () to index temporal persistence, () to index feedback strength, succeeded by canonical divisive normalization, all initialized identically—and train on image classification tasks. Serendipitously, ResNet, DenseNet, ConvNeXt, ventral stream model CORnet-S, transformer model Swin-Tiny, despite their varied architectures, converge on the same final structure: All concentrate Iconic persistence at the sensory input and at the categorical readout—the two ends we argue double as network's artificial counterparts to dissociating early and late intra-cortex, the two telltale components of . Dependent on natural image structure, this particular emergence finds no expression when non-ecological or phase-scrambled images are used during learning. Furthermore, the putatively V1-correspondent early gain is subsumed when fixed Gabor filters are prepended and, while the late stage jump is preserved, its divergence is amplified with increasing task complexity. Across architectures, yield relative clean gains exceeding 8%, mean corruption gains 18% and greater resilience across adversarial attack budgets—a rarely seen simultaneous improvement that through its robustness profile hints at increased alignment with the visual cortex, a finding we further corroborate with representational-similarity tests.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.