acceptodds
Under review as a conference paper at ICLR 2027

NORMFOLD: Function-Identifiable Spectral Adaptation at LayerNorm Interfaces

Abstract

Spectral adapters derive trainable coordinates from a stored pretrained matrix, but storage coordinates need not be uniquely determined by the function. We expose this ambiguity at a LayerNorm followed by an affine map. The complete local invariant is , where and . We prove that two parameter tuples with the same can induce spectral tangent spaces of different dimensions, so scale folding and offset preservation alone do not make spectral adaptation well defined on the pretrained function. NORMFOLD instead decomposes the unique minimum-Frobenius representative , preserves , and ties exactly repeated singular blocks. Its adapted family and matched optimizer trajectory are therefore determined by the local function in exact arithmetic. A controlled 9,120-fit study spans two CLIP checkpoints, three datasets, two shot counts, five support seeds, seven nonzero rewrite strengths, and five rewrite directions. Across all 2,100 matched nonzero task-direction conditions, NORMFOLD has lower final-logit discrepancy than the closest offset-preserving control. Its mean discrepancy remains between and , while that control ranges from to . NORMFOLD changes 9 of 10.88 million repeated paired predictions, compared with 27,145 for the control. Swin-T, seen-to-unseen transfer, fresh-image replay, merged export, and finite-precision analyses support the same function-identifiability account. The claim concerns whether the learning rule is representation independent, not whether NORMFOLD improves task accuracy over every spectral baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.