Understanding Prompt Learning as Low-Rank Manifold Alignment
Abstract
Under a linear shared-latent model with frozen linear encoders and conditionally zero-mean text noise outside the shared subspace, we show that the population contrastive objective admits an optimal prompt operator with the rank smaller than the shared latent dimension, for any positive temperature. For nonlinear representations, we establish a latent factorization of the centered cross-covariance operator, with a rank bound under low-dimensional conditional-mean structure and a spectral-tail bound when that structure is approximate. HSIC provides a one-way dependence guarantee for centered trace alignment, rather than an equivalent objective. We also derive a generalization bound for explicitly rank-constrained prompt classes. Synthetic diagnostics distinguish exact trace-objective recovery from the nonzero spectral tails of normalized InfoNCE, while real-CLIP measurements provide operator-level evidence of spectral concentration. Inspired by the shared-coordinate structure, LR-MaPLe couples visual and textual prompts through shared factorization codes. Across 10 benchmarks and three seeds, it achieves accuracy comparable to MaPLe with approximately fewer trainable parameters. The factorization width is an architectural choice, not an estimate of the latent dimension.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.