SSC-LoRA: Semantic Spectral Control Low-Rank Adaptation for Generalizable 3D Gaze Estimation
Abstract
Accurately predicting 3D gaze directions from single images remains challenging due to the absence of unified, physically grounded representations and limited diverse annotated data. Although visual foundation models (VFMs) provide strong transferable priors, their inherent visual sensitivity misaligns substantially with downstream regression tasks. Specifically, gaze-irrelevant appearance attributes dominate feature learning and overshadow subtle gaze variations that are critical for precise prediction. Large-scale pretraining prioritizes generic visual capabilities without explicitly encoding fine-grained gaze sensitivity. When appearance patterns spuriously correlate with gaze signals during pretraining, such sensitivity imbalance induces shortcut learning in downstream adaptation and degrades cross-domain generalization. To address these issues, we propose a novel parameter-efficient fine-tuning framework, Semantic Spectral Control LoRA (SSC-LoRA). Built upon vanilla LoRA, SSC-LoRA integrates gradient-based semantic identification (SI) and spectral control (SC) for task-aligned adaptation. Specifically, SI first decomposes pretrained weights into singular components and ranks their responsiveness to gaze-relevant and gaze-irrelevant objectives. SC further suppresses weight updates from gaze-irrelevant spectral bases while reinforcing gaze-specific feature learning. Both modules operate inherently within LoRA’s low-rank space without introducing additional trainable parameters. Evaluated on five gaze estimation benchmarks with CLIP-B/L and DINOv2-B/L backbones, SSC-LoRA consistently outperforms most SOTA methods in both in-domain fitting and cross-domain generalization. Extensive ablation analyses validate that the performance bottleneck stems from gaze-irrelevant-sensitive adaptation, and confirm that our spectral control mechanism effectively recalibrates update allocation across task-relevant and task-irrelevant feature subspaces.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.