Pixel-Level Tasks Need State-Aware Compensation and Adaptation
Abstract
Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of pre-trained models to downstream tasks by updating only a small subset of task-specific parameters, but still requires full model computation during inference. Although token reduction (TR) improves inference efficiency by assigning different computational paths to tokens, it introduces partial information loss as a consequence of reduced token computation. This challenge is even more pronounced in pixel-level tasks, which require dense and fine-grained spatial information for accurate prediction. Existing TR-PEFT methods address this challenge by increasing adaptation capacity or introducing explicit compensation pathways. However, these designs typically do not explicitly account for the distinct token states arising from TR-induced computational paths, resulting in a state-agnostic formulation. Our analysis reveals a design gap in state-agnostic approaches under TR, with observable effects on compensation effectiveness and adaptation differentiation across tokens. To address this gap, we propose a state-aware approach that leverages TR-induced token states to regulate compensation and adaptation accordingly. Our approach effectively improves compensation and adaptation, achieving better performance with reduced computational overhead across diverse pixel-level tasks. We further demonstrate competitive or superior performance compared with strong baseline methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.