acceptodds
Under review as a conference paper at ICLR 2027

GEAR: Response-Aware Test-Time Compute Allocation for Frozen Quantized Spiking Transformers

Abstract

Frozen quantized spiking classifiers cannot change their weights at deployment, but they can spend extra computation on selected inputs. We study when an alternative view helps rather than merely changes a prediction. GEAR routes low-margin images from a frozen 4-bit QSD-Transformer through nested high-resolution and flipped views, then fuses their predictions. Its empirical response curve models a fixed action's net repair utility—repairs minus newly introduced errors—along the cheap-view margin rank. On ImageNet validation and five ImageNet-derived shifts, a fixed 26:8:4 policy gains 0.79–2.46 percentage points without target-set retuning. A labeled target-like probe calibrates the curve and operating width, evaluated on disjoint held-out rows; geometry interventions show why a source-optimal crop need not transfer. Tests on a second spiking family, synthetic corruptions, detection, and segmentation clarify its scope: margin routing helps when escalation corrects errors concentrated among uncertain inputs, but not when refinement mainly benefits confident ones. The response profile is conditioned on the chosen action and deployment distribution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.