Risk-Controlled Post-Compression Verification for Visual Token Reduction with Exact Dense Fallback
Abstract
When a system trusts a compressed visual prediction on a particular input, how often does that choice alter the prediction of a frozen dense teacher? Aggregate accuracy and analytical compute do not answer this deployment-contract question. We study post-compression verification: a frozen vision transformer runs a token-merged continuation, accepts its output only when a calibrated confidence rule passes, and otherwise restores an exact dense continuation from an uncompressed activation saved before the first merge. Six merge ratios, one hundred thresholds, and four targets are calibrated jointly as one finite family of 2,400 candidates with per-candidate error probability 2.08 × 10^-5. Across four dataset–backbone combinations and three seeds, the 1% target accepts 79.5%–95.4% of examples; the largest per-seed calibration upper bound is 0.991%, while observed conditional dense-teacher disagreement is 0.42%–0.55%. A separately implemented program that does not import the production selector recomputes every reported policy and result. The selected policies use 0.74–0.82 of the analytical transformer compute. On the measured hardware platform, wall-clock gains appear only at batch one and reverse at larger batches, underscoring that analytical compute reduction is not a hardware-speedup guarantee. The study therefore reports four non-interchangeable layers of evidence: certified conditional disagreement, analytical transformer cost, compute saved by exact prefix reuse, and measured end-to-end latency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.