CellGain: One-Shot Answer-Metric Pricing for Hard-Budget Visual Token Allocation
Abstract
Fine-detail visual question answering presents a discrete acquisition problem before expensive encoding: a page lattice offers 36 cheap 16-token refinements and 36 detailed 64-token refinements, but a hard cap admits only a question-dependent subset. CellGain assigns these 72 choices their expected answer-metric gain across partial visual states. Offline reveal-and-redecode experiments shape the target with the residual between privileged dense and matched partial-view regressors; deployment compresses that computation into one query-conditioned map from a 22M-parameter scout and compiles the map into a feasible mixed-scale set. With InternVL2-8B and 1,008 optional fine tokens, CellGain reaches 82.7/52.1/74.8/63.8 on DocVQA, InfographicVQA, ChartQA, and ScreenQA—within 0.4/0.7/0.4/0.9 points of dense 2,304-token inference. At the same mean allocation, it exceeds Query-Window by 0.8/1.2/1.7/1.3 points; all four Holm-corrected paired intervals exclude zero. Under an identical cell interface, residual supervision leads observed utility, dense-teacher-only, and non-privileged pairwise targets, while the original map retains correlation with remaining realized values after 16 reveals and leaves only 0.1–0.3 points to sequential re-scoring. The resulting 56.2% reduction in optional fine tokens lowers model-side latency from 912 to 468 ms/image, and retraining the valuation mechanism on Qwen2-VL-7B-Instruct preserves the four-task advantage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.