SABR: Structure-Aware Budgeted Retention under Fixed Token Budgets
Abstract
Selecting relevant fragments is not sufficient when a context budget separates evi- dence needed together. We introduce Structure-Aware Budgeted Retention (SABR), a training-free selector combining frozen semantic scores, local coordination, and exact generator-token packing. We compare SABR with a strong frozen reranker at identical per-question context and prompt lengths, separating selection quality from full-context quality and online cost. On a previously inspected reserve of 1,384 questions across four tasks and three generators, SABR improves SQuAD F1 at a 70% context cap by 3.18, 4.19, and 4.23 points; the latter two paired 95% intervals exclude zero. The four-task macro difference is +0.92 [−0.08, 1.91]. A separate 896-question cohort yields a positive SQuAD estimate but does not confirm an ag- gregate gain. A matched-token sentence-unit pilot increases answer-string presence on SQuAD and mean intact-support fraction on HotpotQA by 10.94 and 13.80 percentage points. Timing studies show that exact implementation optimizations reduce selection overhead, but do not establish an overall speedup over full context. These findings identify a concrete regime in which coordinated retention improves fixed-budget selection, while distinguishing retained evidence, answer quality, and computational efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.