Pairwise Tail-Span Authorization for Resource-Constrained Filtered Search
Abstract
Across a 50M-vector filtered-retrieval stress grid, the fastest recall-feasible action changes on 364/940 adjacent edges in the 5–15% selectivity band, versus 351/3,860 edges elsewhere. These reversals turn a plausible top-ranked plan into lost generation slack when its fanout or SSD traffic produces a latency-tail error. RAVEL makes execution of that winner a pair-conditioned authorization decision: held-out recall tables and RAM, SSD, and latency caps define feasible P1–P4 actions; calibrated median-to-P95 spans characterize the leading pair; and ambiguous requests take a regime-local fallback. On dual AMD EPYC 7543 CPUs with 128GB RAM and a 1TB SSD budget, RAVEL reaches 37ms P95 and 56ms P99 at a common 0.95 Recall@20 target, improving on GateANN's 44/70ms by 16%/20% and Qdrant's 68/101ms by 46%/45%. With actions, requests, recall tables, provisional ranks, and fallback implementation fixed, the executed-policy mis-pick rate falls from 12.6% for winner-only dispatch to 4.9%. At 80% authorization coverage, RAVEL reduces selective P95 excess latency from RCPS-Tail's 14.6ms [12.8,16.4] to 9.6ms [8.3,10.9]. The enterprise-selected gate also reaches 38.9ms P95 on 356 public band queries, compared with 41.8ms for RCPS-Tail and 58.7ms for the native policy. Layering this measurable tail-risk decision over existing kernels preserves retrieval latency budget for downstream generation without requiring an index replacement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.