acceptodds
Under review as a conference paper at ICLR 2027

DepthLens: Measuring How Evidence Assembly Changes Useful Retrieval Depth

Abstract

A RAG score observed at one retrieval depth does not identify how long additional evidence remains useful: the assembly operator can change the entire depth–utility response. We define useful-depth shift as the paired displacement of the fitted low-return transition under an assembly replacement, and instantiate it in DepthLens by crossing seven depths with assembly operators on a fixed retriever–reranker–generator stack and resampling matched queries through the complete estimator. On HotpotQA, 2WikiMultiHopQA, MuSiQue, and FEVER, a Manual Ledger that carries entities, supporting evidence, and unresolved hops moves this transition from 7.4–16.2 passages to 3.4–5.2, a 2.2–3.1\(\times\) shift; Qwen2.5-72B-Instruct reproduces 2.1–2.9\(\times\). In the operator ladder, layout-matched and random-schema controls retain CoT-like transitions, whereas a frozen automatic scaffold recovers most of the movement; the Manual-over-Auto residual rises across pooled hop strata, and the Manual-over-CoT gap widens with distractor rate. The measured displacement aligns with held-out action quality: with Manual Ledger fixed, DepthLens reaches 88.9–93.1% near-SLO pair accuracy, stays within 0.5–0.9 primary-score points (EM; FEVER label accuracy) of the feasible oracle, and improves on the matched Learned-\(k_r\) comparator by 0.9–2.9 points. Joint depth–assembly selection then gains 3.2–12.2 points over matched depth-only adaptation at 1.48–1.68 s P95, showing that retrieval evaluation must cross depth with evidence assembly to expose both scientific attribution and efficient operating points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.