PACO-PO: Policy-Adaptive Cost-Constrained Policy Optimization for Joint Reasoning Effort and Retrieval Depth in Budgeted Single-Turn RAG
Abstract
Single-turn retrieval-augmented question answering must decide, for every query, how hard to reason and how deeply to retrieve, and providers charge both the generated reasoning and the retrieved passages to one token budget. However, existing controllers tune the two levers apart, length control for reasoning effort and retrieval control for depth, so they miss that strong retrieval substitutes for reasoning while weak retrieval demands more of it, and cost-aware training bills reasoning alone, through a reward weight that must be swept or a per-query target length. We propose PACO-PO, a controller of roughly 0.2M parameters that, with the base LLM and retriever frozen, treats reasoning effort and retrieval depth as one coupled budgeted decision and learns it as a factorized joint effort–depth policy whose depth head is conditioned on the selected effort. The policy reads a zero-cost retrieval preview, the retriever's score profile and the passage lengths, to judge whether retrieval already suffices before any passage enters the context, so the choice itself spends no LLM tokens. Training casts the choice as a one-step constrained Markov decision process optimized by group-relative policy optimization, where a single Lagrangian dual variable prices the token budget, so one training run replaces a reward-weight sweep. On HotpotQA, Natural Questions, and MATH-500 with retrieval, PACO-PO comes within one point of the best fixed configuration (46.9% versus 47.6% macro accuracy) at 64% of its token cost, exceeds the best budget-matched single-axis controller by 2.9 points, and keeps the macro training-time expected cost within 3% of the budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.