Pay to Reveal: Adaptive Information Acquisition under Communication Budgets
Abstract
Retrieval-augmentedgeneration(RAG)groundslanguage-modelanswersinexter nalcorpora,andmulti-hopquestionsoftenmakelaterretrievaldependonevidence foundearlier.Existingiterativesystemstypicallyquerybody-derivedindexesor rerankretrievedcandidatepools.Westudyacomplementarydecision:acontroller canenumeratedocumenttitles,buteachbodymustbeacquiredunderalimited readbudget.WeformalizethissettingasPaytoReveal.Apolicymayacquireat mostKdocuments;eachacquisitionreturnsaquestion-conditionedexcerptcapped atLunderacharacter-basedtokenestimate;andaccumulatedexcerptsguidelater choices.WeproposeBM25-RF,aniterativepseudo-relevance-feedbackpolicy thatexpandsthequerywithacquiredexcerptsandreranksunopenedtitleswith BM25.WealsoproposeResidual-CMG,whichaddstothestatictitle-BM25score thedifferencebetweenasupport-documentclassifier’sacquired-statescoreandits empty-statescoreatthesamestep.Theclassifieristrainedonlyonthetrainingsplit. Underanexplicitexpected-progresscondition,afinite-Kanalysisrelatesper-step selectiongapstofinalsupport-documentrecall.Weevaluatesixmethodsunderone protocolon22,246retrievalquestionsand3,000Qwen3-4B-Instruct-2507answer questionsfrom2WikiMultiHopQA,MuSiQue,andHotpotQAFullWiki.Relative toStatic-BM25,BM25-RFchangesanswerF1by+.067,+.049,and−.028across thethreedatasets.Addingthestatictitle-BM25scoretothelearnedscoreimproves supportAPoverDirect-CMGonallthreedatasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.