When Retrieval Endpoint Scores Miss the Value of Feedback
Abstract
Equal blind and known-condition retrieval scores do not certify equal partial-feedback value. An exact average-precision construction matches dimension, action count, and endpoints while reversing the preferred question. A natural-image submenu census finds disjoint optimal one-question sets in 13 of 1,920 cells; six also favor a paid question over stopping, and one native-rank rational-AP case verifies the reversal. This evaluator-only existence result does not establish a deployable selector. Across the same learned libraries, endpoint-preserving deletion preserves every optimal tree in 1,887 cells, including 318 of 351 with positive allowed-probe risk. Publicly predicted endpoint cores show no established replanning advantage over matched menus. At exactly two answers, SOURCE-calibrated planning improves single-ranker AP over SOURCE-selected controls, while fusion gains remain unestablished. Endpoint insufficiency, retained-action inadequacy, and deployed-policy failure are distinct diagnoses. These exposed-population results support evaluation under explicit information and cost assumptions, not a universal repair or serving speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.