Escalate Only When Necessary: Marginal-Value-Aware Feedback Acquisition for Retrosynthetic Planning
Abstract
Retrosynthetic planning decomposes a target molecule into building blocks, and recent LLM-based planners enrich this search with external feedback. Such feedback, however, varies substantially in fidelity and acquisition cost, creating a trade-off between decision reliability and resource consumption. We formulate feedback acquisition as adaptive escalation, where a planner progressively acquires higher-fidelity feedback only when its expected benefit justifies the additional cost. To learn such policies, we propose Marginal-Value-Aware Advantage Estimation (MAVE). MAVE models the reward-cost relationship across successive escalation levels, using its slope to estimate the current marginal value and its curvature to capture how this value evolves with further resource investment. MAVE incorporates both signals into advantage estimation, enabling marginal-value-aware credit assignment for just-enough escalation policy while discouraging premature stopping and over-escalation. Across three bencharks, MAVE achieves the highest success rates with hierarchical feedback, while reducing average feedback cost by 64.9% under the yield-only setting. Code is available at https://anonymous.4open.science/r/MAVE-retro/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.