acceptodds
Under review as a conference paper at ICLR 2027

Learning Retrieval Plans from Sibling Execution Outcomes

Abstract

Multi-hop question answering requires agents to retrieve evidence and use earlier findings to guide later searches and reasoning. Yet different workflows can yield the same correct answer, so answer labels alone do not define a unique target for supervising intermediate actions. To address this problem, we propose SIPPO (Sibling-Informed Planner Preference Optimization), which learns executable workflows through execution-grounded supervision and final-answer rewards. A shared role-conditioned policy generates a directed acyclic graph whose dependencies resolve against node memory, connecting actions to states from earlier execution. We construct execution-grounded supervision by executing teacher-proposed workflows on multi-hop question-answering data, verifying evidence support and answer correctness, and pairing each recorded action with the state available at that step. We instantiate RLOO at workflow granularity (Workflow-RLOO), assigning advantages from final-answer Exact Match (EM) rewards to every generated role action in a workflow. Sibling Execution Outcomes are the terminal EM rewards for on-policy workflows sampled for the same question. Planner Sibling Contrast (PSC) uses reward differences to compare planner responses from EM-correct and EM-incorrect workflows, without additional intermediate labels. Across four multi-hop benchmarks comprising 22,517 answerable questions, SIPPO achieves the highest macro EM and macro F1 in the main comparison in both backbone settings. Its macro EM/F1 are 0.3362/0.4212 on Qwen2.5-3B and 0.3813/0.4701 on Qwen3-4B. Adding PSC to Workflow-RLOO improves macro EM by 2.72 and 7.01 percentage points, respectively; the paired 95% confidence interval for the 3B gain excludes zero. The 3B gain accompanies fewer generated tokens and more retrieval calls. Frozen-policy diagnostics show that PSC gradients respond to the terminal EM reward labels of same-question workflows.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.