FlowWeaver: Weaving Reasoning Paths with Cross-Trajectory Flow for Fine-Grained Credit Assignment
Abstract
A language model may make hundreds of token-level decisions before receiving a single outcome reward, while responses sampled for the same prompt can follow different reasoning paths and reach different outcomes. Turning this cross-trajectory variation into fine-grained credit remains challenging: within-trajectory attribution can identify important tokens but evaluates each response largely in isolation, whereas cross-trajectory methods exploit outcome variation at coarser granularity or rely on shared prefixes or states. We introduce FlowWeaver, which weaves sampled reasoning paths into a shared attribution graph through cross-trajectory information flow for fine-grained credit assignment signals. We build on attention-based flow within each response, constructing paths from prompt to answer using attention weights and quantifying the share of graph propagation routed through each token. Using cumulative flow as a common coordinate, we connect sampled responses through aligned token positions, forming a shared distribution over graph paths from the prompt to different outcomes. We then condition propagation on each response's observed outcome, measuring how much of the group's flow toward that outcome passes through each token. We use these importance scores to shape token-level credit by reweighting GRPO advantages. RL experiments across mathematical reasoning, general reasoning, and agentic tasks demonstrate the effectiveness of , highlighting outcome-conditioned flow as a fine-grained signal for allocating outcome supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.