I-Flow: Intent-Revealing Flow Optimization for Multi-Turn LLM Reasoning
Abstract
Multi-turn reasoning enables large language models (LLMs) to tackle complex problems through structured reasoning across turns. However, we uncover a critical failure mode of final-reward reinforcement learning: optimizing only for final correctness can collapse the intended hierarchy and drive diverse reasoning toward short and repetitive patterns. To address this limitation, we propose I-Flow, which structures each reasoning turn around a turn-level *intent* and introduces a Generative Flow Network (GFlowNet) optimization for the resulting intent hierarchy. Our bootstrap-free subtrajectory objective propagates terminal supervision across intent sequences, encouraging diverse high-reward reasoning structures without step-wise bootstrapping. We further replay successful intent structures to preserve effective reasoning patterns throughout training. Experiments on mathematical reasoning and code generation demonstrate consistent performance gains while effectively mitigating hierarchy collapse and reasoning simplification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.