acceptodds
Under review as a conference paper at ICLR 2027

I-Flow: Intent-Revealing Flow Optimization for Multi-Turn LLM Reasoning

Abstract

Multi-turn reasoning enables large language models (LLMs) to tackle complex problems through structured reasoning across turns. However, we uncover a critical failure mode of final-reward reinforcement learning: optimizing only for final correctness can collapse the intended hierarchy and drive diverse reasoning toward short and repetitive patterns. To address this limitation, we propose I-Flow, which structures each reasoning turn around a turn-level *intent* and introduces a Generative Flow Network (GFlowNet) optimization for the resulting intent hierarchy. Our bootstrap-free subtrajectory objective propagates terminal supervision across intent sequences, encouraging diverse high-reward reasoning structures without step-wise bootstrapping. We further replay successful intent structures to preserve effective reasoning patterns throughout training. Experiments on mathematical reasoning and code generation demonstrate consistent performance gains while effectively mitigating hierarchy collapse and reasoning simplification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.