acceptodds
Under review as a conference paper at ICLR 2027

AIPO: Learning to Reason from Active Interaction

Abstract

Recent advances in large language models (LLMs) have demonstrated strong reasoning capabilities, largely stimulated by reinforcement learning with verifiable rewards (RLVR). However, the exploration of existing RLVR algorithms remains largely constrained by the knowledge and reasoning strategies already accessible to the policy model. Although recent methods introduce external expert demonstrations to broaden exploration, they typically rely on complete trajectory-level guidance, which can be sample-inefficient, information-sparse, and insufficiently adaptive to intermediate reasoning bottlenecks. Inspired by collaborative multi-agent systems, we propose AIPO, an enhanced reinforcement learning framework that improves LLM reasoning through active multi-agent interaction during exploration. Specifically, when encountering reasoning bottlenecks, AIPO enables the policy model to proactively consult three functional collaborative agents, namely the , , and , thereby obtaining fine-grained and state-dependent guidance during rollout. The resulting mixed-policy trajectories expose the policy to reasoning directions that may be difficult to discover through isolated on-policy exploration. To learn effectively from collaborator-provided tokens, we further introduce a corrected importance sampling coefficient together with a lower-bound clipping strategy to mitigate off-policy discrepancy and vanishing gradients. After training, the policy model reasons independently without relying on collaborative agents. Extensive experiments across mathematical, scientific, coding, and puzzle reasoning benchmarks show that AIPO consistently improves reasoning performance and generalizes across different policy models, collaborator backbones, and RLVR algorithms. Additional out-of-distribution and training-dynamics analyses indicate that active interaction broadens the policy model's , enabling it to solve additional problems independently under the same inference-time sampling budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.