acceptodds
Under review as a conference paper at ICLR 2027

Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers

Abstract

Long-horizon LLM agents succeed or fail through intermediate information-gathering turns, yet training signals are usually observed only at the final answer. Turn-level shaping methods reward turns that increase the likelihood of a gold answer, but require answer supervision or stable task-specific verifiers that are often unavailable in agentic settings. Label-free reinforcement learning (RL) methods extract self-signals from output distributions, but mostly at the answer or trajectory level and therefore cannot attribute value to intermediate turns. We propose *Self-Induced Outcome Potential* (SIOP), which treats semantic clusters of final answers as latent future outcome states for potential-based turn-level credit assignment. For each query, we sample multiple rollouts, cluster their final answers into semantic outcome modes, and construct a reliability-aware target distribution over these latent states. Each turn is then rewarded by how much it increases support for reliable outcome states, and these rewards are propagated through turn-level advantages rather than a single rollout-level advantage broadcast to all turns. We formalize the framework and establish its connection to supervised turn-level shaping in the gold-answer limit. We evaluate SIOP on seven search-augmented agentic reasoning benchmarks with Qwen3 and Phi-4-mini backbones, and on multi-turn tool-integrated reasoning. SIOP achieves the best average performance among verifier-free methods at two model scales while approaching a gold-supervised outcome baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.