OCPO: Observed Continuation Policy Optimization for Social Agents
Abstract
Social intelligence simulation challenges language agents to navigate open-ended, multi-turn interactions while balancing social goals and interpersonal relationships. However, training signals are typically available only at the dialogue level, making credit assignment to individual utterances difficult and resulting in sparse, delayed supervision. Existing methods often obtain finer-grained feedback through additional dialogue rollouts, external evaluation, or separately trained reward models, making supervision costly to acquire. To address this, we propose Observed Continuation Policy Optimization (OCPO), a framework that turns immediate partner responses recorded in high-quality dialogue trajectories into local policy supervision. OCPO scores policy-sampled utterances by their compatibility with the same observed partner continuation. It uses sensitivity-based token weighting to emphasize tokens sensitive to candidate utterances and mean-only advantage centering to preserve reward-difference magnitudes in sequence-level updates. By reusing recorded continuations, OCPO acquires rewards without additional full-dialogue rollouts, a separately trained reward model, or external evaluator calls during policy optimization. Experiments on the benchmark social intelligence environment demonstrate that OCPO consistently outperforms large reasoning models and existing optimization methods. Notably, OCPO cuts GPU-hours by 65% compared with previous social-policy optimization pipelines, offering an effective and cost-efficient paradigm for building social agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.