acceptodds
Under review as a conference paper at ICLR 2027

Which States Should Be Compared? Behavioral Quotients for Interactive Policy Optimization

Abstract

Group-relative policy optimization converts rewards into advantages by normalizing within a reference group. For interactive agents, the appropriate reference group is unclear: exact observation identity splits situations with the same observed behavior, while broad pooling can merge distinctions that matter for control. We study this choice in graph-based path credit. From the action-labeled transition graph of a rollout batch, Behavioral Quotient Policy Optimization (BQPO) uses stable partition refinement to build a terminal-preserving quotient without parsing text, pixels, or task fields. With the unit edge costs used in our experiments, the quotient preserves observed success reachability, shortest-path credit, and action ordering. It therefore changes the group used for normalization, not the raw path target. Because equal path values do not make all samples interchangeable, an exact-first router keeps exact groups when local evidence identifies a clear continuation. It uses quotient groups when exact evidence is insufficient and, for other ambiguous cases, when a nontrivial class supplies complete action comparisons. On WebShop, ALFWorld, and Sokoban, BQPO improves the corresponding GraphGPO result in all seven reported aggregate comparisons, covering 1.5B and 7B language models and a 3B vision-language model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.