acceptodds
Under review as a conference paper at ICLR 2027

Jacobian Policy Optimization: Geometry-Aware Learning from Structured Rewards

Abstract

Reinforcement learning for language-model reasoning increasingly uses multiple reward signals to evaluate the consequences and properties of generated responses. Training a shared policy to respond to different priorities requires learning behaviors whose outcomes reflect the requested emphasis among these signals. We introduce **Jacobian Policy Optimization (JPO)**, a preference-conditioned reinforcement learning method that uses objective-specific gradient geometry to construct policy updates. JPO independently normalizes each reward dimension, constructs separate clipped policy losses, and differentiates them to form a Jacobian of policy gradients. Preference-weighted UPGrad modifies conflicting gradient components before aggregation, allowing both the requested preference and the relationships among learning signals to determine the update. We evaluate JPO in two complementary settings. On MaZE, where rewards describe completion, resource collection, and hazard avoidance, JPO achieves substantially higher matched-preference utility and stronger preference specialization than GRPO and GDPO on held-out mazes under the training preference set. On Level-5 MATH, where rewards assess correctness, conciseness, and rigour, JPO achieves the highest individual-sample accuracy and Pass@8 among the evaluated methods, while remaining competitive on Majority@8. It also achieves the highest Majority@8 after Recursive Self-Aggregation. These results support geometry-aware multi-reward policy optimization as an approach to improving both preference-conditioned behavior and reasoning performance under multiple quality criteria.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.