acceptodds
Under review as a conference paper at ICLR 2027

CABLE: Correlation-Aware Bargaining for Multi-Reward Post-Training

Abstract

Multi-reward post-training can maximize easy proxy rewards while task accuracy declines. We introduce CABLE (Correlation-Aware Bargaining with Liveness and Equalization), which combines correlation-based reward weighting with explicit control of normalization. It estimates interactions between objectives from token-logit correlations and combines their learning signals before a single actor backward pass. CABLE-L preserves reward-signal amplitudes; CABLE-LE adjusts weights using recent reward levels and adds entropy regularization. Across three training runs on Open-R1 coding with a 1.5B model, CABLE-LE reaches 12.56% LiveCodeBench-v6 pass@1 and a 36.37% four-suite average, improving on GAPO by 0.87 and 1.10 percentage points, respectively. On Math-M3, CABLE-L reaches a 52.21% five-benchmark average, compared with 51.32% for GRPO and 51.18% for the base policy. It also reduces mean response length from the base policy's 9,085 tokens to 3,203 tokens. On a matched actor-core workload, CABLE is over 4.2× faster than a parameter-gradient implementation of the same allocation rule.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.