TEMPO: Multi-timescale Value Learning and Policy Credit for Lanuage Models
Abstract
Reinforcement learning for language models must turn response-level rewards into credit for individual tokens. Inspired by heterogeneous temporal value signals in the brain, TEMPO trains a multi-timescale critic, normalizes each horizon’s advantages by their root-mean-square amplitude, and combines them for policy optimization. We analyze the reward kernel induced by this construction and distinguish temporal prediction from actor readout. In the original Qwen3-14B comparison, TEMPO improves greedy accuracy on all five mathematical bench- marks, including MATH500 from 78.4% to 84.0% and MATH5000 from 79.40% to 85.18%. On CommonsenseQA, Qwen2.5-7B-Instruct reaches 83.95 ± 0.54% validation accuracy across three training seeds, compared with 82.69 ± 0.91% for PPO, while generating substantially shorter answers. Controlled Qwen3-1.7B ex- periments then isolate the readout choice: with the same four-horizon target design, long-horizon readout achieves 72.2% MATH500 accuracy versus 64.8% for RMS averaging; a duplicated long-horizon control achieves 72.8%. Shared-pool sampled evaluation and logged advantage geometry clarify this separation. The results show why amplitude balance and temporal diversity must be evaluated alongside the reward weighting supplied to the actor, providing a concrete account of when a multi-timescale construction changes language-model policy learning. Code is available at https://anonymous.4open.science/r/TEMPO-5D50.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.