acceptodds
Under review as a conference paper at ICLR 2027

TFPO: Token-Level Objective Fusion for Stable Preference Alignment

Abstract

Preference optimization is usually applied uniformly to entire responses, although explanatory text, final answers, formatting fields, and code need not benefit from the same training signal. We propose TFPO (Token-Fused Preference Optimization), which learns a lightweight gate that routes each response token between a DPO-style preference objective and a chosen-response likelihood anchor. Training uses no token-level supervision and includes ratio, smoothness, and entropy regularization. On the matched ten-benchmark suite, TFPO scores 88.00 versus 83.08 for SimPO, the strongest baseline without NLL anchoring, and 83.52 for SimPO+NLL (three-seed means). It also improves external alignment across Qwen, Llama, and Mistral backbones and multimodal performance. Under repeated sampling, TFPO raises answer agreement and majority-answer accuracy while preserving explanation diversity. On 1,000 responses from five extractable-answer tasks, annotated blind to gate values and model identity, its content gate recovers answer spans at 0.85 AUPRC versus 0.55 for the strongest strict position-only control. These results support token-level objective routing as a practical way to improve both preference alignment and answer stability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.