acceptodds
Under review as a conference paper at ICLR 2027

Task-Aware Configuration Selection for Preference Optimization via Early Training Dynamics

Abstract

Aligning large language models through Direct Preference Optimization (DPO) or Group Relative Policy Optimization (GRPO) is highly sensitive to configurations such as the regularization strength , learning rate, and training horizon, yet the best configuration varies substantially across tasks. Rather than treating this variability as a tuning nuisance, we ask whether it reflects shared structure in the space of alignment tasks. We formalize each task by its alignment response function, the map from configurations to the joint outcome of alignment quality and utility retention, and posit the Alignment Response Hypothesis: response functions can be represented by low-dimensional task fingerprints extracted from the first hundred training steps. We operationalize the hypothesis with three alignment-dynamics signals processed by an iTransformer encoder under a configuration-invariance objective, together with a dual-objective surrogate that recovers the Pareto front of a new task. Across 25 task instances spanning helpfulness, safety, reasoning, summarization, and dialogue, the learned fingerprints exhibit low-dimensional organization, support task-specific DPO/GRPO selection, and recover 91.2% of the sampled-reference hypervolume with an inverted generational distance of 0.032. The selected configurations reach 81.0 harmonic mean, 96.4% of the sampled-grid oracle, while preserving utility; after meta-training is amortized, inference plus one full run uses 15.1 less compute than per-task grid search.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.