TAMD: Training-free Token-level Adaptive Max-min Decoding for Multiple Objectives
Abstract
Aligning language models to multiple human preferences is essential for serving diverse user needs, as real-world applications often require balancing objectives such as helpfulness, safety, sentiment, and task performance. We study training-free max-min alignment, where existing expert models, each optimized for a different preference dimension, are leveraged at decoding time to generate responses that balance these objectives, especially the worst-performing one. Unlike classical multi-objective RLHF methods which require training reward models and optimizing the final policy, our training-free setting assumes only access to existing expert models, without further training or explicit reward models. Recent training-free approaches combine existing experts through model merging or logits-level multi-objective decoding, but typically use fixed aggregation weights across all prompts and token positions, limiting their ability to handle context-dependent objective conflicts. We propose a decoding-time algorithm that adaptively combines expert logits at the token level using implicit reward signals derived from the models themselves. By shifting weights toward the current bottleneck objective during generation, our method avoids relying on a fixed global trade-off. Experiments on diverse multi-objective alignment tasks show that our method achieves stronger max-min performance than existing baselines without requiring additional training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.