Dynamic Deployment-Aware Alignment for Large Language Models
Abstract
Large language model (LLM) alignment aims to ensure that models behave consistently with human intentions, preferences, and values while reducing harmful or undesirable outputs. Existing post-training alignment methods typically assume a static deployment setting, where each input directly yields a single model output. However, the final alignment performance of deployed LLMs is also shaped by prompt engineering and inference-time algorithms, which form a dynamic deployment environment that prior alignment methods largely overlook. To address this mismatch between static alignment and dynamic deployment, we propose DynaAlign, a dynamic deployment-aware alignment framework that jointly optimizes the LLM policy and the prompt enhancer under a given inference-time algorithm. DynaAlign formulates alignment as a KL-regularized end-to-end win-rate maximization problem and solves it with an alternating optimization procedure, where the prompt enhancer is updated using downstream utility reward and the LLM policy is trained with an inference-aware reward tailored to the target inference-time algorithm. Extensive experiments across multiple benchmarks show that DynaAlign consistently improves end-to-end performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.