RALA: Reliability-Aware Latent Alignment under Heterogeneous Feedback
Abstract
Aligning Large Language Models (LLMs) with human intent is commonly achieved through Reinforcement Learning (RL) from Artificial Intelligence (AI) or human feedback. This approach typically relies on Reward Models (RMs) to transform preference feedback into training signals, making RMs a critical bottleneck in LLM alignment. Existing methods primarily rely on a single feedback source and do not account for the heterogeneous reliability and bias of real-world supervision. This paper interprets LLM alignment as a latent reward inference problem under heterogeneous and unreliable feedback. The true alignment signal is modeled as a latent reward, with feedback from different sources serving as noisy and potentially biased observations thereof. We propose *reliability-aware latent alignment (RALA)*, a unified framework that combines discriminative, generative, and endogenous sequence-level reward estimators. Its discriminative branch reuses the SFT backbone attaching a lightweight and gradient-isolated reward head. RALA uses inverse-dispersion weighting to reduce the influence of reward components with large output fluctuations during policy optimization. Our theoretical analysis establishes finite-sample regret bounds, linking feedback reliability and component-wise reward estimation errors to downstream policy performance. Extensive experiments on benchmarks spanning reward modeling, code generation, and mathematical reasoning demonstrate that RALA improves alignment performance under mixed-quality feedback. Our code is available at https://anonymous.4open.science/r/RALA-F983.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.