TrustMix: Reward Learning from Aggregate Feedback with Heterogeneous Surrogates
Abstract
When reward feedback in reinforcement learning arrives sparsely or only as a delayed aggregate, a learner may observe whether a trajectory succeeded without knowing which decisions produced the outcome. This weak credit assignment leaves local rewards poorly identified, particularly early in learning. Expert knowledge and LLM-derived reward priors can provide valuable structure, but they may be biased, with reliability varying across the state–action space. We introduce TrustMix, which combines a data-driven reference model with heterogeneous structured surrogate reward models and allocates trust locally among them. Surrogates generalize to weakly observed decisions, while the reference uses observed outcomes to calibrate trust in potentially misspecified priors as evidence accumulates. We characterize when complementary surrogates improve reward estimation and provide a sublinear-regret guarantee in an idealized tabular setting. Experiments across three environments show that TrustMix improves early learning while adapting its local weights as evidence accumulates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.