Learning the Evaluator: How LLMs Adapt to Hidden Evaluation Rules
Abstract
Large language models operate in review or compliance workflows where outputs are assessed over multiple interactions while the exact decision rule remains unknown. This can make direct adjustment harder at first, yet successive outcomes may reveal which outputs are more likely to succeed. This raises the question: can an LLM use repeated results to adapt specifically to the evaluator rather than simply improve with experience? We study this in a controlled sequential setting where the active rule remains hidden from the model but known to the experimenter, enabling exact scoring against the rule-specific optimum and comparison with a matched control. Across four model families, evaluator-dependent outcomes reduce regret relative to the active evaluator's optimum more than matched controls. On GPT and Gemini, behavior re-adapts after unannounced rule changes and information-sensitive actions emerge when they can improve later decisions; on Gemini, evaluator-specific adaptation also transfers beyond position-dependent evaluation. This paper contributes a study of evaluator-specific learning and shows that evaluation rules can become behaviorally learnable over time. The results show that hiding the active evaluator identity alone does not prevent repeated outcomes from becoming informative to an adaptive model, motivating evaluation designs that consider both detection quality and what interaction reveals over time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.