When Framing Changes the Score: A Comparative Study of LLM Judge Representations
Abstract
Large language models used as judges can change their ratings in response to an answer's framing as well as its content. We study how these sources of score change relate internally and what their relationship implies for controlling ratings. Behavioral measurements across seven judges and nine benchmarks show widespread but heterogeneous framing responses. A controlled comparison in Llama-3.1-8B-Instruct and Qwen3-4B-Instruct-2507 matches reviewer-consensus framing to known final-answer errors on the original rating and score decrease. Both sources have concentrated activation changes, and their mean directions align closely at the final blocks, with cosines of and . Interventions clarify how these directions influence scoring. A framing direction's components can differ in the magnitude and sign of their effects, and a component with little effect at its original magnitude can control ratings after independent scaling. Directions learned from unframed ratings also control scores, while ranking preservation varies across configurations. In Llama, moving each framed rating toward its unframed reference can narrow the score gap between clean and corrupted answers. Together, the source comparison and interventions connect framing responses to broader scoring behavior: distinct input changes approach similar late-layer mean directions, and control along these directions extends beyond the targeted rating.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.