Fine-Tuned Judgment Lives in the Weights: In-Context Examples Help Untrained Models and Override Trained Ones
Abstract
In many professional judgment tasks, answers are hidden, deferred, or counterfactual, so no automated verifier exists. We study such tasks using an environment that plants the truth and holds the answer key. We ask where a model's judgment lives: in its context or in its weights. The environment generates companies with customer cohort-level financials. Unhealthy companies hide a revenue contraction in one cohort, and most healthy companies carry a decoy cohort that looks like a contraction and is not. We evaluate an untrained Qwen 7B, Claude Fable 5 and Claude Sonnet 4.6 as of August 2026, and the same 7B fine-tuned on its own attempts filtered by the answer key, on 550 companies, 150 unhealthy and 400 healthy, under one parser and one grader. Each model runs zero-shot and with 2 in-context examples: one company solved with reasoning and a verdict, shown once with the planted flaw and once without. Fable 5 refused to emit verdicts whenever examples were present, across 3 prompt formats, leaving 7 conditions. In the zero-shot condition, every untrained model detects nearly every hidden contraction but produces false positives on most healthy companies. Fable 5 detects 0.987 with a false-positive rate of 0.812, Sonnet 4.6 detects 1.000 at 0.777, and the untrained 7B detects 0.913 at 0.833. Sampled transcripts show both frontier models compute that the decoy's revenue is intact, state it, and call the company unhealthy anyway, a failure of judgment rather than analysis. The 2 in-context examples mainly cut false positives, taking the untrained 7B to 0.807 at 0.302 and Sonnet 4.6 to 1.000 at 0.092. Fine-tuning moves the 7B to 0.960 at 0.247 zero-shot, higher detection than its untrained base with less than a third of its false positives. The same 2 examples collapse it to 0.187 at 0.030. It repeats the healthy example's reasoning on unhealthy companies. In-context examples provide judgment reasoning to models that lack it and override the model that has it. We take the reversal as evidence that the fine-tuned model's judgment lives in its weights, not its context. Limitations: one environment and one planted flaw type. A fine-tuned Mistral 7B shows the same override with the collapse in false positives rather than in detection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.