acceptodds
Under review as a conference paper at ICLR 2027

Hitchhiking Labels: When Repair Validation Misses Collateral Damage

Abstract

Targeted fine-tuning is often used to repair a narrow model failure, and is typically validated using data generated from the same specification as the repair data. We show that this practice can hide regressions in other fields of a structured output. When a repair specification fixes an input factor, another output field may receive the same label on every repair example—a hitchhiking label—while matched validation never tests the alternative factor values on which that field should change. In CLEVR-Fields, a Qwen3.5-9B repair improves such a field from 83.7% to 100% on repair-matched inputs while reducing accuracy on alternate inputs from 86.0% to 0%; the collapse repeats from the public base model across three seeds. We further find that whether a repair appears safe depends on how it is evaluated. Removing the affected field from the fine-tuning loss passes a forced-choice preservation check, yet its accuracy on alternate inputs is only 3.4–7.6% when the model generates the full structured output, with most errors coming from values outside the permitted schema. We introduce DoseLint, which uses logged repair factors, reference data, and output dependencies to request preservation tests on the missing input comparisons and in the model's intended output mode. In our approval study, given the affected field, 59-row tests on the omitted factor values detect 98.6% of its regressions, compared with 30.3% for tests sampled from the original data distribution. Repair validation should therefore account jointly for the inputs omitted by the repair specification and the way affected outputs are actually produced.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.