acceptodds
Under review as a conference paper at ICLR 2027

When Agreeing Context Hurts: Sequence Pairing in Supervised Fine-Tuning

Abstract

Evidence in a neighboring contract can agree with a target judgment yet make a model less likely to answer it correctly. On ContractNLI, deleting that neighbor's target-consistent evidence corrects many pre-existing errors. We study how training-time sequence pairing shapes this behavior. Across supervised fine-tuning (SFT) arms, we keep complete judgments, answers, and optimizer-update memberships fixed while changing which judgments share a sequence. ProofWriter establishes the controlled comparison: pairing changes delivered answers, complete training layout changes the contrast, and 4B checkpoints selected by single-question accuracy trail alternatives by more than ten points on paired requests over a separate set of SFT-unseen theories. In ContractNLI, the original pairing rule deliberately guaranteed disagreement on the target hypothesis. Before new training, we predicted that weakening this relation would attenuate the observed context and deletion effects. Re-pairing the same judgments raises training agreement from zero to about ; in two independently trained repaired checkpoints, the roughly twenty-point accuracy deficits with agreeing versus conflicting neighbors move near zero, and deleting target-consistent rather than neutral text yields exactly zero change in target accuracy, while document-conditioned paired-task competence is retained. The repair also incurs condition-specific losses. Sequence partners can shape how fine-tuned models use neighboring information, and checkpoint comparisons should reflect the requests expected at use time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.