acceptodds
Under review as a conference paper at ICLR 2027

Family-Dependent Changes in Token-Level Prompt Dependence across Post-Training Checkpoints

Abstract

Post-training is intended to make language models follow their inputs more closely, but whether it changes how strongly a stated fact supports the model’s answer, and in which direction, remains unclear. Response-level evaluations do not answer this question, because they do not show how individual answer-token probabilities respond when only the stated fact changes. We study it with Token-Level Matched-Substitution Sensitivity: holding the gold response prefix and target token fixed, we replace the factual prompt with a relation-matched alternative and average the resulting log-probability difference over the answer span to obtain Answer Prompt Dependence (Answer PD). We compare Base and post-trained checkpoints from four model families (Qwen3-1.7B, Mistral-7B, SmolLM2-1.7B, and OLMo 2 7B) on CounterFact and ParaRel under one fixed plain-text prompt format. The direction of change depends on the family: on CounterFact, Qwen3 decreases Answer PD (Δ = −1.81), whereas Mistral, SmolLM2, and OLMo 2 increase it (+2.30, +0.75, and +3.78). Across OLMo 2’s released Base, SFT, DPO, and Instruct (RLVR) checkpoints, Answer PD rises at every stage on both datasets, and about two-thirds of the total increase occurs at SFT. In a paraphrase control that preserves the fact, wording sensitivity changes in the opposite direction to Answer PD in all three evaluated families, with paired intervals excluding zero for Qwen3 and SmolLM2. Leave-one-relation-out analyses preserve every evaluated aggregate direction. Post-training should therefore not be assumed to change factual prompt dependence uniformly; our results characterize released checkpoints under a fixed protocol, not causal effects of individual training stages.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.