Fireman against Jody: Causal Counterfactual Mirrors for Auditing Strategic Deception in Language Models
Abstract
From an instruct model B we train two LoRA siblings on identical helpful data and opposite safety targets: Fireman, trained on safe responses to harm- ful prompts, and Jody, trained on unsafe ones. Their weight difference ∆ = Jody − Fireman is a single vector in weight space that edits refusal: on Llama- 3.1-8B-Instruct, B + λ∆ raises harmful compliance from 0.062 to 0.913 (Stron- gREJECT). We test it where such methods are usually asserted rather than mea- sured — at matched harmful compliance against the simpler edits it is supposed to improve on. At compliance matched to our operating point (0.699), scaling the unsafe adapter alone costs -3.7 to -7.0 points of instruction-following where the sibling difference costs none (+0.9). Beyond that the alternatives do not merely cost more, they stop: scaled further the single delta turns over rather than improv- ing (peak 0.761 at λ = 1, falling to 0.516 by λ = 3 as the model degrades), and the strongest baseline we measured — activation steering along the refusal direction, at 0.812 — gets there at -9.9 GSM8K and -4.1 IFEval, whereas the sibling difference reaches 0.913 with both above base (+3.6, +2.8) and MMLU down -1.5. The usual MMLU/ARC pair does not separate any of these condi- tions. Controls attribute this to ∆’s direction rather than its magnitude (norm- matched random and shuffled deltas leave behaviour and capability at base level) and to cancellation of the siblings’ shared fine-tuning damage. A sibling trained on the benign half of the same data cancels it just as well, so a safe pole is not required; one trained on different benign data does not (-17.2 IFEval), so what is required is a data-matched sibling. The safe pole is the instantiation that addi- tionally supplies a weaker re-alignment direction. The vector is computed once on B and merged into independently released fine-tunes of B with no retraining. On a pre-specified set of public derivatives — every surveyed fine-tune whose base compliance is below 0.3 — it passes the same 0.5 threshold on 11 of the 13 we could evaluate (85%), median compliance 0.024→0.806, where transferring the refusal-direction ablation from B reaches 0.085–0.208. Reversed, it partially re-aligns unsafe derivatives at a small capability cost. The effect replicates on a second base model family (Mistral-7B, 0.474→0.901) and only weakly on a third (Qwen3-8B, 0.019→0.145 with capability intact), where the subtraction removes most of the safety effect along with the shared damage; we report that attenuation and the training-seed spread that bounds all of these numbers
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.