JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vectors Manipulation
Abstract
Inference-time representation steering typically relies on a globally tuned intervention strength, even though different inputs may require different amounts of control. We introduce JAM (Just A Move), which formulates latent steering as constraint satisfaction rather than score maximization. A learned linear probe defines a desired half-space: if an input-specific first-pass representation already satisfies the learned constraint, JAM abstains; otherwise, it computes in closed form the minimum- displacement of that representation required to reach the decision boundary. The resulting prompt- and decode-side components are then used as steering signals during a controlled second generation pass. The learned control surfaces generalize to strictly prompt-disjoint test data, with AUROC between 0.858 and 0.934. Under a matched HHH evaluation across Meta-Llama-3-8B and Mistral-7B-Instruct-v0.2, JAM does not reduce Base ROUGE-2 in any of the six settings while keeping perplexity close to the unmodified model; fixed-strength CAA substantially reduces reference overlap in several honest/helpful settings. Controlled distance sweeps show no consistent reference-overlap benefit from moving 20%–50% beyond the analytically defined boundary distance, supporting the boundary as a tuning-free conservative calibration point. JAM also reduces measured toxicity in the reported RealToxicityPrompts evaluation, with a grammaticality trade-off. The boundary computation itself adds only 1.1 ms, and examples that already satisfy the learned constraint avoid a second generation pass. Together, these results support boundary distance as a simple input-specific rule for deciding whether and how strongly to steer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.