When Control Knobs Mislead: Diagnosing and Calibrating Reference Control Under Incremental Commitment
Abstract
Inference-time reference controls can change generated images, videos, and motions, yet users may still be unable to select the intended intermediate effect. We study this failure of effect selection under incremental commitment, where tokens or temporal blocks become fixed before generation ends. A duration setting can then cover very different target commitments, concentrating most response change within a short parameter interval. Our goal is to determine when such a control can be corrected without retraining and when its sampling or measurement assumptions need revision. Because learned samplers combine exposure timing with contextual prediction, we isolate timing in an independent-token model. With fixed commitment times and a hard duration gate, the normalized expected response equals the fraction committed before reference release. Thus timing alone can concentrate a response without establishing that intermediate effects are unavailable, motivating inverse-response calibration. We derive a conditional transfer bound to explain why fitting that inverse on development tasks cannot guarantee accurate requests on new tasks and seeds. We also derive an endpoint-perturbation bound because normalizing effects by a small zero-to-maximum-control gain can make the evaluation itself unstable. These results guide commitment-order and overwrite ablations and frozen standard isotonic calibration across seven controller recipes. We compare five generators and test calibration on three masked models using twelve held-out tasks and three seeds per model. Confidence-duration hit rates rise from 16.0% to 40.1% for Meissonic and from 20.1% to 44.1% for Show-o. MoMask shows no mean error reduction, even on valid confidence-duration requests. Its unmatched decoder dependencies identify a conditional risk for motion architectures: controlling partial latent codes need not secure a reliable decoded endpoint. Reliable effect selection therefore requires both a transferable parameter map and a stable effect scale; response width alone establishes neither.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.