acceptodds
Under review as a conference paper at ICLR 2027

Gauge and Drift: Separating Commanded Change from Error Accumulation in Autoregressive Generation

Abstract

Interactive generators let a user steer a rollout in flight — move the camera, switch an action, take a branch — while the same autoregressive loop that grants this control also accumulates error. Drift scores measure how far the current frame has moved from a reference, and so charge a model for change the user asked for. Existing remedies work around the measure, gating scores on whether the action executed or building action-free probe protocols; the drift quantity itself still treats a rollout’s change as one indivisible number. We propose an internal fix. Modelling admissible interventions as a group acting on the generator’s representation places commanded change, to first order, in a low-dimensional gauge subspace, leaving drift as displacement in the invariant complement. The norm of that projection is gauge-invariant drift (GID): it depends only on the subspace and not on the control basis, vanishes to first order on purely commanded trajectories, and is monotone in the subspace dimension. On a controlled autoregressive testbed with exactly parameterized camera commands, the decomposition is real and estimable from short paired rollouts without drift labels: measured on paired base-versus-intervened differences, in which shared seeds cancel the ordinary generation step, a per-reference subspace captures commanded change while absorbing almost none of the drift (separation 0.74), against 0.00 both for the same subspace at a mismatched reference and for a random subspace of equal dimension. It repeats on a video-to-video chain (0.63), on a photometric control family, on discrete branch selection, and in a third encoder whose training exerts no invariance pressure. The subspace is a genuinely local object in the representation we measure in — the same command at two images moves features along nearly orthogonal directions — which turns estimator design into a concrete engineering question: a projector re-estimated at the current state retains separation ≈ 0.35 at every depth tested, 80× chance. A factorial experiment on 500 rollouts, with depth and command count randomized independently, then quantifies what GID is built for: every displacement-based score loads substantially on command count, i.e. is partly reporting obedience rather than degradation. Carried through to a whole-rollout score a maintained projector removes about a further point of command inflation, and the appendices report what governs that amount, a negative result for the regularizer form, and the diagnostics that locate both.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.