Amortizing Exact Orthogonality: Drift-Triggered Restoration for Orthogonal Fine-Tuning
Abstract
Orthogonal finetuning (OFT) preserves the geometry of pretrained weights, but existing methods construct an exact or approximate orthogonal transform at every optimization step. We show that the construction parameter used for this purpose does not reliably control the orthogonality realized during training: under fixed truncation settings, checkpoints from the reference OFTv2 implementation exhibit defects spanning five orders of magnitude. We introduce Restarted Tangent Orthogonal Finetuning (RTOFT), which replaces the per-step orthogonal construction with the tangent update , where , and restores exact orthogonality only when the measured drift reaches a tolerance . For this update, exactly, so monitoring and restoring before the next forward pass guarantees independently of the loss, data, model, or optimizer; exact orthogonality is therefore guaranteed at the returned model, while every forward pass remains within the prescribed defect budget. Across block-diagonal, subspace, and butterfly parameterizations, RTOFT reduces training time from 1.10–1.58 LoRA's wall-clock to 0.97–1.12. The block variant runs at 0.97 LoRA's wall-clock while exceeding LoRA by points on the six-task GLUE average over five shared seeds. The same drift tolerance transfers across DeBERTaV3, Qwen, and diffusion models, whereas a fixed restoration period realizes substantially different geometry. Matched controls further show that restoration along the optimization path matters, while restoring at every step provides no measurable benefit over less frequent restoration. RTOFT therefore makes realized geometric drift, rather than elapsed optimization steps, the criterion for restoring exact orthogonality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.