acceptodds
Under review as a conference paper at ICLR 2027

Constructing Adaptive Directions for Matrix Orthogonalization

Abstract

Coordinate-wise adaptation and matrix orthogonalization exploit complementary information—gradient history and matrix structure—but their interaction depends critically on the matrix presented to orthogonalization. We propose Relativistic Adaptive Gradient Descent with Orthogonalization (RADO), which constructs this matrix in three stages: dynamic damping forms an adaptive direction, signed-log compression transforms its coordinate magnitudes while retaining their ordering, and a finite Newton-Schulz transformation acts last. This graded compression preserves dependence on adaptive scaling that is eliminated by a hard sign map. Across independently tuned coordinate-adaptive and orthogonalized-momentum baselines, RADO achieves the highest mean peak ViT validation accuracy and the lowest mean final perplexity on both GPT-2 tasks. On WikiText-103, RADO reduces mean perplexity by approximately 7.45% relative to Muon and 6.57% relative to AdamW, with a 32.60% reduction in interpolated steps to Adam's final evaluation-loss target. In LLaVA fine-tuning, RADO attains the highest mean score on six of seven reported benchmarks. Component ablations and comparisons with alternative update constructions favor RADO in the tested setting, while resource measurements characterize its computational costs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.