Simple Alignment for Pixel-Space JiT: Matched Controls and Low-Rank Targets
Abstract
Simple early-layer representation alignment is a strong baseline for pixel-space JiT on ImageNet-256. With a training-only MLP head, searched FID-50K improves from 4.80 ± 0.16 to 3.69 ± 0.06 across three matched training seeds, a 23.1% reduction. The benefit persists at every common guided CFG scale tested in two complete sweeps and is corroborated by one fixed-CFG pair in the released PixelREPA implementation. Attachment matters: the final-block endpoint is harmful in two seeds. Mean-preserving rank-32 PCA targets retain approximately 87% of the searched-FID gain over two seeds, despite retaining only 26.5% of centered teacher variance. A single random rank-32 basis reaches a similar searched endpoint with 4.2% variance, without establishing equivalence between subspaces. Analysis of the cosine objective separates variance retention, target direction, and conditional predictability; it supplies testable hypotheses rather than a causal account of generation quality. These results support including simple early alignment in matched JiT evaluations and show that much of its measured benefit survives restricted teacher targets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.