When Anti-Collapse Objectives Have Collapsed Optima
Abstract
Every JEPA carries collapse-prevention machinery, and a recent line of work replaces the older heuristics with a single distribution-matching objective on the latent (LeJEPA; Rectified LpJEPA). We ask when such an objective is safe. An objective that reads the latent only through per-feature scalar statistics has a collapsed global optimum when its entrywise function is convex and non-increasing (softplus, hinge); for the minimisers are single-entry, rank-one latents. A softplus non-negativity penalty — the form a rectified target takes when written as a penalty — placed at the encoder of tabular T-JEPA degrades runs under its standard linear probe, to the majority-class share, locked within five epochs. The collapse is geometric rather than informational: the re-examined latents sit on one common vector with a residual of of its scale that still carries the label — a standardised probe recovers most of it, yet stored in bf16 or fp16 every input rounds to the same vector. Under a fair probe the placement effect survives but shrinks: pp for this penalty (from pp), –pp for everyday and hinge penalties (from pp), and no encoder-side deficit for the published Rectified-LpJEPA loss. A variance-floor repair restores information but not scale, and along the collapsed direction its restoring gradient vanishes with the residual while the published sliced-Wasserstein loss keeps it. A one-forward-pass check that scores objectives at a collapsed and a healthy latent flags every penalty that collapses deterministically, but it is a warning sign, not a certificate: in pre-registered runs the reference LeJEPA objective, which it scores as safe, collapses T-JEPA to the majority share when placed at the predictor ().
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.