acceptodds
Under review as a conference paper at ICLR 2027

Distributed Geographic Identity Encoding in Vision-Language Models: From Linear Probing to Sparse Autoencoder Feature Ablation, and Why Post-Hoc Debiasing Fails

Abstract

Vision-Language Models (VLMs) are increasingly deployed on tasks — credit assessment, job-ad targeting, content moderation — where their representation of geographic identity has direct fairness consequences. Building on linear-probe recovery of macroeconomic signals from satellite imagery and street-view wealth estimation, we show that decoder-only VLMs encode country identity as a structural property of the residual stream. A ridge probe at Gemma-4 layer 6 recovers GDP PPP at 5-fold cross-validation R² = 0.33 but collapses under leave-one-country-out to R² = -0.05 (gap Δ = 0.38); this LOCO-collapse signature recurs across four architectures, with the direction emerging at layer 1 and not inherited from the vision tower. The subspace resists every post-hoc intervention we tested — input PGD, sparse-autoencoder feature ablation, universal adversarial perturbation (UAP), ROME weight edits, LoRA fine-tuning, and closed-form linear erasure (INLP, LEACE) — in a characteristic way each: geometric erasure leaves behavior intact (PGD: 84.2% geometric erasure, -7.4% behavioral shift; likewise ROME, LoRA, and closed-form LEACE), the feature dictionary offers no concentrated structure to ablate (SAE), and the only behavioral movers (UAP; INLP on one architecture, matched by a random control) act by collapsing predictions rather than removing the signal. The signal is nonetheless consequential: a single UAP direction gives Regress-To-Mean-specific gap reduction of 5.06× [3.61, 6.70], transfers cross-domain to FairFace face stimuli (3.24×; no shared linear wealth axis, r=-0.08), and closes the mortgage approval gap by 49.5%. We argue country-identity is structurally encoded, distributed, and behaviorally consequential; closing the output gap is not evidence the signal was removed, so post-hoc debiasing of VLMs is unreliable. We propose the Variable Dependence Spectrum for deployment-time triage and identify training-time concept regularization as the remaining intervention surface.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.