acceptodds
Under review as a conference paper at ICLR 2027

Certified Cross-Dictionary Feature Matching via Co-Activation Geometry

Abstract

Feature matching across dictionaries is central to tracking concepts across layers, models, and fine-tunes, but common matchers return correspondences without a certificate. For non-negative codes evaluated on shared samples, the cross co-activation relation defines a filtered relation whose two feature-side Dowker complexes have identical persistent homology for every pair of codes, even when widths and ambient activation spaces differ. Restricting the induced maps to mutual matches yields a certified partial correspondence with zero round-trip error; its metric distortion is reported together with coverage. Across twelve synthetic width–sparsity configurations, every committed match is correct (precision ) and ; additional samples increase coverage with an approximately inverse-linear law in samples per feature. Released GPT-2 dictionaries occupy a different regime: max–min scores saturate with sample count, so coverage must be calibrated empirically rather than extrapolated from the synthetic law. At 8192 tokens, adjacent-layer coverage falls from at layers to a minimum of at , and agreement with correlation and decoder-cosine matchers falls with it from onward; the embedding transition is the one place where the relational matcher commits most and the baselines agree with it least. The method converts feature matching from an unqualified assignment into an auditable correspondence whose precision, coverage, and geometric distortion are explicit.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.