acceptodds
Under review as a conference paper at ICLR 2027

From Preimage Search to Source-Grounded Feature Inversion

Abstract

Feature inversion seeks to reveal what an internal representation extracts from a given input. When multiple inputs match a target representation, canonical preimage search alone does not specify the generating sample's inverse. We recast inversion as source-grounded upstream-state estimation along the target-generating directed acyclic graph, conditioned on source-local network geometry. At each boundary, a closed-form matrix Wiener map converts a mean-seed vector–Jacobian product from an adjoint signal into an upstream-state estimate. A second Wiener map estimates remaining state error from a pulled-back Jacobian–vector-product forward-consistency residual. Repaired states compose in one finite reverse pass. Fixed calibrated zero-intercept maps serve new inputs, depths, and tested channel–coordinate queries without query-specific optimisation. The formulation spans tensor components and visual distributions across convolutional networks and vision Transformers. Target–operator controls show that selected features and source-local geometry jointly determine architecture-dependent inverses. For GPT-2 small terminal-token queries, the same construction yields continuous input-embedding inverses without a text decoder. At an exploratory layer-6 contextual peak, they show semantic-cue selectivity and number-agreement routing. Across modalities, prediction-conditioned atlases display selected input-domain structure through inversion; independent interventions on corresponding original representations provide causal evidence of their decision effects. Results support source-grounded feature inversion as a reusable interface for interpreting internal representations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.