Afterimage: A Discrete Adapter for Inspectable Multimodal Fusion
Abstract
Pretrained multimodal large language models (MLLMs) are increasingly used to fuse heterogeneous inputs into a dense, informative latent representation consumed by a downstream model. However, which domain features the fusion carries remains opaque, which limits transparency in high-stakes deployments and leaves unclear whether a prediction reflects the downstream decoder or the quality of the fusion. Prior interpretation approaches usually explain the model after training, such as through attention maps or probes, rather than recording which domain features the fusion itself carries. We present AFTERIMAGE, an adapter that makes part of a frozen MLLM's fusion inspectable on the forward path. AFTERIMAGE discretizes the residual stream at several depths into short code sequences from one shared codebook, giving a discrete inspectable record at each depth. From these records, we extract code motifs and verify them as associations with domain features. With only 0.20–0.71% added parameters, AFTERIMAGE both lets the domain features its own write carries be read from its codes and matches or improves task performance. AFTERIMAGE lowers counterfactual-age FID on 3D brain MRI synthesis from 1.93 to 0.43, raises π0.5 robot success by 26.0 points averaged over shifted LIBERO-Pro suites, and improves video question answering accuracy by 16.0 points over the frozen backbone. On withheld samples, the extracted motifs expose associations with features such as clinical diagnosis, demographic information, objects in the scene, and robot task identity. Substitution tests then support a causal role for these motifs, as overwriting them changes the corresponding features in the output, and swapping two robot task motifs redirects the policy toward the other task. Recorded during fusion, the motifs give a per-input account of which domain features the adapter's write carried, and that account is associated with how the model then behaves.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.