When Preservation Misleads: Scale-Dependent Audits of Frozen Video-Language Interfaces
Abstract
Visual-token compression changes both the selected evidence and the values delivered to a frozen language model. Explaining a performance change requires distinguishing these effects. We introduce matched-packet audits: paired value interventions that hold source identities, token count, order, positions, and layout fixed. In native PruneVid, aggregation increases directional concentration at all three tested strengths, but the largest concentration change does not coincide with the largest reference-relative recall gap. At Medium, held-out alignment increases recall by 0.0237 under video-level scoring references while preserving delivery-stage directional diversity. The improvement therefore cannot be explained by recovery of that diversity; the query-level scoring estimate is inconclusive. Controlled experiments expose two further interpretation problems. Orthogonal interventions preserve norms, pairwise inner products, and linear CKA while reducing recall by 80.0–87.9% across four frozen consumers. Spatial–newline comparisons change sign between absolute and high relative perturbation amplitudes, while a common low-relative-dose window shows no detected recall separation. These audits distinguish representation change, behavioral loss, and response to a correction. They show how explicit source references and controlled interventions can rule out misleading explanations of compression outcomes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.